RDRepDex AI
Text Cleanup

Unicode Inspector

Inspect every character in your text with its Unicode code point and category.

Free · Runs in your browser · Nothing is uploaded

Paste text above to inspect every character.

The Unicode Inspector shows the actual characters inside pasted text: visible letters, spaces, symbols, emoji, non-ASCII punctuation, and invisible marks. It is for debugging text that looks normal but behaves as if something hidden is inside it.

What a Unicode inspection tells you

Every character has a code point. The ordinary capital A is U+0041. A zero-width space is U+200B. A non-breaking space is U+00A0. The inspector lists those code points so you can see what the text really contains instead of relying on appearance.

This is especially useful when two strings look identical but do not match. The difference may be a curly quote instead of a straight quote, a non-breaking space instead of a regular space, or a hidden direction mark between letters.

Using the unusual-character filter

For long text, the full character table can be noisy. The filter shows only non-ASCII and invisible characters, which is usually where copy-paste bugs live. If the filtered table is empty, the issue is probably not hidden Unicode.

If the table shows invisible characters, move to the Invisible Character Remover. If it shows curly punctuation or non-breaking spaces, use the AI Text Cleaner or Smart Quote Converter depending on how broad the cleanup needs to be.

Where unusual characters are legitimate

Do not automatically delete every non-ASCII character. Names, accented words, emoji, mathematical symbols, currency marks, and non-English scripts are all legitimate Unicode. The point of inspection is to identify surprises, not to force every text into plain English ASCII.

Unicode Inspector: choosing the right approach

The Unicode Inspector helps you inspect text characters and their code points when appearance alone is misleading. Checking Unicode values is useful for identifying hidden characters, unusual spaces, and lookalike symbols before deciding whether to remove or replace them.

Preserve the words while removing unwanted characters or spacing. A clean paragraph should keep its punctuation, paragraph boundaries, and intended meaning.

A character can be present without a visible glyph. Other characters look alike but have different code points. Visual inspection alone cannot distinguish a normal space from every spacing character or a Latin letter from a similar-looking letter in another script.

Decide what would count as a successful result before you begin. For invisible Unicode characters and character identity, that means checking the relevant constraint rather than choosing whatever looks most polished. Keep a short note of the intended audience or destination, any details that must remain exact, and the change you actually need. This gives you a basis for comparing alternatives and prevents a later edit from quietly changing the task.

Preparing input for Unicode Inspector

Before using Unicode Inspector, prepare a representative sample. Separate ordinary prose from code, tables, addresses, and verse. Save the original before changing whitespace or Unicode characters. Copy one troublesome paragraph first so you can inspect the exact change.

Work with the smallest complete sample that still represents the real task. An isolated word may hide a problem that only appears in a sentence or a list, while a large unrelated block can make the result harder to inspect. Include the difficult cases from your actual material and keep a separate copy of the original. If the source has several independent parts, process one part first and confirm the approach before continuing.

Separate requirements from examples. An example shows the kind of material you have; a requirement says what the result must preserve or accomplish. For invisible Unicode characters and character identity, write down any non-negotiable detail before changing the input. If two requirements conflict, such as keeping every detail while sharply reducing length, decide which matters more instead of expecting an automatic result to resolve that tradeoff reliably.

Text Cleanup use case: a practical brief

A paragraph copied from a PDF may contain extra spaces, a line break after every visual line, and a hyphen splitting a word. These are three different problems. Fix spacing first, inspect the line structure, and repair split words only after checking the original document.

Use this as a planning scenario, not a claim about an output already produced by Unicode Inspector. The useful part is the constraint: identify what the person needs, which information is available, and what would make a result unsuitable. Replace the scenario's details with your own before using it. A brief that names a concrete situation is easier to evaluate than a request for something simply better, more interesting, or more professional.

Compare two possible approaches to the same brief. One might prioritize speed or brevity; another might preserve more context or structure. Keep the underlying facts identical during that comparison. Otherwise, a result can appear stronger merely because it introduced a new claim. Choose the version that meets the actual task, then make a separate pass for presentation and tone.

A focused exercise for invisible Unicode characters and character identity

Inspect a short string that looks correct but fails to match another copy. Compare character positions and code points where the tool exposes them. Keep a copy before removing characters, and check the resulting word in its original language.

Look for zero-width characters, direction marks, and non-breaking spaces without assuming that all of them are unwanted. Some writing systems need joining behavior, and a blanket removal can change presentation or meaning.

Run the exercise on a copy and compare each meaningful part of the result with its source. When a change surprises you, isolate that part rather than adding more unrelated input. This makes it easier to tell whether the issue comes from the source, a selected option, or an expectation that belongs to another type of tool. Keep the exercise small enough to check manually before trusting it as a repeatable workflow.

How to review unicode inspector results

Review the unicode inspector result against the original task. Compare the first and last words of each paragraph, inspect quotation marks, and check whether list items still occupy separate lines. Paste the cleaned paragraph into the intended editor because that editor may add its own formatting.

Use two passes. First check correctness: does the result preserve the necessary facts, values, relationships, or boundaries? Then check usefulness: does it fit the person and place it is intended for? Keeping those questions separate helps you avoid accepting a fluent but inaccurate draft or rejecting a technically correct result only because it still needs ordinary presentation work.

Inspect the difficult part of the input first. A long result can look convincing at the beginning while mishandling a special case farther down. Compare that special case directly with the original, then check the surrounding material. For invisible Unicode characters and character identity, a short manual check is often more informative than repeatedly running the same input and hoping that another result will resolve the uncertainty.

Common mistakes with copied text and formatting cleanup

Removing every break can merge headings with body text. Removing every invisible character can change scripts that use joining characters. Treating cleanup as rewriting can hide a factual error that was already present in the source.

Invisible text is not secure hidden storage. Removing unusual Unicode does not prove that a document is free of tracking, attribution, or every possible watermark.

When the result is unsuitable, identify the failure before retrying. Was the source incomplete, the instruction ambiguous, the selected format inappropriate, or the task outside this tool's purpose? Change one relevant detail and compare again. Changing the entire brief at once makes it harder to learn which correction helped and can introduce a new problem into material that was already correct.

Using Unicode Inspector in a repeatable workflow

For a CMS article, clean the body copy before adding heading styles and links. For an email, clean the message separately from the signature. For spreadsheet imports, preserve the separators that identify columns.

Keep the source, the chosen settings, and the reviewed result together when you repeat this task. A simple note is enough; the important part is being able to explain why the final version was accepted. If another person will use the output, include the assumptions they need to know rather than passing along an unexplained result. This is especially useful when several people edit the same content at different stages.

Repeat a check when the input changes in a meaningful way. A workflow that worked for a short English paragraph may need another review for a structured list, unusual characters, a different audience, or a stricter destination. Reusing a process saves time, but reusing an old conclusion without checking the new conditions can create avoidable errors.

Worked examples

Input
A B
Result
A: U+0041, non-breaking space: U+00A0, B: U+0042

A non-breaking space looks like a normal space but has a different code point.

Frequently asked questions

Is every non-ASCII character a problem?+

No. Non-ASCII includes normal accented letters, symbols, and emoji. Focus on unexpected characters, especially invisible marks and spacing characters.

Why can identical-looking text fail an exact match?+

The strings may contain different code points, combining characters, or invisible characters. Compare the underlying characters and normalize only what your task permits. Appearance alone is not evidence that two strings are identical.

Why does copied text look different in another editor?+

The destination can interpret characters, line endings, and paragraph styles differently. A plain-text tool handles the text it receives, while the editor may add its own visual spacing. Compare the stored characters with the editor's formatting settings before deciding which part needs correction.

Should I keep a copy before cleaning a long document?+

Yes. Keep the original separately, especially when working with quotations, structured lists, or unusual characters. Some transformations discard information. An unchanged copy lets you compare the result and recover formatting that turns out to be meaningful after the cleanup.

Can I process prose and source code together?+

It is safer to separate them. Prose often tolerates normalized spacing, but code can depend on indentation, quotation marks, and exact characters. Clean only the passage that needs the chosen operation and inspect code with tools appropriate to its language and format.

How should I handle text copied from a two-column PDF?+

Check reading order before cleaning characters. A PDF selection may interleave columns or include headers and page numbers. Removing spaces or breaks cannot reconstruct a scrambled argument. Compare the copied passage with the page and rebuild the correct sequence first.

Will plain-text cleanup retain bold text and links?+

Do not assume so. Bold styling, link destinations, and document layout are separate from ordinary text characters. Keep the formatted original if those details matter, then apply the required styles in the destination after the text itself has been reviewed.

What should I do with a quoted passage?+

Preserve an exact source copy and follow the quotation rules relevant to your publication. Even small punctuation or spacing changes can matter in a quotation. If you create a normalized working version, distinguish it from the authoritative original rather than silently replacing it.

Why does an apparent blank line remain after cleanup?+

The gap may come from paragraph margins or line-height settings rather than blank characters. Inspect the destination editor's formatting. A character-based tool cannot remove spacing that exists only in a document style or a webpage's layout rules.

How can I tell whether a cleanup pass changed meaning?+

Compare names, numbers, punctuation around clauses, and paragraph boundaries. Read the result aloud and compare any changed sentence with the original. A mechanical operation should be evaluated by the changes it actually makes, not by how neat the final block looks.

Should I apply every cleanup operation at once?+

Start with the operation that solves the visible problem. Applying several transformations at once makes unexpected changes harder to trace. Work on a copy, check the result after each meaningful step, and stop when the text meets the requirements of its destination.

How do I report a reproducible formatting problem?+

Use a short non-sensitive sample, describe the chosen operation, and show the expected result beside what appeared. Include the source and destination applications when relevant. A minimal example with the troublesome character is more useful than a large private document.

What should I prepare before using Unicode Inspector?+

Separate ordinary prose from code, tables, addresses, and verse. Save the original before changing whitespace or Unicode characters. Copy one troublesome paragraph first so you can inspect the exact change.

What does a useful unicode inspector brief look like?+

A paragraph copied from a PDF may contain extra spaces, a line break after every visual line, and a hyphen splitting a word. These are three different problems. Fix spacing first, inspect the line structure, and repair split words only after checking the original document.

How should I check the result from Unicode Inspector?+

Compare the first and last words of each paragraph, inspect quotation marks, and check whether list items still occupy separate lines. Paste the cleaned paragraph into the intended editor because that editor may add its own formatting.

More Text Cleanup tools

Browse all 293 tools →