Text Metadata Remover — Check and Remove Hidden Text Characters Free
Wording lifted out of a document, a web page, a PDF or an AI chat window can arrive carrying code points you cannot see, from zero-width spaces to directional controls. This page shows you every one of them and takes them out on your own device, and your writing is never uploaded.
- Check & Clean In One Pass
- Emoji Stay Whole
- 100% Free & Private
Hidden Code Points
- U+200B:
- Zero width space
- U+202E:
- RTL override
- U+00A0:
- No-break space
- U+FEFF:
- Byte order mark
Same words,
nothing hidden
Paste Your Text to Reveal What Is Hiding
Which families should be taken out?
This additionally rewrites Cyrillic or Greek letters that imitate Latin ones, curly quotes, long dashes and ellipses. It also drops the protection that keeps emoji sequences together, so pictographs may come out altered. Leave it off unless you want maximum stripping.
Working…
Looking at a whole file rather than a passage of wording? The Document Metadata Remover handles Word, Excel and PowerPoint properties instead.
Characters Removed
- Zero-Width Characters
- Bidirectional Controls
- Unicode Tag Characters
- Unusual Whitespace
Hidden Character Report
Nothing pasted yet.
What "Metadata" Means in Plain Text
Almost none in the file; plenty in the characters.
A .txt or .md file is about as bare as a
document gets: no author field, no revision counter, no thumbnail. All
it holds is the characters you typed.
That is why people assume plain text is automatically safe, and it is where the assumption goes wrong. The payload is not wrapped around the words — it can sit between them. Unicode defines thousands of code points, and some take up no width at all.
Dropped into a sentence, such a code point is genuinely invisible: the line looks ordinary, the letter count quietly disagrees, and a search for a word containing one fails. Five families cover most of what turns up:
-
Zero-width characters. The zero-width space at
U+200B, the non-joiner and joiner atU+200CandU+200D, the word joiner atU+2060and the byte order mark atU+FEFFrender as nothing at all. -
Bidirectional controls.
U+202AtoU+202E, plus the isolates fromU+2066toU+2069, exist so Arabic or Hebrew can mix with Latin script. They change the order characters are displayed in. -
Unicode tag characters. The block from
U+E0000toU+E007Fmirrors the ASCII alphabet invisibly, so a readable message can be spelled out inside one that shows nothing. -
Variation selectors.
U+FE00toU+FE0Fand the supplement atU+E0100pick between alternate glyph forms, and can be chained to encode data. - Unusual whitespace. Non-breaking spaces, narrow no-break spaces, figure spaces and soft hyphens look like ordinary gaps but behave differently.
None of this is a flaw in Unicode, because each family has a real typographic or linguistic job. The difficulty is that a character invented to join a ligature works just as well as a marker nobody can see, and you cannot audit what you cannot perceive.
Where Hidden Text Characters Come From
Mostly ordinary copying, not sabotage.
The usual source is harmless. Most invisible characters arrive because software put them there for a sensible reason that stopped applying the moment you moved the words elsewhere.
Word processors and page layout
Word and Google Docs insert non-breaking spaces to stop a figure splitting across lines, and soft hyphens to mark where a long word may break. Lift the paragraph into a form field or database and they travel with it, serving no purpose and upsetting matching.
Web pages and PDFs
HTML authors write non-breaking spaces as entities, and selecting text in a browser captures them as real characters. PDFs are worse: the format stores glyph positions rather than words, so extraction tools often leave odd spaces or stray joiners behind.
Copy buttons and code samples
A "copy to clipboard" control hands over whatever the page author put in the hidden source, which may include zero-width markers used for syntax highlighting. Snippets from an editor can carry a byte order mark at the start, a classic cause of a script that will not run.
Chatbot and assistant output
Text from an AI assistant is still just text and rarely contains anything exotic. It does tend to carry typographic punctuation — curly quotes, long dashes, single-character ellipses — plus the odd non-breaking space. Those are visible characters, not hidden ones, which is why the stricter pass handles them separately.
Deliberate insertion is rarer, but real. Because a zero-width sequence survives copying, it can mark which recipient of a document later leaked it. Treat that as one possibility rather than the default explanation for anything the scan reports.
Why Hidden Text Characters Are a Privacy and Security Risk
Four documented problems, without the hype.
Marking a copy so it can be traced
A pattern of zero-width characters spread through a paragraph behaves like a serial number. Give each reader a different pattern and whichever copy reappears identifies its source. It survives most copying, and nothing on screen reveals it.
Misleading display order
Bidirectional overrides tell the renderer to lay characters out backwards. Researchers published this as Trojan Source in 2021: code can be arranged so a reviewer reads one thing while the compiler reads another, because a comment marker sits where the eye never puts it. The same trick works on a file name or a link.
Instructions smuggled inside ordinary text
Tag characters from the U+E0000 block spell out ASCII
invisibly. Pasted into a prompt, such a sequence reaches the language
model even though the person pasting it saw nothing unusual. It is a
recognised prompt-injection route, and the same mechanism can hide a
short message inside an innocent-looking line.
Quiet breakage in ordinary work
The dullest consequence is the most common. A soft hyphen in a part number stops it matching in a spreadsheet lookup. A non-breaking space in an email address fails validation with a baffling error. A byte order mark atop a config file triggers a complaint about character one.
Two honest limits. Removing these characters does not make a passage untraceable, because wording, phrasing and timing identify a document perfectly well on their own. And cleaning changes nothing you can actually read — the visible sentences come out as they went in.
How to Check and Remove Hidden Characters From Text (3 Steps)
Under a minute, start to finish.
-
Paste or load the text
Put your writing into the box at the top of this page, or use the second button to open a
.txt,.md,.csvor.jsonfrom your device. Either way the content is read into the page and goes no further. The sample button loads a short passage with planted characters. -
Read what the scan found
The report counts the total, breaks it down by family, then lists every distinct character with its code point, Unicode name and frequency. Above the table a preview of your own text marks each position, so you can tell whether a non-breaking space is doing a real job in a price or simply arrived by accident.
-
Copy or save the cleaned text
Untick any family you would rather keep and the report refreshes at once. When the result looks right, take it with the copy button or save it as a plain text file. Spacing characters fold back to an ordinary space so words do not run together; invisible ones are simply dropped.
Text Metadata Checker vs Text Metadata Remover
Two jobs, deliberately one page.
People search for both phrases, and they describe two halves of one task. Checking answers "is there anything in here?" Removing answers "take it out." Splitting them across two addresses would mean two thin pages competing with each other, so both live here.
The checking half is the report, deliberately read-only: code points, names, counts, and a preview showing where each sits. That matters because not every hit is a problem — a non-breaking space between a number and its currency is correct typography.
The removing half is the output box, which only acts on the families you ticked. Because the scan and the rewrite run together, the before and after counts sit side by side. Nothing is decided for you.
If your question is about a file rather than a passage of wording, the metadata viewer is the better starting point: it inspects photographs, PDFs and Office documents without altering them.
Does "Paste as Plain Text" Remove Hidden Characters?
It strips styling, not code points.
This misconception is worth clearing up, because the habit is so widespread. Pasting without formatting does something real, and it is not what most people think it is.
Copying from a rich editor puts several representations on the clipboard: a styled one carrying fonts, colours and sizes, and a stripped one holding only characters. Asking for the plain variant takes the second, so typefaces and heading levels vanish.
Invisible code points are not styling. A zero-width space is a character in the same sense that the letter k is, so it belongs to the stripped representation too and crosses over untouched. The same goes for directional controls, tag characters, non-breaking spaces and soft hyphens.
Round-tripping through a basic editor behaves the same way. Notepad, TextEdit in plain mode and most code editors store every one of these characters faithfully, because that is their purpose. Some can reveal them with a "show invisibles" setting, but revealing is not removing.
So use a formatting-free paste for what it is good at, and treat it as unrelated to cleaning. The only reliable answer comes from inspecting the code points, which is what the scan above does.
Text Metadata Remover FAQ
Short answers about cleaning text.
They are real Unicode code points that occupy no visible width, or that only tell software how to arrange what is around them. Zero-width spaces, joiners, word joiners, byte order marks, soft hyphens and directional controls all qualify. Your eyes see ordinary writing, yet the stored sequence contains more than the letters you can read.
Paste the passage into the box near the top of this page and press the check button. Zero-width entries are picked up straight away and listed by code point, and the cleaned version appears beside the report. You can then copy it or save it as a text file, with no account needed.
No. The scan and the rewrite both happen in JavaScript inside the tab you already have open, so your wording stays on your own machine. We never receive it and have no way to read it. If you want proof, open your browser network panel and watch while you clean a passage.
Not in the default mode. Many emoji are built from several code points glued together by joiners and variation selectors, so the text is first split into grapheme clusters and those glue characters are protected. Family emoji, hearts, flags and skin tones survive untouched. Only the stricter optional pass may alter them.
Usually not. Pasting without formatting drops fonts, colours and sizes, because those live in styling data rather than in the characters. Invisible code points are part of the text itself, so they ride along into the plain result. Treat a formatting-free paste as a styling fix, never as a cleaning step.
Yes, with no sign-up, no trial window and no cap on how many passages you run through it. The page is supported by the adverts around the tool rather than by charging you. Because the work happens in your own browser, there is no server cost for us to pass on to you.
Look Before You Paste It Somewhere Permanent
One scan tells you where you stand.
Invisible characters are an ordinary consequence of moving words between programs, and usually they are merely untidy. They deserve attention when the passage is heading into a contract, an article, a configuration file or a prompt, where an unseen marker lasts.
Scroll back up and run the passage you were about to send. The report answers in a second, and if there is nothing you get a plain confirmation rather than vague reassurance. Either way the words come back exactly as you wrote them.
Check my text now Read the privacy guide
Other kinds of file have their own pages: the image metadata remover for photographs, the EXIF data remover for camera fields, the GPS location remover for coordinates, the PDF metadata remover, the document metadata remover for Office properties, the video metadata remover for clips, and the batch metadata remover for a whole folder. Our privacy policy, terms of use and disclaimer explain how the site operates.