Characters the detector identifies
Copied text can contain zero-width spaces, non-breaking spaces, a byte-order mark, soft hyphens, control characters, tabs, mixed line endings, or repeated ordinary spaces. They can affect search, validation, wrapping, and source control even when the text looks normal.
| Marker | Character | Default handling |
|---|---|---|
| ZWSP | Zero width space | Review before removal |
| ZWJ / ZWNJ | Joiner / non-joiner | Keep; potentially semantic |
| NBSP | Non-breaking space | Replace in ordinary prose |
| BOM | Stray byte-order mark | Remove |
| BIDI | Direction control | Keep; manually review |
| 2×SPACE | Repeated U+0020 | Collapse in prose |
Why review matters
A zero-width joiner is part of many family emoji and Indic-script combinations. A non-joiner can distinguish words in Persian. Directional controls help mixed right-to-left and left-to-right text display correctly. Text Harmonizer therefore reconstructs a clean copy from selected actions instead of mutating the source automatically.
How invisible mode works
Ordinary spaces appear as middle dots. Special characters appear as labeled tokens with their Unicode name, code point, source index, and suggested action. The exact source remains available throughout the workflow.