打开 PDFnoted 网页版(无需下载),点击「上传文件」,支持 JPG/PNG/PDF 格式。建议单页分辨率 ≥200 DPI;若为多页扫描件,PDFnoted 自动批量处理全部页面。
在工具栏选择「OCR 扫描件识别」,语言默认为「中文+英文」(适配老档案中混合出现的旧式标点、繁体字与英文印章)。点击「开始识别」,PDFnoted 在 10–60 秒内完成文字提取与位置映射。
识别完成后,预览文字层叠加效果(绿色高亮显示识别区域)。点击「导出为可搜索 PDF」生成带隐藏文本层的标准 PDF,支持全文搜索、复制粘贴;也可选「导出为 TXT」获取纯文本用于归档索引。
Go to PDFnoted’s web app (no install required), click ‘Upload File’, and select your scanned archive (PDF/JPG/PNG). For best OCR accuracy, use scans ≥200 DPI. Multi-page PDFs are processed page-by-page automatically.
Click ‘OCR Scan Recognition’ in the toolbar. Default language is ‘Chinese + English’—optimized for archival documents with mixed traditional characters, legacy punctuation, and English stamps/seals. Click ‘Start OCR’; PDFnoted maps text positions and extracts content in 10–60 seconds.
Preview the overlay: green highlights show recognized text regions. Click ‘Export as Searchable PDF’ to generate a standard PDF with invisible text layer—fully searchable and copyable. Or choose ‘Export as TXT’ for clean plain text ideal for cataloging and metadata indexing.
PDFnoted 主要针对印刷体文字优化,对清晰铅印/胶印档案识别率超 95%;手写体和严重褪色内容暂不支持。建议先用扫描仪增强对比度,再上传。
可以。PDFnoted 导出的可搜索 PDF 符合 ISO 32000 标准,兼容 Adobe Acrobat Reader、Foxit、Edge 等所有主流阅读器,支持 Ctrl+F 全文检索。
支持。单次上传最大 200MB(约 300 页 A4 扫描件),PDFnoted 自动分页 OCR 并保持原始顺序;如需处理更多,可分批上传后合并导出。
PDFnoted is optimized for printed text—95%+ accuracy on clear letterpress/offset prints. Handwriting and severely faded ink are not supported. Pre-scan with contrast enhancement (e.g., ‘Text Mode’ on scanner) for best results.
Yes. PDFnoted’s searchable PDFs comply with ISO 32000 standards and work seamlessly in Adobe Reader, Foxit, Microsoft Edge, etc.—Ctrl+F works instantly across all pages.
Yes. Upload up to 200 MB per session (~300 A4 pages). PDFnoted processes pages sequentially while preserving order. For larger archives, upload in batches and merge outputs manually.