Have you ever wondered what happens when you combine Microsoft Research’s Stroke Width Transform algorithm with Google’s open-source Tesseract OCR engine? No? Well, neither had I until I recently stumbled upon something called Project Naptha. Still confused? That’s okay—take a look at it yourself first (and then come back to read on!) here: Project Naptha.
Project Naptha is a Chrome plugin (and soon to be available for Firefox) that extracts text from images found on the web, allowing you to edit it directly. Have you ever come across a photo on a website where text has been stamped across it, perhaps the author’s name or initials in the corner? Well, now you can remove that text—or even edit it—without needing any complex software or tools beyond your web browser! This may not be a new concept (Adobe Photoshop users have been doing this for years), but what’s revolutionary here is how easily anyone can do it now, with no specialized knowledge.
While one obvious (and potentially problematic) application could be removing text from copyrighted images, there are positive uses too. For instance, text in images—particularly common in e-commerce—often contains valuable information, such as product details, but extracting that text typically meant retyping it manually. Now, thanks to Project Naptha, it’s as simple as copying and pasting.
Want to try it out for yourself? Visit Project Naptha and give it a go!