Heh, I read the title and thought this was going to be about someone porting legacy publications into Latex code. Anything from Newton's gravity to Marie Curie radium.
Which would be nice - some seminal publications from not all that long ago are only available as terrible quality scans.
I just asked GPT6-Astra to transcribe a page from Newton's Principia to LaTeX. It did an amazing job outputing a XeLaTeX doc with a mixture of text and TikZ. Looks really good.
This is an excellent idea. I was curious about using DeepSeek OCR for exactly this purpose. But a tricky question is if we could do some sort of looping or something "energy based" and use classical search to find optimal parameters (LaTeX settings) to minimize the error (pixel difference). Me knowing I would get obsessed with the second half is what's keeping me from the first half. Maybe a vision JEPA would be good. If I had API credits to burn, I'd copy paste our two comments and see how far Fable gets.
Rather Claudish(?)-sounding README (and even the examples!) aside, this definitely has that early 20th/late 19th-century aesthetic. I think 99% of it is due to the choice of font.
Heh, I read the title and thought this was going to be about someone porting legacy publications into Latex code. Anything from Newton's gravity to Marie Curie radium.
Which would be nice - some seminal publications from not all that long ago are only available as terrible quality scans.
I just asked GPT6-Astra to transcribe a page from Newton's Principia to LaTeX. It did an amazing job outputing a XeLaTeX doc with a mixture of text and TikZ. Looks really good.
Did you do a pixel diff against the original? The output could look plausible at a glance but contain hard-to-spot errors or inaccuracies.
Unchecked OCR can lead to amusing results: https://news.ycombinator.com/item?id=18374359
This is an excellent idea. I was curious about using DeepSeek OCR for exactly this purpose. But a tricky question is if we could do some sort of looping or something "energy based" and use classical search to find optimal parameters (LaTeX settings) to minimize the error (pixel difference). Me knowing I would get obsessed with the second half is what's keeping me from the first half. Maybe a vision JEPA would be good. If I had API credits to burn, I'd copy paste our two comments and see how far Fable gets.
Rather Claudish(?)-sounding README (and even the examples!) aside, this definitely has that early 20th/late 19th-century aesthetic. I think 99% of it is due to the choice of font.