- Archaeological News
-
The Xiaojing Pavilion at the Xi’an Beilin Museum, also known as the Stele Forest, in Xi’an, China.
Image Credit:
蒋亦炯, CC BY-SA 3.0, via Wikimedia Commons
AI Reveals Ancient Chinese Inscriptions Hidden by Centuries of Damage
A new artificial intelligence system is showing how severely weathered Chinese stone inscriptions can be digitally restored without sacrificing the shape, rhythm and visual character of the original writing.
Ancient steles have preserved historical records and calligraphy for centuries, but their carved surfaces inevitably suffer from erosion, cracking, flaking and chemical deterioration. Rubbings made from these monuments can preserve valuable information, yet they often carry their own problems: faded strokes, uneven ink, creases and heavy background noise. Once the damage becomes severe, even advanced image-enhancement systems can struggle to tell where a character ends and the damaged stone begins.
That creates a difficult balance. Conventional image-restoration models may sharpen an inscription but accidentally join unrelated marks or break strokes that belong together, leaving characters that look cleaner while becoming less readable. Systems focused mainly on language can take the opposite approach, using context to reconstruct likely characters but losing the distinctive calligraphy and physical appearance of the original inscription.
The newly developed framework, known as SteleSR, aims to bridge that gap by considering both visual structure and textual meaning. Instead of treating the inscription simply as a damaged photograph, the system pays particular attention to the edges and continuity of individual strokes while also using textual information to keep reconstructed characters recognizable.
To develop the approach, researchers assembled a benchmark called SISTR from historical inscription rubbings held by the National Library of China. The material spans the Ming and Qing dynasties as well as the Republic of China, providing a wide range of calligraphic styles, engraving depths and states of preservation. The collection included 721 relatively clear images and 2,086 heavily degraded examples affected by problems including weathering, cracks and weak contrast.
Training such a system presents another problem: perfectly matched pairs showing the same inscription before and after centuries of damage simply don’t exist. The researchers therefore created artificial training examples in stages, simulating erosion, surface loss, blur, ink irregularities and other forms of deterioration. A second set of training images used large vision models to imitate the more unpredictable appearance of genuinely weathered historical rubbings.
This two-stage process first teaches the system to separate the underlying character structure from a damaged or noisy background. It then exposes the model to more realistic patterns of wear so that it can adapt to material found in actual historical collections.
SteleSR adds two important forms of guidance to existing image super-resolution models. The first concentrates on preserving stroke edges and continuity, helping prevent cracks or random marks from being mistaken for parts of a character. The second uses textual information from an optical character-recognition system, encouraging the restored image to remain consistent with recognizable writing rather than simply becoming visually sharper.
Tests across several established image-restoration systems showed consistent improvements. With the strongest-performing model, sequence-level character recognition rose from 0.6530 to 0.7210 after SteleSR training. Structural measurements also indicated fewer scattered noise fragments, fewer false connections and better continuity in the restored strokes.
The visual results support those measurements. In heavily degraded examples, conventional systems often left blurred, disconnected or incorrectly joined strokes, while the specialised framework produced cleaner character structures and reduced interference from damaged backgrounds. The improvement wasn’t simply cosmetic; the inscriptions also became easier for recognition software to read.
For cultural heritage, that distinction matters. A digitally restored inscription shouldn’t merely produce a plausible version of what the text might have said. It should remain tied to the surviving monument, preserving the width, geometry and texture of its strokes as closely as possible. The new framework is designed with that principle at its core.
There are still limitations. The system relies partly on a pretrained character-recognition model, which may struggle with very rare character forms or unusual historical scripts, while creating realistic simulated damage with large vision models requires substantial computing resources. Future development is expected to extend the approach to a wider range of writing systems and historical materials while reducing its computational demands. The challenge, then, is no longer simply making an ancient inscription look clearer. It is recovering lost information without allowing the digital restoration to overwrite the character of the original artifact.
Published on: 07-10-2026
Edited by: Abdulmnam Samakie
Source: npj Heritage Science