This invention belongs to the field of
text processing technology, and in particular to a method and
system for text
structure recognition and
recovery based on structural features and
spatial distribution analysis. The method establishes a shared slice set recording
line length and indentation features, maps structural feature vectors containing line slice skeleton information and constructs a structural
fingerprint database, establishes structural node families through similarity clustering, performs collaborative arbitration based on confidence weights calculated from structural length consistency and node density stability to identify structural nodes, calculates text structure impulses and identifies structurally abnormal regions, and then completes structural nodes within abnormal regions through digital inheritance or feature homology
verification to recover the logical sequence. This achieves text
layout based on the
physical structure of the text, solving the problem of text structure discontinuity caused by
layout variations or missing identifiers. It does not rely on specific language corpora, and the shared slice
pool mechanism avoids repeated scanning of text, significantly improving the efficiency of text structure reconstruction.