Ancient book content integrity detection method and system based on semantic structure features

The ancient book content integrity detection system based on semantic structure features solves the problem of semantic integrity assessment at the structural level of ancient book content, realizes multi-level automated and standardized analysis of ancient book content, and improves the efficiency and value of intelligent utilization of ancient book documents.

CN121997934APending Publication Date: 2026-05-08SICHUAN AGRI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SICHUAN AGRI UNIV
Filing Date
2026-01-23
Publication Date
2026-05-08

AI Technical Summary

Technical Problem

Existing methods for detecting the integrity of ancient texts have significant shortcomings at the semantic structure level. They are unable to effectively determine the connection between paragraphs and sentences, whether the temporal logic is disordered, or whether the terminology lacks definition. This makes it difficult to achieve a comprehensive quality assessment of the content and limits the intelligent use of ancient texts.

Method used

An ancient book content integrity detection system based on semantic structure features is adopted. The system extracts chapter and paragraph structure feature data through the data acquisition unit, and generates multi-level integrity feedback signals by combining chapter integrity analysis, paragraph integrity analysis and comprehensive analysis units, so as to realize the automated and standardized evaluation of ancient book content.

Benefits of technology

It enables multi-level automated and standardized analysis of ancient texts, accurately judges the integrity of chapters and paragraphs, supports the classified use of ancient texts with different integrity levels in education, research and cultural dissemination scenarios, and improves the efficiency and value of intelligent use of ancient texts.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121997934A_ABST
    Figure CN121997934A_ABST
Patent Text Reader

Abstract

The invention discloses a semantic structure feature-based ancient book content integrity detection method and system, and relates to the technical field of information processing, the system comprises a data acquisition unit, a chapter integrity analysis unit, a paragraph integrity analysis unit, an integrity comprehensive analysis unit, a result feedback unit and a display terminal; structural features of chapters and paragraphs are accurately extracted by utilizing a semantic structure recognition technology, and the content integrity of the ancient books is comprehensively detected from the aspects of chapters and paragraphs in combination with an integrity analysis model, so that the defect that a traditional means depends on manual proofreading or simple page number comparison and keyword extraction, and the efficiency is high is effectively overcome. According to the method, the problems of typesetting confusion, chapter page missing, quotation omission and fuzzy paraphrasing which are common in the ancient books are difficult to deal with, automatic and multi-level analysis of the structural content and semantic level content integrity of the ancient books is realized, and the detection and recognition efficiency of the content integrity of the ancient books is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of information processing technology, specifically to a method and system for detecting the integrity of ancient book content based on semantic structure features. Background Technology

[0002] With the deep integration of digital technology and cultural preservation, information processing technology, especially natural language processing, is becoming an important force in promoting the intelligent inheritance of cultural heritage. Among numerous cultural data resources, ancient books and documents, with their unique historical value, linguistic style, and cultural inheritance status, have become the core objects in the current digital humanities, smart libraries, and cultural knowledge reconstruction. However, to achieve in-depth utilization of ancient books, relying solely on scanning and archiving or keyword retrieval is clearly insufficient. The structural integrity detection of content is gradually becoming a key indicator for measuring high-quality digital ancient book resources.

[0003] However, existing methods for detecting the integrity of ancient texts still have significant shortcomings in assessing the semantic integrity at the structural level. Traditional methods often rely on manual proofreading or simple page number comparison and keyword extraction, which are insufficient to address common issues in ancient texts such as disordered layout, missing pages, omitted citations, and ambiguous interpretations. In particular, they are ineffective in determining whether an ancient text possesses complete learning and inheritance value. More importantly, in the absence of semantic structural understanding, it is difficult to determine the connection between paragraphs and sentences, whether the temporal logic is disordered, and whether terminology lacks definition, thus hindering a comprehensive quality assessment of the content. These technological gaps directly limit the intelligent utilization of ancient text content and reduce the application effectiveness of digital documents in education, research, and cultural dissemination scenarios. Summary of the Invention

[0004] To address the shortcomings of existing technologies, this invention provides a method and system for detecting the integrity of ancient book content based on semantic structural features, thus solving the problems mentioned in the background technology.

[0005] To achieve the above objectives, the present invention provides the following technical solution: a system for detecting the integrity of ancient text content based on semantic structure features, comprising a data acquisition unit, a chapter integrity analysis unit, a paragraph / sentence integrity analysis unit, a comprehensive integrity analysis unit, a result feedback unit, and a display terminal; the data acquisition unit is used to extract chapter structure feature data information and paragraph / sentence structure feature data information of the current ancient text based on the indexed current ancient text content and in conjunction with semantic structure feature recognition technology, and send them to the chapter integrity analysis unit and the paragraph / sentence integrity analysis unit respectively; characterized in that: the chapter integrity analysis unit is used to receive chapter structure feature data information, perform chapter content integrity analysis processing, and generate a high-level chapter integrity signal, a mid-level chapter integrity signal, and a low-level chapter integrity signal, and send them to the comprehensive integrity analysis unit; The paragraph and sentence integrity analysis unit is used to receive paragraph and sentence structure feature data information and perform paragraph and sentence content integrity analysis processing. Based on this, it generates superior signals, medium signals, and poor signals of ancient book paragraph and sentence content integrity, and sends them to the integrity comprehensive analysis unit. The integrity comprehensive analysis unit is used to perform comprehensive integrity analysis on the received chapter integrity level determination signal and paragraph integrity level determination signal, thereby generating a first-level comprehensive integrity feedback signal, a second-level comprehensive integrity feedback signal and a third-level comprehensive integrity feedback signal, and sending them to the result feedback unit. The result feedback unit is used to receive the comprehensive integrity feedback signal of the corresponding level, perform early warning analysis and processing, and send it to the display terminal for display description in the form of text description.

[0006] Preferably, the specific process of the chapter content integrity analysis and processing includes: The semantic structure features of the current ancient text content in the index are identified. The total difference in page numbers, the total number of missing citations, and the frequency of keyword recurrence in each chapter are normalized from the extracted chapter structure feature data. The completeness of each chapter in the current ancient text is analyzed, and the chapter completeness coefficient of each chapter is determined. Specifically: xzj i =1 / (a1×zym i +a2×nqs i +a3×pfx i In the formula, xzj i zym represents the chapter integrity coefficient of the i-th chapter in the current ancient text. i nqs i and PFX i These represent the total difference in page number comparison, the total number of missing citations, and the frequency of keyword recurrence for the i-th chapter in the current ancient text, respectively. Here, / represents a division sign, and a1, a2, and a3 represent the weight values ​​of the total difference in page number comparison, the total number of missing citations, and the frequency of keyword recurrence for the chapter, respectively. The specific values ​​are set by the personnel in this technical field. A numerical model analysis was performed on the chapter integrity coefficient of each chapter of the current ancient text. A two-dimensional coordinate system was established with the chapter integrity coefficient value as the vertical axis and the chapter number as the horizontal axis. The chapter integrity coefficient of each chapter of the current ancient text was marked in the two-dimensional coordinate system in the form of anchor points, and a chapter integrity anchor point marking diagram of the current ancient text was established. By using a pre-set reference line in a two-dimensional coordinate system, chapters with integrity coefficients at or above the reference line are marked as complete chapters, and the number of complete chapters is counted. Chapters with integrity coefficients below the reference line are marked as incomplete chapters, and the number of incomplete chapters is counted. The number of complete chapters is compared and analyzed with the number of incomplete chapters to generate a corresponding completeness-based ranking signal for each chapter. The specific comparison process includes: If the number of complete chapters is greater than the number of incomplete chapters, a high-level complete chapter signal is generated; if the number of complete chapters is equal to the number of incomplete chapters, a medium-level complete chapter signal is generated; and if the number of complete chapters is less than the number of incomplete chapters, a low-level complete chapter signal is generated.

[0007] Preferably, the specific process of the paragraph and sentence content integrity analysis includes: The semantic structure features of the current ancient text content in the index are identified. The total number of missing sentence / segment titles, the total number of time mismatches, and the total number of missing terminology explanations in the extracted sentence / segment structure feature data of the current ancient text are normalized. The sentence / segment integrity of the current ancient text is analyzed to determine the sentence / segment integrity coefficient of the current ancient text. Specifically, xjd=(b1+b2+b3) / (b1×cbt+b2×ccp+b3×csy); where xjd represents the sentence / segment integrity coefficient of the current ancient text, cbt, ccp, and csy represent the total number of missing sentence / segment titles, the total number of time mismatches, and the total number of missing terminology explanations, respectively. / represents a division sign, and b1, b2, and b3 represent the weight values ​​of the total number of missing sentence / segment titles, the total number of time mismatches, and the total number of missing terminology explanations, respectively. The specific values ​​are set by the personnel in this technical field. Pre-set integrity gradient reference thresholds CZ1 and CZ2 for the current ancient text's sentence integrity coefficient, where CZ1 < CZ2. Then, compare and analyze the current ancient text's sentence integrity coefficient with the preset integrity gradient reference thresholds CZ1 and CZ2. The specific comparison process includes: If the sentence integrity coefficient of the current ancient text is greater than the integrity gradient reference threshold CZ2, a signal of excellent sentence integrity is generated. If the sentence integrity coefficient of the current ancient text is between the integrity gradient reference thresholds CZ1 and CZ2, a signal of medium sentence integrity is generated. If the sentence integrity coefficient of the current ancient text is less than the integrity gradient reference threshold CZ1, a signal of poor sentence integrity is generated.

[0008] Preferably, the specific process of comprehensive analysis and processing of the integrity of chapters and paragraphs includes: Based on the received corresponding chapter integrity level determination signals, the chapter high-level integrity signal, the chapter medium-level integrity signal, and the chapter low-level integrity signal are labeled as W1, W2, and W3, respectively; Based on the received corresponding paragraph and sentence integrity level judgment signals, the signals of excellent integrity of ancient book paragraphs and sentences, medium integrity of ancient book paragraphs and sentences, and poor integrity of ancient book paragraphs and sentences are respectively labeled as M1, M2, and M3; The two types of labeled signals are integrated and analyzed. The specific analysis process includes: If two of the two types of annotation signals acquired simultaneously contain the digit "1", a Level 1 integrated integrity feedback signal is generated. If one of the two types of annotation signals acquired simultaneously contains the digit "1", a Level 2 integrated integrity feedback signal is generated. If neither of the two types of annotation signals acquired simultaneously contains the digit "1", a Level 3 integrated integrity feedback signal is generated.

[0009] Preferably, the specific process of the early warning analysis and processing includes: When a Level 1 comprehensive integrity feedback signal is received, a text description is generated stating "The chapters and paragraphs of the current ancient text are relatively complete and can be directly used for subsequent learning and reference," and sent to the display terminal for display explanation; When a Level 2 comprehensive integrity feedback signal is received, a text message is generated stating "The current chapters or paragraphs of the ancient text are incomplete and will not be used for subsequent learning or reference," and this message is sent to the display terminal for display explanation. When a Level 3 comprehensive integrity feedback signal is received, a text message is generated stating "The current chapters and paragraphs of the ancient text are incomplete and cannot be used for subsequent learning and reference," and then sent to the display terminal for display explanation.

[0010] A method for detecting the integrity of ancient book content based on semantic structure features includes the following steps: S1. Based on the current ancient text content indexed and combined with semantic structure feature recognition technology, extract the chapter structure feature data information and paragraph structure feature data information of the current ancient text. S2. Receive chapter structure feature data information and perform chapter content integrity analysis and processing to generate chapter high-level integrity signal, chapter medium-level integrity signal and chapter low-level integrity signal; S3. Receive paragraph and sentence structure feature data information, and perform paragraph and sentence content integrity analysis and processing to generate superior signals, medium signals, and poor signals of ancient book paragraph and sentence content integrity. S4. Perform a comprehensive analysis and processing of the integrity of the received complete chapter and paragraph / sentence determination signals, and generate a first-level comprehensive integrity feedback signal, a second-level comprehensive integrity feedback signal, and a third-level comprehensive integrity feedback signal accordingly. S5. Receive the comprehensive integrity feedback signal of the corresponding level, perform early warning analysis and processing, and send it to the display terminal for display description in the form of text description.

[0011] This invention provides a method and system for detecting the integrity of ancient book content based on semantic structural features, which has the following beneficial effects: (1) By constructing a method for detecting the integrity of ancient book content based on semantic structure features, a closed-loop technical process has been formed in terms of structure extraction, content evaluation and result feedback. This has enabled the automation, standardization and multi-level analysis and feedback of the integrity of ancient book content at the structural and semantic levels. The semantic structure recognition technology is used to accurately extract the structural features of chapters and paragraphs. Combined with the integrity analysis model, the chapter content is evaluated with high precision from the dimensions of chapter page number comparison, missing citations and keyword recurrence frequency. At the same time, parameters such as missing titles, time mismatch and terminology definition are introduced at the paragraph level, which effectively overcomes the limitations of relying solely on keywords or page numbers in traditional methods. In the comprehensive analysis stage, multi-level judgment signals of chapters and paragraphs are cross-compared to form first-level, second-level and third-level integrity feedback results, which effectively support the classification and utilization of ancient books with different integrity levels in education, research, learning and digital dissemination scenarios. This not only fundamentally solves the problems of storing but not using and using but not accurately in the digitization of traditional ancient books, but also improves the efficiency and value of intelligent utilization of ancient books.

[0012] (2) In the traditional process of digitizing ancient books, the problems of missing chapters, disordered chapters, and missing citations have long been difficult to identify automatically. It often requires manual reading and comparison of multiple versions, which is time-consuming, laborious, and easily affected by subjective judgment. To address this problem, this solution introduces a modeling and extraction mechanism for chapter structure feature data information. It performs unified normalization analysis on the dimensions of structural page number differences, total number of missing citations, and keyword recurrence frequency of each chapter in ancient books, and calculates the chapter integrity coefficient. Furthermore, it constructs a two-dimensional coordinate annotation map and reference line comparison model to accurately determine the completeness of each chapter. By comparing the number of complete chapters with the number of incomplete chapters, it automatically generates a chapter integrity level signal, which effectively solves the problems of difficult chapter structure identification and difficult automatic judgment of missing pages and citations. This lays the foundation for subsequent paragraph semantic judgment and full-text quality assessment, and fills the gap in the current technical means for intelligent judgment of chapter structure.

[0013] (3) By designing an independent identification pathway for the structural feature data information of paragraphs and sentences, three rare parameters with great semantic depth are introduced: the number of missing paragraph titles, the total number of time mismatches, and the total number of missing terminology explanations. These parameters are then fused and normalized to construct a paragraph integrity coefficient model. By comparing it with a dual threshold reference, the model can quickly determine whether the content is of superior, medium, or poor integrity level, accurately reflecting the content coherence and expression completeness between paragraphs and sentences, thus avoiding the problem of misjudgment based solely on shallow rules such as word count or punctuation. Attached Figure Description

[0014] Figure 1 This is a block diagram of an ancient book content integrity detection system based on semantic structure features according to the present invention; Figure 2 This is a schematic diagram of the process for a method to detect the integrity of ancient book content based on semantic structure features according to the present invention. Detailed Implementation

[0015] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0016] Example 1 Please see Figure 1 This invention provides a system for detecting the integrity of ancient book content based on semantic structure features, including a data acquisition unit, a chapter integrity analysis unit, a paragraph / sentence integrity analysis unit, a comprehensive integrity analysis unit, a result feedback unit, and a display terminal; The data acquisition unit is used to extract the chapter structure feature data information and paragraph structure feature data information of the current ancient book text based on the indexed current ancient book text content and combined with semantic structure feature recognition technology, and send them to the chapter integrity analysis unit and paragraph integrity analysis unit respectively. The chapter integrity analysis unit receives chapter structure feature data and performs chapter content integrity analysis, thereby generating a high-level chapter integrity signal, a mid-level chapter integrity signal, and a low-level chapter integrity signal, which are then sent to the integrity comprehensive analysis unit. The specific process includes: The semantic structure features of the current ancient text content in the index are identified. The total difference in page numbers, the total number of missing citations, and the frequency of keyword recurrence in each chapter are normalized from the extracted chapter structure feature data. The completeness of each chapter in the current ancient text is analyzed, and the chapter completeness coefficient of each chapter is determined. Specifically: xzji =1 / (a1×zym i +a2×nqs i +a3×pfx i In the formula, xzj i zym represents the chapter integrity coefficient of the i-th chapter in the current ancient text. i nqs i and PFX i These represent the total difference in page number comparison, the total number of missing citations, and the frequency of keyword recurrence for the i-th chapter in the current ancient text, respectively. Here, / represents a division sign, and a1, a2, and a3 represent the weight values ​​of the total difference in page number comparison, the total number of missing citations, and the frequency of keyword recurrence for the chapter, respectively. The specific values ​​are set by the personnel in this technical field. It should be noted that the total difference in chapter page number comparison represents the cumulative deviation between the current page number order of each chapter in the ancient book and the standard index page number, reflecting the completeness and accuracy of the chapter arrangement; the total number of missing citations represents the total number of citations, classical references, or reference statements that should appear in the chapter but do not, measuring the completeness of the text citation system; and the frequency of recurrence of chapter keywords represents the frequency of repetition of core keywords within the specified semantic range in the chapter, used to assess the redundancy of the chapter's thematic expression.

[0017] A numerical model analysis was performed on the chapter integrity coefficient of each chapter of the current ancient text. A two-dimensional coordinate system was established with the chapter integrity coefficient value as the vertical axis and the chapter number as the horizontal axis. The chapter integrity coefficient of each chapter of the current ancient text was marked in the two-dimensional coordinate system in the form of anchor points, and a chapter integrity anchor point marking diagram of the current ancient text was established. By using a pre-set reference line in a two-dimensional coordinate system, chapters with integrity coefficients at or above the reference line are marked as complete chapters, and the number of complete chapters is counted. Chapters with integrity coefficients below the reference line are marked as incomplete chapters, and the number of incomplete chapters is counted. The number of complete chapters is compared and analyzed with the number of incomplete chapters to generate a corresponding completeness-based ranking signal for each chapter. The specific comparison process includes: If the number of complete chapters is greater than the number of incomplete chapters, a high-level complete chapter signal is generated; if the number of complete chapters is equal to the number of incomplete chapters, a medium-level complete chapter signal is generated; and if the number of complete chapters is less than the number of incomplete chapters, a low-level complete chapter signal is generated.

[0018] The paragraph and sentence integrity analysis unit receives paragraph and sentence structure feature data and performs paragraph and sentence content integrity analysis processing. Based on this, it generates signals of excellent, medium, and poor integrity of ancient text paragraphs and sentences, and sends them to the integrity comprehensive analysis unit. The specific process includes: The semantic structure features of the current ancient text content in the index are identified. The total number of missing sentence / segment titles, the total number of time mismatches, and the total number of missing terminology explanations in the extracted sentence / segment structure feature data of the current ancient text are normalized. The sentence / segment integrity of the current ancient text is analyzed to determine the sentence / segment integrity coefficient of the current ancient text. Specifically, xjd=(b1+b2+b3) / (b1×cbt+b2×ccp+b3×csy); where xjd represents the sentence / segment integrity coefficient of the current ancient text, cbt, ccp, and csy represent the total number of missing sentence / segment titles, the total number of time mismatches, and the total number of missing terminology explanations, respectively. / represents a division sign, and b1, b2, and b3 represent the weight values ​​of the total number of missing sentence / segment titles, the total number of time mismatches, and the total number of missing terminology explanations, respectively. The specific values ​​are set by the personnel in this technical field. It should be noted that the total number of missing sentence / paragraph titles represents the cumulative number of times that a paragraph or sentence in an ancient text should have a title or section marker but is actually missing, reflecting the degree of incompleteness of structural guidance information; the total number of time mismatches represents the cumulative number of times that the time descriptions in the text have logical inconsistencies, disordered order, or contradictory dates, used to measure the accuracy of the content's time logic; and the total number of missing terminology explanations represents the cumulative number of times that professional terms and historical terms appear in the text without accompanying explanations, used to assess the clarity of knowledge expression.

[0019] Pre-set integrity gradient reference thresholds CZ1 and CZ2 for the current ancient text's sentence integrity coefficient, where CZ1 < CZ2. Then, compare and analyze the current ancient text's sentence integrity coefficient with the preset integrity gradient reference thresholds CZ1 and CZ2. The specific comparison process includes: If the sentence integrity coefficient of the current ancient text is greater than the integrity gradient reference threshold CZ2, a signal of excellent sentence integrity is generated. If the sentence integrity coefficient of the current ancient text is between the integrity gradient reference thresholds CZ1 and CZ2, a signal of medium sentence integrity is generated. If the sentence integrity coefficient of the current ancient text is less than the integrity gradient reference threshold CZ1, a signal of poor sentence integrity is generated.

[0020] The integrity comprehensive analysis unit is used to perform comprehensive integrity analysis on the received chapter integrity level determination signal and paragraph integrity level determination signal, thereby generating a first-level comprehensive integrity feedback signal, a second-level comprehensive integrity feedback signal, and a third-level comprehensive integrity feedback signal, and sending them to the result feedback unit. The specific process includes: Based on the received corresponding chapter integrity level determination signals, the chapter high-level integrity signal, the chapter medium-level integrity signal, and the chapter low-level integrity signal are labeled as W1, W2, and W3, respectively; Based on the received corresponding paragraph and sentence integrity level judgment signals, the signals of excellent integrity of ancient book paragraphs and sentences, medium integrity of ancient book paragraphs and sentences, and poor integrity of ancient book paragraphs and sentences are respectively labeled as M1, M2, and M3; The two types of labeled signals are integrated and analyzed. The specific analysis process includes: If two of the two types of annotation signals acquired simultaneously contain the digit "1", a Level 1 integrated integrity feedback signal is generated. If one of the two types of annotation signals acquired simultaneously contains the digit "1", a Level 2 integrated integrity feedback signal is generated. If neither of the two types of annotation signals acquired simultaneously contains the digit "1", a Level 3 integrated integrity feedback signal is generated.

[0021] The result feedback unit is used to receive the comprehensive integrity feedback signal of the corresponding level, perform early warning analysis and processing, and send it to the display terminal for display description in the form of text. The specific process includes: When a Level 1 comprehensive integrity feedback signal is received, a text description is generated stating "The chapters and paragraphs of the current ancient text are relatively complete and can be directly used for subsequent learning and reference," and sent to the display terminal for display explanation; When a Level 2 comprehensive integrity feedback signal is received, a text message is generated stating "The current chapters or paragraphs of the ancient text are incomplete and will not be used for subsequent learning or reference," and this message is sent to the display terminal for display explanation. When a Level 3 comprehensive integrity feedback signal is received, a text message is generated stating "The current chapters and paragraphs of the ancient text are incomplete and cannot be used for subsequent learning and reference," and then sent to the display terminal for display explanation.

[0022] Example 2 Please see Figure 2 A method for detecting the integrity of ancient book content based on semantic structure features includes the following steps: S1. Based on the current ancient text content indexed and combined with semantic structure feature recognition technology, extract the chapter structure feature data information and paragraph structure feature data information of the current ancient text. S2. Receive chapter structure feature data information and perform chapter content integrity analysis and processing to generate chapter high-level integrity signal, chapter medium-level integrity signal and chapter low-level integrity signal; S3. Receive paragraph and sentence structure feature data information, and perform paragraph and sentence content integrity analysis and processing to generate superior signals, medium signals, and poor signals of ancient book paragraph and sentence content integrity. S4. Perform a comprehensive analysis and processing of the integrity of the received complete chapter and paragraph / sentence determination signals, and generate a first-level comprehensive integrity feedback signal, a second-level comprehensive integrity feedback signal, and a third-level comprehensive integrity feedback signal accordingly. S5. Receive the comprehensive integrity feedback signal of the corresponding level, perform early warning analysis and processing, and send it to the display terminal for display description in the form of text description.

[0023] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. A system for detecting the integrity of ancient text content based on semantic structure features, comprising a data acquisition unit, a chapter integrity analysis unit, a paragraph / sentence integrity analysis unit, a comprehensive integrity analysis unit, a result feedback unit, and a display terminal; the data acquisition unit is used to extract chapter structure feature data information and paragraph / sentence structure feature data information of the current ancient text based on the indexed current ancient text content and in conjunction with semantic structure feature recognition technology, and send them respectively to the chapter integrity analysis unit and the paragraph / sentence integrity analysis unit; characterized in that: The chapter integrity analysis unit is used to receive chapter structure feature data information and perform chapter content integrity analysis processing, thereby generating a high-level chapter integrity signal, a medium-level chapter integrity signal and a low-level chapter integrity signal, and sending them to the integrity comprehensive analysis unit. The paragraph and sentence integrity analysis unit is used to receive paragraph and sentence structure feature data information and perform paragraph and sentence content integrity analysis processing. Based on this, it generates superior signals, medium signals, and poor signals of ancient book paragraph and sentence content integrity, and sends them to the integrity comprehensive analysis unit. The integrity comprehensive analysis unit is used to perform comprehensive integrity analysis on the received chapter integrity level determination signal and paragraph integrity level determination signal, thereby generating a first-level comprehensive integrity feedback signal, a second-level comprehensive integrity feedback signal and a third-level comprehensive integrity feedback signal, and sending them to the result feedback unit. The result feedback unit is used to receive the comprehensive integrity feedback signal of the corresponding level, perform early warning analysis and processing, and send it to the display terminal for display description in the form of text description.

2. The ancient book content integrity detection system based on semantic structure features according to claim 1, characterized in that: The specific process for analyzing and processing the integrity of the chapter content includes: The semantic structure features of the current ancient text content in the index are identified. The total difference in page numbers, the total number of missing citations, and the frequency of keyword recurrence in each chapter are normalized from the extracted chapter structure feature data. The completeness of each chapter in the current ancient text is analyzed, and the chapter completeness coefficient of each chapter is determined. Specifically: xzj i =1 / (a1×zym i +a2×nqs i +a3×pfx i In the formula, xzj i zym represents the chapter integrity coefficient of the i-th chapter in the current ancient text. i nqs i and PFX i These represent the total difference in page number comparison, the total number of missing citations, and the frequency of keyword recurrence for the i-th chapter in the current ancient text, respectively. Here, / represents a division sign, and a1, a2, and a3 represent the weight values ​​of the total difference in page number comparison, the total number of missing citations, and the frequency of keyword recurrence for the chapter, respectively. The specific values ​​are set by the personnel in this technical field. A numerical model analysis was performed on the chapter integrity coefficient of each chapter of the current ancient text. A two-dimensional coordinate system was established with the chapter integrity coefficient value as the vertical axis and the chapter number as the horizontal axis. The chapter integrity coefficient of each chapter of the current ancient text was marked in the two-dimensional coordinate system in the form of anchor points, and a chapter integrity anchor point marking diagram of the current ancient text was established. By using a pre-set reference line in a two-dimensional coordinate system, chapters with integrity coefficients at or above the reference line are marked as complete chapters, and the number of complete chapters is counted. Chapters with integrity coefficients below the reference line are marked as incomplete chapters, and the number of incomplete chapters is counted. The number of complete chapters is compared and analyzed with the number of incomplete chapters to generate a corresponding completeness-based ranking signal for each chapter. The specific comparison process includes: If the number of complete chapters is greater than the number of incomplete chapters, a high-level complete chapter signal is generated; if the number of complete chapters is equal to the number of incomplete chapters, a medium-level complete chapter signal is generated; and if the number of complete chapters is less than the number of incomplete chapters, a low-level complete chapter signal is generated.

3. The ancient book content integrity detection system based on semantic structure features according to claim 1, characterized in that: The specific process of analyzing and processing the integrity of paragraph and sentence content includes: The semantic structure features of the current ancient text content in the index are identified. The total number of missing sentence / segment titles, the total number of time mismatches, and the total number of missing terminology explanations in the extracted sentence / segment structure feature data of the current ancient text are normalized. The sentence / segment integrity of the current ancient text is analyzed to determine the sentence / segment integrity coefficient of the current ancient text. Specifically, xjd=(b1+b2+b3) / (b1×cbt+b2×ccp+b3×csy); where xjd represents the sentence / segment integrity coefficient of the current ancient text, cbt, ccp, and csy represent the total number of missing sentence / segment titles, the total number of time mismatches, and the total number of missing terminology explanations, respectively. / represents a division sign, and b1, b2, and b3 represent the weight values ​​of the total number of missing sentence / segment titles, the total number of time mismatches, and the total number of missing terminology explanations, respectively. The specific values ​​are set by the personnel in this technical field. Pre-set integrity gradient reference thresholds CZ1 and CZ2 for the current ancient text's sentence integrity coefficient, where CZ1 < CZ2. Then, compare and analyze the current ancient text's sentence integrity coefficient with the preset integrity gradient reference thresholds CZ1 and CZ2. The specific comparison process includes: If the sentence integrity coefficient of the current ancient text is greater than the integrity gradient reference threshold CZ2, a signal of excellent sentence integrity is generated. If the sentence integrity coefficient of the current ancient text is between the integrity gradient reference thresholds CZ1 and CZ2, a signal of medium sentence integrity is generated. If the sentence integrity coefficient of the current ancient text is less than the integrity gradient reference threshold CZ1, a signal of poor sentence integrity is generated.

4. The ancient book content integrity detection system based on semantic structure features according to claim 1, characterized in that: The specific process of comprehensive analysis and processing of the integrity of the chapters and paragraphs includes: Based on the received corresponding chapter integrity level determination signals, the chapter high-level integrity signal, the chapter medium-level integrity signal, and the chapter low-level integrity signal are labeled as W1, W2, and W3, respectively; Based on the received corresponding paragraph and sentence integrity level judgment signals, the signals of excellent integrity of ancient book paragraphs and sentences, medium integrity of ancient book paragraphs and sentences, and poor integrity of ancient book paragraphs and sentences are respectively labeled as M1, M2, and M3; The two types of labeled signals are integrated and analyzed. The specific analysis process includes: If two of the two types of annotation signals acquired simultaneously contain the digit "1", a Level 1 integrated integrity feedback signal is generated. If one of the two types of annotation signals acquired simultaneously contains the digit "1", a Level 2 integrated integrity feedback signal is generated. If neither of the two types of annotation signals acquired simultaneously contains the digit "1", a Level 3 integrated integrity feedback signal is generated.

5. The ancient book content integrity detection system based on semantic structure features according to claim 1, characterized in that: The specific process of the early warning analysis and processing includes: When a Level 1 comprehensive integrity feedback signal is received, a text description is generated stating "The chapters and paragraphs of the current ancient text are relatively complete and can be directly used for subsequent learning and reference," and then sent to the display terminal for display explanation. When a Level 2 comprehensive integrity feedback signal is received, a text message is generated stating "The current chapter or paragraph of the ancient text is incomplete and will not be used for subsequent learning or reference at this time," and sent to the display terminal for display explanation. When a Level 3 comprehensive integrity feedback signal is received, a text message is generated stating "The current chapters and paragraphs of the ancient text are incomplete and cannot be used for subsequent learning and reference," and then sent to the display terminal for display explanation.

6. A method for detecting the integrity of ancient book content based on semantic structure features, wherein the method is applied to the ancient book content integrity detection system based on semantic structure features as described in any one of claims 1 to 5, characterized in that: Includes the following steps: S1. Based on the current ancient text content indexed and combined with semantic structure feature recognition technology, extract the chapter structure feature data information and paragraph structure feature data information of the current ancient text. S2. Receive chapter structure feature data information and perform chapter content integrity analysis and processing to generate chapter high-level integrity signal, chapter medium-level integrity signal and chapter low-level integrity signal; S3. Receive paragraph and sentence structure feature data information, and perform paragraph and sentence content integrity analysis and processing to generate superior signals, medium signals, and poor signals of ancient book paragraph and sentence content integrity. S4. Perform a comprehensive analysis and processing of the integrity of the received complete chapter and paragraph / sentence determination signals, and generate a first-level comprehensive integrity feedback signal, a second-level comprehensive integrity feedback signal, and a third-level comprehensive integrity feedback signal accordingly. S5. Receive the comprehensive integrity feedback signal of the corresponding level, perform early warning analysis and processing, and send it to the display terminal for display description in the form of text description.