A method for automatic processing of Chinese corpus cleaning and labeling

By combining text morpheme-cell network and self-attention hypergraph network model, the problem of the disconnect between Chinese corpus cleaning and annotation of text traces and annotation clues is solved, achieving high accuracy and consistency in corpus processing, which is suitable for digital reading and large language model training.

CN122489679APending Publication Date: 2026-07-31GUANGXI WEISHUFANG DIGITAL TECHNOLOGY CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
GUANGXI WEISHUFANG DIGITAL TECHNOLOGY CO LTD
Filing Date
2026-05-25
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

Existing Chinese corpus cleaning and annotation methods struggle to preserve textual traces and annotation clues in complex corpus scenarios, resulting in a disconnect between cleaning and annotation results. Furthermore, traditional methods are prone to misjudging or accidentally deleting valid corpus data during deduplication and annotation.

Method used

We employ a method that combines a morpheme-cell network with an improved self-attention hypergraph network model. By encapsulating corpus granules, crimping morpheme structures, and cross-layer labeling associations, we introduce a break-bridge module and a break-labeling mechanism to construct a corpus structure that carries textual continuity, source traces, and labeling clues. We then perform dimension-by-dimensional subtraction and closure processing.

Benefits of technology

It improves the accuracy of Chinese corpus cleaning and the consistency of annotation, forming a stable, high-quality corpus suitable for large-scale corpus construction, and reduces the problem of accidental deletion or retention.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122489679A_ABST
    Figure CN122489679A_ABST
Patent Text Reader

Abstract

This invention discloses an automated method for cleaning and annotating Chinese corpora, comprising the following steps: S1, organizing multi-source Chinese corpora and encapsulating corpus fragments with source traces; S2, segmenting corpus fragments and pressing them together into a morpheme structure; S3, labeling cross-layer cleaning-labeling associations and connecting them into a morpheme-cavity network; S4, feeding the data into an improved self-attention hypergraph network model, where a morpheme embedding module forms morpheme embedding states; S5, a local hyperedge extraction module connects local hyperedge groups; S6, a break bridging module, combined with a break-labeling locking mechanism, forms break-resonance hyperedge features; S7, a global hyperedge attention module aggregates features to create a Chinese corpus with cleaning-labeling closure markers. This invention uses a morpheme-cavity network and an improved self-attention hypergraph network model to achieve automatic cleaning, annotation locking, and quality closure management of Chinese corpora.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of natural language processing and corpus data processing technology, and in particular to an automated method for cleaning and annotating Chinese corpora. Background Technology

[0002] Existing Chinese corpus cleaning and annotation methods typically employ a linear process: first, format standardization, noise filtering, and sentence segmentation are performed; then, deduplication is achieved through similarity calculation, keyword matching, or semantic vector comparison; and finally, labels are added using classification models or manual methods. While such methods can complete the basic cleaning, the cleaning results, source records, and annotation results are fragmented, making it difficult to preserve textual traces formed from different sources, versions, and contexts.

[0003] With the increasing demand for high-quality Chinese corpora in digital reading, smart education, and large language model training, existing methods are clearly insufficient in complex corpus scenarios. Common issues in Chinese texts include reprinting and rewriting, excerpting and splicing, line breaks and misjoints, reuse of template sentences, and tag conflicts. Traditional similarity-based deduplication easily misclassifies common sentence structures as duplicate content and tends to remove rewritten segments with version value. Existing annotation models typically output labels directly after inputting the entire text, lacking internal checks for textual coherence, source traces, and misaligned annotation clues, making it difficult to promptly offload and review erroneous annotations.

[0004] Furthermore, existing hypergraph attention or text classification models mostly adopt a transmission path where local features directly enter global attention, lacking a bridging processing structure for corpus breaks, which affects the accuracy of Chinese corpus cleaning, annotation consistency, and quality closure management.

[0005] Therefore, how to provide an automated method for cleaning and annotating Chinese corpora is a problem that urgently needs to be solved by those skilled in the art. Summary of the Invention

[0006] One objective of this invention is to propose an automated method for cleaning and annotating Chinese corpora. This invention employs a combination of a morpheme-cell network and an improved self-attention hypergraph network model. Through corpus granular encapsulation, morpheme structure compression, and cross-layer cleaning-annotation association, a corpus structure carrying textual continuity, source traces, and annotation clues is constructed. Furthermore, a break-edge bridging module and a break-edge locking mechanism are introduced to perform dimensional subtraction, co-positioning, misalignment stripping, and closure latching on local hyperedge groups, achieving collaborative processing of noise disconnection, source misalignment, annotation suspension, and version differences. Based on this, a global hyperedge attention module aggregates break-edge resonance features to form a Chinese corpus with cleaning-annotation closure markers. This method boasts advantages such as high cleaning accuracy, strong annotation consistency, stable quality closure management, and applicability to large-scale Chinese corpus construction.

[0007] An automated method for cleaning and annotating Chinese corpora according to an embodiment of the present invention includes the following steps: S1. Rectify the text format of multi-source Chinese corpus, remove noisy content and cut sentence boundaries, and encapsulate the remaining text into corpus fragments with source traces; S2. Segment the textual hierarchy within the corpus fragments and compress the textual relationships, source traces, and annotation clues into a textual structure; S3. Determine the cross-layer clearing associations along the morphological structure and connect the cross-layer clearing associations into a morphological cell network; S4. Feed the text cell cavity network into the improved self-attention hypergraph network model. The improved self-attention hypergraph network model includes a text embedding module, a local hyperedge extraction module, a break bridging module, and a global hyperedge attention module. The text embedding module extracts the connecting chain positions in the text cell cavity network and arranges them into text embedding states. S5. The local superedge extraction module extracts local chain windows along the receiving chain position of the element embedding state, and delimits the clearing associations that can be closed within the local chain window into local superedge groups. S6. The fracture bridging module accepts the local super-edge group and embeds the fracture locking mechanism. It disassembles the super-edge pieces of the same boundary in the local super-edge group dimension by dimension, snaps the closable fracture back to the adjacent chain position, peels the misaligned fracture into the core fracture, and encapsulates it as the fracture resonance super-edge feature with a closed latch. S7. The global superedge attention module collects fracture resonance superedge features along the superedge correlation, and reloads the cleared closure markers and unclosed chain positions into the corresponding corpus particles to create a Chinese corpus.

[0008] Optionally, S1 specifically includes: S11. Read the original text and source traces of multi-source Chinese corpus, and paste the source traces to the text boundaries and paragraph positions of the original text to form the original text with traces; S12. Reorganize the text form of the original text with traces, and incorporate the format misalignment, layout residue and line break misconnection caused by cross sources into the same text continuation chain to form a reorganized text. S13. Remove noise fragments that cannot be connected to the sentence boundary along the text continuation chain of the consolidated text, and retain the source traces of adjacent positions of the noise fragments to form the purified text. S14. Cut the sentence boundaries according to the semantic pauses and paragraph connections of the purified text, and encapsulate the cut-off preserved text and source traces into corpus fragments.

[0009] Optionally, S2 specifically includes: S21. Cut the text fragments along the semantic pauses and contextual connections of the corpus fragments, and merge fragments that cannot be connected independently into adjacent text fragments to form continuous text fragments. S22. Arrange the successor links according to the sequential position of continuous morphemes in the corpus particles, and mark the break, segment crossing and source transition positions as the link boundaries; S23. Extract semantic clues from continuous text fragments to participate in annotation judgment, embed the semantic clues into the corresponding chain positions, and paste the source traces back to the chain position boundary. S24. Encapsulate the continuous fragments of text that have completed chain position arrangement, semantic clue embedding, and source trace re-attachment into a text structure.

[0010] Optionally, S3 specifically includes: S31. Search for the locations where text continuity is broken, source traces are suspended, and annotation clues are mismatched along the continuity chain of the text structure, and merge the found chain into a cleared chain to be connected. S32. Using the chain to be linked as the center, pull the adjacent text structure, check the text continuation direction, source attribution relationship and annotation placement relationship of adjacent chain positions, and fold the chain positions that can complement each other into cross-layer clearing association. S33. Link the cross-layer clearing markers back to the corresponding element structure, and press the incomplete chain positions into the break position to form an element cell network.

[0011] Optionally, S4 specifically includes: S41. Expand the network links along the receiving direction of the morpheme cavity network, and gather the morpheme structures, cross-layer clearing associations and break points at the same receiving position into the links to be embedded. S42. Based on the chain position boundary of the chain piece to be embedded, the text is sorted and preserved. The source trace is pasted back to the start and end boundaries of the text. The location of the embedded text is marked. The chain position slot is formed with the boundary constraints. S43. The chain slot is compared with the adjacent chain slot in terms of the direction of continuation. The continuous content is retained in the chain slot. The broken content caused by line breakage, semantic suspension and source separation is pressed into the corresponding chain slot boundary to form a chain slot breakage groove. S44. The content that is completed and continued in the chain slot is unfolded into a text base piece in the order of continuation, and the source traces that fit the boundary and the marking clues that fall into the belonging position are folded into an attachment piece. S45. The text substrate is converted into a chain bit basis vector, the attachment is converted into a boundary attachment vector, and the boundary attachment vector is superimposed on the corresponding boundary of the chain bit basis vector to form a chain bit embedding vector. S46. The chain position break groove is converted into a break embedding piece. The break embedding piece is then attached back to the break boundary of the chain position embedding vector. The chain position embedding vector is arranged in the order of the chain positions in the network to form a textual embedding state.

[0012] Optionally, S5 specifically includes: S51. The local super-edge extraction module receives the element embedding state, opens the link embedding vector along the link position sequence in the network, and gathers the current link position, the preceding link position and the following link position into a local link window. The break boundary and break embedding piece of the link embedding vector are brought into the local link window along with the link position. S52. Compare the text continuation direction along the preceding and following boundaries within the local chain window. Chain position embedding vectors with the same continuation direction are joined together along the boundary to form a clean text super-edge piece. Chain position embedding vectors with broken continuation directions do not participate in the joining and leave a break mark at the break boundary. S53. The source trace pulls the homologous chain position embedding vector back to the break mark position. When the source boundary can fit the break boundary, the homologous chain position embedding vector is folded into the source trace super edge piece. When the source boundary cannot fit the break boundary, the original break mark connection is retained in the break mark position. S54. When the labeling clue guides the same labeling chain embedding vector to get closer to the text belonging position, when the labeling belonging can fall into the corresponding text boundary, the same labeling chain embedding vector is folded into the labeling hyperedge piece; when the labeling belonging deviates from the corresponding text boundary, the deviated part is labeled as the labeling dangling piece. S55, the clean text super edge piece, the source trace super edge piece and the symbolic super edge piece are arranged in the same position along the same chain position boundary. The break mark retention, the original break mark connection and the symbolic suspended piece are pressed into the corresponding boundary positions to form the super edge piece to be bridged. S56. The superedge pieces to be bridged are encapsulated into local superedge groups according to the sequence of morpheme embedding states.

[0013] Optionally, S6 specifically includes: S61, the break bridging module accepts the local super-edge group, cuts off the super-edge piece to be bridged at the bridging chain position, and reads the chain position boundary, break mark position, original break mark connection and marked suspended piece carried by the super-edge piece to be bridged. S62. Arrange the super-edge pieces to be bridged in the same position according to the same chain position boundary. Place the clean text super-edge piece into the main bridging slot, place the source trace super-edge piece into the source side slot, and place the label super-edge piece into the label side slot to form a three-slot bridging piece. S63. The source side groove moves back towards the bridging main groove along the source trace, and the label side groove moves closer to the bridging main groove along the text belonging position. The three-groove bridging pieces where the source trace and the text belonging position fall into the same chain position boundary are pressed together to form a co-position bridging piece. The source side groove that deviates from the chain position boundary is peeled out to form a source trace broken piece, and the label side groove that deviates from the text belonging position is peeled out to form a label suspended piece. S64. The bridging main slot in the same position bridging piece is degraded dimensionally with the source side slot and the label side slot respectively. The degradation result is pressed into the same chain position boundary to form the source text break piece and the label text break piece, and then stacked as a break bridging piece. S65, the fault mark locking mechanism completes the snapping, peeling and latching processes for the fault bridging piece, the source fault disconnect piece and the mark suspension piece, and outputs the closed bridging piece and the verification fault piece; S66. The closed bridging pieces are rearranged according to the original acceptance sequence of the local super-edge group, and the broken piece is reattached to the corresponding chain position boundary to form a broken resonance super-edge feature.

[0014] Optionally, S65 specifically includes: S651, the fracture bridging module has an embedded fracture locking mechanism that is connected in series in the order of fracture differential unit, co-positioning engagement unit, misalignment peeling unit and closing latch unit, and supports fracture bridging piece, source fracture break piece and mark suspension piece. S652. The fracture differential unit unfolds the source text fracture piece and the label text fracture piece along the chain position boundary of the fracture bridging piece, reads the direction state of the bridging main groove and the source side groove after the difference is removed in the source text fracture piece, reads the direction state of the bridging main groove and the label side groove after the difference is removed in the label text fracture piece, presses the fracture pieces with the same direction state in the same chain position boundary into the same position differential groove, and presses the fracture pieces with opposite direction state or staggered chain position boundaries into the misaligned temporary storage groove. S653, the corresponding snap-fit ​​unit pulls the corresponding differential slot to fit the adjacent chain position boundary. When the text receiving direction is continuous, the source return direction falls back to the same source trace, and the label placement direction is embedded in the same text belonging position, the corresponding differential slot is snapped back to the original chain position boundary and formed into a closable break piece. When it cannot fit in any direction, the corresponding differential slot is transferred to the misalignment temporary storage slot. S654, the misalignment stripping unit receives the misalignment temporary storage slot, the source trace break piece, and the marked suspended piece. It compares the source break position with the marked suspended position along the chain position boundary. When the source break position and the marked suspended position fall on the same chain position boundary, they are pressed together to form a verification break piece. When the source break position can be backed up but the marked suspended position cannot be embedded, the marked suspended piece is stripped. When the marked suspended position can be embedded but the source break position cannot be backed up, the source trace break piece is stripped. S655, the closing latch unit reattaches the closable break plate to the original chain position boundary of the break bridging plate, and presses the reattached source text break plate, label text break plate and the bridging main groove, source side groove and label side groove in the same position bridging plate into a closed bridging groove, and the closed bridging groove is connected in series along the original receiving direction to form a closed bridging plate. S656. When verifying the chain position boundary of the closed latch unit, if there are residual misalignment temporary storage slots, source trace broken pieces, or marked suspended pieces within the chain position boundary, peel the residual part into the broken piece for verification. If there are no misalignment residues within the chain position boundary, retain the corresponding chain position, source back-attachment relationship, and marked placement relationship in the closed bridging piece.

[0015] Optionally, S7 specifically includes: S71. The global super-edge attention module receives the fracture resonance super-edge feature, disassembles the closed bridging piece and the core fracture piece along the chain position boundary, arranges the closed bridging pieces into a global attention chain, and retains the core fracture piece at the corresponding chain position boundary. S72. Select closed bridging pieces one by one along the global attention chain as target super-edge pieces, read the bearing direction, source back-attachment relationship and label placement relationship of the target super-edge pieces, and compare them with the corresponding relationship of the remaining closed bridging pieces to form super-edge bonding marks. S73. Closed bridging pieces with the same traction direction as the super-edge bonding mark converge towards the target super-edge piece. Closed bridging pieces that can complete the receiving closure and source return are incorporated into the clearing closure piece. Closed bridging pieces that cannot complete the bonding retain the original chain position boundary. S74. The verification fracture piece is reattached along the original chain position boundary to the adjacent clearing and closing piece. The verification fracture piece does not participate in the annotation locking, and the fracture position, source boundary and annotation placement offset are retained. S75. The clearing and closing pieces are backfilled into the corresponding textual structures. The marking clues that have completed the closure are locked as clearing and closing marks. The chain position boundaries that have not completed the closure receive the verification break pieces and are marked as verification chain positions. S76. The morpheme structure with clearing and closing marks is reassembled according to the corpus particle boundaries, and the verification chain position is attached along with the corresponding source trace to form a Chinese corpus with clearing and closing marks.

[0016] The beneficial effects of this invention are: This invention constructs a morpheme-cell network to unify textual relationships, source traces, and annotation clues from multi-source Chinese corpora into a single corpus structure, thus eliminating text cleaning, source tracing, and annotation as separate processing steps. Compared to traditional methods that rely solely on format cleaning, similarity deduplication, and overall classification annotation, this invention preserves source boundaries and version traces at the corpus granular level, structurally addressing issues such as line breaks, noise residue, excerpt rewriting, and annotation mismatches, reducing the problem of erroneously deleted or retained duplicate data.

[0017] Furthermore, this invention introduces an improved self-attention hypergraph network model, setting up a break-bridge module between the local hyperedge extraction module and the global hyperedge attention module, and using a break-labeling mechanism to perform dimension-by-dimensional subtraction, co-positioning, misalignment stripping, and closure latching on local hyperedge groups. This processing can distinguish between closable and non-closable breaks before global attention convergence, preventing corpora with misaligned sources, suspended labels, and semantically disconnected corpora from directly participating in label locking, thereby improving the accuracy of Chinese corpus cleaning results and the consistency of labeling results.

[0018] This invention also aggregates fracture resonance hyperedge features through a global hyperedge attention module, and reassembles the quality closure markers and unclosed chain positions into the corresponding corpus granules to form a Chinese corpus with quality closure markers. This method can provide a more stable and high-quality Chinese corpus foundation for digital reading, smart education, and large language model training, and has the advantages of high automation, clear quality control, and suitability for large-scale corpus construction. Attached Figure Description

[0019] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings: Figure 1 This is an overall flowchart of an automated Chinese corpus cleaning and annotation method proposed in this invention; Figure 2 This is a schematic diagram of the structure of the improved self-attention hypergraph network model in the automated processing method for cleaning and annotating Chinese corpus proposed in this invention; Figure 3 This is a schematic diagram of the broken tagging mechanism of an automated Chinese corpus cleaning and annotation method proposed in this invention. Detailed Implementation

[0020] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams, illustrating only the basic structure of the invention, and therefore only show the components relevant to the invention.

[0021] refer to Figures 1-3 An automated method for cleaning and annotating Chinese corpora includes the following steps: S1. Rectify the text format of multi-source Chinese corpus, remove noisy content and cut sentence boundaries, and encapsulate the remaining text into corpus fragments with source traces; S2. Segment the textual hierarchy within the corpus fragments and compress the textual relationships, source traces, and annotation clues into a textual structure; S3. Determine the cross-layer clearing associations along the morphological structure and connect the cross-layer clearing associations into a morphological cell network; S4. Feed the text cell cavity network into the improved self-attention hypergraph network model. The improved self-attention hypergraph network model includes a text embedding module, a local hyperedge extraction module, a break bridging module, and a global hyperedge attention module. The text embedding module extracts the connecting chain positions in the text cell cavity network and arranges them into text embedding states. S5. The local superedge extraction module extracts local chain windows along the receiving chain position of the element embedding state, and delimits the clearing associations that can be closed within the local chain window into local superedge groups. S6. The fracture bridging module accepts the local super-edge group and embeds the fracture locking mechanism. It disassembles the super-edge pieces of the same boundary in the local super-edge group dimension by dimension, snaps the closable fracture back to the adjacent chain position, peels the misaligned fracture into the core fracture, and encapsulates it as the fracture resonance super-edge feature with a closed latch. S7. The global superedge attention module collects fracture resonance superedge features along the superedge correlation, and reloads the cleared closure markers and unclosed chain positions into the corresponding corpus particles to create a Chinese corpus.

[0022] In this embodiment, S1 specifically includes: The multi-source Chinese corpus is derived from text libraries of digital reading platforms, educational assessment material libraries, manually compiled imported texts, and excerpts from published content. The original text consists of a title, body, paragraphs, and the order in which they were imported. Source traces are composed of the source platform number, import batch number, file name, chapter position, paragraph number, and original line number. During reading, the text is uniformly converted to UTF-8 encoding. Text that cannot be converted, contains only control characters, or has an empty body is recorded as an exception. The exception record saves the original text number, source trace, reason for the exception, and import batch number, but is not included in the corpus granular encapsulation. The start and end positions of the text are marked as text boundaries, the paragraph sequence numbers are marked as paragraph positions, the source platform number and import batch number are appended to the text boundaries, and the chapter position, paragraph number, and original line number are appended to the paragraph positions, forming the original text with traces.

[0023] After the original text with traces is entered into the text formatting and simplification process, full-width characters are converted to half-width characters, consecutive spaces, tabs, and redundant line breaks are compressed into a single paragraph interval, and headers, footers, page numbers, webpage navigation, download prompts, and copyright notices are separated from the text continuation chain. Titles falling into the middle of the body text, paragraph order not continuous with the original line numbers, and table fields being merged into the body text are marked as format misalignment; the same short sentence repeated more than 5 times in the same import batch and not participating in the continuation of adjacent text is marked as format residue; when the previous line does not end with a period, question mark, exclamation mark, semicolon, or colon, and the next line begins with Chinese characters, numbers, or a left bracket, it is judged as a line break misconnection and merged into the same text continuation chain to form simplified text.

[0024] The cleaned text is scanned along the text continuation chain to identify noise patches. A text patch containing fewer than two visible Chinese characters, and where the number of consecutive symbol characters divided by the total number of visible characters is greater than half, is identified as a symbol noise patch. A text patch repeated more than five times in the same import batch, and where, after removing stop words, it shares fewer than one word with adjacent text patches, is identified as a template noise patch. Patches that cannot be converted to UTF-8 characters or contain consecutive invisible characters constitute garbled text. Stop words are selected using a built-in functional vocabulary, including particles, modal particles, conjunctions, prepositions, and high-frequency meaningless adverbs. Stop words are only removed during the counting of shared words and do not alter the cleaned text content. Noise patches are stripped from the text continuation chain, and the source traces of adjacent positions of the noise patch are preserved at the cleaned text boundary, forming the cleaned text.

[0025] Text purification uses semantic pauses and paragraph continuation to define sentence boundaries. Semantic pauses are determined by periods, question marks, exclamation marks, semicolons, colons, paragraph breaks, and the end of heading fields. Paragraph continuation is determined by three conditions: consecutive paragraph numbers, at least one common word after removing stop words, and a pronoun phrase at the beginning of the current paragraph and a noun phrase at the end of the previous paragraph. Meeting any one of these conditions confirms continuation. Noun phrases consist of consecutive nouns, proper nouns, and quantifier-noun structures. Pronouns include personal pronouns, demonstrative pronouns, and quantifier pronouns. Text segments with a length of 8 Chinese characters or more are retained. If the segment length is less than 8 Chinese characters and meets the paragraph continuation requirement, it is merged into the adjacent segment. If the segment length is less than 8 Chinese characters and does not meet the paragraph continuation requirement, it is moved to the noise stripping location. Retained text, text boundaries, paragraph positions, source trace indexes, and noise stripping locations are concatenated in the same record and encapsulated as corpus granular segments.

[0026] In this embodiment, S2 specifically includes: The corpus slices inherit the preserved text, text boundaries, paragraph positions, and source trace indexes obtained from the previous processing. First, periods, question marks, exclamation marks, semicolons, colons, paragraph breaks, and the ends of heading fields within the preserved text are read as semantic pauses. Adjacent text slices are marked as context-sequential positions if their paragraph numbers are consecutive, or if they share at least one word after removing stop words, or if the current text slice begins with a personal pronoun, demonstrative pronoun, or quantifier pronoun and the previous text slice ends with a noun phrase. Stop words are retained from the built-in functional vocabulary, only removed during the common word count. Noun phrases consist of consecutive nouns, proper nouns, and quantifier-plus-noun structures. Semantic pauses and context-sequential positions jointly define the segmentation boundaries, and the preserved text is segmented into morpheme slices. A morpheme slice represents a continuous text unit that can participate in cleaning and annotation judgments.

[0027] After segmenting the text fragments, the number of Chinese characters in each fragment is counted, and the number of shared words and connection direction between the fragment and adjacent fragments are checked. Fragments with 8 or more Chinese characters that can be connected to any preceding or following fragment are retained as main text fragments; fragments with fewer than 8 Chinese characters, lacking noun phrases, predicates, or unable to form a modifying relationship with preceding or following main text fragments are marked as fragments that cannot be connected independently. Fragments that share at least one word with the preceding main text fragment are merged into the preceding main text fragment; fragments that share at least one word with the following main text fragment and whose preceding main text fragment does not meet the merging condition are merged into the following main text fragment; fragments that cannot be connected to either preceding or following main text fragments are pushed into a broken link position. After merging, the main text fragments are arranged according to the original paragraph position and text sequence to form continuous fragments.

[0028] When consecutive text fragments are arranged in a chain position, each consecutive text fragment corresponds to a successor chain position. The successor chain position records the sequential position of the consecutive text fragment, its preceding object, its succeeding object, and the text fragment to which it belongs. Adjacent consecutive text fragments that cannot satisfy the contextual succession condition are marked as break positions; paragraph numbers that are not consecutive but whose source platform number, file name, and chapter position are consistent are marked as cross-paragraph positions; changes in the source platform number, file name, or chapter position are marked as source transition positions. Break positions, cross-paragraph positions, and source transition positions are collectively marked as chain position boundaries, which retain text boundaries, paragraph positions, and source trace indexes.

[0029] Semantic cues are extracted from continuous fragments of text. Core nouns are extracted from noun phrases, and predicates are extracted from verbs and adjectives. During keyword screening, words are first segmented and stop words are removed. The frequency of candidate words is obtained by dividing the number of occurrences of candidate words by the total number of words in the continuous fragment. Candidate words located in the title, first sentence, or end of a paragraph are marked with a position tag. Candidate words are arranged from high to low frequency. If there are fewer than 3 candidate words, all are retained. If there are more than 3 candidate words, the first 3 candidate words with position tags are retained. The built-in annotation vocabulary includes a sentiment vocabulary, a difficulty hint vocabulary, and a target audience vocabulary. When a word item completely matches the vocabulary, it is written into the annotation cues. Unmatched words are retained in the text fragments to continue participating in the continuation judgment. Semantic cues are embedded into the text location of the corresponding continuation link, and source traces are pasted back to the link boundary. The continuous fragments of text that have completed link arrangement, semantic cue embedding, and source trace pasting are encapsulated into a fragment structure.

[0030] In this embodiment, S3 specifically includes: The text structure inherits the continuous text fragments, successor links, link boundaries, semantic clues, and source trace indexes formed in the previous processing. When scanning along successor links from front to back, the current link, the preceding link, and the following link are read simultaneously each time. If the number of shared words between the current link and the preceding link after removing stop words is 0, and the paragraph numbers are not consecutive, it is marked as a break in the forward text succession. If the number of shared words between the current link and the following link after removing stop words is 0, and the paragraph number of the following link cannot continue from the current link, it is marked as a break in the backward text succession. The number of shared words is obtained by counting the intersection of the terms within the two links. Stop words are taken from the functional vocabulary list in the previous processing.

[0031] Source trace hangover is determined by the positional relationship between the source trace index and the chain position boundary. If the source platform number, file name, or chapter position exists in the current chain position but does not fall within the start and end boundaries of the current chain position, or is only attached to the boundary of an adjacent chain position without pointing to the current chain position, it is marked as a source trace hangover. Labeling clue mismatch is determined by semantic clues and text location. If the subject word, sentiment word, difficulty hint word, or target audience word is outside the text range of the current chain position, it is marked as a labeling clue location mismatch. If both high-difficulty hint words and target audience words for young readers appear simultaneously within the same chain position, or if the same subject word is pointed to different text locations by adjacent chains, it is marked as a labeling clue conflict mismatch. The chain position hit by the scan retains the break type, source trace index, and labeling clue position, and is merged with hit chains within a distance of no more than two adjacent chains into a cleared chain awaiting connection.

[0032] After the text continuation chain is formed, the first and last links serve as the traction boundaries. Links in adjacent text structures that can fill the gaps are included in the same verification scope. The text continuation direction is confirmed by any one of the following conditions: at least one shared word, valid pronoun continuation, or consecutive paragraph numbers. The source attribution relationship is confirmed by at least two of the following: source platform number, file name, and chapter position being consistent. The annotation placement relationship is confirmed by the annotation clue falling into the text attribution position of adjacent links without any annotation clue conflicts. If at least two of the following conditions are met—text continuation direction, source attribution relationship, and annotation placement relationship—and there are no conflicts in the source platform numbers, then adjacent links are deemed sufficient to complete the text continuation chain. A source platform number conflict refers to the existence of two different source platform numbers within the same link boundary, where the file names and chapter positions corresponding to the two source platform numbers cannot form a consistent relationship.

[0033] Links that can be supplemented are folded into cross-layer classification associations. Folding does not change the original text order; only the text continuation relationship, source attribution relationship, and annotation placement relationship are pushed into the same association record. Cross-layer classification associations are reverted to the link boundaries of the corresponding text structure, with the revert position consistent with the hit link in the classification chain to be connected. Links that do not meet the supplementation conditions are pushed into the break position, which records the break type, reason for non-supplementation, source trace index, and annotation clue position. When connecting to form a text cell network, the receiving link is used as the network node, the cross-layer classification association as the closed connection, and the break position as the unclosed connection. The text receiving state, source alignment state, and annotation placement state are respectively attached to the corresponding link, forming a text cell network that can continue to enter the text embedding module.

[0034] In this embodiment, S4 specifically includes: An improved self-attention hypergraph network model is developed to complement the morpheme-cell network. The original self-attention hypergraph network model typically consists of node embedding, self-attention interactions within hyperedges, and hyperedge discrimination output. The node embedding part converts hypergraph nodes into general vectors, the hyperedge attention part calculates the correlation between node vectors within the same hyperedge, and the hyperedge output part determines the higher-order relationship between nodes and hyperedges. This structure is suitable for general hyperedge prediction, but in Chinese corpus cleaning and annotation scenarios, corpus chain positions, source traces, annotation clues, and breakpoint locations cannot be directly pushed into the unified attention weight as ordinary node attributes. Otherwise, similar common sentence structures, misconnected sources, and dangling annotations can easily be mixed within the same hyperedge attention path, making it impossible to distinguish between text continuity breaks and fillable breaks before global aggregation. Therefore, this invention modifies the node embedding position into a morpheme embedding module, so that the morpheme-cell network is first organized into morpheme embedding states with chain position boundaries, breakpoint boundaries, and attachment boundaries, and then handed over to the local hyperedge extraction module for processing.

[0035] The morpheme embedding module consists of a chain position expansion structure, a chain position slot forming structure, a break slot allocation structure, a substrate mapping structure, an attachment vector overlay structure, and an embedding state arrangement structure, all sequentially connected. The chain position expansion structure reads the receiving chain positions within the morpheme cell network. These receiving chain positions are formed by the previous processing step and record the position of consecutive morpheme pieces in the corpus particles, their preceding and following objects, and source trace indices. The chain position slot forming structure pushes the retained text, source traces, and annotation clues at the same receiving position into the same chain position range. The same receiving position is determined by the text boundary, paragraph position, and source trace index. The break slot allocation structure receives broken content that cannot be continued and pushes it into the chain position break slot. The substrate mapping structure converts the continued text content into chain position basis vectors. The attachment vector overlay structure converts source traces and annotation clues into boundary attachment vectors and overlays them onto the corresponding boundaries of the chain position basis vectors. The embedding state arrangement structure arranges the chain position embedding vectors according to the chain position order within the network, forming the morpheme embedding state.

[0036] When the chain position unfolding structure is working, the chain positions within the network are first unfolded along the receiving direction of the element-cell network. The receiving direction is determined by the ascending order of the receiving chain position numbers. The chain positions within the network inherit the text receiving state, source alignment state, and annotation placement state from the element-cell network. When each chain position within the network is read, the element structure falling into the same receiving position, the cross-layer annotation association, and the break position are extracted simultaneously. The element structure provides continuous element pieces and chain position boundaries, the cross-layer annotation association provides the closed connection between the current chain position and adjacent chain positions, and the break position provides the incomplete break boundary and the reason for the incompleteness. After the above contents are gathered, a chain piece to be embedded is formed. The chain piece to be embedded does not change the original text order and is only used as the chain position organization object before vectorization.

[0037] When processing the chain slot forming structure to be embedded in the chain piece, the chain position boundary is first determined by the text start character offset, text end character offset, paragraph number, and source trace index. After removing the noise stripping positions already marked in the previous processing, the text is organized into continuous text segments according to the chain position boundaries. The source trace is pasted back to the start and end boundaries of the continuous text segments. The start boundary records the source platform number and import batch number, and the end boundary records the file name, chapter position, and paragraph number. The text belonging position of the annotation clue is embedded, and the text belonging position is determined by the start character offset and end character offset of the term containing the annotation clue in the continuous text segment. After the continuous text segments, source traces, and annotation clues are completed and positioned at the boundaries, the chain slot is formed according to the boundary constraints. The left boundary of the chain slot saves the preceding relationship, and the right boundary of the chain slot saves the following relationship.

[0038] The break-slot allocation structure compares the continuation direction of a link slot with adjacent link slots. During comparison, the right boundary of the current link slot and the left boundary of the following link slot are read first, followed by the left boundary of the current link slot and the right boundary of the preceding link slot. If any one of the following three conditions is met between the two boundaries: continuous paragraph numbers, at least one common word after removing stop words, and the pronoun at the beginning of the current link slot pointing to a noun phrase at the end of the preceding link slot, the corresponding content is retained within the link slot and marked as continuous content. Line breaks and misconnections are determined by the end and beginning of the line. If the preceding link slot does not end with a period, question mark, exclamation mark, semicolon, or colon, and the following link slot begins with Chinese characters, numbers, or a left bracket with consistent source traces, the broken content is pushed into the corresponding link slot boundary. Semantic dangling is determined by a common word count of 0, invalid pronoun continuation, and discontinuous paragraph numbers. Source separation is determined by the source platform number, file name, or chapter position not falling within the start and end boundaries of the current link slot. Content that is broken, semantically suspended, or whose source is detached from its corresponding link is pushed into the link break slot. The link break slot records the broken content, the break boundary, the break type, and the source trace index.

[0039] The substrate mapping structure expands the completed continuation content within the chain slots into text substrates according to the continuation order. The text substrates are first segmented into words, retaining complete terms, while stop words are only removed from the common word count. Each term uses a 128-dimensional word vector trained on the corpus. The chain slot basis vector is obtained by accumulating all word vectors in the text substrate dimension-wise and dividing by the number of terms. Out-of-vocabulary (OV) words are not directly discarded; they are first segmented into single characters, and the 128-dimensional character vectors corresponding to each single character are read and averaged dimension-wise to obtain the OV word vector. If a single character still cannot match the character vector table, the term vector is set to a 128-dimensional all-zero vector, and an OV word marker is retained at the chain slot boundary. When the number of terms is 0, the chain slot basis vector is set to a 128-dimensional all-zero vector, and an empty segment marker is retained at the chain slot boundary. Source traces are folded into source attachments. The source platform number, import batch number, file name, chapter position, and paragraph number are converted into field numbers. Each field number reads its corresponding 128-dimensional embedding vector. The source vector is formed by averaging all field embedding vectors dimension by dimension. Annotation clues are folded into annotation attachments. Keywords, sentiment words, difficulty hint words, and target audience words are converted into 128-dimensional annotation vectors in the same way. Items that do not match the annotation thesaurus are not included in the annotation attachments but remain in the text base.

[0040] The attached vector overlay structure attaches the source vector to the start and end boundaries of the chain position and the annotation vector to the text location. During overlay, the chain position base vector stores the main dimension of the text content, while the source and annotation vectors are written to the boundary dimension; the boundary dimension is recorded separately from the main dimension, and the source traces and annotation clues do not cover the text content. After the chain position base vector, source vector, and annotation vector are overlaid at the boundary, a chain position embedding vector is formed. The broken content in the chain position break slot is converted into a break embedding piece using the same 128-dimensional mapping method. The break embedding piece is pasted back to the break boundary of the chain position embedding vector. The pasting position is determined by the break boundary recorded by the chain position break slot, and the break embedding piece maintains a one-to-one correspondence with the original break position.

[0041] The embedding state arrangement rearranges the link embedding vectors according to the intranet link order of the morphological cavity network. Precedence and succession relationships are preserved between adjacent link embedding vectors, and each link embedding vector carries its corresponding boundary attachment vector and break embedding piece. After arrangement, a morphological embedding state is formed, which includes link embedding vectors, boundary attachment vectors, break embedding pieces, precedence relationships, succession relationships, and the intranet link order. When the local hyperedge extraction module reads the morphological embedding state, it can directly identify link boundaries, break boundaries, source re-attachment positions, and labeled attribution positions.

[0042] Compared to the original attention-based hypergraph network model, the improvement of this invention lies in transforming the original general node embedding structure into a morpheme embedding module with chain position boundary constraints. The original node embedding only forms ordinary node vectors, failing to preserve text continuation direction, source / reply position, annotation attribution position, and break boundaries in the Chinese corpus. The improved morpheme embedding module completes chain position slot shaping, break slot allocation, and attached vector overlay before vector formation, ensuring that the morpheme embedding state entering the local hyperedge extraction module has readable continuation and break boundaries. Through this structural modification, the local hyperedge extraction module can directly extract clean text hyperedge pieces, source / break hyperedge pieces, and semantic hyperedge pieces from the morpheme embedding state, while the break bridging module can identify closable and verified breaks before global attention convergence, thereby reducing interference from common sentence structure similarity, source misalignment, and annotation dangling on the cleaning and annotation results.

[0043] In this embodiment, S5 specifically includes: The local hyperedge extraction module takes over the element embedding state. The element embedding state includes chain embedding vectors, boundary attachment vectors, break embedding pieces, preceding relationships, following relationships, and the order of chain positions within the network. The local hyperedge extraction module consists of a chain window opening structure, a boundary continuation structure, a source trace backfit structure, a semantic proximity structure, a co-position arrangement structure, and a hyperedge group encapsulation structure, all connected in sequence. The chain window opening structure places the chain embedding vectors into the local processing range; the boundary continuation structure determines whether a clean text connection can be formed between chain positions; the source trace backfit structure processes the replacement of broken positions by source traces; the semantic proximity structure processes the alignment of annotation clues with the text's location; the co-position arrangement structure presses clean text, source trace, and semantic hyperedge pieces onto the same chain boundary; and the hyperedge group encapsulation structure arranges the local results into local hyperedge groups.

[0044] After reading the morpheme embedding state through the open chain window structure, the chain embedding vectors are opened sequentially along the chain positions within the network. The order of the chain positions within the network is determined by the successor chain position number from smallest to largest. Each chain embedding vector carries the current chain position number, the preceding chain position number, the following chain position number, the break boundary, and the break-line embedding piece. When processing any current chain position, the preceding chain position, the current chain position, and the following chain position are collapsed into a local chain window. When the current chain position is at the beginning of a corpus particle, the preceding chain position has an empty boundary; when it is at the end of a corpus particle, the following chain position has an empty boundary. Empty boundaries only record the missing state and do not fill in text content. The break boundary and break-line embedding piece of the chain embedding vector are brought into the local chain window along with the chain position position, serving as the judgment objects for the boundary continuation structure.

[0045] The boundary continuation structure reads the preceding and following boundaries within the local chain window. The preceding boundary is the connection position between the right boundary of the preceding chain position and the left boundary of the current chain position, and the following boundary is the connection position between the right boundary of the current chain position and the left boundary of the following chain position. The text continuation direction is determined by three conditions: the paragraph numbers are consecutive; the number of common words between the two boundary texts after removing stop words is not less than one; and the pronoun at the beginning of the latter boundary can point to the noun phrase at the end of the former boundary. When any one of the conditions is true, the continuation direction is marked as consistent, and the corresponding chain position embedding vectors are joined along the boundary to form a clean text hyperedge piece; when none of the three conditions are true, the continuation direction is marked as broken, the corresponding chain position embedding vectors do not participate in the clean text joining, and a break mark is left at the break boundary. The break mark storage saves the chain position number, break boundary, break mark type, and break mark embedding piece.

[0046] The source trace back-fitting structure searches for homologous link embedding vectors centered on the break trace retention position. Homologous link embedding vectors are confirmed through source platform number, file name, and chapter position; if at least two of these three items are consistent, they are considered homologous. The search scope is initially limited to the same corpus grain; if no candidate link is found within the same corpus grain, it is expanded to the next 10 links within the same file name and chapter position. When multiple candidate links simultaneously meet the homologous condition, the candidate link with the smallest paragraph number distance is selected first; if the paragraph number distances are the same, the candidate link with the smallest original line number distance is selected. The source boundary can fit the break boundary, meaning the source trace index of the candidate link falls within the start and end boundaries of the break trace retention record, and the source platform number does not conflict. Source platform number conflict means that different source platform numbers exist within the same break trace retention position, and the file name and chapter position are not aligned. When the fitting condition is met, the homologous link embedding vector is folded into the source trace super-edge piece; when the fitting condition is not met, the original break trace connection is retained in the break trace retention position. The original break mark connection record includes the break boundary, source trace index, and reason for incomplete bonding.

[0047] The semantic proximity structure searches for co-labeled chain embedding vectors centered on the text's location. Co-labeled chain embedding vectors are confirmed through the labeling clue category, labeling clue term, and the built-in labeling dictionary. A vector is considered co-labeled if any of the following conditions are met: identical category, identical term, or from the same labeling dictionary. The text's location is determined by the start and end character offsets of the term containing the labeling clue. A label's location must fall within the corresponding text boundary, meaning both the start and end character offsets of the labeling clue are within the chain text boundary. A label's location exceeding the chain text boundary, falling into an empty boundary, or only being attached to an adjacent chain is considered deviating from the corresponding text boundary. Co-labeled chain embedding vectors that have successfully fallen within the boundary are folded into semantic hyperedge pieces; labeling clues that deviate from the corresponding text boundary are designated as semantic dangling pieces. Semantic dangling pieces store the labeling clue category, labeling term, dangling position, and original text location.

[0048] The co-position arrangement structure uses the same chain position boundary as the arrangement coordinate. The chain position boundary is jointly defined by the chain position number, left boundary, right boundary, and source trace index. The clean text super-edge piece is placed in the main position of the boundary slot, the source trace super-edge piece is attached to the source side boundary, and the label super-edge piece is attached to the text belonging position; the break mark retention piece, the original break mark connection piece, and the label hanging piece are pressed into the corresponding boundary positions. When pressing, the break mark retention piece is attached to the broken boundary, the original break mark connection piece is attached to the source side boundary, and the label hanging piece is attached to the label hanging position. After the co-position arrangement is completed, the clean text connection result, the source back-alignment result, the label close result, and the break mark object that has not been fully attached are retained in the same boundary slot, forming the super-edge pieces to be bridged.

[0049] The superedge group encapsulation structure arranges the superedge pieces to be bridged along the inheritance sequence of the element embedding state. Each superedge piece to be bridged retains the local chain window number, chain position boundary, clean text superedge piece, source trace superedge piece, marked superedge piece, break trace reservation, original break trace connection, and marked floating piece; the preceding and following relationships are retained between adjacent superedge pieces to be bridged. After encapsulation, a local superedge group is formed, and the corresponding boundary, break trace reservation, original break trace connection, and marked floating piece in the local superedge group are the direct objects read by the break trace bridging module.

[0050] Compared to the original self-attention hypergraph network model, the improvement of this invention lies in transforming the ordinary node aggregation process before hyperedge construction into a local hyperedge extraction module. The original model typically organizes node vectors directly into hyperedges before entering self-attention interaction, lacking local encapsulation for the chain boundaries, source boundaries, and annotation attribution positions of the Chinese corpus. The improved local hyperedge extraction module, before entering the break-line bridging module, first forms clean text hyperedge pieces, source-line hyperedge pieces, and semantic hyperedge pieces, while retaining break-line positions, original break-line connections, and semantically suspended pieces. Through this structural modification, text continuation, source supplementation, and annotation proximity in local corpus fragments can be arranged in the same position within the same chain boundary, providing clear local objects for the dimensional subtraction, co-positioning, and misalignment removal of the break-line bridging module, reducing interference from common sentence structure similarity, source boundary drift, and suspended annotation cues on global hyperedge attention.

[0051] In this embodiment, S6 specifically includes: The fracture bridging module accepts local super-edge groups, which originate from the local super-edge extraction module. Internally, it retains the super-edge pieces to be bridged, chain position boundaries, fracture reservation positions, original fracture connections, and labeled suspended pieces. The fracture bridging module consists of a bridging chain position interception structure, a co-position arrangement structure, a three-slot pressing structure, a fracture separation structure, a label locking access structure, and a rearranged encapsulation structure connected in sequence. The bridging link interception structure is responsible for retrieving the super-edge pieces to be bridged from the local super-edge group; the co-position arrangement structure is responsible for placing the super-edge pieces under the same link boundary into the same processing position; the three-slot pressing structure is responsible for determining whether the three types of super-edge pieces—clean text, source trace, and label definition—can fall into the same link boundary; the break mark splitting structure is responsible for splitting the co-position bridging pieces into source text break mark pieces and label text break mark pieces; the label locking access structure hands over the break mark bridging pieces, source trace break pieces, and label definition suspended pieces to the break mark locking mechanism; and the rearrangement encapsulation structure organizes the closed bridging pieces and the verified break mark pieces into break mark resonance super-edge features.

[0052] When processing local hyperedge groups, the bridging link interception structure reads the hyperedge pieces to be bridged one by one according to the inheritance order of the element embedding state. The link boundary carried by the hyperedge piece to be bridged is determined by the link number, left boundary, right boundary, and source trace index. The left boundary is the starting character offset of the current link in the text, and the right boundary is the ending character offset of the current link in the text. The break mark retention record records the break boundary, break mark type, and break mark embedding piece. The break boundary is determined by the character offset before and after the break position. The original break mark connection record records the source index and boundary position of the source trace that is not fully attached. The label hanging piece records the label clue category, label term, label hanging position, and original text belonging position. Hyperedge pieces to be bridged with the same link number and overlapping left and right boundaries are intercepted in the same bridging link. Hyperedge pieces to be bridged with the same link number but overlapping boundaries maintain the original inheritance order and are transferred to the adjacent bridging link to avoid different text pieces being mistakenly merged.

[0053] The co-position arrangement structure uses the same chain position boundary as the arrangement coordinate. The clean text super-edge piece is placed in the bridging main slot, the source trace super-edge piece in the source side slot, and the label super-edge piece in the label side slot, forming a three-slot bridging piece. The bridging main slot stores the clean text vector after text continuation, the source side slot stores the source trace vector after the source trace is back, and the label side slot stores the label vector after the label clue is close. When the clean text super-edge piece is missing, the current super-edge piece to be bridged cannot provide the text principal position; the original chain position boundary is retained, and the process transitions to the verification of the broken trace piece's forming path. When the source trace super-edge piece is missing, the original broken trace connection is written into the source side slot. When the label super-edge piece is missing, the label side slot writes the label suspended piece. The three-slot bridging piece retains the co-position relationship between the bridging main slot, the source side slot, and the label side slot, serving as the processing object for the three-slot pressing structure.

[0054] The three-groove pressing structure processes the source side groove first, then the annotation side groove. The source side groove moves back towards the bridging main groove along the source trace. If the source trace falls within the same chain position boundary, it means the source platform number is consistent, and at least two of the file name, chapter position, and paragraph number are consistent with the corresponding boundary of the bridging main groove. If the source platform number is inconsistent, or if fewer than two of the file name, chapter position, and paragraph number are consistent, the source trace is considered to have deviated from the chain position boundary. The annotation side groove moves closer to the bridging main groove along the text belonging position. If the text belonging position falls within the same chain position boundary, the start and end character offsets of the annotation line are both within the text boundary of the bridging main groove. If the start or end character offset exceeds the text boundary of the bridging main groove, the text belonging position is considered to have deviated from the chain position boundary. When both the source trace and the text location fall within the same chain position boundary, the three-groove bridging piece is pressed together to form a co-position bridging piece; when the source trace deviates from the chain position boundary, the source side groove is peeled out to form a source trace disconnected piece; when the text location deviates from the chain position boundary, the label side groove is peeled out to form a label suspended piece.

[0055] When processing the bridging fragments in the fractured text structure, the vectors within the bridging main slot, source side slot, and label side slot are read. The vector dimensions are consistent with the element embedding state, and 128 dimensions are used in implementation. During source text fracture, the values ​​of the bridging main slot from dimension 1 to 128 are subtracted from the corresponding dimensions of the source side slot, and the difference and direction state are saved for each dimension. When the difference is greater than 0, the direction state is recorded as main slot biased; when the difference is less than 0, the direction state is recorded as source side biased; when the difference is equal to 0, the direction state is recorded as co-position consistent. The source text fracture results of the 128 dimensions are pressed into the same chain position boundary to form the source text fracture fragment. During label text fracture, the values ​​of the bridging main slot from dimension 1 to 128 are subtracted from the corresponding dimensions of the label side slot, and the difference and direction state are saved in the same way to form the label text fracture fragment. Source text break fragments and label text break fragments are stacked along the same chain position boundary. The original dimensional order is not changed during stacking. Source text break fragments are arranged on the source side, and label text break fragments are arranged on the label side, forming a break bridging fragment.

[0056] The tagging access structure uses the break bridging piece as the main processing object. It attaches the source break piece and the label-suspended piece to the corresponding chain position boundary. The break tagging mechanism compares the break direction along the chain position boundary, locking break pieces with the same direction that can be reattached to adjacent chain positions as closable break pieces. Break pieces with deviated source boundaries or unresolved label suspensions are peeled off as verification break pieces, and the bridging boundary is locked in place by a closing latch. After processing, the closable break piece is pressed against the corresponding three-slot bridging piece to form a closed bridging piece. The verification break piece retains the break boundary, source trace index, label suspension position, and reason for non-closure.

[0057] After receiving the closed bridging piece and the verification fracture piece, the rearranged encapsulation structure first arranges the closed bridging pieces according to the original acceptance order of the local superedge group. The original acceptance order is determined by the chain position number from smallest to largest. When the chain position numbers are the same, the closed bridging piece with the smaller left boundary character offset is arranged first. The verification fracture piece is pasted back to the corresponding chain position boundary. The pasting position is determined by the break boundary or the marked floating position saved by the verification fracture piece. When both break boundaries and marked floating positions exist, the break boundary is pasted back first, and then the marked floating position is attached to the marked side of the same chain position boundary. During encapsulation, the closed bridging piece serves as the resonant master piece, and the verification fracture piece serves as the unclosed mark. Both are written into the same fracture resonant superedge feature. The fracture resonance hyperedge feature preserves the chain position number, closed bridging vector, verified fracture position, reason for non-closure, and original local hyperedge group index. The closed bridging vector is obtained by averaging the vectors of the bridging main groove, source side groove, and labeled side groove in the closed bridging piece dimension by dimension. When the source side groove or labeled side groove is missing, only the actual groove vector is averaged, and the missing groove mark is retained.

[0058] Compared to the original self-attention hypergraph network model, the improvement of this invention lies in changing the transmission path from local hyperedge features directly entering global self-attention to first being bridged by a break-bridge module. The original model directly calculates global relevance after local hyperedges are formed, and similar common sentence structures, source boundary offsets, and suspended labels easily enter the attention weights simultaneously. The improved break-bridge module completes co-positioning, three-slot compression, dimension-by-dimensional subtraction, and label locking before global aggregation, encapsulating closable breaks into closed bridging pieces, and retaining source disconnections and suspended labels as verification break pieces. Through this structural modification, the global hyperedge attention module receives not ordinary local hyperedges, but hyperedge features processed by break-resonance, which reduces the interference of misaligned sources and incorrect labels on the final corpus label clearing and closure.

[0059] In this embodiment, S65 specifically includes: The break mark locking mechanism is located within the break mark bridging module, receiving break mark bridging pieces, source break mark pieces, and label hanging pieces. The break mark locking mechanism is connected in series in the following order: break mark differential unit, co-positioning engagement unit, misalignment peeling unit, and closing latch unit. The break mark differential unit is responsible for separating the source text break mark pieces and label break mark pieces in the break mark bridging piece and pressing break marks with consistent orientation into the co-position differential slot; the co-positioning engagement unit is responsible for checking whether the co-position differential slot can be reattached to the adjacent chain position boundary; the misalignment peeling unit is responsible for handling break marks, source break mark pieces, and label hanging pieces that cannot be reattached; the closing latch unit is responsible for pressing the closable break mark pieces back to the original chain position boundary and locking them as closed bridging pieces.

[0060] The fracture difference unit first expands the source text fracture piece and the label text fracture piece along the chain position boundary of the fracture bridging piece. The chain position boundary is jointly defined by the chain position number, the left boundary character offset, the right boundary character offset, and the source trace index; the source text fracture piece comes from the dimension-by-dimensional subtraction of the bridging main slot and the source side slot, and the label text fracture piece comes from the dimension-by-dimensional subtraction of the bridging main slot and the label side slot. Each fracture piece retains 128 dimensions of difference and direction status, including main slot height, side slot height, and same position consistency. The fracture difference unit compares the source text fracture piece and the label text fracture piece within the same chain position boundary dimension by dimension. When two direction statuses within the same dimension are both main slot height, both side slot height, or both same position consistency, that dimension is recorded as direction consistency; when two direction statuses within the same dimension are opposite, or when the two fracture pieces are not on the same chain position boundary, that dimension is recorded as direction offset. When the number of dimensions with consistent orientation divided by the number of valid dimensions reaches two-thirds, the corresponding source text fragments and target text fragments are pushed into the same-position difference slot; if it does not reach two-thirds, they are pushed into the misalignment temporary storage slot. Valid dimensions are those where both source text fragments and target text fragments have values. Dimensions lacking values ​​in either fragment are not included in the proportional calculation, and missing dimension markers are retained in the misalignment temporary storage slot.

[0061] After receiving the corresponding differential slot, the corresponding interlocking unit attempts to fit along the original chain boundary, both the preceding and following chain positions. The text continuation direction is determined by three conditions: continuous paragraph numbers, at least one common word after removing stop words, and the starting pronoun of the current chain position pointing to the end of a noun phrase in the adjacent chain position. Meeting any one of these conditions indicates a continuous text continuation direction. The source backtracking direction is determined by the source platform number, file name, chapter position, and paragraph number. If the source platform number is consistent, and at least two of the file name, chapter position, and paragraph number are consistent, it is considered a return to the same source trace. The annotation placement direction is determined by the character offset of the annotation clue. If both the starting and ending character offsets of the annotation clue are within the text boundaries of adjacent chain positions, it is considered an embedding into the same text belonging position. When the text receiving direction, the source return direction, and the annotation placement direction are all met simultaneously, the corresponding differential slot snaps back to the original chain position boundary and forms a closable break piece; if any direction cannot be fitted, the corresponding differential slot moves into the misalignment temporary storage slot and retains the direction type that cannot be fitted.

[0062] After receiving the misalignment temporary storage slot, the source trace break piece, and the labeling overhang piece, the misalignment stripping unit first reads the direction misalignment dimension, missing dimension markers, and non-fitting direction types from the misalignment temporary storage slot. Then, it reads the source break position from the source trace break piece and the label overhang position from the labeling overhang piece. The source break position is determined by the break boundary and source trace index within the source trace break piece, while the label overhang position is determined by the offset of the label term character and the original text's location within the labeling overhang piece. When the source break position and the labeled dangling position fall on the same chain boundary, they are pressed together to form a verification break piece, retaining the reasons for the source break and the labeled dangling position. When the source break position can be backed to an adjacent source trace but the labeled dangling position cannot be embedded in the text belonging position, the source connection is retained at the original chain boundary, and the labeled dangling piece is peeled into the verification break piece. When the labeled dangling position can be embedded in the text belonging position but the source break position cannot be backed to, the labeled connection is retained at the text belonging position, and the source trace break piece is peeled into the verification break piece. When neither the source break position nor the labeled dangling position can be corrected, the source trace break piece and the labeled dangling piece are jointly merged into the verification break piece. The verification break piece records the chain boundary, direction misalignment dimension, missing dimension marker, source break position, labeled dangling position, and reason for non-closure.

[0063] After receiving the closable break piece, the closing latch unit reattaches it to the original chain position boundary of the break bridging piece. During reattachment, the source text break piece and the label text break piece are restored first. Then, the source text break piece is pressed between the bridging main slot and the source side slot, and the label text break piece is pressed between the bridging main slot and the label side slot. The bridging main slot retains the main text position, the source side slot conforms to the source side boundary, and the label side slot conforms to the text's location. After pressing, if there are no misaligned temporary slots, source break pieces, or label hanging pieces within the same chain position boundary, that chain position boundary is locked as a closed bridging slot; the closed bridging slots are connected in series along the original receiving direction to form a closed bridging piece. The original receiving direction is determined by the chain position number from smallest to largest. When the chain position numbers are the same, the closed bridging slot with the smaller left boundary character offset is arranged first.

[0064] After the closed bridging piece is formed, the closed latch unit verifies the co-location link, source back-docking relationship, and annotation placement relationship. The co-location link requires that the bridging main slot, source side slot, and annotation side slot share the same link number, left boundary character offset, and right boundary character offset. The source back-docking relationship requires that the source trace index retained by the source side slot is consistent with the source trace index within the boundary of the bridging main slot. The annotation placement relationship requires that the annotation clue of the annotation side slot remains within the text's belonging position. When all three relationships are true, the closed bridging piece remains locked; if any relationship is false, the remaining portion is peeled off from the closed bridging piece and incorporated into the verification break piece. After verification, the closed bridging piece saves the locked co-location link, source back-docking relationship, and annotation placement relationship, while the verification break piece saves the unlocked link boundaries, direction status, source disconnection position, annotation suspension position, and reason for non-closure.

[0065] Compared to the original attention-based hypergraph network model, the improvement of this invention lies in setting a break-marking mechanism within the break bridging module. This allows local hyperedge groups to complete break direction verification, in-situ reconnection, misalignment removal, and closure locking before entering the global hyperedge attention module. The original model performs unified attention calculation on the correlation between nodes within the hyperedge, failing to separately handle source breaks, suspended labels, and fillable breaks. The improved break-marking mechanism first uses a break differential unit to retain the directional state of the source text break and the label text break, then uses an in-situ reconnection unit to determine whether the break can be reconnected to adjacent chain positions, subsequently uses a misalignment removal unit to convert non-closable parts into verification break pieces, and finally uses a closure latching unit to lock the completed bridging boundary. Through this mechanism, only closed bridging relationships and clearly marked verification breaks are retained in the break resonance hyperedge features, reducing erroneous aggregation after source misalignment and suspended labels enter the global attention module.

[0066] In this embodiment, S7 specifically includes: The global hyperedge attention module inherits the fracture resonance hyperedge feature, which consists of closed bridging pieces, verified fracture pieces, chain position numbers, closed bridging vectors, verified fracture positions, reasons for non-closure, and the original local hyperedge group index. The global hyperedge attention module is composed of a feature chain splitting structure, a target hyperedge selection structure, a fitting comparison structure, a closure clustering structure, a verification back-fitting structure, and a corpus reloading structure, all sequentially connected. The feature chain splitting structure is responsible for extracting closed bridging pieces and verified fracture pieces along the chain position boundaries; the target hyperedge selection structure is responsible for determining the target hyperedge pieces one by one in the global attention chain; the fitting comparison structure is responsible for calculating the fitting state between the target hyperedge pieces and the remaining closed bridging pieces; the closure clustering structure is responsible for incorporating closable objects into the cleared closed pieces; the verification back-fitting structure is responsible for fixing unclosed objects to the original chain position boundaries; and the corpus reloading structure is responsible for reloading the cleared closure markers and verification chain positions into the corpus pieces.

[0067] After reading the fracture resonance hyperedge features using the feature-based chain-breaking structure, the source chain segment is first located using the original local hyperedge group index. Then, the closed bridging piece and the verification fracture piece are separated using the chain position number and chain position boundary. The chain position boundary is jointly defined by the left boundary character offset, the right boundary character offset, and the source trace index. The left boundary character offset indicates the starting character position of the chain position within the corpus particle, and the right boundary character offset indicates the ending character position of the chain position within the corpus particle. Closed bridging pieces are arranged in ascending order of chain position number to form a global attention chain. When the chain position numbers are the same, the closed bridging piece with the smaller left boundary character offset is placed first. Verification fracture pieces are retained at the corresponding chain position boundary and do not enter the global attention chain. The global attention chain only carries closed bridging pieces that have completed the closure latch. Verification fracture pieces retain the fracture position, source boundary, annotation placement offset, and reason for non-closure.

[0068] The target superedge selection structure reads closed bridging pieces sequentially along the global attention chain. Each time, the current closed bridging piece is designated as the target superedge piece, and the remaining closed bridging pieces are used as comparison superedge pieces in turn. Both the target and comparison superedge pieces read the connection direction, source backlink relationship, and annotation placement relationship. The connection direction is determined by the preceding chain position number, the following chain position number, the left boundary character offset, and the right boundary character offset; the source backlink relationship is determined by the source trace index, the source platform number, the file name, and the chapter position; the annotation placement relationship is determined by the annotation clue category, the annotation term, the start character offset, and the end character offset. Closed bridging pieces that have already been compared with the current target superedge piece are not re-entered into the current comparison queue to avoid duplicate counting of the same superedge relationship.

[0069] The alignment structure forms an alignment mark between the target hyperedge and the alignment hyperedge. The continuation alignment value is calculated based on the chain position adjacency and text boundary continuity. The continuation alignment value is 1 if the following chain position number of the target hyperedge equals the chain position number of the alignment hyperedge, or if the right boundary character offset of the target hyperedge is continuous with the left boundary character offset of the alignment hyperedge; otherwise, it is 0. The source alignment value is calculated based on the source backlink relationship. It is 1 if the source platform number is consistent and the file name and chapter position are also consistent; otherwise, it is 0. The annotation alignment value is calculated based on the annotation placement relationship. It is 1 if the annotation clue category is consistent and the annotation term falls into the adjacent text's belonging position; otherwise, it is 0. The super-edge bonding value is obtained by multiplying the receiving bonding value by 0.4, the source bonding value by 0.3, and the label bonding value by 0.3, and then adding them together. The weights of the three items are added together to get 1. When the super-edge bonding value reaches 0.7, an effective super-edge bonding mark is formed. When it is lower than 0.7, the original chain position boundary is retained when comparing the super-edge pieces.

[0070] After reading the valid super-edge bonding marks, the closed bridging structure gathers the closed bridging pieces with consistent orientations toward the target super-edge piece. Consistent orientation means that the target super-edge piece and the comparison super-edge piece can connect in their bearing directions, there is no conflict in the source platform number in the source back-docking relationship, and there is no text attribution shift in the annotation placement relationship. When the bearing directions can connect and the source back-docking relationship does not conflict, the comparison super-edge piece is merged into the clearing closed piece where the target super-edge piece is located; when the annotation placement relationship shifts, the comparison super-edge piece does not enter the clearing closed piece, and the original chain position boundary retains the reason for non-bonding. The clearing closed piece records the target super-edge piece number, the merged super-edge piece number, the valid super-edge bonding marks, the annotation clues, and the clearing closure value. The clearing closure value is obtained by averaging all valid super-edge bonding values ​​within the clearing closed piece; when the clearing closure value reaches 0.7 and no verification break piece covers the corresponding annotation clue, the clearing closed piece enters the annotation locking range.

[0071] The review and re-paste structure handles review breakpoints that are not within the annotation locking range. Review breakpoints are re-pasted along the original chain boundary to between adjacent cleared closing points. The re-paste position is determined by the breakpoint location within the review breakpoint, the source boundary, and the annotation placement offset. When both the breakpoint location and the source boundary exist, the breakpoint location is re-pasted first, followed by the source boundary. When an annotation placement offset exists, it is attached to the annotation side of the same chain boundary. Review breakpoints do not participate in annotation locking. When a review breakpoint covers an annotation clue, the corresponding annotation clue remains in a pending review state. When a review breakpoint only covers the source boundary, the corresponding source trace remains in a pending verification state, and closed text connections are not removed.

[0072] The corpus reloading structure reloads the cleared closure fragments into the corresponding morpheme structures. During reloading, the morpheme structure is first located using the original local hyperedge group index, and then the receiving chain position is located using the chain position number. Annotation clues within the cleared closure fragment that meet the locking conditions are written into the cleared closure marker. The cleared closure marker records the annotation clue, cleared closure value, closed chain position, and source trace index. The boundary of an incompletely closed chain position receives a verification fragment and is marked as a verification chain position. The verification chain position records the verification fragment, the reason for incomplete closure, the source boundary, the annotation placement offset, and the manual verification entry point. Morpheme structures with cleared closure markers are reloaded according to the corpus fragment boundaries, and verification chains are attached along with the corresponding source traces, forming the Chinese corpus. Each corpus fragment in the Chinese corpus stores the original source trace, morpheme structure index, cleared closure marker, and verification chain position.

[0073] Compared to the original attention hypergraph network model, the improvement of this invention lies in that the global hyperedge attention module does not directly calculate unified attention weights for ordinary local hyperedges, but only globally aggregates the resonant hyperedge features output by the break-bridge module. The original model tends to mix and aggregate closed hyperedges, source misaligned hyperedges, and labeled suspended hyperedges during the global attention stage. The improved global hyperedge attention module first splits closed bridging pieces and verification break-edge pieces, then constructs a global attention chain using closed bridging pieces, and limits the label locking range using verification break-edge pieces. Through this structural modification, the label closure markers only come from hyperedge pieces that have completed closure, source relocation, and label placement. Chain positions that have not completed closure are reloaded as verification chain positions, thus allowing the Chinese corpus to simultaneously retain content that can be directly added to the database and break-edge positions that need verification, reducing the risk of erroneous labels being directly added to the database.

[0074] Example 1: To verify the feasibility of this invention in practice, it was applied to the construction process of a Chinese corpus for a digital reading and smart education platform. The corpus to be added to the platform mainly consists of book excerpts, online reading content, educational assessment materials, and manually compiled texts. The corpus contains issues such as cross-source reprints, excerpts from the same text, rewritten and reused texts, line breaks and misjoints, repeated template prompts, and inconsistent tags. Existing processing methods primarily relied on rule cleaning, similarity deduplication, and manual review. When faced with semantically similar but different versions of text, valuable excerpts were easily mistakenly deleted. Furthermore, when faced with texts with similar main text but different reading levels, themes, or target audiences, it was difficult to promptly identify annotation conflicts. This resulted in a mixture of duplicate, mislabeled, and unreviewed content in the corpus, affecting the stability of intelligent reading leveling, content recommendation, and model training data.

[0075] In actual processing, the incoming Chinese corpus is first formatted, noise removed, and sentence boundary segmented, and the retained text is encapsulated into corpus fragments with source traces. Then, the corpus fragments are segmented at the morpheme level, and the text succession relationships, source traces, and annotation clues are compressed into a morpheme structure. Cross-layer clearing and labeling associations are then marked along the morpheme structure, and the morpheme cavity network is formed. Subsequently, the morpheme cavity network is fed into an improved self-attention hypergraph network model. The morpheme embedding module extracts succession chain positions and arranges them into morpheme embedding states. The local hyperedge extraction module latches local hyperedge groups along succession chain positions. Before the local hyperedge groups enter global hyperedge attention, the break bridging module performs dimension-wise subtraction on the hyperedge fragments of the same position boundary, and latches the closable breaks back to the adjacent chain positions through the break locking labeling mechanism, and peels off misaligned breaks into core breaks, forming break resonance hyperedge features. Finally, the global hyperedge attention module aggregates fracture resonance hyperedge features along hyperedge correlations, and reassembles the cleared closure markers and unclosed chain positions into the corresponding corpus fragments to form a Chinese corpus.

[0076] To ensure data authenticity, Chinese corpora from the same batch of corpus construction were selected as verification samples, totaling 120,000 records. Among them, 27,860 records were manually verified to have issues such as duplication, near-rewriting, source misalignment, line breaks, or tag conflicts, while 92,140 records met the conditions for direct inclusion in the database. The platform's original rule-based cleaning method, manual annotation method, similarity deduplication combined with machine annotation method, and the method of this invention were compared. The accuracy rate of corpus cleaning, annotation consistency rate, effective corpus retention rate, and accuracy rate triggered by manual review were statistically analyzed. The results are shown in Table 1 below. Table 1 Comparison of Chinese Corpus Cleaning and Annotation Quality

[0077] As can be seen from the data in Table 1 above, the method of the present invention outperforms the rule-based cleaning and manual annotation methods, as well as the similarity deduplication combined with machine annotation methods, in terms of cleaning accuracy, annotation consistency rate, effective corpus retention rate, and accuracy rate triggered by manual review. The rule-based cleaning and manual annotation method achieves a cleaning accuracy of 88.46%, an annotation consistency rate of 84.73%, an effective corpus retention rate of 86.21%, a manual review trigger accuracy of 79.58%, and an effective corpus deletion error rate of 7.84%. The average processing time per 10,000 corpora is 41.7 minutes, indicating that fixed rules and human experience are difficult to adapt to large-scale complex corpora. The similarity deduplication combined with machine annotation method improves the cleaning accuracy to 92.31%, the annotation consistency rate to 89.64%, and reduces the average processing time per 10,000 corpora to 24.6 minutes. However, the effective corpus deletion error rate remains at 5.26%, and the manual review trigger accuracy rate is 85.42%, indicating that its differentiation of approximate rewriting, source misalignment, and annotation conflicts is still insufficient. The method of this invention achieves a cleaning accuracy of 96.82%, a labeling consistency rate of 94.76%, an effective corpus retention rate of 95.33%, a manual review trigger accuracy of 92.68%, and reduces the effective corpus deletion error rate to 2.18%. The average processing time per 10,000 corpus entries is reduced to 16.9 minutes, indicating that this invention can balance processing efficiency, corpus retention, and quality control.

[0078] This embodiment achieves collaborative processing of Chinese corpus cleaning, source connection, and annotation locking by constructing a morpheme-cell network and combining it with an improved self-attention hypergraph network model. The morpheme-cell network compresses text connection relationships, source traces, and annotation clues into the same structure, enabling line breaks, source separation, annotation dangling, and version differences to form a structural basis before model processing. The local hyperedge extraction module delimits local hyperedge groups from the morpheme embedding state, and the break bridging module performs dimension-wise subtraction of hyperedge pieces with corresponding boundaries before global hyperedge attention aggregation, and uses a break tag locking mechanism to lock closed breaks back to the chain position, and strips misaligned breaks into verification breaks, thereby reducing misjudgment of texts with similar common sentence patterns and direct input of conflicting corpora into the database. The global hyperedge attention module aggregates break resonance hyperedge features, reassembles the cleaned and closed tags and unclosed chain positions back into the corresponding corpus pieces, so that the final Chinese corpus has the ability to retain effective corpora, verify abnormal corpora, and close annotation quality. The overall solution maintains stable processing capabilities even in complex Chinese corpus scenarios with multiple sources, cross-versions, and annotation conflicts. It can accurately distinguish between valid corpus, abnormal breaks, and annotation conflicts, and has good automated processing stability and application value for building large-scale corpora.

[0079] The above are merely preferred embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.

Claims

1. An automated method for cleaning and annotating Chinese corpora, characterized in that, Includes the following steps: S1. Rectify the text format of multi-source Chinese corpus, remove noisy content and cut sentence boundaries, and encapsulate the remaining text into corpus fragments with source traces; S2. Segment the textual hierarchy within the corpus fragments and compress the textual relationships, source traces, and annotation clues into a textual structure; S3. Determine the cross-layer clearing associations along the morphological structure and connect the cross-layer clearing associations into a morphological cell network; S4. Feed the text cell cavity network into the improved self-attention hypergraph network model. The improved self-attention hypergraph network model includes a text embedding module, a local hyperedge extraction module, a break bridging module, and a global hyperedge attention module. The text embedding module extracts the connecting chain positions in the text cell cavity network and arranges them into text embedding states. S5. The local superedge extraction module extracts local chain windows along the receiving chain position of the element embedding state, and delimits the clearing associations that can be closed within the local chain window into local superedge groups. S6. The fracture bridging module accepts the local super-edge group and embeds the fracture locking mechanism. It disassembles the super-edge pieces of the same boundary in the local super-edge group dimension by dimension, snaps the closable fracture back to the adjacent chain position, peels the misaligned fracture into the core fracture, and encapsulates it as the fracture resonance super-edge feature with a closed latch. S7. The global superedge attention module collects fracture resonance superedge features along the superedge correlation, and reloads the cleared closure markers and unclosed chain positions into the corresponding corpus particles to create a Chinese corpus.

2. The automated processing method for cleaning and annotating Chinese corpora according to claim 1, characterized in that, S1 specifically includes: S11. Read the original text and source traces of multi-source Chinese corpus, and paste the source traces to the text boundaries and paragraph positions of the original text to form the original text with traces; S12. Reorganize the text form of the original text with traces, and incorporate the format misalignment, layout residue and line break misconnection caused by cross sources into the same text continuation chain to form a reorganized text. S13. Remove noise fragments that cannot be connected to the sentence boundary along the text continuation chain of the consolidated text, and retain the source traces of adjacent positions of the noise fragments to form the purified text. S14. Cut the sentence boundaries according to the semantic pauses and paragraph connections of the purified text, and encapsulate the cut-off preserved text and source traces into corpus fragments.

3. The automated processing method for cleaning and annotating Chinese corpora according to claim 1, characterized in that, S2 specifically includes: S21. Cut the text fragments along the semantic pauses and contextual connections of the corpus fragments, and merge fragments that cannot be connected independently into adjacent text fragments to form continuous text fragments. S22. Arrange the successor links according to the sequential position of continuous morphemes in the corpus particles, and mark the break, segment crossing and source transition positions as the link boundaries; S23. Extract semantic clues from continuous text fragments to participate in annotation judgment, embed the semantic clues into the corresponding chain positions, and paste the source traces back to the chain position boundary. S24. Encapsulate the continuous fragments of text that have completed chain position arrangement, semantic clue embedding, and source trace re-attachment into a text structure.

4. The automated processing method for cleaning and annotating Chinese corpora according to claim 1, characterized in that, S3 specifically includes: S31. Search for the locations where text continuity is broken, source traces are suspended, and annotation clues are mismatched along the continuity chain of the text structure, and merge the found chain into a cleared chain to be connected. S32. Using the chain to be linked as the center, pull the adjacent text structure, check the text continuation direction, source attribution relationship and annotation placement relationship of adjacent chain positions, and fold the chain positions that can complement each other into cross-layer clearing association. S33. Link the cross-layer clearing markers back to the corresponding element structure, and press the incomplete chain positions into the break position to form an element cell network.

5. The automated processing method for cleaning and annotating Chinese corpora according to claim 1, characterized in that, S4 specifically includes: S41. Expand the network links along the receiving direction of the morpheme cavity network, and gather the morpheme structures, cross-layer clearing associations and break points at the same receiving position into the links to be embedded. S42. Based on the chain position boundary of the chain piece to be embedded, the text is sorted and preserved. The source trace is pasted back to the start and end boundaries of the text. The location of the embedded text is marked. The chain position slot is formed with the boundary constraints. S43. The chain slot is compared with the adjacent chain slot in terms of the direction of continuation. The continuous content is retained in the chain slot. The broken content caused by line breakage, semantic suspension and source separation is pressed into the corresponding chain slot boundary to form a chain slot breakage groove. S44. The content that is completed and continued in the chain slot is unfolded into a text base piece in the order of continuation, and the source traces that fit the boundary and the marking clues that fall into the belonging position are folded into an attachment piece. S45. The text substrate is converted into a chain bit basis vector, the attachment is converted into a boundary attachment vector, and the boundary attachment vector is superimposed on the corresponding boundary of the chain bit basis vector to form a chain bit embedding vector. S46. The chain position break groove is converted into a break embedding piece. The break embedding piece is then attached back to the break boundary of the chain position embedding vector. The chain position embedding vector is arranged in the order of the chain positions in the network to form a textual embedding state.

6. The automated processing method for cleaning and annotating Chinese corpora according to claim 1, characterized in that, S5 specifically includes: S51. The local super-edge extraction module receives the element embedding state, opens the link embedding vector along the link position sequence in the network, and gathers the current link position, the preceding link position and the following link position into a local link window. The break boundary and break embedding piece of the link embedding vector are brought into the local link window along with the link position. S52. Compare the text continuation direction along the preceding and following boundaries within the local chain window. Chain position embedding vectors with the same continuation direction are joined together along the boundary to form a clean text super-edge piece. Chain position embedding vectors with broken continuation directions do not participate in the joining and leave a break mark at the break boundary. S53. The source trace pulls the homologous chain position embedding vector back to the break mark position. When the source boundary can fit the break boundary, the homologous chain position embedding vector is folded into the source trace super edge piece. When the source boundary cannot fit the break boundary, the original break mark connection is retained in the break mark position. S54. When the labeling clue guides the same labeling chain embedding vector to get closer to the text belonging position, when the labeling belonging can fall into the corresponding text boundary, the same labeling chain embedding vector is folded into the labeling hyperedge piece; when the labeling belonging deviates from the corresponding text boundary, the deviated part is labeled as the labeling dangling piece. S55, the clean text super edge piece, the source trace super edge piece and the symbolic super edge piece are arranged in the same position along the same chain position boundary. The break mark retention, the original break mark connection and the symbolic suspended piece are pressed into the corresponding boundary positions to form the super edge piece to be bridged. S56. The superedge pieces to be bridged are encapsulated into local superedge groups according to the sequence of morpheme embedding states.

7. The automated processing method for cleaning and annotating Chinese corpora according to claim 1, characterized in that, S6 specifically includes: S61, the break bridging module accepts the local super-edge group, cuts off the super-edge piece to be bridged at the bridging chain position, and reads the chain position boundary, break mark position, original break mark connection and marked suspended piece carried by the super-edge piece to be bridged. S62. Arrange the super-edge pieces to be bridged in the same position according to the same chain position boundary. Place the clean text super-edge piece into the main bridging slot, place the source trace super-edge piece into the source side slot, and place the label super-edge piece into the label side slot to form a three-slot bridging piece. S63. The source side groove moves back towards the bridging main groove along the source trace, and the label side groove moves closer to the bridging main groove along the text belonging position. The three-groove bridging pieces where the source trace and the text belonging position fall into the same chain position boundary are pressed together to form a co-position bridging piece. The source side groove that deviates from the chain position boundary is peeled out to form a source trace broken piece, and the label side groove that deviates from the text belonging position is peeled out to form a label suspended piece. S64. The bridging main slot in the same position bridging piece is degraded dimensionally with the source side slot and the label side slot respectively. The degradation result is pressed into the same chain position boundary to form the source text break piece and the label text break piece, and then stacked as a break bridging piece. S65, the fault mark locking mechanism completes the snapping, peeling and latching processes for the fault bridging piece, the source fault disconnect piece and the mark suspension piece, and outputs the closed bridging piece and the verification fault piece; S66. The closed bridging pieces are rearranged according to the original acceptance sequence of the local super-edge group, and the broken piece is reattached to the corresponding chain position boundary to form a broken resonance super-edge feature.

8. The automated processing method for cleaning and annotating Chinese corpora according to claim 7, characterized in that, Specifically, S65 includes: S651, the fracture bridging module has an embedded fracture locking mechanism that is connected in series in the order of fracture differential unit, co-positioning engagement unit, misalignment peeling unit and closing latch unit, and supports fracture bridging piece, source fracture break piece and mark suspension piece. S652. The fracture differential unit unfolds the source text fracture piece and the label text fracture piece along the chain position boundary of the fracture bridging piece, reads the direction state of the bridging main groove and the source side groove after the difference is removed in the source text fracture piece, reads the direction state of the bridging main groove and the label side groove after the difference is removed in the label text fracture piece, presses the fracture pieces with the same direction state in the same chain position boundary into the same position differential groove, and presses the fracture pieces with opposite direction state or staggered chain position boundaries into the misaligned temporary storage groove. S653, the corresponding snap-fit ​​unit pulls the corresponding differential slot to fit the adjacent chain position boundary. When the text receiving direction is continuous, the source return direction falls back to the same source trace, and the label placement direction is embedded in the same text belonging position, the corresponding differential slot is snapped back to the original chain position boundary and formed into a closable break piece. When it cannot fit in any direction, the corresponding differential slot is transferred to the misalignment temporary storage slot. S654, the misalignment stripping unit receives the misalignment temporary storage slot, the source trace break piece, and the marked suspended piece. It compares the source break position with the marked suspended position along the chain position boundary. When the source break position and the marked suspended position fall on the same chain position boundary, they are pressed together to form a verification break piece. When the source break position can be backed up but the marked suspended position cannot be embedded, the marked suspended piece is stripped. When the marked suspended position can be embedded but the source break position cannot be backed up, the source trace break piece is stripped. S655, the closing latch unit reattaches the closable break plate to the original chain position boundary of the break bridging plate, and presses the reattached source text break plate, label text break plate and the bridging main groove, source side groove and label side groove in the same position bridging plate into a closed bridging groove, and the closed bridging groove is connected in series along the original receiving direction to form a closed bridging plate. S656. When verifying the chain position boundary of the closed latch unit, if there are residual misalignment temporary storage slots, source trace broken pieces, or marked suspended pieces within the chain position boundary, peel the residual part into the broken piece for verification. If there are no misalignment residues within the chain position boundary, retain the corresponding chain position, source back-attachment relationship, and marked placement relationship in the closed bridging piece.

9. The automated processing method for cleaning and annotating Chinese corpora according to claim 1, characterized in that, Specifically, S7 includes: S71. The global super-edge attention module receives the fracture resonance super-edge feature, disassembles the closed bridging piece and the core fracture piece along the chain position boundary, arranges the closed bridging pieces into a global attention chain, and retains the core fracture piece at the corresponding chain position boundary. S72. Select closed bridging pieces one by one along the global attention chain as target super-edge pieces, read the bearing direction, source back-attachment relationship and label placement relationship of the target super-edge pieces, and compare them with the corresponding relationship of the remaining closed bridging pieces to form super-edge bonding marks. S73. Closed bridging pieces with the same traction direction as the super-edge bonding mark converge towards the target super-edge piece. Closed bridging pieces that can complete the receiving closure and source return are incorporated into the clearing closure piece. Closed bridging pieces that cannot complete the bonding retain the original chain position boundary. S74. The verification fracture piece is reattached along the original chain position boundary to the adjacent clearing and closing piece. The verification fracture piece does not participate in the annotation locking, and the fracture position, source boundary and annotation placement offset are retained. S75. The clearing and closing pieces are backfilled into the corresponding textual structures. The marking clues that have completed the closure are locked as clearing and closing marks. The chain position boundaries that have not completed the closure receive the verification break pieces and are marked as verification chain positions. S76. The morpheme structure with clearing and closing marks is reassembled according to the corpus particle boundaries, and the verification chain position is attached along with the corresponding source trace to form a Chinese corpus with clearing and closing marks.