A Text Data Cleaning and Calibration System Based on a Large Language Model

CN122779063APending Publication Date: 2026-09-18BEIJING XINRUIXIANGTONG TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611221377.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-08-12
Publication Date
2026-09-18

AI Technical Summary

Technical Problem

现有文本清洗方式通常采用字符匹配、规则替换或编辑距离比对,只能处理形式较为明确的差异,面对同义改写、句式变化和跨词元缺失时,难以准确圈定需要校准的文本范围

Benefits of technology

1、通过在Qwen3-8B模型末级解码块中增设参照键投射、参照值投射和局部键值重构结构,将差异片段对应的原文键值从末级注意力读取范围中移除,并将参照文本对应的参照键值接入左右锚点之间,降低错误原文键值对校准词元预测产生的直接干扰,提高校准文本片段与参照文本之间的语义一致性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122779063A_ABST
    Figure CN122779063A_ABST
Patent Text Reader

Abstract

This invention discloses a text data cleaning and calibration system based on a large language model, belonging to the field of text data processing technology. It includes: a text acquisition module for acquiring the text to be cleaned and a reference text; a difference localization module for forming difference fragments; an anchor point construction module for forming a double-anchor calibration patch; an input encoding module for forming a calibration input sequence; a model calibration module for inputting the calibration input sequence into an improved Qwen3-8B model to generate calibration text fragments; a text splicing module for embedding the calibration text fragments between the left and right anchor points to form candidate calibration texts; and a result verification module for performing calibration consistency verification on the calibration text fragments in the candidate calibration texts to form text data cleaning and calibration results. This invention can reduce interference from erroneous original texts and improve the accuracy of text calibration and the continuity of contextual coherence.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of text data processing technology, and in particular to a text data cleaning and calibration system based on a large language model. Background Technology

[0002] As enterprise knowledge bases, business archives, and historical documents continue to accumulate, textual data is prone to problems such as missing characters, duplicate content, misspelled terms, and inconsistencies between local semantics and standard text. Existing text cleaning methods typically employ character matching, rule replacement, or edit distance comparison, which can only handle relatively clear differences. When faced with synonymous rewriting, sentence structure variations, and cross-lexical missing terms, it is difficult to accurately define the text range requiring calibration. Some solutions utilize large language models to regenerate entire text segments, which can improve semantic consistency, but it is easy to rewrite the originally correct content and cause a shift in the connection between the generated text and the original context. At the same time, the differing original text may still continue to participate in the model's attention calculation, interfering with the calibration results; when the generated results lack reference semantic verification and boundary continuity verification, it is also difficult to promptly identify calibration segments that do not conform to the reference content or contextual continuity.

[0003] Therefore, how to provide a text data cleaning and calibration system based on a large language model is a problem that urgently needs to be solved by those skilled in the art. Summary of the Invention

[0004] One objective of this invention is to propose a text data cleaning and calibration system based on a large language model. This invention sets up a reference key-value branch and a local key-value reconstruction structure in the final-level decoding block of the Qwen3-8B model. This removes the original text key-values ​​corresponding to the differing fragments from the final-level attention reading range and connects the reference key-values ​​corresponding to the reference text between the left and right anchor points, reducing the direct interference of erroneous original texts on the calibration term prediction. The calibration position is defined by the left and right anchor points, and reference semantic verification and bilateral continuation verification are performed on the generated calibration text fragments, ensuring that the calibration text fragments simultaneously satisfy the consistency of reference content and the continuity of the original text context. If the verification fails, the original text content of the differing fragments is restored, avoiding local calibration from damaging non-discrepant content. The model structure modification is concentrated in the final-level decoding block. The first to 35th decoding blocks maintain the original structure and parameters, and a local key-value cache and a generated key-value cache are used to separately store static reference content and incrementally generated content, narrowing the model parameter update range and reducing redundant key-value calculations.

[0005] A text data cleaning and calibration system based on a large language model according to an embodiment of the present invention includes: The text acquisition module is used to acquire the text to be cleaned and the reference text, preserving the paragraph boundaries and original character order of the text to be cleaned; The difference localization module is used to perform semantic alignment between the text to be cleaned and the reference text, merge continuous difference positions and shrink difference boundaries to form difference fragments; The anchor point construction module is used to extract adjacent content that corresponds to the semantics of the reference text from both sides of the difference fragment in the text to be cleaned, set left anchor points and right anchor points respectively, and assemble the left anchor points, difference fragments and right anchor points according to the original character position order to form a double anchor calibration patch. The input encoding module is used to extract the reference segment corresponding to the difference segment from the reference text, perform partition encoding on the dual-anchor calibration piece and the reference segment, and form a calibration input sequence; The model calibration module is used to input the calibration input sequence into the improved Qwen3-8B model. The improved Qwen3-8B model adds a reference key-value branch and a local key-value reconstruction structure in the key-value projection path of grouped query attention to generate calibration text fragments. The text splicing module is used to embed the calibration text fragment between the left anchor point and the right anchor point, replace the original text character range corresponding to the difference fragment, and splice the original text content other than the double anchor calibration fragment according to the original character position order to form the candidate calibration text. The result verification module is used to perform calibration consistency verification on calibration text fragments in candidate calibration texts, and generate text data cleaning and calibration results.

[0006] Optionally, the difference localization module specifically comprises: The semantically corresponding text segments are identified from the text to be cleaned and the reference text, and word matching is performed on the semantically corresponding text segments to form a matching word sequence. In the context of texts where the reference text contains words but the text to be cleaned is missing words, placeholders are configured for missing words. Placeholders for words with different word content, words present on one side of the text to be cleaned, or words with missing placeholders are marked as difference positions. The adjacent differences are merged along the original character order of the text to be cleaned, and the words at both ends of the merged interval that are re-aligned with the reference text are excluded to form the difference fragments.

[0007] Optionally, the anchor point construction module specifically includes: Starting from the beginning of the difference segment and moving left and right respectively, check the corresponding word sequence one by one. Extract the words in the text to be cleaned that are adjacent to the difference segment and maintain semantic correspondence as left anchor point and right anchor point respectively. Arrange the left anchor point, difference fragment, and right anchor point according to the original character position order of the text to be cleaned, and record the original character range corresponding to the left anchor point and the right anchor point to form a double anchor calibration patch.

[0008] Optionally, the input encoding module specifically includes: The reference fragment is formed by extracting the text range corresponding to the difference fragment from the semantically corresponding text segment of the reference text; Lexicalization is performed on the dual-anchor calibration piece and the reference fragment respectively to form a dual-anchor lexical sequence and a reference lexical sequence; Configure segment identifiers to distinguish the left anchor word, the difference segment word, the right anchor word, and the reference word. Associate the double anchor word sequence with the original position index and the reference word sequence with the corresponding position index. Connect the calibration start identifier at the end of the reference word sequence and arrange them in the order of double anchor word sequence, reference word sequence, and calibration start identifier to form the calibration input sequence.

[0009] Optionally, the lexicalization process performed on the dual-anchor calibration piece and the reference fragment specifically includes: The word segmentation rules of the Qwen3-8B model are used to perform word segmentation on the left anchor point, the difference segment, the right anchor point and the reference segment respectively; The left anchor word, the difference fragment word, and the right anchor word are respectively bound to the corresponding original character range in the text to be cleaned, and arranged according to the original character position to form a double anchor word sequence; Arrange the reference words according to the character order of the reference segment, and record the position index between each reference word and the corresponding position of the difference segment to form a reference word sequence.

[0010] Optionally, the improved Qwen3-8B model includes a lexical embedding layer, multiple sequentially connected decoding blocks, an output normalization layer, and a lexical prediction layer; The decoding block includes a group query attention sublayer and a feedforward network sublayer. The input ends of the group query attention sublayer and the feedforward network sublayer are configured with normalization processing, and the corresponding input hidden states are passed through the residual path respectively. In a series of sequentially connected decoding blocks, the final-level decoding block connected to the output normalization layer is set as the modified decoding block. The final-level decoding block adds reference key projection and reference value projection on the side of the key projection and value projection of the grouped query attention sub-layer, and sets a local key-value reconstruction structure between the original key-value projection result, the reference key-value projection result and the attention operation. The lexical embedding layer maps each lexical in the calibration input sequence to an input hidden state, forming an input hidden state sequence. The input hidden state sequence is passed layer by layer to the final decoding block along the connection order of the decoding blocks. The reference key projection, reference value projection, and local key-value reconstruction structure process the input hidden state of the final-level decoding block to form a local key-value sequence; In the word-by-word generation process of the calibration text fragment, the final-level decoding block uses local key-value sequences to complete group query attention processing and feedforward processing, forming the calibration hidden state of the current prediction position; The output normalization layer and the lexical prediction layer process the calibration hidden state at the current prediction position in turn, predict the current calibration lexical, and then continue the current calibration lexical into the calibration input sequence and enter the next prediction position until the end marker is predicted. The calibration text fragment is then restored according to the lexical generation order.

[0011] Optionally, the reference key projection, reference value projection, and local key-value reconstruction structure process the input hidden state of the final-level decoding block to form a local key-value sequence, specifically as follows: According to the segment identifier, extract the left anchor point hidden state, the difference segment hidden state, the right anchor point hidden state, and the reference word hidden state from the input hidden state sequence of the final decoding block respectively. The original key projection and value projection of the grouped query attention sublayer are used to process the hidden states of left anchor points, difference fragments, and right anchor points, forming the original text key sequence and original text value sequence; the reference key projection and reference value projection are used to process the hidden states of reference terms, forming the reference key sequence and reference value sequence. Remove the original key subsequences and original value subsequences corresponding to the difference segments from the original key sequence and original value sequence, and retain the original key sequence and original value subsequences corresponding to the left anchor and the right anchor; When a reference segment contains reference terms, the reference key sequence and reference value sequence are arranged according to their position indices. When the same position index corresponds to multiple reference terms, they are arranged according to the character order of the multiple reference terms in the reference segment. The arranged reference key sequence is then connected between the left anchor text key sequence and the right anchor text key sequence, and the arranged reference value sequence is connected between the left anchor text value sequence and the right anchor text value sequence. When a reference segment does not contain reference terms, the left anchor text key sequence and the right anchor text key sequence are connected, and the left anchor text value sequence and the right anchor text value sequence are also connected. Configure the local position index according to the arrangement order of the left anchor point, reference fragment and right anchor point, perform RMSNorm processing on each key vector in the local key sequence along the key value head state dimension, perform rotation position encoding according to the local position index, and pair the processed local key sequence with the local value sequence to form a local key value sequence.

[0012] Optionally, the final-level decoding block uses local key-value sequences to complete group query attention processing and feedforward processing, forming the calibration hidden state of the current prediction position, specifically: The first prediction position obtains the input hidden state of the final-level decoder block corresponding to the calibration start identifier, and other prediction positions obtain the input hidden state of the final-level decoder block corresponding to the latest generated calibration term. The obtained input hidden states of the final-level decoder block are processed by pre-RMSNorm and query projection to form query vectors corresponding to multiple query heads. To calibrate the starting identifier and each generated calibration term, a generated position index is sequentially configured to follow the local position index. RMSNorm processing is performed on each query vector along the query header state dimension, and rotation position encoding is performed according to the generated position index corresponding to the current predicted position. Obtain the calibration start identifier and the input hidden state of the final-level decoding block corresponding to the calibration tokens generated before the current prediction position. Use key projection and value projection to form the generated key sequence and generated value sequence. Perform RMSNorm processing on each key vector in the generated key sequence along the key value head state dimension, and perform rotation position encoding according to the corresponding generated position index. Pair the processed generated key sequence with the generated value sequence and continue to the local key value sequence. Based on the grouping relationship between the query header and the key value header, multiple query headers are assigned to corresponding query groups, and query vectors in the same query group share the key sequence and value sequence of the corresponding key value header; Calculate the dot product between each query vector and each key vector in the corresponding key sequence. Multiply the dot product result by the reciprocal of the square root of the query header dimension to form a scaled dot product. Perform exponential normalization on all scaled dot products corresponding to the same query vector to form the attention weights corresponding to each key value position. Each attention weight is multiplied with its corresponding value vector dimension by dimension. The multiplication results corresponding to the same query vector are accumulated position by position to form the attention results of each query head. The attention results of multiple query heads are concatenated and output projection is performed. The output projection result is added in the same position to the input hidden state of the final-level decoded block corresponding to the current prediction position to form the attention hidden state of the current prediction position. Normalize the attention hidden state, and input the normalization result into the gated projection path and the rising projection path of the feedforward network sub-layer respectively. Use the SiLU activation function to process the gated projection result, multiply the activated gated projection result with the rising projection result dimension by dimension, and restore it to the input hidden state dimension of the final decoding block through the falling projection path. The descent projection result is added in the same position to the attention hidden state to form the output hidden state of the final-level decoder block at the current prediction position, and the output hidden state of the final-level decoder block at the current prediction position is determined as the calibration hidden state.

[0013] Optionally, the text concatenation module specifically includes: Embed the calibration text fragment between the left and right anchor points, and replace the original text character range corresponding to the difference fragment; The candidate calibration text is formed by sequentially concatenating the original text before the double anchor calibration patch, the left anchor point, the calibration text fragment, the right anchor point, and the original text after the double anchor calibration patch according to the original character position order.

[0014] Optionally, the result verification module specifically includes: Perform semantic alignment between the calibration text fragments in the candidate calibration text and the reference fragments, and verify the semantic connection between the calibration text fragments and the left and right anchor points; When the calibration text fragment maintains semantic correspondence with the reference fragment and semantic continuity with the left and right anchor points, the calibration text fragment is retained; When the calibration text fragment does not maintain semantic correspondence with the reference fragment or does not maintain semantic continuity with the left and right anchor points, the calibration text fragment is replaced with the original text content corresponding to the difference fragment, thus forming the text data cleaning and calibration results.

[0015] The beneficial effects of this invention are: 1. By adding reference key projection, reference value projection and local key value reconstruction structure to the final decoding block of the Qwen3-8B model, the original text key value corresponding to the difference segment is removed from the final attention reading range, and the reference key value corresponding to the reference text is connected between the left and right anchor points. This reduces the direct interference of the erroneous original text key value on the calibration lexical prediction and improves the semantic consistency between the calibration text segment and the reference text.

[0016] 2. By constructing left and right anchor points around the difference fragments, and performing reference semantic verification and bilateral continuation verification after the calibration text fragments are generated, the calibration text fragments are simultaneously constrained by the reference content and the original text context; when the verification fails, the original text content of the difference fragments is restored, reducing semantic breaks and non-difference content rewriting caused by local calibration.

[0017] 3. The model structure modification is concentrated on the last-level decoding block. The structure and parameters of the first to the 35th decoding blocks remain unchanged. The dual anchor reference content and the generated word key values ​​are stored through local key value caching and generated key value caching, respectively, which reduces the full model modification and repeated key value reconstruction, and reduces the training range of the model and the amount of repeated calculations in the word-by-word generation process. Attached Figure Description

[0018] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings: Figure 1 This is a structural diagram of a text data cleaning and calibration system based on a large language model proposed in this invention; Figure 2 This is a schematic diagram of the improved Qwen3-8B model and local key-value reconstruction structure of a text data cleaning and calibration system based on a large language model proposed in this invention. Detailed Implementation

[0019] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams, illustrating only the basic structure of the invention, and therefore only show the components relevant to the invention.

[0020] refer to Figure 1 and Figure 2 A text data cleaning and calibration system based on a large language model includes: The text acquisition module is used to acquire the text to be cleaned and the reference text, preserving the paragraph boundaries and original character order of the text to be cleaned; The difference localization module is used to perform semantic alignment between the text to be cleaned and the reference text, merge continuous difference positions and shrink difference boundaries to form difference fragments; The anchor point construction module is used to extract adjacent content that corresponds to the semantics of the reference text from both sides of the difference fragment in the text to be cleaned, set left anchor points and right anchor points respectively, and assemble the left anchor points, difference fragments and right anchor points according to the original character position order to form a double anchor calibration patch. The input encoding module is used to extract the reference segment corresponding to the difference segment from the reference text, perform partition encoding on the dual-anchor calibration piece and the reference segment, and form a calibration input sequence; The model calibration module is used to input the calibration input sequence into the improved Qwen3-8B model. The improved Qwen3-8B model adds a reference key-value branch and a local key-value reconstruction structure in the key-value projection path of grouped query attention to generate calibration text fragments. The text splicing module is used to embed the calibration text fragment between the left anchor point and the right anchor point, replace the original text character range corresponding to the difference fragment, and splice the original text content other than the double anchor calibration fragment according to the original character position order to form the candidate calibration text. The result verification module is used to perform calibration consistency verification on calibration text fragments in candidate calibration texts, and generate text data cleaning and calibration results.

[0021] In this embodiment, the difference localization module specifically includes: The semantically corresponding text segments are identified from the text to be cleaned and the reference text, and word matching is performed on the semantically corresponding text segments to form a matching word sequence. In the context of texts where the reference text contains words but the text to be cleaned is missing words, placeholders are configured for missing words. Placeholders for words with different word content, words present on one side of the text to be cleaned, or words with missing placeholders are marked as difference positions. The adjacent differences are merged along the original character order of the text to be cleaned, and the words at both ends of the merged interval that are re-aligned with the reference text are excluded to form the difference fragments.

[0022] In specific implementation, sentence units are first segmented respectively according to paragraph boundaries in the text to be cleaned and the reference text, and the arrangement order of each sentence unit in the original text is maintained. The fast tokenizer matching the Qwen3-8B model is called to perform token encoding on each sentence unit respectively, the token encoding is input into the token embedding layer and the 1st to 35th decoding blocks, and the hidden states corresponding to each non-padding token position are obtained; the hidden states of the same sentence unit are averaged dimension by dimension along token positions, and vector length normalization is performed on the averaging result to form the sentence semantic vector of the corresponding sentence unit. The cosine similarity between the sentence semantic vector of the text to be cleaned and the sentence semantic vector of the reference text is calculated. Sentence units with a cosine similarity reaching 0.78 enter the candidate correspondence range; when the same sentence unit has multiple candidate correspondence ranges, the candidate correspondence range with the maximum cosine similarity is retained, and the arrangement direction of adjacent candidate correspondence ranges is limited to be consistent with the sentence order of the original text. When there are unmatched sentence units between two adjacent selected sentence correspondence pairs, the unmatched sentence units located between the two sentence correspondence pairs are merged into the text segment jointly defined by the two sentence correspondence pairs; when there are unmatched sentence units at the paragraph start or paragraph end, the unmatched sentence units are merged into the text segment defined by the closest sentence correspondence pair. The continuous text range including the corresponding sentence units and unmatched sentence units is determined as a semantically corresponding text segment.

[0023] Using the fast tokenizer matching the same Qwen3-8B model as that in the input encoding stage, token segmentation is performed on the two semantically corresponding text segments respectively, and minimum edit distance alignment is performed on the token segmentation results. An alignment operation with the same token content is recorded as a match, an alignment operation with different token content is recorded as a substitution, an alignment operation with tokens only on the side of the text to be cleaned is recorded as a deletion, and an alignment operation with tokens only on the side of the reference text is recorded as an insertion. Backtracking the alignment path according to the minimum cumulative edit distance, the tokens in the two text segments are arranged into position-by-position corresponding aligned token sequences.

[0024] For insertion operations, a missing placeholder is configured between two adjacent original characters in the text to be cleaned. The missing placeholder does not carry any text character, and the missing placeholder records the end sequence of the previous original character and the start sequence of the subsequent original character; when the insertion operation is located at the start of the paragraph, the missing placeholder records the insertion position before the first character; when the insertion operation is located at the end of the paragraph, the missing placeholder records the insertion position after the last character. Positions corresponding to substitution operations, positions corresponding to deletion operations, and positions where missing placeholders are configured are marked as difference positions, and positions corresponding to matching operations remain as consistent positions.

[0025] The differences are examined along the direction of the sequence of corresponding words, and directly adjacent differences are merged into the same difference interval. When a difference interval contains both words to be cleaned and missing placeholders, the original text range is determined according to the starting character position of the first word to be cleaned and the ending character position of the last word to be cleaned. When the difference interval contains only missing placeholders, the positions between adjacent characters of the missing placeholder records are determined as the zero-length original text range.

[0026] For each difference interval, along with the two identical words on each side, local word alignment is performed again. Starting from the left end of the difference interval, the words in the text to be cleaned and the words in the reference text are checked position by position, and edge words with identical content after realignment are deleted. Then, the same check is performed from the right end of the difference interval until the first alignment position with different content, unilateral presence, or missing placeholders is reached at both ends. The text content to be cleaned between the first and last difference positions after shrinking is determined as the difference segment, and the starting character position, ending character position, and missing placeholder positions corresponding to the difference segment are recorded.

[0027] In this embodiment, the anchor point construction module specifically includes: Starting from the beginning of the difference segment and moving left and right respectively, check the corresponding word sequence one by one. Extract the words in the text to be cleaned that are adjacent to the difference segment and maintain semantic correspondence as left anchor point and right anchor point respectively. Arrange the left anchor point, difference fragment, and right anchor point according to the original character position order of the text to be cleaned, and record the original character range corresponding to the left anchor point and the right anchor point to form a double anchor calibration patch.

[0028] In the specific implementation process, the starting character position, ending character position, and missing placeholder position corresponding to the difference segment are read, and the first and last difference positions corresponding to the difference segment are located in the corresponding word sequence. When the difference segment only contains missing placeholders, the insertion position of the missing placeholder between adjacent original characters is used as the starting and ending positions of the difference segment; when the difference segment contains words of the text to be cleaned, the original character range of the difference segment is limited by the starting character position of the foremost word and the ending character position of the last word.

[0029] Starting from the position preceding the first discrepancy position, check the sequence of paired words one by one to the left. If the current paired position contains both the word to be cleaned and the reference word, and is recorded as a consistent position during the discrepancy localization process, include the word to be cleaned in the left anchor point; continue checking adjacent paired positions to the left until a discrepancy position, missing placeholder, paragraph boundary, or the left anchor point already contains 8 words to be cleaned is encountered. Rearrange the words to be cleaned included in the left anchor point according to the original character order, and record the starting character order of the first word and the ending character order of the last word as the original character range corresponding to the left anchor point.

[0030] Starting from the position following the last difference position, check the paired word sequence position by position to the right. If the current paired position contains both the word to be cleaned and the reference word, and the current paired position is recorded as a consistent position, include the word to be cleaned in the right anchor point; continue checking adjacent paired positions to the right until a difference position, missing placeholder, paragraph boundary, or the right anchor point already contains 8 words to be cleaned is encountered. Arrange the words to be cleaned included in the right anchor point according to the original character order, and record the starting character order of the first word and the ending character order of the last word as the original character range corresponding to the right anchor point.

[0031] When the difference segment is located at the beginning of a paragraph, the position of the paragraph's starting character is recorded as the zero-length original text character range of the left anchor point; when the difference segment is located at the end of a paragraph, the position after the paragraph's ending character is recorded as the zero-length original text character range of the right anchor point. The zero-length original text character range only records the anchor position to ensure that the difference segment still has a clear bilateral position boundary when it is at the beginning or end of a paragraph.

[0032] Following the order of the original text character ranges for the left anchor point, the difference fragment, and the right anchor point, the corresponding original text content is extracted from the text to be cleaned. The left and right anchor points retain their original text characters, while the difference fragments retain the original text content or missing placeholder positions determined during the difference positioning stage. The left anchor point, difference fragment, and right anchor point are then continuously assembled according to their original character order to form a dual-anchor calibration patch. The dual-anchor calibration patch simultaneously records the original text character ranges corresponding to each of the left anchor point, difference fragment, and right anchor point.

[0033] In this embodiment, the input encoding module is specifically as follows: The reference fragment is formed by extracting the text range corresponding to the difference fragment from the semantically corresponding text segment of the reference text; Lexicalization is performed on the dual-anchor calibration piece and the reference fragment respectively to form a dual-anchor lexical sequence and a reference lexical sequence; Configuring mutually distinguishable segment identifiers for the left anchor token, the difference fragment token, the right anchor token and the reference token, associating the dual-anchor token sequence with an original position index, associating the reference token sequence with an aligned position index, adding a calibration start identifier at the end of the reference token sequence, and arranging them in the order of the dual-anchor token sequence, the reference token sequence and the calibration start identifier to form a calibration input sequence.

[0034] In this embodiment, tokenization processing is performed on the dual-anchor calibration fragment and the reference fragment, specifically: Adopting the token splitting rule of the Qwen3-8B model to perform token splitting on the left anchor, the difference fragment, the right anchor and the reference fragment respectively; Binding the left anchor token, the difference fragment token and the right anchor token respectively to the corresponding original character intervals in the text to be cleaned, and arranging them according to the original character order to form a dual-anchor token sequence; Arranging the reference tokens according to the character arrangement order of the reference fragment, and recording the aligned position index between each reference token and the corresponding position of the difference fragment to form a reference token sequence.

[0035] In a specific implementation process, the aligned token sequence, the first difference position and the last difference position saved during the difference positioning process are read. Each aligned position in the aligned token sequence records the token on the side of the text to be cleaned, the token on the side of the reference text, and the character intervals corresponding to the tokens on both sides respectively. Tokens located on the reference text side between the first difference position and the last difference position are extracted, and the starting character order of the foremost reference text side token to the ending character order of the rearmost reference text side token is determined as the reference text interception range, and corresponding characters are intercepted from the reference text to form a reference fragment. When there is no token on the reference text side within the difference range, the reference text insertion order corresponding to the first difference position is recorded as the starting order and ending order of the reference fragment to form a zero-length reference fragment. The reference fragment and the difference fragment share the first difference position and the last difference position, so as to avoid re-dividing the corresponding range.

[0036] After the text to be cleaned and the reference text are read, they are continuously numbered from 0 according to Unicode code points, and no full-width half-width conversion, case conversion or Unicode normalization processing is performed. The fast tokenizer matching the Qwen3-8B model is called, automatic addition of start tokens, end tokens and other special tokens is disabled, and character offset return is enabled. The tokenizer returns token codes and the character start offset and character end offset of each token in the input text for the input text, and the character end offset points to the next character position at the end of the character interval covered by the token.

[0037] The word segmenter is invoked with the left anchor point, the difference segment, the right anchor point, and the reference segment as independent inputs. The left anchor point, the difference segment, and the right anchor point are segmented independently, and words are prohibited from crossing the boundary between the left anchor point and the difference segment, or between the difference segment and the right anchor point. The word segmentation results are concatenated in the order of left anchor point words, difference segment words, and right anchor words to form a double-anchor word sequence. When the difference segment is within the zero-length original text range, the difference segment word sequence is empty; when the reference segment is within the zero-length reference range, the reference word sequence is empty.

[0038] The character offsets of the left-side anchor words are added to the starting character position of the left anchor in the text to be cleaned. Similarly, the character offsets of the difference fragment words are added to the starting character position of the difference fragment in the text to be cleaned. The character offsets of the right-side anchor words are added to the starting character position of the right anchor in the text to be cleaned. This yields the starting and ending character positions of each double-anchor word within the entire text to be cleaned. These starting and ending character positions are used as the original position indices for the corresponding double-anchor words. When the difference fragment is within the zero-length original text range, missing placeholders retain the insertion order between adjacent original text characters; missing placeholders are not converted into text words.

[0039] Reference terms are arranged according to the character order in the reference segment. The character offset of each reference term is added to the starting character position of the reference segment in the reference text to obtain the character range of each reference term within the entire reference text. The number of overlapping characters between the reference term character range and the corresponding character range in the reference text is calculated sequentially. The position with the largest number of overlapping characters is recorded as the corresponding position index of the reference term. When multiple corresponding positions have the same number of overlapping characters, the position that appears first is recorded. When a reference term corresponds to a missing placeholder, the corresponding position of the missing placeholder is recorded. When multiple reference terms correspond to the same missing placeholder, the same corresponding position index is used for all reference terms, while maintaining the order of the reference terms in the reference segment.

[0040] A segment identifier of 0 is assigned to the left anchor lexicon, a segment identifier of 1 is assigned to the difference segment lexicon, a segment identifier of 2 is assigned to the right anchor lexicon, and a segment identifier of 3 is assigned to the reference lexicon. The segment identifiers are stored independently as integer sequences, neither converted to segment embeddings nor added to lexicon embeddings. Each segment identifier corresponds one-to-one with a lexicon sequence position. The final-level decoding block extracts the left anchor hidden state, the difference segment hidden state, the right anchor hidden state, and the reference lexicon hidden state from the input hidden state sequence according to the segment identifiers.

[0041] A special lexical unit "<|calib_start|>" is added to the lexical table accompanying the Qwen3-8B model. This special lexical unit is registered as an additional special lexical unit that does not participate in ordinary text segmentation, and the lexical parameter rows of the lexical embedding layer and the lexical prediction layer are expanded simultaneously. The lexical code corresponding to the special lexical unit serves as the calibration start identifier, with the segment identifier configured as 4, and both the original position index and the corresponding position index configured as -1. When the reference lexical sequence is not empty, the calibration start identifier is appended to the end of the reference lexical sequence; when the reference lexical sequence is empty, the calibration start identifier is directly appended to the end of the double-anchor lexical sequence.

[0042] All lexical codes are concatenated in the order of the double-anchor lexical sequence, the reference lexical sequence, and the calibration start identifier to form a lexical code sequence. Corresponding segment identifiers are concatenated in the same order to form a segment identifier sequence. The original positional index is written to the position corresponding to the double-anchor lexical, and -1 is written to the positional index corresponding to the reference lexical; the original positional index is written to the position corresponding to the reference lexical, and -1 is written to the original positional index. Both types of indices at the position corresponding to the calibration start identifier are written to -1. The lexical code sequence, segment identifier sequence, original positional index sequence, and positional index sequence maintain the same length and together constitute the calibration input sequence. The lexical code sequence enters the lexical embedding layer; the segment identifier sequence and positional index sequence are bound to the input hidden state sequence position by position and passed to the final decoding block; the original positional index sequence is retained until text concatenation processing.

[0043] In this embodiment, the improved Qwen3-8B model includes a lexical embedding layer, multiple sequentially connected decoding blocks, an output normalization layer, and a lexical prediction layer. The decoding block includes a group query attention sublayer and a feedforward network sublayer. The input ends of the group query attention sublayer and the feedforward network sublayer are configured with normalization processing, and the corresponding input hidden states are passed through the residual path respectively. In a series of sequentially connected decoding blocks, the final-level decoding block connected to the output normalization layer is set as the modified decoding block. The final-level decoding block adds reference key projection and reference value projection on the side of the key projection and value projection of the grouped query attention sub-layer, and sets a local key-value reconstruction structure between the original key-value projection result, the reference key-value projection result and the attention operation. The lexical embedding layer maps each lexical in the calibration input sequence to an input hidden state, forming an input hidden state sequence. The input hidden state sequence is passed layer by layer to the final decoding block along the connection order of the decoding blocks. The reference key projection, reference value projection, and local key-value reconstruction structure process the input hidden state of the final-level decoding block to form a local key-value sequence; In the word-by-word generation process of the calibration text fragment, the final-level decoding block uses local key-value sequences to complete group query attention processing and feedforward processing, forming the calibration hidden state of the current prediction position; The output normalization layer and the lexical prediction layer process the calibration hidden state at the current prediction position in turn, predict the current calibration lexical, and then continue the current calibration lexical into the calibration input sequence and enter the next prediction position until the end marker is predicted. The calibration text fragment is then restored according to the lexical generation order.

[0044] In this embodiment, the reference key projection, reference value projection, and local key-value reconstruction structure process the input hidden state of the final-level decoding block to form a local key-value sequence, specifically as follows: According to the segment identifier, extract the left anchor point hidden state, the difference segment hidden state, the right anchor point hidden state, and the reference word hidden state from the input hidden state sequence of the final decoding block respectively. The original key projection and value projection of the grouped query attention sublayer are used to process the hidden states of left anchor points, difference fragments, and right anchor points, forming the original text key sequence and original text value sequence; the reference key projection and reference value projection are used to process the hidden states of reference terms, forming the reference key sequence and reference value sequence. Remove the original key subsequences and original value subsequences corresponding to the difference segments from the original key sequence and original value sequence, and retain the original key sequence and original value subsequences corresponding to the left anchor and the right anchor; When a reference segment contains reference terms, the reference key sequence and reference value sequence are arranged according to their position indices. When the same position index corresponds to multiple reference terms, they are arranged according to the character order of the multiple reference terms in the reference segment. The arranged reference key sequence is then connected between the left anchor text key sequence and the right anchor text key sequence, and the arranged reference value sequence is connected between the left anchor text value sequence and the right anchor text value sequence. When a reference segment does not contain reference terms, the left anchor text key sequence and the right anchor text key sequence are connected, and the left anchor text value sequence and the right anchor text value sequence are also connected. Configure the local position index according to the arrangement order of the left anchor point, reference fragment and right anchor point, perform RMSNorm processing on each key vector in the local key sequence along the key value head state dimension, perform rotation position encoding according to the local position index, and pair the processed local key sequence with the local value sequence to form a local key value sequence.

[0045] In this embodiment, the final-level decoding block uses local key-value sequences to complete group query attention processing and feedforward processing, forming the calibration hidden state of the current prediction position, specifically: The first prediction position obtains the input hidden state of the final-level decoder block corresponding to the calibration start identifier, and other prediction positions obtain the input hidden state of the final-level decoder block corresponding to the latest generated calibration term. The obtained input hidden states of the final-level decoder block are processed by pre-RMSNorm and query projection to form query vectors corresponding to multiple query heads. To calibrate the starting identifier and each generated calibration term, a generated position index is sequentially configured to follow the local position index. RMSNorm processing is performed on each query vector along the query header state dimension, and rotation position encoding is performed according to the generated position index corresponding to the current predicted position. Obtain the calibration start identifier and the input hidden state of the final-level decoding block corresponding to the calibration tokens generated before the current prediction position. Use key projection and value projection to form the generated key sequence and generated value sequence. Perform RMSNorm processing on each key vector in the generated key sequence along the key value head state dimension, and perform rotation position encoding according to the corresponding generated position index. Pair the processed generated key sequence with the generated value sequence and continue to the local key value sequence. Based on the grouping relationship between the query header and the key value header, multiple query headers are assigned to corresponding query groups, and query vectors in the same query group share the key sequence and value sequence of the corresponding key value header; Calculate the dot product between each query vector and each key vector in the corresponding key sequence. Multiply the dot product result by the reciprocal of the square root of the query header dimension to form a scaled dot product. Perform exponential normalization on all scaled dot products corresponding to the same query vector to form the attention weights corresponding to each key value position. Each attention weight is multiplied with its corresponding value vector dimension by dimension. The multiplication results corresponding to the same query vector are accumulated position by position to form the attention results of each query head. The attention results of multiple query heads are concatenated and output projection is performed. The output projection result is added in the same position to the input hidden state of the final-level decoded block corresponding to the current prediction position to form the attention hidden state of the current prediction position. Normalize the attention hidden state, and input the normalization result into the gated projection path and the rising projection path of the feedforward network sub-layer respectively. Use the SiLU activation function to process the gated projection result, multiply the activated gated projection result with the rising projection result dimension by dimension, and restore it to the input hidden state dimension of the final decoding block through the falling projection path. The descent projection result is added in the same position to the attention hidden state to form the output hidden state of the final-level decoder block at the current prediction position, and the output hidden state of the final-level decoder block at the current prediction position is determined as the calibration hidden state.

[0046] In the specific implementation, the improved Qwen3-8B model retains the original Qwen3-8B model's lexical embedding layer, 36 sequentially connected decoding blocks, output normalization layer, and lexical prediction layer. The lexical embedding dimension and the hidden state dimension of the decoding blocks are both 4096, and the intermediate dimension of the feedforward network is 12288. The grouped query attention sublayer includes 32 query heads and 8 key-value heads, with every 4 query heads sharing 1 key-value head, and each head containing 128 state components. The first to 35th decoding blocks retain the original structure. The 36th decoding block, adjacent to the output normalization layer, serves as the final decoding block, adding reference key projection and reference value projection alongside the original key projection and value projection, and connecting a local key-value reconstruction structure between the original key-value projection result, the reference key-value projection result, and the attention operation.

[0047] Before the improvement, both the double-anchor lexicons and the reference lexicons underwent the same key projection and value projection, and the original text key values ​​corresponding to the difference fragments remained within the attention reading range. When generating calibration text, the query vector simultaneously reads the difference original text, the reference content, and the context on both sides. Errors already present in the difference original text may continue to affect the prediction of calibration lexicons. After the model modification, the first 35 decoding blocks continue to extract general semantic information from the calibration input sequence. The final decoding block then separates the hidden states corresponding to the left anchor point, the difference fragment, the right anchor point, and the reference lexicons according to the segment identifier. The reference lexicons form a reference key-value sequence through independent reference key projection and reference value projection. The local key-value reconstruction structure removes the original text key values ​​corresponding to the difference fragments and connects the reference key values ​​between the original text key values ​​of the left and right anchor points, so that the final decoding block directly generates calibration text fragments around the left and right anchor points and the reference content.

[0048] The token encodings in the calibration input sequence are fed into the token embedding layer and mapped to the input hidden state sequence according to the token arrangement order. The input hidden state sequence is then fed into the first to the 35th decoding blocks. Each decoding block first performs RMSNorm processing on the input hidden state. RMSNorm calculates the average of the squared values ​​of each state component of the input hidden state, adds 10^-6 to the average, takes the square root, divides each state component by the square root result, and multiplies it by the corresponding scaling parameter to form the normalized hidden state. The normalized hidden state is projected into query vectors, key vectors, and value vectors. RMSNorm processing is performed along the state dimension of each query head and each key-value head for each query vector, and rotational position encoding is performed according to the corresponding token positions. Grouped query attention is performed according to the grouping relationship between the query head and the key-value head. The attention result is projected onto the output and added position by position to the input hidden state of the decoding block to form the attention hidden state. The attention hidden state is processed again by RMSnorm, and then gated projection, rising projection, SiLU activation, dimension-wise multiplication, and falling projection are performed sequentially. The feedforward processing result is added to the attention hidden state position by position to form the output hidden state of the current decoding block. The output hidden state of the previous decoding block is used as the input hidden state of the next decoding block. The 35th decoding block outputs the final-level input hidden state sequence, which maintains a position-wise correspondence with the segment identifier sequence.

[0049] The final-level decoding block extracts hidden states at different positions according to segment identifiers. Segment identifier 0 corresponds to the left anchor hidden state, segment identifier 1 corresponds to the difference segment hidden state, segment identifier 2 corresponds to the right anchor hidden state, segment identifier 3 corresponds to the reference term hidden state, and segment identifier 4 corresponds to the calibration start identifier hidden state. The original key projection and value projection of the final-level decoding block process the left anchor hidden state, the difference segment hidden state, and the right anchor hidden state, respectively, forming the original key sequence and the original value sequence; the newly added reference key projection and reference value projection only process the reference term hidden state, forming the reference key sequence and the reference value sequence. The input dimension of the reference key projection parameter matrix and the reference value projection parameter matrix are both 4096, and the output dimension is both 1024, corresponding to 8 key-value heads and 128-dimensional states of each key-value head.

[0050] The local key-value reconstruction structure locates the original key sequence and original value subsequence corresponding to the difference segment based on segment identifier 1, and removes both types of subsequences from the final attention reading range. The original key values ​​corresponding to the left and right anchors remain unchanged. When the reference segment contains reference terms, the reference key vector and reference value vector are arranged in ascending order according to the alignment index; when multiple reference terms have the same alignment index, they are arranged according to the character order of the reference terms in the reference segment; the arranged reference key sequence is connected between the original key sequence of the left anchor and the original key sequence of the right anchor, and the arranged reference value sequence is connected between the original value subsequence of the left anchor and the original value subsequence of the right anchor; when the reference segment is empty, the original key sequence of the left anchor and the original key sequence of the right anchor are connected, and the original value subsequence of the left anchor and the original value subsequence of the right anchor are connected.

[0051] Local positional indices are sequentially arranged starting from 0. Left-side anchor terms occupy the first index, reference terms occupy the middle index, and right-side anchor terms occupy the last index. The local key sequence is split into eight key-value headers. The key vectors in each header undergo RMSNorm processing and rotational position encoding according to the local positional index. The local value sequence retains the state components formed by the original value projection or reference value projection. The processed local key sequence and local value sequence are paired position-by-position to form the local key-value sequence. The local key-value sequence is stored in a separately configured local key-value cache in the final decoding block, and is not written to the original key-value cache called by the first to 35th decoding blocks. For the same dual-anchor calibration piece, the local key-value cache is established only once at the start of calibration text fragment generation and remains unchanged at each prediction position.

[0052] The end marker uses the existing end lexicon in the Qwen3-8B model lexicon table. When processing a new dual-anchor calibration piece, the local key-value cache and generated key-value cache stored in the final-level decoding block are first cleared, and then a local key-value cache is established according to the current dual-anchor calibration piece; the local key-values ​​and generated key-values ​​corresponding to the previous dual-anchor calibration piece do not participate in the lexicon prediction of the current dual-anchor calibration piece. Before performing exponential normalization, the prediction scores corresponding to the calibration start marker and other special lexicons (excluding the end marker) are set to negative infinity, making the probability of the lexicons corresponding to the above special lexicons zero.

[0053] When the calibration input sequence first enters the improved Qwen3-8B model, the lexical embedding layer and the first through 35th decoding blocks process all the lexical units in the calibration input sequence at once, and write the key and value sequences formed by each decoding block into the corresponding original key-value buffers. The 35th decoding block outputs the final-level input hidden state corresponding to each input lexical unit. The final-level decoding block establishes a local key-value buffer based on the final-level input hidden state and retains the final-level input hidden state corresponding to the calibration start identifier.

[0054] The first prediction position obtains the final-level input hidden state corresponding to the calibration start identifier, while other prediction positions obtain the final-level input hidden state corresponding to the latest generated calibration lexical. Pre-RMSNorm processing is performed on the final-level input hidden state corresponding to the current prediction position, and then query vectors corresponding to 32 query heads are formed through query projection. Each query vector undergoes RMSNorm processing along the state dimension within the head, and rotational position encoding is performed according to the generation position index corresponding to the current prediction position. After the text lexical is generated at the current prediction position, it enters the lexical embedding layer, and the original key-value cache from the 1st to the 35th decoding blocks is called to calculate layer by layer, forming the final-level input hidden state corresponding to the next prediction position. Each prediction position only calculates the hidden state corresponding to the newly accessed lexical, without repeatedly calculating lexicals that have already been processed in the calibration input sequence.

[0055] The generation sequence index of the calibration start identifier follows the local sequence index, and each generated calibration term occupies an increasing generation sequence index according to its generation order. The final-level decoding block uses the original key projection and value projection to process the final-level input hidden state corresponding to the current prediction position, forming the current generated key vector and the current generated value vector; the current generated key vector undergoes RMSNorm processing and rotation position encoding according to the current generation sequence index. The processed current generated key vector and current generated value vector are written to the generation key-value buffer separately configured in the final-level decoding block. The first prediction position writes the key-value pair corresponding to the calibration start identifier into the generation key-value buffer, and other prediction positions sequentially add the key-value pair corresponding to the latest generated calibration term to the generation key-value buffer.

[0056] The final-level decoding block connects the local key-value buffer and the generated key-value buffer at each prediction position to form the complete key sequence and complete value sequence actually read at the current prediction position. The local key-value buffer provides key-value pairs corresponding to the left anchor, reference fragment, and right anchor, while the generated key-value buffer provides the calibration start identifier and key-value pairs corresponding to the already generated calibration terms.

[0057] The final-level decoding block configures the attention read range for the current prediction position. All key-value positions in the local key-value cache are set to readable positions, as are the calibration start flag and calibration terminology key-value positions already written before the current prediction position in the generated key-value cache. During training, when multiple target calibration terms enter the model in parallel, the target terminology key-value positions following the current target calibration terminology are masked to prevent the current prediction position from reading target terminology information that has not yet been generated.

[0058] The 32 query headers are divided into 8 query groups of 4, with each group sharing a single key-value header. Each query vector is multiplied dimension-wise with all key vectors in its corresponding key-value header. The sum of these 128 multiplications is then multiplied by the reciprocal of the square root of 128 to obtain the scaled dot product between the query vector and each key-value position. Exponential normalization is applied to all scaled dot products corresponding to the same query vector to form attention weights. Each attention weight is then multiplied dimension-wise with its corresponding value vector and accumulated along the key-value positions to form the attention result for the current query header.

[0059] The attention results from the 32 query heads are concatenated in query head order to form a 4096-dimensional attention convergence vector. This vector, after output projection, is added dimension-by-dimensionally to the input hidden state of the final-level decoder block at the current prediction position, forming the attention hidden state. After normalization, the attention hidden state is entered into both the gated projection path and the ascending projection path. Both paths map the 4096-dimensional state to a 12288-dimensional intermediate state. The gated projection result, after SiLU activation processing, is multiplied dimension-by-dimensionally with the ascending projection result, and then restored to 4096 dimensions via the descending projection path. The descending projection result is added dimension-by-dimensionally to the attention hidden state, forming the output hidden state of the final-level decoder block at the current prediction position. This output hidden state serves as the calibration hidden state for the current prediction position.

[0060] The output normalization layer processes the calibration hidden state, and the lexical prediction layer maps the normalization result to the prediction score corresponding to each lexical in the lexical table, and forms the lexical probability through exponential normalization. During the inference phase, the lexical with the highest probability is selected as the current calibration lexical. Generation terminates when the current calibration lexical is the end marker; when the current calibration lexical is a text lexical, the text lexical is appended to the calibration start marker and proceeds to the next prediction position. Generation terminates when the generated length reaches the number of reference lexicals plus 32 without obtaining an end marker, and the number of generated lexicals in a single calibration text fragment does not exceed 512. All generated lexicals before the end marker are decoded in the generation order to form the calibration text fragment.

[0061] The model training uses cleaned text, reference text, and manually verified target calibration text as training samples. Each training sample sequentially undergoes difference localization, dual anchor construction, and input encoding to form a calibration input sequence. The target calibration text is encoded using the same lexical segmenter as the calibration input sequence, with an end marker added to the end of the target calibration lexical sequence. Model training employs a teacher-mandated approach. The first target calibration lexical uses the calibration start marker as its preceding input lexical, and the remaining target calibration lexicals use the actual preceding target calibration lexical as their input. Each input lexical sequentially passes through the lexical embedding layer and the 1st to 35th decoding blocks to form the corresponding final-level input hidden state. The final-level decoding blocks use the same local key-value cache and update the generated key-value cache position by position according to the arrangement of the target calibration lexicals. Each prediction position can only read the local key-value cache and the generated key values ​​before the current prediction position. The lexical prediction layer outputs the prediction result corresponding to the current target calibration lexical.

[0062] The left anchor term, difference segment term, right anchor term, reference term, and calibration start identifier in the calibration input sequence are not included in the generation loss calculation, and their corresponding training labels are set to -100. The target calibration term and end identifier are written with the corresponding real term encodings. Cross-entropy loss is only calculated at the target calibration term and end identifier positions. The reference key projection parameter is initialized by copying the original key projection parameter of the final decoding block, and the reference value projection parameter is initialized by copying the original value projection parameter of the final decoding block. The parameters of the first to the 35th decoding blocks remain unchanged, and the training parameters for the reference key projection, reference value projection, final decoding block, and calibration start identifier, as well as the parameters of the output normalization layer and the term prediction layer, are maintained. The RMSNORM scaling parameters in the first to the 35th decoding blocks remain unchanged, while the RMSNORM scaling parameters in the final decoding block are updated along with the final decoding block, and the RMSNORM scaling parameters in the output normalization layer are updated along with the output normalization layer. The lexical embedding parameter matrix is ​​separately assigned to an optimization parameter group with a weight decay coefficient of 0. A gradient mask is set for the parameter rows of the lexical embedding parameter matrix, retaining only the gradient of the parameter row corresponding to the calibration start identifier, and setting the gradient of the parameter rows corresponding to the other lexical terms to 0.

[0063] Training employed the AdamW optimizer with a maximum learning rate of 2 × 10⁻⁵, a batch size of 4 training samples, a gradient accumulation count of 8, and a maximum gradient norm of 1.0. Training consisted of 3 epochs. The total number of optimizer updates across these 3 epochs was used as the total number of optimization steps. Within the first 5% of these steps, the learning rate was linearly increased from 0 to 2 × 10⁻⁵, and then linearly decreased back to 0 within the remaining steps. After each epoch, the average cross-entropy of the target calibration tokens was calculated on the validation samples, and the model parameters corresponding to the lowest average cross-entropy were saved.

[0064] In this embodiment, the text concatenation module specifically includes: Embed the calibration text fragment between the left and right anchor points, and replace the original text character range corresponding to the difference fragment; The candidate calibration text is formed by sequentially concatenating the original text before the double anchor calibration patch, the left anchor point, the calibration text fragment, the right anchor point, and the original text after the double anchor calibration patch according to the original character position order.

[0065] In the specific implementation process, when the text to be cleaned contains multiple discrepancy segments, these segments are arranged in descending order of their starting character positions within the original text character range. Then, starting from the end of the text to be cleaned and moving towards the beginning, calibration text segment generation, text concatenation, and calibration consistency verification are performed sequentially. If the current discrepancy segment passes the calibration consistency verification, it is retained; if it fails, the original text content corresponding to that segment is restored. After processing the discrepancy segments on the right, the original character order corresponding to the unprocessed discrepancy segments on the left remains unchanged.

[0066] Read the original character ranges of the left anchor point, the original character range of the difference segment, and the original character range of the right anchor point recorded in the text to be cleaned, the text fragment to be calibrated, and the dual-anchor calibration patch. Each original character range is recorded in a left-closed, right-open manner. The starting character position corresponds to the first character in the range, and the ending character position corresponds to the position after the last character in the range. When the starting character position and the ending character position are the same, the character range is a zero-length character range.

[0067] Based on the starting character order of the original text character range of the left anchor point, the original text content before the double-anchor calibration patch is extracted from the text to be cleaned; the original text content of the left anchor point is extracted based on the original text character range of the left anchor point; the original text content of the right anchor point is extracted based on the original text character range of the right anchor point; and the remaining text content of the text to be cleaned is extracted starting from the ending character order of the original text character range of the right anchor point, forming the original text content after the double-anchor calibration patch. The above extraction process directly calls the original text character order recorded during the anchor point construction process to avoid positional shifts when the same text appears repeatedly in the text to be cleaned.

[0068] When the original text character range of the difference fragment contains original text characters, the original text content between the starting and ending character positions of the difference fragment is deleted, and the calibration text fragment is written to the deletion position. When the original text character range of the difference fragment is a zero-length character range, the original text characters in the text to be cleaned are not deleted, and the calibration text fragment is inserted between the adjacent original text characters corresponding to the starting character position of the difference fragment. When the calibration text fragment is empty, the writing position remains empty, and the original text content contained in the original text character range of the difference fragment is deleted, thus completing the calibration for the text deletion type.

[0069] When the difference fragment is located at the beginning of a paragraph and the left anchor point is a zero-length original text character range, the original text content before the double-anchor calibration piece is directly connected to the calibration text fragment; when the difference fragment is located at the end of a paragraph and the right anchor point is a zero-length original text character range, the calibration text fragment is directly connected to the original text content after the double-anchor calibration piece. The splicing process does not change the original line breaks, spaces, tabs, and paragraph separators in the text to be cleaned; the character content outside the original text character range of the difference fragment remains unchanged.

[0070] The candidate calibration text is formed by sequentially connecting the original text before the double-anchor calibration patch, the original text of the left anchor point, the calibration text fragment, the original text of the right anchor point, and the original text after the double-anchor calibration patch according to the original character position order. The original text of the left anchor point and the original text of the right anchor point are each extracted only once from the text to be cleaned. The original text content that is replaced within the character range of the difference fragment is no longer retained to avoid duplicate writing of anchor content or difference content.

[0071] The sum of the length of the original text preceding the double-anchor calibration patch and the length of the original text at the left anchor point is used as the starting character position of the calibration text fragment in the candidate calibration text. The sum of the starting character position and the number of characters in the calibration text fragment is used as the ending character position of the calibration text fragment in the candidate calibration text. When the calibration text fragment is empty, the starting character position and the ending character position are the same.

[0072] In this embodiment, the result verification module specifically includes: Perform semantic alignment between the calibration text fragments in the candidate calibration text and the reference fragments, and verify the semantic connection between the calibration text fragments and the left and right anchor points; When the calibration text fragment maintains semantic correspondence with the reference fragment and semantic continuity with the left and right anchor points, the calibration text fragment is retained; When the calibration text fragment does not maintain semantic correspondence with the reference fragment or does not maintain semantic continuity with the left and right anchor points, the calibration text fragment is replaced with the original text content corresponding to the difference fragment, thus forming the text data cleaning and calibration results.

[0073] In the specific implementation process, the following steps are taken: read the candidate calibration text, the starting and ending character positions of the calibration text fragment within the candidate calibration text, the reference fragment, the original text content of the left anchor point, the original text content of the right anchor point, and the original text content corresponding to the difference fragment. Based on the starting and ending character positions of the calibration text fragment, the calibration text fragment to be verified is extracted from the candidate calibration text. The text content of the calibration text fragment is not searched again to avoid verification position shifts when identical characters exist in the candidate calibration text.

[0074] Token encoding is performed on the calibration text segment and the reference segment respectively by using the tokenizer matched with the Qwen3-8B model. When both the calibration text segment and the reference segment are not empty, the two sets of token encodings are respectively input into the token embedding layer and the 1st to 35th decoding blocks to obtain hidden states corresponding to each non-padding token position; the hidden states of the calibration text segment and the reference segment are averaged dimension-wise along the token positions respectively, and vector length normalization is performed on the averaged results to form a calibration semantic vector and a reference semantic vector.

[0075] Calculate the cosine similarity between the calibration semantic vector and the reference semantic vector. When both the calibration text segment and the reference segment are not empty, if the cosine similarity is not lower than 0.85, the semantic correspondence verification is recorded as passed; if the cosine similarity is lower than 0.85, the semantic correspondence verification is recorded as not passed. When both the calibration text segment and the reference segment are empty, the semantic correspondence verification is recorded as passed, the left semantic coherence verification and the right semantic coherence verification are not performed, and both of the two semantic coherence verifications are recorded as passed. When only one of the calibration text segment and the reference segment is empty, the semantic correspondence verification is recorded as not passed, the left semantic coherence verification and the right semantic coherence verification are not performed, and both of the two semantic coherence verifications are recorded as passed.

[0076] When both the calibration text segment and the reference segment are not empty, the semantic coherence relationship between the calibration text segment and the left anchor point, and between the calibration text segment and the right anchor point is checked respectively. When performing semantic coherence verification, the reference key projection, reference value projection and local key-value reconstruction structure are disabled, the last-level decoding block uses the original key projection and value projection to process all token hidden states in the coherence sequence, and performs causal attention calculation according to the token arrangement order. Each token to be verified can only read the token key-values arranged before the token to be verified, and the token prediction layer outputs the prediction probability of the token to be verified based on this. The semantic coherence verification process only performs forward calculation and does not update the parameters of the improved Qwen3-8B model. When the left anchor point is not empty, intercept at most 8 tokens from the end of the left anchor point and at most 8 tokens from the start end of the calibration text segment, and form a left coherence sequence in the order that the tokens of the left anchor point come first and the tokens of the calibration text segment come after. When the right anchor point is not empty, intercept at most 8 tokens from the end of the calibration text segment and at most 8 tokens from the start end of the right anchor point, and form a right coherence sequence in the order that the tokens of the calibration text segment come first and the tokens of the right anchor point come after.

[0077] The left-hand continuation sequence is input into the original causal language prediction path of the Qwen3-8B model. The predicted probabilities of each word at the beginning of the calibration text segment are read, and the average negative log probability of up to four words at the beginning of the calibration text segment is calculated to form the left-hand continuation loss. Similarly, the left-hand anchor terminology and the terminology at the beginning of the reference segment are combined to form the left-hand reference sequence, and the average negative log probability of up to four words at the beginning of the reference segment is calculated to form the left-hand reference loss. The left-hand semantic continuation verification passes if the left-hand continuation loss does not exceed the sum of the left-hand reference loss and 0.30; otherwise, the left-hand semantic continuation verification fails.

[0078] The right-hand continuation sequence is input into the original causal language prediction path of the Qwen3-8B model. The predicted probabilities of each word at the beginning of the right-hand anchor are read, and the average negative log probability of up to four words at the beginning of the right-hand anchor is calculated to form the right-hand continuation loss. Similarly, the right-hand reference sequence is formed by combining the words at the end of the reference segment and the words at the beginning of the right-hand anchor, and the average negative log probability of up to four words at the beginning of the right-hand anchor is calculated to form the right-hand reference loss. The right-hand semantic continuation verification passes if the right-hand continuation loss does not exceed the sum of the right-hand reference loss and 0.30; otherwise, the right-hand semantic continuation verification fails.

[0079] When the left anchor point is a zero-length original text character range, the left semantic continuity check is not performed, and the left semantic continuity check is recorded as passed. When the right anchor point is a zero-length original text character range, the right semantic continuity check is not performed, and the right semantic continuity check is recorded as passed. When both the calibration text fragment and the reference fragment are empty and both the left and right anchor points contain original text characters, the deletion continuity sequence is formed in the order of left anchor point words first and right anchor point words last, and the original text continuity sequence is formed in the order of left anchor point words, difference fragment original text words, and right anchor point words. The deletion continuity sequence and the original text continuity sequence are input into the original causal language prediction path, respectively, and the average negative log probability of the maximum of 4 words at the beginning of the right anchor point in the two sets of sequences is calculated to form the deletion continuity loss and the original text continuity loss. If the loss following deletion does not exceed the sum of the original loss and 0.30, the direct connection verification of the left and right anchor points passes; if the loss following deletion exceeds the sum of the original loss and 0.30, the direct connection verification of the left and right anchor points fails. If both the calibration text segment and the reference segment are empty and the left or right anchor point is a zero-length original character range, the direct connection verification of the left and right anchor points is not performed, and the verification is recorded as passed.

[0080] If both the calibration text fragment and the reference fragment are empty, and the direct connection verification of the left and right anchor points passes, the deletion result is retained. If the calibration text fragment and the reference fragment are not simultaneously empty, and the calibration text fragment maintains semantic correspondence with the reference fragment, and both the left and right semantic connection verifications pass, the calibration text fragment is retained. When the corresponding retention conditions are met, the candidate calibration text is determined as the text data cleaning and calibration result.

[0081] If both the calibration text fragment and the reference fragment are empty and the direct connection verification between the left and right anchor points fails, or if the calibration text fragment and the reference fragment are not simultaneously empty and any one of the semantic correspondence verification, left semantic connection verification, or right semantic connection verification fails, the calibration text fragment is deleted according to the starting and ending character positions of the calibration text fragment in the candidate calibration text, and the original text content corresponding to the difference fragment is written to the deletion position. When the original text character range of the difference fragment is a zero-length character range, no text characters are written after deleting the calibration text fragment; when the original text content of the difference fragment is empty, the deletion position remains empty. After fragment-level recovery is completed, the candidate text content before the calibration text fragment, the original text content corresponding to the difference fragment, and the candidate text content after the calibration text fragment are connected according to the original character position to form the text data cleaning and calibration result.

[0082] Example 1: To verify the feasibility of this invention in practice, it was applied to a scenario of organizing historical documents in an enterprise knowledge base. The text to be cleaned consisted of historical business documents obtained during the knowledge base migration process, while the reference text was a standard document reviewed and published by business personnel. The historical business documents contained issues such as missing text, misspelled terms, duplicate content, and inconsistencies between some sentences and the standard document. Therefore, it was necessary to perform local calibration on the discrepancies while preserving the paragraph structure and non-discrepancy content of the historical documents.

[0083] This embodiment collected 12,000 paragraph-level text pairs, each pair consisting of a text to be cleaned and a reference text, totaling 4.86 million characters. 18,742 discrepancies were manually identified, of which content replacement differences accounted for 58.6%, redundant content in the text to be cleaned accounted for 21.3%, and missing content in the text to be cleaned accounted for 20.1%. The training, validation, and test sets were divided in an 8:1:1 ratio. The training set contained 9,600 text pairs, while the validation and test sets each contained 1,200 text pairs. Each discrepancy in the training set was paired with a manually verified target calibration text segment.

[0084] Number the text to be cleaned and the reference text consecutively according to Unicode code points, and save paragraph boundaries, line breaks and the original character order corresponding to each character. Split sentence units according to paragraph boundaries, input the sentence units into the token embedding layer of the Qwen3-8B model and from the 1st decoding block to the 35th decoding block, average the hidden states corresponding to each non-padding token position dimension by dimension and perform vector length normalization to form sentence semantic vectors. Calculate the cosine similarity between the sentence semantic vectors of the text to be cleaned and the reference text, set sentence units whose cosine similarity is not less than 0.78 and whose arrangement direction is consistent as sentence corresponding pairs, and merge unmatched sentence units between adjacent sentence corresponding pairs into corresponding text segments to form semantically corresponding text segments.

[0085] Process the semantically corresponding text segments with the fast tokenizer supporting the Qwen3-8B model, and align the tokens of the text to be cleaned and the tokens of the reference text by using the minimum edit distance algorithm. Mark positions corresponding to substitution operations, deletion operations and insertion operations as difference positions, and merge consecutive difference positions into difference intervals. Perform local alignment again on the difference interval and two consistent tokens on each side of the difference interval, delete the re-aligned consistent tokens at both ends of the difference interval to form difference fragments, and record the original character interval corresponding to the difference fragments.

[0086] Intercept consecutively consistent tokens of the text to be cleaned from the left side and the right side of the difference fragment respectively, with the number of interceptions on each side not exceeding 8 tokens, to form a left anchor and a right anchor respectively. When the difference fragment is located at a paragraph boundary, configure a zero-length anchor at the corresponding paragraph boundary. Assemble the left anchor, the difference fragment and the right anchor according to the original character order to form a dual-anchor calibration slice.

[0087] Intercept a reference fragment sharing the same alignment position range with the difference fragment from the reference text. Process the left anchor, the difference fragment, the right anchor and the reference fragment respectively with the fast tokenizer supporting the Qwen3-8B model, configure segment identifiers 0, 1, 2 and 3 for the four types of tokens in sequence, and connect a calibration start identifier with segment identifier 4 at the end of the reference tokens. Form a calibration input sequence in the order of the dual-anchor token sequence, the reference token sequence and the calibration start identifier.

[0088] Improve the Qwen3-8B model by adding reference key projection and reference value projection beside the original key projection and value projection of the 36th decoding block, and setting a local key-value reconstruction structure. The hidden states corresponding to the left anchor, the difference fragment and the right anchor pass through the original key projection and value projection to form original text key values, and the hidden states of reference tokens pass through the reference key projection and reference value projection to form reference key values.

[0089] The local key-value reconstruction structure removes the original text keys corresponding to the differencing segments, retains the original text keys corresponding to the left and right anchor points, and arranges the reference keys according to their position indices. When the same position contains multiple reference terms, they are arranged according to the character order of the reference terms in the reference segments. The arranged reference keys are then inserted between the original text keys of the left and right anchor points to form a local key-value sequence, which is written to the local key-value cache of the final decoding block.

[0090] The query vector for the first prediction position is formed using the latent state of the final-level input corresponding to the calibration start identifier. For other prediction positions, the query vector is formed using the latent state of the latest generated calibration term. The key values ​​corresponding to the calibration start identifier and the generated calibration terms are sequentially written to the generated key value cache. The final-level decoding block connects the local key value cache and the generated key value cache at each prediction position, performing grouped query attention processing, feedforward processing, output normalization, and term prediction. Generation terminates when the number of generated terms reaches the number of reference terms plus 32, with a maximum of 512 generated terms.

[0091] Model training employs a teacher-mandated approach, with the training label corresponding to the input position set to -100, and the actual word encodings written at the positions corresponding to the target calibration word and the end word. The reference key projection parameters and reference value projection parameters are copied from the original key projection parameters and value projection parameters of the 36th decoding block as initial values. The parameters for the 1st to 35th decoding blocks remain unchanged. Training is performed on the reference key projection, reference value projection, the embedding parameters of the word corresponding to the 36th decoding block and the calibration start identifier, as well as the parameters of the output normalization layer and the word prediction layer. Training uses the AdamW optimizer with a maximum learning rate of 2×10^-5, a batch size of 4 samples, 8 gradient accumulation iterations, and 3 training epochs.

[0092] After the model outputs calibration text fragments, text replacement is performed according to the original character range of the difference fragments. When there are multiple difference fragments in the text to be cleaned, they are arranged in descending order of the starting character position of the difference fragments, and calibration text fragment generation, text splicing, and calibration consistency verification are performed sequentially from the end of the text to the beginning.

[0093] The calibration consistency verification includes reference semantic verification and left-right boundary continuation verification. The latent states of the lexical units in the calibration text segment and the reference segment are averaged dimension-wise and vector length normalized. The cosine similarity between the two segments is calculated; a cosine similarity of at least 0.85 indicates that the semantic verification is passed. The average negative log probability of the continuation sequence formed by the calibration text segment and its left and right anchors, and the average negative log probability of the reference sequence formed by the reference segment and its left and right anchors, are calculated separately. The boundary continuation verification is passed when the continuation loss does not exceed the sum of the reference loss and 0.30. If all verifications pass, the calibration text segment is retained; if any verification fails, the original text content corresponding to the differing segment is restored.

[0094] To verify the processing effect of this embodiment, the rule difference replacement method, the original Qwen3-8B whole segment generation method and the method of this system were compared on the same test set. The results are shown in the table below.

[0095] Table 1. Test Results of Different Text Calibration Methods

[0096] The calibration fragment complete accuracy rate represents the proportion of discrepancies where the output calibration text fragment is completely identical to the manually verified result in character content, out of the total number of discrepancies tested. The boundary acceptance pass rate represents the proportion where the acceptance loss between the calibration text fragment and both the left and right anchor points meets the judgment criteria. The manual correction per thousand fragments represents the number of fragments that still require manual adjustment after processing 1000 discrepancies.

[0097] As shown in Table 1, the calibration segment accuracy of this system's method reaches 91.5%, an improvement of 7.7 percentage points compared to the original Qwen3-8B method; the boundary acceptance rate reaches 97.0%, and the number of manual corrections per thousand discrepancies is reduced to 79. After replacing the original text key values ​​of the discrepancies with reference key values, the model can generate calibration text segments while preserving the left and right anchor context, reducing the direct impact of erroneous original text on the final attention calculation, and reducing the semantic break between the calibration text segments and adjacent original text.

[0098] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.

Claims

1. A text data cleaning and calibration system based on a large language model, characterized in that, include: The text acquisition module is used to acquire the text to be cleaned and the reference text, preserving the paragraph boundaries and original character order of the text to be cleaned; The difference localization module is used to perform semantic alignment between the text to be cleaned and the reference text, merge continuous difference positions and shrink difference boundaries to form difference fragments; The anchor point construction module is used to extract adjacent content that corresponds to the semantics of the reference text from both sides of the difference fragment in the text to be cleaned, set left anchor points and right anchor points respectively, and assemble the left anchor points, difference fragments and right anchor points according to the original character position order to form a double anchor calibration patch. The input encoding module is used to extract the reference segment corresponding to the difference segment from the reference text, perform partition encoding on the dual-anchor calibration piece and the reference segment, and form a calibration input sequence; The model calibration module is used to input the calibration input sequence into the improved Qwen3-8B model. The improved Qwen3-8B model adds a reference key-value branch and a local key-value reconstruction structure in the key-value projection path of grouped query attention to generate calibration text fragments. The text splicing module is used to embed the calibration text fragment between the left anchor point and the right anchor point, replace the original text character range corresponding to the difference fragment, and splice the original text content other than the double anchor calibration fragment according to the original character position order to form the candidate calibration text. The result verification module is used to perform calibration consistency verification on calibration text fragments in candidate calibration texts, and generate text data cleaning and calibration results.

2. The text data cleaning and calibration system based on a large language model according to claim 1, characterized in that, The difference localization module specifically comprises: The semantically corresponding text segments are identified from the text to be cleaned and the reference text, and word matching is performed on the semantically corresponding text segments to form a matching word sequence. In the context of texts where the reference text contains words but the text to be cleaned is missing words, placeholders are configured for missing words. Placeholders for words with different word content, words present on one side of the text to be cleaned, or words with missing placeholders are marked as difference positions. The adjacent differences are merged along the original character order of the text to be cleaned, and the words at both ends of the merged interval that are re-aligned with the reference text are excluded to form the difference fragments.

3. The text data cleaning and calibration system based on a large language model according to claim 1, characterized in that, The anchor point construction module is specifically as follows: Starting from the beginning of the difference segment and moving left and right respectively, check the corresponding word sequence one by one. Extract the words in the text to be cleaned that are adjacent to the difference segment and maintain semantic correspondence as left anchor point and right anchor point respectively. Arrange the left anchor point, difference fragment, and right anchor point according to the original character position of the text to be cleaned, and record the original character range corresponding to the left anchor point and the right anchor point to form a double anchor calibration patch.

4. The text data cleaning and calibration system based on a large language model according to claim 1, characterized in that, The input encoding module is specifically: The reference fragment is formed by extracting the text range corresponding to the difference fragment from the semantically corresponding text segment of the reference text; Lexicalization is performed on the dual-anchor calibration piece and the reference fragment respectively to form a dual-anchor lexical sequence and a reference lexical sequence; Configure segment identifiers to distinguish the left anchor word, the difference segment word, the right anchor word, and the reference word. Associate the double anchor word sequence with the original position index and the reference word sequence with the corresponding position index. Connect the calibration start identifier at the end of the reference word sequence and arrange them in the order of double anchor word sequence, reference word sequence, and calibration start identifier to form the calibration input sequence.

5. A text data cleaning and calibration system based on a large language model according to claim 4, characterized in that, The lexicalization process performed on the dual-anchor calibration patch and the reference fragment is as follows: The word segmentation rules of the Qwen3-8B model are used to perform word segmentation on the left anchor point, the difference segment, the right anchor point and the reference segment respectively; The left anchor word, the difference fragment word, and the right anchor word are respectively bound to the corresponding original character range in the text to be cleaned, and arranged according to the original character position to form a double anchor word sequence; Arrange the reference words according to the character order of the reference segment, and record the position index between each reference word and the corresponding position of the difference segment to form a reference word sequence.

6. A text data cleaning and calibration system based on a large language model according to claim 5, characterized in that, The improved Qwen3-8B model includes a lexical embedding layer, multiple sequentially connected decoding blocks, an output normalization layer, and a lexical prediction layer. The decoding block includes a group query attention sublayer and a feedforward network sublayer. The input ends of the group query attention sublayer and the feedforward network sublayer are configured with normalization processing, and the corresponding input hidden states are passed through the residual path respectively. In a series of sequentially connected decoding blocks, the final-level decoding block connected to the output normalization layer is set as the modified decoding block. The final-level decoding block adds reference key projection and reference value projection on the side of the key projection and value projection of the grouped query attention sub-layer, and sets a local key-value reconstruction structure between the original key-value projection result, the reference key-value projection result and the attention operation. The lexical embedding layer maps each lexical in the calibration input sequence to an input hidden state, forming an input hidden state sequence. The input hidden state sequence is passed layer by layer to the final decoding block along the connection order of the decoding blocks. The reference key projection, reference value projection, and local key-value reconstruction structure process the input hidden state of the final-level decoding block to form a local key-value sequence; In the word-by-word generation process of the calibration text fragment, the final decoding block uses local key-value sequences to complete group query attention processing and feedforward processing, forming the calibration hidden state of the current prediction position; The output normalization layer and the lexical prediction layer process the calibration hidden state at the current prediction position in turn, predict the current calibration lexical, and then continue the current calibration lexical into the calibration input sequence and enter the next prediction position until the end marker is predicted. The calibration text fragment is then restored according to the lexical generation order.

7. A text data cleaning and calibration system based on a large language model according to claim 6, characterized in that, The reference key projection, reference value projection, and local key-value reconstruction structure process the input hidden state of the final-level decoding block to form a local key-value sequence, specifically as follows: According to the segment identifier, extract the left anchor point hidden state, the difference segment hidden state, the right anchor point hidden state, and the reference word hidden state from the input hidden state sequence of the final decoding block respectively. The original key projection and value projection of the grouped query attention sublayer are used to process the hidden states of left anchor points, difference fragments, and right anchor points, forming the original text key sequence and original text value sequence; the reference key projection and reference value projection are used to process the hidden states of reference terms, forming the reference key sequence and reference value sequence. Remove the original key subsequences and original value subsequences corresponding to the difference segments from the original key sequence and original value sequence, and retain the original key sequence and original value subsequences corresponding to the left anchor and the right anchor; When a reference segment contains reference terms, the reference key sequence and reference value sequence are arranged according to their position indices. When the same position index corresponds to multiple reference terms, they are arranged according to the character order of the multiple reference terms in the reference segment. The arranged reference key sequence is then connected between the left anchor text key sequence and the right anchor text key sequence, and the arranged reference value sequence is connected between the left anchor text value sequence and the right anchor text value sequence. When a reference segment does not contain reference terms, the left anchor text key sequence and the right anchor text key sequence are connected, and the left anchor text value sequence and the right anchor text value sequence are also connected. Configure the local position index according to the arrangement order of the left anchor point, reference fragment and right anchor point, perform RMSNorm processing on each key vector in the local key sequence along the key value head state dimension, perform rotation position encoding according to the local position index, and pair the processed local key sequence with the local value sequence to form a local key value sequence.

8. A text data cleaning and calibration system based on a large language model according to claim 6, characterized in that, The final-level decoding block uses local key-value sequences to complete group query attention processing and feedforward processing, forming the calibration hidden state of the current prediction position, specifically: The first prediction position obtains the input hidden state of the final-level decoder block corresponding to the calibration start identifier, and other prediction positions obtain the input hidden state of the final-level decoder block corresponding to the latest generated calibration term. The obtained input hidden states of the final-level decoder block are processed by pre-RMSNorm and query projection to form query vectors corresponding to multiple query heads. To calibrate the starting identifier and each generated calibration term, a generated position index is sequentially configured to follow the local position index. RMSNorm processing is performed on each query vector along the query header state dimension, and rotation position encoding is performed according to the generated position index corresponding to the current predicted position. Obtain the calibration start identifier and the input hidden state of the final-level decoding block corresponding to the calibration tokens generated before the current prediction position. Use key projection and value projection to form the generated key sequence and generated value sequence. Perform RMSNorm processing on each key vector in the generated key sequence along the key value head state dimension, and perform rotation position encoding according to the corresponding generated position index. Pair the processed generated key sequence with the generated value sequence and continue to the local key value sequence. Based on the grouping relationship between the query header and the key value header, multiple query headers are assigned to corresponding query groups, and query vectors in the same query group share the key sequence and value sequence of the corresponding key value header; Calculate the dot product between each query vector and each key vector in the corresponding key sequence. Multiply the dot product result by the reciprocal of the square root of the query header dimension to form a scaled dot product. Perform exponential normalization on all scaled dot products corresponding to the same query vector to form the attention weights corresponding to each key value position. Each attention weight is multiplied with its corresponding value vector dimension by dimension. The multiplication results corresponding to the same query vector are accumulated position by position to form the attention results of each query head. The attention results of multiple query heads are concatenated and output projection is performed. The output projection result is added in the same position to the input hidden state of the final-level decoded block corresponding to the current prediction position to form the attention hidden state of the current prediction position. Normalize the attention hidden state, and input the normalization result into the gated projection path and the rising projection path of the feedforward network sub-layer respectively. Use the SiLU activation function to process the gated projection result, multiply the activated gated projection result with the rising projection result dimension by dimension, and restore it to the input hidden state dimension of the final decoding block through the falling projection path. The descent projection result is added in the same position to the attention hidden state to form the output hidden state of the final-level decoder block at the current prediction position, and the output hidden state of the final-level decoder block at the current prediction position is determined as the calibration hidden state.

9. A text data cleaning and calibration system based on a large language model according to claim 1, characterized in that, The text concatenation module is specifically as follows: Embed the calibration text fragment between the left and right anchor points, and replace the original text character range corresponding to the difference fragment; The candidate calibration text is formed by sequentially concatenating the original text before the double anchor calibration patch, the left anchor point, the calibration text fragment, the right anchor point, and the original text after the double anchor calibration patch according to the original character position order.

10. A text data cleaning and calibration system based on a large language model according to claim 1, characterized in that, The result verification module is specifically as follows: Perform semantic alignment between the calibration text fragments in the candidate calibration text and the reference fragments, and verify the semantic connection between the calibration text fragments and the left and right anchor points; When the calibration text fragment maintains semantic correspondence with the reference fragment and semantic continuity with the left and right anchor points, the calibration text fragment is retained; When the calibration text fragment does not maintain semantic correspondence with the reference fragment or does not maintain semantic continuity with the left and right anchor points, the calibration text fragment is replaced with the original text content corresponding to the difference fragment, thus forming the text data cleaning and calibration results.