Chinese spelling error correction method based on multi-modal enhancement
Through the multimodal enhanced Chinese spelling error correction method, combined with semantic, auditory and visual features, the Transformer encoder and MogaNet are used to fusion of feature, and a self-reflection mechanism is introduced, which solves the problem of insufficient integration of information in the existing technology and achieves efficient and accurate Chinese spelling error correction.
Patent Information
- Application Number
- CN202511082138.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-04
- Publication Date
- 2025-09-02
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing Chinese spelling error correction technology is difficult to effectively integrate speech and image information, and the performance improvement is limited by the scale of pre-trained models.
The multimodal enhancement method is adopted to extract feature vectors through semantic, auditory and visual perception modules, combine large language models for error correction, use Transformer encoder and MogaNet for feature fusion, and introduce a self-reflection mechanism to optimize the error correction process.
It significantly improves the accuracy and efficiency of Chinese spelling error correction, reduces training costs, avoids unconstrained polishing, and improves the error correction performance of the model.
Smart Images

Figure CN120579543A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of natural language processing, and in particular relates to a Chinese spelling correction method based on multimodal enhancement. Background Art
[0002] Chinese spelling correction technology is based on a fusion of linguistic rules, statistical models, and deep learning. Early models relied on manual rules and statistical methods to address common misspellings and lexical errors. With technological advancements, deep learning has become a core technology. Language models pre-trained on large-scale corpora significantly improve correction accuracy through their powerful contextual semantic understanding capabilities. Furthermore, to address the unique challenges of homophones and similar characters in Chinese, existing models integrate multimodal features such as glyph shape and pronunciation, and integrate with external knowledge bases to enhance adaptability to complex semantics and domain transfer.
[0003] Existing Chinese spelling correction technologies, such as the BERT model, struggle to integrate speech and image information when handling complex homophones and similar characters. Furthermore, model training tends to rely on character alignment rather than semantic information. Furthermore, many existing Chinese spelling correction methods are based on pre-trained models, and the scale of these pre-trained models limits the improvement of correction capabilities. Summary of the Invention
[0004] The technical problem to be solved by the present invention is to overcome the defects of the existing Chinese spelling correction technology in the art, that is, it is difficult to effectively integrate voice and image information and the performance improvement is limited by the scale of the pre-trained model, thereby providing a Chinese spelling correction method based on multimodal enhancement.
[0005] The present invention discloses a Chinese spelling error correction method based on multimodal enhancement, comprising the following steps: inputting a text sequence to be corrected into a semantic perception module to obtain a semantic feature vector; the semantic perception module maps the text to be corrected into a character embedding code, and feeds the character embedding code and the corresponding position code into a stacked multi-layer Transformer encoder to output a semantic feature vector; inputting the text sequence to be corrected into an auditory perception module to obtain an auditory feature vector; the auditory perception module converts the text sequence to be corrected into a Chinese pinyin sequence, concatenates and maps the text to be corrected and the Chinese pinyin sequence to form a character embedding code, and outputs the character embedding code and the corresponding position code into a stacked multi-layer Transformer encoder to output a semantic feature vector; The encoding is sent to a stacked multi-layer Transformer encoder to output an auditory feature vector; the text sequence to be corrected is input into a visual perception module to obtain a visual feature vector; the visual feature vector forms a character image of the text sequence to be corrected through font rendering, and is input into MogaNet to output a visual feature vector; the semantic feature vector, the auditory feature vector and the visual feature vector are fused and input into an error detection network to obtain a detection label, and the error detection label marks the error position in the text sequence to be corrected; the text sequence to be corrected and the error detection label are input into a revision model to obtain a corrected text sequence; the revision model is a large language model.
[0006] Furthermore, the semantic perception module includes a BERT pre-training model, which obtains the text sequence to be corrected and outputs a semantic feature vector.
[0007] Furthermore, each pinyin in the pinyin sequence of the auditory perception module includes an initial consonant and a final vowel.
[0008] Furthermore, the visual perception module includes a MogaNet, which inputs a character image and outputs a visual feature vector; the MogaNet includes a coding layer, a Moga block and a linear layer, each of the Moga blocks includes a spatial aggregation layer and a channel aggregation layer connected in sequence; the coding layer is used to adjust the resolution of the character image and embed it into a specific dimension; the spatial aggregation layer is used to fuse spatial context information of different scales, and the channel aggregation layer is used to adaptively redistribute channel weights; the linear layer is used to receive the feature vector output by the Moga block and map it to a preset category.
[0009] Furthermore, in the spatial aggregation layer, a 1x1 convolution and a scaling factor To reweight the complementary interaction part, it is expressed as: ; ; in represents the feature vector of the original input of the spatial aggregation layer, GAP(·) represents the global average pooling, is the scaling factor, GELU (·) represents the activation function, Represents the output of the spatial aggregation layer.
[0010] Furthermore, in the channel aggregation layer, it is expressed as: ; ; ; Among them, The feature vector representing the original input of the channel aggregation layer, represents normalization, DW (·) represents depth convolution, CA (·) represents channel aggregation, Represents the scaling factor, downscaling the projection by channels and GELU activation function to collect and redistribute information of channel dimensions, Represents the output of the channel aggregation layer.
[0011] Furthermore, the training method of the auditory perception module includes: dividing the original expected text into sentences and filtering sentences whose length is outside a preset range; randomly sampling the filtered sentences to obtain the original text sequence; replacing some characters with words with similar pronunciations, randomly replacing some characters, and keeping some characters unchanged to obtain a replacement text sequence; obtaining the initial consonants and finals of the replacement text sequence and splicing them with the replacement text sequence, using the original text sequence as a label to obtain a training data set; replacing the unused tokens in the BERT model with the obtained initial consonants and finals, inputting the training data set into the BERT model to obtain the characters and corresponding pinyin embedding representations; obtaining the restored text sequence based on the character and pinyin embedding representations, comparing it with the label to calculate the loss and then updating the model parameters; training until the model converges.
[0012] Furthermore, the error detection module combines the semantic feature vector, the auditory feature vector, and the visual feature vector through a fusion function, inputs them into a classifier to generate an error probability, and determines the error detection label through an activation function, which is expressed as: ; ; ; in, represents the semantic feature vector, represents the auditory feature vector, represents the visual feature vector, represents the trainable parameters, Indicates bias, represents the softmax activation function, represents the error probability distribution of the corresponding position, Indicates the error detection tag.
[0013] Furthermore, the Chinese spelling error correction method also includes the following steps: inputting the text sequence containing auditory errors, visual errors and random errors and their corresponding corrected text sequences into the large language model respectively, and inputting the prompt words describing the input content into the large language model; inputting the text sequence containing no errors into the large language model, and inputting the prompt words describing the input content into the large language model; inputting the text sequence to be corrected and the error detection mark into the large language model to obtain the corrected text sequence.
[0014] Furthermore, the text sequence to be corrected and the corresponding correction text sequence are input into the large language model, and the prompt word question is asked whether the semantic information is effectively utilized. When the large language model responds affirmatively, the correction text sequence is determined as the final correction text sequence.
[0015] Beneficial effects: The present invention discloses a Chinese spelling correction method based on multimodal enhancement and contextual learning. By constructing semantic, visual, and auditory perception modules, it can extract semantic, visual, and auditory perception feature vectors in the text to be corrected, realize deep fusion of multimodal information error detection, and significantly improve the error detection effect. After obtaining the error detection label through the error detection network, the error correction task is converted into a fill-in-the-blank task, and the powerful semantic understanding ability of the large language model is used for error correction. In combination with contextual learning, task-related examples are provided to the large model, so that it can quickly adapt to the Chinese spelling correction task and avoid unconstrained polishing. At the same time, a self-reflection mechanism is proposed, so that the large model can reflect on the results after correction many times, thereby improving the performance of the model correction. At the same time, the present invention is based on the training and construction of pre-trained models and large language models, which avoids the use of large-scale corpus for complex training and reduces the application cost. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0017] Figure 1 It is a schematic block diagram of the method flow of the present invention; Figure 2 Schematic diagram of the algorithm flow of the present invention; Figure 3 This is a flow chart of the visual perception module of the present invention; Figure 4 This is a schematic diagram of the spatial aggregation layer flow of the visual perception module of the present invention; Figure 5 Schematic diagram of the channel aggregation layer process of the visual perception module of the present invention; DETAILED DESCRIPTION
[0018] To make the above-mentioned objects, features, and advantages of the present application more clearly understood, the specific embodiments of the present application are described in detail below with reference to the accompanying drawings. The following description sets forth many specific details to facilitate a full understanding of the present application. However, the present application can be implemented in many other ways than those described herein, and those skilled in the art can make similar improvements without violating the scope of the present application. Therefore, the present application is not limited to the specific embodiments disclosed below.
[0019] Reference Figure 1 and Figure 2 As shown, this embodiment provides a Chinese spelling correction method based on multimodal enhancement, comprising the following steps: Step S1.1: Input the text sequence to be corrected into a semantic perception module to obtain a semantic feature vector; the semantic perception module maps the text to be corrected into a character embedding code, and feeds the character embedding code and the corresponding position code into a stacked multi-layer Transformer encoder to output a semantic feature vector; Step S1.2: Input the text sequence to be corrected into the auditory perception module to obtain an auditory feature vector; the auditory perception module converts the text sequence to be corrected into a Chinese phonetic sequence, concatenates and maps the text to be corrected and the Chinese phonetic sequence to form a character embedding code, and feeds the character embedding code and the corresponding position code into a stacked multi-layer Transformer encoder to output an auditory feature vector; Step S1.3: Input the text sequence to be corrected into the visual perception module to obtain a visual feature vector; the visual feature vector is converted into a character image by font rendering of the text sequence to be corrected, and is input into MogaNet to output a visual feature vector; Step S2: fusing the semantic feature vector, the auditory feature vector, and the visual feature vector and inputting the result into an error detection network to obtain a detection label, wherein the error detection label marks the error position in the text sequence to be corrected; Step S3: Input the text sequence to be corrected and the error detection mark into the revision model to obtain the error-corrected text sequence; the revision model is a large language model.
[0020] Specifically, for a given task training sample set ,in Represents the input text, which is a sequence of text that may contain spelling errors. represents the corresponding target output, that is, the corrected text sequence, is the total number of training sample sets. In this embodiment, the input text Indicated by A text sequence consisting of characters , target output Can be expressed as characters of the corrected output text The modeling goal of the Chinese spelling correction task is to build an automated analysis and prediction function , to achieve accurate correction of the input target text.
[0021] In step S1.1, the semantic perception module includes a BERT pre-training model, which obtains the text sequence to be corrected and outputs a semantic feature vector.
[0022] For a given text sequence , the semantic feature vector encoded by the BERT pre-training model is represented as ,The model uses a multi-layer bidirectional self-attention mechanism to encode contextual information.
[0023] ; in, Indicates character embedding encoding, Represents positional encoding, i and j represent the elements in the i-th row and j-th column of the matrix respectively.
[0024] ; Here, i represents the position index in the sequence, and j represents the dimension index of the encoding vector, which is used to control the generation of codes of different dimensions. This method generates a unique code for each position in the sequence. By alternating the sine and cosine functions, the codes are made continuous and periodic, which helps the model capture the order and relative position information of the sequence.
[0025] In the step S1.2, each pinyin in the pinyin sequence of the auditory perception module includes an initial consonant and a final vowel.
[0026] The training method of the auditory perception module includes: Step S1.2.1: Segment the original expected text into sentences and filter out sentences whose lengths are outside a preset range; Step S1.2.2: Randomly sample the filtered sentences to obtain the original text sequence; Step S1.2.3: Replace some characters with words that have similar pronunciations, randomly replace some characters, and keep some characters unchanged to obtain a replacement text sequence; Step S1.2.4: Obtain the initials and finals of the replacement text sequence and concatenate them with the replacement text sequence, using the original text sequence as a label to obtain a training dataset; Step S1.2.5: Replace the unused tokens in the BERT model with the obtained initials and finals, and input the training dataset into the BERT model to obtain the character and corresponding pinyin embedding representation; Step S1.2.6: Based on the character and pinyin embedding representation, the recovered text sequence is obtained, compared with the label and the loss is calculated, and the model parameters are updated; training is performed until the model converges.
[0027] Reference Figure 4~Figure 5 As shown, in the step S1.3, the visual perception module includes a MogaNet, which inputs a character image and outputs a visual feature vector; the MogaNet includes a coding layer, a Moga block and a linear layer, and each of the Moga blocks includes a spatial aggregation layer and a channel aggregation layer connected in sequence; the coding layer is used to adjust the resolution of the character image and embed it into a specific dimension; the spatial aggregation layer is used to fuse spatial context information of different scales, and the channel aggregation layer is used to adaptively redistribute channel weights; the linear layer is used to receive the feature vector output by the Moga block and map it to a preset category. Its task is to make a final classification decision based on the learned features.
[0028] Specifically, in the spatial aggregation layer, a 1x1 convolution and a scaling factor To reweight the complementary interaction part, it is expressed as: ; ; in represents the feature vector of the original input of the spatial aggregation layer, GAP(·) represents the global average pooling, is the scaling factor, GELU (·) represents the activation function, Represents the output of the spatial aggregation layer.
[0029] Furthermore, in the channel aggregation layer, it is expressed as: ; ; ; Among them, The feature vector representing the original input of the channel aggregation layer, represents normalization, DW (·) represents depth convolution, CA (·) represents channel aggregation, Represents the scaling factor, downscaling the projection by channels and GELU activation function to collect and redistribute information of channel dimensions, Represents the output of the channel aggregation layer.
[0030] In step S2, the error detection module combines the semantic feature vector, the auditory feature vector, and the visual feature vector through a fusion function, inputs them into a classifier to generate an error probability, and determines the error detection label through an activation function, which is expressed as: ; ; ; in, represents the semantic feature vector, represents the auditory feature vector, represents the visual feature vector, represents the trainable parameters, Indicates bias, represents the softmax activation function, represents the error probability distribution of the corresponding position, The error detection module loads the pre-trained weights of the semantic perception module, the auditory perception module, and the visual perception module. The classifier adds a fully connected layer after the final feature vector, connecting it to each category, and finally applying a softmax function to obtain the final probability distribution.
[0031] In step S3, the Chinese spelling error correction method further includes the following steps: inputting a text sequence containing auditory errors, visual errors, and random errors, and their corresponding corrected text sequences, into a large language model, and simultaneously inputting a prompt word describing the input content into the large language model; inputting a text sequence containing no errors into the large language model, and simultaneously inputting a prompt word describing the input content into the large language model; and inputting the text sequence to be corrected and an error detection tag into the large language model to obtain a corrected text sequence. The semantic information is the semantics of the input text sequence to be corrected.
[0032] In this embodiment, one way to construct the prompt word is as follows: Taking "food safety and low price have nothing to do with each other" as an example, the correct expression is "food safety has nothing to do with low price", and the corresponding error detection module's error detection label should be "00001001000000". Converting it into a mask representation should result in "food safety [ unknow ]Low price[ unknow ] has nothing to do with it.”
[0033] Error correction prompt word design: - Role: Natural Language Processing Expert and Chinese Spelling Correction Engineer Background: A user needs to perform spelling correction on a text. The model has detected possible error locations and converted them into masked representations. The user hopes to leverage the semantic understanding capabilities of the large model, combining the glyph and phonetic information of the characters in the error location, to correct the error. The user provides four pairs of input sentences and their corrected versions as examples. These sentences contain visual errors, auditory errors, random errors, and correct sentences.
[0034] - Profile: You are an expert with many years of experience in natural language processing, with extensive research and practical experience in text error correction, semantic understanding, and word shape and phonetic analysis. You are adept at efficiently addressing text spelling errors by leveraging advanced language models and algorithms, combined with human intuition.
[0035] Constraints: Correction results should conform to Chinese grammatical and semantic standards, maintaining the overall semantic coherence of the original text. Furthermore, the correction process should be efficient and accurate, avoiding the introduction of new errors. The output format should be concise and clear, containing only complete corrected sentences.
[0036] - OutputFormat: Only output the corrected complete sentence without other information.
[0037] - Workflow: 1. Receive the original input text and the masked representation of the text.
[0038] 2. Analyze the contextual semantic information of the mask position and combine the glyph and phonetic features to determine the possible correct character.
[0039] 3. Correct each mask position and generate the corrected text.
[0040] - Examples: - Example 1 (auditory error): Original input: He studies hard every day, hoping to get good grades Mask said: He studies hard every day and hopes to get good grades Output: He studies hard every day, hoping to get good grades - Example 2 (visual error): Original input: I like apples and bananas Mask said: I like eating apples and fragrance [unknow] Output: I like to eat apples and bananas - Example 3 (random error): Original input: The content of this book is very exciting, I recommend it to every friend Mask said: The content of this book is very exciting, I recommend it to every [unknow] friend Output: The content of this book is very exciting, I recommend it to every friend - Example 4 (correct sentence): Original input: Spring is here, everything is revived Mask means: Spring is coming, everything is revived Output: Spring is here, everything is revived The Chinese spelling correction method also includes a self-reflection mechanism: the text sequence to be corrected and the corresponding correction text sequence are input into the large language model, and the prompt word asks whether the semantic information is effectively utilized. When the large language model responds affirmatively, the correction text sequence is determined as the final correction text sequence.
[0041] Specifically, the large language model is asked to reflect on a question in the prompt word: whether the added rich semantic information was effectively utilized in the error correction process. Only when the answer to this question is "true" is the answer output as the final error correction result. Otherwise, the current conversation is added to the historical conversation and provided to the model as context. If the semantic information is not utilized in the conversation, a reply request is sent to the model again. In this method's self-reflection mechanism, we set the maximum number of loops to 3. If no answer is obtained after requesting the model three times, the model is judged to be unable to correct the sentence, and the original input sentence without introspection is used as the answer.
[0042] In this embodiment, a prompt word for the self-reflection mechanism is expressed as: - Role: Chinese spelling correction self-reflection mechanism engineer Background: A user needs to perform spelling correction on a text and wants to implement a self-reflection mechanism to ensure that semantic information is fully utilized during the correction process. By re-inputting the corrected sentence and the original input sentence into the model, the model needs to reflect on whether it effectively utilized semantic information for correction. If the answer is "true", the final correction result is output; otherwise, the current conversation is added to the historical context and the model is requested to correct the error again, with a maximum number of cycles of three.
[0043] - Profile: You are an expert with extensive experience in natural language processing and model optimization. You excel at designing and implementing self-reflection mechanisms to improve the accuracy of model correction and the efficiency of semantic utilization. You can ensure that models fully utilize contextual and semantic information during error correction through logical analysis and algorithm optimization.
[0044] Constraints: Correction results must conform to Chinese grammatical and semantic standards and maintain the overall semantic coherence of the original text. The maximum number of self-reflection cycles is 3. If no valid result is obtained after 3 attempts, the original input sentence without introspection is used as the answer.
[0045] - OutputFormat: Outputs the final error correction result. If the self-reflection mechanism fails, the original input sentence is output.
[0046] - Workflow: 1. Receive the original input text, the corrected text, and the number of iterations.
[0047] 2. The model determines whether semantic information has been effectively used for error correction. If the answer is "true," the model outputs the corrected text. Otherwise, the model adds the current conversation to the historical context and requests correction again.
[0048] 3. If no valid result is obtained after 3 cycles, the original input sentence is output.
[0049] This embodiment provides a Chinese spelling error correction method based on multimodal enhancement. By constructing semantic, visual, and auditory perception modules, it can extract semantic, visual, and auditory perception feature vectors in the text to be corrected, realize deep fusion of multimodal information error detection, and significantly improve the error detection effect. After obtaining the error detection label through the error detection network, the error correction task is converted into a fill-in-the-blank task, and the powerful semantic understanding ability of the large language model is used for error correction. In combination with context learning, task-related examples are provided to the large model, so that it can quickly adapt to the Chinese spelling error correction task and avoid unconstrained polishing. At the same time, a self-reflection mechanism is proposed, so that the large model can reflect on the results after correction many times, thereby improving the performance of the model error correction. At the same time, the present invention is based on the training and construction of pre-trained models and large language models, avoiding the use of large-scale corpus for complex training and reducing application costs.
[0050] The technical features of the above-mentioned embodiments can be combined arbitrarily. In order to make the description concise, not all possible combinations of the technical features in the above-mentioned embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0051] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the patent application. It should be noted that a person of ordinary skill in the art may make various modifications and improvements without departing from the spirit of the present application, and these modifications and improvements fall within the scope of protection of the present application. Therefore, the scope of protection of the present patent application shall be determined by the appended claims.
Claims
1. A Chinese spelling error correction method based on multimodal enhancement, characterized in that: The following steps are involved: Input the text sequence to be corrected into the semantic perception module to obtain the semantic feature vector; The semantic perception module maps the text to be corrected into a character embedding code, and feeds the character embedding code and the corresponding position code into a stacked multi-layer Transformer encoder to output a semantic feature vector; Inputting the text sequence to be corrected into the auditory perception module to obtain an auditory feature vector; the auditory perception module converts the text sequence to be corrected into a Chinese pinyin sequence, concatenates and maps the text to be corrected and the Chinese pinyin sequence to form a character embedding code, and feeds the character embedding code and the corresponding position code into a stacked multi-layer Transformer encoder to output an auditory feature vector; Input the text sequence to be corrected into the visual perception module to obtain a visual feature vector; the visual feature vector forms a character image by font rendering of the text sequence to be corrected, and inputs it into MogaNet to output a visual feature vector; The semantic feature vector, the auditory feature vector, and the visual feature vector are fused and input into an error detection network to obtain a detection label, wherein the error detection label marks the error position in the text sequence to be corrected; The text sequence to be corrected and the error detection mark are input into the revision model to obtain the error-corrected text sequence; the revision model is a large language model.
2. The Chinese spelling error correction method based on multimodal enhancement according to claim 1, characterized in that: The semantic perception module includes a BERT pre-training model, which obtains the text sequence to be corrected and outputs a semantic feature vector.
3. The Chinese spelling error correction method based on multimodal enhancement according to claim 1, characterized in that: Each pinyin in the pinyin sequence of the auditory perception module includes an initial consonant and a final vowel.
4. The Chinese spelling error correction method based on multimodal enhancement according to claim 1, characterized in that: The visual perception module includes a MogaNet, which inputs a character image and outputs a visual feature vector; the MogaNet includes a coding layer, a Moga block and a linear layer, each of the Moga blocks includes a spatial aggregation layer and a channel aggregation layer connected in sequence; the coding layer is used to adjust the resolution of the character image; the spatial aggregation layer is used to fuse spatial context information of different scales, and the channel aggregation layer is used to adaptively redistribute channel weights; the linear layer is used to receive the feature vector output by the Moga block and map it to a preset category.
5. The Chinese spelling error correction method based on multimodal enhancement according to claim 4, characterized in that: In the spatial aggregation layer, a 1x1 convolution and a scaling factor To reweight the complementary interaction part, it is expressed as: ; ; in represents the feature vector of the original input of the spatial aggregation layer, GAP(·) represents the global average pooling, is the scaling factor, GELU (·) represents the activation function, Represents the output of the spatial aggregation layer.
6. The Chinese spelling error correction method based on multimodal enhancement according to claim 4, characterized in that: In the channel aggregation layer, it is expressed as: ; ; ; Among them, The feature vector representing the original input of the channel aggregation layer, represents normalization, DW (·) represents depth convolution, CA (·) represents channel aggregation, Represents the scaling factor, downscaling the projection by channels and GELU activation function to collect and redistribute information of channel dimensions, Represents the output of the channel aggregation layer.
7. The Chinese spelling error correction method based on multimodal enhancement according to claim 1, characterized in that: The training method of the auditory perception module includes: dividing the original expected text into sentences and filtering sentences whose lengths are outside a preset range; randomly sampling the filtered sentences to obtain an original text sequence; replacing some characters with words with similar pronunciations, randomly replacing some characters, and keeping some characters unchanged to obtain a replacement text sequence; obtaining the initial consonants and finals of the replacement text sequence and splicing them with the replacement text sequence, using the original text sequence as a label to obtain a training data set; replacing the unused tokens in the BERT model with the obtained initial consonants and finals, inputting the training data set into the BERT model to obtain character and corresponding pinyin embedding representations; obtaining a restored text sequence based on the character and pinyin embedding representations, comparing it with the label to calculate the loss and then updating the model parameters; training until the model converges.
8. The Chinese spelling error correction method based on multimodal enhancement according to claim 1, characterized in that: The error detection module combines the semantic feature vector, auditory feature vector, and visual feature vector through a fusion function, inputs them into a classifier to generate error probability, and determines the error detection label through an activation function, which is expressed as: ; ; ; in, represents the semantic feature vector, represents the auditory feature vector, represents the visual feature vector, represents the trainable parameters, Indicates bias, represents the softmax activation function, represents the error probability distribution of the corresponding position, Indicates the error detection tag.
9. The Chinese spelling error correction method based on multimodal enhancement according to claim 1, characterized in that: The Chinese spelling error correction method also includes the following steps: inputting a text sequence containing auditory errors, visual errors, and random errors and their corresponding corrected text sequences into a large language model, and inputting a prompt word describing the input content into the large language model; inputting a text sequence containing no errors into the large language model, and inputting a prompt word describing the input content into the large language model; inputting the text sequence to be corrected and the error detection mark into the large language model to obtain a corrected text sequence.
10. The Chinese spelling error correction method based on multimodal enhancement according to claim 9, characterized in that: The text sequence to be corrected and the corresponding correction text sequence are input into the large language model. The prompt word asks whether the semantic information is effectively utilized. When the large language model responds affirmatively, the correction text sequence is determined as the final correction text sequence.
Citation Information
Patent Citations
Chinese text spelling checking method, system and device and storage medium
CN117009914A
Error correction method and system for large language model fusing rich semantic information
CN117852528A
Chinese text correction method and device based on spelling check and computer equipment
CN119886116A
Text error correction method and device based on large model, equipment and medium
CN120373293A