Chinese medical text spelling error correction method and system of multi-scale SoftMasked-ChineseBERT model

Through the multi-scale SoftMasked-ChineseBERT model and joint detection model, the problem of low accuracy of spelling error detection in Chinese medical texts is solved, more efficient error detection and correction is achieved, and the quality and processing efficiency of medical texts are improved.

CN120012765APending Publication Date: 2025-05-16XI'AN UNIVERSITY OF ARCHITECTURE AND TECHNOLOGY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510069843.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-16
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The existing Chinese medical text error correction methods have low accuracy in spelling error detection, which affects the error correction efficiency.

Method used

The multi-scale SoftMasked-ChineseBERT model is used to split the text into a single-language sentence set through a multi-lingual step splitting algorithm, combining the joint detection model and the correction model to detect the wrong characters and correct them.

Benefits of technology

Improves the accuracy and efficiency of typo detection, especially when handling professional terms and complex contexts, improving the quality and processing efficiency of medical texts.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120012765A_ABST
    Figure CN120012765A_ABST
Patent Text Reader

Abstract

The invention discloses a Chinese medical text spelling error correction method and a Chinese medical text spelling error correction system based on a multi-scale SoftMasked-ChineseBERT (SoftMasked-ChineseBERT) model. The method comprises the following steps: splitting a Chinese medical text to be corrected into a monolingual step sentence set by adopting a multilingual step splitting algorithm; the monolingual step sentence set is input into a joint detection model to obtain an embedded sequence, the joint detection model comprises three sub-models, the embedded sequence is input into the three sub-models to obtain label sequences respectively, the three label sequences are subjected to weighted summation to obtain a character error probability sequence, and error characters of the embedded sequence are detected; inputting the embedded sequence into SoftMasked, and based on the character error probability sequence, shielding semantic features of error characters in the embedded sequence to obtain a fusion feature sequence; and inputting the fusion feature sequence into the correction model to obtain a correction character, and replacing the error character in the embedded sequence with the correction character. According to the method, the NLP technology and medical knowledge are combined, potential errors in the medical text can be effectively detected and corrected, accurate information transmission is guaranteed, and misdiagnosis is avoided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of natural language processing, and in particular relates to a Chinese medical text spelling correction method, system, medium and device of a multi-scale SoftMasked-ChineseBERT model. Background Art

[0002] With the rapid advancement of medical informatization, the digital processing of medical data has become an important part of the medical industry. These data include patient medical records, test reports, prescription information, etc., which are usually entered and managed through systems such as electronic health records (EHR) and electronic medical records (EMR). However, due to the inevitability of manual input, there are a large number of spelling errors in Chinese medical data, such as typos, miswritten homophones, missing words, etc., which pose challenges to the accuracy and completeness of the data. These errors not only affect clinical decisions and treatment results, but may also reduce the accuracy of medical data mining and artificial intelligence-assisted decision-making systems, thereby affecting the training and prediction capabilities of the model. In medical text analysis, spelling errors are particularly complex. The unique polyphones, homophones, and similar characters in Chinese make it difficult to effectively identify spelling errors through traditional spelling check technology. For example, "pneumonia" may be written as "pulmonary research", "multiple" may be written as "deprivation", or "diagnosis" may be written as "symptom determination". These spelling errors will have an adverse effect on the recognition of key information such as disease and drug names, increasing the risk of misdiagnosis.

[0003] In recent years, with the development of artificial intelligence technology, especially natural language processing (NLP) and deep learning, new solutions have been found for the automatic detection and correction of spelling errors. These technologies can not only identify common errors, but also perform semantic understanding in combination with context, thereby achieving more accurate error correction. Language models trained based on large-scale medical text corpora can effectively identify spelling errors in medical terms and provide intelligent spelling correction tools for the medical industry. However, there are still some challenges in the processing of Chinese medical data with existing technologies. First, medical data contains a large number of professional terms, abbreviations and dialects, which leads to more diverse manifestations of spelling errors, and general spelling correction algorithms are difficult to deal with these field-specific errors. Secondly, spelling errors in Chinese medical texts are usually greatly affected by context, and how to use context information to improve the accuracy of error correction is still a technical problem. Therefore, this technology combines deep learning with natural language processing methods to propose an innovative Chinese medical data spelling error detection and correction technology. This technology can not only identify and correct common spelling errors, but also provide high-precision correction for specific errors in the medical field, thereby improving the quality and processing efficiency of medical data and providing support for the further development of intelligent medicine.

[0004] In recent years, spelling correction methods based on pre-trained language models such as BERT have gained wide attention. BERT captures contextual information through bidirectional training and can effectively learn semantic features, which is crucial for spelling correction. In particular, when dealing with characters with similar spellings but different semantics, the pre-trained model can significantly improve the accuracy of correction. The FASpell model proposed by Hong et al. uses BERT as a denoising autoencoder to generate candidate characters and select the most likely corrected characters by calculating the similarity between characters. Although this method uses semantic information to improve the correction effect, it mainly focuses on the semantic similarity of characters and ignores the visual and phonetic similarity of characters. Liu et al. pointed out that the main causes of Chinese spelling errors can be divided into phonetic similarity and visual similarity. About 83% of spelling errors are caused by phonetic similarity, and 48% of errors are related to character shape similarity. Characters with similar pronunciations (such as "hungry" and "goose") or similar shapes (such as "and" and "world") are prone to spelling errors. Therefore, in addition to semantic information, phonetic and glyph similarity are also very important in spelling correction. In order to solve the problem of relying only on semantic information, some studies in recent years have tried to integrate speech and visual information into spelling correction models. SpellGCN model: This model uses a graph convolutional network (GCN) to model the similarity of glyphs and pinyin respectively, and combines BERT for feature initialization. REALISE model: Combines GRU and convolutional neural network (CNN) to extract glyph and phonetic features and enhance spelling correction capabilities.

[0005] PHMOSpell model: Through the VGG19 convolutional neural network and neural TTS model, features are extracted from both the shape and pronunciation of characters to improve the error correction effect. These methods have made some progress by introducing speech and visual information, but they still face the problem of feature inequality, because the spelling correction model and the pre-trained language model rely on different data sources, resulting in inconsistent learned feature types and distributions. Cui et al. proposed a Chinese spelling correction method based on ChineseBert. This method solves the problem of multimodal data inequality by adopting the ChineseBert pre-trained model, and by improving the embedding layer of Bert, the Bert model can fully modify characters that may have spelling errors.

[0006] The current mainstream method of correcting Chinese spelling errors has the following problems.

[0007] For long texts with too many characters, an end-to-end approach is used to input the original text and obtain the target text through the model. There are several challenges in processing long texts. Traditional sequence models (such as RNN, LSTM) will encounter the problems of "gradient disappearance" or "gradient explosion", which makes it difficult to capture long-distance dependencies. Even for the Transformer based on the self-attention mechanism, it is difficult to capture all the information at once when processing long texts due to excessive computational and memory consumption. In addition, long texts often contain redundant or irrelevant information. How to effectively extract useful information and ignore irrelevant parts is a key issue. Long texts may involve multiple topics or viewpoints. The model needs to recognize and flexibly handle these changes to maintain a global understanding, especially when the topic shifts.

[0008] Studies have shown that accurately detecting misspelled characters in text is crucial for effectively correcting errors. However, most traditional spelling correction methods fail to adequately address this problem. Although some methods, such as Chinese spelling correction technology based on ChineseBERT, have improved the ability to detect spelling errors to a certain extent, their detection results are still unsatisfactory, and the methods generally have the problem of being coarse, making it difficult to correct spelling errors in a fine and efficient manner. Therefore, how to improve the accurate detection of misspelled characters remains a key challenge in current spelling correction research. In recent years, deep learning methods, such as models based on BERT, LSTM, CNN, and Transformer, have shown strong potential for error correction. These models are able to automatically learn complex error patterns and effectively handle context-dependent spelling errors. However, in specific fields, especially in the medical field, the error correction effect of general models is not ideal. Summary of the invention

[0009] The present invention provides a multi-scale SoftMasked-ChineseBERT Chinese medical text spelling correction method, system, medium and device to solve the problem that the existing Chinese medical text spelling correction method has low accuracy in spelling error detection and affects the correction efficiency.

[0010] The present invention provides a Chinese medical text spelling correction method based on a multi-scale SoftMasked-ChineseBERT model, comprising:

[0011] The Chinese medical text to be corrected is split into a set of monolingual step sentences using a multilingual step splitting algorithm;

[0012] The monolingual step sentence set is input into the joint detection model to obtain an embedding sequence. The joint detection model includes three sub-models. The embedding sequence is input into the three sub-models to obtain label sequences respectively. The three label sequences are weighted summed to obtain a character error probability sequence to detect the wrong characters in the embedding sequence.

[0013] The embedded sequence is input into SoftMasked, and the semantic features of the erroneous characters in the embedded sequence are masked based on the character error probability sequence to obtain a fused feature sequence;

[0014] The fused feature sequence is input into the correction model to obtain the corrected characters, and the corrected characters are used to replace the incorrect characters in the embedded sequence.

[0015] The method of splitting the Chinese medical text to be corrected into a set of monolingual step sentences using a multilingual step splitting algorithm includes:

[0016] Divide the text to be corrected into several complete sentences using periods as boundaries;

[0017] The LTP platform is used to analyze the syntax of each complete sentence, determine the words in the complete sentence, determine the core words, and analyze the parallel structure in the complete sentence. When there is no parallel structure in the complete sentence, it is a monolingual step sentence and is directly output; when there is a parallel structure in the complete sentence, it is a multilingual step sentence and the sentence is split according to the core words;

[0018] When the parent node of the parallel structure is the root node, the two parallel core words are used as the dividing points to split the complete sentence into monolingual step sentences. When the parent node of the parallel structure is not the root node, the monolingual step sentences are directly output. All the output monolingual step sentences constitute a monolingual step set.

[0019] The joint detection model consists of three sub-models, namely BiGRU, TextCNN and DPCNN.

[0020] The monolingual step sentence set is input into the joint detection model to obtain an embedded sequence. The joint detection model includes three sub-models. The embedded sequence is input into the three sub-models to obtain label sequences respectively. The three label sequences are weighted summed to obtain a character error probability sequence. The error characters of the embedded sequence are detected, including:

[0021] Set the Chinese character sequence of each sentence in the monolingual sentence set, extract features of the Chinese character sequence through the joint detection model, and obtain the corresponding embedding sequence;

[0022] The joint detection model consists of three sub-models, namely BiGRU, TextCNN, and DPCNN. The embedded sequences are input into BiGRU, TextCNN, and DPCNN respectively to obtain three sets of label sequences.

[0023] The three sets of label sequences are weighted and summed to obtain the character error probability sequence, which is expressed as:

[0024] G i =w1·G BiGRUi +w2·G TextCNNi+w3·G DPCNNi

[0025] Among them, w1, w2, and w3 are the weights of the three sub-models, G BiGRUi Represents the character x output by BiGRU i Is it the probability of error? G TextCNNi Represents the character x output by TextCNN i Is it the probability of error? G DPCNNi Indicates DPCNN output character x i Is it the probability of error? G i For character x i The probability of being wrong.

[0026] The embedded sequence is input into SoftMasked, and the semantic features of the erroneous characters in the embedded sequence are masked based on the character error probability sequence to obtain a fused feature sequence, specifically:

[0027] Input the embedded sequence into SoftMasked, mask the semantic features of the erroneous characters in the embedded sequence according to the character error probability sequence, replace them with mask characters, retain the character's glyph vector and pronunciation vector, and obtain a fused feature sequence;

[0028] The fusion features are:

[0029]

[0030] in, For character x i The fusion features, e wi is the semantic vector; e si is the glyph vector; e pi is the word vector, e mask is the mask vector, G i For character x i The probability of being wrong.

[0031] The fused feature sequence is input into the correction model to obtain the corrected characters, and the corrected characters are used to replace the erroneous characters in the embedded sequence, specifically:

[0032] The fused feature sequence is input into the correction model. For the correct characters in the embedded sequence, the original input will be retained. For the wrong characters in the embedded sequence, the correction model generates a set of candidate characters for each wrong character, and uses the last layer of BERT to output the probability of error correction for each candidate character, expressed as P c (y i = j | X) = softmax(Wh′ i +b)[j]

[0033] Where: P c (y i =j|X) is the character x in the Chinese character sequence i The probability of being corrected as candidate character j, W and b are parameters in the correction model, h′ i Represents character x i The hidden state after the linear transformation of ;

[0034] The softmax function is used to calculate h′ i Calculate h′ i The calculation formula is as follows:

[0035]

[0036] in, is the hidden state of the last layer of the correction model; e i Represents character x i The embedding vector of

[0037] The correction model selects the candidate character with the highest probability as the correction character and replaces the incorrect character in the embedded sequence with the correction character.

[0038] The multi-scale SoftMasked-ChineseBert model includes a joint detection model and a correction model. The training of the multi-scale SoftMasked-ChineseBert model includes:

[0039] The multi-scale SoftMasked-ChineseBert model is trained by combining the objective functions of the detection model and the correction model. The loss function of the multi-scale SoftMasked-ChineseBert model is:

[0040] L=λ·L c +(1-λ)·L d

[0041] Among them, L d is the loss of the joint detection model, L c To correct the loss of the model, λ is the weight factor, and the value range of λ is [0,1].

[0042] The present invention also provides a Chinese medical text spelling correction system based on a multi-scale SoftMasked-ChineseBERT model, comprising:

[0043] A text splitting module is used to split the Chinese medical text to be corrected into a set of monolingual step sentences using a multilingual step splitting algorithm;

[0044] The error character detection module is used to input the monolingual step sentence set into the joint detection model to obtain an embedded sequence. The joint detection model includes three sub-models. The embedded sequence is input into the three sub-models to obtain label sequences respectively. The three label sequences are weighted summed to obtain a character error probability sequence to detect the error characters in the embedded sequence.

[0045] The error character masking module is used to input the embedded sequence into SoftMasked, and based on the character error probability sequence, the semantic features of the error characters in the embedded sequence are masked to obtain a fused feature sequence;

[0046] The error character correction module is used to input the fused feature sequence into the correction model to obtain the corrected characters, and replace the error characters in the embedded sequence with the corrected characters.

[0047] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the Chinese medical text spelling correction method based on the multi-scale SoftMasked-ChineseBERT model as described in any one of claims 1 to 7 when executing the computer program.

[0048] The present invention also provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the steps of the Chinese medical text spelling correction method based on the multi-scale SoftMasked-ChineseBERT model as described in any one of claims 1 to 7.

[0049] In order to achieve the above object, the present invention provides the following technical solutions:

[0050] Compared with the prior art, the present invention has the following beneficial effects:

[0051] The present invention adopts a multi-language step splitting algorithm to split the text to be corrected into a set of single-language step sentences, and the single-language step sentences can improve the accuracy of error detection; the joint detection model integrates multiple models, and through multi-level context information capture and local feature extraction, it can capture different levels of erroneous characters, and enhance the accuracy and robustness of character error recognition. When correcting erroneous characters, SoftMasked is used to shield the semantic features of erroneous characters, while retaining the glyph and pronunciation information; the correction model can better rely on the context for correction, thereby improving the accuracy and efficiency of correction. Through multi-level context capture, feature fusion and precise correction, the present invention can better adapt to the language features in medical texts and improve the accuracy and robustness of medical text processing. The present invention improves the efficiency and accuracy of spelling correction through multi-language step splitting, multi-model fusion detection model and correction model, especially showing excellent performance in the processing of professional terms and complex contexts. The present invention provides effective technical support for improving the quality of Chinese medical texts, and has strong practical application value and broad application prospects.

[0052] Furthermore, the joint detection model integrates three sub-models: BiGRU, TextCNN, and DPCNN. These sub-models support character error recognition from different angles through bidirectional information flow, convolution operations, and deep feature extraction. Through weighted fusion output, they comprehensively evaluate character errors and improve the efficiency and accuracy of spelling correction.

[0053] Furthermore, when optimizing the model, by jointly optimizing the detection and correction tasks, the redundant operations between error detection and correction are reduced, thus improving the overall processing efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for use in the embodiments are briefly introduced below. It should be understood that the following drawings only show certain embodiments of the present invention and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without creative work.

[0055] Figure 1 It is a flow chart of the Chinese medical text spelling correction method based on the multi-scale SoftMasked-ChineseBert model of the present invention;

[0056] Figure 2 Dependency parsing diagram for the example;

[0057] Figure 3 This is the syntax analysis matrix diagram of Example 1;

[0058] Figure 4It is a flowchart of the multi-language step splitting algorithm of the present invention;

[0059] Figure 5 Schematic diagram of the multi-scale SoftMasked-ChineseBert model;

[0060] Figure 6 This is a case diagram of error correction of the multi-scale SoftMasked-ChineseBert model;

[0061] Figure 7 It is a schematic diagram of the structure of a Chinese medical text spelling correction system based on a multi-scale SoftMasked-ChineseBert model according to a preferred embodiment of the present invention;

[0062] Figure 8 It is a schematic diagram of the structure of an electronic device according to a preferred embodiment of the present invention. DETAILED DESCRIPTION

[0063] In order to enable those skilled in the art to understand the characteristics and effects of the present invention, the following is a general description and definition of the terms and expressions mentioned in the specification and claims. Unless otherwise specified, all technical and scientific terms used in the text are the common meanings understood by those skilled in the art for the present invention. In the event of a conflict, the definition in this specification shall prevail.

[0064] The theories or mechanisms described and disclosed herein, whether correct or incorrect, should not limit the scope of the present invention in any way, that is, the present invention can be implemented without being limited by any specific theory or mechanism.

[0065] Herein, all features such as values, quantities, contents and concentrations defined in the form of numerical ranges or percentage ranges are for simplicity and convenience only. Accordingly, the description of numerical ranges or percentage ranges should be considered to have included and specifically disclosed all possible secondary ranges and individual values ​​within the range (including integers and fractions).

[0066] In this document, unless otherwise specified, “includes,” “including,” “contains,” “has,” or similar terms cover the meanings of “consisting of” and “mainly consisting of,” for example, “A includes a” covers the meanings of “A includes a and other” and “A only includes a.”

[0067] In this document, in order to make the description concise, not all possible combinations of various technical features in various embodiments or examples are described. Therefore, as long as there is no contradiction in the combination of these technical features, the various technical features in various embodiments or examples can be combined arbitrarily, and all possible combinations should be considered to be within the scope of this specification.

[0068] like Figure 1 and Figure 5 As shown, the present invention proposes a Chinese medical text spelling correction method based on a multi-scale SoftMasked-ChineseBERT model, comprising the following steps:

[0069] The Chinese medical text to be corrected is split into a set of monolingual step sentences using a multilingual step splitting algorithm;

[0070] The monolingual step sentence set is input into the joint detection model to obtain an embedding sequence. The joint detection model includes three sub-models. The embedding sequence is input into the three sub-models to obtain label sequences respectively. The three label sequences are weighted summed to obtain a character error probability sequence, and the wrong characters in the embedded sequence are detected.

[0071] The embedded sequence is input into SoftMasked, and the semantic features of the erroneous characters in the embedded sequence are masked based on the character error probability sequence to obtain a fused feature sequence;

[0072] The fused feature sequence is input into the correction model to obtain the corrected characters, and the corrected characters are used to replace the incorrect characters in the embedded sequence.

[0073] The Chinese medical text spelling correction algorithm based on the multi-scale SoftMasked-ChineseBert model provided by the present invention aims to improve the accuracy of spelling, grammar and word usage of medical texts; by combining NLP technology with medical knowledge, the algorithm can effectively detect and correct potential errors in medical texts, ensure accurate information transmission, and avoid misdiagnosis.

[0074] Among them, Figure 4 As shown, the Chinese medical text to be corrected is split into a set of monolingual step sentences using a multilingual step splitting algorithm, specifically:

[0075] Divide the text to be corrected into several complete sentences using periods as boundaries;

[0076] The LTP platform (language technology platform) is used to analyze the syntax of each complete sentence, determine the words in the complete sentence, determine the core words, and analyze the parallel structure (COO) in the complete sentence. When there is no COO in the complete sentence, it is a monolingual sentence and does not need to be split and is directly output; when there is COO in the complete sentence, it is a multilingual sentence and the sentence is split according to the core words;

[0077] When the parent node of COO is the root node (HED), the two parallel core words are used as the splitting points to split the complete sentence into monolingual step sentences. When the parent node of COO is not the root node, the monolingual step sentences are directly output. All the output monolingual step sentences constitute a monolingual step set.

[0078] It should be noted that the multi-step splitting algorithm proposed in the present invention is based on periods and dependency syntax. The use of dependency grammar analysis plays a vital role in natural language processing based on artificial intelligence. Through the Chinese grammatical rules, the structure of the sentence is clarified and the dependency relationship between words is revealed, so as to achieve the purpose of reducing the semantic complexity of the sentence. The core goal of dependency syntactic analysis is to analyze the mutual relationship between words, such as "subject, predicate, object", "attributive, adverbial, complement", etc., to find the core words in the sentence. The result of dependency analysis of the sentence "The patient underwent chest CT at the county hospital and showed pulmonary interstitial fibrosis, suspected of tuberculosis" using the LTP platform is shown as follows Figure 2 exhibit.

[0079] Punctuation marks play a vital role in language understanding, and each punctuation mark has its unique function. For example, a comma indicates that a sentence has not yet ended, and a short pause enhances semantic clarity; while a period indicates the end of semantics. The present invention first divides the abstract into multiple semantically complete sentences by a period. Although the sentences separated by a period are semantically complete, they may contain multiple semantics. Studies have shown that complex semantics may affect the processing effect of natural language models. Therefore, the present invention combines periods and dependency syntactic analysis to process each Chinese medical data text.

[0080] The relationship between sentences, such as parallel relationships, can be determined by analyzing the dependency relationship of sentence components. The structure of each clause is independent and has a logical semantic relationship, so the method of splitting multi-step structures combined with punctuation and syntactic analysis is more effective.

[0081] Embodiment 1:

[0082] For example, "The patient was diagnosed with suppurative appendicitis and paralytic ileus, and underwent intestinal relaxation and appendectomy under general anesthesia at 18:30." The result matrix of dependency syntactic analysis is as follows Figure 3 As shown. Figure 3 , HED is the root node, and the core word "diagnosis" corresponds to HED. In the continuous sentence, a COO structure is found, whose parent node is HED, forming a parallel relationship, so "diagnosis" and "conduct" are the two central words of the sentence, connected by commas. According to syntactic analysis, this sentence is a multi-step structure sentence, which can be divided into two single-step structures (singstep). Each parallel clause contains its own SBV and VOB, so it can independently form a monolingual step structure. In the end, the two monolingual step structures are: "The patient was diagnosed with suppurative appendicitis and paralytic ileus" and "Enterolysis and appendectomy were performed under general anesthesia at 18:30".

[0083] The multi-scale SoftMasked-ChineseBert model proposed in the present invention includes two parts: a joint detection model and a correction model.

[0084] It should be noted that the joint detection model is a sequential binary tagging model, which converts the error character recognition task into a sequential binary tagging problem. The goal of the character error recognition task is to determine whether each character in the input text sequence is wrong. The joint detection model first identifies potential error characters, shields their semantic features, and only retains the glyph and pronunciation information that is useful for prediction. Then, the denoised features are input into the correction model for correction. Since the correction model has removed the semantic interference of the error characters during input, and the glyph and pronunciation features have played a limiting role in the prediction, the model can more effectively solve the problem of spelling errors.

[0085] The monolingual step sentence set is input into the joint detection model to obtain an embedded sequence. The joint detection model includes three sub-models. The embedded sequence is input into the three sub-models to obtain label sequences respectively. The three label sequences are weighted summed to obtain a character error probability sequence, and the error characters of the embedded sequence are detected, specifically:

[0086] Assume that the Chinese character sequence X of each sentence in the monolingual sentence set is {x1, x2, ..., x n},}x i is the i-th character in the sentence. The joint detection model is used to extract features from the Chinese character sequence X and obtain the corresponding embedding sequence E = (e1, e2, ...., e n ), where e i Represents character x i The embedding vector, e i Contains the word embedding information for this character.

[0087] The joint detection model consists of three sub-models, namely BiGRU, TextCNN, and DPCNN. The embedded sequence E is input into BiGRU, TextCNN, and DPCNN respectively to obtain three sets of label sequences. The label sequence output by BiGRU is G BiGRU =(G BiGRU1 ,G BiGRU2 ,…,G BiGRUn ), where G BiGRUi Represents the character x output by BiGRU i Is it the probability of error? The label sequence output by TextCNN is G TextCNN =(G TextCNN1 ,G TextCNN2 ,…,G TextCNNn ), G TextCNNi Represents the character x output by TextCNN i Is it the probability of error? The label sequence output by DPCNN is G DPCNN =(G DPCNN1 ,G DPCNN2 ,…,GDPCNNn ), G DPCNNi Indicates DPCNN output character x i The probability of being wrong.

[0088] The three sets of label sequences are weighted and summed to obtain the character error probability sequence G = (G1, G2, ..., G n ), the formula is:

[0089] G i =w1·G BiGRUi +w2·G TextCNNi +w3·G DPCNNi

[0090] Among them, w1, w2, w3 represent the weights of the three sub-models, G i For character x i The probability of being wrong.

[0091] It should be noted that the character error probability sequence fuses the output information of the three sub-models, enabling the model to more comprehensively capture the semantic information of each character in the context, thereby improving the subsequent processing effect.

[0092] The embedded sequence is input into SoftMasked, and the semantic features of the erroneous characters in the embedded sequence are masked based on the character error probability sequence to obtain a fused feature sequence, specifically:

[0093] Input the embedded sequence E into SoftMasked, and mask the semantic features of the erroneous characters in the embedded sequence E according to the character error probability sequence G, and replace them with mask characters, retaining the characters x i The glyph vector and the pronunciation vector of the word can be used to provide more accurate assistance when performing error correction to obtain a fused feature sequence in Represents character x i The fusion features are expressed as follows:

[0094]

[0095] in, For character x i The fusion features, e wi is the semantic vector; e si is the glyph vector; e pi is the word vector, e mask is the mask vector, G i For character x i The probability of being wrong.

[0096] It should be noted that after identifying misspelled characters, the feature fusion strategy combines the detected character error probability with the word vector information of the character to generate a fused feature vector containing the character error probability and semantic information. The fused feature vector can not only express the semantic relationship of the characters, but also help the correction model understand the potential meaning in the context when correcting errors, thereby providing a more accurate basis for subsequent error correction. The purpose of this feature fusion strategy is to ensure that in the process of identifying and correcting incorrect characters, the correction model is not limited to relying solely on shielded semantic information, but takes into account the glyph and phonetic features, so that the correction model takes into account the morphology and pronunciation of the characters when correcting errors, thereby improving the accuracy of error correction.

[0097] The fused feature sequence is input into the correction model to obtain the corrected characters, and the corrected characters are used to replace the erroneous characters in the embedded sequence, specifically:

[0098] The fused feature sequence E′ is input into the correction model. For the correct characters in the embedded sequence, the original input will be retained. For the wrong characters in the embedded sequence, the correction model generates a set of candidate characters for each wrong character, and uses the last layer of BERT to output the probability of error correction for each candidate character, which is expressed as:

[0099] P c (y i = j | X) = softmax(Wh′ i +b)[j]

[0100] Where: P c (y i =j|X) is the character x in the Chinese character sequence X i The probability of being corrected as candidate character j, W and b are parameters in the correction model, h′ i Represents character x i The hidden state after the linear transformation of , which is obtained through linear combination and residual connection;

[0101] The softmax function is used to calculate h′ i Calculate h′ i The calculation formula is as follows:

[0102]

[0103] in, is the hidden state of the last layer of the correction model; e i Represents character x i The embedding vector of , which provides additional information about the character;

[0104] The correction model selects the candidate character with the highest probability as x iThe correction character is used to replace the incorrect character in the embedded sequence.

[0105] It should be noted that the correction model is based on the SoftMasked-ChineseBERT model, combined with a multi-head self-attention mechanism and a feedforward neural network to perform error correction at the character level. The correction model performs spelling correction based on a fused feature sequence, in which the semantic features of the incorrect characters have been masked. Therefore, the correction model can only rely on contextual information and the retained pronunciation and glyphs to predict these masked parts.

[0106] The multi-scale SoftMasked-ChineseBert model optimizes the detection and correction probability of each character through end-to-end training, combined with error detection and correction goals. The training data generates error sequences through confusion tables. The optimization goal of the model is to improve the error detection and correction performance at the same time through weighted loss functions, and finally achieve high-accuracy character correction.

[0107] The multi-scale SoftMasked-ChineseBert model includes a joint detection model and a correction model. The training of the multi-scale SoftMasked-ChineseBert model includes:

[0108] The multi-scale SoftMasked-ChineseBert model is trained by combining the objective functions of the detection model and the correction model. The loss function of the multi-scale SoftMasked-ChineseBert model is:

[0109] L=λ·L c +(1-λ)·L d

[0110] Among them, L d is the loss of the joint detection model, L c To correct the loss of the model, λ is the weight factor, and the value range of λ is [0,1].

[0111] In order to optimize the training process of the joint detection model, the joint detection model adopts a binary cross entropy loss function, which is applicable to the binary classification task of each character (determining whether the character is wrong).

[0112] The loss function of the joint detection model is the average of all character losses, expressed as:

[0113]

[0114] Each character x i The loss function is:

[0115]

[0116] Among them, gi is the character x i The true label (0 or 1), p i For character x i is the probability of error;

[0117] The joint detection model includes three sub-models. The weights of the three sub-models are w1, w2, and w3. w1, w2, and w3 will participate in the back-propagation process and will be automatically adjusted according to the gradient of the loss function during training.

[0118] The loss function of the calibration model is expressed as:

[0119]

[0120] Among them, p(y i |x i ) is the character x i Corrected to y i probability.

[0121] Finally, the objective functions of the detection network and the correction network are combined through linear combination to optimize the multi-scale SoftMasked-ChineseBert model.

[0122] Example 2

[0123] This embodiment uses the official training data of SIGHAN2013 and the MCSCSet training data published by Fudan University as training sets. SIGHAN2013 is a general Chinese spelling correction dataset, and MCSCSet is a specific Chinese spelling correction dataset in the medical field. The statistical results of the dataset training are shown in Table 1.

[0124] Table 1 Dataset statistics

[0125]

[0126] Tables 2 and 3 respectively give the evaluation results of the sentence-level detection level and correction level of the model of the present invention and four comparative models on the training set SIGHAN2013 and the training set MCSCSet.

[0127] Table 2 Comparison results of SIGHAN2013 experiments

[0128]

[0129] Table 3 MCSCSet experimental comparison results

[0130]

[0131] On the SIGHAN2013 dataset, the model of the present invention achieved excellent results at both the detection and correction levels. At the detection level, the precision of the model of the present invention is 87.6, the recall rate is 85.4, and the F1 is 86.6, which is significantly better than other models. At the correction level, the precision of the model of the present invention is 86.9, the recall rate is 83.0, and the F1 is 84.9, demonstrating its powerful error correction ability. Compared with other models, especially SpellGCN and BERT, the model of the present invention can effectively improve the recall rate, thereby ensuring more comprehensive error recognition and correction while maintaining a high precision rate. On the MCSCSet dataset in the medical field, the model of the present invention also performs well. At the detection level, the precision rate is 87.4, the recall rate is 88.1, and the F1 is 87.8, far exceeding other models. Especially in terms of recall rate, the performance of the model of the present invention is very outstanding, and more erroneous characters can be detected. At the correction level, the precision is 81.6, the recall is 81.3, and the F1 is 81.3, showing its good performance in specific tasks in the medical field. Compared with BERT, MedBERT and SoftBERT, the F1 value of the model of the present invention is always in a leading position, especially in the improvement of the recall rate, showing its professionalism in text processing in the medical field, and can more accurately identify and correct spelling errors in specific fields. The performance on the MCSCSet dataset verifies the superiority of the model of the present invention in the medical field. Texts in the medical field usually contain a large number of professional terms, abbreviations and industry-specific languages, which are often misidentified or improperly corrected in traditional general models. The model of the present invention has demonstrated good adaptability and professionalism in text processing in the medical field through its powerful error detection and correction capabilities. In particular, the improvement in the recall rate ensures that more specific errors in the medical field are detected and corrected in a timely manner, thereby improving the practical application value of the model.

[0132] Through experimental comparison on the SIGHAN2013 and MCSCSet datasets, the model of the present invention has shown significant advantages in the task of Chinese character error detection and correction. Especially in terms of recall rate and F1 value, the model of the present invention has shown high accuracy and robustness both on general text and specific datasets in the medical field. This is due to the multi-sub-model integration strategy adopted by the model in the detection phase, and the powerful semantic representation ability of SoftMasked-ChineseBERT in the correction phase. In the application of the medical field, the model of the present invention has demonstrated its unique professionalism, can effectively handle spelling errors in medical texts, and has broad application potential.

[0133] In order to more intuitively demonstrate the performance of the multi-scale SoftMasked-ChineseBert model proposed in the present invention for the Chinese medical field spelling correction task, including the advantages and disadvantages of the model in the continuous error problems of pronunciation, glyphs, medical proper nouns, and complex semantics, this embodiment gives the prediction results of five cases, such as Figure 6 As shown in the figure, in case 1, "Tang Niu Bing Ketone Orthoacidosis Insulin Treatment Method", "Tang" and "Tang", "Zheng" and "Zheng" are similar in both shape and pronunciation, and the model can correct them correctly. In case 2, "Liao Gaoxin is a drug commonly used to treat mental failure and certain types of arrhythmia, but it has some contraindications and situations that need attention", "Liao" and "Di", "Li" and "Li" are similar in pronunciation, and the model can correct them correctly. From the above two cases, it can be seen that the multi-scale SoftMasked-ChineseBert model has a better performance in Chinese pronunciation and spelling correction. In case 3, "Symptoms of Jin Run Breast Smelling Cancer Lotus Node Metastasis", the three characters "Jin", "Xing" and "Lian" appear in succession. The model can also correct this kind of continuous error problem correctly according to the context of the text. The original Bert model adopts a random masking strategy, which makes most of the characters that are easy to misspell unable to be trained, and in the problem of continuous errors, the wrong characters will affect the semantic expression of other wrong characters. "Continue to monitor blood pressure, adjust antihypertensive drug Enalapril, and use painkillers to relieve headache symptoms." "Enalapril" is a very professional term in the medical field, and the general Chinese spelling error correction model performs poorly for such medical professional terms. Since the MCSCSe dataset has sufficient medical field data, the multi-scale SoftMasked-ChineseBert model can also handle this problem well. Case 5 "When seeking medical treatment for paraganglioma of the bladder, patients need to pay attention to bring CT, urodynamic examination, etc." "Attached" and "flow" were not correctly corrected. The semantics of Case 5 is relatively complex and ambiguous, and the model's ability to handle continuous errors in this case is relatively insufficient. Through the display and analysis of the above cases, it can be seen that the multi-scale SoftMasked-ChineseBert model performs better in the two datasets. The model improves the detection rate of incorrect characters through the multi-scale joint detection model and utilizes the processing advantages of ChineseBert in Chinese character shapes and the soft masking strategy, and has a better performance in the spelling error correction task in the Chinese medical field.

[0134] like Figure 7 As shown, another object of the present invention is to propose a Chinese medical text spelling correction system based on a multi-scale SoftMasked-ChineseBERT model, comprising:

[0135] A text splitting module is used to split the Chinese medical text to be corrected into a set of monolingual step sentences using a multilingual step splitting algorithm;

[0136] The error character detection module is used to input the monolingual step sentence set into the joint detection model to obtain an embedded sequence. The joint detection model includes three sub-models. The embedded sequence is input into the three sub-models to obtain label sequences respectively. The three label sequences are weighted summed to obtain a character error probability sequence to detect the error characters in the embedded sequence.

[0137] The error character masking module is used to input the embedded sequence into SoftMasked, and based on the character error probability sequence, the semantic features of the error characters in the embedded sequence are masked to obtain a fused feature sequence;

[0138] The error character correction module is used to input the fused feature sequence into the correction model to obtain the corrected characters, and replace the error characters in the embedded sequence with the corrected characters.

[0139] like Figure 8 As shown, the third object of the present invention is to provide an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the steps of the Chinese medical text spelling correction method based on the multi-scale SoftMasked-ChineseBERT model are implemented.

[0140] The Chinese medical text spelling correction method based on the multi-scale SoftMasked-ChineseBERT model includes:

[0141] The Chinese medical text to be corrected is split into a set of monolingual step sentences using a multilingual step splitting algorithm;

[0142] The monolingual step sentence set is input into the joint detection model to obtain an embedding sequence. The joint detection model includes three sub-models. The embedding sequence is input into the three sub-models to obtain label sequences respectively. The three label sequences are weighted summed to obtain a character error probability sequence to detect the wrong characters in the embedding sequence.

[0143] The embedded sequence is input into SoftMasked, and the semantic features of the erroneous characters in the embedded sequence are masked based on the character error probability sequence to obtain a fused feature sequence;

[0144] The fused feature sequence is input into the correction model to obtain the corrected characters, and the corrected characters are used to replace the incorrect characters in the embedded sequence.

[0145] The fourth object of the present invention is to provide a computer-readable storage medium, which stores a computer program, and when the computer program is executed by a processor, it implements the steps of the Chinese medical text spelling correction method based on the multi-scale SoftMasked-ChineseBERT model.

[0146] The Chinese medical text spelling correction method based on the multi-scale SoftMasked-ChineseBERT model includes:

[0147] The Chinese medical text to be corrected is split into a set of monolingual step sentences using a multilingual step splitting algorithm;

[0148] The monolingual step sentence set is input into the joint detection model to obtain an embedding sequence. The joint detection model includes three sub-models. The embedding sequence is input into the three sub-models to obtain label sequences respectively. The three label sequences are weighted summed to obtain a character error probability sequence to detect the wrong characters in the embedding sequence.

[0149] The embedded sequence is input into SoftMasked, and the semantic features of the erroneous characters in the embedded sequence are masked based on the character error probability sequence to obtain a fused feature sequence;

[0150] The fused feature sequence is input into the correction model to obtain the corrected characters, and the corrected characters are used to replace the incorrect characters in the embedded sequence.

[0151] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, systems, or computer program products. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0152] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0153] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0154] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process in the computer or other programmable device. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0155] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, ordinary technicians in the relevant field should understand that the specific implementation methods of the present invention can still be modified or replaced by equivalents. Any modification or equivalent replacement that does not depart from the spirit and scope of the present invention should be covered within the scope of protection of the claims of the present invention.

Claims

1. A Chinese medical text spelling correction method based on a multi-scale SoftMasked-ChineseBERT model, characterized in that: include: The Chinese medical text to be corrected is split into a set of monolingual step sentences using a multilingual step splitting algorithm; The monolingual step sentence set is input into the joint detection model to obtain an embedding sequence. The joint detection model includes three sub-models. The embedding sequence is input into the three sub-models to obtain label sequences respectively. The three label sequences are weighted summed to obtain a character error probability sequence to detect the wrong characters in the embedding sequence. The embedded sequence is input into SoftMasked, and the semantic features of the erroneous characters in the embedded sequence are masked based on the character error probability sequence to obtain a fused feature sequence; The fused feature sequence is input into the correction model to obtain the corrected characters, and the corrected characters are used to replace the incorrect characters in the embedded sequence.

2. According to claim 1, a Chinese medical text spelling correction method based on a multi-scale SoftMasked-ChineseBERT model is characterized in that: The method of splitting the Chinese medical text to be corrected into a set of monolingual step sentences using a multilingual step splitting algorithm includes: Divide the text to be corrected into several complete sentences using periods as boundaries; The LTP platform is used to analyze the syntax of each complete sentence, determine the words in the complete sentence, determine the core words, and analyze the parallel structure in the complete sentence. When there is no parallel structure in the complete sentence, it is a monolingual step sentence and is directly output; when there is a parallel structure in the complete sentence, it is a multilingual step sentence and the sentence is split according to the core words; When the parent node of the parallel structure is the root node, the two parallel core words are used as the dividing points to split the complete sentence into monolingual step sentences. When the parent node of the parallel structure is not the root node, the monolingual step sentences are directly output. All the output monolingual step sentences constitute a monolingual step set.

3. According to claim 1, a Chinese medical text spelling correction method based on a multi-scale SoftMasked-ChineseBERT model is characterized in that: The joint detection model consists of three sub-models, namely BiGRU, TextCNN and DPCNN.

4. A Chinese medical text spelling correction method based on a multi-scale SoftMasked-ChineseBERT model according to claim 3, characterized in that: The monolingual step sentence set is input into the joint detection model to obtain an embedded sequence. The joint detection model includes three sub-models. The embedded sequence is input into the three sub-models to obtain label sequences respectively. The three label sequences are weighted summed to obtain a character error probability sequence. The error characters of the embedded sequence are detected, including: Set the Chinese character sequence of each sentence in the monolingual sentence set, extract features of the Chinese character sequence through the joint detection model, and obtain the corresponding embedding sequence; The joint detection model consists of three sub-models, namely BiGRU, TextCNN, and DPCNN. The embedded sequences are input into BiGRU, TextCNN, and DPCNN respectively to obtain three sets of label sequences. The three sets of label sequences are weighted and summed to obtain the character error probability sequence, which is expressed as: G i =w1·G BiGRUi +w2·G TextCNNi +w3·G DPCNNi Among them, w1, w2, and w3 are the weights of the three sub-models, G BiGRUi Represents the character x output by BiGRU i Is it the probability of error? G TextCNNi Represents the character x output by TextCNN i Is it the probability of error? G DPCNNi Indicates DPCNN output character x i Is it the probability of error? G i For character x i The probability of being wrong.

5. According to claim 1, a Chinese medical text spelling correction method based on a multi-scale SoftMasked-ChineseBERT model is characterized in that: The embedded sequence is input into SoftMasked, and the semantic features of the erroneous characters in the embedded sequence are masked based on the character error probability sequence to obtain a fused feature sequence, specifically: Input the embedded sequence into SoftMasked, mask the semantic features of the erroneous characters in the embedded sequence according to the character error probability sequence, replace them with mask characters, retain the character's glyph vector and pronunciation vector, and obtain a fused feature sequence; The fusion features are: in, For character x i The fusion features, e wi is the semantic vector; e si is the glyph vector; e pi is the word vector, e mask is the mask vector, G i For character x i The probability of being wrong.

6. A Chinese medical text spelling correction method based on a multi-scale SoftMasked-ChineseBERT model according to claim 1, characterized in that: The fused feature sequence is input into the correction model to obtain the corrected characters, and the corrected characters are used to replace the erroneous characters in the embedded sequence, specifically: The fused feature sequence is input into the correction model. For the correct characters in the embedded sequence, the original input will be retained. For the wrong characters in the embedded sequence, the correction model generates a set of candidate characters for each wrong character, and uses the last layer of BERT to output the probability of error correction for each candidate character, which is expressed as P c (y i =j|X)=softmax(Wh′ i +b)[j] Where: P c (y i =j|X) is the character x in the Chinese character sequence i The probability of being corrected as candidate character j, W and b are parameters in the correction model, h′ i Represents character x i The hidden state after the linear transformation of ; The softmax function is used to calculate h′ i Calculate h′ i The calculation formula is as follows: in, is the hidden state of the last layer of the correction model; e i Represents character x i The embedding vector of The correction model selects the candidate character with the highest probability as the correction character and replaces the incorrect character in the embedded sequence with the correction character.

7. A Chinese medical text spelling correction method based on a multi-scale SoftMasked-ChineseBERT model according to claim 1, characterized in that: The multi-scale SoftMasked-ChineseBert model includes a joint detection model and a correction model. The training of the multi-scale SoftMasked-ChineseBert model includes: The multi-scale SoftMasked-ChineseBert model is trained by combining the objective functions of the detection model and the correction model. The loss function of the multi-scale SoftMasked-ChineseBert model is: L=λ·L c +(1-λ)·L d Among them, L d is the loss of the joint detection model, L c To correct the loss of the model, λ is the weight factor, and the value range of λ is [0,1].

8. A Chinese medical text spelling correction system based on a multi-scale SoftMasked-ChineseBERT model, characterized in that: include: A text splitting module is used to split the Chinese medical text to be corrected into a set of monolingual step sentences using a multilingual step splitting algorithm; The error character detection module is used to input the monolingual step sentence set into the joint detection model to obtain an embedded sequence. The joint detection model includes three sub-models. The embedded sequence is input into the three sub-models to obtain label sequences respectively. The three label sequences are weighted summed to obtain a character error probability sequence to detect the error characters in the embedded sequence. The error character masking module is used to input the embedded sequence into SoftMasked, and based on the character error probability sequence, the semantic features of the error characters in the embedded sequence are masked to obtain a fused feature sequence; The error character correction module is used to input the fused feature sequence into the correction model to obtain the corrected characters, and replace the error characters in the embedded sequence with the corrected characters.

9. An electronic device, characterized in that: The invention comprises a memory, a processor and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the Chinese medical text spelling correction method based on the multi-scale SoftMasked-ChineseBERT model described in any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of the Chinese medical text spelling correction method based on the multi-scale SoftMasked-ChineseBERT model described in any one of claims 1 to 7.