Text error correction method and system for insurance input terms and related equipment
By constructing a dedicated lexicon and training model for insurance clauses, and combining multiple error correction methods, the problems of low accuracy and high missed detection rate of text error correction in insurance clause scenarios were solved, achieving a highly efficient text error correction effect.
Patent Information
- Application Number
- CN202410558596.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-08
- Publication Date
- 2025-11-14
AI Technical Summary
Existing text correction methods are ineffective in insurance policy scenarios due to a lack of dedicated lexicons and training data, resulting in low accuracy and high false negative rates.
By constructing a dedicated vocabulary for insurance terms, training a ternary statistical language model and a deep learning model, and combining homophone, similar-looking, and similar-sounding character error correction methods, word-level and character-level error correction is performed, and a deep learning model is used for supplementary error detection.
It significantly improved the accuracy of text correction in insurance clause scenarios, reduced the rate of missed detections, and enhanced the convenience of intelligent search services for insurance clauses.
Smart Images

Figure CN120950663A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of natural language processing technology, and in particular to a text correction method, system, and related equipment for insurance input clauses. Background Technology
[0002] Text correction technology plays a crucial role in the field of natural language processing, contributing to improved efficiency and accuracy in both everyday life and professional settings. Text correction methods primarily include rule-based and statistical learning-based approaches. Rule-based methods typically rely on predefined language models and are only suitable for simple text correction. Statistical learning-based methods, on the other hand, require extensive corpora to learn the statistical patterns of language. This approach demands a high degree of comprehensiveness in the corpus; otherwise, the accuracy of correction will be compromised.
[0003] The methods described above do not have specific vocabulary databases developed for particular scenarios such as insurance clauses, nor are they trained using corpora specific to these scenarios. Furthermore, these general methods have not been screened for specific scenarios to determine which methods perform better in particular situations. Therefore, when these general text correction methods are directly applied to specific scenarios like insurance clauses, they may not achieve the desired results. Summary of the Invention
[0004] The purpose of this invention is to provide a text correction method, system, and related equipment based on insurance input terms to solve at least one of the above-mentioned problems.
[0005] To achieve the above objectives, embodiments of the present invention provide a text correction method for insurance input clauses, comprising the following steps:
[0006] S1: Preprocessing and expansion of the insurance clause terminology database;
[0007] S101: Obtain the insurance terms text, perform word segmentation preprocessing to form an insurance terms text dictionary, and convert the words in the insurance terms text dictionary into pinyin to form an insurance terms pinyin dictionary;
[0008] S102: Obtain a general terminology and dictionary, merge them with the insurance clause text terminology and the insurance clause pinyin terminology, remove punctuation marks and English letters and numbers, and form a comprehensive insurance clause terminology;
[0009] Furthermore, it acquires general-purpose thesaurus and dictionaries, including dictionaries containing stroke order, Wubi input method, stroke count, and pronunciation information for Chinese characters.
[0010] S103: Narrow the scope of the comprehensive insurance clause terminology to Chinese characters appearing in the insurance clause text to form an insurance clause terminology;
[0011] This invention improves the comprehensiveness and accuracy of the terminology database for insurance clauses in various application scenarios by preprocessing the database.
[0012] S2: Statistical language model self-training;
[0013] This includes: training a ternary statistical language model using the aforementioned insurance terms lexicon to obtain an insurance terms statistical language model; and / or obtaining a deep learning model and pre-training it on labeled data with a sample size of tens of millions to obtain an insurance terms deep learning model.
[0014] This invention trains a ternary statistical language model using an insurance clause lexicon, resulting in an insurance clause statistical language model that is more suitable for the application scenarios of insurance clauses and can significantly improve the accuracy of the language model.
[0015] S3: Check for errors in the insurance policy name input;
[0016] S301: Perform word segmentation preprocessing on the name of the insurance input clause, compare it with the insurance clause thesaurus, and mark the no-matching item as an error word to achieve word-level error detection;
[0017] S302: Using the statistical language model of the insurance terms, score the name of the insurance input terms, calculate the average score of each character in the insurance input terms, and mark characters with low scores as typos to achieve character-level error detection;
[0018] Furthermore, step S3 also includes: S303: supplementing error detection with a deep learning model of insurance terms to reduce the false negative rate.
[0019] S4: Correct the name of the insurance policy entry;
[0020] The detected typos are recalled, and new characters are matched and replaced in the insurance clause thesaurus. Then, the multiple replacement options are scored by the insurance clause statistical language model, and the option with the highest score is selected to complete the text correction of the insurance input clause.
[0021] Furthermore, in step S4, when matching new text in the insurance terms thesaurus for replacement, at least one of the following methods is used:
[0022] a) Homophone correction method: Match homophones and replace them to complete word-level error correction;
[0023] b) Similar-looking character correction method: Match similar-looking characters and replace them to complete character-level error correction;
[0024] c) Method for correcting similar-sounding characters: Match similar-sounding characters and replace them to complete character-level error correction.
[0025] Furthermore, step S4 also includes: performing text correction using a deep learning model of insurance terms and outputting a correction result to achieve supplementary correction and improve the accuracy of correction.
[0026] This invention employs a deep learning model for insurance clauses, pre-trained on labeled data with tens of millions of samples, to supplement the error detection and correction of the statistical language model for insurance clauses, thereby further reducing the false negative rate and improving the accuracy of error correction.
[0027] Another embodiment of the present invention provides a text correction system for insurance input terms, comprising:
[0028] The system includes an insurance clause terminology preprocessing module, a statistical language model self-training module, an insurance input clause name error detection module, and an insurance input clause name error correction module.
[0029] The insurance terms terminology preprocessing module includes:
[0030] The insurance terms acquisition module is used to obtain the text of insurance terms;
[0031] The word segmentation module is used to segment the obtained insurance clause text into words to form an insurance clause text lexicon.
[0032] The Pinyin conversion module is used to convert words in the insurance clause text dictionary into Pinyin, forming an insurance clause Pinyin dictionary.
[0033] The thesaurus expansion module is used to acquire a general thesaurus and dictionary, merge them with the insurance clause text thesaurus and the insurance clause pinyin thesaurus, and remove punctuation marks and English letters and numbers to form a comprehensive insurance clause thesaurus;
[0034] The terminology filtering module is used to narrow down the expanded terminology to Chinese characters appearing in the insurance terms text, thus forming an insurance terms terminology.
[0035] The statistical language model self-training module is used to train a ternary statistical language model to obtain a statistical language model for insurance terms; and / or to obtain a deep learning model by pre-training labeled data with tens of millions of samples to obtain a deep learning model for insurance terms.
[0036] Furthermore, the deep learning model for the insurance terms is used to supplement error detection to reduce the false negative rate.
[0037] Furthermore, the deep learning model for the insurance terms is used for text correction and outputs a correction result to achieve supplementary correction, thereby improving the accuracy of the correction.
[0038] The insurance policy term name error detection module includes a thesaurus comparison module and a score calculation module; among them,
[0039] The terminology comparison module performs word segmentation preprocessing on the name of the insurance input clause and compares it with the insurance clause terminology database. Clauses that do not find a match are marked as erroneous words, thus achieving word-level error detection.
[0040] The score calculation module uses the statistical language model of the insurance terms to score the name of the insurance input terms, calculates the average score of each character in the insurance input terms, and marks characters with low scores as typos, thus realizing character-level error detection.
[0041] The insurance input clause name error correction module recalls detected typos, matches new characters in the insurance clause thesaurus for replacement, and then scores the multiple replacement options using the insurance clause statistical language model, selecting the option with the highest score to complete the text error correction of the insurance input clause.
[0042] The insurance input clause name error correction module includes at least one of the following: a homophone error correction module, a similar-looking character error correction module, and a similar-sounding character error correction module; wherein...
[0043] The homophone correction module is used to match homophones and replace them, thus completing word-level error correction.
[0044] The similar-looking character correction module is used to match similar-looking characters and replace them, thus completing character-level error correction.
[0045] The phonetic similarity correction module is used to match similar-sounding characters and replace them, thus completing character-level error correction.
[0046] Embodiments of the present invention also provide a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the steps of the above-described text correction method.
[0047] Embodiments of the present invention also provide a computer device, including a readable storage medium, a processor, and a computer program stored on the readable storage medium and executable on the processor, wherein the computer program, when executed by the processor, implements the steps of the above-described text correction method.
[0048] The technical solution of this invention is also applicable to other related scenarios, requiring only the expansion of the lexicon and retraining of the language model and deep learning model.
[0049] Beneficial effects
[0050] This invention improves the comprehensiveness and accuracy of the insurance clause lexicon by preprocessing it for various application scenarios. By training a ternary statistical language model using the insurance clause lexicon, the resulting statistical language model is more suitable for insurance clause applications, significantly improving the accuracy of the language model. Simultaneously, a deep learning model for insurance clauses, pre-trained on labeled data with tens of millions of samples, supplements the error detection and correction capabilities of the statistical language model, further reducing the false negative rate and improving the accuracy of error correction. In summary, this invention significantly improves text correction in insurance clause scenarios. Applied to intelligent insurance clause retrieval services, it can identify and correct incorrectly entered clause names, improving retrieval convenience. The technical solution of this invention is also applicable to other related scenarios, requiring only expansion of the lexicon and retraining of the language model and deep learning model. Of course, implementing any method or system of this invention does not necessarily require achieving all the advantages described above simultaneously. Attached Figure Description
[0051] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0052] Figure 1 A flowchart illustrating a text correction method for insurance input terms provided in an embodiment of the present invention;
[0053] Figure 2 A flowchart illustrating an error detection method for insurance clause names provided in an embodiment of the present invention;
[0054] Figure 3 This is a flowchart illustrating a method for correcting the name of insurance input clauses, as provided in an embodiment of the present invention.
[0055] Figure 4 This is a schematic diagram illustrating the composition of a text correction system for insurance input clauses provided in an embodiment of the present invention. Detailed Implementation
[0056] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art based on this application are within the scope of protection of the present invention.
[0057] To achieve the above objectives, embodiments of the present invention provide a text correction method, system, and related equipment for insurance input clauses. The text correction method for insurance input clauses will be described in detail below. The steps in the following method embodiments can be executed in a logical order; the step numbers or the order in which the steps are described do not limit the execution order of the steps.
[0058] Figure 1 This is a flowchart illustrating a text correction method for insurance input terms provided in an embodiment of the present invention. The method includes:
[0059] S1: Insurance clause terminology preprocessing.
[0060] S101: Obtain the insurance terms text, perform word segmentation preprocessing to form an insurance terms text lexicon, and convert the words in the insurance terms text lexicon into pinyin to form an insurance terms pinyin lexicon.
[0061] For example, insurance policy texts may contain long sentences. First, the insurance policy text is preprocessed using long_cut word segmentation. long_cut is an algorithm or tool for segmenting Chinese text, used to break continuous text into meaningful words (similar to tokenization in English). The set of words obtained after word segmentation constitutes the insurance policy text lexicon, which contains insurance-specific terms and keywords.
[0062] Furthermore, the words in the insurance terms text dictionary are converted into pinyin form, for example, "employer" is converted to "guzhu", "work injury" is converted to "gongshang", and "all risks" is converted to "yiqiexian", resulting in an insurance terms pinyin dictionary. This insurance terms pinyin dictionary can be used to provide more accurate results when processing pinyin input or pinyin retrieval.
[0063] S102: Obtain a general terminology and dictionary, merge them with the insurance clause text terminology and the insurance clause pinyin terminology, remove punctuation marks and English letters and numbers, and form a comprehensive insurance clause terminology.
[0064] For example, a general-purpose thesaurus is acquired, containing widely used vocabulary. This general-purpose thesaurus is then merged with existing thesauruses (insurance policy text thesaurus, insurance policy pinyin thesaurus) to expand its coverage, making it include both specialized insurance terminology and a broader range of commonly used vocabulary.
[0065] In one embodiment of the present invention, to further optimize the lexicon, a dictionary containing stroke order, Wubi input method, stroke count, and pronunciation information for tens of thousands of Chinese characters is obtained. This information is very useful for applications such as input methods that require retrieval based on the structure of Chinese characters. Insurance policy text may contain Chinese characters, English letters, Arabic numerals, Chinese punctuation, and English punctuation. Removing English letters, numbers, and punctuation marks from the text yields a comprehensive insurance policy lexicon, helping the model to focus more on the characteristics of the Chinese language, thereby improving the efficiency and accuracy of subsequent processing.
[0066] S103: Narrow the scope of the comprehensive insurance terms lexicon to Chinese characters appearing in the insurance terms text to form an insurance terms lexicon.
[0067] The final step in the insurance clause terminology preprocessing is to filter out the Chinese characters that actually appear in the insurance clause text from the comprehensive insurance clause terminology database, and delete Chinese characters that do not appear in the insurance clause text, as well as their stroke order, Wubi input method, number of strokes, and pronunciation information. This reduces redundancy in the terminology database and ensures that it better meets the actual needs of insurance clause text processing.
[0068] S2: Statistical language model self-training.
[0069] This includes: training a ternary statistical language model using the aforementioned insurance terms lexicon to obtain an insurance terms statistical language model; and / or obtaining a deep learning model and pre-training it on labeled data with a sample size of tens of millions to obtain an insurance terms deep learning model.
[0070] On one hand, a ternary (3-gram) statistical language model is trained using preprocessed insurance clause lexicon text data. By training the ternary statistical language model using the insurance clause lexicon, the resulting model is more suitable for the application scenarios of insurance clauses and can significantly improve the accuracy of the language model. For example, in natural language processing, an n-gram model predicts the probability distribution of the next word or tag by analyzing the frequency of co-occurrence of n consecutive words or tags in text. The training process can be performed in a Linux environment. After training, the model is saved as an Arps format KenLM language model file. KenLM is a high-efficiency language model library commonly used for natural language processing tasks. It can handle large amounts of text data and generate high-performance language models. It should be noted that the training process can also be performed on other operating systems. The method, apparatus, computer equipment, and storage medium of this invention are not limited to a specific operating system, as long as they meet the training requirements.
[0071] On the other hand, in one embodiment of the present invention, the open-source macbert4csc deep learning model can be obtained. macbert4csc is a pre-trained deep learning model based on the Transformer architecture, specifically designed for Chinese text. This model has been pre-trained on a large-scale corpus and is capable of understanding and generating language. On datasets with tens of millions of samples, this model can be pre-trained on labeled data to obtain a deep learning model for insurance clauses, making it better suited to the specific linguistic characteristics and needs of the insurance field, such as the understanding and analysis of insurance clause texts. Using the above method, both text correction and other natural language processing tasks will be more accurate and effective.
[0072] S3: Check for errors in the insurance policy name input.
[0073] Figure 2 A flowchart illustrating an error detection method for insurance clause names provided in an embodiment of the present invention includes:
[0074] S301: Perform word segmentation preprocessing on the name of the insurance input clause, compare it with the insurance clause thesaurus, and mark the no-matching item as an error word to achieve word-level error detection.
[0075] For example, when a clause name is entered, the system will perform word segmentation and compare the segmented words with the insurance clause thesaurus formed above. If a word does not find a match in the insurance clause thesaurus, then the word will be marked as an error word, thus completing word-level error detection.
[0076] S302: Using the statistical language model of the insurance terms, the name of the insurance input terms is scored, the average score of each character in the insurance input terms is calculated, and characters with low scores are marked as typos, thus achieving character-level error detection.
[0077] For example, the insurance clause statistical language model includes a score calculation module that scores the names of the input insurance clauses, using an n-gram insurance clause statistical language model to calculate the average score for each word. After inputting a clause name, if a word in the input clause name has a low score in the n-gram statistical language model, it is marked as a typo.
[0078] S303: Supplement error detection by using a deep learning model of insurance terms to reduce the false negative rate.
[0079] For example, a deep learning model could be the Macbert deep learning model. The Macbert model is a type of model that utilizes deep learning techniques, typically pre-trained on large amounts of text data to learn representations of words, sentences, and documents. It can capture complex features and patterns of language and is suitable for various natural language processing tasks, including text classification, information extraction, and error detection. Preferably, open-source pre-trained deep learning models can be used. Because these models have already been trained on large-scale corpora, there is no need to train the model from scratch, and they can be directly applied to specific error detection tasks.
[0080] In a specific embodiment, how is error detection performed on the input insurance clause name? Taking an insurance clause named "Employer's Liability Insurance" as an example, first, this name is input into the error detection system. In the lexicon comparison module, the system first segments the phrase "Employer's Liability Insurance," possibly into "Employer," "Liability," and "Insurance." Next, the system compares these words with an insurance clause lexicon and a general domain lexicon. If all words find a match, then at the word level, the clause name is error-free; otherwise, words without a match are marked as errors. Next, in the score calculation module, the system evaluates the score of each character in the input clause name "Employer's Liability Insurance" in an n-gram statistical language model. If a character has a low score, it is marked as a misspelling. Finally, a deep learning model for insurance clauses is used for supplementary error detection. By using these three methods in parallel, potential errors in insurance clause names can be identified more accurately, effectively reducing the missed detection rate and thus improving the accuracy and professionalism of the text.
[0081] Figure 3 A flowchart illustrating an insurance clause name correction method provided in this embodiment of the invention includes:
[0082] S4: Correct the name of the insurance policy clause.
[0083] The detected typos are recalled, and new characters are matched and replaced in the insurance clause thesaurus. Then, the multiple replacement options are scored by the insurance clause statistical language model, and the option with the highest score is selected to complete the text correction of the insurance input clause.
[0084] On the one hand, an unsupervised approach is used to recall candidate characters (detected misspellings).
[0085] Under normal circumstances, the unsupervised approach refers to a method of training a model without using labeled data. In the context of text error correction, the unsupervised approach means that there is no need for a dataset containing examples of correct and incorrect text pairs. Instead, certain rules or heuristics are directly applied to detect and correct errors.
[0086] In one embodiment of the present invention, an unsupervised text error correction method utilizes a pre-constructed word bank and dictionary to discover and correct errors in the text. Specifically, first, a word bank (possibly a word bank for insurance terms) and a dictionary (possibly including one or more dictionaries) are prepared in advance. The word bank contains accurate vocabulary in a specific domain (such as insurance terms), while the dictionary provides broader language knowledge, including the stroke order of Chinese characters, Wubi, strokes, pronunciation, etc.
[0087] Match the Chinese text of the insurance input clause name with the pre-prepared insurance clause word bank. In a specific embodiment, the matching methods include homophone, shape similarity, and sound similarity matching. These three matching methods can be used individually or in parallel. Here, "homophone" refers to characters with the same pronunciation but different meanings; "shape similarity" refers to characters with similar shapes that are easily confused; "sound similarity" refers to characters with similar pronunciations. These three types of characters are all common sources of typos. Next, form multiple options after replacing the detected typos, and score the multiple options after replacement through the insurance clause statistical language model, and select the option with a high score to complete the text error correction of the insurance input clause. For example, when a possible typo is detected in the text, the system will try to replace the typo with a matching character to generate multiple possible correct versions. For example, if the character "保" in the original text is miswritten as the homophone "宝", the system will replace it back with the correct "保" character. Among all the possible correct options, the system needs to evaluate which option is the most reasonable. This is usually accomplished by the insurance clause statistical language model considering the context information and the probability of the candidate words appearing in the given text.
[0088] The homophone error correction method matches homophones for replacement to complete word-level error correction.
[0089] Perform homophone error correction on the detected error words. Specifically, recall the detected error words and retrieve and replace homophones in the insurance clause word bank to complete word-level error correction.
[0090] The shape similarity error correction method matches shape-similar characters for replacement to complete character-level error correction.
[0091] The detected misspelled characters are corrected using similar-looking characters. Specifically, based on the characteristics of Chinese characters, the similarity between two characters can be calculated using stroke count, stroke order, Wubi stroke count, and four-corner encoding. In one embodiment, cosine similarity can be selected when choosing the algorithm for calculating similarity. In another embodiment, edit distance can also be selected. Furthermore, other algorithms capable of calculating similarity can be used as algorithms for this step without departing from the principles of this invention. Between cosine similarity and edit distance, cosine similarity is preferred in terms of speed and correction accuracy. In one embodiment, stroke order and Wubi stroke count are selected. For each detected misspelled character, the cosine similarity of stroke order and Wubi stroke count is calculated for each character in the character set, and the 20 characters with the highest similarity scores are selected for replacement. There are many possible replacements. For example, if three typos are detected, and each typo recalls 20 similar-looking characters, 20*20*20 results can be generated. To improve speed, while ensuring applicability to the scenario, it can be adjusted to correct only half of the typos. Preferably, the meta-Cartesian product method can be used to handle the above combination problem.
[0092] The method for correcting similar-sounding characters involves matching similar-sounding characters and replacing them, thus achieving character-level error correction.
[0093] The detected misspelled words are corrected using similar-sounding characters. Specifically, in the similarity correction method, the detected misspelled words are corrected by sound similarity, the pronunciation similarity is calculated, and the 20 characters with the highest similarity scores are selected for replacement. There are many possible replacements. For example, if the similarity correction detects 3 misspelled words, and each misspelled word recalls 20 similar-looking characters, 20*20*20 results can be generated. To improve speed, while ensuring applicability to the scenario, it can be adjusted to correct only half of the misspelled words. Preferably, the meta-Cartesian product method can be used to handle the above combination problem.
[0094] On the other hand, a supervised approach is adopted, using deep learning models for error correction.
[0095] Supervised methods refer to the approach of training a model using labeled data. In one embodiment of this invention, the supervised method involves using a pre-trained deep learning model of insurance terms, such as the Macbert4csc deep learning model, to perform text correction and output a corrected result. Specifically, when using the Macbert4csc deep learning model for error correction, when an erroneous text is input, the model can directly output a corrected text, without needing to generate multiple candidate options as in unsupervised methods. The Macbert4csc deep learning model is trained on labeled data. The labeled data consists of short sentences, and the typos in these sentences have been clearly marked, with correct replacements provided.
[0096] The annotated corpus is formed by annotating errors in short sentences, for example
[0097] [{"id":"A2-0003-1",
[0098] "original_text":"But I can't go because something's come up!"
[0099] "wrong_ids":
[17] ,
[0100] "correct_text":"But I can't go because I have something to do!"}]
[0101] To identify the more correct option in this scenario after replacement, a comprehensive judgment can be made using three scores: name matching degree, character similarity, and statistical language model score. The three scores are multiplied together, and the three with the highest scores are selected as the criteria for determining correctness.
[0102] This method can significantly improve the text correction effect in insurance clause scenarios. When applied to intelligent insurance clause retrieval services, it can identify and correct the clause names entered incorrectly by users, thereby improving the convenience of retrieval.
[0103] The technical solutions covered in this method are also applicable to other insurance-related scenarios; they can be applied simply by expanding the lexicon and retraining the language model and deep learning model.
[0104] Corresponding to the above method embodiments, this invention also provides a text correction system based on insurance clauses.
[0105] Figure 4 This is a schematic diagram illustrating the composition of a text correction system for insurance input terms provided in an embodiment of the present invention. The system includes:
[0106] Insurance terms and conditions terminology preprocessing module.
[0107] The insurance clause terminology preprocessing module includes an insurance clause acquisition module, used to acquire insurance clause text.
[0108] The word segmentation module is used to segment the acquired insurance clause text into words, forming an insurance clause text lexicon. For example, the insurance clause text may contain long sentences. First, the insurance clause text is preprocessed using long_cut word segmentation. The word set obtained after word segmentation constitutes a specific lexicon for insurance clauses, which includes insurance-specific terms and keywords.
[0109] The Pinyin conversion module converts words in the insurance policy text dictionary into Pinyin, creating a Pinyin dictionary for insurance policy terms. For example, it converts words in the insurance policy dictionary into Pinyin form, such as "employer" to "guzhu", "work injury" to "gongshang", and "all risks" to "yiqiexian". This Pinyin dictionary can be used to provide more accurate results when processing Pinyin input or Pinyin retrieval.
[0110] The dictionary expansion module is used to acquire a general-purpose dictionary and a general-purpose lexicon, and merge them with the insurance clause text dictionary and the insurance clause pinyin dictionary. After removing punctuation marks and alphanumeric characters, a comprehensive insurance clause dictionary is formed. In one embodiment, the dictionary expansion module is used to acquire a general-domain dictionary, which can be obtained from open-source channels. This general-domain dictionary may contain widely used vocabulary. This general-domain dictionary is then merged with the existing insurance clause text dictionary and insurance clause pinyin dictionary to expand the dictionary's coverage, making it include both insurance-related professional terminology and a wider range of commonly used vocabulary. To further optimize the dictionary, the dictionary expansion module can also acquire a dictionary containing stroke order, Wubi input method, stroke count, and pronunciation information for tens of thousands of Chinese characters. This information is very useful for applications such as input methods that require retrieval based on the structure of Chinese characters.
[0111] The dictionary filtering module is used to narrow down the expanded dictionary to include only Chinese characters appearing within the insurance terms text, thus forming an insurance terms dictionary. In one embodiment, the dictionary filtering module can be used to filter out the Chinese characters that actually appear in the insurance terms text from the dictionary, thereby reducing redundancy in the dictionary and ensuring that it better meets the actual needs of insurance terms text processing.
[0112] The statistical language model self-training module is used to train a ternary statistical language model to obtain a statistical language model for insurance terms; and to obtain a deep learning model by pre-training labeled data with tens of millions of samples to obtain a deep learning model for insurance terms.
[0113] On one hand, a ternary (3-gram) statistical language model is trained using preprocessed insurance clause lexicon text data. By training the ternary statistical language model using the insurance clause lexicon, the resulting model is more suitable for the application scenarios of insurance clauses and can significantly improve the accuracy of the language model. For example, in natural language processing, an n-gram model predicts the probability distribution of the next word or tag by analyzing the frequency of co-occurrence of n consecutive words or tags in text. The training process can be performed in a Linux environment. After training, the model is saved as an Arps format KenLM language model file. KenLM is a high-efficiency language model library commonly used for natural language processing tasks. It can handle large amounts of text data and generate high-performance language models. It should be noted that the training process can also be performed on other operating systems. The method, apparatus, computer equipment, and storage medium of this invention are not limited to a specific operating system, as long as they meet the training requirements.
[0114] On the other hand, in one embodiment of the present invention, the open-source macbert4csc deep learning model can be obtained. macbert4csc is a pre-trained deep learning model based on the Transformer architecture, specifically designed for Chinese text. This model has been pre-trained on a large-scale corpus and is capable of understanding and generating language. On datasets with tens of millions of samples, this model can be pre-trained on labeled data to obtain a deep learning model for insurance clauses, making it better suited to the specific linguistic characteristics and needs of the insurance field, such as the understanding and analysis of insurance clause texts. Using the above method, both text correction and other natural language processing tasks will be more accurate and effective.
[0115] The insurance policy name error detection module includes a thesaurus comparison module and a score calculation module.
[0116] The terminology comparison module preprocesses the names of the input insurance clauses by segmenting them into words, then compares them with the insurance clause terminology database. Clauses without a matching term are marked as errors, achieving word-level error detection. For example, when a clause name is entered, the system segments it into words and compares the segmented words with the aforementioned insurance clause terminology database. If a word does not find a matching term in the database, it is marked as an error, thus completing word-level error detection.
[0117] The score calculation module uses the statistical language model of the insurance terms to score the names of the input insurance terms, calculating the average score for each character in the input terms. Characters with low scores are marked as misspelled, achieving character-level error detection. For example, the statistical language model of the insurance terms includes a score calculation module that scores the names of the input insurance terms, using an n-gram statistical language model to calculate the average score for each character. After inputting a term name, if a character in the input term name has a low score in the n-gram statistical language model, it is marked as a misspelled character.
[0118] Supplemental error detection is performed using a deep learning model of insurance clauses to reduce the false negative rate.
[0119] For example, a deep learning model could be the Macbert deep learning model. The Macbert model is a type of model that utilizes deep learning techniques, typically pre-trained on large amounts of text data to learn representations of words, sentences, and documents. It can capture complex features and patterns of language and is suitable for various natural language processing tasks, including text classification, information extraction, and error detection. Preferably, open-source pre-trained deep learning models can be used. Because these models have already been trained on large-scale corpora, there is no need to train the model from scratch, and they can be directly applied to specific error detection tasks.
[0120] In a specific embodiment, how is error detection performed on the input insurance clause name? Taking an insurance clause named "Employer's Liability Insurance" as an example, first, this name is input into the error detection system. In the lexicon comparison module, the system first segments the phrase "Employer's Liability Insurance," possibly into "Employer," "Liability," and "Insurance." Next, the system compares these words with an insurance clause lexicon and a general domain lexicon. If all words find a match, then at the word level, the clause name is error-free; otherwise, words without a match are marked as errors. Next, in the score calculation module, the system evaluates the score of each character in the input clause name "Employer's Liability Insurance" in an n-gram statistical language model. If a character has a low score, it is marked as a misspelling. Finally, a deep learning model for insurance clauses is used for supplementary error detection. By using these three methods in parallel, potential errors in insurance clause names can be identified more accurately, effectively reducing the missed detection rate and thus improving the accuracy and professionalism of the text.
[0121] The insurance input clause name error correction module recalls the detected typos, matches new words in the insurance clause lexicon for replacement, and then scores multiple options after replacement through the insurance clause statistical language model, selects the option with a high score, and completes the text error correction of the insurance input clause.
[0122] The insurance input clause name error correction module can adopt an unsupervised method to recall candidate words (detected typos).
[0123] Generally, the unsupervised method means not using labeled data for model training. In the scenario of text error correction, the unsupervised method means that there is no need to have a dataset containing examples of correct and incorrect text corresponding to each other, but directly applying certain rules or heuristic methods to detect and correct errors.
[0124] In an embodiment of the present invention, an unsupervised text error correction system uses a pre-constructed lexicon and dictionary to discover and correct errors in the text. Specifically, first, a lexicon (possibly an insurance clause word lexicon) and a dictionary (possibly including one or more dictionaries) are prepared in advance. The lexicon contains accurate vocabulary in a specific field (such as insurance clauses), while the dictionary provides broader language knowledge, including the stroke order of Chinese characters, Wubi, strokes, pronunciation, etc.
[0125] Match the Chinese characters of the insurance input clause name with the pre-prepared insurance clause lexicon. In a specific embodiment, the matching methods include homophone, shape similarity, and sound similarity matching. These three matching methods can be used alone or in parallel. Here, "homophone" refers to characters with the same pronunciation but different meanings; "shape similarity" refers to characters with similar shapes that are easily confused; "sound similarity" refers to characters with similar pronunciations. These three types of characters are all common sources of typos. Next, form multiple options after replacing the detected typos, score the multiple options after replacement through the insurance clause statistical language model, select the option with a high score, and complete the text error correction of the insurance input clause. For example, when a possible typo is detected in the text, the system will try to replace the typo with a matched character to generate multiple possible correct versions. For example, if the character "保" in the original text is miswritten as the homophone "宝", the system will replace it back with the correct "保" character. Among all possible correct options, the system needs to evaluate which option is the most reasonable. This is usually accomplished by the insurance clause statistical language model considering context information and the probability of candidate words appearing in the given text.
[0126] The homophone error correction module matches homophones for replacement and completes word-level error correction.
[0127] The system performs homophone correction on detected erroneous words. Specifically, it recalls the detected erroneous words, searches for homophones in the insurance policy terminology database, and replaces them, thus completing word-level error correction.
[0128] The similar-looking character correction module matches similar-looking characters and replaces them, completing character-level error correction.
[0129] The detected misspelled characters are corrected using similar-looking characters. Specifically, based on the characteristics of Chinese characters, the similarity between two characters can be calculated using stroke count, stroke order, Wubi stroke count, and four-corner encoding. In one embodiment, cosine similarity can be selected when choosing the algorithm for calculating similarity. In another embodiment, edit distance can also be selected. Furthermore, other algorithms capable of calculating similarity can be used as algorithms for this step without departing from the principles of this invention. Between cosine similarity and edit distance, cosine similarity is preferred in terms of speed and correction accuracy. In one embodiment, stroke order and Wubi stroke count are selected. For each detected misspelled character, the cosine similarity of stroke order and Wubi stroke count is calculated for each character in the character set, and the 20 characters with the highest similarity scores are selected for replacement. There are many possible replacements. For example, if three typos are detected, and each typo recalls 20 similar-looking characters, 20*20*20 results can be generated. To improve speed, while ensuring applicability to the scenario, it can be adjusted to correct only half of the typos. Preferably, the meta-Cartesian product method can be used to handle the above combination problem.
[0130] The phonetic similarity correction module matches similar-sounding characters and replaces them, completing character-level error correction.
[0131] The detected misspellings are corrected using similar-sounding characters. Specifically, by calculating the pronunciation similarity, the 20 characters with the highest similarity scores are selected for replacement. There are many possible replacements. For example, if the pronunciation correction detects 3 misspellings, and each misspelling recalls 20 similar-looking characters, 20*20*20 results can be generated. To improve speed, while ensuring applicability to the scenario, this can be adjusted to correct only half of the misspellings. Preferably, the meta-Cartesian product method can be used to handle the above combination problem.
[0132] The insurance policy name error correction module can also adopt a supervised approach, using a deep learning model for error correction.
[0133] Supervised training refers to training a model using labeled data. In one embodiment of this invention, the supervised training method involves using a pre-trained deep learning model of insurance terms, such as the Macbert4csc deep learning model, to perform text correction and output a corrected result. Specifically, when using the Macbert4csc deep learning model for correction, when given erroneous text, the model can directly output a corrected text, without needing to generate multiple candidate options as in unsupervised methods. The Macbert4csc deep learning model is trained on labeled data. The labeled data consists of short sentences, and typos in these sentences have been explicitly marked, with correct replacements provided.
[0134] The annotated corpus is formed by annotating errors in short sentences, for example
[0135] [{"id":"A2-0003-1",
[0136] "original_text":"But I can't go because something's come up!"
[0137] "wrong_ids":
[17] ,
[0138] "correct_text":"But I can't go because I have something to do!"}]
[0139] To identify the more correct option in this scenario after replacement, a comprehensive judgment can be made using three scores: name matching degree, character similarity, and statistical language model score. The three scores are multiplied together, and the three with the highest scores are selected as the criteria for determining correctness.
[0140] This system can significantly improve the text correction effect in insurance clause scenarios. When applied to intelligent insurance clause retrieval services, it can identify and correct the clause names entered incorrectly by users, thereby improving the convenience of retrieval.
[0141] The technical solutions covered in this system are also applicable to other insurance-related scenarios; they can be applied simply by expanding the lexicon and retraining the language model and deep learning model.
[0142] Embodiments of the present invention also provide a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the steps of the above-described text correction method.
[0143] Embodiments of the present invention also provide a computer device, including a readable storage medium, a processor, and a computer program stored on the readable storage medium and executable on the processor, wherein the computer program, when executed by the processor, implements the steps of the above-described text correction method.
[0144] In the above embodiments, implementation can be achieved entirely or partially through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented entirely or partially in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of the present invention are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid state disk (SSD)).
[0145] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0146] The various embodiments in this specification are described in a related manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the system embodiments, client embodiments, server embodiments, computer-readable storage medium embodiments, and computer program product embodiments are basically similar to the method embodiments, and therefore the descriptions are relatively simple; relevant parts can be referred to the descriptions of the method embodiments.
[0147] The above description is merely a preferred embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention are included within the scope of protection of the present invention.
Claims
1. A method for text correction of insurance input clauses, characterized in that, Includes the following steps: S1: Insurance Clause Terminology Preprocessing; S101: Obtain the insurance terms text, perform word segmentation preprocessing to form an insurance terms text dictionary, and convert the words in the insurance terms text dictionary into pinyin to form an insurance terms pinyin dictionary; S102: Obtain a general terminology and dictionary, merge them with the insurance clause text terminology and the insurance clause pinyin terminology, remove punctuation marks and English letters and numbers, and form a comprehensive insurance clause terminology; S103: Narrow the scope of the comprehensive insurance clause terminology to Chinese characters appearing in the insurance clause text to form an insurance clause terminology; S2: Self-training of statistical language models; This includes: training a ternary statistical language model using the aforementioned insurance terms lexicon to obtain an insurance terms statistical language model; and / or obtaining a deep learning model and pre-training it on labeled data with a sample size of tens of millions to obtain an insurance terms deep learning model. S3: Check for errors in the insurance policy name input; S301: Perform word segmentation preprocessing on the name of the insurance input clause, compare it with the insurance clause thesaurus, and mark the no-matching item as an error word to achieve word-level error detection; S302: Using the statistical language model of the insurance terms, score the name of the insurance input terms, calculate the average score of each character in the insurance input terms, and mark characters with low scores as typos to achieve character-level error detection; S4: Correct the name of the insurance policy entry; The detected typos are recalled, and new characters are matched and replaced in the insurance clause thesaurus. Then, the multiple replacement options are scored by the insurance clause statistical language model, and the option with the highest score is selected to complete the text correction of the insurance input clause.
2. The method according to claim 1, characterized in that, Step S3 also includes: S303: Supplementing error detection with a deep learning model of insurance terms to reduce the false negative rate.
3. The method according to claim 1, characterized in that, In step S4, when matching new text in the insurance terms terminology database for replacement, at least one of the following methods is used: a) Homophone correction method: Match homophones and replace them to complete word-level error correction; b) Similar-looking character correction method: Match similar-looking characters and replace them to complete character-level error correction; c) Method for correcting similar-sounding characters: Match similar-sounding characters and replace them to complete character-level error correction.
4. The method according to claim 3, characterized in that, Step S4 also includes: performing text correction using a deep learning model of insurance terms and outputting a correction result to achieve supplementary correction and improve the accuracy of correction.
5. A text correction system for insurance input clauses, characterized in that, include: The system includes an insurance clause terminology preprocessing module, a statistical language model self-training module, an insurance input clause name error detection module, and an insurance input clause name error correction module. The insurance terms terminology preprocessing module includes: The insurance terms acquisition module is used to obtain the text of insurance terms; The word segmentation module is used to segment the obtained insurance clause text into words to form an insurance clause text lexicon. The Pinyin conversion module is used to convert words in the insurance clause text dictionary into Pinyin, forming an insurance clause Pinyin dictionary. The thesaurus expansion module is used to acquire a general thesaurus and dictionary, merge them with the insurance clause text thesaurus and the insurance clause pinyin thesaurus, and remove punctuation marks and English letters and numbers to form a comprehensive insurance clause thesaurus; The terminology filtering module is used to narrow down the expanded terminology to Chinese characters appearing in the insurance terms text, thus forming an insurance terms terminology. The statistical language model self-training module is used to train a ternary statistical language model to obtain a statistical language model for insurance terms; and / or to obtain a deep learning model by pre-training labeled data with tens of millions of samples to obtain a deep learning model for insurance terms. The insurance policy term name error detection module includes a thesaurus comparison module and a score calculation module; among them, The terminology comparison module performs word segmentation preprocessing on the name of the insurance input clause and compares it with the insurance clause terminology database. Clauses that do not find a match are marked as erroneous words, thus achieving word-level error detection. The score calculation module uses the statistical language model of the insurance terms to score the name of the insurance input terms, calculates the average score of each character in the insurance input terms, and marks characters with low scores as typos, thus realizing character-level error detection. The insurance input clause name error correction module recalls detected typos, matches new characters in the insurance clause thesaurus for replacement, and then scores the multiple replacement options using the insurance clause statistical language model, selecting the option with the highest score to complete the text error correction of the insurance input clause.
6. The system according to claim 5, characterized in that, The deep learning model for the insurance terms is used to supplement error detection and reduce the false negative rate.
7. The system according to claim 5, characterized in that, The insurance input clause name error correction module includes at least one of the following: a homophone error correction module, a similar-looking character error correction module, and a similar-sounding character error correction module; wherein... The homophone correction module is used to match homophones and replace them, thus completing word-level error correction. The similar-looking character correction module is used to match similar-looking characters and replace them, thus completing character-level error correction. The phonetic similarity correction module is used to match similar-sounding characters and replace them, thus completing character-level error correction.
8. The system according to claim 7, characterized in that, The deep learning model for the insurance terms is used for text correction and outputs a correction result to achieve supplementary correction and improve the accuracy of correction.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 4.
10. A computer device comprising a readable storage medium, a processor, and a computer program stored on the readable storage medium and executable on the processor, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 4.