Specific field spelling error correction corpus construction method and device based on confusion set

By constructing a domain-specific spelling error correction corpus method based on obfuscated sets, high-quality spelling error data are generated using speech recognition and network crawling technology, and combining pre-trained language models to adjust the attention mechanism, the problem of scarcity of spelling error correction data in specific fields is solved and the model's error correction performance is improved.

CN120387443APending Publication Date: 2025-07-29YUNNAN POWER GRID CO LTD ELECTRIC POWER RES INST
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510394227.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-31
Publication Date
2025-07-29

AI Technical Summary

Technical Problem

Traditional spelling error correction methods based on dictionary matching and rules are limited in specific fields, while machine learning-based methods require a large amount of labeled data and are difficult to obtain, resulting in limited application scope of spelling error correction technology in specific fields.

Method used

Use the speech recognition model to generate pseudo-data to build spelling error obfuscation sets, combine it with network crawlers to obtain monolingual corpus, adjust the attention mechanism through pre-training language models, enhance the vocabulary weight of obfuscation sets, iterative training to screen spelling error-correcting corpus, and finally fine-tune the model.

Benefits of technology

It effectively solves the problem of scarce spelling error correction data in specific fields, enhances the application effect of the model in specific fields, reduces the cost of manual data construction, and improves error correction performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120387443A_ABST
    Figure CN120387443A_ABST
Patent Text Reader

Abstract

The invention discloses a specific field spelling error correction corpus construction method and device based on a confusion set, and the method comprises the steps: recognizing the voice input of a specific field into a preliminary text result through a voice recognition model, and comparing the preliminary text result with a real label to obtain pseudo data; constructing a confusion set based on the pseudo data, sorting each group of words in the confusion set according to word frequencies, and reserving the first n words; obtaining a monolingual corpus in a specific field, and generating a spelling error correction corpus in combination with the confusion set; inputting the data into a pre-training language model for training, enhancing the weight of vocabularies in a confusion set by adjusting the attention mechanism of the model, and screening spelling error correction corpora which are distributed in a preset difference with spelling errors of a real corpus data set through iterative training; and performing fine adjustment on the model by using the screened spelling error correction corpus until a final spelling error correction model is obtained. According to the method, high-quality spelling error data can be generated by fully utilizing knowledge in a specific field and characteristics of a confusion set, so that the performance of a spelling error correction model in the specific field is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of natural language processing, and in particular, to a method and device for constructing a domain-specific spelling correction corpus based on a confusion set. Background Art

[0002] Domain-specific spelling correction technology plays an important role in the field of natural language processing. Its main task is to identify and correct spelling mistakes in text. In application scenarios in specific fields such as medicine and law, the importance of spelling correction is particularly prominent. Texts in these fields often contain a large number of professional terms and abbreviations, and follow specific term usage habits and spelling rules, which pose additional challenges to spelling correction. Traditional dictionary-matching and rule-based methods have limited effectiveness in processing texts in these fields, while machine learning-based methods, although having made significant progress, usually require a large amount of labeled data, which is difficult to obtain in specific domains, thus limiting the application scope of these methods. Summary of the Invention

[0003] Based on this, it is necessary to address the above problems and propose a method and device for constructing a domain-specific spelling correction corpus based on a confusion set.

[0004] An embodiment of the present application provides a method for constructing a domain-specific spelling correction corpus based on a confusion set, the method comprising:

[0005] Using a speech recognition model to recognize a domain-specific speech input as a preliminary text result, comparing the preliminary text result with a true label to obtain pseudo data containing spelling mistakes;

[0006] Constructing a spelling mistake confusion set for the domain based on the pseudo data, sorting each group of words in the confusion set by word frequency and retaining the top n high-frequency words;

[0007] Using web crawler technology to obtain a monolingual corpus for the domain, combining the confusion set and the monolingual corpus to generate a spelling correction corpus containing various spelling mistakes;

[0008] Inputting the spelling correction corpus into a pre-trained language model for training. During the training process, by adjusting the attention mechanism of the model, enhancing the weights of the words in the confusion set, and through iterative training, screening out a domain-specific spelling correction corpus whose spelling mistake distribution is within a preset difference from the true corpus dataset;

[0009] Using the screened domain-specific spelling correction corpus to fine-tune the pre-trained language model until a final spelling correction model is obtained.

[0010] In some embodiments, using a speech recognition model to recognize a speech input in a specific domain as a preliminary text result, and comparing the preliminary text result with a true label to obtain pseudo data containing spelling errors, includes:

[0011] Collect text data in the specific domain and convert it into speech data;

[0012] Use the speech recognition model to recognize the speech data to obtain the preliminary text result;

[0013] Compare the preliminary text result with the text data to construct a spelling correction data set, and compare the correction data set with the true label to obtain the pseudo data containing spelling errors.

[0014] In some embodiments, constructing the spelling error confusion set in the specific domain based on the pseudo data, and sorting each group of words in the confusion set by word frequency and retaining the top n high-frequency words, includes:

[0015] Identify the spelling errors in the pseudo data, pair the wrong words with the correct words corresponding to the wrong words, and construct the spelling error confusion set in the specific domain;

[0016] Perform word frequency statistics on each group of words in the spelling error confusion set in the specific domain, and sort all groups of words according to word frequency;

[0017] Retain the top n high-frequency words from all the sorted groups of words.

[0018] In some embodiments, using web crawler technology to obtain the monolingual corpus in the specific domain, includes:

[0019] Use web crawler technology to obtain online text resources in the specific domain, and the online text resources are real corpus data sets;

[0020] Segment the text data in the collected online text resources and divide it into independent sentences;

[0021] Preprocess the divided independent sentences;

[0022] Delete the sentences containing website information or unrecognizable encoding information from the preprocessed sentences, and standardize the text format to obtain the monolingual corpus.

[0023] In some embodiments, combining the confusion set and the monolingual corpus to generate a spelling correction corpus containing various spelling errors, includes:

[0024] Take each sentence in the monolingual corpus as the source sentence, and screen all the words contained in the source sentence that have appeared in the confusion set as the corresponding candidate replacement words;

[0025] Randomly select several words from the source sentence, and screen the highest-frequency words with the same spelling error type as each randomly selected word in the confusion set for replacement and modification to generate sentences containing multiple spelling errors;

[0026] Pair the modified sentence with the source sentence to form the spelling correction corpus containing multiple spelling errors.

[0027] In some embodiments, the pre-trained language model adopts a Transformer architecture, and a weight matrix is added to the attention layer of the pre-trained language model to adjust the attention weights of the words in the confusion set.

[0028] In some embodiments, the training objective of the pre-trained language model is to minimize the cross-entropy loss function, which is used to evaluate the difference between the error distribution of the prediction dataset and the error distribution of the real dataset;

[0029] The specific form of the minimizing cross-entropy loss function is θ is the trainable model parameter, x is the source sentence, y = {y1, y2,..., y n} is the correct sentence with n words, and y <t = {y1, y2,..., y t-1} is the word visible at time step t.

[0030] The embodiment of the present application also provides a device for constructing a domain-specific spelling correction corpus based on a confusion set. The device for constructing a domain-specific spelling correction corpus based on a confusion set includes:

[0031] A pseudo-data generation module, configured to use a speech recognition model to recognize a specific-domain speech input as a preliminary text result, compare the preliminary text result with a real label, and obtain pseudo-data containing spelling errors;

[0032] A confusion set construction module, configured to construct the domain-specific spelling error confusion set based on the pseudo-data, sort each group of words in the confusion set by word frequency, and retain the top n high-frequency words;

[0033] An error correction corpus generation module, configured to use web crawler technology to obtain the monolingual corpus of the specific domain, and combine the confusion set and the monolingual corpus to generate a spelling correction corpus containing multiple spelling errors;

[0034] The error correction corpus screening module is used to input the spelling error correction corpus into a pre-trained language model for training. During the training process, by adjusting the attention mechanism of the model, the weights of the words in the confusion set are enhanced, and through iterative training, a spelling error correction corpus in a specific domain whose spelling error distribution is within a preset difference from the real corpus dataset is screened out;

[0035] The model training module is used to fine-tune the pre-trained language model using the screened spelling error correction corpus in a specific domain until the final spelling error correction model is obtained.

[0036] An embodiment of the present application also provides a computer device, including a memory and a processor. The memory stores a computer program. When the computer program is executed by the processor, the processor performs the following steps:

[0037] Using a speech recognition model to recognize the speech input in a specific domain as a preliminary text result, comparing the preliminary text result with the real label to obtain pseudo-data containing spelling errors;

[0038] Based on the pseudo-data, constructing a spelling error confusion set in the specific domain, sorting each group of words in the confusion set by word frequency and retaining the first n high-frequency words;

[0039] Using web crawler technology to obtain the monolingual corpus in the specific domain, combining the confusion set and the monolingual corpus to generate a spelling error correction corpus containing various spelling errors;

[0040] Inputting the spelling error correction corpus into a pre-trained language model for training. During the training process, by adjusting the attention mechanism of the model, the weights of the words in the confusion set are enhanced, and through iterative training, a spelling error correction corpus in a specific domain whose spelling error distribution is within a preset difference from the real corpus dataset is screened out;

[0041] Using the screened spelling error correction corpus in a specific domain to fine-tune the pre-trained language model until the final spelling error correction model is obtained.

[0042] An embodiment of the present application also provides a computer-readable storage medium, storing a computer program. When the computer program is executed by a processor, the processor performs the following steps:

[0043] Using a speech recognition model to recognize the speech input in a specific domain as a preliminary text result, comparing the preliminary text result with the real label to obtain pseudo-data containing spelling errors;

[0044] Based on the pseudo-data, constructing a spelling error confusion set in the specific domain, sorting each group of words in the confusion set by word frequency and retaining the first n high-frequency words;

[0045] Use web crawler technology to obtain the monolingual corpus in the specific field, and combine the confusion set and the monolingual corpus to generate a spelling correction corpus containing various spelling mistakes.

[0046] Input the spelling correction corpus into a pre-trained language model for training. During the training process, by adjusting the attention mechanism of the model, enhance the weights of the words in the confusion set, and through iterative training, screen out the spelling correction corpus in the specific field whose spelling mistakes distribution is within a preset difference from the real corpus dataset.

[0047] Use the screened spelling correction corpus in the specific field to fine-tune the pre-trained language model until the final spelling correction model is obtained.

[0048] Adopting the embodiments of the present application has the following beneficial effects:

[0049] In the method for constructing a spelling correction corpus in a specific field based on a confusion set provided by the embodiments of the present application, by using an existing speech recognition model to generate preliminary speech recognition results, combining real labels and spelling correction characteristics, post-processing these results to simulate spelling mistakes in real scenarios. Then, based on these pseudo-data, construct a spelling mistake confusion set, and sort each group of words in the confusion set according to word frequency, retaining the top n high-frequency words to ensure the effectiveness of the confusion set. Then, use web crawler technology to obtain the monolingual corpus in the specific field, and screen and preprocess it to ensure the quality of the data. Design a spelling correction data algorithm constructed using the confusion set, generate a correction corpus containing diverse spelling mistakes by integrating the monolingual data obtained by the crawler. On this basis, we adopt a pre-trained language model, enhance the model's ability to correct spelling mistakes by integrating the knowledge of the confusion set, in the attention mechanism of the model, enhance the attention to the words in the confusion set, so that the model can more effectively identify and correct these spelling mistakes. Through repeated iterative training, screen out the spelling correction corpus in the specific field whose error distribution is more similar to that in the real dataset. The present invention effectively utilizes the confusion set and the pre-trained language model, solves the problem of scarce spelling correction data in the specific field, and enhances the application effect of the model in the specific field. Description of the Drawings

[0050] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0051] Among them:

[0052] Figure 1Schematic flowchart of a method for constructing a domain-specific spelling correction corpus based on a confusion set in an embodiment;

[0053] Figure 2 Structure diagram of a device for constructing a domain-specific spelling correction corpus based on a confusion set in an embodiment;

[0054] Figure 3 Schematic structural diagram of a computer device in an embodiment;

[0055] Figure 4 Schematic structural diagram of a computer-readable storage medium in an embodiment. Detailed implementation manners

[0056] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.

[0057] In an embodiment of the present application, a method for constructing a domain-specific spelling correction corpus based on a confusion set is provided. Please refer to Figure 1 , Figure 1 Schematic flowchart of a method for constructing a domain-specific spelling correction corpus based on a confusion set in an embodiment; the method for constructing a domain-specific spelling correction corpus based on a confusion set includes steps S1 to S5.

[0058] Step S1, using a speech recognition model to recognize the speech input in a specific domain as a preliminary text result, and comparing the preliminary text result with a true label to obtain pseudo data containing spelling errors;

[0059] In some implementation manners, the step of using a speech recognition model to recognize the speech input in a specific domain as a preliminary text result, and comparing the preliminary text result with a true label to obtain pseudo data containing spelling errors includes:

[0060] Collecting the text data in the specific domain and converting it into speech data;

[0061] Using the speech recognition model to recognize the speech data to obtain the preliminary text result;

[0062] Comparing the preliminary text result with the text data to construct a spelling correction data set, and comparing the correction data set with the true label to obtain the pseudo data containing spelling errors.

[0063] Specifically, first, use the existing speech recognition model to recognize the speech input in a specific domain to obtain a preliminary recognition result. Then, combine the characteristics of spelling correction, compare the recognition result with the true label, and perform post-processing to simulate spelling mistakes in the real scenario. For example, collect the plain text data in a specific domain and use tools such as ttskit to convert it into speech data; use the existing speech recognition model to convert the speech data into text data and form a spelling correction data set with the original text; perform post-processing on the data set, such as deleting sentence pairs with unequal lengths, screening out sentence pairs with large differences using pypinyin and distance editing algorithms, removing duplicates, denoising, etc. to ensure that the data conforms to the real scenario.

[0064] Step S2, construct a spelling error confusion set for the specific domain based on the pseudo data, sort each group of words in the confusion set according to the word frequency, and retain the first n high-frequency words.

[0065] In some embodiments, the constructing a spelling error confusion set for the specific domain based on the pseudo data, sorting each group of words in the confusion set according to the word frequency, and retaining the first n high-frequency words includes:

[0066] Identify the spelling mistakes in the pseudo data, pair the wrong words with the correct words corresponding to the wrong words, and construct a spelling error confusion set for the specific domain.

[0067] Perform word frequency statistics on each phrase in the spelling error confusion set for the specific domain, and sort all phrases according to the word frequency.

[0068] Retain the first n high-frequency words from all the sorted phrases.

[0069] Specifically, use the pseudo data generated in step S1 to construct a spelling error confusion set, sort each group of words in the confusion set according to the word frequency, and retain the first n high-frequency words to ensure the effectiveness of the confusion set.

[0070] First, extract all words from the processed text data, that is, the pseudo data, to generate a vocabulary list, which should include the terms and common words in this domain.

[0071] Secondly, define the spelling error rules:

[0072] Homophone error: For example, misspelling "tomorrow" as "mingtian".

[0073] Similar-looking character error: For example, misspelling "gong" as "wu".

[0074] Pinyin input method error: For example, misspelling "information" as "xinxi".

[0075] Typo: For example, misspelling "data" as "shuju".

[0076] Then, according to the above rules, generate a set of misspelled forms for each word in the vocabulary to form a preliminary confusion set; perform word frequency statistics on the preprocessed text data to calculate the frequency of each word appearing in the text. For each misspelled word in the confusion set, calculate its frequency of appearance in the text. Sort the misspelled forms of each word according to the word frequency, and preferentially retain the misspelled words with higher frequencies of appearance. To ensure the effectiveness of the confusion set, only retain the top n words with the highest word frequencies in each group of misspelled words.

[0077] Step S3: Use web crawler technology to obtain the monolingual corpus of the specific domain. Combine the confusion set and the monolingual corpus to generate a spelling correction corpus containing multiple spelling mistakes.

[0078] In some embodiments, use web crawler technology to obtain the online text resources of the specific domain, and the online text resources are real corpus datasets.

[0079] Segment the text data in the collected online text resources and divide it into independent sentences.

[0080] Preprocess the divided independent sentences.

[0081] Delete the sentences containing website information or unrecognizable encoding information from the preprocessed sentences, and standardize the text format to obtain a monolingual corpus.

[0082] Specifically, use web crawler technology to obtain the monolingual corpus of the specific domain from the specific domain, and screen and preprocess these crawled data to ensure the quality and relevance of the data. The preprocessing steps include removing noise, filtering irrelevant content, and standardizing the text format, etc.

[0083] First, use web crawlers to collect online texts in multiple fields such as electricity, agriculture, news, and finance, and collect industry standards, industry guidelines, etc. of the specific domain, so as to obtain a large-scale unlabeled data of the specific domain.

[0084] Secondly, preprocess these sentences, such as removing non-alphabetic symbols and common abbreviations, deleting special symbols, deleting short and useless texts, and removing sentences that are too long or too short to conform to the real scenario distribution.

[0085] Finally, clean the preprocessed sentences, remove the sentences containing website information and unrecognizable encoding, and what is finally obtained is a large-scale and standardized monolingual corpus of the specific domain.

[0086] In some embodiments, the combination of the confusion set and the monolingual corpus to generate a spelling correction corpus containing multiple spelling mistakes includes:

[0087] Take each sentence in the monolingual corpus as the source sentence, and screen all the words that appear in the confusion set included in the source sentence as the corresponding candidate replacement words;

[0088] Randomly select several words from the source sentence, and screen the highest-frequency words with the same spelling error type as each randomly selected word in the confusion set for replacement and modification to generate sentences containing multiple spelling errors;

[0089] Pair the modified sentence with the source sentence to form the spelling correction corpus containing multiple spelling errors.

[0090] Specifically, design an algorithm for constructing spelling correction data using a confusion set, combine with monolingual data to generate an error correction corpus covering multiple spelling errors. Using the constructed confusion set and pseudo-data, high-quality spelling correction data can be generated. For example, check each word in the sentence s = {w1, w2, w3, …, w m}, find all the words that appear in the confusion set as candidate replacement words, and then randomly select n words (1 < n < m) from the original sentence, and replace them according to the highest-frequency words of the same type of error in the confusion set. For example, for the sentence "The development of information science and technology", if the wrong form of "information" in the confusion set is "new west" and the wrong form of "development" is "departure station", then after replacement, it may generate "New west science and technology departure station".

[0091] Step S4, input the spelling correction corpus into a pre-trained language model for training. During the training process, by adjusting the attention mechanism of the model, enhance the weight of the words in the confusion set, and through iterative training, screen out the spelling correction corpus of a specific domain whose spelling error distribution is within a preset difference from the spelling error distribution of the real corpus dataset;

[0092] In some embodiments, the pre-trained language model adopts a Transformer architecture, and a weight matrix is added to the attention layer of the pre-trained language model to adjust the attention weights of the words in the confusion set.

[0093] In some embodiments, the training objective of the pre-trained language model is to minimize the cross-entropy loss function, which is used to evaluate the difference between the error distribution of the prediction dataset and the error distribution of the real dataset;

[0094] The minimization of the cross-entropy loss function is specifically θ is the trainable model parameter, x is the source sentence, y = {y1, y2, …, y n} is the correct sentence with n words, y <t = {y1, y2, …, y t-1}{are the words visible at time step t.

[0095] Specifically, spelling correction can be regarded as a sequence-to-sequence task, and its model structure usually adopts the popular Transformer architecture. The Transformer consists of two main components: an encoder and a decoder. The encoder uses a multi-head attention mechanism to capture the context representation of each word in the source sentence, while the decoder has a similar structure, with an added masked multi-head self-attention module to more effectively model the generated word information. In this application, an additional weight matrix is added to the attention layer of the model to adjust the attention weights of the words in the confusion set. During each forward pass, when the input sequence contains words in the confusion set, the attention weights of these words are dynamically adjusted. The specific operation is to boost the weights of these words by looking up the words in the confusion set in the input sequence when calculating the attention distribution. In this way, when calculating the key-value pairs, the keys corresponding to the words in the confusion set will be given higher weights, making the model more likely to focus on these words when generating queries.

[0096] During the training process of the spelling correction task, the goal of the pre-trained language model is to minimize the cross-entropy loss function. The cross-entropy loss function is a commonly used loss function for evaluating the difference between the predicted distribution and the true distribution. In spelling correction, we hope that the corrected results output by the model are as close as possible to the actual correct spellings.

[0097]

[0098]

[0099] θ are the trainable model parameters, x is the source sentence, y = {y1, y2,..., y n} is the correct sentence with n words, y <t = {y1, y2,..., y t-1} are the words visible at time step t; perform iterative training on the constructed domain-specific error correction corpus, and further select the high-score data from it. These screened data will be added to the final domain-specific spelling correction corpus.

[0100] Step S5, use the screened domain-specific spelling correction corpus to fine-tune the pre-trained language model until the final spelling correction model is obtained.

[0101] Adopting the technical solution of this embodiment can solve the problem of insufficient spelling correction corpus in a specific field. The technical solution of this embodiment obtains spelling correction data by using a speech recognition model and constructs a confusion set of a certain scale with it. Compared with the traditional method, this application designs a method for constructing a specific field spelling correction corpus based on the confusion set, which greatly reduces the manpower and time costs of manual data construction. By adopting the method of iterative verification, this application can screen out high-quality error correction pseudo-data and use it in the model pre-training stage, greatly improving the error correction performance of the model.

[0102] The following combines specific embodiments to further verify and explain the method for constructing a specific field spelling correction corpus based on the confusion set of this application, and verify whether the specific field spelling correction corpus based on the confusion set can improve the error correction performance of the model.

[0103] This application selects an existing model to construct a certain scale of pseudo-data as the experimental data of the present invention, and at the same time uses the processed specific field corpus crawled from the Internet as the unlabeled data of the present invention. The pseudo-data constructed by this method will be used in the pre-training stage of the model.

[0104] The commonly used accuracy (Precision), recall (Recall), and F-value (F-measure) are used as the evaluation indicators of the specific field spelling correction model. The specific calculation methods are as follows:

[0105]

[0106]

[0107] In order to verify the effect of the method proposed in this application on improving the performance of the specific field spelling correction model, this application selects a mainstream specific field spelling correction model as the benchmark model. In this verification, this application uses the open-source bart-large-chinese model for pre-training and fine-tuning, and selects NaSGEC (including three specific fields) as the test set. The corpus scale involved is shown in Table 1:

[0108] Among them, the basic corpus for generating homophonic error corpus comes from HSK and Lang8, and the basic corpus for generating shape-similar error corpus comes from three fields: Chinese social media, publicly available Chinese papers on the Internet, and Chinese examinations.

[0109] Table 1: Corpus sources and corpus scales involved in the verification part

[0110]

[0111] The final experimental results of the model on the error corpus are shown in Table 2 below.

[0112] Table 2: Model performance on the training data of the confusion set, where C-A is used as the pre-training phase

[0113]

[0114] The final experimental results of the model on the homograph error corpus are shown in Table 2. The experimental results show that by training the model with the corpus generated based on the confusion set, that is, the final spelling correction model, the error correction ability of the baseline model can be improved. The method of this application also shows good results in the experiment, proving its effectiveness. In addition, the experiment also found that combining the homograph error and homophone error corpus can further enhance the generalization ability of the error correction model, which indicates that improving the quality of the corpus helps the model better capture the common knowledge between different languages and make more full use of the knowledge of the language model.

[0115] In the embodiment of this application, a device for constructing a specific domain spelling correction corpus based on a confusion set is provided. Please refer to Figure 2 , Figure 2 which is a structural diagram of the device for constructing a specific domain spelling correction corpus based on a confusion set in an embodiment. The device for constructing a specific domain spelling correction corpus based on a confusion set includes: a pseudo-data generation module 201, a confusion set construction module 202, an error correction corpus generation module 203, an error correction corpus screening module 204, and a model training module 205;

[0116] Among them, the pseudo-data generation module 201 is used to use a speech recognition model to recognize the speech input of a specific domain as a preliminary text result, compare the preliminary text result with the true label, and obtain pseudo-data containing spelling errors;

[0117] The confusion set construction module 202 is used to construct the spelling error confusion set of the specific domain based on the pseudo-data, sort each group of words in the confusion set according to the word frequency, and retain the top n high-frequency words;

[0118] The error correction corpus generation module 203 is used to use web crawler technology to obtain the monolingual corpus of the specific domain, and combine the confusion set and the monolingual corpus to generate a spelling correction corpus containing various spelling errors;

[0119] The error correction corpus screening module 204 is used to input the spelling correction corpus into a pre-trained language model for training. During the training process, by adjusting the attention mechanism of the model, the weight of the words in the confusion set is enhanced, and through iterative training, a specific domain spelling correction corpus whose spelling error distribution is within a preset difference from the true corpus dataset is screened out;

[0120] The model training module 205 is used to fine-tune the pre-trained language model using the filtered spelling correction corpus in a specific domain until the final spelling correction model is obtained.

[0121] In some embodiments, the pseudo-data generation module 201 is further configured to collect the text data in the specific domain and convert it into speech data;

[0122] Use the speech recognition model to recognize the speech data to obtain the preliminary text result;

[0123] Compare the preliminary text result with the text data to construct a spelling correction data set, and compare the correction data set with the true label to obtain the pseudo-data containing spelling errors.

[0124] In some embodiments, the confusion set construction module 202 is further configured to identify the spelling errors in the pseudo-data, pair the wrong words with the correct words corresponding to the wrong words, and construct the spelling error confusion set in the specific domain;

[0125] Perform word frequency statistics on each phrase in the spelling error confusion set in the specific domain, and sort all phrases according to the word frequency;

[0126] Retain the top n high-frequency words from all the sorted phrases.

[0127] In some embodiments, the error correction corpus generation module 203 is further configured to use web crawler technology to obtain the online text resources in the specific domain, and the online text resources are real corpus data sets;

[0128] Segment the text data in the collected online text resources and divide it into independent sentences;

[0129] Preprocess the divided independent sentences;

[0130] Delete the sentences containing website information or unrecognizable encoding information from the preprocessed sentences, and standardize the text format to obtain a monolingual corpus.

[0131] In some embodiments, the error correction corpus screening module 204 is further configured to use each sentence in the monolingual corpus as a source sentence, and screen all the words that appear in the confusion set contained in the source sentence as the corresponding candidate replacement words;

[0132] Randomly select several words from the source sentence, and screen the highest frequency words with the same spelling error type as each randomly selected word in the confusion set for replacement and modification to generate sentences containing multiple spelling errors;

[0133] Pair the modified sentence with the source sentence to form the spelling correction corpus containing multiple spelling mistakes.

[0134] In some embodiments, the model training module 205 is further configured to determine that the pre-trained language model adopts a Transformer architecture, and add a weight matrix to the attention layer of the pre-trained language model to adjust the attention weights of the words in the confusion set.

[0135] In some embodiments, the model training module 205 is further configured to determine that the training objective of the pre-trained language model is to minimize the cross-entropy loss function, which is used to evaluate the difference between the error distribution of the prediction data set and the error distribution of the true data set.

[0136] The specific form of the cross-entropy loss function is θ are trainable model parameters, x is the source sentence, y = {y1, y2, …, y n} is the correct sentence with n words, and y <t = {y1, y2, …, y t-1} are the words visible at time step t.

[0137] For other details of the implementation of each module in the apparatus for constructing a domain-specific spelling correction corpus based on a confusion set, reference may be made to the description in the above-provided method for constructing a domain-specific spelling correction corpus based on a confusion set, which will not be elaborated here.

[0138] In an embodiment of the present application, a computer device is provided. Please refer to Figure 3 , Figure 3 which is a schematic structural diagram of a computer device in an embodiment. The device includes a memory 301 and a processor 302. The memory 301 stores a computer program. When the computer program is executed by the processor 302, the processor 302 is caused to perform the following steps:

[0139] Use a speech recognition model to recognize the speech input in a specific domain as a preliminary text result, compare the preliminary text result with the true label to obtain pseudo data containing spelling mistakes.

[0140] Construct the domain-specific spelling mistake confusion set based on the pseudo data, sort each group of words in the confusion set according to word frequency, and retain the top n high-frequency words.

[0141] Use web crawler technology to obtain the monolingual corpus in the specific domain, and combine the confusion set and the monolingual corpus to generate a spelling correction corpus containing multiple spelling mistakes.

[0142] Input the spelling correction corpus into a pre-trained language model for training. During the training process, by adjusting the attention mechanism of the model, enhance the weights of the words in the confusion set, and through iterative training, filter out a spelling correction corpus for a specific domain whose spelling errors are within a preset difference from the real corpus dataset.

[0143] Use the filtered spelling correction corpus for the specific domain to fine-tune the pre-trained language model until the final spelling correction model is obtained.

[0144] Among them, the processor 302 can also be called a CPU (Central Processing Unit), and the processor 302 may be an integrated circuit chip with signal processing capabilities; the processor 302 can also be a general-purpose processor, DSP (Digital Signal Process), ASIC (Application Specific Integrated Circuit), FPGA (Field Programmable Gata Array), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. Among them, the general-purpose processor can be a microprocessor or the processor 302 can also be any conventional processor, etc.

[0145] In an embodiment of the present application, a computer-readable storage medium is provided. Please refer to Figure 4 , Figure 4 FIG. is a schematic structural diagram of a computer-readable storage medium in an embodiment. A readable computer program 401 is stored on the storage medium; among them, the computer program 401 can be stored in the above storage medium in the form of a software product, including several instructions to enable a computer device (which can be a personal computer, a service machine, or a network device, etc.) or a processor to perform the following steps:

[0146] Use a speech recognition model to recognize the speech input in a specific domain as a preliminary text result, compare the preliminary text result with the real label to obtain pseudo data containing spelling errors;

[0147] Based on the pseudo data, construct a spelling error confusion set for the specific domain, sort each group of words in the confusion set according to word frequency, and retain the top n high-frequency words;

[0148] Use web crawler technology to obtain the monolingual corpus for the specific domain, and combine the confusion set and the monolingual corpus to generate a spelling correction corpus containing various spelling errors;

[0149] Input the spelling correction corpus into a pre-trained language model for training. During the training process, by adjusting the attention mechanism of the model, enhance the weights of the words in the confusion set, and through iterative training, screen out a spelling correction corpus for a specific domain whose spelling errors are within a preset difference from the real corpus dataset.

[0150] Use the screened spelling correction corpus for a specific domain to fine-tune the pre-trained language model until the final spelling correction model is obtained.

[0151] The aforementioned storage medium includes: various media such as USB flash drives, external hard drives, magnetic disks or optical discs, ROM (Read-Only Memory), RAM (Random Access Memory), etc. that can store program codes, or terminal devices such as computers, servers, mobile phones, tablets, etc.

[0152] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The program can be stored in a non-volatile computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the various embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memories (ROMs), programmable ROMs (PROMs), electrically programmable ROMs (EPROMs), electrically erasable programmable ROMs (EEPROMs), or flash memories. Volatile memories can include random access memories (RAMs) or external cache memories. By way of illustration and not limitation, RAMs are available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0153] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.

[0154] The above-described embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all fall within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the appended claims.

Claims

1. A method for constructing a domain-specific spelling correction corpus based on a confusion set, characterized in that, Including: Using a speech recognition model to recognize the speech input in a specific field as a preliminary text result, comparing the preliminary text result with the true label to obtain pseudo data containing spelling mistakes; Constructing a spelling mistake confusion set for the specific field based on the pseudo data, sorting each group of words in the confusion set by word frequency and retaining the top n high-frequency words; Using web crawler technology to obtain the monolingual corpus of the specific field, combining the confusion set and the monolingual corpus to generate a spelling correction corpus containing various spelling mistakes; Inputting the spelling correction corpus into a pre-trained language model for training. During the training process, by adjusting the attention mechanism of the model, enhancing the weight of the words in the confusion set, and through iterative training, screening out a spelling correction corpus for the specific field whose spelling mistake distribution is within a preset difference from the true corpus dataset; Using the screened spelling correction corpus for the specific field to fine-tune the pre-trained language model until the final spelling correction model is obtained.

2. The method for constructing a specific domain spelling correction corpus based on a confusion set according to claim 1, wherein The step of using a speech recognition model to recognize the speech input in a specific field as a preliminary text result, comparing the preliminary text result with the true label to obtain pseudo data containing spelling mistakes includes: Collecting the text data in the specific field and converting it into speech data; Using the speech recognition model to recognize the speech data to obtain the preliminary text result; Comparing the preliminary text result with the text data to construct a spelling correction dataset, and comparing the correction dataset with the true label to obtain the pseudo data containing spelling mistakes.

3. The method for constructing a domain-specific spelling correction corpus based on a confusion set according to claim 1, wherein The step of constructing a spelling mistake confusion set for the specific field based on the pseudo data, sorting each group of words in the confusion set by word frequency and retaining the top n high-frequency words includes: Identifying the spelling mistakes in the pseudo data, pairing the wrong words with the corresponding correct words to construct a spelling mistake confusion set for the specific field; Performing word frequency statistics on each phrase in the spelling mistake confusion set for the specific field, and sorting all phrases according to word frequency; Retaining the top n high-frequency words from all the sorted phrases.

4. The method for constructing a specific domain spelling correction corpus based on a confusion set according to claim 1, wherein, The step of using web crawler technology to obtain the monolingual corpus of the specific field includes: Using web crawler technology to obtain the online text resources in the specific field, and the online text resources are the true corpus dataset; Segmenting the text data in the collected online text resources and dividing it into independent sentences; Preprocessing the divided independent sentences; Deleting the sentences containing website information or unrecognizable encoding information from the preprocessed sentences, and standardizing the text format to obtain the monolingual corpus.

5. The method for constructing a domain-specific spelling correction corpus based on a confusion set according to claim 1, wherein The step of combining the confusion set and the monolingual corpus to generate a spelling correction corpus containing various spelling mistakes includes: Taking each sentence in the monolingual corpus as the source sentence, screening all the words in the source sentence that appear in the confusion set as the corresponding candidate replacement words; Randomly selecting several words from the source sentence, screening the highest frequency words with the same spelling mistake type as each randomly selected word in the confusion set for replacement and modification to generate sentences containing various spelling mistakes; Pair the modified sentence with the source sentence to form the spelling correction corpus containing multiple spelling mistakes.

6. The method for constructing a domain-specific spelling correction corpus based on a confusion set according to claim 1, characterized in that, The pre-trained language model adopts a Transformer architecture, and a weight matrix is added to the attention layer of the pre-trained language model to adjust the attention weights of the words in the confusion set.

7. The method for constructing a specific domain spelling correction corpus based on a confusion set according to claim 1, wherein The training objective of the pre-trained language model is to minimize the cross-entropy loss function, which is used to evaluate the difference between the error distribution of the predicted dataset and the error distribution of the true dataset. The specific minimization cross-entropy loss function is θ are trainable model parameters, x is the source sentence, and y = {y1, y2, …, y n} is the correct sentence with n words, and y <t = {y1, y2, …, y t-1} are the words visible at time step t.

8. A specific domain spelling correction corpus construction device based on a confusion set, characterized in that, including: A pseudo-data generation module, which uses a speech recognition model to recognize the speech input in a specific domain as a preliminary text result, compares the preliminary text result with the true label, and obtains pseudo-data containing spelling mistakes. A confusion set construction module, which constructs the spelling mistake confusion set in the specific domain based on the pseudo-data, sorts each group of words in the confusion set according to word frequency, and retains the first n high-frequency words. An error correction corpus generation module, which uses web crawler technology to obtain the monolingual corpus in the specific domain, combines the confusion set and the monolingual corpus, and generates a spelling correction corpus containing multiple spelling mistakes. An error correction corpus screening module, which inputs the spelling correction corpus into the pre-trained language model for training. During the training process, by adjusting the attention mechanism of the model, the weights of the words in the confusion set are enhanced, and through iterative training, a spelling correction corpus in the specific domain whose spelling mistake distribution is within a preset difference from the true corpus dataset is screened out. A model training module, which uses the screened spelling correction corpus in the specific domain to fine-tune the pre-trained language model until the final spelling correction model is obtained.

9. A computer device, including a memory and a processor, where the memory stores a computer program, and when the computer program is executed by the processor, the processor executes the steps of the method according to any one of claims 1 to 7.

10. A computer-readable storage medium, storing a computer program, and when the computer program is executed by a processor, the processor executes the steps of the method according to any one of claims 1 to 7.