A text error detection method, device, electronic device and storage medium
By constructing a semantic unknown detection model, using the training set of standard statement text and wrong statement text, accurately detecting semantic unknown errors in the text, solving the problem of ineffective detection of semantic unknown errors in the existing technology and improving the accuracy of text detection.
Patent Information
- Application Number
- CN202010558690.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-06-18
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2040-06-18
AI Technical Summary
The prior art is difficult to effectively detect semantic errors in text, resulting in the inability to accurately understand the user's true intention of expressing it and unable to provide meaningful feedback.
By obtaining standard statement text and adding noise to generate error statement text, building a training set, using the initial model training to obtain a semantic unknown detection model, and then performing text error detection on the detection statement text.
It realizes accurate detection of semantic errors in text, improves the accuracy of text detection, can more accurately understand the user's true intentions and provide meaningful feedback.
Smart Images

Figure CN113822073B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of text detection, and particularly relates to a method and device for text error detection, an electronic device, and a storage medium. Background Art
[0002] In interactive grammar feedback teaching, detecting errors in the text input by users can improve the teaching quality. The semantic ambiguity error in the text is an error that causes ambiguity in the text. In the related art, when the user's input is a semantically ambiguous sentence, it is directly determined that a grammar error is detected and grammar correction is performed. This correction method often results in changing the wrong place to an expression that is still wrong. The above method cannot understand the true expression intention of the user and provide meaningful feedback.
[0003] Therefore, how to detect semantic ambiguity errors in text and improve the accuracy of text detection is a technical problem that those skilled in the art need to solve currently. Summary of the Invention
[0004] The purpose of the present application is to provide a method and device for text error detection, an electronic device, and a storage medium, which can detect semantic ambiguity errors in text and improve the accuracy of text detection.
[0005] To solve the above technical problem, the present application provides a method for text error detection, which includes:
[0006] Obtain a standard statement text and add the standard statement text as a first sample to a training set;
[0007] Add noise to the standard statement text to obtain an error statement text, and add the error statement text as a second sample to the training set;
[0008] Use the samples in the training set to train an initial model to obtain a semantic ambiguity detection model;
[0009] Perform a text error detection operation on the statement text to be detected through the semantic ambiguity detection model.
[0010] Optionally, adding noise to the standard statement text to obtain an error statement text includes:
[0011] Perform a noise addition operation on the words in the standard statement text based on a preset probability P to obtain the error statement text;
[0012] Wherein, during the process of performing the noise addition operation, the probability that a word is subjected to the noise addition operation is P, and the probability that a word is not subjected to the noise addition operation is 1 - P.
[0013] Optionally, the noise addition operation includes any one or a combination of any several of a word addition operation, a word deletion operation, a word replacement operation, and a word order swapping operation.
[0014] Optionally, performing the word replacement operation on the words in the standard statement text includes:
[0015] Selecting a first target word from a word list and replacing the word in the standard statement text with the first target word; wherein, the word list includes all words in the standard statement text.
[0016] Optionally, before performing the noise addition operation on the words in the standard statement text based on a preset probability P, it further includes:
[0017] Querying a second target word in the standard statement text;
[0018] Performing a corresponding noise addition operation on the second target word according to a word noise addition mapping relation table;
[0019] Correspondingly, performing the noise addition operation on the words in the standard statement text based on the preset probability P includes:
[0020] Performing a noise addition operation on the words in the standard statement text except the second target word based on the preset probability P.
[0021] Optionally, after training an initial model with samples in the training set to obtain a semantic ambiguity detection model, it further includes:
[0022] Performing a text error detection operation on the statement text with known classification using the semantic ambiguity detection model;
[0023] Determining the classification accuracy of the semantic ambiguity detection model according to the detection result;
[0024] If the classification accuracy is less than a preset value, adjust the preset probability P and perform the noise addition operation again using the adjusted preset probability P to obtain a new error statement text, so as to train the initial model with samples in the training set of the standard statement text and the new error statement text to obtain a new semantic ambiguity detection model.
[0025] Optionally, adding noise to the standard statement text to obtain an error statement text includes:
[0026] Dividing the standard statement into N statement groups in units of sentences;
[0027] Perform a noise addition operation on the words in the $i$-th sentence group based on a preset probability $P_i$ to obtain the incorrect sentence text; where $0 \lt i \leq N$, during the process of performing the noise addition operation on the words in the $i$-th sentence group, the probability that a word is subjected to the noise addition operation is $P_i$, and the probability that a word is not subjected to the noise addition operation is $1 - P_i$.
[0028] This application also provides a text error detection device, which includes:
[0029] A first sample acquisition module, configured to acquire a standard sentence text and add the standard sentence text as a first sample to the training set;
[0030] A second sample acquisition module, configured to add noise to the standard sentence text to obtain an incorrect sentence text and add the incorrect sentence text as a second sample to the training set;
[0031] A model training module, configured to train an initial model using the samples in the training set to obtain a semantic ambiguity detection model;
[0032] An error detection module, configured to perform a text error detection operation on the sentence text to be detected through the semantic ambiguity detection model.
[0033] This application also provides a storage medium, on which a computer program is stored, and when the computer program is executed, the steps performed by the above text error detection method are implemented.
[0034] This application also provides an electronic device, including a memory and a processor, where a computer program is stored in the memory, and when the processor calls the computer program in the memory, the steps performed by the above text error detection method are implemented.
[0035] This application provides a text error detection method, including acquiring a standard sentence text and adding the standard sentence text as a first sample to the training set; adding noise to the standard sentence text to obtain an incorrect sentence text and adding the incorrect sentence text as a second sample to the training set; training an initial model using the samples in the training set to obtain a semantic ambiguity detection model; performing a text error detection operation on the sentence text to be detected through the semantic ambiguity detection model.
[0036] After obtaining the standard statement text, this application adds noise to the standard statement text to obtain an incorrect statement text, so as to simulate the text errors that occur during the real input process of users. In this embodiment, the standard statement text and the incorrect statement text are respectively used as positive examples and negative examples to train an initial model to obtain a semantic ambiguity detection model. The semantic ambiguity detection model has the ability to detect semantic ambiguity errors in text. Therefore, the detection of semantic ambiguity can be achieved through the semantic ambiguity detection model. It can be seen that this application can detect semantic ambiguity errors in text and improve the accuracy of text detection. This application also provides a text error detection device, an electronic device, and a storage medium, which have the above beneficial effects and will not be elaborated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] To more clearly illustrate the embodiments of this application, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings in the following description are only some embodiments of this application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0038] Figure 1 It is a flowchart of a text error detection method provided by an embodiment of this application;
[0039] Figure 2 It is a flowchart of a semantic ambiguity detection method based on a neural network provided by an embodiment of this application;
[0040] Figure 3 It is a schematic diagram of the construction process of a semantic ambiguity detection model provided by an embodiment of this application;
[0041] Figure 4 It is a schematic diagram of the structure of a text error detection device provided by an embodiment of this application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0042] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the following will clearly and completely describe the technical solutions in the embodiments of this application with reference to the drawings in the embodiments of this application. Obviously, the described embodiments are some, but not all, of the embodiments of this application. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of this application.
[0043] Please refer to the following Figure 1 , Figure 1 It is a flowchart of a text error detection method provided by an embodiment of this application.
[0044] The specific steps may include:
[0045] S101: Obtain the standard statement text and add the standard statement text as the first sample to the training set;
[0046] Among them, in this step, multiple standard statement texts can be obtained as normal corpora. The normal corpus is a text without grammar problems or semantic ambiguity errors. To improve the training effect of the model, the standard statement texts selected in this embodiment can include sentences with multiple total word counts. This embodiment does not limit the language type of the standard statement text. This embodiment can obtain standard statement texts in one or more languages. For example, this embodiment can obtain multiple standard statement texts all in English, or can obtain multiple standard statement texts in English and Chinese.
[0047] Based on obtaining the standard statement text, this step can also add the standard statement text as the first sample to the training set. Specifically, this embodiment can also add negative example labels to all the first samples so as to perform training based on the sample labels in the subsequent machine learning model training process.
[0048] S102: Add noise to the standard statement text to obtain the error statement text and add the error statement text as the second sample to the training set;
[0049] Among them, based on obtaining the standard statement text in S101, this step simulates the errors existing in the user input text by adding noise to the standard statement text, and then obtains the error statement text. Specifically, the operation of adding noise can include any one operation or a combination of several operations among word addition operation, word deletion operation, word replacement operation, and word order exchange operation. The word addition operation means: adding a word before or after a certain word in the standard statement text; the word deletion operation means: deleting a certain word in the standard statement text; the word replacement operation means: replacing a certain word in the standard statement text with another word; the word order exchange operation means: swapping the positions of two words belonging to the same sentence in the standard statement text. Based on obtaining the error statement text, this step can also add the error statement text as the second sample to the training set. Specifically, this embodiment can also add positive example labels to all the second samples so as to perform training based on the sample labels in the subsequent machine learning model training process.
[0050] S103: Use the samples in the training set to train the initial model to obtain a semantic ambiguity detection model;
[0051] S104: Perform text error detection operations on the statement text to be detected through the semantic ambiguity detection model.
[0052] Among them, this step uses the standard sentence text and the erroneous sentence text in the training set to train the initial model to obtain a semantic ambiguity detection model with semantic ambiguity detection capabilities. The above initial model can be a two-classification model, which can learn the text features in the standard sentence text and the erroneous sentence text based on neural networks such as convolutional neural networks, recurrent neural networks, transformers, and finally output the probability of judging the input sentence as "semantically unclear" and "non-semantically unclear". After the model training is completed, for the input sentence, the model can derive the probability of classifying it as "semantically unclear" and "non-semantically unclear". When the probability of "semantically unclear" is greater than a certain threshold, the model judges it as semantically unclear.
[0053] After obtaining the standard sentence text, this embodiment adds noise to the standard sentence text to obtain the error sentence text, so as to simulate the text errors that occur during the real input process of the user. This embodiment uses the standard sentence text and the error sentence text as positive examples and negative examples to train the initial model to obtain the semantic ambiguity detection model. The semantic ambiguity detection model has the ability to detect semantic ambiguity errors in the text. Therefore, the semantic ambiguity detection model can realize semantic ambiguity detection. It can be seen that this embodiment can detect semantic ambiguity errors in the text and improve the accuracy of text detection.
[0054] As for Figure 1 In a further description of the corresponding embodiment, the process of adding noise to the standard sentence text to obtain the erroneous sentence text in S102 may be: performing a noise adding operation on the words in the standard sentence text based on a preset probability P to obtain the erroneous sentence text. In the above-mentioned process of performing the noise adding operation, the probability that the word is subjected to the noise adding operation is P, and the probability that the word is not subjected to the noise adding operation is 1-P.
[0055] Furthermore, if the noise adding operation can include any one of a word adding operation, a word deleting operation, a word replacing operation and a word order exchanging operation, then this embodiment can perform operations such as adding words, deleting words, replacing words, exchanging word order, etc. on the standard sentence text randomly in a certain proportion.
[0056] Specifically, when performing the noise addition operation on a standard sentence text in this embodiment, it first determines whether each word needs to perform the noise addition operation one by one. If a certain word needs to perform the noise addition operation, a noise addition operation is randomly selected from the word addition operation, word deletion operation, word replacement operation, and word order swapping operation to execute. The probabilities of selecting the word addition operation, word deletion operation, word replacement operation, and word order swapping operation can be preset, and the default values of the selection probabilities of a certain operation can be the same. The main basic operations of the noise addition method may include adding words, deleting words, replacing words, and swapping word orders. On this basis, various changes can be made to enhance the effectiveness of the generated data.
[0057] As a feasible implementation manner, in this embodiment, performing the word replacement operation on the words in the standard sentence text includes: selecting a first target word from the word list and replacing the word to be replaced in the standard sentence text with the first target word; where the word list includes all the words in the standard sentence text.
[0058] As a feasible implementation manner, before performing the noise addition operation on the words in the standard sentence text based on the preset probability P, the second target word in the standard sentence text can also be queried, and the corresponding noise addition operation is performed on the second target word according to the word noise addition mapping relationship table. In this embodiment, a specific noise addition operation can be performed on the second target word, and the corresponding relationship between the second target word and the noise addition operation can be recorded in the word noise addition mapping relationship table. For example, the word deletion operation can be performed on the second target word "to", and the word replacement operation can be performed on the second target word "do", replacing it with "doing". After performing the corresponding noise addition operation in the word noise addition mapping relationship table on the second target word, the process of performing the noise addition operation on the words in the standard sentence text based on the preset probability P can be: performing the noise addition operation on the words in the standard sentence text except the second target word based on the preset probability P.
[0059] The value of the preset probability P in the above embodiment can be any number greater than 0 and less than 1. For each word in the standard sentence text, there is a probability of P to add noise to it, and there is also a probability of (1-P) not to perform any operation on it. The noise adding operation can be a random one of the four operations of word adding operation, word deleting operation, word replacing operation and word order exchange operation. For example, the standard sentence "I want to eat an apple.", its translation is "I want to eat an apple." Randomly select 0.5 proportion of words for modification, such as deleting "want", randomly replacing "apple" with "is", and adding a word "morning" after "eat". After adding noise, the sentence becomes "I to eat morning is", which is a sentence belonging to the semantically unclear category. The higher the value of the preset probability P, the more random operations will be performed for each sentence, and the corresponding generated erroneous sentence text will be more "chaotic"; the lower the preset probability P, the closer the generated corpus is to the original corpus. In actual operation, if the P value is too low, the generated corpus will not meet the "non-semantic ambiguity". For example, if only one word in a sentence is modified, it is likely that the semantics of the sentence will still be clear after the modification. On the other hand, if the P value is too high, although the generated corpus meets the "semantic ambiguity", such generated data is far from the actual user data, and the trained model cannot identify the semantically unclear sentences input by the real user. For example, if P=9.999 is taken to generate corpus, and then the generated corpus is used as the training corpus of the "semantic ambiguity" category, the trained model can identify such sentences as "semantic ambiguity", but in fact the real semantically unclear sentences of the user are far different from the semantically unclear sentences generated by this method. The model cannot distinguish well when encountering semantically unclear sentences input by the user. Therefore, the purpose of adjusting the p value is to make the generated semantically unclear data simulate the real semantic ambiguity problem of the user as much as possible.
[0060] As a feasible implementation method, after the initial model is trained using the samples in the training set to obtain the semantic ambiguity detection model, the semantic ambiguity detection model can also be used to perform text error detection operations on sentence texts of known classifications; the classification accuracy of the semantic ambiguity detection model is determined based on the detection results; if the classification accuracy is less than a preset value, the preset probability P is adjusted, and the noise addition operation is re-executed using the adjusted preset probability P to obtain a new erroneous sentence text, so that the initial model can be trained using the samples in the standard sentence text and the new erroneous sentence text training set to obtain a new semantic ambiguity detection model.
[0061] Of course, the selection of the preset probability P can also be adjusted according to actual needs or experience. Different variation methods will have different impacts on the results, which can be adjusted according to experience or experimental results. Specifically, the above P1, P2, P3,..Pn can be used as the hyperparameters of the model, and then the optimal parameters can be selected according to the performance of the actual model on the dataset. For example, the corpus is randomly shuffled and divided into ten equal parts, and operations are performed with probabilities of 0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, and 1.0 respectively. These corpora are combined as training data. Multiple groups of experiments are conducted, and the P value of a certain part of these ten parts is modified respectively. Finally, a set of hyperparameters with the best effect on the test set is selected.
[0062] As a feasible implementation manner, the process of adding noise to the standard statement text to obtain the error statement text may include: dividing the standard statement into N statement groups in units of sentences; performing noise addition operations on the words in the i-th statement group based on the preset probability Pi to obtain the error statement text; where 0 < i ≤ N, during the process of performing noise addition operations on the words in the i-th statement group, the probability that a word is subjected to a noise addition operation is Pi, and the probability that a word is not subjected to a noise addition operation is 1 - Pi. For example, noise addition operations can be performed on a part of the standard statement text with probability P1 in the above manner, a part with probability P2 in the above manner, and the remaining part with probability P3 in the above manner. The above process can add noise to multiple standard statement texts according to different probabilities to obtain multiple error statement texts with different degrees of confusion, thereby truly simulating the error distribution situation of users when actually writing texts and improving the accuracy of text detection.
[0063] Furthermore, in this embodiment, there may also be an operation of determining the topological structure (i.e., the original model) for detecting semantic ambiguity. The topological structure of the detection model can be a neural network, which can be a convolutional neural network (CNN), or a recurrent neural network (RNN), or a neural network model with an attention mechanism (Transformer), etc. The last layer of the model is two-dimensional, corresponding to the two probability values of the binary classification task respectively. Taking the convolutional neural network as an example, for the input sentence, through the changes of convolutional operations and pooling operations, the last layer of the final network is two-dimensional, corresponding to the probability of the semantic ambiguity category and the probability of non-semantic ambiguity respectively. The hyperparameters of the above network are adjustable. After obtaining the standard statement text and the error statement text, the generated corpus containing two categories is used for model training, and finally a converged semantic ambiguity detection model is obtained.
[0064] The following illustrates the process described in the above embodiments through examples in actual applications. Please refer to Figure 2 and Figure 3 ,Figure 2 A flowchart of a semantic ambiguity detection method based on a neural network provided by an embodiment of the present application, Figure 3 is a schematic diagram of the construction process of a semantic ambiguity detection model provided by an embodiment of the present application. The semantic ambiguity detection method may include the following steps:
[0065] Step 1, pre-construct a semantic ambiguity detection model for automatically detecting whether a statement is semantically ambiguous.
[0066] Specifically, in this step, a certain number of standard statement texts in English can be collected first as negative examples in the training set; then these standard statements can be made into semantically ambiguous sentences by adding noise to them as positive examples in the training set. Then determine the topological structure of the model and train the model based on the training corpus to obtain the final detection model.
[0067] The process of constructing the semantic ambiguity detection model may include: obtaining standard statement texts, and generating positive and negative examples required for training the model based on the above standard statement texts. Among them, the negative examples, that is, the category data of "non-semantically ambiguous", are composed of standard statement texts; the positive examples, that is, the "semantically ambiguous" category, are generated by adding noise to the standard corpus. The method of adding noise includes, but is not limited to, randomly adding words, deleting words, replacing words, swapping word orders, etc. to the sentence at a certain ratio.
[0068] Step 2, obtain the English statement written by the user;
[0069] Here, the manner of obtaining the English statement written by the user in this step is not limited, and the obtained English statement is information in text form.
[0070] Step 3, input the English statement written by the user into the detection model to obtain the probability that the sentence belongs to the semantically ambiguous category. If the probability is greater than the preset value, it is determined that there is a semantically ambiguous error in the English statement written by the user.
[0071] The above method can model complex features such as the semantics of sentences, and thus can obtain accurate detection results in identifying semantically ambiguous problems. The above solution is based on an unlabeled correct text corpus, adds noise to the text corpus to generate semantically ambiguous data, and thus generates positive and negative examples required by the classifier. The data acquisition difficulty is small and the data volume is large. The above method is used to judge whether the sentences in English writing are semantically ambiguous, especially for sentences that cannot parse the normal sentence structure and are completely incomprehensible. Based on the pre-constructed text classifier model (semantic ambiguity detection model), it can identify whether a sentence is semantically ambiguous, and thus can timely feedback to the user to meet the user's need for correct expression.
[0072] Please refer to Figure 4 , Figure 4Schematic diagram of the structure of a text error detection device provided by an embodiment of the present application;
[0073] The device may include:
[0074] A first sample acquisition module 100, configured to acquire a standard statement text and add the standard statement text as a first sample to a training set;
[0075] A second sample acquisition module 200, configured to add noise to the standard statement text to obtain an error statement text and add the error statement text as a second sample to the training set;
[0076] A model training module 300, configured to train an initial model using the samples in the training set to obtain a semantic ambiguity detection model;
[0077] An error detection module 400, configured to perform a text error detection operation on a to-be-detected statement text through the semantic ambiguity detection model.
[0078] In this embodiment, after obtaining the standard statement text, noise is added to the standard statement text to obtain an error statement text, so as to simulate text errors existing in the real input process of a user. In this embodiment, the standard statement text and the error statement text are used as positive examples and negative examples respectively to train an initial model to obtain a semantic ambiguity detection model. The semantic ambiguity detection model has the ability to detect semantic ambiguity errors in text. Therefore, semantic ambiguity can be detected through the semantic ambiguity detection model. It can be seen that this embodiment can detect semantic ambiguity errors in text and improve the accuracy of text detection.
[0079] Further, the second sample acquisition module 200 is configured to perform a noise addition operation on the words in the standard statement text based on a preset probability P to obtain the error statement text;
[0080] Wherein, during the process of performing the noise addition operation, the probability that a word is subjected to the noise addition operation is P, and the probability that a word is not subjected to the noise addition operation is 1 - P.
[0081] Further, the noise addition operation includes any one operation or a combination of any several operations among word addition operation, word deletion operation, word replacement operation, and word order exchange operation.
[0082] Further, the second sample acquisition module 200 includes:
[0083] A word replacement unit, configured to select a first target word from a word list and replace the word in the standard statement text with the first target word; wherein, the word list includes all words in the standard statement text.
[0084] Further, it further includes:
[0085] A word query module, configured to query a second target word in the standard statement text before performing a noise addition operation on the words in the standard statement text based on a preset probability P; and perform a corresponding noise addition operation on the second target word according to a word noise addition mapping relation table.
[0086] Correspondingly, a second sample acquisition module 200 is configured to perform a noise addition operation on words in the standard statement text other than the second target word based on the preset probability P.
[0087] Further, it further includes:
[0088] A re-noising module, configured to, after training an initial model with samples in the training set to obtain a semantic ambiguity detection model, perform a text error detection operation on a statement text with a known classification by using the semantic ambiguity detection model; and further configured to determine the classification accuracy of the semantic ambiguity detection model according to the detection result; if the classification accuracy is less than a preset value, adjust the preset probability P, and re-perform the noise addition operation by using the adjusted preset probability P to obtain a new error statement text, so as to train the initial model with samples in the standard statement text and the new error statement text training set to obtain a new semantic ambiguity detection model.
[0089] Further, the second sample acquisition module 200 includes:
[0090] A statement group division unit, configured to divide the standard statement into N statement groups in units of sentences;
[0091] A grouped noise addition unit, configured to perform a noise addition operation on words in the i-th statement group based on a preset probability Pi to obtain the error statement text; where 0 < i ≤ N, and during the process of performing the noise addition operation on words in the i-th statement group, the probability that a word is subjected to the noise addition operation is Pi, and the probability that a word is not subjected to the noise addition operation is 1 - Pi.
[0092] Since the embodiments in the apparatus part correspond to the embodiments in the method part, for the embodiments in the apparatus part, please refer to the description of the embodiments in the method part, and details are not described herein again.
[0093] The present application further provides a storage medium, on which a computer program is stored, and when the computer program is executed, the steps provided in the above embodiments can be implemented. The storage medium may include: various media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc that can store program codes.
[0094] The present application also provides an electronic device, which may include a memory and a processor. When the processor calls the computer program stored in the memory, the steps provided in the above embodiments can be implemented. Of course, the electronic device may also include various network interfaces, power supplies and other components.
[0095] The various embodiments in the specification are described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. For the same or similar parts among the various embodiments, reference can be made to each other. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the descriptions are relatively simple. For the relevant parts, reference can be made to the description in the method part. It should be noted that for those of ordinary skill in the art in the technical field of the present application, without departing from the principle of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the protection scope of the claims of the present application.
[0096] It should also be noted that in this specification, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variation thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article or device including the said element.
Claims
1. A method for text error detection, characterized in that, comprising: Obtain a standard statement text, and add the standard statement text as a first sample to the training set; Add noise to the standard statement text to obtain an error statement text, and add the error statement text as a second sample to the training set; Use the samples in the training set to train an initial model to obtain a semantic ambiguity detection model; Perform text error detection operations on the text of the statement to be detected through the semantic ambiguity detection model; Among them, adding noise to the standard statement text to obtain an error statement text includes: Divide the standard statement into N statement groups in units of sentences; Perform noise addition operations on the words in the i-th statement group based on a preset probability Pi to obtain the error statement text; where 0 < i ≤ N, during the process of performing noise addition operations on the words in the i-th statement group, the probability that a word is subjected to a noise addition operation is Pi, and the probability that a word is not subjected to a noise addition operation is 1 - Pi; Among them, after using the samples in the training set to train an initial model to obtain a semantic ambiguity detection model, it further includes: Use the semantic ambiguity detection model to perform text error detection operations on the text of the statement with known classification; Determine the classification accuracy of the semantic ambiguity detection model according to the detection results; If the classification accuracy is less than the preset value, adjust the preset probability, and use the adjusted preset probability to re-perform the noise addition operation to obtain a new error statement text, so as to use the samples in the training set of the standard statement text and the new error statement text to train an initial model to obtain a new semantic ambiguity detection model.
2. The text error detection method according to claim 1, characterized in that, The noise addition operation includes any one or a combination of any several operations among word addition operation, word deletion operation, word replacement operation, and word order exchange operation.
3. The text error detection method according to claim 2, characterized in that, Performing the word replacement operation on the words in the standard statement text includes: Select a first target word from the word list, and replace the word in the standard statement text with the first target word; where the word list includes all the words in the standard statement text.
4. The text error detection method according to claim 2, characterized in that, Before performing the noise addition operation on the words in the standard statement text based on a preset probability P, it further includes: Query the second target word in the standard statement text; Perform the corresponding noise addition operation on the second target word according to the word noise addition mapping table; Correspondingly, performing the noise addition operation on the words in the standard statement text based on a preset probability P includes: Perform the noise addition operation on the words in the standard statement text except the second target word based on the preset probability P.
5. A text error detection device, characterized in that, comprising: A first sample acquisition module, configured to obtain a standard statement text, and add the standard statement text as a first sample to the training set; A second sample acquisition module, configured to add noise to the standard statement text to obtain an incorrect statement text, and use the incorrect statement text as a second sample to be added to the training set; A model training module, configured to train an initial model using the samples in the training set to obtain a semantic ambiguity detection model; An error detection module, configured to perform a text error detection operation on the statement text to be detected through the semantic ambiguity detection model; The second sample acquisition module includes: A statement group division unit, configured to divide the standard statement into N statement groups in units of sentences; A grouped noise addition unit, configured to perform a noise addition operation on the words in the i-th statement group based on a preset probability Pi to obtain the incorrect statement text; where 0 < i ≤ N, during the process of performing the noise addition operation on the words in the i-th statement group, the probability that a word is subjected to the noise addition operation is Pi, and the probability that a word is not subjected to the noise addition operation is 1 - Pi; The text error detection device further includes: A re-noise addition module, configured to, after training an initial model using the samples in the training set to obtain a semantic ambiguity detection model, perform a text error detection operation on the statement text with known classification through the semantic ambiguity detection model; and further configured to determine the classification accuracy of the semantic ambiguity detection model according to the detection result; if the classification accuracy is less than a preset value, adjust the preset probability, and re-perform the noise addition operation using the adjusted preset probability to obtain a new incorrect statement text, so as to train the initial model using the samples in the training set of the standard statement text and the new incorrect statement text to obtain a new semantic ambiguity detection model.
6. An electronic device, characterized in that it includes a memory and a processor, a computer program is stored in the memory, and when the processor calls the computer program in the memory, the steps of the text error detection method according to any one of claims 1 to 4 are implemented.
7. A storage medium, characterized in that computer-executable instructions are stored in the storage medium, and when the computer-executable instructions are loaded and executed by a processor, the steps of the text error detection method according to any one of the above claims 1 to 4 are implemented.
Citation Information
Patent Citations
KR20200044208A