Speech error recognition method, device, electronic device and storage medium

By combining the pre-trained language model and the syntactic dependency classification model, the context semantics and syntactic dependency relationship of word segmentation are extracted, and the problem of low recognition accuracy of semantic sexually transmitted sentences in the prior art is solved, and the accurate recognition of semantic and syntactic problems is achieved.

CN114154497BActive Publication Date: 2025-07-22IFLYTEK CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202111467935.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-03
Publication Date
2025-07-22
Estimated Expiration
2041-12-03

AI Technical Summary

Technical Problem

The recognition accuracy of semantic sexually transmitted sentences in the prior art is low, and traditional methods such as BERT models are difficult to accurately identify sentences with semantic problems such as unclear expression and illogicality.

Method used

By combining the pre-trained language model and the syntactic dependency classification model, the context semantic and syntactic dependency relationships of word segments are extracted, and semantic information and syntactic information are fused for wrong sentence recognition.

Benefits of technology

It improves the recognition accuracy of semantic sexually transmitted sentences, and can accurately identify problems such as improper word order, incomplete components, improper coordination, confusing structure, unclear expressions, and illogical expressions, providing more accurate results for identifying sentences.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114154497B_ABST
    Figure CN114154497B_ABST
Patent Text Reader

Abstract

The present invention provides a method, apparatus, electronic device and storage medium for identifying language errors. The method includes: determining a sentence to be identified; extracting the token representations of each token in the sentence to be identified; performing language error identification on the sentence to be identified based on the token representations of each token in the sentence to be identified and the syntactic structure of the sentence to be identified; the token representation is used to characterize the context semantics of the corresponding token and the syntactic dependency relationship between the corresponding token and the remaining tokens in the sentence to be identified. The method, apparatus, electronic device and storage medium for identifying language errors provided by the present invention can identify syntactic structure problems and semantic problems in the sentence to be identified by combining semantic information and syntactic information, and thus accurately obtain the language error identification result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of natural language processing, and in particular, to a method, device, electronic device and storage medium for identifying language errors. Background Art

[0002] During text input, the input text often has language errors due to various reasons. For example, spelling mistakes, improper collocations, incomplete components, etc. may all lead to problems such as grammar errors and unclear semantics in the text.

[0003] Currently, the language representation model BERT (Bidirectional Encoder Representations from Transformers) is mostly used to identify language errors in the sentence to be identified. However, the above method has a reduced recognition accuracy for semantic language errors. Summary of the Invention

[0004] The present invention provides a method, device, electronic device and storage medium for identifying language errors, so as to solve the defect of low recognition accuracy for semantic language errors in the prior art.

[0005] The present invention provides a method for identifying language errors, including:

[0006] Determine the sentence to be identified;

[0007] Extract the token representations of each token in the sentence to be identified;

[0008] Based on the token representations of each token in the sentence to be identified and the syntactic structure of the sentence to be identified, identify language errors in the sentence to be identified; the token representation is used to characterize the context semantics of the corresponding token and the syntactic dependency relationship between the corresponding token and the remaining tokens in the sentence to be identified.

[0009] According to the method for identifying language errors provided by the present invention, the extracting the token representations of each token in the sentence to be identified includes:

[0010] Based on a pre-trained language model, extract the token representations of each token in the sentence to be identified;

[0011] The pre-trained language model is obtained by training an initial language model in combination with a syntactic dependency relationship classification model using a first sample sentence and the syntactic dependency relationship labels between the tokens in the first sample sentence.

[0012] According to the method for identifying language errors provided by the present invention, the pre-trained language model is obtained by training based on the following steps:

[0013] Input the first sample sentence into the initial language model to obtain the predicted token representations of each token in the first sample sentence output by the initial language model;

[0014] Input the predicted token representations of each token in the first sample sentence into the syntactic dependency relationship classification model to obtain the predicted syntactic dependency relationships between the tokens in the first sample sentence output by the syntactic dependency relationship classification model;

[0015] Based on the predicted syntactic dependency relationships between the tokens in the first sample sentence and the syntactic dependency relationship labels between the tokens in the first sample sentence, jointly train the initial language model and the syntactic dependency relationship classification model to obtain the pre-trained language model.

[0016] According to a method for identifying language errors provided by the present invention, the first sample sentence includes a sample masked sentence and a first original sample sentence, and the sample masked sentence is obtained by performing token masking on a second original sample sentence;

[0017] The step of jointly training the initial language model and the syntactic dependency relationship classification model based on the predicted syntactic dependency relationships between the tokens in the first sample sentence and the syntactic dependency relationship labels between the tokens in the first sample sentence to obtain the pre-trained language model includes:

[0018] Based on the predicted syntactic dependency relationships between the tokens in the first original sample sentence and the syntactic dependency relationship labels between the tokens in the first original sample sentence, and the predicted sentence corresponding to the sample masked sentence and the second original sample sentence, jointly train the initial language model, the syntactic dependency relationship classification model, and the sentence prediction model to obtain the pre-trained language model;

[0019] The sentence prediction model is used to predict the sample masked sentence to obtain the predicted sentence corresponding to the sample masked sentence.

[0020] According to a method for identifying language errors provided by the present invention, the step of identifying language errors in the sentence to be identified based on the token representations of each token in the sentence to be identified and the syntactic structure of the sentence to be identified includes:

[0021] Based on the token representations of each token in the sentence to be identified and the syntactic structure of the sentence to be identified, determine the syntactic representation of the sentence to be identified;

[0022] Based on the fusion weights, fuse the token representations of each token in the sentence to be identified and the syntactic representation of the sentence to be identified to obtain a fused sentence representation; the fusion weights are determined based on the token representations of each token in the sentence to be identified;

[0023] Based on the fused sentence representation, perform error recognition on the sentence to be recognized.

[0024] According to an error recognition method provided by the present invention, extracting the token representations of each token in the sentence to be recognized, and based on the token representations of each token in the sentence to be recognized and the syntactic structure of the sentence to be recognized, performing error recognition on the sentence to be recognized, including:

[0025] Based on an error recognition model, extracting the token representations of each token in the sentence to be recognized, and based on the token representations of each token in the sentence to be recognized and the syntactic structure of the sentence to be recognized, performing error recognition on the sentence to be recognized;

[0026] The error recognition model is trained based on a pre-trained language model, using a second sample sentence and the error label of the second sample sentence. The pre-trained language model is used to extract token representations, and the pre-trained language model is trained based on an initial language model, using a first sample sentence and the syntactic dependency relationship labels between each token in the first sample sentence, in combination with a syntactic dependency relationship classification model.

[0027] According to an error recognition method provided by the present invention, the syntactic structure is determined based on the following steps:

[0028] Perform syntactic analysis on the sentence to be recognized to obtain the syntactic dependency relationships between each token in the sentence to be recognized;

[0029] Based on the syntactic dependency relationships between each token, construct a syntactic dependency relationship structure tree representing the syntactic dependency relationships between each token and other tokens in the sentence to be recognized as the syntactic structure.

[0030] The present invention also provides an error recognition device, including:

[0031] A determination unit for determining a sentence to be recognized;

[0032] An extraction unit for extracting the token representations of each token in the sentence to be recognized;

[0033] An identification unit for performing error recognition on the sentence to be recognized based on the token representations of each token in the sentence to be recognized and the syntactic structure of the sentence to be recognized; the token representation is used to represent the context semantics of the corresponding token and the syntactic dependency relationship between the corresponding token and the remaining tokens in the sentence to be recognized.

[0034] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps of any one of the above-mentioned speech error recognition methods are implemented.

[0035] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of any one of the above-mentioned speech error recognition methods are implemented.

[0036] The present invention also provides a computer program product, including a computer program. When the computer program is executed by a processor, the steps of any one of the above-mentioned speech error recognition methods are implemented.

[0037] The speech error recognition method, device, electronic device, and storage medium provided by the present invention can identify speech errors in a sentence to be recognized by combining the context semantics of corresponding words in each word segmentation representation, the syntactic dependencies between the corresponding words and the remaining words in the sentence to be recognized, and the syntactic dependencies between words in the syntactic structure, so as to accurately obtain the speech error recognition result by combining semantic information and syntactic information. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0039] Figure 1 is a flowchart of the speech error recognition method provided by the present invention;

[0040] Figure 2 is one of the flowcharts of the pre-trained language model training method provided by the present invention;

[0041] Figure 3 is another flowchart of the pre-trained language model training method provided by the present invention;

[0042] Figure 4 is a flowchart of the implementation manner of step 130 of the speech error recognition method provided by the present invention;

[0043] Figure 5 is a flowchart of the speech error recognition model training method provided by the present invention;

[0044] Figure 6 is a flowchart of the syntactic structure determination method provided by the present invention;

[0045] Figure 7 It is a schematic structural diagram of the disease sentence recognition device provided by the present invention;

[0046] Figure 8 It is a schematic structural diagram of the electronic device provided by the present invention. Specific embodiments

[0047] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below with reference to the accompanying drawings in the present invention. Apparently, the described embodiments are some but not all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention without making creative efforts shall fall within the protection scope of the present invention.

[0048] Diseased sentences usually include structural diseased sentences and semantic diseased sentences. Structural diseased sentences refer to sentences with syntactic structure problems such as improper word order, incomplete components, improper collocations, and chaotic structures. Semantic diseased sentences refer to sentences with semantic problems such as unclear meaning and illogicality.

[0049] Currently, the language representation model BERT is mostly used to identify diseased sentences in the sentence to be recognized. This method can be used to identify syntactic structure problems such as improper word order, incomplete components, improper collocations, and chaotic structures in sentences. However, for semantic diseased sentences with semantic problems such as unclear meaning and illogicality, the recognition accuracy is relatively low. For example, for the sentence to be recognized "His hometown is a person from Fuzhou City, Fujian Province", using the language representation model BERT in the traditional method to detect that it has no syntactic structure problems such as improper word order, incomplete components, improper collocations, and chaotic structures. However, according to the context semantic information of this sentence, "hometown" cannot be "a person from Fuzhou City", that is, this sentence has an illogical semantic problem, that is, this sentence is actually a semantic diseased sentence, but the diseased sentence result cannot be accurately recognized by the traditional method.

[0050] In view of this, the present invention provides a method for recognizing diseased sentences. Figure 1 It is a schematic flow diagram of the method for recognizing diseased sentences provided by the present invention. As Figure 1 shown, this method includes the following steps:

[0051] Step 110, determine the sentence to be recognized.

[0052] Here, the sentence to be recognized is the sentence that needs to be recognized for diseased sentences. The sentence to be recognized can be directly input by the user, or obtained by transcribing the collected audio, or obtained by collecting an image through an image acquisition device such as a scanner, a mobile phone, a camera, etc. and performing OCR character recognition on the image. The embodiments of the present invention do not make specific limitations on this.

[0053] Step 120: Extract the token representations of each token in the sentence to be recognized; the token representation is used to characterize the context semantics of the corresponding token and the syntactic dependency relationship between the corresponding token and the remaining tokens in the sentence to be recognized.

[0054] Specifically, the sentence to be recognized includes multiple tokens. The context semantics of any token includes the information of the token in the semantic level in the sentence to be recognized, and the syntactic dependency relationship between any token and the remaining tokens in the sentence to be recognized includes the relationship between the token and the remaining tokens in the syntactic level. The token representations of each token are used to characterize the above two kinds of information, that is, the token representations of each token combine the information of the corresponding token in the semantic level and the relationship in the syntactic level. Among them, the token representation of each token can be a hidden layer vector containing the context information and syntactic dependency relationship of the corresponding token, or a fusion form of a feature vector representing the context information of the corresponding token and a feature vector representing the syntactic dependency relationship between the corresponding token and other tokens. The embodiments of the present invention do not make specific limitations on this.

[0055] Optionally, four syntactic dependency relationships, namely "father-child relationship", "child-father relationship", "brother relationship" and "no direct relationship", can be set. For example, for the sentence to be recognized "His hometown is Fuzhou City, Fujian Province", the relationship between "he" and "hometown" is a father-child relationship, the relationship between "is" and "city" is a child-father relationship, the relationship between "hometown" and "city" is a brother relationship, and there is no direct relationship between "city" and "he".

[0056] In addition, when extracting the token representations of each token in the sentence to be recognized, the sentence to be recognized can be input into a pre-trained language model, and the pre-trained language model can mine the context semantic information of each token and the syntactic dependency relationship between each token and the remaining tokens, so as to accurately obtain the token representation that characterizes the context semantic information of the corresponding token and the syntactic dependency relationship between the corresponding token and the remaining tokens. Among them, the pre-trained language model can be trained by applying sample sentences and the syntactic dependency relationship labels between the tokens in the sample sentences on the basis of the existing initial language representation model BERT, in combination with a syntactic dependency relationship classification model.

[0057] Step 130: Based on the token representations of each token in the sentence to be recognized and the syntactic structure of the sentence to be recognized, perform disease recognition on the sentence to be recognized.

[0058] Specifically, the syntactic structure of the sentence to be recognized can be obtained by performing syntactic analysis on the sentence to be recognized. The syntactic analysis can specifically reveal its syntactic structure by analyzing the dependency relationship between each participle in the sentence to be recognized. For example, the sentence to be recognized "help me turn on the silent fan in the bedroom" can be known through syntactic analysis. The participles "help", "me", "turn on", "bedroom", "fan", and "silent wind" are verb v, pronoun r, verb v, noun n, noun n, and noun n respectively. Among them, "help" is the core relationship HED in the text, "me" is the auxiliary word DBL of "help", there is a verb-object relationship VOB between "help" and "turn on", there is a verb-object relationship VOB between "turn on" and "silent wind", there is a subject-predicate relationship ATT between "fan" and "silent wind", and there is a subject-predicate relationship ATT between "bedroom" and "fan". It can be seen from this that the syntactic structure of the sentence to be recognized can represent the syntactic dependency relationship between each participle in the sentence to be recognized at the level of the part of speech of each participle and the syntactic structure between each participle.

[0059] Since the participle representation represents the contextual semantics of the corresponding participle and the syntactic dependency between the corresponding participle and the other participles in the sentence to be recognized from the semantic and syntactic levels, and the syntactic structure represents the syntactic dependency between the participles from the level of the part of speech of each participle and the syntactic structure between the participles, it is possible to determine whether there are syntactic structure problems such as improper word order, incomplete components, improper collocation, and chaotic structure in the sentence to be recognized based on the participle representation and the syntactic dependency between the participles in the syntactic structure. At the same time, combined with the contextual semantics of each participle in the participle representation, it is possible to determine whether there are semantic problems such as unclear meaning and illogicality in the sentence to be recognized, thereby accurately identifying linguistic errors in the sentence to be recognized.

[0060] Compared with the traditional method that uses the language representation model BERT to only identify structurally incorrect sentences, the embodiment of the present invention integrates the contextual semantics of the corresponding segmentation in each segmentation representation and the syntactic dependency relationship between the corresponding segmentation and the other segmentations in the sentence to be identified, as well as the syntactic dependency relationship between each segmentation in the syntactic structure, so that it can combine semantic information and syntactic information to identify incorrect sentences for syntactic structure problems and semantic problems in the sentence to be identified, and thus accurately obtain incorrect sentence recognition results.

[0061] The method for identifying grammatical errors provided by the embodiment of the present invention combines the contextual semantics of the corresponding participle in each participle representation and the syntactic dependency relationship between the corresponding participle and the remaining participles in the sentence to be identified, as well as the syntactic dependency relationship between the participles in the syntactic structure, so that it can combine semantic information and syntactic information to identify grammatical structure problems and semantic problems in the sentence to be identified, thereby accurately obtaining the error sentence identification result.

[0062] Based on the above embodiment, step 120 includes:

[0063] Based on a pre-trained language model, extract the token representations of each token in the sentence to be recognized;

[0064] The pre-trained language model is obtained by jointly training an initial language model with a first sample sentence and the syntactic dependency relationship labels between the tokens in the first sample sentence using a syntactic dependency relationship classification model.

[0065] Specifically, the initial language model can be a BERT model or other models that have semantic extraction capabilities after pre-training. The pre-trained model refers to the initial language model that has completed joint training. After obtaining the pre-trained model, the sentence to be recognized can be input into the pre-trained language model to obtain the token representations of each token in the sentence to be recognized.

[0066] Among them, during the process of jointly training with the first sample sentence and the syntactic dependency relationship labels between the tokens in the first sample sentence using the syntactic dependency relationship classification model, the initial language model can not only learn the context semantic information of each token in the first sample sentence but also learn the syntactic dependency relationships between each token and other tokens in the first sample sentence, so that the pre-trained language model obtained by training can obtain token representations for characterizing the context semantics of the corresponding token and the syntactic dependency relationships between the corresponding token and the remaining tokens in the sentence to be recognized.

[0067] It can be seen that the embodiment of the present invention is based on an initial language model, applies a first sample sentence and the syntactic dependency relationship labels between the tokens in the first sample sentence, and jointly trains a pre-trained language model with a syntactic dependency relationship classification model, so that token representations for characterizing the context semantics of the corresponding token and the syntactic dependency relationships between the corresponding token and the remaining tokens in the sentence to be recognized can be obtained based on the pre-trained language model. Furthermore, when identifying language errors in the sentence to be recognized, the context semantic information of the corresponding token and the syntactic dependency relationship information between the corresponding token and the remaining tokens in the sentence to be recognized can be provided from the semantic level and the syntactic level, and the language error recognition result can be accurately obtained.

[0068] Based on any of the above embodiments, Figure 2 is a schematic flowchart of the pre-trained language model training method provided by the present invention. As Figure 2 shown, the pre-trained language model is obtained by training based on the following steps:

[0069] Step 210: Input the first sample sentence into the initial language model to obtain the predicted token representations of each token in the first sample sentence output by the initial language model;

[0070] Step 220: Input the predicted token representations of each token in the first sample sentence into the syntactic dependency relationship classification model to obtain the predicted syntactic dependency relationships between the tokens in the first sample sentence output by the syntactic dependency relationship classification model;

[0071] Step 230: Based on the predicted syntactic dependency relationships between the word segments in the first sample sentence and the syntactic dependency relationship labels between the word segments in the first sample sentence, jointly train the initial language model and the syntactic dependency relationship classification model to obtain a pre-trained language model.

[0072] Specifically, after inputting the first sample sentence into the initial language model, the initial language model can respectively obtain the representation vectors of the word segments in the first sample sentence, and then splice the representation vector of any word segment with the representation vectors of other word segments to obtain a predicted word segment representation for characterizing the context semantics of the corresponding word segment in the first sample sentence and the syntactic dependency relationship between the corresponding word segment and the remaining word segments.

[0073] After obtaining the predicted word segment representation, input the predicted word segment representation into the syntactic dependency relationship classification model, and the syntactic dependency relationship classification model predicts the syntactic dependency relationship between the corresponding word segment and the remaining word segments in the first sample sentence to obtain the predicted syntactic dependency relationships between the word segments in the first sample sentence.

[0074] Then, based on the predicted syntactic dependency relationships between the word segments in the first sample sentence and the syntactic dependency relationship labels between the word segments in the first sample sentence, jointly train the initial language model and the syntactic dependency relationship classification model. During the joint training process, the initial language model can not only learn the context semantic information of the word segments in the first sample sentence, but also learn the syntactic dependency relationships between the word segments in the first sample sentence and other word segments, so that the trained pre-trained language model can obtain a word segment representation for characterizing the context semantics of the corresponding word segment and the syntactic dependency relationship between the corresponding word segment and the remaining word segments in the sentence to be recognized.

[0075] It can be seen that the embodiment of the present invention jointly trains the initial language model and the syntactic dependency relationship classification model based on the predicted syntactic dependency relationships between the word segments in the first sample sentence and the syntactic dependency relationship labels between the word segments in the first sample sentence to obtain a pre-trained language model, so that a word segment representation for characterizing the context semantics of the corresponding word segment and the syntactic dependency relationship between the corresponding word segment and the remaining word segments in the sentence to be recognized can be obtained based on the pre-trained language model. Furthermore, when identifying language errors in the sentence to be recognized, the context semantic information of the corresponding word segment and the syntactic dependency relationship information between the corresponding word segment and the remaining word segments in the sentence to be recognized can be provided from the semantic level and the syntactic level, and the language error recognition result can be accurately obtained.

[0076] It can be understood that the syntactic dependency relationship labels between the word segments in the first sample sentence can be obtained through dependency syntactic analysis by a syntactic analysis tool (such as LTP). Thus, based on the predicted syntactic dependency relationships between the word segments in the first sample sentence and the syntactic dependency relationship labels between the word segments in the first sample sentence, the initial language model and the syntactic dependency relationship classification model can be jointly trained to obtain a pre-trained language model.

[0077] Based on any of the above embodiments, the first sample sentence includes a sample masked sentence and a first original sample sentence, and the sample masked sentence is obtained by performing word segmentation masking on a second original sample sentence; wherein, step 230 includes:

[0078] Based on the predicted syntactic dependency relationships between the word segments in the first original sample sentence and the syntactic dependency relationship labels between the word segments in the first original sample sentence, and the predicted sentence corresponding to the sample masked sentence and the second original sample sentence, the initial language model, the syntactic dependency relationship classification model, and the sentence prediction model are jointly trained to obtain a pre-trained language model;

[0079] The sentence prediction model is used to predict the sample masked sentence to obtain the predicted sentence corresponding to the sample masked sentence.

[0080] Here, the first original sample sentence and the second original sample sentence can be the original corpus obtained from a public dataset, and the sample masked sentence is obtained by performing word segmentation masking on the second original sample sentence. For example, the second original sample sentence can be segmented, then 15% of the word segments are selected as the word segments to be masked, 80% of the word segments to be masked are masked (MASK), 10% of the word segments to be masked are replaced with other word segments, and the remaining 10% of the word segments to be masked remain unchanged.

[0081] After obtaining the sample masked sentence, the sample masked sentence is input into the initial language model to obtain the predicted word segment representation of the sample masked sentence, and the predicted word segment representation of the sample masked sentence is input into the sentence prediction model. The sentence prediction model restores the sample masked sentence to further learn the syntactic dependency relationships between each word segment and other word segments to obtain the predicted sentence corresponding to the sample masked sentence. At the same time, the first original sample sentence is input into the initial language model to obtain the predicted word segment representation of the first original sample sentence output by the initial language model, and the predicted word segment representation of the first original sample sentence is input into the syntactic dependency relationship classification model. The syntactic dependency relationship classification model predicts the syntactic dependency relationships between the corresponding word segments and the remaining word segments in the first original sample sentence to obtain the predicted syntactic dependency relationships between the word segments in the first original sample sentence.

[0082] Then, based on the predicted syntactic dependency relationships among the word segments in the first original sample sentence, the syntactic dependency relationship labels among the word segments in the first original sample sentence, the predicted sentence corresponding to the sample masked sentence, and the second original sample sentence, the initial language model, the syntactic dependency relationship classification model, and the sentence prediction model are jointly trained to obtain a pre-trained language model. Thus, a word segment representation for characterizing the context semantics of the corresponding word segment and the syntactic dependency relationship between the corresponding word segment and the remaining word segments in the sentence to be recognized can be obtained based on the pre-trained language model. Furthermore, when identifying language errors in the sentence to be recognized, the context semantic information of the corresponding word segment and the syntactic dependency relationship information between the corresponding word segment and the remaining word segments in the sentence to be recognized can be provided from the semantic level and the syntactic level, and the language error recognition result can be accurately obtained.

[0083] Figure 3 It is the second flowchart of the pre-trained language model training method provided by the present invention. As Figure 3 shown, the pre-trained language model is secondarily pre-trained based on the existing initial language model BERT. The pre-training tasks include SRP (Syntax Relation Prediction) and MLM (Masked Language Model). The specific training process is as follows: Input the first original sample sentence into the initial language model BERT to obtain the predicted word segment representation of the first original sample sentence, and input the predicted word segment representation of the first original sample sentence into the syntactic dependency relationship classification model SRP. The syntactic dependency relationship classification model SRP predicts the syntactic dependency relationship between the corresponding word segment and the remaining word segments in the first original sample sentence to obtain the predicted syntactic dependency relationships among the word segments in the first original sample sentence. Perform word segment masking processing on the second original sample sentence to obtain a sample masked sentence. Input the sample masked sentence into the initial language model BERT to obtain the predicted word segment representation of the sample masked sentence, and input the predicted word segment representation of the sample masked sentence into the sentence prediction model MLM. The sentence prediction model MLM restores the sample masked sentence to further learn the syntactic dependency relationships between each word segment and other word segments to obtain the predicted sentence corresponding to the sample masked sentence. Then, based on the predicted syntactic dependency relationships among the word segments in the first original sample sentence, the syntactic dependency relationship labels among the word segments in the first original sample sentence, the predicted sentence corresponding to the sample masked sentence, and the second original sample sentence, the initial language model, the syntactic dependency relationship classification model, and the sentence prediction model are jointly trained to obtain a pre-trained language model.

[0084] Based on any of the above embodiments, Figure 4 It is the flowchart of the implementation manner of step 130 of the language error recognition method provided by the present invention. As Figure 4 shown, step 130 includes:

[0085] Step 131: Determine the syntactic representation of the sentence to be recognized based on the token representations of each token in the sentence to be recognized and the syntactic structure of the sentence to be recognized.

[0086] Step 132: Based on the fusion weights, fuse the token representations of each token in the sentence to be recognized and the syntactic representation of the sentence to be recognized to obtain a fused sentence representation; the fusion weights are determined based on the token representations of each token in the sentence to be recognized.

[0087] Step 133: Identify grammar errors in the sentence to be recognized based on the fused sentence representation.

[0088] Specifically, since the token representation characterizes the context semantics of the corresponding token and the syntactic dependency relationship between the corresponding token and the remaining tokens in the sentence to be recognized from the semantic and syntactic levels, and the syntactic structure characterizes the syntactic dependency relationship between each token from the part-of-speech of each token and the syntactic structure between each token, the syntactic representation of the sentence to be recognized obtained based on these two contains the semantic-level information and syntactic-level information of the sentence to be recognized.

[0089] Thus, it can be seen that both the token representations of each token in the sentence to be recognized and the syntactic representation of the sentence to be recognized can characterize the context semantics of the corresponding token and the syntactic dependency relationship between the corresponding token and the remaining tokens in the sentence to be recognized from the semantic and syntactic levels. In addition, the token representations of each token in the sentence to be recognized can characterize the importance degree of each token in the sentence to be recognized at the semantic and syntactic levels. Furthermore, during the process of grammar error identification, the fusion weights can be determined based on the token representations of each token in the sentence to be recognized, and then, based on the fusion weights, the token representations of each token in the sentence to be recognized and the syntactic representation of the sentence to be recognized are fused to obtain a fused sentence representation. Thus, based on the syntactic dependency relationship information between each token in the fused sentence representation, it can be determined whether there are syntactic structure problems such as improper word order, incomplete components, improper collocations, and chaotic structures in the sentence to be recognized. At the same time, by combining the context semantic information of each token in the fused sentence representation, it can be determined whether there are semantic problems such as unclear meaning and illogicality in the sentence to be recognized. Furthermore, grammar errors in the sentence to be recognized can be accurately identified.

[0090] Based on any of the above embodiments, extract the token representations of each token in the sentence to be recognized, and based on the token representations of each token in the sentence to be recognized and the syntactic structure of the sentence to be recognized, identify grammar errors in the sentence to be recognized, including:

[0091] Based on a grammar error identification model, extract the token representations of each token in the sentence to be recognized, and based on the token representations of each token in the sentence to be recognized and the syntactic structure of the sentence to be recognized, identify grammar errors in the sentence to be recognized.

[0092] The error recognition model is based on a pre-trained language model, which is trained using the second sample sentence and the error labels of the second sample sentence. The pre-trained language model is used to extract word segmentation representations. The pre-trained language model is based on the initial language model, which is trained using the first sample sentence and the syntactic dependency labels between each word segment in the first sample sentence, and a joint syntactic dependency classification model.

[0093] Specifically, the error recognition model can be obtained by joint training based on a pre-trained language model and a syntactic information model. The pre-trained language model is used to extract the word segmentation representation of each word in the sentence to be recognized. The syntactic information model is used to perform error recognition on the sentence to be recognized based on the word segmentation representation of each word in the sentence to be recognized and the syntactic structure of the sentence to be recognized, and obtain the error recognition result. Among them, the pre-trained language model can be obtained by using the initial language model as described in the above embodiment, applying the first sample sentence and the syntactic dependency relationship labels between each word in the first sample sentence, and jointly training with the syntactic dependency relationship classification model. The syntactic information model can be a Tree-LSTM model, and the Tree-LSTM model can include a SyntaxTree-LSTM model, a Child-Sum Tree-LSTMs model, an N-ary Tree LSTM model, and the like.

[0094] Figure 5 : is a flow chart of the method for training a speech error recognition model provided by the present invention, such as Figure 5 As shown, the second sample sentence is input into the Embedding layer of the pre-trained language model, and each word segmentation vector of the second sample sentence is extracted. Then, each word segmentation vector is input into the Transformer layer of the pre-trained language model to obtain the predicted word segmentation representation of each word in the second sample sentence. Then, the Pooling Layer of the pre-trained language model averages the predicted word segmentation representation of each word in the second sample sentence to obtain the sentence representation x of the second sample sentence. At the same time, the predicted word segmentation representation of each word in the second sample sentence and the syntactic structure of the second sample sentence are input into the syntactic information model SyntaxTree-LSTM to obtain the syntactic representation H of the second sample sentence. Then, the Highway Gate layer fuses the sentence representation x of the second sample sentence and the syntactic representation H of the second sample sentence based on the fusion weight to obtain the fused sentence representation v of the second sample sentence. Then, the Classification layer determines the predicted error recognition result of the second sample sentence based on the fused sentence representation v of the second sample sentence, and iterative training is performed based on the predicted error recognition result of the second sample sentence and the error label of the second sample sentence. The fusion weight is determined based on the sentence representation x of the second sample sentence.

[0095] It can be seen that, based on the initial language model, the embodiments of the present invention apply the first sample sentence and the syntactic dependency relationship tags between each word segment in the first sample sentence, and jointly train a pre-trained language model with the syntactic dependency relationship classification model. Thus, a word segment representation for characterizing the context semantics of the corresponding word segment and the syntactic dependency relationship between the corresponding word segment and the remaining word segments in the sentence to be recognized can be obtained based on the pre-trained language model. Furthermore, a language disorder recognition model trained based on the pre-trained language model, the second sample sentence, and the language disorder tags of the second sample sentence can provide context semantic information of the corresponding word segment and syntactic dependency relationship information between the corresponding word segment and the remaining word segments in the sentence to be recognized from the semantic level and the syntactic level, and accurately obtain the language disorder recognition result.

[0096] Based on any of the above embodiments, Figure 6 is a schematic flowchart of the syntactic structure determination method provided by the present invention. As Figure 6 shown, the syntactic structure is determined based on the following steps:

[0097] Step 610: Perform syntactic analysis on the sentence to be recognized to obtain the syntactic dependency relationships between each word segment in the sentence to be recognized;

[0098] Step 620: Based on the syntactic dependency relationships between each word segment, construct a syntactic dependency relationship structure tree representing the syntactic dependency relationships between each word segment and other word segments in the sentence to be recognized as the syntactic structure.

[0099] Specifically, through the syntactic dependency relationships between each word segment obtained by syntactic analysis, and then based on the syntactic dependency relationships between each word segment, construct a syntactic dependency relationship structure tree representing the syntactic dependency relationships between each word segment and other word segments in the sentence to be recognized as the syntactic structure. Specifically, based on the syntactic dependency relationships between each word segment, it can be determined whether there is a syntactic dependency relationship between one word segment and the remaining word segments, and then a structure tree representing the syntactic dependency relationships between each character in this word segment and each character in the remaining word segments can be generated, and thus the syntactic structure in the form of a structure tree can be obtained.

[0100] Based on any of the above embodiments, the present invention further provides a language disorder recognition method, which includes:

[0101] Perform syntactic analysis on the sentence to be recognized to obtain the syntactic structure of the sentence to be recognized, then input the sentence to be recognized into the pre-trained language model of the language disorder recognition model to obtain the word segment representations of each word segment in the sentence to be recognized output by the pre-trained language model, and take the average of the word segment representations of each word segment in the sentence to be recognized to obtain the sentence representation of the sentence to be recognized. Among them, the pre-trained language model is obtained by performing secondary pre-training on the existing initial language model BERT, and the pre-training tasks include SRP and MLM.

[0102] Meanwhile, the token representations of each token in the sentence to be recognized and the syntactic structure of the sentence to be recognized are input into the syntactic information model Tree-LSTM to obtain the syntactic representation of the sentence to be recognized containing syntactic information.

[0103] Then, based on the fusion weights, the sentence representation of the sentence to be recognized and the syntactic representation of the sentence to be recognized are fused to obtain a fused sentence representation, and the sentence to be recognized is identified for language errors based on the fused sentence representation to obtain a language error recognition result.

[0104] Among them, the language error recognition model is obtained by jointly training the pre-trained language model and the syntactic information model Tree-LSTM using the second sample sentence and the language error label of the second sample sentence. The pre-trained language model is used to extract token representations.

[0105] Next, the language error recognition device provided by the present invention will be described. The language error recognition device described below can be mutually corresponding and referred to the language error recognition method described above.

[0106] Based on any of the above embodiments, the present invention provides a language error recognition device. Figure 7 is a structural schematic diagram of the language error recognition device provided by the present invention, as Figure 7 shown, the device includes:

[0107] A determination unit 710, configured to determine a sentence to be recognized;

[0108] An extraction unit 720, configured to extract the token representations of each token in the sentence to be recognized;

[0109] An identification unit 730, configured to identify language errors in the sentence to be recognized based on the token representations of each token in the sentence to be recognized and the syntactic structure of the sentence to be recognized; the token representation is used to characterize the context semantics of the corresponding token and the syntactic dependency relationship between the corresponding token and the remaining tokens in the sentence to be recognized.

[0110] Based on any of the above embodiments, the extraction unit 720 is configured to:

[0111] Extract the token representations of each token in the sentence to be recognized based on a pre-trained language model;

[0112] The pre-trained language model is obtained by jointly training an initial language model using the first sample sentence and the syntactic dependency relationship labels between the tokens in the first sample sentence and a syntactic dependency relationship classification model.

[0113] Based on any of the above embodiments, the device further includes:

[0114] A word segmentation representation unit, configured to input the first sample sentence into the initial language model, and obtain the predicted word segmentation representations of the word segments in the first sample sentence output by the initial language model;

[0115] A relationship determination unit, configured to input the predicted word segmentation representations of the word segments in the first sample sentence into the syntactic dependency relationship classification model, and obtain the predicted syntactic dependency relationships between the word segments in the first sample sentence output by the syntactic dependency relationship classification model;

[0116] A model training unit, configured to jointly train the initial language model and the syntactic dependency relationship classification model based on the predicted syntactic dependency relationships between the word segments in the first sample sentence and the syntactic dependency relationship labels between the word segments in the first sample sentence, to obtain the pre-trained language model.

[0117] Based on any of the above embodiments, the first sample sentence includes a sample masked sentence and a first original sample sentence, and the sample masked sentence is obtained by performing word segmentation masking on a second original sample sentence;

[0118] The model training unit is configured to:

[0119] Based on the predicted syntactic dependency relationships between the word segments in the first original sample sentence and the syntactic dependency relationship labels between the word segments in the first original sample sentence, and the predicted sentence corresponding to the sample masked sentence and the second original sample sentence, jointly train the initial language model, the syntactic dependency relationship classification model and the sentence prediction model, to obtain the pre-trained language model;

[0120] The sentence prediction model is configured to predict the sample masked sentence to obtain the predicted sentence corresponding to the sample masked sentence.

[0121] Based on any of the above embodiments, the recognition unit 730 includes:

[0122] A syntactic representation unit, configured to determine the syntactic representation of the sentence to be recognized based on the word segmentation representations of the word segments in the sentence to be recognized and the syntactic structure of the sentence to be recognized;

[0123] A fusion unit, configured to fuse the word segmentation representations of the word segments in the sentence to be recognized and the syntactic representation of the sentence to be recognized based on a fusion weight to obtain a fused sentence representation; the fusion weight is determined based on the word segmentation representations of the word segments in the sentence to be recognized;

[0124] A language error recognition unit, configured to perform language error recognition on the sentence to be recognized based on the fused sentence representation.

[0125] Based on any of the above embodiments, the extraction unit 720 and the recognition unit 730 are configured to:

[0126] Based on the language error recognition model, extract the token representations of each token in the sentence to be recognized, and perform language error recognition on the sentence to be recognized based on the token representations of each token in the sentence to be recognized and the syntactic structure of the sentence to be recognized;

[0127] The language error recognition model is trained based on a pre-trained language model, applying the second sample sentence and the language error labels of the second sample sentence. The pre-trained language model is used to extract token representations, and the pre-trained language model is trained based on an initial language model, applying the first sample sentence and the syntactic dependency relationship labels between each token in the first sample sentence, in combination with a syntactic dependency relationship classification model.

[0128] Based on any of the above embodiments, the device further includes:

[0129] A syntactic analysis unit, configured to perform syntactic analysis on the sentence to be recognized to obtain the syntactic dependency relationships between each token in the sentence to be recognized;

[0130] A construction unit, configured to construct a syntactic dependency relationship structure tree representing the syntactic dependency relationships between each token in the sentence to be recognized and other tokens based on the syntactic dependency relationships between each token, as the syntactic structure.

[0131] Figure 8 It is a schematic structural diagram of an electronic device provided by the present invention. As Figure 8 shown, the electronic device may include: a processor 810, a memory 820, a communication interface 830, and a communication bus 840. Among them, the processor 810, the memory 820, and the communication interface 830 communicate with each other through the communication bus 840. The processor 810 can call the logical instructions in the memory 820 to execute the language error recognition method, which includes: determining the sentence to be recognized; extracting the token representations of each token in the sentence to be recognized; performing language error recognition on the sentence to be recognized based on the token representations of each token in the sentence to be recognized and the syntactic structure of the sentence to be recognized; the token representation is used to characterize the context semantics of the corresponding token and the syntactic dependency relationship between the corresponding token and the remaining tokens in the sentence to be recognized.

[0132] In addition, the logic instructions in the above-mentioned memory 820 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when it is sold or used as an independent product. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk and other media that can store program codes.

[0133] On the other hand, the present invention also provides a computer program product, which includes a computer program stored on a non-transitory computer-readable storage medium, and the computer program includes program instructions. When the program instructions are executed by a computer, the computer can execute the language error recognition method provided by the above-mentioned methods, and the method includes: determining a sentence to be recognized; extracting the participle representation of each participle in the sentence to be recognized; based on the participle representation of each participle in the sentence to be recognized and the syntactic structure of the sentence to be recognized, performing language error recognition on the sentence to be recognized; the participle representation is used to characterize the contextual semantics of the corresponding participle and the syntactic dependency relationship between the corresponding participle and the remaining participles in the sentence to be recognized.

[0134] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to execute the above-mentioned error recognition methods, the methods comprising: determining a sentence to be recognized; extracting the participle representation of each participle in the sentence to be recognized; based on the participle representation of each participle in the sentence to be recognized and the syntactic structure of the sentence to be recognized, identifying errors in the sentence to be recognized; the participle representation is used to characterize the contextual semantics of the corresponding participle and the syntactic dependency relationship between the corresponding participle and the remaining participles in the sentence to be recognized.

[0135] The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without paying creative labor.

[0136] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0137] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for identifying language disorders, characterized in that, Including: Determine the statement to be recognized; Extract the token representations of each token in the statement to be recognized; Based on the token representations of each token in the statement to be recognized and the syntactic structure of the statement to be recognized, perform disease recognition on the statement to be recognized; the token representation is used to characterize the context semantics of the corresponding token and the syntactic dependency relationship between the corresponding token and the remaining tokens in the statement to be recognized; the syntactic dependency relationship includes a father-child relationship, a child-father relationship, a sibling relationship, and no direct relationship, and the syntactic structure characterizes the syntactic dependency relationship between the tokens in the statement to be recognized from the part-of-speech of each token and the syntactic structure between the tokens; The extracting the token representations of each token in the statement to be recognized includes: Based on a pre-trained language model, extract the token representations of each token in the statement to be recognized; The pre-trained language model is trained by jointly applying a first sample statement and the syntactic dependency relationship labels between the tokens in the first sample statement to an initial language model and a syntactic dependency relationship classification model.

2. The method for identifying language errors according to claim 1, wherein The pre-trained language model is trained based on the following steps: Input the first sample statement into the initial language model to obtain the predicted token representations of each token in the first sample statement output by the initial language model; Input the predicted token representations of each token in the first sample statement into the syntactic dependency relationship classification model to obtain the predicted syntactic dependency relationships between the tokens in the first sample statement output by the syntactic dependency relationship classification model; Based on the predicted syntactic dependency relationships between the tokens in the first sample statement and the syntactic dependency relationship labels between the tokens in the first sample statement, jointly train the initial language model and the syntactic dependency relationship classification model to obtain the pre-trained language model.

3. The method for identifying language errors according to claim 2, wherein The first sample statement includes a sample masked statement and a first original sample statement, and the sample masked statement is obtained by performing token masking on a second original sample statement; The jointly training the initial language model and the syntactic dependency relationship classification model based on the predicted syntactic dependency relationships between the tokens in the first sample statement and the syntactic dependency relationship labels between the tokens in the first sample statement to obtain the pre-trained language model includes: Based on the predicted syntactic dependency relationships between the tokens in the first original sample statement and the syntactic dependency relationship labels between the tokens in the first original sample statement, and the predicted statement corresponding to the sample masked statement and the second original sample statement, jointly train the initial language model, the syntactic dependency relationship classification model, and a sentence prediction model to obtain the pre-trained language model; The sentence prediction model is used to predict the sample masked statement to obtain the predicted statement corresponding to the sample masked statement.

4. The method for identifying language errors according to claim 1, wherein The performing disease recognition on the statement to be recognized based on the token representations of each token in the statement to be recognized and the syntactic structure of the statement to be recognized includes: Determine the syntactic representation of the sentence to be recognized based on the token representations of each token in the sentence to be recognized and the syntactic structure of the sentence to be recognized; Based on the fusion weights, fuse the token representations of each token in the sentence to be recognized and the syntactic representation of the sentence to be recognized to obtain a fused sentence representation; the fusion weights are determined based on the token representations of each token in the sentence to be recognized; Based on the fused sentence representation, identify language errors in the sentence to be recognized.

5. The method for identifying language errors according to claim 1, characterized in that, The extraction of the token representations of each token in the sentence to be recognized, and the identification of language errors in the sentence to be recognized based on the token representations of each token in the sentence to be recognized and the syntactic structure of the sentence to be recognized, include: Based on a language error recognition model, extract the token representations of each token in the sentence to be recognized, and identify language errors in the sentence to be recognized based on the token representations of each token in the sentence to be recognized and the syntactic structure of the sentence to be recognized; The language error recognition model is trained based on a pre-trained language model, using a second sample sentence and the language error labels of the second sample sentence. The pre-trained language model is used to extract token representations and is trained based on an initial language model, using a first sample sentence and the syntactic dependency relationship labels between each token in the first sample sentence, in combination with a syntactic dependency relationship classification model.

6. The method for identifying language errors according to any one of claims 1 to 5, characterized in that, The syntactic structure is determined based on the following steps: Perform syntactic analysis on the sentence to be recognized to obtain the syntactic dependency relationships between each token in the sentence to be recognized; Based on the syntactic dependency relationships between each token, construct a syntactic dependency relationship structure tree representing the syntactic dependency relationships between each token in the sentence to be recognized and other tokens as the syntactic structure.

7. An improper grammar recognition device, characterized in that, Include: A determination unit for determining the sentence to be recognized; An extraction unit for extracting the token representations of each token in the sentence to be recognized; An identification unit for identifying language errors in the sentence to be recognized based on the token representations of each token in the sentence to be recognized and the syntactic structure of the sentence to be recognized; the token representation is used to represent the context semantics of the corresponding token and the syntactic dependency relationship between the corresponding token and the remaining tokens in the sentence to be recognized; the syntactic dependency relationship includes a father-child relationship, a child-father relationship, a sibling relationship, and no direct relationship, and the syntactic structure represents the syntactic dependency relationships between each token in the sentence to be recognized from the aspects of the part-of-speech of each token and the syntactic structure between each token; The extraction of the token representations of each token in the sentence to be recognized includes: Based on a pre-trained language model, extract the token representations of each token in the sentence to be recognized; The pre-trained language model is trained based on an initial language model, using a first sample sentence and the syntactic dependency relationship labels between each token in the first sample sentence, in combination with a syntactic dependency relationship classification model.

8. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps of the language error recognition method according to any one of claims 1 to 6.

9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the language error recognition method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Deep neural network and multi-tag classification based wrong sentence detection method

    CN105045779A

  • Semantic comprehension method and device, electronic equipment and storage medium

    CN112560497A

  • Grammar error correction method and device, electronic equipment and storage medium

    CN112686030A