English writing grammar error detection system based on deep learning
By using the lightweight gradient hoist model for grammatical error positioning and transformer model for grammatical error classification in the English writing grammatical error detection system, multiple problems of grammatical error positioning and classification in traditional systems are solved, and higher accuracy and context consistency are achieved.
Patent Information
- Application Number
- CN202510631623.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-16
- Publication Date
- 2025-06-13
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The traditional English writing grammar error detection system has many problems in grammatical error positioning and classification, including the limitations of a single language model, the insufficient completeness of rule methods, the insufficient analytical ability of statistical methods to depend on long-distance and complex syntactic structures, and the ignoring of the inherent association and error-related effects of syntactic structures when classification of grammatical errors.
The lightweight gradient hoist model is used for syntax error positioning, and the accuracy and interpretability of error positioning are improved through adaptive fusion rule conflicts and statistical mode exceptions, and combined with inverse frequency weighting strategies. At the same time, the transformer model is used as the grammatical error classification model to model the dependence relationship between words and introduce graph structure constraints to capture the impact of local grammatical errors on the global syntax, and the positioning probability is injected into the model as a structured prior to realizing collaborative optimization of positioning and classification.
It improves the accuracy and interpretability of grammatical error positioning, improves the accuracy and context consistency of grammatical error classification, and can more effectively deal with complex long sentences and close dependencies.
Smart Images

Figure CN120146036A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of text grammatical error detection, and specifically to an English writing grammatical error detection system based on deep learning. Background Art
[0002] The English writing grammar error detection system is a system that automatically identifies grammatical errors, spelling errors, and improper grammatical structures in English writing based on technologies such as natural language processing and machine learning. It provides real-time feedback and correction suggestions by analyzing the structure and grammatical rules of the text, playing a key role in improving writing efficiency, reducing the workload of manual modification, and promoting the development of language technology.
[0003] However, the grammatical error location technology in traditional English writing grammatical error detection systems has technical problems such as being based on a single language model, the rule-based method being limited by the completeness of preset rules and being difficult to cover flexible language phenomena, and the statistical method being insufficient in parsing long-distance dependencies and complex syntactic structures. When classifying grammatical errors, traditional English writing grammatical error detection systems often treat grammatical errors as independent labels, ignoring the inherent correlation of syntactic structures in English grammar, and are prone to misjudgment of errors with close dependencies. They also lack consideration of the collateral effects of errors and are unable to handle technical problems involving complex and long sentences. Summary of the invention
[0004] In view of the above situation, in order to overcome the defects of the prior art, the present invention provides an English writing grammatical error detection system based on deep learning. In view of the technical problems that the grammatical error location technology in the traditional English writing grammatical error detection system is mostly based on a single language model, the rule method is limited by the completeness of the preset rules and is difficult to cover flexible language phenomena, and the statistical method is insufficient in parsing long-distance dependencies and complex syntactic structures, this solution creatively adopts a lightweight gradient boosting machine model as a grammatical error location model, and through adaptive fusion of rule conflicts and statistical pattern anomalies, combined with the inverse frequency weighting strategy, it ensures the effective detection of low-frequency grammatical errors while improving the accuracy of error location. degree and explainability; when traditional English writing grammatical error detection systems classify grammatical errors, they often treat grammatical errors as independent labels, ignore the inherent correlation of syntactic structures in English grammar, easily misjudge errors with close dependencies, and lack consideration of the collateral effects of errors, making it difficult to handle technical problems with complex long sentences. This solution creatively uses the Transformer model as a grammatical error classification model. By modeling the dependency relationship between words and introducing graph structure constraints, it can capture the impact of local grammatical errors on global syntax. At the same time, the positioning probability is injected into the model as a structured prior, realizing the coordinated optimization of positioning and classification, and improving the accuracy of classification and context consistency.
[0005] The technical solution adopted by the present invention is as follows: The English writing grammar error detection system based on deep learning provided by the present invention includes an original data acquisition module, a data preliminary processing module, a grammar error location model construction module, a grammar error classification model construction module, and a grammar error detection module;
[0006] The original data acquisition module obtains an original error detection data set through data acquisition;
[0007] The data preliminary processing module adopts data preliminary processing methods such as text data trimming, text vectorization, data standardization, and data set segmentation to obtain a data set to be detected, a preliminary detection training set, and a preliminary detection test set;
[0008] The grammar error location model construction module constructs a lightweight gradient boosting machine model for grammar error location as the grammar error location model, and processes the preliminary detection training set and the preliminary detection test set to obtain a detection training set and a detection test set;
[0009] The grammar error classification model construction module constructs a transformer model as the grammar error classification model. The transformer model is specifically a hybrid neural network model based on the standard transformer architecture and integrating graph attention mechanism and constraint injection;
[0010] The grammar error detection module performs grammar error detection by using the grammar error location model and the grammar error classification model to obtain a reference result for English writing grammar error detection.
[0011] Further, in the original data acquisition module, the original error detection data set specifically includes a historical original data set and a current original data set. Both the historical original data set and the current original data set include English writing text data and general corpus text data. The historical original data set further includes historical grammar error label data, grammar dependency relationship label data, and grammar rule data.
[0012] Further, in the data preliminary processing module, the text data trimming is specifically to remove HTML tags, special symbols, and non-text content in the original data and unify the text format. The text vectorization is specifically to convert the text data into a structured vector by using the BERT model. The data standardization is specifically to perform data standardization on the original data by using the Z-Score standardization method. The data set segmentation is used to segment the data set;
[0013] Perform preliminary data processing on the current original dataset through the above-mentioned text data trimming, text vectorization, and data normalization to obtain a dataset to be detected. Perform preliminary data processing on the historical original dataset through the above-mentioned text data trimming, text vectorization, data normalization, and dataset segmentation to obtain a preliminary detection training set and a preliminary detection test set.
[0014] Furthermore, in the grammar error location model construction module, it is used to construct a model required for preliminary grammar error location. Specifically, a lightweight gradient boosting machine model is constructed as the grammar error location model, and the preliminary detection training set and the preliminary detection test set are used as the input of the grammar error location model to obtain a detection training set and a detection test set.
[0015] The grammar error location model construction module specifically includes fused feature construction, obtaining word weights, model construction and training, obtaining grammar error probabilities, and grammar error location.
[0016] The fused feature construction is used to construct a multi-dimensional grammar feature representation. Specifically, a fused feature vector is obtained by combining predefined grammar rules and a statistical model. The content includes:
[0017] Rule feature extraction is used to capture explicit grammar rule conflicts. Specifically, a rule conflict score matrix is obtained through a predefined grammar rule set.
[0018] Statistical feature generation is used to capture implicit language pattern anomalies. Specifically, a statistical anomaly score vector is obtained through an n-gram language model. The formula used is as follows:
[0019] ;
[0020] In the formula, represents the average negative log probability of the language model within the window centered on the a-th word. le represents the window radius. represents the probability of predicting the c-th word based on the previous le words calculated by the n-gram language model. Sv represents the statistical anomaly score vector. represents the average negative log probability of the language model within the window centered on the first word. represents the average negative log probability of the language model within the window centered on the second word. represents the average negative log probability of the language model within the window centered on the A-th word. A represents the total number of words, and T represents the transpose operation.
[0021] Dynamic feature fusion is used to optimize the feature combination weights. Specifically, the rule features and statistical features are fused through a gating mechanism to obtain the final fused feature.
[0022] The obtained word weights are used to handle the problem of data imbalance, specifically by inverse frequency weighting to obtain the adjusted word weights;
[0023] The model is constructed and trained to construct a lightweight gradient boosting machine model as a grammar error localization model. Specifically, a lightweight gradient boosting machine model is constructed with a decision tree model as the basic learner, the model is trained based on the preliminary detection training set, and the model performance is verified based on the preliminary detection test set;
[0024] The obtained grammar error probability is used to visualize potential grammar error locations. Specifically, it is inferred through the grammar error localization model to obtain the position-level grammar error probability. The formula used is as follows:
[0025] ;
[0026] In the formula, represents the integrated grammar error probability of the a-th word, Num represents the number of basic learners, represents the prediction probability of the num-th basic learner for the a-th word, H represents the position-level grammar error probability, represents the integrated grammar error probability of the first word, represents the integrated grammar error probability of the second word, represents the integrated grammar error probability of the A-th word;
[0027] The grammar error localization is used to perform preliminary grammar error localization. Specifically, the preliminary detection training set and the preliminary detection test set are used as the inputs of the grammar error localization model to obtain the position-level grammar error probabilities based on the preliminary detection training set and the preliminary detection test set respectively, and they are incorporated into the preliminary detection training set and the preliminary detection test set respectively to obtain the detection training set and the detection test set.
[0028] Furthermore, in the grammar error classification model construction module, it is used to construct the model required for grammar error classification. Specifically, a transformer model is constructed as the grammar error classification model;
[0029] The grammar error classification model construction module specifically includes dependency edge prediction, constructing node features, designing a graph attention layer, obtaining the model output, and constructing and training the model;
[0030] The dependency edge prediction is used to capture grammar dependencies and determine the type and direction of the dependency relationships between words. Specifically, the input English composition text data is processed through a bidirectional long short-term memory network to obtain the dependency comprehensive features as the edge features for subsequent graph convolution operations. The formula used is as follows:
[0031] ;
[0032] In the formula, represents the dependency feature of the ath word processed by the bidirectional long short-term memory network, represents the operation function of the bidirectional long short-term memory network, represents the input English composition text data, represents the score of the ath word as the head node of the dependency relationship with the dth word, represents the hyperbolic tangent function, represents the learnable head node score vector, represents the learnable head node score weight matrix, represents the dependency feature of the dth word processed by the bidirectional long short-term memory network, represents the probability that there is a dependency relationship t between the ath word and the dth word, represents the softmax function, represents the learnable dependency type weight matrix, and pr represents the predefined grammar rule embedding, represents the comprehensive dependency feature between the ath word and the dth word;
[0033] The constructed node feature is used to encode the part-of-speech attribute of the word and serve as the node feature for subsequent graph convolution operations. Specifically, by looking up the part-of-speech table of the word and performing encoding, the part-of-speech feature of the word is obtained as the node feature;
[0034] The designed graph attention layer is used to fuse the syntactic structure information. Specifically, through the improved graph attention mechanism, with the comprehensive dependency feature as the edge feature and the part-of-speech feature of the word as the node feature, graph convolution propagation is performed to obtain the graph attention output feature, which includes:
[0035] Construct an adjacency matrix, which is used to construct an adjacency matrix. Specifically, based on whether there is a dependency relationship between words and the direction of the dependency relationship, an adjacency matrix is constructed. If the dependency relationship exists, the corresponding position is 1; if the dependency relationship does not exist, the corresponding position is 0;
[0036] Design an improved graph attention mechanism to focus on the key area to obtain the context-aware representation. Specifically, a mask matrix guided by the syntactic error position is constructed based on the position-level syntactic error probability, combined with the gating mechanism. The formula used for the designed improved graph attention mechanism is as follows:
[0037] ;
[0038] In the formula, Qu represents the improved graph attention query, Ke represents the improved graph attention key, and Va represents the improved graph attention value, represents the improved graph attention query transformation matrix, represents the improved graph attention key transformation matrix, Denote the improved graph attention value transformation matrix, Denote the improved graph attention input, and Mh denotes the mask matrix guided by the syntax error position, Denote the elements of the mask matrix guided by the syntax error position, Denote the integrated syntax error probability of the d-th word, and Fa denotes the output feature of the improved graph attention mechanism, Denote the dimension of the improved graph attention key;
[0039] Graph convolution propagation is used to aggregate domain information, specifically, node update features are obtained through message passing;
[0040] The obtaining of the model output is used to inject constraints and obtain the model output. Specifically, the position-level syntax error probability is injected as a constraint into the graph attention output feature, and after being processed by the transformer model, the model output syntax error classification result is obtained;
[0041] The constructing and training of the model are specifically carried out by the dependency edge prediction, the constructing of node features, the designing of the graph attention layer, and the obtaining of the model output to construct the transformer model, training the model based on the detection training set, and verifying the model performance based on the detection test set to obtain the transformer model as the syntax error classification model.
[0042] Further, in the syntax error detection module, specifically, the syntax error localization model is used to perform preliminary syntax error localization on the dataset to be detected, and the obtained position-level syntax error probability and the dataset to be detected are used as the input of the syntax error classification model to obtain the syntax error classification result. And based on the position-level syntax error probability and the syntax error classification result, the English writing syntax error detection reference result is comprehensively obtained. The English writing syntax error detection reference result specifically includes the position-level syntax error probability output by the syntax error localization model with the dataset to be detected as the input, and the syntax error classification result output by the syntax error classification model with the position-level syntax error probability obtained based on the dataset to be detected and the dataset to be detected as the input.
[0043] The beneficial effects achieved by the present invention by adopting the above scheme are as follows:
[0044] (1) In order to address the technical problems of grammatical error location technology in traditional English writing grammatical error detection systems, such as the fact that most of them are based on a single language model, the rule-based method is limited by the completeness of the preset rules and is difficult to cover flexible language phenomena, and the statistical method is insufficient in parsing long-distance dependencies and complex syntactic structures, this solution creatively adopts a lightweight gradient boosting machine model as a grammatical error location model. By adaptively fusing rule conflicts and statistical pattern anomalies and combining an inverse frequency weighting strategy, it ensures the effective detection of low-frequency grammatical errors while improving the accuracy and interpretability of error location.
[0045] (2) In view of the technical problem that traditional English writing grammatical error detection systems often classify grammatical errors as independent labels, ignoring the intrinsic correlation of syntactic structures in English grammar, and easily misjudging errors with close dependencies, and lack consideration of the collateral effects of errors, it is difficult to handle complex long sentences. This solution creatively uses the Transformer model as a grammatical error classification model. By modeling the dependency relationship between words and introducing graph structure constraints, it can capture the impact of local grammatical errors on global syntax. At the same time, the location probability is injected into the model as a structured prior, realizing the coordinated optimization of location and classification, and improving the accuracy of classification and context consistency. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] Figure 1 A schematic diagram of a module of an English writing grammatical error detection system based on deep learning provided by the present invention;
[0047] Figure 2 This is a flowchart of the data preliminary processing module;
[0048] Figure 3 Schematic diagram of the process of building modules for the grammatical error localization model;
[0049] Figure 4 Schematic diagram of the process of building modules for the grammatical error classification model.
[0050] The accompanying drawings are used to provide further understanding of the present invention and constitute a part of the specification. They are used to explain the present invention together with the embodiments of the present invention and do not constitute a limitation of the present invention. DETAILED DESCRIPTION
[0051] The technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments; based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.
[0052] In the description of the present invention, it should be understood that the terms "upper", "lower", "front", "rear", "left", "right", "top", "bottom", "inner", "outer", etc. indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings. These are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus should not be construed as a limitation to the present invention.
[0053] Example 1. Refer to Figure 1 , the English writing grammar error detection system based on deep learning provided by the present invention includes a raw data acquisition module, a data preliminary processing module, a grammar error location model construction module, a grammar error classification model construction module, and a grammar error detection module;
[0054] The raw data acquisition module obtains a raw data set for error detection by performing data acquisition;
[0055] The data preliminary processing module uses data preliminary processing methods such as text data trimming, text vectorization, data standardization, and data set segmentation to obtain a data set to be detected, a preliminary detection training set, and a preliminary detection test set;
[0056] The grammar error location model construction module constructs a lightweight gradient boosting machine model for grammar error location as the grammar error location model, and processes the preliminary detection training set and the preliminary detection test set to obtain a detection training set and a detection test set;
[0057] The grammar error classification model construction module constructs a transformer model as the grammar error classification model;
[0058] The grammar error detection module performs grammar error detection by using the grammar error location model and the grammar error classification model to obtain a reference result for English writing grammar error detection.
[0059] Example 2. Refer to Figure 1 , in the raw data acquisition module, the raw data set for error detection specifically includes a historical raw data set and a current raw data set. Both the historical raw data set and the current raw data set include English writing text data and general corpus text data. The historical raw data set also includes historical grammar error label data, grammar dependency relationship label data, and grammar rule data;
[0060] The general corpus text data is used to obtain statistical language features, specifically, it is a pure corpus resource obtained by acquiring canonical texts. The historical grammar error label data specifically includes grammar error type labels and grammar error location labels. The grammatical dependency relationship label data is specifically the grammatical subordination relationship labels between words in a sentence. The grammar rule data is specifically a predefined grammar rule set for representing grammar rules obtained based on the LanguageTool open-source library.
[0061] Embodiment 3, refer to Figure 1 and Figure 2 , based on the above embodiment, in the data preliminary processing module, the text data trimming is specifically to remove HTML tags, special symbols, and non-text content from the original data and unify the text format. The text vectorization is specifically to convert the text data into a structured vector using the BERT model. The data standardization is specifically to perform data standardization on the original data using the Z-Score standardization method. The dataset splitting is used to split the dataset;
[0062] Perform preliminary data processing on the current original dataset through the text data trimming, the text vectorization, and the data standardization to obtain the dataset to be detected. Perform preliminary data processing on the historical original dataset through the text data trimming, the text vectorization, the data standardization, and the dataset splitting to obtain the preliminary detection training set and the preliminary detection test set.
[0063] Embodiment 4, refer to Figure 1 and Figure 3 , based on the above embodiment, in the grammar error location model construction module, it is used to construct the model required for preliminary grammar error location. Specifically, a lightweight gradient boosting machine model is constructed as the grammar error location model, and the preliminary detection training set and the preliminary detection test set are used as the input of the grammar error location model to obtain the detection training set and the detection test set;
[0064] The grammar error location model construction module specifically includes fused feature construction, obtaining word weights, model construction and training, obtaining grammar error probabilities, and grammar error location;
[0065] The fused feature construction is used to construct a multi-dimensional grammatical feature representation. Specifically, a fused feature vector is obtained by combining predefined grammar rules with a statistical model. The content includes:
[0066] Rule feature extraction is used to capture explicit grammar rule conflicts. Specifically, a rule conflict score matrix is obtained through a predefined grammar rule set. The rule conflict score matrix is expressed as follows:
[0067] ;
[0068] Wherein, Rm represents the rule conflict score matrix, represents the conflict score of the first word under the first rule, represents the conflict score of the first word under the Bth rule, represents the conflict score of the Ath word under the first rule, represents the conflict score of the Ath word under the Bth rule, A represents the total number of words, and B represents the total number of rules;
[0069] Statistical feature generation, which is used to capture implicit language pattern anomalies. Specifically, through the n-gram language model, a statistical anomaly score vector is obtained. The formula used is as follows:
[0070] ;
[0071] Wherein, represents the average negative log probability of the language model within the window centered on the ath word, le represents the window radius, represents the probability of predicting the cth word based on the first le words calculated by the n-gram language model, Sv represents the statistical anomaly score vector, represents the average negative log probability of the language model within the window centered on the first word, represents the average negative log probability of the language model within the window centered on the second word, represents the average negative log probability of the language model within the window centered on the Ath word, and T represents the transpose operation;
[0072] Dynamic feature fusion, which is used to optimize the feature combination weights. Specifically, through the gating mechanism, the rule features and statistical features are fused to obtain the final fused features. The formula used is as follows:
[0073] ;
[0074] Wherein, Ft represents the final fused feature, represents the hyperbolic tangent function, represents the sigmoid function, represents the learnable rule weight matrix, represents the learnable statistical weight matrix, represents element-wise multiplication;
[0075] The obtaining of word weights is used to handle the data imbalance problem. Specifically, through inverse frequency weighting, the adjusted word weights are obtained. The formula used is as follows:
[0076] ;
[0077] In the formula, represents the word weight of the a-th word after adjustment, Ca represents the total number of word categories, represents the total number of words in the ca-th category;
[0078] The model is constructed and trained to construct a lightweight gradient boosting machine model as a grammar error localization model. Specifically, a lightweight gradient boosting machine model is constructed with a decision tree model as the basic learner, the model is trained based on the preliminary detection training set, and the model performance is verified based on the preliminary detection test set. The objective function is designed as follows:
[0079] ;
[0080] In the formula, represents the objective function value, represents the true label of the a-th word, represents the predicted probability of the a-th word, represents the regularization coefficient, represents the model parameter, represents calculating the Euclidean norm, represents the learnable probability prediction weight matrix, represents the final fused feature of the a-th word;
[0081] The obtaining of the grammar error probability is used to visualize potential grammar error positions. Specifically, through model inference, the position-level grammar error probability is obtained. The formula used is as follows:
[0082] ;
[0083] In the formula, represents the integrated grammar error probability of the a-th word, Num represents the number of basic learners, represents the predicted probability of the num-th basic learner for the a-th word, H represents the position-level grammar error probability, represents the integrated grammar error probability of the first word, represents the integrated grammar error probability of the second word, represents the integrated grammar error probability of the A-th word;
[0084] The grammar error localization is used to perform preliminary grammar error localization. Specifically, the preliminary detection training set and the preliminary detection test set are used as the inputs of the grammar error localization model, and the position-level grammar error probabilities based on the preliminary detection training set and the preliminary detection test set are obtained respectively, and are incorporated into the preliminary detection training set and the preliminary detection test set respectively to obtain the detection training set and the detection test set.
[0085] By performing the above operations, in view of the technical problems that the grammar error location technology in the traditional English writing grammar error detection system is mostly based on a single language model, the rule method is limited by the completeness of the preset rules and is difficult to cover flexible language phenomena, and the statistical method has insufficient parsing ability for long-distance dependencies and complex syntactic structures, this solution creatively adopts a lightweight gradient boosting machine model as the grammar error location model. By adaptively fusing rule conflicts and statistical pattern anomalies and combining the inverse frequency weighting strategy, while ensuring the effective detection of low-frequency grammar errors, the accuracy and interpretability of error location are improved.
[0086] Example 5, refer to Figure 1 and Figure 4 , based on the above example, in the grammar error classification model construction module, it is used to construct the model required for grammar error classification, specifically to construct a transformer model as the grammar error classification model;
[0087] The grammar error classification model construction module specifically includes dependency edge prediction, constructing node features, designing a graph attention layer, obtaining model output, and constructing and training the model;
[0088] The dependency edge prediction is used to capture grammar dependencies and determine the type and direction of the dependency relationship between words. Specifically, the input English writing text data is processed by a bidirectional long short-term memory network to obtain the dependency comprehensive feature, which is used as the edge feature for subsequent graph convolution operations. The formula used is as follows:
[0089] ;
[0090] In the formula, represents the dependency feature of the a-th word processed by the bidirectional long short-term memory network, represents the operation function of the bidirectional long short-term memory network, represents the input English writing text data, represents the score of the a-th word as the head node of the dependency relationship of the d-th word, represents the learnable head node score vector, represents the learnable head node score weight matrix, represents the dependency feature of the d-th word processed by the bidirectional long short-term memory network, represents the probability that there is a dependency relationship t between the a-th word and the d-th word, represents the softmax function, represents the learnable dependency type weight matrix, pr represents the predefined grammar rule embedding, represents the dependency comprehensive feature between the a-th word and the d-th word;
[0091] The above-mentioned structural node features are used to encode the word part-of-speech attributes as node features for subsequent graph convolution operations. Specifically, by looking up the word part-of-speech table and performing encoding, word part-of-speech features are obtained as node features. The formula used is as follows:
[0092] ;
[0093] In the formula, represents the word part-of-speech feature of the a-th word, represents the learnable part-of-speech weight matrix, represents the word part-of-speech label encoding of the a-th word;
[0094] The above-mentioned design graph attention layer is used to fuse syntactic structure information. Specifically, through an improved graph attention mechanism, with the dependency comprehensive feature as the edge feature and the word part-of-speech feature as the node feature, graph convolution propagation is performed to obtain the graph attention output feature. The content includes:
[0095] Construct an adjacency matrix, which is used to construct an adjacency matrix. Specifically, based on whether there is a dependency relationship between words and the direction of the dependency relationship, an adjacency matrix is constructed. If the dependency relationship exists, the corresponding position is 1; if the dependency relationship does not exist, the corresponding position is 0;
[0096] Design an improved graph attention mechanism to focus on key regions to obtain context-aware representations. Specifically, a mask matrix guided by the syntactic error position is constructed based on the position-level syntactic error probability, combined with a gating mechanism. The formula for the above-mentioned designed improved graph attention mechanism is as follows:
[0097] ;
[0098] In the formula, Qu represents the improved graph attention query, Ke represents the improved graph attention key, Va represents the improved graph attention value, represents the improved graph attention query transformation matrix, represents the improved graph attention key transformation matrix, represents the improved graph attention value transformation matrix, represents the improved graph attention input, Mh represents the mask matrix guided by the syntactic error position, represents the element of the mask matrix guided by the syntactic error position, represents the integrated syntactic error probability of the d-th word, Fa represents the output feature of the improved graph attention mechanism, represents the dimension of the improved graph attention key;
[0099] Graph convolution propagation is used to aggregate domain information. Specifically, through message passing, node update features are obtained. The formula used is as follows:
[0100] ;
[0101] In the formula, represents the update feature of the l+1th layer node, represents the graph convolution operation function, Am represents the adjacency matrix, represents the update feature of the l-th layer node, represents the learnable graph convolution weight matrix of layer l, represents the initial node feature, and Fp represents the word part-of-speech feature;
[0102] The model output is obtained to inject constraints and obtain the model output. Specifically, the position-level grammatical error probability is used as the constraint injection graph attention output feature. After being processed by the transformer model, the model output grammatical error classification result is obtained. The formula used is as follows:
[0103] ;
[0104] In the formula, Represents the grammatical error classification result, represents the model output weight matrix, represents the transformer operation function, Fg represents the graph attention output feature, Represents the model output bias term;
[0105] The constructing and training the model specifically comprises constructing a transformer model through the dependency edge prediction, the construction node features, the design graph attention layer and the acquisition model output, training the model based on the detection training set, verifying the model performance based on the detection test set, and obtaining the transformer model as a grammatical error classification model;
[0106] Further, Table 1 is a model performance comparison table of the transformer model described in this scheme and the traditional model. As shown in the table, the transformer model adopts dependency edge prediction, constructs node features and designs graph attention layers to model the dependency relationship between words and introduces graph structure constraints, which solves the problem that traditional methods are prone to misjudgment of errors with tight dependency relationships, lack of consideration of the collateral effects of errors, and difficulty in processing complex long sentences. The model performance is better than that of traditional models. The traditional models specifically include support vector machine models, traditional transformer models and long short-term memory network models;
[0107] Table 1 shows the model accuracy of the transformer model and the traditional model on a dataset containing five types of English writing grammatical errors, where SVA represents the consistency error between the subject and the predicate verb in number and person, TENSE represents the error of improper or inconsistent use of verb tense, ARTICLE represents the error of article use, PREP represents the error of preposition use, and PRON represents the error of pronoun use;
[0108] Table 1 Comparison of model performance between the transformer model described in this scheme and the traditional model
[0109]
[0110] By performing the above operations, the traditional English writing grammatical error detection system often treats grammatical errors as independent labels when classifying grammatical errors, ignoring the inherent correlation of syntactic structures in English grammar, easily misjudging errors with close dependencies, and lacks consideration of the collateral effects of errors, making it difficult to handle technical problems with complex long sentences. This solution creatively uses the transformer model as a grammatical error classification model. By modeling the dependency relationship between words and introducing graph structure constraints, it can capture the impact of local grammatical errors on global syntax. At the same time, the positioning probability is injected into the model as a structured prior, realizing the coordinated optimization of positioning and classification, and improving the accuracy of classification and context consistency.
[0111] Example 6, see Figure 1 This embodiment is based on the above embodiment. In the grammatical error detection module, specifically, the grammatical error localization model is used to perform preliminary grammatical error localization on the data set to be detected. The obtained position-level grammatical error probability and the data set to be detected are used together as inputs of a grammatical error classification model to obtain a grammatical error classification result. Based on the position-level grammatical error probability and the grammatical error classification result, a reference result for grammatical error detection in English writing is obtained. The reference result for grammatical error detection in English writing specifically includes the position-level grammatical error probability output by the grammatical error localization model with the data set to be detected as input, and the grammatical error classification result output by the grammatical error classification model with the position-level grammatical error probability obtained based on the data set to be detected and the data set to be detected as inputs.
[0112] It should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device.
[0113] While the embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that many changes, modifications, substitutions and variations can be made to the embodiments without departing from the principles and spirit of the invention.
[0114] The above description of the present invention and its implementation manners is not restrictive. What is shown in the drawings is only one of the implementation manners of the present invention, and the actual structure is not limited thereto. Generally speaking, if those of ordinary skill in the art are inspired by it and design, without creative efforts, structural manners and embodiments similar to the technical solution without departing from the gist of the present invention, they shall fall within the protection scope of the present invention.
Claims
1. A deep learning-based English writing grammatical error detection system, characterized by: The system includes an original data acquisition module, a data preliminary processing module, a grammatical error location model construction module, a grammatical error classification model construction module and a grammatical error detection module; The original data acquisition module acquires an error detection original data set by performing data acquisition, and the error detection original data set specifically includes a historical original data set and a current original data set; The data preliminary processing module adopts a data preliminary processing method of text data trimming, text vectorization, data standardization and data set segmentation to obtain a data set to be detected, a preliminary detection training set and a preliminary detection test set; The grammatical error localization model construction module is used to construct a model required for preliminary grammatical error localization, specifically to construct a lightweight gradient boosting machine model as a grammatical error localization model, and use the preliminary detection training set and the preliminary detection test set as inputs of the grammatical error localization model to obtain a detection training set and a detection test set, and the specific contents include fusion feature construction, obtaining word weights, model construction and training, obtaining grammatical error probability and grammatical error localization; The fusion feature construction is used to construct a multi-dimensional grammatical feature representation, specifically by combining predefined grammatical rules with statistical models to obtain a fusion feature vector; The grammatical error classification model construction module is used to construct a model required for grammatical error classification, specifically to construct a transformer model as a grammatical error classification model, and the transformer model is specifically a hybrid neural network model based on a standard transformer architecture and integrating a graph attention mechanism and constraint injection; The grammatical error detection module performs grammatical error detection by adopting the grammatical error location model and the grammatical error classification model to obtain a reference result of English writing grammatical error detection.
2. The English writing grammatical error detection system based on deep learning according to claim 1, characterized in that: The fusion feature structure includes: Rule feature extraction is used to capture explicit grammar rule conflicts. Specifically, a rule conflict score matrix is obtained through a predefined grammar rule set. Statistical feature generation is used to capture implicit language pattern anomalies. Specifically, the statistical anomaly score vector is obtained through the n-gram language model. The formula used is as follows: ; In the formula, represents the average negative logarithmic probability of the language model in the window centered on the ath word, le represents the window radius, represents the probability of predicting the cth word based on the first le words calculated by the n-gram language model, Sv represents the statistical anomaly score vector, represents the average negative log probability of the language model in the window centered on the first word, represents the average language model negative log probability in the window centered on the second word, represents the average negative log probability of the language model in the window centered on the Ath word, A represents the total number of words, and T represents the transposition operation; Dynamic feature fusion is used to optimize the feature combination weights. Specifically, it fuses rule features and statistical features through a gating mechanism to obtain the final fused features. The word weights are obtained to deal with the data imbalance problem, specifically by inverse frequency weighting to obtain adjusted word weights; The model is constructed and trained to construct a lightweight gradient boosting machine model as a grammatical error localization model, specifically, a lightweight gradient boosting machine model is constructed with a decision tree model as a basic learner, the model is trained based on the preliminary detection training set, and the model performance is verified based on the preliminary detection test set; The grammatical error probability is obtained to visualize the potential grammatical error position, specifically, the position-level grammatical error probability is obtained by inferring the grammatical error location model, and the formula used is as follows: ; In the formula, represents the probability of ensemble grammatical errors of the ath word, Num represents the number of basic learners, represents the prediction probability of the num-th basic learner for the a-th word, H represents the position-level grammatical error probability, represents the integrated grammatical error probability of the first word, represents the integrated grammatical error probability of the second word, represents the integrated grammatical error probability of the Ath word; The grammatical error location is used to perform preliminary grammatical error location, specifically using the preliminary detection training set and the preliminary detection test set as inputs of the grammatical error location model, respectively obtaining the position-level grammatical error probability based on the preliminary detection training set and the position-level grammatical error probability based on the preliminary detection test set, and merging them into the preliminary detection training set and the preliminary detection test set, respectively, to obtain the detection training set and the detection test set.
3. The English writing grammatical error detection system based on deep learning according to claim 1 is characterized by: The grammatical error classification model construction module specifically includes dependency edge prediction, node feature construction, graph attention layer design, model output acquisition, and model construction and training.
4. The English writing grammatical error detection system based on deep learning according to claim 3 is characterized by: The dependency edge prediction is used to capture grammatical dependencies and determine the type and direction of dependency relationships between words. Specifically, the input English writing text data is processed through a bidirectional long short-term memory network to obtain dependency comprehensive features as edge features for subsequent graph convolution operations. The formula used is as follows: ; In the formula, represents the dependency features of the ath word processed by the bidirectional long short-term memory network, Represents the bidirectional long short-term memory network operation function, Indicates input English writing text data, represents the score of the a-th word as the head node of the d-th word dependency relationship, represents the hyperbolic tangent function, represents the learnable head node score vector, represents the learnable head node score weight matrix, represents the dependency features of the dth word processed by the bidirectional long short-term memory network, represents the probability that there is a dependency relationship t between the a-th word and the d-th word, represents the softmax function, represents the learnable dependency type weight matrix, pr represents the predefined grammatical rule embedding, Represents the comprehensive dependency features between the a-th word and the d-th word; The constructed node features are used to encode word part-of-speech attributes as node features for subsequent graph convolution operations, specifically by searching a word part-of-speech table and encoding to obtain word part-of-speech features as node features; The designed graph attention layer is used to integrate grammatical structure information. Specifically, through the improved graph attention mechanism, the dependency comprehensive features are used as edge features, and the word part-of-speech features are used as node features. Graph convolution propagation is performed to obtain graph attention output features, including: Constructing an adjacency matrix, which is used to construct an adjacency matrix. Specifically, the adjacency matrix is constructed based on whether there is a dependency relationship between words and the direction of the dependency relationship. If the dependency relationship exists, the corresponding position is 1, and if the dependency relationship does not exist, the corresponding position is 0; An improved graph attention mechanism is designed to focus on key areas to obtain context-aware representation. Specifically, a mask matrix guided by the position-level grammatical error probability is constructed. Combined with the gating mechanism, the improved graph attention mechanism is designed. The formula used is as follows: ; Where Qu represents the improved graph attention query, Ke represents the improved graph attention key, and Va represents the improved graph attention value. represents the improved graph attention query transformation matrix, represents the improved graph attention key transformation matrix, Represents the improved graph attention value transformation matrix, represents the improved graph attention input, Mh represents the mask matrix guided by the grammatical error position, The elements of the mask matrix representing the syntax error position guide, represents the integrated grammatical error probability of the dth word, Fa represents the output feature of the improved graph attention mechanism, represents the dimension of the improved graph attention key; Graph convolution propagation is used to aggregate domain information, specifically to obtain node update features through message passing; The obtaining of model output is used to inject constraints and obtain model output, specifically injecting position-level grammatical error probability as a constraint into the graph attention output feature, and obtaining the model output grammatical error classification result after being processed by the transformer model; The constructing and training model specifically involves constructing a transformer model through the dependency edge prediction, the construction node features, the design graph attention layer and the acquisition model output, training the model based on the detection training set, and verifying the model performance based on the detection test set to obtain the transformer model as a grammatical error classification model.
5. The English writing grammatical error detection system based on deep learning according to claim 1, characterized in that: In the original data collection module, both the historical original data set and the current original data set include English writing text data and general corpus text data. The historical original data set also includes historical grammatical error label data, grammatical dependency label data and grammatical rule data.
6. The English writing grammatical error detection system based on deep learning according to claim 1, characterized in that: In the data preliminary processing module, the text data trimming is specifically to remove HTML tags, special symbols and non-text content in the original data and unify the text format, the text vectorization is specifically to convert the text data into a structured vector using the BERT model, the data standardization is specifically to standardize the original data using the Z-Score standardization method, and the data set segmentation is used to segment the data set; The current original data set is preliminarily processed through the text data trimming, the text vectorization and the data standardization to obtain the data set to be detected, and the historical original data set is preliminarily processed through the text data trimming, the text vectorization, the data standardization and the data set segmentation to obtain a preliminary detection training set and a preliminary detection test set.
7. The English writing grammatical error detection system based on deep learning according to claim 1, characterized in that: In the grammatical error detection module, specifically, the grammatical error localization model is used to perform preliminary grammatical error localization on the data set to be detected, and the obtained position-level grammatical error probability is used together with the data set to be detected as the input of the grammatical error classification model to obtain a grammatical error classification result, and based on the position-level grammatical error probability and the grammatical error classification result, a reference result of grammatical error detection in English writing is obtained comprehensively. The reference result of grammatical error detection in English writing specifically includes the position-level grammatical error probability output by the grammatical error localization model with the data set to be detected as input, and the grammatical error classification result output by the grammatical error classification model with the position-level grammatical error probability obtained based on the data set to be detected and the data set to be detected as input.
Citation Information
Patent Citations
Policy text analysis method based on policy text classification and key information identification
CN115310425A
Text error correction method and device based on intention consistency and medium
CN116136957A
BERT and LightGBM integrated mental health intelligent monitoring method and system
CN118919023A
Method and system for automatic search and correction of errors in texts in natural language
RU2785207C1
Cited By
Digital auxiliary repairing system for historical building
CN120764047A
A digitalization-assisted repair system for historic buildings
CN120764047B
Machine learning grammar rule generation
US12572734B1