A method, system, storage medium and terminal for automatic error correction of NOTAM text
By constructing a knowledge graph of pronunciation and glyphs and combining it with the CKBERT model for multi-hop knowledge comparative learning, the problem of insufficient expressiveness of the existing NOTAM text correction model is solved, and more efficient correction performance is achieved, which is suitable for automatic error correction of NOTAM text.
Patent Information
- Application Number
- CN202510940538.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-09
- Publication Date
- 2025-09-23
- Estimated Expiration
- 2045-07-09
AI Technical Summary
The existing NOTAM text error correction model has deficiencies in expressiveness and error correction performance, especially when processing Chinese NOTAM text. It lacks fine-grained expression of Chinese character pronunciation and glyphs, and has limited graph structure and scalability, resulting in low error correction efficiency and high error rate.
Construct a knowledge graph of pronunciation and glyphs, combine it with the CKBERT masked language model, and through multi-hop knowledge comparison learning tasks, integrate pronunciation similarity and glyph similarity to build a similar vocabulary. Use the training model to correct the error model and perform error correction by extracting the navigation notice text.
It improves the error correction efficiency of NOTAM texts, reduces the error rate, and enhances the error correction performance of the model, making it suitable for the development of digital intelligence.
Smart Images

Figure CN120430300B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of text processing, and in particular to a method, system, storage medium and terminal for automatic error correction of NOTAM text. Background Art
[0002] With the rapid development of the aviation industry, the volume of data required to process civil aviation intelligence has increased exponentially, and the complexity of its content has also increased, making it even more challenging to process. Notices to airmen (NOTAMs) are a crucial information transmission mechanism in civil aviation, communicating important developments affecting flight safety to air traffic controllers and pilots to prevent potential hazards. Item E (the main text of the NOTAM) in NOTAMs is written in plain language and abbreviated characters. Typos in Item E of Chinese NOTAMs can cause semantic changes and seriously impact flight safety. The unique structure of Chinese characters in both pronunciation and shape presents unique challenges in the compilation and publication of Chinese NOTAMs. With the rapid growth in the number of NOTAMs, current manual processing of NOTAM text is inefficient and error-prone, making existing processing methods inadequate for aviation development.
[0003] The development and application of deep learning have driven further research in the field of text correction. Currently, Chinese text correction has evolved from focusing solely on character correction to studying the semantic relationship between Chinese character pronunciation and words and sentences. Due to the complexity of character features and the diversity of semantics, Chinese text correction is challenging. Some existing technologies have proposed an NMT-based GEC model that dynamically adds random masks to the original source sentences during training, improving the generalization ability of the grammatical correction model. Other technologies have proposed splitting Chinese characters into phonetic and character structures, constructing a hybrid graph-based and rule-based model to improve the efficiency of Chinese spelling checking.
[0004] The emergence of knowledge graphs has brought new approaches to Chinese spelling correction and improved the performance of Chinese text correction models. For example, some methods construct knowledge graphs based on the pronunciation and glyphs of Chinese characters, use the Node2Vec model to generate pronunciation and glyph vectors, combine them with CNN to determine the similarity between pronunciation and glyphs, and then train them using the BERT model. Other methods propose using an improved MacBERT model to correct Chinese NOTAM text errors, using synonyms to mask the original text to improve the performance of text correction models.
[0005] However, the model expression ability and error correction performance of existing text error correction models can be further improved. For example, the improved MacBERT model lacks an enhancement mechanism for relationship modeling and contrastive training. Its embedding method is only based on the combination of BERT embedding and CNN, and has not formed an end-to-end fusion network. At the same time, it does not include features such as tone in terms of pronunciation, does not achieve fine-grained data structure expression, and has relatively limited graph structure and scalability. Therefore, there is an urgent need to improve the existing text error correction model to further improve the error correction performance of navigation notice text. Summary of the Invention
[0006] The purpose of the present invention is to overcome the technical problems existing in the prior art and provide a method, system, storage medium and terminal for automatically correcting errors in navigation notice texts.
[0007] The object of the present invention is achieved through the following technical solutions:
[0008] In a first aspect, a method for automatically correcting errors in NOTAM text is provided, comprising the following steps:
[0009] S1. Extract Chinese characters from item E of the NOTAM text and build a Chinese character database based on common Chinese characters.
[0010] S2. Calculate the pronunciation similarity of characters based on the Chinese character database and construct a pronunciation knowledge graph; calculate the glyph similarity of characters based on the Chinese character database and construct a glyph knowledge graph;
[0011] S3, fusing the pronunciation similarity and the shape similarity to obtain a similar word library;
[0012] S4. Build and train an automatic error correction model based on the similar word library. The automatic error correction model replaces the mask with similar words in the masked language model of CKBERT, and constructs positive and negative samples to be added to the multi-hop knowledge comparison learning task;
[0013] S5. Use the trained automatic error correction model to correct the navigation notice text.
[0014] In some embodiments, extracting Chinese characters from item E of the NOTAM text includes:
[0015] De-noise the original data of NOTAM text;
[0016] Save the Chinese telegraph code of item E and the corresponding Chinese characters as a mapping file, and use regular expressions to standardize the telegraph code into Chinese characters;
[0017] Then extract the structured information of the text of item E to generate a new file.
[0018] In some embodiments, the calculating of the pronunciation similarity of characters based on the Chinese character database and constructing a pronunciation knowledge graph includes:
[0019] The Cypher language in the Neo4j tool is used to establish the nodes, attributes and relationships of the phonetic features to form a phonetic knowledge graph;
[0020] The continuous bag-of-words model is used as the character embedding model for training to obtain the word-phonetic vector;
[0021] The pronunciation vector is used as one input of the RGCN model, and the pronunciation knowledge graph is used as another input of the RGCN model, and the RGCN model is jointly trained in an unsupervised manner.
[0022] In some embodiments, the unsupervised training of the RGCN model includes:
[0023] The node feature update formula of RGCN is
[0024] ,in For nodes In the The characteristics of the layer, σ is the activation function (ReLU), For the relationship Next node The set of neighbor nodes of is the normalization coefficient, is the set of all relationship types in the graph, For nodes In the Layer characteristics, For nodes In the Layer characteristics, For the Corresponding relationship type in the layer The weight matrix, It is The self-loop weight matrix in the layer;
[0025] The mean square error is used to calculate the loss function value. The formula is as follows:
[0026] ,in is the number of nodes, is the similarity, is the generated similarity.
[0027] In some embodiments, the calculating of glyph similarity based on the Chinese character database and constructing a glyph knowledge graph includes:
[0028] The Chinese characters in the Chinese character database are split into radicals and classified according to the radicals; and the Glyce model is used to calculate the similarity.
[0029] In some embodiments, before building the automatic error correction model, the method further includes:
[0030] Mark the text in item E as incorrect;
[0031] The incorrectly annotated text is segmented and annotated with entities and relationships using the LTP tool;
[0032] The text after entity and relationship annotation is converted into JSON file format and processed into triples according to the knowledge graph.
[0033] In some embodiments, the loss function of the masked language model is
[0034] ,in, is the loss value of the masked language model, To cover the token collection, For the collection of uncovered tokens, is the number of tokens to be covered, Expressing sentences The index of the masked token, For the A masked token, For the The predicted probability of the masked token, represents the parameter set of the model;
[0035] The loss function used in the multi-hop knowledge contrastive learning task is
[0036] ,in, represents the loss value of the multi-hop knowledge contrast learning task, represents a positive sample triple, represents a negative sample triplet, is the hidden representation of the entity, is the temperature coefficient, is the cosine similarity of the positive sample group, which is used to indicate the feature similarity between the entity and the positive sample. is the cosine similarity of the negative sample group, which is used to indicate the feature similarity between the entity and the negative sample.
[0037] In a second aspect, a system for automatically correcting NOTAM text errors is provided, comprising:
[0038] A Chinese character database construction module is used to extract Chinese characters from item E of the NOTAM text and construct a Chinese character database based on common Chinese characters;
[0039] A knowledge graph construction module, which calculates the pronunciation similarity of characters based on the Chinese character database and constructs a pronunciation knowledge graph, and calculates the glyph similarity of characters based on the Chinese character database and constructs a glyph knowledge graph;
[0040] A similar word library construction module is used to integrate the pronunciation similarity and the shape similarity to obtain a similar word library;
[0041] An automatic error correction model construction and training module is used to construct and train an automatic error correction model based on the similar word library. The automatic error correction model uses similar words to replace masks in the masked language model of CKBERT, and constructs positive and negative samples to be added to the multi-hop knowledge comparison learning task;
[0042] The error correction module is used to correct the navigation notice text using the trained automatic error correction model.
[0043] In a third aspect, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method for automatically correcting navigation notice texts described in the first aspect is implemented.
[0044] In a fourth aspect, a terminal is provided, comprising a memory and a processor, wherein the memory stores computer instructions that can be executed on the processor, and when the processor executes the computer instructions, the method for automatically correcting navigation notice texts described in the first aspect is executed.
[0045] It should be further explained that the technical features corresponding to the above embodiments can be combined or replaced with each other to form a new technical solution if there is no conflict.
[0046] Compared with the prior art, the present invention has the following beneficial effects:
[0047] 1. This paper proposes an improved model based on the CKBERT framework and introduces a multi-hop knowledge comparison learning mechanism. During the pre-training phase, it not only relies on context masks but also further optimizes the model's expressiveness by constructing positive and negative samples. This improves the model's error correction performance, increases the error correction efficiency of NOTAM text, and reduces the error rate, providing technical support for the development of digital intelligence.
[0048] 2. This invention combines the phonetic vectors obtained from CBOW model training with the RGCN model to form CEV-RGCN. This improves the model's understanding of differences in phonetic structure, such as initials and finals. It also provides detailed supplementary information for graph structure feature data, improving the expressiveness of features in the graph data. This allows the model to better capture phonetic characteristics while ensuring training efficiency and accuracy. This allows for deep embedding of complex phonetic graphs, forming an end-to-end fusion network.
[0049] 3. This paper uses Neo4j and Cypher language to construct a structured triple graph, with nodes covering features such as initials, finals, final tones and radicals, thus achieving fine-grained data structure expression.
[0050] 4. The present invention uses cosine similarity to evaluate the fusion vector and adds the mean square error loss function to control the accuracy during the training process. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] Figure 1 This is a flow chart of a method for automatically correcting errors in NOTAM texts according to the present invention;
[0052] Figure 2 This is the original structure diagram of the NOTAM text of the present invention;
[0053] Figure 3 This is a schematic diagram of the text of Item E of the Chinese NOTAM of the present invention;
[0054] Figure 4 Schematic diagram of the pronunciation knowledge map of the present invention;
[0055] Figure 5 Schematic diagram of the CEV-RGCN model structure of the present invention;
[0056] Figure 6 Schematic diagram of the Glyce-BERT model structure of the present invention;
[0057] Figure 7 Schematic diagram of the glyph knowledge graph of the present invention;
[0058] Figure 8 A schematic diagram of entities and relationships of the present invention;
[0059] Figure 9 Schematic diagram of the Ctc-CKBERT model structure of the present invention;
[0060] Figure 10 This is a three-dimensional graph comparing the F1 values of the models of the present invention. DETAILED DESCRIPTION
[0061] The technical solutions of the present invention are described clearly and completely below in conjunction with the accompanying drawings. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. The components of the embodiments of the present application generally described and shown in the drawings herein can be arranged and designed in various different configurations. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.
[0062] It should be noted that the defects existing in the solutions in the above-mentioned prior art are the results obtained by the inventor after practice and careful research. Therefore, the discovery process of the above-mentioned problems and the solutions proposed in the embodiments of this application below for the above-mentioned problems should be the contributions made by the inventor to this application in the process of invention and creation, and should not be understood as technical contents known to technical personnel in this field.
[0063] In response to the technical problems pointed out in the background technology, the embodiments provided by the present invention are as follows:
[0064] In an exemplary embodiment, referring to Figure 1 A method for automatically correcting errors in NOTAM texts comprises the following steps:
[0065] S1. Extract Chinese characters from item E of the NOTAM text and build a Chinese character database based on common Chinese characters.
[0066] S2. Calculate the pronunciation similarity of characters based on the Chinese character database and construct a pronunciation knowledge graph; calculate the glyph similarity of characters based on the Chinese character database and construct a glyph knowledge graph;
[0067] S3, fusing the pronunciation similarity and the shape similarity to obtain a similar word library;
[0068] S4. Build and train an automatic error correction model based on the similar word library. The automatic error correction model replaces the mask with similar words in the masked language model of CKBERT, and constructs positive and negative samples to be added to the multi-hop knowledge comparison learning task;
[0069] S5. Use the trained automatic error correction model to correct the navigation notice text.
[0070] The NOTAM samples from the Civil Aviation Information Center from September 2021 to September 2023 are used as the original data for the Chinese NOTAM text correction in this invention. The text contains item Q: limiting item, item A: place of occurrence, item B: effective time, item C: expiration time, item D: segment time, item E: navigation notice text, item F: upper limit and item G: lower limit. Its structure is as follows: Figure 2The Chinese characters used in NOTAM are mainly concentrated in the text of item E. Due to the professional nature of NOTAM text, item E adopts Chinese Commercial Code as the editing and issuing standard. The present invention extracts the text of item E of NOTAM and translates it into Chinese as shown in the following figure: Figure 3 shown.
[0071] In step S1, the original NOTAM data is first subjected to noise reduction to remove line breaks and redundant information from the first four lines of data. The Chinese telegraph code and the corresponding Chinese characters are then stored as a mapping file, and the telegraph code is standardized into Chinese characters using regular expressions. Structured information such as the text in item E is then extracted to generate a new file. Finally, each piece of information is categorized, and the time and longitude and latitude are converted into a common format for easy understanding, resulting in a data file containing the date, longitude, latitude, text, and airport. The present invention constructs a database containing 11,000 Chinese characters based on the extracted Chinese characters from the text in item E of the Chinese NOTAM and in combination with commonly used Chinese characters in daily use, for constructing a knowledge graph and improving subsequent models.
[0072] In step S2, the pronunciation similarity of characters is calculated based on the Chinese character database and a pronunciation knowledge graph is constructed, which specifically includes:
[0073] 1. Data Processing
[0074] We correctly annotated the pronunciations of 11,000 Chinese characters in the Chinese character database, splitting the characters' pronunciations according to the relationships in Table 1. We then divided the characters into initials, finals, and final tones, annotated them according to similar types, and constructed a character pronunciation database. An example of character pronunciation splitting is shown in Table 1.
[0075] Table 1 Example of word-phonetic splitting
[0076]
[0077] 2. Pronunciation Knowledge Graph
[0078] Knowledge graphs describe concepts, entities, and their relationships in the form of structured triples. This paper uses pronunciations, initials, finals, and final tones as nodes, and similarity types as edges. The Cypher language within the Neo4j tool is used to establish nodes, attributes, and relationships for pronunciation features. Cypher is typically constructed using three code structures: node creation, relationship assignment, and query. Table 2 shows some examples of Cypher code.
[0079] Table 2 Cypher language examples
[0080]
[0081] Import the pronunciation database into Neo4j in batches to form a pronunciation knowledge graph. You can query any data in the knowledge graph through the query code. Some examples of the pronunciation knowledge graph are as follows: Figure 4 As shown in the figure, due to the large number of nodes and edges in the knowledge graph, in order to avoid graph redundancy, the example figure only shows some relationships of one tone.
[0082] 3. Phonetic Similarity Model
[0083] In order to improve the accuracy of similarity calculation of word pronunciation, the present invention adopts the continuous bag of words (CBOW) model in word2vec as a character embedding model for training to obtain word pronunciation vectors, and the obtained word pronunciation vectors are used as an input of the RGCN model. The word pronunciation knowledge graph is then used as graph structure data as another input of the RGCN model, and the RGCN model is jointly trained in an unsupervised manner. The present invention adopts a common formula for similarity calculation, namely the cosine similarity formula. RGCN is used to process heterogeneous graphs of the knowledge graph type, where the nodes in the graph are word pronunciations and the edges are the relationships between word pronunciations. During the training process, the RGCN model uses the word pronunciation nodes in the word pronunciation knowledge graph as each other's context, combines the edge relationships between nodes with the feature information in the word pronunciation vectors, and comprehensively learns the word pronunciation features. In order to output the similarity of word pronunciation and simultaneously evaluate the performance of the model, the present invention adds similarity calculation after the output layer of the model.
[0084] Specifically, the phonetic vector of the character embedding is obtained after training the phonetic database through the CBOW model. The vector data and the graph structure data are used as the input layer of the RGCN model for training. The phonetic vector is input into the RGCN model as a supplementary vector, and the improved RGCN model with the character embedding vector as the supplementary vector is obtained, which is customized as CEV-RGCN. The structure of the CEV-RGCN model is as follows Figure 5 shown.
[0085] The CBOW model is divided into input layer, projection layer and output layer. The CBOW model takes the pronunciation of characters as the smallest language learning unit, and each unit is mapped to an independent embedding vector. and The context vector is accumulated as shown below:
[0086]
[0087] in is the size of the context window. Indicates the central word The context vector of Corresponding to the first The word embedding vector is generated by the cumulative vector input from the projection layer. Vector. The present invention uses the Sigmoid function to generate vectors. As shown in the following formula:
[0088]
[0089] in, for The probability of the vector appearing, is the weight matrix of the output layer, is the embedding vector of the target word, is the matching degree between the context and the target word, represents the dot product, Represents the output layer vector of the target word, each item is a Sigmoid function whose output is a probability value. This is a well-known Sigmoid function calculation and will not be detailed here. After calculating the pronunciation vector, it is input into the input layer of the RGCN model and combined with the graph data of the knowledge graph for training. During the learning process, the RGCN node feature update formula is as follows:
[0090]
[0091] in For nodes In the The characteristics of the layer, σ is the activation function (ReLU), For the relationship Next node The set of neighbor nodes of is the normalization coefficient, is the set of all relationship types in the graph, For nodes In the Layer characteristics, For nodes In the Layer characteristics, For the Corresponding relationship type in the layer The weight matrix, It is The self-loop weight matrix in the layer is used to maintain the node's own characteristics.
[0092] Before the downstream task, the node vector with weights needs to be added to the word vector, as shown in the following formula:
[0093]
[0094] in is the weighted fusion coefficient, The present invention uses the mean square error (MSE) to calculate the loss function value to ensure that the CEV-RGCN model can effectively retain the characteristic information of the pronunciation, as shown in the following formula:
[0095]
[0096] in is the number of nodes, is the similarity, is the generated similarity.
[0097] 4. Model Experiment
[0098] The pronunciation database is input into the CBOW model for training to generate pronunciation vectors. These vectors are then used as supplementary vectors and the graph data from the knowledge graph as the input layer for training the CEV-RGCN model. The data is divided into a 6:2:2 ratio for training, testing, and validation sets. Table 3 shows examples of pronunciation similarity calculation results. Table 4 shows pronunciation similarity calculated using the RGCN model using only graph data as input.
[0099] Table 3 Phonetic similarity
[0100]
[0101] Table 4: Grapheme-phonetic similarity (RGCN)
[0102]
[0103] The performance parameters of the model training are shown in Table 5.
[0104] Table 5 Model performance parameters
[0105]
[0106] The phonetic similarity results in Tables 3 and 4 show that the improved CEV-RGCN model achieves higher similarity scores than the RGCN model. The RGCN model scores "bào" and "bāo" at 0.81, while the improved model achieves a score of 0.87. Regarding the negative similarity results, the RGCN model scores "zhī" and "bā" at -0.63, while the improved model scores at -0.68. Higher scores indicate poorer discrimination ability in negative similarity.
[0107] Comparing the performance parameters of the RGCN model with those of the present invention's model, the improved CEV-RGCN model outperforms the RGCN model in both accuracy and precision in calculating word-phonetic similarity. Finally, the resulting word-phonetic similarity is marked in the knowledge graph to facilitate subsequent Chinese text error correction.
[0108] In step S2, the similarity of glyphs is calculated based on the Chinese character database and a glyph knowledge graph is constructed, which specifically includes:
[0109] 1. Data Processing
[0110] The present invention splits Chinese characters in the Chinese character database into radicals, classifies the characters according to the radicals, and constructs a glyph database to enhance character features and improve model calculation efficiency. The Chinese character splitting is shown in Table 6.
[0111] Table 6 Examples of glyph features
[0112]
[0113] The glyph structure contains more complex information such as strokes and radicals. In order to effectively identify the characteristics of glyphs, the present invention uses 16 fonts such as Traditional Chinese, Chinese Lishu and Songti, and a total of 176,000 glyph training images for each font of the 11,000 Chinese characters in the glyph database to enrich the characteristics of Chinese characters and meet the requirements of the model.
[0114] 2. Glyph Similarity Model
[0115] In natural language processing tasks, leveraging Chinese character glyph information can improve the accuracy of Chinese text processing. This paper employs the Tianzege-CNN model for training. After inputting a Chinese character image, a convolutional layer captures the character's pictographic features. A maximum pooling layer then reduces the image resolution to a 2×2 grid, capturing information about the glyph's structure.
[0116] Furthermore, based on the open source Glyce toolkit, Glyce-BERT was trained using the provided MSR dataset, and its accuracy, recall rate, and F1 value reached 98.2%, 98.3%, and 98.3%. The similarity task was used as the downstream task of the model to calculate the similarity of Chinese characters. The Glyce-BERT model structure is as follows: Figure 6 shown.
[0117] 3. Model Experiment
[0118] The image was fed into the Glyce-BERT model as the input layer to obtain the model's performance parameters, and then similarity was calculated. The obtained performance parameters and glyph similarity are shown in Tables 7 and 8.
[0119] Table 7 Glyce model performance parameters
[0120]
[0121] Table 8 Glyph similarity
[0122]
[0123] 4. Chinese Character Shape Knowledge Graph
[0124] Since a model different from the pronunciation model is adopted, the similarity of Chinese character shapes is obtained after training, and the Chinese character shape knowledge graph is constructed according to the shape similarity. In order to better display the relationship between Chinese character shape structures, the present invention splits Chinese characters and uses each as a node. The Chinese character shape knowledge graph is as Figure 7 shown. The structures of pronunciation and shape are presented in the form of a knowledge graph, intuitively showing the relationship between Chinese characters.
[0125] CKBERT is a model proposed by the Alibaba Cloud team for NLP. This model includes a Linguistic-aware Masked Language Model (LMLM) and a Contrastive Multi-hop Relation Modeling (CMRM) task. The combination of knowledge graph information and conventional data can effectively inject relationship and language knowledge into CKBERT, and it is superior to benchmark NLP tasks and different language training models in terms of performance.
[0126] In step S4, an automatic error correction model is constructed and trained according to the similarity dictionary, specifically including:[[]]
[0127] 1. Data Preparation and Processing
[0128] Since the text of item E in Chinese NOTAM contains Chinese characters, letters, symbols, etc., the present invention deletes all characters other than Chinese characters in the text of item E to ensure that there is no interference from other factors during subsequent Chinese error correction. The processing results are shown in Table 9.
[0129] Table 9 Comparison of the Text of Item E Before and After Processing
[0130]
[0131] The processed text of item E is organized into a Chinese database of the text of item E, and then error marking is performed as shown in Table 10.
[0132] Table 10 Example of Error Marking
[0133]
[0134] After marking, the LTP tool is used to segment the data and mark entities and relationships as Figure 8 shown. Then it is converted into a JSON file format suitable for the model, and the data is processed into triple form according to the knowledge graph, such as "foot - bag - run". The Chinese character knowledge graph is constructed with positive and negative samples in a ratio of 1:3 for model training. Examples of positive and negative samples are shown in Table 11.
[0135] Table 11 Examples of Positive and Negative Samples
[0136]
[0137] The previously obtained pronunciation and glyph similarity are combined to create a similarity database, which is then incorporated into the knowledge graph to simplify the subsequent model training process. Specifically, the pronunciation and glyph similarity between characters are used to construct a triplet of the form "character A - character B - similarity type (sound / shape / double)". The specific similarity value is annotated in the "similarity type", thus generating a similarity database that integrates sound and shape information. This is used for candidate character generation and context judgment in the Chinese error correction model.
[0138] 2. Model Building
[0139] The present invention makes improvements in the LMLM of the CKBER model. The similar word library constructed in the previous experiment is used to calculate the similar word package using the word2vec model, and the similar word package is used as a replacement for the mask token [MASK] to cover the original word, which enhances the consistency of the context within the scope and solves the performance difference caused by the model using the same token in different text processing. Then, the positive and negative samples in the CMRM task in the model are changed to the structure of Chinese characters to further improve the performance of the model in error correction. Thus, the CKBERT (Ctc-CKBERT) model suitable for Chinese text error correction is obtained, and the model structure is as follows: Figure 9 shown.
[0140] After importing the similar word package into the model, the processed text data is input into the model. Used to express sentences The index of the masked token in To cover the token collection, is the set of uncovered tokens. The LMLM loss function is as follows:
[0141]
[0142] in, is the loss value of the masked language model, For the A masked token, is the loss value, is the number of tokens to be covered, For the The predicted probability of a masked token. Represents the set of parameters for the model.
[0143] Then, the information of the knowledge graph and the constructed positive and negative samples are used through the CMRM task to bring similar words closer together and alienate the wrong relationships in the negative samples. and negative samples , which is used to express the entity's perception of the context. As shown in the following formula:
[0144]
[0145] in is the hidden representation of the entity, is the nonlinear activation function RELU, is the weight matrix, is the self-attention pooling operator, Normalize the layers to improve the generalization ability of the model. represents the hidden representation of node i, Represents the hidden representation of node j. The role of the self-attention pooling operator is to extract all hidden representations Representative hidden representation features in .
[0146] CMRM uses InfoNCE as the loss function as shown below:
[0147]
[0148] in, Represents the loss value, and the positive sample triplet is , the negative sample triplet is , is a set of positive and negative samples, is the temperature coefficient, used to adjust the distribution smoothness, is the cosine similarity of the positive sample group, which is used to indicate the feature similarity between the entity and the positive sample. is the cosine similarity of the negative sample group, which is used to indicate the feature similarity between the entity and the negative sample.
[0149] The model jointly trains error detection and error correction in downstream tasks. First, the error probability calculation formula based on the detection network is used to calculate the error probability of each character as shown below:
[0150]
[0151] in is the sigmiod function, For the The probability that a character is an error character, For the The hidden state of the detection network for characters, To detect the parameters of the network, is the bias term. The error probability of each detected character is between [0, 1]. Based on the value, the erroneous character is selected for correction in the subsequent error correction. The error correction probability is shown as follows:
[0152]
[0153] Use CKBERT to normalize the output value, The component of is expressed as the probability that the character is corrected. For the The hidden state of the detection network for characters, is the bias term, is the weight matrix for linear transformation of input features.
[0154] It is also necessary to calculate the probability of each character being covered by the model and dynamically adjust the embedding based on the error probability as follows:
[0155]
[0156] in and are mask embedding and original embedding respectively, is the probability of a character being covered. Joint training can improve the efficiency of model detection and error correction, and has better error correction performance for Chinese NOTAM text.
[0157] After the LMLM and CMRM tasks, the model selects some similar words as mask candidates and selects the correct word as a replacement after considering the context information and similarity.
[0158] The model's multi-layer self-attention mechanism can capture contextual dependencies and enhance the model's ability to understand semantics. The sinusoidal positional encoding preserves the order of characters and reduces semantic confusion caused by replacement.
[0159] The dynamic masking strategy of the LMLM task can adjust the probability of masking according to the complexity of the context, avoiding the destruction of semantics due to excessive replacement density. Its combination with the CMRM task strengthens the semantic relevance in the error correction process. When dealing with similar synonym replacements, it can select the correct replacement based on the learning of NOTAM data.
[0160] The model's dependency syntax constructs a dependency tree to analyze subordinate relationships between phrases, such as subject-verb and verb-object relationships. It then associates other phrases with the core phrase through dependency relationships, ensuring sentence structure. Semantic dependency learns the deep semantic connections between different words, unaffected by syntactic structure, and directly annotates semantic relationships using dependency arcs. This combination of mechanisms ensures that the model preserves the semantic information level and structural integrity of sentences during training.
[0161] During the replacement process, some highly similar characters may prevent the model from accurately interpreting the replaced sentence, leading to ambiguity. This method uses named entity annotation in the text, allowing the model to learn and analyze the collocation logic of phrases. For example, the collocation for "ban (stop / zhi) landing" in NOTAM is "ban", but after learning, it correctly selects "stop". Furthermore, English characters are removed during data processing, eliminating some of the risk of ambiguity.
[0162] 3. Model Experiment and Analysis
[0163] During training, the processed Chinese NOTAM text data was randomly divided into a training set, a test set, and a validation set with a 6:2:2 ratio. The constructed positive and negative samples were input into a multi-hop contrastive learning task, and similar word bags were used as markers to mask the original text. The experiment was set to 20 epochs and a batch size of 64. Due to the large amount of data, the number of iterations was set to 5000.
[0164] The present invention uses the CEV-RGCN model and the Tianzege-CNN model to extract pronunciation and glyph features, respectively. The cosine similarity formula is then used to calculate the similarity between the pronunciations and glyphs of the characters. The closer the similarity value calculated by the cosine similarity formula is to 1, the more similar the pronunciations or glyphs of the characters are; conversely, the less similar they are.
[0165] Because the aviation field requires high accuracy for NOTAM text, we selected characters with both phonetic and glyph similarities between 0.9 and 1.0 to improve error correction performance. To verify the rationale for this range, we constructed three new similarity word lists from the similarity word list, selecting characters with phonetic and glyph similarities between 0.7-0.79, 0.8-0.89, and 0.9-1.0. These were used to replace the [MASK] token in the mask for training. The training results were compared with those of the unmodified CKBERT model, the MacBERT model, and the BERT model. The comparison results are shown in Table 12.
[0166] Table 12 Model comparison
[0167]
[0168] Table 12 shows that the Ctc-CKBERT model has the best performance parameters when the similarity of the masked replacement words is 0.9-1. Compared with other models, it can be concluded that the Ctc-CKBERT model outperforms other models in all performance parameters.
[0169] The F1 values of each model and the Ctc-CKBERT model in the iteration process when using a similar vocabulary with a similarity of 0.9-1 are as follows Figure 10 The F1 value is the harmonic mean of precision and accuracy, and is a comprehensive performance indicator used to evaluate models. As can be seen from the figure, the F1 value of the improved model of the present invention grows rapidly and fluctuates less during training.
[0170] The above results show that the improved Ctc-CKBERT model based on the knowledge graph and the similar vocabulary of Chinese NOTAM injected into CKBERT has better performance than other models in processing E-item text error correction of Chinese NOTAM.
[0171] In another exemplary embodiment, based on the same inventive concept as the method embodiment, a system for automatically correcting NOTAM text errors is provided, comprising:
[0172] A Chinese character database construction module is used to extract Chinese characters from item E of the NOTAM text and construct a Chinese character database based on common Chinese characters;
[0173] A knowledge graph construction module, which calculates the pronunciation similarity of characters based on the Chinese character database and constructs a pronunciation knowledge graph, and calculates the glyph similarity of characters based on the Chinese character database and constructs a glyph knowledge graph;
[0174] A similar word library construction module is used to integrate the pronunciation similarity and the shape similarity to obtain a similar word library;
[0175] An automatic error correction model construction and training module is used to construct and train an automatic error correction model based on the similar word library. The automatic error correction model uses similar words to replace masks in the masked language model of CKBERT, and constructs positive and negative samples to be added to the multi-hop knowledge comparison learning task;
[0176] The error correction module is used to correct the navigation notice text using the trained automatic error correction model.
[0177] In another exemplary embodiment, based on the same inventive concept as the method embodiment, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores a computer program. When executed by a processor, the computer program implements the automatic error correction method for NOTAM text provided in the embodiment of the present invention. Based on this understanding, the technical solution of this embodiment, or the portion that contributes to the prior art, or the portion of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the various embodiments of the method of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0178] In another exemplary embodiment, based on the same inventive concept as the method embodiment, a terminal is provided, including a memory and a processor, wherein the memory stores computer instructions that can be executed on the processor, and when the processor executes the computer instructions, the automatic error correction method for the navigation notice text provided in the embodiment of the present invention is executed.
[0179] The processor may be a single-core or multi-core central processing unit or a specific integrated circuit, or one or more integrated circuits configured to implement the present invention.
[0180] Embodiments of the subject matter and functional operations described in this specification may be implemented in: tangibly embodied computer software or firmware, computer hardware including the structures disclosed in this specification and their structural equivalents, or a combination of one or more thereof. Embodiments of the subject matter described in this specification may be implemented as one or more computer programs, i.e., one or more modules of computer program instructions encoded on a tangible, non-transitory program carrier for execution by a data processing apparatus or to control the operation of the data processing apparatus. Alternatively or in addition, the program instructions may be encoded on an artificially generated propagated signal, such as a machine-generated electrical, optical, or electromagnetic signal, that is generated to encode and transmit information to a suitable receiver apparatus for execution by the data processing apparatus.
[0181] The processes and logic flows described in this specification can be performed by one or more programmable computers executing one or more computer programs to perform the corresponding functions by operating on input data and generating output. The processes and logic flows can also be performed by, and apparatus can be implemented as, special purpose logic circuitry, such as an FPGA (field programmable gate array) or an ASIC (application-specific integrated circuit).
[0182] Processors suitable for executing computer programs include, for example, general-purpose and / or special-purpose microprocessors, or any other type of central processing unit. Typically, a central processing unit will receive instructions and data from a read-only memory and / or random access memory. The basic components of a computer include a central processing unit for implementing or executing instructions and one or more memory devices for storing instructions and data. Typically, a computer will also include one or more mass storage devices for storing data, such as magnetic disks, magneto-optical disks, or optical disks, or the computer will be operably coupled to such a mass storage device to receive data from it or to transmit data to it, or both. However, a computer does not necessarily have such a device. In addition, a computer can be embedded in another device, such as a mobile phone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a global positioning system (GPS) receiver, or a portable storage device such as a universal serial bus (USB) flash drive, to name a few.
[0183] It should be understood that each box in the flow chart or block diagram can represent a module, program segment or a part of code, and the part of the module, program segment or code comprises one or more executable instructions for realizing the logical function of the provision. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a sequence different from that marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs the function or action of the provision, or can be implemented with a combination of dedicated hardware and computer instructions.
[0184] The above specific implementation methods are detailed descriptions of the present invention. It cannot be considered that the specific implementation methods of the present invention are limited to these descriptions. For ordinary technicians in the technical field to which the present invention belongs, they can make several simple deductions and substitutions without departing from the concept of the present invention, which should be regarded as falling within the scope of protection of the present invention.
Claims
1. A method for automatically correcting errors in NOTAM text, characterized in that: The following steps are involved: S1. Extract Chinese characters from item E of the NOTAM text and build a Chinese character database based on common Chinese characters. S2. Calculating the pronunciation similarity of characters based on the Chinese character database and constructing a pronunciation knowledge graph; Calculating glyph similarity based on the Chinese character database and constructing a glyph knowledge graph; The calculating of the pronunciation similarity based on the Chinese character database and constructing the pronunciation knowledge graph includes: The Cypher language in the Neo4j tool is used to establish the nodes, attributes, and relationships of the phonetic features to form a phonetic knowledge graph; The continuous bag-of-words model is used as the character embedding model for training to obtain the word-phonetic vector; The word pronunciation vector is used as one input of the RGCN model, and the word pronunciation knowledge graph is used as another input of the RGCN model, and the RGCN model is jointly trained in an unsupervised manner; The calculating of glyph similarity based on the Chinese character database and constructing a glyph knowledge graph includes: Split the Chinese characters in the Chinese character database into radicals and classify them according to the radicals; and use the Glyce model to calculate similarity; S3, fusing the pronunciation similarity and the shape similarity to obtain a similar word library; S4. Build and train an automatic error correction model based on the similar word library. The automatic error correction model replaces the mask with similar words in the masked language model of CKBERT, and constructs positive and negative samples to be added to the multi-hop knowledge comparison learning task; S5. Use the trained automatic error correction model to correct the navigation notice text.
2. The automatic error correction method for NOTAM text according to claim 1, characterized in that: The extraction of Chinese characters in item E of the NOTAM text includes: De-noising the original data of NOTAM text; Save the Chinese telegraph code of item E and the corresponding Chinese characters as a mapping file, and use regular expressions to standardize the telegraph code into Chinese characters; Then extract the structured information of the text of item E to generate a new file.
3. The automatic error correction method for NOTAM text according to claim 1, characterized in that: The unsupervised training of the RGCN model includes: The node feature update formula of RGCN is ,in For nodes In the The characteristics of the layer, σ is the activation function (ReLU), For the relationship Next node The set of neighbor nodes of is the normalization coefficient, is the set of all relationship types in the graph, For nodes In the Layer characteristics, For nodes In the Layer characteristics, For the Corresponding relationship type in the layer The weight matrix, It is The self-loop weight matrix in the layer; The mean square error is used to calculate the loss function value. The formula is as follows: ,in is the number of nodes, is the similarity, is the generated similarity.
4. The automatic error correction method for NOTAM text according to claim 1, characterized in that: Before building the automatic error correction model, it also includes: Mark the text in item E as incorrect; The incorrectly annotated text is segmented and annotated with entities and relationships using the LTP tool; The text after entity and relationship annotation is converted into JSON file format and processed into triples according to the knowledge graph.
5. The automatic error correction method for NOTAM text according to claim 1, characterized in that: The loss function of the masked language model is ,in is the loss value of the masked language model, To cover the token collection, A collection of uncovered tokens. is the number of tokens to be covered, Expressing sentences The index of the masked token, For the A masked token, For the The predicted probability of the masked token, represents the parameter set of the model; The loss function used in the multi-hop knowledge contrastive learning task is ,in, represents the loss value of the multi-hop knowledge contrast learning task, represents a positive sample triple, represents a negative sample triplet, is the hidden representation of the entity, is the temperature coefficient, is the cosine similarity of the positive sample group, which is used to indicate the feature similarity between the entity and the positive sample. is the cosine similarity of the negative sample group, which is used to indicate the feature similarity between the entity and the negative sample.
6. An automatic error correction system for NOTAM text, characterized in that: include: A Chinese character database construction module is used to extract Chinese characters from item E of the NOTAM text and construct a Chinese character database based on common Chinese characters; A knowledge graph construction module, which calculates the pronunciation similarity of characters based on the Chinese character database and constructs a pronunciation knowledge graph, and calculates the glyph similarity of characters based on the Chinese character database and constructs a glyph knowledge graph; The calculating of the pronunciation similarity based on the Chinese character database and constructing the pronunciation knowledge graph includes: The Cypher language in the Neo4j tool is used to establish the nodes, attributes and relationships of the phonetic features to form a phonetic knowledge graph; The continuous bag-of-words model is used as the character embedding model for training to obtain the word-phonetic vector; The word pronunciation vector is used as one input of the RGCN model, and the word pronunciation knowledge graph is used as another input of the RGCN model, and the RGCN model is jointly trained in an unsupervised manner; The calculating of glyph similarity based on the Chinese character database and constructing a glyph knowledge graph includes: Split the Chinese characters in the Chinese character database into radicals and classify them according to the radicals; and use the Glyce model to calculate similarity; A similar word library construction module is used to integrate the pronunciation similarity and the shape similarity to obtain a similar word library; An automatic error correction model construction and training module is used to construct and train an automatic error correction model based on the similar word library. The automatic error correction model uses similar words to replace masks in the masked language model of CKBERT, and constructs positive and negative samples to be added to the multi-hop knowledge comparison learning task; The error correction module is used to correct the navigation notice text using the trained automatic error correction model.
7. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method for automatically correcting errors in a NOTAM text as described in any one of claims 1 to 5 is implemented.
8. A terminal comprising a memory and a processor, wherein the memory stores computer instructions that can be executed on the processor, characterized in that: When the processor runs the computer instructions, it executes the automatic error correction method for navigation notice text described in any one of claims 1-5.
Citation Information
Patent Citations
Navigation notification text processing method, computer program product and terminal
CN118332138A
Natural language grammatical content communication method during speech recognition involves generating supplementary text for grammar message using permutations of symbols
DE10015859A1