Semantic error elimination method and system based on knowledge meta-model
By introducing a knowledge element model-based method and graph attention network in the semantic error elimination task, the problems of insufficient accuracy and poor flexibility of error elimination in the prior art are solved, and higher error elimination accuracy and effect are achieved.
Patent Information
- Application Number
- CN202411791909.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-06
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2044-12-06
AI Technical Summary
The prior art has problems of insufficient accuracy and poor flexibility in the semantic error elimination task, especially when dealing with target words in phrase form, the noise information has a great impact, resulting in poor error elimination effect.
Using a semantic error elimination method based on knowledge element model, a hybrid encoder algorithm is constructed by introducing structured knowledge and graph attention networks to reduce the impact of noise information and improve the error elimination accuracy of target words in phrase form.
The error elimination accuracy of the target words in the phrase form is improved, the overall error elimination effect of the method is enhanced, and the flexibility and noise information influence problems in the prior art are overcome.
Smart Images

Figure CN119938919A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of semantic error elimination, and in particular to a semantic error elimination method and system based on a knowledge meta-model. Background Art
[0002] Natural language is ambiguous in most cases, but when a word appears in a specific context, the meaning of the target word is limited by the context, so the meaning of the target word can be interpreted according to the context in which it appears; the goal of error elimination is to find the exact meaning of ambiguous words in a specific context. Once the scope of the context is determined, the choice of error elimination method becomes the key to semantic error elimination.
[0003] There are two major categories of methods that work well for semantic error elimination: supervised methods and knowledge-based methods. Supervised methods treat semantic error elimination as a classification problem, use a large amount of annotated semantic data as annotation data for semantic error elimination, classify specific words into the meaning of dictionary definitions, and use methods such as machine learning and neural networks to achieve semantic error elimination. Knowledge-based methods use existing vocabulary resources, such as knowledge bases, semantic networks, etc. as knowledge bases for error elimination, and then determine the meaning of words based on context. Context representations derived from neural language models have been shown to effectively represent the differences between different meanings of the same word. However, these context representations are not easily connected to semantic networks, so it is easy to ignore the information of the knowledge base itself. Compared with knowledge-based methods, supervised methods are currently able to produce better error elimination accuracy results, but they are not as flexible as knowledge-based methods in the overall semantic error elimination task. Supervised methods use short text datasets, which makes most corpora less information and difficult to distinguish semantics in different scenarios, which limits the application and development speed of supervised methods. At the same time, in the task of semantic error elimination, the connection between words and phrases in the input text may affect the error elimination effect of the model. When the error elimination target word is a word in a phrase, the connection between the phrase and the word may lead to errors in word meaning judgment, thus affecting the error elimination effect. Summary of the invention
[0004] The purpose of the present invention is to overcome the deficiencies in the prior art and provide a semantic error elimination method and system based on a knowledge meta-model, which introduces structured knowledge from knowledge meta-models to supplement broader semantic information, uses the hierarchical structure of contextual input text to describe the meaning of words and phrases, and constructs a hybrid encoder algorithm based on context and interpretation, introduces a graph attention network to reduce noise information in the contextual input text, thereby improving the error elimination accuracy of phrase-form target words, and ultimately improving the error elimination effect of the method.
[0005] The present invention first discloses a semantic error elimination method based on a knowledge meta-model, and the specific steps are as follows:
[0006] Step 1: Obtain the target word information to be eliminated, build a knowledge element embedding matrix based on the relationship between words and semantics in WordNet, and embed the knowledge element into the annotation vector representation of the target word through the adjacency matrix to obtain the structured information and annotation data of the knowledge element;
[0007] Step 2: construct a hierarchical structure of phrases and their related words based on the contextual input text;
[0008] Step 3: construct a hybrid encoder, which consists of two parts, including a context encoder and a paraphrase encoder, wherein the context input text with a hierarchical structure constructed in step 2 is fed into the context encoder; the structured information and annotation data of the knowledge element extracted in step 1 are input into the paraphrase encoder; the text feature vector and the structural feature vector are extracted from the output results of the context encoder and the paraphrase encoder respectively;
[0009] Step 4: The text feature vector extracted from the context encoder is weighted and aggregated by the graph attention network to obtain a weighted aggregate feature vector, thereby eliminating the influence of noisy neighboring nodes;
[0010] Step 5: Obtain the final error elimination score by performing a dot product operation on the weighted aggregate feature vector output in step 4 and the structural feature vector output from the interpretation encoder.
[0011] Furthermore, in order to better utilize the structured information and annotation data of the knowledge element, the method further includes:
[0012] The constructed knowledge element embedding matrix A is dot-producted with the annotation vector representation matrix of the target word to be error-eliminated, its weight matrix, and the bias vector to obtain the final knowledge element vector matrix Q;
[0013] Q=(HO+b q )A T +(HO+b q )
[0014] Where Q represents the obtained knowledge element vector matrix; H represents the annotation vector representation matrix formed by all annotation data of the target word to be error-eliminated; O represents the weight matrix of the target word to be error-eliminated in the knowledge element; b q Represents the bias vector.
[0015] Further, step 2 is specifically as follows:
[0016] The sentence c in the context is divided into any two parts; in each iteration, one of the parts is selected and divided into smaller partitions until the minimum segmentation unit of the sentence is reached; after each iteration, the correlation between the new partitions is evaluated and the correlation set is updated; finally, a correlation set and a hierarchical set are obtained; wherein the correlation set contains all text partitions generated in each division and the number of correlation links between words and phrases between the partitions; the hierarchical set contains all partitions of the sentence c in the context in each division.
[0017] Further, step 3 is specifically as follows:
[0018] The context input text is processed through a hierarchy and then input into the context encoder. When the target word for error elimination is a single word, start and end symbols are added to the input text and then input into the pre-trained language model encoder. When the target word for error elimination is a phrase, the average of the vectors of each word in the phrase is used to represent the phrase.
[0019]
[0020] Among them, j and q represent the first word and the last word in the partition span of this phrase respectively; T c represents the context encoder; [l] means putting each word into the context encoder in turn; r ws represents the text feature vector obtained from the input text after being processed by the context encoder;
[0021] In the paraphrase encoder, the annotation data g of the target word s to be error-free s is used as input; the paraphrase encoder generates the paraphrase vector g of the target word s to be error-corrected (s) , as follows:
[0022] g (s) =T g (g s )[0]
[0023] Among them, T g represents the paraphrase encoder;
[0024] For each synonym of the target word, the vector score is calculated based on the structured information of the knowledge element, and the interpretation vector score q s It is obtained by adding the vector score of the target word to be error-eliminated to the weight of the edge between the target word and all synonyms;
[0025] q s =z s +∑ s'∈V|<s',s>∈E w(<s',s> )·Z s'
[0026] in,<s',s> represents the edge between each target word s and synonym s'; q s represents the weighted combination of the vector scores of all target words; w(<s',s> ) represents the weight of the edge between the target word s to be error-eliminated and all its synonyms s'; V represents the set of all vertices in the knowledge element, and E represents all edges; Z s' Represents the vector matrix representation of synonym s'; z s The vector matrix representation of the target word s to be error-eliminated is:
[0027] z s =h T g (s) +b T g (s)
[0028] Among them, h T and b T denote the transpose of the annotation vector representation matrix and the bias vector, respectively;
[0029] The interpretation vector score q s It is concatenated with the knowledge element vector matrix Q to obtain the structural feature vector output by the interpretation encoder.
[0030] Further, step 4 is specifically as follows:
[0031] The graph attention network is divided into two steps for weighting and aggregation: the first step is to calculate the attention coefficient. For the target word to be error-eliminated, the similarity coefficient between it and the neighboring nodes is calculated one by one; then the attention coefficient is calculated through normalization operation; the second step is to obtain a new feature vector by weighted summing of the attention coefficients, and perform feature enhancement transformation on it to obtain the weighted aggregated feature vector finally output by the encoder.
[0032] Further, the specific process of calculating the attention coefficient is:
[0033] The specific process of calculating the attention coefficient is:
[0034] For the target word s to be error-eliminated, the similarity coefficient between it and its neighbor node j is calculated one by one. The similarity coefficient is calculated as:
[0035]
[0036] Among them, r ws represents the text feature vector of the target word s obtained from the context encoder output; r wj Represents the text feature vector of neighbor node j; Wr represents the set of all neighboring nodes of the target word to be error-eliminated;ws Indicates r ws The value obtained by linear transformation; h represents the value obtained by linear transformation; Wr ws ‖Wr wj ] indicates that in r ws and r wj The concatenation operation is performed after feature transformation; LeakyReLU represents the activation function.
[0037] Furthermore, a new feature vector is obtained by weighted summing of the attention coefficients:
[0038] According to the similarity coefficient α sj , calculate the attention coefficient β through normalization operation sj , the specific calculation process is as follows:
[0039]
[0040] By weighted summation of attention coefficient β sj To obtain a new feature vector, specifically:
[0041]
[0042] in, represents the new feature vector output by the graph attention network for each word s after fusing the neighborhood information; σ(·) represents the activation function.
[0043] Furthermore, the feature enhancement transformation is specifically:
[0044] New feature representation obtained by weighted summation After feature enhancement transformation, we obtain It is the weighted aggregate feature vector finally output by the encoder. The specific process is:
[0045]
[0046] Among them, K represents the number of heads in the multi-head attention mechanism; W k is the weight matrix of the kth attention head; Represents the attention coefficient of the target word s and the neighbor node j in the kth attention head.
[0047] On the other hand, the present invention also provides a semantic error elimination system based on a knowledge meta-model, the system comprising:
[0048] at least one processor; and a memory in communication with the at least one processor; wherein,
[0049] The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute any one of the semantic error elimination methods for a knowledge meta-model described above.
[0050] On the other hand, the present invention also provides a computer storage medium, which stores a semantic error elimination method based on a knowledge meta-model, and when the method is executed by at least one processor, it implements any of the above-mentioned semantic error elimination methods based on a knowledge meta-model.
[0051] The present invention relates to a semantic error elimination method and system based on a knowledge meta-model, and proposes a hybrid encoder semantic error elimination method that combines a knowledge meta-model and a text hierarchy. The method introduces structured knowledge from knowledge meta-models to supplement broader semantic information, uses the hierarchy of context input text to describe the meaning of words and phrases, and constructs a hybrid encoder algorithm based on context and interpretation. The graph attention network is introduced to reduce noise information in the context input text, thereby improving the error elimination accuracy of phrase-form target words, and ultimately improving the error elimination effect of the method. The present invention also processes words and phrases in manually annotated context input texts, making the data set more suitable for the application of the semantic error elimination model in this article. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] Figure 1 is a schematic diagram of a hybrid encoder;
[0053] Figure 2 A schematic diagram of the structure of a semantic error elimination system based on a knowledge meta-model is provided in one embodiment of the present invention. DETAILED DESCRIPTION
[0054] The detailed description set forth below in conjunction with the accompanying drawings is intended to be a description of various exemplary embodiments of the present invention, and is not intended to represent the only embodiment that can practice the present invention. For the purpose of providing a thorough understanding of the present invention, the detailed description includes specific details. However, it is apparent to those skilled in the art that the present invention can be practiced without these specific details. In some instances, well-known structures and components are shown in block diagram form to avoid blurring the concept of the present invention.
[0055] In the description of the present invention, it should be noted that the embodiments described in the present invention are only part of the embodiments of the present invention, rather than all of the embodiments; based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in the field without making creative work are within the scope of protection of the present invention.
[0056] The terms "first", "second", etc. in the specification and claims of this document and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of this document described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, device, product or equipment that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or equipment.
[0057] The detailed description set forth below in conjunction with the accompanying drawings is intended to be a description of various exemplary embodiments of the present invention, and is not intended to represent the only embodiment that can practice the present invention. For the purpose of providing a thorough understanding of the present invention, the detailed description includes specific details. However, it is apparent to those skilled in the art that the present invention can be practiced without these specific details. In some instances, well-known structures and components are shown in block diagram form to avoid blurring the concept of the present invention.
[0058] Supervised semantic error elimination methods are very time-consuming and labor-intensive when acquiring large amounts of annotated data, so the present invention proposes to embed the structured information in knowledge elements into the annotated data and use existing knowledge metadata to alleviate the difficulty of acquiring large amounts of data. The supervised semantic error elimination method regards semantic error elimination as a label classification problem, similar to the vocabulary annotation task and named entity recognition task in natural language processing. To achieve semantic error elimination, it is necessary to generate a vector representation of the target word in a given context and use it to generate a probability distribution of all its possible semantics. In the present invention, knowledge elements are embedded into the annotated data in the form of an adjacency matrix.
[0059] In the field of natural language processing, semantic errors often lead to difficulties in understanding and information extraction. To address this problem, supervised semantic error removal methods provide an efficient solution by treating semantic error removal as a labeled classification problem. However, traditional methods are time-consuming and labor-intensive when acquiring large amounts of annotated data, which limits the widespread application of this method.
[0060] To address this challenge, this paper proposes a novel method to alleviate the difficulty of obtaining large amounts of data by embedding the structured information in knowledge metadata into annotation data. This process actually utilizes existing knowledge metadata and converts its structured information into a format that can be used for annotation.
[0061] 1. Semantic error elimination is regarded as a classification problem:
[0062] The method of the present invention transforms semantic error elimination into a tag classification task, which is similar to the common vocabulary annotation task and named entity recognition task in natural language processing. Such a setting can be processed with the help of existing machine learning algorithms.
[0063] 2. Generate vector representation:
[0064] The key to eliminating semantic errors is to generate a vector representation of the target word in a given context, which can effectively capture the semantic information of the target word. By building a semantic vector model, a probability distribution of all possible semantics of the target word can be generated.
[0065] 3. Embedding of knowledge elements:
[0066] In the present invention, knowledge elements are embedded into the annotation data in the form of an adjacency matrix. The use of the adjacency matrix enables structured information to be intuitively and efficiently integrated into data processing, provides rich contextual information, and thus improves the accuracy of semantic error elimination.
[0067] Example 1
[0068] In the present invention, the knowledge element embedding matrix A is constructed based on the relationship between words and semantics in WordNet. Since there is a subordinate relationship between words, the superior-subordinate relationship between words is considered when constructing the knowledge element embedding matrix A. The knowledge element embedding matrix is constructed by experimenting with only the superordinate level, only the subordinate level, and both the hypernym and the hyponym.
[0069] Step 1: Structured information embedding of knowledge elements
[0070] In order to embed knowledge elements, the present invention performs a dot product operation on the constructed knowledge element embedding matrix A, the annotation vector representation matrix of the target word to be error eliminated, its weight matrix and the bias vector to obtain the final knowledge element vector matrix Q, which is specifically:
[0071] Q=(HO+b q )A T +(HO+b q )
[0072] Where Q represents the obtained knowledge element vector matrix; since each word to be error-eliminated has multiple annotations, each annotation vector has a vector representation, H represents the annotation vector representation matrix formed by all annotation data of the target word to be error-eliminated, and O represents the weight matrix of the target word to be error-eliminated in the knowledge element; b q Represents the bias vector.
[0073] Step 2: Text Hierarchy
[0074] When selecting a target word from an input text, the target word may be a single word or a phrase, but a phrase usually contains multiple words, which often affects the judgment of the meaning of the words in the phrase. Therefore, the present invention recursively detects the degree of interaction between words and phrases, and then divides a larger text partition into smaller text partitions according to the degree of influence to build a hierarchical structure.
[0075] For the word meaning error elimination task, the input is a sentence c containing a set of words s to be error eliminated. Define c as: c = c0, c1, ..., c i ,…,c n Among them, c i is the target word to be corrected at position i in the sentence. In addition, partition is defined as a sentence partition of length n, The words that represent the endpoints of each sentence partition. The basic process of partitioning is to divide a given larger partition into Divided into two smaller partitions, and Where j represents the split point, i <j<j+1。
[0076] Starting from the entire sentence, the sentence c in the context is first divided into two parts. In the next iteration, it selects one of the two parts and further divides it into smaller partitions until the minimum segmentation unit of the sentence is reached. The last step of each iteration is to evaluate the correlation between the new partitions and update the correlation set R; the final output of the algorithm includes the correlation set R and the hierarchy set H. The correlation set R contains all the text partitions generated in each partition and the number of correlation links between words and phrases between partitions, while the hierarchy set H contains all the partitions of the sentence c in the context in each partition.
[0077] Step 3: Hybrid Encoder
[0078] Depend on Figure 1 It can be seen that the hybrid encoder structure encodes the context of the target word and its interpretation respectively, and each encoder is initialized using a pre-trained language model. Therefore, the input of the encoder requires the start and end symbols unique to the language model. The part of the hybrid encoder structure that processes the context is called the context encoder, and the part that processes the semantic interpretation is called the interpretation encoder.
[0079] The context input text is processed through a hierarchy and then input into the context encoder. When the target word for error elimination is a single word, start and end symbols are added to the input text and then input into the pre-trained language model encoder. When the target word for error elimination is a phrase, the average vector value of each word in the phrase is used to represent the phrase, specifically:
[0080]
[0081] Among them, j and q represent the first word and the last word in the partition span of this phrase respectively; T c represents the context encoder; [l] means putting each word into the context encoder in turn; r ws represents the text feature vector obtained from the input text after being processed by the context encoder;
[0082] In the paraphrase encoder, the annotation data g of the target word s to be error-free s is used as input; the paraphrase encoder generates the paraphrase vector g of the target word s to be error-corrected (s) , as follows:
[0083] g (s) =T g (g s )[0]
[0084] Among them, T g represents the paraphrase encoder; the first vector output by the paraphrase encoder (corresponding to the start token of the input) is considered as the global vector;
[0085] To achieve semantic error elimination, it is necessary to generate a vector representation for the target word to be eliminated in a given context and use it to generate the probability distribution of all possible semantics of the word. In order to be able to use the information in the knowledge element, it is denoted as G =<V,E,w> ; where V represents the set of all vertices in the knowledge element, E represents all edges, and w represents the weight of the edge between two vertices.
[0086] For each synonym of the target word, the vector score is calculated based on the structured information of the knowledge element, and the interpretation vector score q s It is obtained by adding the vector matrix of the target word to be error-eliminated to the weights of the edges between the target word and all synonyms;
[0087] q s =z s +Σ s'∈V|<s',s>∈E w(<s',s> )·Z s'
[0088] in,<s',s> represents the edge between each target word s and synonym s'; q s represents the weighted combination of the vector scores of all target words; w(<s',s> ) represents the weight of the edge between the target word s to be error-eliminated and all its synonyms s'; Z s' The vector matrix representing the synonym s'; z s The vector matrix representing the target word s to be error-eliminated is:
[0089] z s =h T g (s) +b T g (s)
[0090] Among them, h T and b T denote the transpose of the annotation vector representation matrix and the bias vector, respectively;
[0091] The interpretation vector score q s and the knowledge element vector matrix Q to obtain the structural feature vector r output by the interpretation encoder gs .
[0092] Step 4: Weighted aggregation of neighbor nodes
[0093] Since there are many neighboring words in the context of the target word to be error-eliminated, the noise information in these neighboring words will affect the accuracy of the model's error elimination of the word. This makes the output of the context encoder model not robust enough, so this paper uses a graph attention network to weighted aggregate neighboring nodes to eliminate the impact of noisy neighbors on accuracy; the graph attention network used in this invention is divided into two main steps for weighting and aggregation. The first step is to calculate the attention coefficient. For the target word s to be error-eliminated, the similarity coefficient between it and the neighbor node j is calculated one by one. The similarity coefficient calculation is specifically as follows:
[0094]
[0095] Among them, r ws represents the text feature vector of the target word s obtained from the context encoder output; r wj Represents the text feature vector of neighbor node j; N s Wr represents the set of all neighboring nodes of the target word to be error-eliminated; ws Indicates r ws The value obtained by linear transformation; h represents the value obtained by linear transformation; Wr ws ‖Wr wj ] indicates that in r ws and r wj Perform concatenation operation after feature transformation; LeakyReLU represents activation function;
[0096] Once the similarity coefficient α is obtained sj , and then the attention coefficient β is calculated by normalization operation sj , the specific calculation process is as follows:
[0097]
[0098] In the first step, we get the attention coefficient β sj After that, in the second step, the weighted sum of attention coefficients β sj To obtain a new eigenvector; the calculation process is:
[0099]
[0100] in, represents the new feature vector output by the graph attention network for each word s after fusing the neighborhood information; σ(·) represents the activation function;
[0101] New eigenvector obtained by weighted summation After feature enhancement transformation, the result is more robust; the final is the weighted aggregate feature vector finally output by the encoder; the calculation process is
[0102]
[0103] Among them, K represents the number of heads in the multi-head attention mechanism; W k is the weight matrix of the kth attention head; Represents the attention coefficient of the target word s and the neighbor node j in the kth attention head.
[0104] Step 5: Error elimination accuracy calculation
[0105] The feature vector finally output by the context encoder is expressed as The final output feature vector of the paraphrase encoder is represented as r gs ; The final error elimination accuracy is achieved by and r gs The calculation process is as follows:
[0106]
[0107] is the semantic score of the target word in a given context; When it is greater than a preset threshold, the correct meaning of the target word is determined, thereby achieving error elimination of the target word.
[0108] The model is trained using the cross entropy loss of the candidate semantic scores of the error-eliminated word s; the loss function used is shown in the following formula:
[0109]
[0110] Among them, (s,g i ) indicates a word and its meaning.
[0111] Experiments have shown that the error elimination accuracy of the model proposed in this invention is better than the latest model in most cases. The main reason why this model is better than other models is that it not only uses a pre-trained language model, but also uses a dual encoder to encode context and interpretation information respectively, and for the context input text information, it has a clear hierarchical representation between words and phrases, which enhances the model's ability to eliminate phrase errors. In addition, the model integrates the structured information in the knowledge element and introduces a graph attention mechanism for the noise data in the input text, making the error elimination accuracy of the model better than other models.
[0112] In order to understand the impact of each module in the model on the accuracy of error elimination, the degree of influence of each module on the model is analyzed by analyzing each part of the error elimination model separately. It can be seen from the experimental simulation that the context and annotation encoders have the greatest impact on the error elimination accuracy of the model and are the main components of the model. On the full dataset, the hierarchical module has a significant reduction in the error elimination accuracy of the model, indicating that the impact of the hierarchical module on the error elimination accuracy of the model is affected by the number of phrases in the dataset, which also has a greater impact on the error elimination accuracy of the model. The knowledge meta-structured information embedding module and the graph attention mechanism module have a certain impact on all datasets.
[0113] Example 2
[0114] like Figure 2 As shown, this embodiment provides a semantic error elimination system based on a knowledge meta-model, and the system may include: a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other through the communication bus.
[0115] Memory, used to store computer programs;
[0116] The processor is used to implement the semantic error elimination method of a knowledge meta-model described in the above embodiment 1 when executing the computer program stored in the memory.
[0117] As an embodiment of the present invention, the communication bus mentioned in the above terminal can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. The communication bus can be divided into an address bus, a data bus, a control bus, etc.
[0118] As an embodiment of the present invention, the memory may include a random access memory (RAM) or a non-volatile memory, such as at least one disk memory. Optionally, the memory may also be at least one storage device located away from the aforementioned processor.
[0119] As an embodiment of the present invention, the above-mentioned processor can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components.
[0120] At the same time, the present invention also provides a computer storage medium, characterized in that: the computer storage medium stores a semantic error elimination method based on a knowledge meta-model, and when the method is executed by at least one processor, it implements any of the above-mentioned semantic error elimination methods based on a knowledge meta-model.
[0121] In summary, the present invention introduces structured knowledge from knowledge elements to supplement broader semantic information, uses the hierarchical structure of context input text to describe the meaning of words and phrases, constructs a hybrid encoder algorithm based on context and interpretation, and introduces a graph attention network to reduce noise information in the context input text, thereby improving the error elimination accuracy of phrase form target words and ultimately improving the error elimination effect of the method.
[0122] In the several embodiments provided herein, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, or it can be an electrical, mechanical or other form of connection.
[0123] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the embodiments of this article.
[0124] In addition, each functional unit in each embodiment of this invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above integrated unit may be implemented in the form of hardware or in the form of software functional unit.
[0125] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this article is essentially or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of this article. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk and other media that can store program codes.
[0126] The above descriptions are merely embodiments of the present invention and are not intended to limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made using the contents of the present invention specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present invention.
Claims
1. A semantic error elimination method based on a knowledge meta-model, the specific steps of which are as follows: Step 1: Obtain the target word information to be eliminated, build a knowledge element embedding matrix based on the relationship between words and semantics in WordNet, and embed the knowledge element into the annotation vector representation of the target word through the adjacency matrix to obtain the structured information and annotation data of the knowledge element; Step 2: construct a hierarchical structure of phrases and their related words based on the contextual input text; Step 3: construct a hybrid encoder, which consists of two parts, including a context encoder and a paraphrase encoder, wherein the context input text with a hierarchical structure constructed in step 2 is fed into the context encoder; the structured information and annotation data of the knowledge element extracted in step 1 are input into the paraphrase encoder; Extracting text feature vectors and structural feature vectors from the output results of the context encoder and the paraphrase encoder respectively; Step 4: The text feature vector extracted from the context encoder is weighted and aggregated by the graph attention network to obtain a weighted aggregate feature vector, thereby eliminating the influence of noisy neighboring nodes; Step 5: Obtain the final error elimination score by performing a dot product operation on the weighted aggregate feature vector output in step 4 and the structural feature vector output from the interpretation encoder.
2. According to the method for eliminating semantic errors based on the knowledge element model of claim 1, in order to better utilize the structured information and annotation data of the knowledge element, the method further comprises: The constructed knowledge element embedding matrix A is dot-producted with the annotation vector representation matrix of the target word to be error-eliminated, its weight matrix, and the bias vector to obtain the final knowledge element vector matrix Q; Q=(HO+b q )A T +(HO+b q ) Where Q represents the obtained knowledge element vector matrix; H represents the annotation vector representation matrix formed by all annotation data of the target word to be error-eliminated; O represents the weight matrix of the target word to be error-eliminated in the knowledge element; b q Represents the bias vector.
3. The semantic error elimination method based on knowledge meta-model according to claim 1, wherein: Step 2 is as follows: Divide the sentence c in the context into any two parts; in each iteration, select one of the parts and divide it into smaller partitions until the minimum segmentation unit of the sentence is reached; After each iteration, the correlation between new partitions is evaluated and the correlation set is updated; finally, a correlation set and a hierarchical set are obtained; the correlation set contains all text partitions generated in each partition and the number of correlation links between words and phrases in each partition; the hierarchical set contains all partitions of sentence c in the context of each partition.
4. According to the method for eliminating semantic errors based on knowledge meta-model as claimed in claim 2, wherein step 3 specifically comprises: The context input text is processed through a hierarchy and then input into the context encoder. When the target word for error elimination is a single word, start and end symbols are added to the input text and then input into the pre-trained language model encoder. When the target word for error elimination is a phrase, the average of the vectors of each word in the phrase is used to represent the phrase. in, j and q represent the first and last words in the span of this phrase respectively; T c represents the context encoder; [l] means putting each word into the context encoder in turn; r ws represents the text feature vector obtained from the input text after being processed by the context encoder; In the paraphrase encoder, the annotation data g of the target word s to be error-free s is used as input; the paraphrase encoder generates the paraphrase vector g of the target word s to be error-corrected (s) , as follows: g (s) =T g (g s )[0] Among them, T g represents the paraphrase encoder; For each synonym of the target word, the vector score is calculated based on the structured information of the knowledge element, and the interpretation vector score q s It is obtained by adding the vector matrix of the target word to be error-eliminated to the weights of the edges between the target word and all synonyms; q s =from s +∑ s'∈V|<s',s>∈E In(<s',s> )·WITH s' in,<s',s> represents the edge between each target word s and synonym s'; q s represents the weighted combination of the vector scores of all target words; w(<s',s> ) represents the weight of the edge between the target word s to be error-eliminated and all its synonyms s'; V represents the set of all vertices in the knowledge element, and E represents all edges; Z s' The vector matrix representing the synonym s'; z s The vector matrix representing the target word s to be error-eliminated is: z s =h T g (s) +b T g (s) Among them, h T and b T denote the transpose of the annotation vector representation matrix and the bias vector, respectively; The interpretation vector fraction q s and the knowledge element vector matrix Q to obtain the structural feature vector r output by the interpretation encoder gs .
5. According to the method for eliminating semantic errors based on knowledge meta-model as claimed in claim 1, step 4 specifically comprises: The graph attention network is divided into two steps for weighting and aggregation: the first step is to calculate the attention coefficient. For the target word to be error-eliminated, the similarity coefficient between it and the neighboring nodes is calculated one by one; then the attention coefficient is calculated through normalization operation; the second step is to obtain a new feature vector by weighted summing of the attention coefficients, and perform feature enhancement transformation on it to obtain the weighted aggregated feature vector finally output by the encoder.
6. According to the semantic error elimination method based on the knowledge meta-model of claim 5, the specific process of calculating the attention coefficient is: For the target word s to be error-eliminated, the similarity coefficient between it and its neighbor node j is calculated one by one. The similarity coefficient is calculated as: in, r ws represents the text feature vector of the target word s obtained from the context encoder output; r wj Represents the text feature vector of the neighbor node j obtained; Wr represents the set of all neighboring nodes of the target word to be error-eliminated; ws Indicates r ws The value obtained by linear transformation; h represents the value obtained by linear transformation; [Wr ws ‖Wr wj ] indicates that in r ws and r wj The concatenation operation is performed after feature transformation; LeakyReLU represents the activation function.
7. According to the semantic error elimination method based on the knowledge meta-model of claim 6, obtaining a new feature vector by weighted summing of attention coefficients is specifically: According to the similarity coefficient α sj , calculate the attention coefficient β through normalization operation sj , the specific calculation process is as follows: By weighted summation of attention coefficient β sj To obtain a new feature vector, specifically: in, represents the new feature vector output by the graph attention network for each word s after fusing the neighborhood information; σ(·) represents the activation function.
8. According to the semantic error elimination method based on knowledge meta-model according to claim 7, the feature enhancement transformation is specifically: New eigenvector obtained by weighted summation After feature enhancement transformation, we get It is the weighted aggregate feature vector finally output by the encoder. The specific process is: in, K represents the number of heads in the multi-head attention mechanism; W k is the weight matrix of the kth attention head; Represents the attention coefficient of the target word s and the neighbor node j in the kth attention head.
9. A semantic error elimination system based on knowledge meta-model, characterized in that: The system comprises: at least one processor; and a memory in communication with the at least one processor; wherein, The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute a semantic error elimination method based on a knowledge meta-model as described in any one of claims 1-8.
10. A computer storage medium, characterized in that: The computer storage medium stores a semantic error elimination method based on a knowledge meta-model, and when the method is executed by at least one processor, the semantic error elimination method based on a knowledge meta-model described in any one of claims 1-8 is implemented.
Citation Information
Patent Citations
Communication field process event semantic disambiguation method based on multi-layer attention
CN114692643A
Knowledge-enhanced word sense disambiguation method and device based on local self-attention
CN115526184A
Semantic error recognition method and device, terminal and medium
CN116956936A
Multi-language visual word sense disambiguation method
CN117610575A
An auto-disambiguation BOT engine for dynamic corpus selection per query
IN201821019827A