A text entailment inference method based on multi-source semantic coupled attention
Patent Information
- Application Number
- CN202410281964.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-12
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2044-03-12
AI Technical Summary
[0005]1)文本语义信息无法对信息融合进行有效的引导,难以发挥文本语义信息的作用;
[0031]1.本发明的模型对于文本语义信息和文本图信息的交互更加可控,并通过权重的稀疏特性增强模型的泛化能力。
Smart Images

Figure SMS_14 
Figure SMS_18 
Figure SMS_21
Abstract
Description
Technical Field
[0001] This invention relates to the field of textual entailment, and more particularly to a textual entailment inference method based on multi-source semantic coupling attention. Background Technology
[0002] Textual entailment, as a fundamental task in natural language inference, provides crucial support for the complex semantic understanding of inference tasks. Its complex semantic understanding forms the basis for multiple text understanding and inference tasks. In textual entailment tasks, text comprises premises and hypotheses; the primary goal is to determine whether the hypothesis can be inferred from the premises. With the increasing richness of social network content, textual entailment tasks have transcended traditional textual semantic reasoning scenarios. Their in-depth understanding of the inferential relationships between texts provides the foundation for diverse tasks such as question answering, text summarization, and machine translation. Therefore, accurately describing the complex interactive relationships between texts, and whether this can be adjusted and optimized through explicit training structures, has become a crucial research question in complex semantic fusion inference.
[0003] Currently, most existing research focuses on textual reasoning using semantic information from premises and assumptions, attempting to extract user intent information from text pairs through complex semantic representation algorithms. Algorithms based on text graph structural information representation mainly construct text graphs to represent and compute the structured information of the text. Multi-view text structural information is primarily computed using methods such as adaptive attention through neighborhood ensembles, multi-graph ensembles, and multi-aggregators. However, with the development of social media, diverse discussions and expressions have made the reasoning process increasingly complex, and relying solely on textual semantic information and text graph information representation is insufficient for representing and computing complex semantics. If the potential interaction relationships between text graphs and textual semantics cannot be accurately constructed, it will be difficult to benefit textual reasoning tasks through semantic fusion. Therefore, based on the concatenation of textual semantic and textual graph information, designing interaction methods to fuse textual semantic and textual graph information has become a feasible fusion framework. However, this remains insufficient for handling the fine-grained complex dependencies inherent in textual implication.
[0004] The following challenges remain in performing complex semantic fusion in text entailment tasks:
[0005] 1) Textual semantic information cannot effectively guide information fusion, making it difficult to leverage the role of textual semantic information;
[0006] 2) The relationship between textual semantic information and textual graph information is difficult to characterize, and the consistency of the fused information cannot be guaranteed;
[0007] 3) The model lacks an explicit mechanism for capturing interactive relationship information in the fusion information, and its ability to accurately capture complex dependency relationships is insufficient.
[0008] Therefore, those skilled in the art are dedicated to developing a text entailment inference method based on multi-source semantic coupling attention. Existing methods for fusing text semantic information and text graph information focus more on information splicing and learning complex interaction relationships through attention mechanisms, lacking characterization of potential associations. To address the current challenges, this invention proposes a latent fusion Transformer mechanism called LaFT (Latent Fusion Transformer). LaFT couples text semantic information and text graph information by using latent topic information as anchor information, and achieves controllable training and learning of association relationships by explicitly constructing an attention weight scaling matrix, thus realizing complex semantic fusion prediction for text entailment tasks. Summary of the Invention
[0009] In view of the above-mentioned deficiencies of the prior art, the technical problem to be solved by the present invention is to realize the coupling of text semantic information and text graph information, ensure the consistency of fused information, and realize complex semantic fusion prediction of text implication tasks.
[0010] To achieve the above objectives, this invention provides a textual implication inference method based on multi-source semantic coupling attention, including a latent fusion Transformer mechanism, deriving a block sparse attention scaling matrix from latent topic similarity, performing semantic calibration within the Transformer, and enabling a fine discussion of premises and assumptions.
[0011] Furthermore, the block sparse attention scaling matrix adjusts the weights of the attention mechanism, enhancing the control of the interaction weights between text semantic information and text graph information.
[0012] Furthermore, extracting potential topic information from the text guides the sampling of text graph neighbors, ensuring meaningful interaction between text and text graph information.
[0013] Furthermore, latent topic vectors are extracted from the semantic information of the text, and neighboring text graph information is filtered based on the similarity of the latent topic vectors of the text graph information.
[0014] Furthermore, this includes potential fusion mechanisms to integrate textual and text-graph information.
[0015] Furthermore, the premises and assumptions in the text semantic information and text graph information are processed separately, and fine-grained interaction relationships are adjusted in a way that involves explicit weight adjustment.
[0016] Furthermore, it includes the following steps:
[0017] Step 1: Parse the grammatical dependencies of text information and construct a grammar-based text graph;
[0018] Step 2: Use HAN as a heterogeneous graph representation algorithm to preserve multi-level structure and textual semantic information;
[0019] Step 3: Use the latent Dirichlet assignment to extract elements and decompose latent topics;
[0020] Step 4: Optimize the text graph representation by scaling the weight matrix within the attention mechanism to achieve controllable scaling of the attention weights;
[0021] Step 5: Construct a block sparse attention scaling matrix and scale the text graph weights based on latent topic weights to increase the information importance of text semantic information and achieve semantic fusion dominated by text semantic information.
[0022] Step 6: Characterize the positional information in the input, encode the positional information, and use a learnable relative position representation;
[0023] Step 7: Train the classifier using the cross-entropy loss function.
[0024] Furthermore, in step 1, the syntactic dependencies of the text information are parsed using the Stanford CoreNLP parser.
[0025] Furthermore, the text graph includes a premise subgraph and a hypothesis subgraph.
[0026] Furthermore, in step 4, a neighbor sampling method is used to optimize the text graph representation.
[0027] In a preferred embodiment of the present invention, the correlation between textual semantic information and textual graph information is difficult to characterize, making it impossible to guarantee the consistency of the fused information. This method proposes and implements a block-sparse attention scaling matrix derived from latent topic similarity, which facilitates semantic calibration within the Transformer, enabling a fine-grained discussion of premises and assumptions. The present invention uses the block-sparse attention scaling matrix to adjust the weights of the attention mechanism, enhancing the accurate control of the interaction weights between textual semantic information and textual graph information.
[0028] Textual semantic information cannot effectively guide information fusion, making it difficult to leverage the role of textual information. This method utilizes latent topic information extracted from text to guide the sampling of graph neighbors, ensuring meaningful interaction between textual and graph information. This invention extracts latent topic vectors from textual semantic information and filters neighboring textual graph information based on the similarity of latent topic vectors between textual and graph information.
[0029] The existing model lacks an explicit mechanism for capturing interactive information relationships, resulting in insufficient ability to accurately capture complex dependencies. This method proposes a latent fusion mechanism, providing a solution for the effective integration of textual implications and achieving efficient integration of textual and graph information. This invention separately processes the premises and assumptions within the textual semantic information and text-graph information, adjusting fine-grained interactive relationships through explicit weight adjustments.
[0030] Compared with the prior art, the present invention has the following obvious substantive features and significant advantages:
[0031] 1. The model of this invention provides more control over the interaction between text semantic information and text graph information, and enhances the generalization ability of the model through the sparsity of weights.
[0032] 2. This invention filters text graph neighbor information through text semantic information, effectively enhancing the consistency of the model when fusing text semantic information and text graph information.
[0033] 3. The fine-grained explicit interaction relationships of this invention improve the efficiency and interpretability of the model for parameter tuning. Detailed Implementation
[0034] This invention proposes a potential fusion Transformer mechanism called LaFT (Latent FusionTransformer). Existing methods for fusing text semantic information and text graph information focus more on information splicing and learning complex interaction relationships through attention mechanisms, but lack characterization of potential correlations.
[0035] The proposed method, LaFT, couples textual semantic information and textual graph information by using latent topic information as anchor information. It achieves controllable training and learning of association relationships by explicitly constructing an attention weight scaling matrix, thus enabling complex semantic fusion prediction for textual entailment tasks.
[0036] Step 1: Using the Stanford CoreNLP parser, various syntactic dependencies in text information can be parsed, thereby constructing a syntactic-based text graph structure. (Syntactic dependency-based text graph) It includes two subgraphs: the premise subgraph. and hypothetical subgraph Word frequency-based text graphs contain word co-occurrence information within a fixed window, incorporating latent topic information and reflecting underlying intent. Similarly, word frequency-based text graphs are... The nodes of a text graph are words, which together form a set of node representations. Adjacency Matrix Includes dependencies between words.
[0037] Step 2: For document d i Text Image Include and There are various types of edges in text graphs, including multiple syntactic dependencies and word co-occurrence relationships. Hierarchical Attention Network (HAN) is used as a heterogeneous graph representation algorithm to preserve multi-level structure and textual semantic information. For graphs... Text graph representation is as follows By swapping the positions of the representations to correspond with the input.
[0038] Step 3: Text Semantic Representation Generated from the last hidden state of BERT. Elements are extracted and latent topics are decomposed using Latent Dirichlet Allocation (LDA). Train an LDA model to distinguish the latent topic information of each document. For document d i =(p i ,h i ), thus obtaining the latent topic representation L i =p(p i ,h i |α,β), where p(p i ,h i |α,β) is the document d i The topic probability distribution, whose topic distribution parameter α and word distribution parameter β are Dirichlet distributions, can be obtained through adaptive training.
[0039] Step 4: The premise-hypothesis pair represents the latent topics, where k is the number of topics with the highest consistency selected from the candidate topics, and d is the representation dimension. Cosine similarity is used to measure premise-hypothesis similarity, calculated as follows:
[0040]
[0041] Where S(X)={s ij Let be the PH similarity matrix of the documents. A neighbor sampling method is used to optimize the graph representation to identify the most relevant neighbors of a given document. For the input... Neighbor sampling representation with L elements The representation is as follows:
[0042]
[0043] S(X)i =TOP L (s i0 ,…,s in )
[0044] Textual information is used to filter document neighbors. However, in the original Transformers, attention interactions are complex and cannot be directed to emphasize the desired relationship. Therefore, the PH (Premise-Hypothesis) similarity matrix S(X) generates latent topic weights W with Z-score normalization. LDA By scaling the matrix within the attention mechanism, the attention weights can be scaled in a controllable manner, making the allocation of attention more reasonable and interpretable. In specific experimental results, the attention scaling matrix based on latent topic weights effectively controls the allocation of model attention weights, improving the model's ability to capture information from different modules and characterize interactive information.
[0045] Step 5: For the input text-graph representation, it can be formatted as X input =(X s ,X g ).in Textual representations of premises and assumptions, while A graph representation of the aggregation. The model input X. input Attention can be formatted as follows:
[0046]
[0047] Where Q = X input W Q K = X input W K and V=X input W V W PH ={w ij} represents the block sparse weight matrix. w ij The calculation is as follows:
[0048]
[0049] W constructed PH The matrix exhibits block sparsity, preserving a degree of randomness in the interaction between textual semantic information and text-graph information to enhance the model's generalization ability in capturing interactive information. Overall W PH The matrix scales the text graph weights based on the weights of potential topics, aiming to increase the information importance of text semantic information and achieve semantic fusion dominated by text semantic information.
[0050] Step 6: To characterize the positional information in the input, this information needs to be encoded using a learnable relative positional representation. This takes into account the interaction between the key, query, and positional representation. This representation vector can be represented as a ij =w j-i ,in This represents the interaction information between elements i and j. Considering sufficient interaction based on the attention weights of positions i and j, it can be formatted as follows:
[0051] e ij = <W Q x i W K x j >=(W Q x i )·(W K x j )+(W Q x i +W K x j )·a ij
[0052] Based on attention weight e ij The attention kernel function is represented as follows:
[0053]
[0054] According to the definition of Transformers, the l-th layer latent fusion Transformer can be formatted using skip connections, as follows:
[0055] X' = X input +LaFA l (X input )
[0056] X l =FFN(X') := ReLU(X'W1)W2
[0057] Step 7: Output the last layer of the model Predicted labels Through formula The calculation is performed. The cross-entropy loss function is used to train the classifier, expressed as... Here, Y is the real label. It is a predicted label.
[0058] The preferred embodiments of the present invention have been described in detail above. It should be understood that those skilled in the art can make numerous modifications and variations based on the concept of the present invention without creative effort. Therefore, all technical solutions that can be obtained by those skilled in the art based on the concept of the present invention through logical analysis, reasoning, or limited experimentation on the basis of existing technology should be within the scope of protection defined by the claims.
Claims
1. A text entailment inference method based on multi-source semantic coupling attention, characterized in that, It includes a Transformer mechanism for potential fusion, derives a block sparse attention scaling matrix from potential topic similarity, performs semantic calibration within the Transformer, and enables a fine discussion of premises and assumptions. The block sparse attention scaling matrix adjusts the weights of the attention mechanism, enhancing the control over the interaction weights of text semantic information and text graph information. Extracting latent topic information from the text guides the sampling of text graph neighbors, ensuring meaningful interaction between text and text graph information. Potential topic vectors are extracted from text semantic information, and neighboring text graph information is filtered based on the similarity of the potential topic vectors in the text graph information. This includes potential fusion mechanisms to integrate text and text-graph information. The premises and assumptions in the text semantic information and text graph information are processed separately, and fine-grained interaction relationships are adjusted by explicit weight adjustment. Includes the following steps: Step 1: Parse the grammatical dependencies of text information and construct a grammar-based text graph; Step 2: Use HAN as a heterogeneous graph representation algorithm to preserve multi-level structure and textual semantic information; Step 3: Use the latent Dirichlet assignment to extract elements and decompose latent topics; Step 4: Optimize the text graph representation by scaling the weight matrix within the attention mechanism to achieve controllable scaling of the attention weights; Step 5: Construct a block sparse attention scaling matrix and scale the text graph weights based on latent topic weights to increase the information importance of text semantic information and achieve semantic fusion dominated by text semantic information. Step 6: Characterize the positional information in the input, encode the positional information, and use a learnable relative position representation; Step 7: Train the classifier using the cross-entropy loss function.
2. The text entailment inference method based on multi-source semantic coupling attention as described in claim 1, characterized in that, Step 1 involves using the Stanford CoreNLP parser to parse the syntactic dependencies of the text information.
3. The text entailment inference method based on multi-source semantic coupling attention as described in claim 1, characterized in that, The text graph includes a premise subgraph and a hypothesis subgraph.
4. The text entailment inference method based on multi-source semantic coupling attention as described in claim 1, characterized in that, Step 4 involves using neighbor sampling to optimize the text graph representation.
Citation Information
Patent Citations
Text implication recognition method and device
CN110390397A
Part-of-speech and self-attention mechanism fused sentiment tendency classification method and system
CN110569508A