Text matching method and device based on fine-grained information consistent reasoning network
By adopting a fine-grained information consistent inference network in text matching, using character-level interaction and cross-attention learning, and automatically selecting important features for consistent inference, the problem of insufficient accuracy in predicting semantic relationships in the existing technology is solved, and the accuracy of text matching is significantly improved.
Patent Information
- Application Number
- CN202510540049.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-27
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-04-27
AI Technical Summary
The existing text matching method has the problem of calculating the alignment of sentence pairs and the actual alignment in the semantic relationship prediction of sentence pairs, resulting in a low correct recognition rate.
Using a method based on fine-grained information consistent inference network, a network composed of basic modules and inference modules is constructed, and global and local features are generated using character-level intersental interactions, and important fine-grained features are automatically selected through cross-attention learning and consistency reasoning to improve sentence prediction of semantic relationships.
It improves the accuracy of sentence prediction for semantic relationships and the accuracy of text matching, and reduces the situation of misidentification.
Smart Images

Figure CN120067303A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of text matching, and particularly to a text matching method and device based on a fine-grained information consistent inference network. Background Art
[0002] The text matching task is one of the very important basic tasks in natural language processing, aiming to study the semantic relevance between two sentences. This task plays an important role in a variety of actual application scenarios, such as dialogue systems, recommendation systems, question-and-answer systems, etc.
[0003] The core goal of the text matching task is to determine the semantic relationship between two sentences. Common tasks include similarity and paraphrase tasks (judging whether questions are equivalent) and inference tasks (judging whether the premise implies the hypothesis). With the wide application of large-scale pre-trained language models such as BERT (Bidirectional Encoder Representations from Transformers), RoBERTa (Robustly Optimized BERT), T5 (Text-to-Text Transfer Transformer), and the GPT (Generative Pretrained Transformer) series, most existing studies have taken pre-trained language models as the main architecture and achieved good results in the text matching task. Traditional methods usually concatenate two sentences together (e.g., [CLS]Sentence A[SEP]Sentence B[SEP]), directly input them into the pre-trained language model, and model the interaction between sentences through the Transformer in the pre-trained language model for classification. Recently, some existing methods have focused on using external knowledge to enhance the semantic representation of the pre-trained model to improve text matching performance. However, although these methods have achieved success, there is still a large gap between the alignment degree of sentence pairs calculated by them and the actual alignment. For example, on the development set of STS-B (Semantic Textual Similarity Benchmark Dataset), more than 50.78% of the relational facts are misidentified, and misidentifications still occur when more than half of the words in the sentence pair are literally the same. Therefore, a natural extension is to infer the semantic relationship of sentence pairs based on fine-grained information, which is particularly important in the text matching task.
[0004] Different from the single-sentence classification task, the sentence pair task faces greater challenges because it needs to consider the relevance of all relevant details between sentences (e.g., who does what, when, where, why). More importantly, the difficulty of relationship prediction for different sentence pairs may vary greatly. Some sentence pairs can be predicted by mainly checking the relevance of different information within the sentence pair, while other sentence pairs require a comprehensive analysis of all details in both sentences to establish the relationship between them. Summary of the Invention
[0005] Based on this, it is necessary to provide a text matching method and device based on a fine-grained information consistent reasoning network for the above technical problems, which can improve the prediction of the semantic relationship between sentence pairs by using the consistency reasoning between global information and more important fine-grained information, and further improve the accuracy of text matching.
[0006] A text matching method based on a fine-grained information consistent reasoning network, the method comprising: Construct a fine-grained information consistent reasoning network composed of a basic module and an inference module; wherein, the basic module encodes and generates the global features and local features of the input sentence pair by using character-level inter-sentence interaction, and calculates a preliminary prediction result of the semantic relationship between the sentence pairs based on the global features; the inference module generates a full information flow containing the text block pair information of all characters in the sentence pair and a different information flow containing the text block pair information of different characters in the sentence pair based on the cross-attention learning of the local features by using sentence-level inter-sentence interaction, and automatically selects the most important fine-grained features from the full information flow and the different information flows, and improves the preliminary prediction result of the semantic relationship between the sentence pairs according to the consistency reasoning between the global features and the fine-grained features to obtain the final prediction result of the semantic relationship between the sentence pairs; Input the text matching data set into the fine-grained information consistent reasoning network for training, and the training objective is to minimize the comprehensive loss of the preliminary prediction, final prediction and consistency reasoning of the semantic relationship between the sentence pairs until a trained fine-grained information consistent reasoning network is obtained to perform the text matching task.
[0007] In one embodiment, the basic module includes an encoder and a classifier based on Transformer; wherein, the encoder based on Transformer selects a cross-encoder, and encodes and generates the global features and local features of the input sentence pair by using the character-level inter-sentence interaction of the cross-encoder, and the cross-encoder is BERT; the classifier is used to calculate a preliminary prediction result of the semantic relationship between the sentence pairs according to the global features.
[0008] In one embodiment, encoding and generating the global features and local features of the input sentence pair by using the character-level inter-sentence interaction of the cross-encoder includes: Given a pair of sentences and , first insert special characters [CLS] and [SEP], and splice them into the encoder input , expressed as: ; where the subscripts and respectively represent the number of characters in sentence A and sentence B, and respectively represent the the character and the th character in sentence B; [CLS] is called the classification token, which is used to represent the start of the input sequence; [SEP] is called the separator token, which is used to separate two sentences or represent the end of a single sentence; Then, the pre-trained language model BERT is selected as the encoder to encode and generate the context representation of the sentence pair , and the expression is: ; Among them, the context representation contains the semantic representation of the special character [CLS] and the semantic representations of all characters in the sentence pair; among them, the semantic representation of the special character [CLS] is regarded as the global feature of the sentence pair and is used to further input into the classifier for the preliminary prediction of the semantic relationship of the sentence pair; the semantic representations of all characters in sentence A and the semantic representations of all characters in sentence B are regarded as the local features of the sentence pair and are used to further input into the inference module to generate and automatically select more important fine-grained features to optimize the preliminary prediction result of the semantic relationship of the sentence pair; and respectively represent the semantic representation of the th character in sentence A and the semantic representation of the th character in sentence B.
[0009] In one embodiment, the classifier calculates the preliminary prediction result of the semantic relationship of the sentence pair based on the global feature, including: Input the global feature into the classifier, and perform the preliminary prediction of the semantic relationship of the sentence pair according to the classifier to obtain the global probability of the semantic relationship of the sentence pair, which is expressed as: ; Among them, is the normalized exponential function, is the learnable parameter, and the superscript T represents the transpose.
[0010] In one embodiment, the inference module includes a cross-attention learning unit and a consistency inference unit; Among them, the cross-attention learning unit utilizes the soft-alignment interaction between two sentences at the sentence level. On the one hand, through cross-attention learning, it obtains a soft-alignment representation of the text block pair information containing all characters in the sentence pair, and uses it as the full information flow to calculate the full-character features of the sentence pair. On the other hand, also through cross-attention learning, it obtains a soft-alignment representation of the text block pair information containing different characters in the sentence pair, and uses it as the different information flow to calculate the different-character features of the sentence pair; The consistency reasoning unit interacts with the classifier in the basic module to calculate the consistency loss between the global probability of the semantic relationship of the sentence pair output by the classifier and the probabilities at different granularities calculated based on the full-character features and different-character features, reasons to determine the local probability of the semantic relationship of the sentence pair, and finally outputs the final probability of the semantic relationship of the sentence pair by performing a weighted sum of the global probability and the local probability.
[0011] In one embodiment, obtaining a soft-alignment representation of the text block pair information containing all characters in the sentence pair through cross-attention learning and using it as the full information flow to calculate the full-character features of the sentence pair includes: According to the semantic representations of all characters in sentence A in the local features and the semantic representations of all characters in sentence B , calculate the similarity matrix of the sentence pair , the similarity matrix The element in the th row and th column is expressed as: ; Among them, represents the similarity function, and the vector dot product calculation is adopted; and respectively represent the semantic representation of the th character in sentence A and the semantic representation of the th character in sentence B; Normalize the similarity matrix to obtain the soft-alignment weights of each character in one sentence to all characters in the other sentence; among them, the soft-alignment weight of each character in sentence A to all characters in sentence B is calculated as follows: ; Among them, is the normalization exponential function; Each element in represents the matching degree of the corresponding character pair; Multiply the semantic representations of all characters in sentence B by the corresponding soft-alignment weights Multiply them to obtain the soft alignment representation of each character in sentence A , which is expressed as: ; Similarly, multiply the semantic representations of all characters in sentence A by the corresponding soft alignment weights to obtain the soft alignment representation of each character in sentence B , and use and as the full information flow; Further, perform average pooling on and , then splice them, and successively pass through a feed - forward neural network and layer normalization processing to obtain the full character features of the sentence pair , which is expressed as: ; Among them, MP, FFN, and LN respectively represent average pooling, feed - forward neural network, and layer normalization processing.
[0012] In one embodiment, obtain the soft alignment representation of the text block pair information containing different characters in the sentence pair through cross - attention learning, and use it as different information flows to calculate and obtain different character features of the sentence pair, including: Construct a 0 - 1 matrix to represent the non - correlation between text block pairs. In the matrix , if a certain text block pair is the same, set the weights of the entire row and entire column where the text pair is located to 0, and the weights of the remaining positions to 1; According to the similarity matrix of the sentence pair and the matrix construct a matrix representing different detailed information of the sentence pair , which is defined as follows: ; Adopt the same calculation method as and , and respectively calculate and obtain the soft alignment representations in sentence A and in sentence B that are affected by different information , and use and as different information flows; Further, perform average pooling on and , then splice them, and successively pass through a feed - forward neural network and layer normalization processing to obtain different character features of the sentence pair.
[0013] In one embodiment, the consistency reasoning unit interacts with the classifier in the basic module, calculates the consistency loss between the global probability of the semantic relationship of the sentence pair output by the classifier and the probabilities of different granularities calculated based on the full character features and different character features, determines the local probability of the semantic relationship of the sentence pair by reasoning, and finally outputs the final probability of the semantic relationship of the sentence pair by weighted summation of the global probability and the local probability, including: The consistency reasoning unit interacts with the classifier in the basic module and uses the full character features to and different character features Calculate the probability of different granularities of semantic relations between sentences and ; Then, the global probability of the semantic relationship between the sentences output by the classifier is calculated by KL divergence. Consistency loss between predictions of different granularities , expressed as: ; if , the local probability of the semantic relationship between sentences Set to , and learn the consistency loss , which means that the details emphasized by different information flows are selected to refine the initial prediction of the sentence's semantic relationship; if ,Will Set to ,and , which means that the full information flow is selected to supplement the missing information of the basic module, thereby optimizing the initial prediction of the semantic relationship of the sentence; Finally, by the global probability With local probability Perform weighted summation and output the final probability of the semantic relationship between the sentences , expressed as: ; in, and Represents a trade-off parameter.
[0014] In one embodiment, the comprehensive loss function of the fine-grained information consistent inference network It is expressed as: ; in, 、 and They represent the trade-off parameters of the losses of each component; Perform preliminary predictions on the basic module to obtain the global probability of the semantic relationship between sentences The cross entropy loss of Make the final prediction for the reasoning module to obtain the final probability of the semantic relationship between the sentences The cross entropy loss of is the consistency loss; For a given training instance, the sentence pair and composition; express The real label.
[0015] A text matching device based on a fine-grained information consistent reasoning network, the device comprising: A network construction module is used to construct a fine-grained information consistent reasoning network composed of a basic module and a reasoning module; wherein the basic module encodes and generates global features and local features of an input sentence pair by utilizing the interactions between sentences at the character level, and calculates the preliminary prediction results of the semantic relationship of the sentence pair based on the global features; the reasoning module utilizes the interactions between sentences at the sentence level, generates a full information stream containing the information of text blocks of all characters in the sentence pair and different information streams containing the information of text blocks of different characters in the sentence pair based on cross-attention learning of local features, and automatically selects the most important fine-grained features from the full information stream and different information streams, improves the preliminary prediction results of the semantic relationship of the sentence pair based on the consistent reasoning between the global features and the fine-grained features, and obtains the final prediction results of the semantic relationship of the sentence pair; The training and task execution module is used to input the text matching dataset into the fine-grained information consistent reasoning network for training. The training goal is to minimize the comprehensive loss of the initial prediction, final prediction and consistent reasoning of the semantic relationship of the sentence until a trained fine-grained information consistent reasoning network is obtained to perform the text matching task.
[0016] The above-mentioned text matching method and device based on the fine-grained information consistent reasoning network realizes the preliminary prediction of the semantic relationship between sentence pairs based on the basic module in the network, and further introduces fine-grained information and consistent reasoning based on the reasoning module in the network to optimize the preliminary prediction of the basic module. Different from the existing method that only considers the overall characteristics of the sentence pair, the reasoning module designed in this application designs an expandable cross-attention learning unit, which enables it to utilize the soft alignment interaction between two sentences to generate fine-grained information. In addition, the reasoning module also automatically selects more important fine-grained information through consistent constraint reasoning to further improve the accuracy of the semantic relationship prediction between sentence pairs, thereby improving the accuracy of text matching. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1Schematic flowchart of a text matching method based on a fine-grained information consistency inference network in an embodiment; Figure 2 Schematic diagram of the overall architecture of a fine-grained information consistency inference network in an embodiment; Figure 3 Schematic flowchart of obtaining the full-character features of a sentence pair through cross-attention learning calculation in an embodiment. Detailed implementation manners
[0018] In order to make the objectives, technical solutions, and advantages of the present application clearer and more understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0019] In one embodiment, as Figure 1 shown, a text matching method based on a fine-grained information consistency inference network is provided, including the following steps: Step S1, constructing a fine-grained information consistency inference network composed of a basic module and an inference module.
[0020] The overall architecture of the fine-grained information consistency inference network is as Figure 2 shown, where the basic module includes an encoder and a classifier based on Transformer.
[0021] In modern text matching methods, sentence pair modeling methods based on pre-trained language models have made significant progress. There are mainly two common architectures: bi-encoders and cross-encoders.
[0022] The bi-encoder architecture first encodes the two sentences separately and then calculates their similarity in the embedding space. Its advantage is high computational efficiency and is particularly suitable for large-scale data. Common bi-encoder models include: (1) SBERT (Sentence-BERT): Fine-tuning BERT through a siamese / triplet network structure for efficient calculation of sentence pair similarity; (2) Enhanced SBERT: Proposing a simple but efficient data augmentation strategy; (3) PromptBERT (Prompt-based BERT): Proposing a sentence embedding method based on prompt learning.
[0023] The cross-encoder architecture processes sentence pairs in one encoding process and directly outputs a classification score to predict the relationship between the sentence pairs. The cross-encoder can model the lexical-level interactions between sentences and usually obtain more accurate results. Common cross-encoder methods include: (1) BERT: Using a Transformer encoder to perform word-level interactions on sentences and output a global feature for classification; (2) To improve the semantic representation of pre-trained language models, SemBERT (Semantic BERT) was proposed, which introduces explicit semantic role annotation into the model through a lightweight fine-tuning method; (3) UERBERT (Universal Pre-trained Encoder BERT): Injecting synonym knowledge directly into the attention mechanism to improve semantic text similarity; (4) SyntaxBERT (Syntactic BERT): Integrating syntactic tree information in a plug-and-play manner in the Transformer model to improve syntactic understanding; (5) DAFA (Domain Adaptation Fine-tuning Approach): Efficiently incorporating dependency structure prior information into the pre-trained model. Since the cross-encoder can directly model the interactions between sentences, it is usually superior to the dual-encoder in performance.
[0024] Based on this, the Transformer-based encoder in the above basic module selects the cross-encoder BERT, uses the character-level inter-sentence interaction of the cross-encoder to encode and generate the global and local features of the input sentence pair, so as to improve the lexical or logical alignment degree of the corresponding detailed information between the two sentences. The classifier in the above basic module is used to calculate the preliminary prediction result of the semantic relationship of the sentence pair according to the global feature.
[0025] Specifically, the data processing process of the Transformer-based encoder includes: Given a pair of sentences and , first insert the special characters [CLS] and [SEP], and concatenate them into the encoder input , which is expressed as: (1) where the subscripts and represent the number of characters in sentence A and sentence B respectively, and represent the th character in sentence A and the th character in sentence B respectively; [CLS] is called the classification token, which is used to represent the start of the input sequence; [SEP] is called the separator token, which is used to separate the two sentences or represent the end of a single sentence.
[0026] Then, select the pre-trained language model BERT as the encoder , and encode to generate the context representation of the sentence pair , the expression is: (2) Among them, the context representation includes the semantic representation of the special character [CLS] and the semantic representations of all characters in the sentence pair; among them, the semantic representation of the special character [CLS] is regarded as the global feature of the sentence pair and is used to further input into the classifier for the preliminary prediction of the semantic relationship of the sentence pair; the semantic representations of all characters in sentence A and the semantic representations of all characters in sentence B are regarded as the local features of the sentence pair and are used to further input into the inference module to generate and automatically select more important fine-grained features to optimize the preliminary prediction result of the semantic relationship of the sentence pair; and respectively represent the semantic representations of the th character in sentence A and the th character in sentence B.
[0027] Specifically, the data processing process of the classifier includes: inputting the global feature into the classifier, and making a preliminary prediction of the semantic relationship of the sentence pair according to the classifier to obtain the global probability of the semantic relationship of the sentence pair , which is expressed as: (3) Among them, is the normalized exponential function, is the learnable parameter, and the superscript T represents the transpose.
[0028] Furthermore, in order to overcome the limitation of the basic module in ignoring important details, the present application introduces an inference module in the fine-grained information consistent inference network to improve the prediction of the semantic relationship of the sentence pair by using the consistency inference between the global information and the more important fine-grained information. Specifically, as Figure 2 shown, the inference module includes a cross-attention learning unit and a consistency inference unit.
[0029] The cross-attention learning unit obtains the soft-alignment interaction between the two sentences at the sentence level. On the one hand, through cross-attention learning, it obtains the soft-alignment representation of the text block pair information containing all characters in the sentence pair and uses it as the full information flow to calculate and obtain the full character features of the sentence pair; on the other hand, also through cross-attention learning, it obtains the soft-alignment representation of the text block pair information containing different characters in the sentence pair and uses it as the different information flow to calculate and obtain the different character features of the sentence pair.
[0030] Specifically, as Figure 3As shown in the figure, a soft alignment representation of text block pairs containing all characters in the sentence pair is obtained through cross-attention learning, and it is used as the full information flow calculation to obtain the full character features of the sentence pair, including: First, according to the semantic representations of all characters in sentence A in the local features and the semantic representations of all characters in sentence B , the similarity matrix of the sentence pair is calculated . The similarity matrix The element in the th row and the th column is expressed as: (4) where represents the similarity function, which is calculated by vector dot product; and respectively represent the semantic representation of the th character in sentence A and the semantic representation of the j th character in sentence B.
[0031] Next, the similarity matrix is normalized to obtain the soft alignment weights of each character in one sentence to all characters in the other sentence. Among them, the soft alignment weights of each character in sentence A to all characters in sentence B are calculated as follows: (5) where is the normalized exponential function; Each element in
[0032] represents the matching degree of the corresponding character pair. Then, the semantic representations of all characters in sentence B are multiplied by the corresponding soft alignment weights to obtain the soft alignment representation of each character in sentence A, which is expressed as: Similarly, following similar steps as above, the semantic representations of all characters in sentence A are multiplied by the corresponding soft alignment weights to obtain the soft alignment representation of each character in sentence B, and and are used as the full information flow. and represent the soft alignment representation of one sentence relative to the other sentence, covering the influence of all detailed information.
[0033] Further, after performing average pooling and concatenation on and and then passing through a feed-forward neural network and layer normalization in sequence, the full character features of the sentence pair are obtained , expressed as: (7) where MP, FFN, and LN represent average pooling, feed-forward neural network, and layer normalization respectively.
[0034] Specifically, through cross-attention learning, a soft alignment representation containing the information of text block pairs with different characters in the sentence pair is obtained, and it is used as different information streams to calculate and obtain the different character features of the sentence pair, including: First, construct a 0-1 matrix to represent the non-correlation between text block pairs. In the matrix , if a certain text block pair is the same, the weights of the entire row and entire column where the text pair is located are set to 0, and the weights of the remaining positions are set to 1.
[0035] Second, according to the similarity matrix of the sentence pair and the matrix , construct a matrix representing different detailed information of the sentence pair, defined as follows: (8) Then, adopt the same calculation method as shown in formulas (5) and (6) for and , and based on , calculate and obtain the soft alignment representations and in sentence A and sentence B affected by different information respectively, and take and as different information streams. In different information streams, focus on those text block pairs that are not literally the same and ignore those that are literally the same. Different text block pairs refer to text blocks that appear in one sentence but not in the other, indicating the differences between the two.
[0036] Further, adopt the calculation formula shown in formula (7), perform average pooling and concatenation on and and then pass through a feed-forward neural network and layer normalization in sequence to obtain the different character features of the sentence pair.
[0037] Considering that the relationship prediction from global features and local features is consistent, the present application improves the prediction by automatically selecting the most relevant fine-grained features from the full information stream and different information streams. During this selection process, it is necessary to ensure that the probabilities of the global features and the selected fine-grained features satisfy the consistency constraint. Specifically, in the inference module, the consistency inference unit interacts with the classifier in the basic module to calculate the consistency loss between the global probability of the semantic relationship of the sentence pair output by the classifier and the probabilities of different granularities calculated based on the full character features and different character features, and infers and determines the local probability of the semantic relationship of the sentence pair. Finally, by performing a weighted sum of the global probability and the local probability, the final probability of the semantic relationship of the sentence pair is output. The data processing process of the consistency inference unit includes: The consistency inference unit interacts with the classifier in the basic module based on the full character features and different character features to calculate the probabilities of different granularities of the semantic relationship of the sentence pair and .
[0038] Subsequently, the global probability of the semantic relationship of the sentence pair output by the classifier is calculated through the KL (Kullback-Leibler) divergence and the consistency loss between different granularity predictions , expressed as: (9) If , the local probability of the semantic relationship of the sentence pair is set to , and the learning consistency loss . This is because the full information stream pays more attention to the same text pair. At this time, it means selecting the detailed information emphasized by different information streams to refine the preliminary prediction of the semantic relationship of the sentence pair; if , is set to , and . At this time, it means selecting the full information stream to supplement the information missing in the basic module, thereby optimizing the preliminary prediction of the semantic relationship of the sentence pair.
[0039] Finally, by performing a weighted sum of the global probability and the local probability , the final probability of the semantic relationship of the sentence pair is output , expressed as: (10) Wherein, and represent trade-off parameters.
[0040] Step S2: Input the text matching dataset into the fine-grained information consistency inference network for training. The training objective is to minimize the comprehensive loss of the preliminary prediction, final prediction, and consistency inference of the semantic relationship between sentence pairs until the trained fine-grained information consistency inference network is obtained to perform the text matching task.
[0041] Specifically, the comprehensive loss function of the fine-grained information consistency inference network is expressed as: (11) where 、 and respectively represent the trade-off parameters of each component loss; is the cross-entropy loss of the global probability of the semantic relationship between sentence pairs obtained by the basic module for preliminary prediction ; is the cross-entropy loss of the final probability of the semantic relationship between sentence pairs obtained by the inference module for final prediction ; is the consistency loss; is the given training instance, composed of the sentence pair and ; represents 's true label.
[0042] In summary, the method proposed in this application constructs a fine-grained information consistency inference network including a basic module and an inference module. Based on the basic module, the preliminary prediction of the semantic relationship between sentence pairs is realized. Based on the inference module, fine-grained information and consistency inference are further introduced to optimize the preliminary prediction of the basic module. Moreover, an extensible cross-attention learning unit is designed in the inference module, enabling it to utilize the soft alignment interaction between two sentences to generate fine-grained information. In addition, in the inference module, more important fine-grained information is automatically selected through consistency constraint inference to further improve the accuracy of the semantic relationship prediction between sentence pairs, thereby improving the correct rate of text matching.
[0043] Furthermore, comparative empirical experiments on six public datasets of GLUE (General Language Understanding Evaluation Benchmark) are conducted to verify the effectiveness of the fine-grained information consistent inference network (hereinafter referred to as ConInf) constructed by the method proposed in this application. The six public datasets are: MRPC (Microsoft Research Paraphrase Corpus), QQP (Quora Question Pairs), STS-B, MNLI-m / mm (Multi-Genre Natural Language Inference (matched / mismatched) dataset), QNLI (Question Natural Language Inference dataset), and RTE (Recognizing Textual Entailment dataset). The performance evaluation parameters include the F1 value of sentence similarity and sentence inference, Pearson correlation coefficient (Pearson), and accuracy (Accuracy, Acc).
[0044] Table 1 Test results of different models on six datasets of GLUE under the BERT-base (basic version of BERT) configuration:
[0045] Table 2 Test results of different models on six datasets of GLUE under the BERT-large (high-performance version of BERT) configuration:
[0046] Table 1 shows the test results of different models on six datasets of GLUE under the BERT-base configuration. The experimental models include UERBERT base , SemBERT base , SyntaxBERT base , DAFA base and ConInf base . Table 2 shows the test results of different models on six datasets of GLUE under the BERT-large configuration. The experimental models include BERT large , Prompt-tuning + BERT large (prompt-tuned BERT), SyntaxBERT large , DAFA large and ConInf large . The values with superscript "+" in the last column of Table 1 and Table 2 represent the average scores of performance parameters on the other five datasets except the QQP dataset. As shown in Table 1, under the BERT-base configuration, the fine-grained information consistent inference network ConInf constructed in this application baseIt is significantly better than all other methods in terms of average score (83.63% v.s. 82.68% and 85.55% v.s. 85.18%), which indicates the effectiveness of the text matching method based on the fine-grained information consistent reasoning network proposed in this application. As shown in Table 2, under the BERT-large configuration, ConInf large is significantly better than BERT large and Prompt-tuning + BERT large in terms of average score (85.60% v.s. 83.39% and 85.60% v.s. 82.18%), and it is consistently better than these basic pre-trained language model methods on six different datasets, indicating the effectiveness of fine-grained information. The experimental results show that the fine-grained information consistency reasoning network constructed in this application has achieved significant improvements on six common datasets of GLUE, reaching an average score of 83.63% (+0.95 absolute improvement) under the BERT-base configuration and an average score of 85.60% (+0.63% absolute improvement) under the BERT-large configuration, demonstrating its effectiveness and advantages in text matching tasks.
[0047] In one embodiment, a text matching device based on a fine-grained information consistent reasoning network is provided, including: A network construction module for constructing a fine-grained information consistent reasoning network composed of a basic module and an inference module; wherein, the basic module encodes and generates the global features and local features of the input sentence pair by utilizing the character-level interaction between sentences, and calculates the preliminary prediction result of the semantic relationship of the sentence pair based on the global features; the inference module generates the full information flow containing the text block pair information of all characters in the sentence pair and the different information flow containing the text block pair information of different characters in the sentence pair through cross-attention learning based on the local features, and automatically selects the most important fine-grained features from the full information flow and the different information flows, and improves the preliminary prediction result of the semantic relationship of the sentence pair according to the consistency reasoning between the global features and the fine-grained features to obtain the final prediction result of the semantic relationship of the sentence pair; A training and task execution module for inputting the text matching dataset into the fine-grained information consistent reasoning network for training, and the training objective is to minimize the comprehensive loss of the preliminary prediction, final prediction and consistency reasoning of the semantic relationship of the sentence pair until the trained fine-grained information consistent reasoning network is obtained to execute the text matching task.
[0048] For the specific limitations of the text matching device based on the fine-grained information consistent reasoning network, reference can be made to the limitations of the text matching method based on the fine-grained information consistent reasoning network in the foregoing text, which will not be elaborated herein. Each module in the above text matching device based on the fine-grained information consistent reasoning network can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in the processor of the computer device in hardware form or independent of it, or stored in the memory of the computer device in software form, so as to facilitate the processor to call and execute the operations corresponding to the above modules.
[0049] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0050] The above-described embodiments only represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the appended claims.
Claims
1. A text matching method based on fine-grained information consistent reasoning network, characterized in that: The method comprises: Construct a fine-grained information consistent reasoning network composed of a basic module and a reasoning module; wherein the basic module encodes and generates global features and local features of an input sentence pair by utilizing the sentence-to-sentence interaction at the character level, and calculates a preliminary prediction result of the semantic relationship of the sentence pair based on the global features; the reasoning module utilizes the sentence-to-sentence interaction at the sentence level, generates a full information stream containing the text block pair information of all characters in the sentence pair and different information streams containing the text block pair information of different characters in the sentence pair based on cross-attention learning of the local features, and automatically selects the most important fine-grained features from the full information stream and the different information streams, improves the preliminary prediction result of the semantic relationship of the sentence pair based on the consistent reasoning between the global features and the fine-grained features, and obtains the final prediction result of the semantic relationship of the sentence pair; The text matching dataset is input into the fine-grained information consistent reasoning network for training, and the training goal is to minimize the comprehensive loss of the initial prediction, final prediction and consistent reasoning of the sentence pair semantic relationship until a trained fine-grained information consistent reasoning network is obtained to perform the text matching task.
2. The method according to claim 1, characterized in that The basic module includes a Transformer-based encoder and a classifier; wherein the Transformer-based encoder selects a cross encoder, utilizes the character-level sentence interaction of the cross encoder, encodes and generates the global features and local features of the input sentence pair, and the cross encoder is BERT; the classifier is used to calculate the preliminary prediction results of the semantic relationship of the sentence pair based on the global features.
3. The method according to claim 2, characterized in that By using the character-level sentence interactions of the cross encoder, the encoding generates the global and local features of the input sentence pair, including: Given a pair of sentences and , first insert the special characters [CLS] and [SEP] and concatenate them as encoder input , expressed as: ; Among them, the subscript and Represents the number of characters in sentence A and sentence B respectively, and Respectively represent the characters and the first characters; [CLS] is called a classification marker, which is used to indicate the beginning of an input sequence; [SEP] is called a separator marker, which is used to separate two sentences or indicate the end of a single sentence; Then, the pre-trained language model BERT is selected as the encoder , encoding the contextual representation of the generated sentence pair , the expression is: ; The context represents Contains the semantic representation of special characters [CLS] and the semantic representation of all characters in the sentence pair; among them, the semantic representation of special characters [CLS] It is regarded as the global feature of the sentence pair and is used to further input the classifier to make a preliminary prediction of the semantic relationship between the sentence pairs; the semantic representation of all characters in sentence A and the semantic representation of all characters in sentence B It is regarded as a local feature of the sentence pair, which is used to further input the reasoning module to generate and automatically select more important fine-grained features to optimize the preliminary prediction results of the semantic relationship of the sentence pair; and Respectively represent the The semantic representation of the characters and the The semantic representation of a character.
4. The method according to claim 3, characterized in that: The classifier calculates preliminary prediction results of the semantic relationship of the sentence pair based on the global features, including: Global Features Input the classifier, make a preliminary prediction of the semantic relationship between the sentence pairs according to the classifier, and obtain the global probability of the semantic relationship between the sentence pairs , expressed as: ; in, is the normalized exponential function, is a learnable parameter, and the superscript T Indicates transpose.
5. The method according to claim 3, characterized in that: The reasoning module includes a cross-attention learning unit and a consistency reasoning unit; The cross-attention learning unit uses sentence-level learning to learn the soft alignment interaction between two sentences. On the one hand, the cross-attention learning unit obtains the soft alignment representation of the text block pair information containing all characters in the sentence pair, and uses it as the full information flow calculation to obtain the full character features of the sentence pair. On the other hand, the cross-attention learning unit also obtains the soft alignment representation of the text block pair information containing different characters in the sentence pair, and uses it as different information flow calculations to obtain different character features of the sentence pair. The consistency reasoning unit interacts with the classifier in the basic module, calculates the consistency loss between the global probability of the semantic relationship of the sentence pair output by the classifier and the probabilities of different granularities calculated based on the full character features and different character features, and determines the local probability of the semantic relationship of the sentence pair by reasoning. Finally, the final probability of the semantic relationship of the sentence pair is output by weighted summing the global probability and the local probability.
6. The method according to claim 5, characterized in that Through cross-attention learning, a soft-aligned representation of the text block pair information containing all characters in the sentence pair is obtained, and it is used as a full information flow calculation to obtain the full character features of the sentence pair, including: Based on the semantic representation of all characters in sentence A in the local features and the semantic representation of all characters in sentence B , calculate the similarity matrix of sentence pairs , similarity matrix The Line Column Elements It is expressed as: ; in, Represents the similarity function, which uses vector dot product calculation; and Respectively represent the The semantic representation of the characters and the The semantic representation of characters; Similarity matrix Normalize and obtain the soft alignment weight of each character in one sentence to all characters in another sentence; among them, the soft alignment weight of each character in sentence A to all characters in sentence B is The calculation is as follows: ; in, is the normalized exponential function; Each element in Indicates the matching degree of corresponding character pairs; The semantic representation of all characters in sentence B And the corresponding soft alignment weight Multiply them together to get the soft alignment representation of each character in sentence A , expressed as: ; Similarly, the semantic representation of all characters in sentence A is And the corresponding soft alignment weight Multiply them together to get the soft alignment representation of each character in sentence B , and and As a full information flow; Further and After average pooling and concatenation, the full character features of the sentence pair are obtained by passing through a feedforward neural network and layer normalization. , expressed as: ; Among them, MP, FFN and LN represent average pooling, feedforward neural network and layer normalization processing respectively.
7. The method according to claim 6, characterized in that Through cross-attention learning, we obtain the soft alignment representation of the text block pair information containing different characters in the sentence pair, and use it as different information flow calculations to obtain different character features of the sentence pair, including: Construct a 0-1 matrix , used to represent the non-correlation between text block pairs, in the matrix In , if a pair of text blocks is the same, the weight of the entire row and column where the text pair is located is set to 0, and the weight of the remaining positions is set to 1; Based on the similarity matrix of sentence pairs With the matrix Construct a matrix representing different details of a sentence pair , The definition is as follows: ; Adopt and and The same calculation method is based on Calculate and obtain the soft alignment representations of sentences A and B that are affected by different information respectively and , and and as different information flows; Further and After average pooling and concatenation, the different character features of the sentence pairs are obtained by passing through a feedforward neural network and layer normalization. .
8. The method according to claim 5, characterized in that The consistency reasoning unit interacts with the classifier in the basic module, calculates the consistency loss between the global probability of the semantic relationship of the sentence pair output by the classifier and the probabilities of different granularities calculated based on the full character features and different character features, determines the local probability of the semantic relationship of the sentence pair by reasoning, and finally outputs the final probability of the semantic relationship of the sentence pair by weighted summation of the global probability and the local probability, including: The consistency reasoning unit interacts with the classifier in the basic module and performs the classification based on the full character features. and different character features Calculate the probability of different granularities of semantic relations between sentences and ; Then, the global probability of the semantic relationship between the sentences output by the classifier is calculated by KL divergence. Consistency loss between predictions of different granularities , expressed as: ; if , the local probability of the semantic relationship between sentences Set to , and learn the consistency loss , which means that the details emphasized by different information flows are selected to refine the initial prediction of the sentence's semantic relationship; if ,Will Set to ,and , which means that the full information flow is selected to supplement the missing information of the basic module, thereby optimizing the initial prediction of the semantic relationship of the sentence; Finally, by the global probability With local probability Perform weighted summation and output the final probability of the semantic relationship between the sentences , expressed as: ; in, and Represents a trade-off parameter.
9. The method according to claim 1, characterized in that: The comprehensive loss function of the fine-grained information consistent inference network It is expressed as: ; in, 、 and They represent the trade-off parameters of the losses of each component; Perform preliminary predictions on the basic module to obtain the global probability of the semantic relationship between sentences The cross entropy loss of Make the final prediction for the reasoning module to obtain the final probability of the semantic relationship between the sentences The cross entropy loss of is the consistency loss; For a given training instance, the sentence pair and composition; express The real label.
10. A text matching device based on a fine-grained information consistent reasoning network, characterized in that: The device comprises: A network construction module, for constructing a fine-grained information consistent reasoning network composed of a basic module and a reasoning module; wherein the basic module encodes and generates global features and local features of an input sentence pair by utilizing the sentence-to-sentence interaction at the character level, and calculates a preliminary prediction result of the semantic relationship of the sentence pair based on the global features; the reasoning module utilizes the sentence-to-sentence interaction at the sentence level, generates a full information stream containing text block pair information of all characters in the sentence pair and different information streams containing text block pair information of different characters in the sentence pair based on cross-attention learning of the local features, and automatically selects the most important fine-grained features from the full information stream and the different information streams, improves the preliminary prediction result of the semantic relationship of the sentence pair based on the consistent reasoning between the global features and the fine-grained features, and obtains the final prediction result of the semantic relationship of the sentence pair; The training and task execution module is used to input the text matching data set into the fine-grained information consistent reasoning network for training. The training goal is to minimize the comprehensive loss of the initial prediction, final prediction and consistency reasoning of the sentence semantic relationship until a trained fine-grained information consistent reasoning network is obtained to perform the text matching task.
Citation Information
Patent Citations
Text implication recognition method and device
CN109165300A
A text implication relation recognition method based on multi-granularity information fusion
CN109299262A
Dual-attention natural language reasoning based on situational awareness
CN109344404A
Systems and methods for processing machine learning language model classification outputs via text block masking
US20230334887A1
Efficient and compact text matching system for sentence pairs
WO2022103440A1