Text Matching Method and Device Based on Fine-Grained Information Consistent Reasoning Network

By constructing a fine-grained information consistent inference network, and using cross-attention learning and consistent inference to optimize sentence-to-semantic relationship prediction, the accuracy problem of sentence-to-semantic relationship prediction in existing methods is solved, and a higher text matching accuracy rate is achieved.

CN120067303BActive Publication Date: 2025-07-08NAT UNIV OF DEFENSE TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510540049.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-27
Publication Date
2025-07-08
Estimated Expiration
2045-04-27

AI Technical Summary

Technical Problem

The existing text matching methods have a gap between the computational alignment and actual alignment in the semantic relationship prediction of sentence pairs. Especially when more than half of the words in the sentence pair are the same literally, it is difficult to effectively use fine-grained information to make accurate predictions.

Method used

A consistent inference network based on fine-grained information is constructed, global and local features are generated through basic modules, and cross-attention learning and consistency inference units are used to automatically select important fine-grained features to optimize sentence prediction of semantic relationships.

Benefits of technology

It improves the accuracy of text matching and significantly improves the accuracy of sentence prediction of semantic relationships, especially in the case of fine-grained differences in sentence pairs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120067303B_ABST
    Figure CN120067303B_ABST
Patent Text Reader

Abstract

The present application relates to a text matching method and device based on a fine-grained information consistent reasoning network. The method includes: constructing a fine-grained information consistent reasoning network composed of a basic module and a reasoning module; wherein, the basic module makes a preliminary prediction of the semantic relationship between sentence pairs by using character-level inter-sentence interaction; the reasoning module generates fine-grained information by using sentence-level inter-sentence interaction, introduces cross-attention learning, and introduces a consistency constraint reasoning to automatically select more important fine-grained information to optimize the preliminary prediction. This method can improve the prediction of the semantic relationship between sentence pairs by using the consistency reasoning between global information and more important fine-grained information, thereby improving the accuracy of text matching.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of text matching, and in particular, to a text matching method and apparatus based on a fine-grained information consistent inference network. Background Art

[0002] The text matching task is one of the very important basic tasks in natural language processing, aiming to study the semantic relevance between two sentences. This task plays an important role in various practical application scenarios, such as dialogue systems, recommendation systems, question-and-answer systems, etc.

[0003] The core goal of the text matching task is to determine the semantic relationship between two sentences. Common tasks include similarity and paraphrase tasks (judging whether questions are equivalent) and inference tasks (judging whether a premise implies a hypothesis). With the wide application of large-scale pre-trained language models such as BERT (Bidirectional Encoder Representations from Transformers), RoBERTa (Robustly Optimized BERT), T5 (Text-to-Text Transfer Transformer), and GPT (Generative Pretrained Transformer) series, most existing studies have used pre-trained language models as the main architecture and achieved good results in the text matching task. Traditional methods usually concatenate two sentences (for example, [CLS]Sentence A[SEP]Sentence B[SEP]), directly input them into the pre-trained language model, and model the interaction between sentences through the Transformer in the pre-trained language model for classification. Recently, some existing methods have focused on using external knowledge to enhance the semantic representation of the pre-trained model to improve text matching performance. However, although these methods have achieved success, there is still a large gap between the alignment degree of sentence pairs calculated by them and the actual alignment. For example, on the development set of STS-B (Semantic Textual Similarity Benchmark Dataset), more than 50.78% of the relational facts are misidentified, and misidentifications still occur when more than half of the words in the sentence pair are literally the same. Therefore, a natural extension is to infer the semantic relationship of sentence pairs based on fine-grained information, which is particularly important in the text matching task.

[0004] Different from the single-sentence classification task, the sentence pair task faces greater challenges because it needs to consider the relevance of all relevant details between sentences (for example, who does what, when, where, and why). More importantly, the difficulty of relationship prediction for different sentence pairs may vary greatly. Some sentence pairs can be predicted by mainly checking the relevance of different information within the sentence pair, while other sentence pairs require a comprehensive analysis of all details in the two sentences to establish the relationship between them. Summary of the Invention

[0005] Based on this, it is necessary to provide a text matching method and apparatus based on a fine-grained information consistency inference network for the above technical problems, which can improve the prediction of the semantic relationship between sentence pairs by leveraging the consistency inference between global information and more important fine-grained information, thereby improving the accuracy of text matching.

[0006] A text matching method based on a fine-grained information consistency inference network, the method comprising:

[0007] Construct a fine-grained information consistency inference network composed of a basic module and an inference module; wherein, the basic module encodes and generates the global features and local features of the input sentence pair by leveraging character-level inter-sentence interaction, and calculates a preliminary prediction result of the semantic relationship between the sentence pair based on the global features; the inference module generates a full information flow containing text block pair information of all characters in the sentence pair and a different information flow containing text block pair information of different characters in the sentence pair through cross-attention learning based on the local features by leveraging sentence-level inter-sentence interaction, and automatically selects the most important fine-grained features from the full information flow and the different information flows, and improves the preliminary prediction result of the semantic relationship between the sentence pair according to the consistency inference between the global features and the fine-grained features to obtain the final prediction result of the semantic relationship between the sentence pair;

[0008] Input a text matching data set into the fine-grained information consistency inference network for training, and the training objective is to minimize the comprehensive loss of the preliminary prediction, final prediction, and consistency inference of the semantic relationship between the sentence pair until a trained fine-grained information consistency inference network is obtained to perform the text matching task.

[0009] In one embodiment, the basic module includes a Transformer-based encoder and a classifier; wherein, the Transformer-based encoder selects a cross-encoder, and encodes and generates the global features and local features of the input sentence pair by leveraging the character-level inter-sentence interaction of the cross-encoder, and the cross-encoder is BERT; the classifier is used to calculate a preliminary prediction result of the semantic relationship between the sentence pair based on the global features.

[0010] In one embodiment, encoding and generating the global features and local features of the input sentence pair by leveraging the character-level inter-sentence interaction of the cross-encoder includes:

[0011] Given a pair of sentences and , first insert special characters [CLS] and [SEP], and splice them into the encoder input , denoted as:

[0012] ;

[0013] wherein, the subscript and respectively represent the number of characters in sentence A and sentence B, and respectively represent the -th character in sentence A and the -th character in sentence B; [CLS] is called the classification marker, which is used to represent the start of the input sequence; [SEP] is called the separator marker, which is used to separate two sentences or represent the end of a single sentence;

[0014] Then, the pre-trained language model BERT is selected as the encoder , and the context representation of the sentence pair is encoded and generated , and the expression is:

[0015] ;

[0016] Among them, the context representation includes the semantic representation of the special character [CLS] and the semantic representations of all characters in the sentence pair; among them, the semantic representation of the special character [CLS] is regarded as the global feature of the sentence pair, which is used to further input into the classifier for the preliminary prediction of the semantic relationship of the sentence pair; the semantic representations of all characters in sentence A and the semantic representations of all characters in sentence B are regarded as the local features of the sentence pair, which are used to further input into the inference module to generate and automatically select more important fine-grained features to optimize the preliminary prediction result of the semantic relationship of the sentence pair; and respectively represent the semantic representation of the -th character in sentence A and the semantic representation of the -th character in sentence B.

[0017] In one embodiment, the classifier calculates the preliminary prediction result of the semantic relationship of the sentence pair based on the global feature, including:

[0018] Input the global feature into the classifier, and according to the classifier, the preliminary prediction of the semantic relationship of the sentence pair is made to obtain the global probability of the semantic relationship of the sentence pair, which is expressed as:

[0019] ;

[0020] Among them, is the normalized exponential function, is the learnable parameter, and the superscript T represents the transpose.

[0021] In one embodiment, the inference module includes a cross-attention learning unit and a consistency inference unit;

[0022] Among them, the cross-attention learning unit obtains a soft-alignment representation of the text block pair information containing all characters in the sentence pair by leveraging the soft-alignment interaction between two sentences at the sentence level. On the one hand, through cross-attention learning, a soft-alignment representation of the text block pair information containing all characters in the sentence pair is obtained and used as the full information flow to calculate the full-character features of the sentence pair. On the other hand, also through cross-attention learning, a soft-alignment representation of the text block pair information containing different characters in the sentence pair is obtained and used as the different information flow to calculate the different-character features of the sentence pair.

[0023] The consistency reasoning unit interacts with the classifier in the basic module to calculate the consistency loss between the global probability of the semantic relationship of the sentence pair output by the classifier and the probabilities at different granularities calculated based on the full-character features and different-character features, and infers and determines the local probability of the semantic relationship of the sentence pair. Finally, by performing a weighted sum of the global probability and the local probability, the final probability of the semantic relationship of the sentence pair is output.

[0024] In one embodiment, obtaining a soft-alignment representation of the text block pair information containing all characters in the sentence pair by cross-attention learning and using it as the full information flow to calculate the full-character features of the sentence pair includes:

[0025] Based on the semantic representations of all characters in sentence A in the local features and the semantic representations of all characters in sentence B , calculate the similarity matrix of the sentence pair , and the element in the th row and th column of the similarity matrix is represented as: ;

[0026] ;

[0027] Among them, represents the similarity function, which is calculated using the vector dot product; and respectively represent the semantic representation of the th character in sentence A and the semantic representation of the th character in sentence B;

[0028] Normalize the similarity matrix to obtain the soft-alignment weights of each character in one sentence to all characters in the other sentence. Among them, the soft-alignment weight of each character in sentence A to all characters in sentence B is calculated as follows:

[0029] ;

[0030] Among them, is the normalized exponential function; each element in represents the matching degree of the corresponding character pair;

[0031] the semantic representations of all characters in sentence B are multiplied by the corresponding soft alignment weights to obtain the soft alignment representation of each character in sentence A , which is expressed as:

[0032] ;

[0033] Similarly, the semantic representations of all characters in sentence A are multiplied by the corresponding soft alignment weights to obtain the soft alignment representation of each character in sentence B , and and are used as the full information flow;

[0034] Further, after average pooling and concatenation of and , they are successively passed through a feed-forward neural network and layer normalization processing to obtain the full character features of the sentence pair , which is expressed as:

[0035] ;

[0036] where MP, FFN, and LN represent average pooling, feed-forward neural network, and layer normalization processing respectively.

[0037] In one embodiment, the soft alignment representation of the text block pair information containing different characters in the sentence pair is obtained through cross-attention learning, and it is used as different information flows to calculate and obtain different character features of the sentence pair, including:

[0038] Construct a 0-1 matrix to represent the non-correlation between text block pairs. In the matrix , if a certain text block pair is the same, the weights of the entire row and entire column where the text pair is located are set to 0, and the weights of the remaining positions are set to 1;

[0039] According to the similarity matrix of the sentence pair and the matrix construct a matrix representing different detailed information of the sentence pair, which is defined as follows:

[0040] ;

[0041] Adopt the same method as and The same calculation method is based on respectively calculate and obtain the soft alignment representations in sentence A and sentence B that are affected by different information and and use and as different information flows;

[0042] Furthermore, after performing average pooling and concatenation on and and then passing through a feed-forward neural network and layer normalization processing in sequence, different character features of the sentence pair are obtained .

[0043] In one embodiment, the consistency inference unit interacts with the classifier in the basic module to calculate the consistency loss between the global probability of the semantic relationship of the sentence pair output by the classifier and the probabilities of different granularities calculated based on the full character features and different character features, and infers and determines the local probability of the semantic relationship of the sentence pair. Finally, by performing weighted summation on the global probability and the local probability, the final probability of the semantic relationship of the sentence pair is output, including:

[0044] The consistency inference unit interacts with the classifier in the basic module and calculates the probabilities of different granularities of the semantic relationship of the sentence pair based on the full character features and different character features ; and ;

[0045] Subsequently, the consistency loss between the global probability of the semantic relationship of the sentence pair output by the classifier and the predictions of different granularities is calculated through KL divergence , expressed as:

[0046] ;

[0047] If , set the local probability of the semantic relationship of the sentence pair to , and learn the consistency loss . At this time, it means selecting the detailed information emphasized by different information flows to refine the preliminary prediction of the semantic relationship of the sentence pair; if , set to , and . At this time, it means selecting the full information flow to supplement the information missing in the basic module, thereby optimizing the preliminary prediction of the semantic relationship of the sentence pair;

[0048] Finally, by performing weighted summation on the global probability and the local probability Perform weighted summation to output the final probability of the semantic relationship of the sentence pair , expressed as:

[0049] ;

[0050] where, and represent the trade-off parameters.

[0051] In one embodiment, the comprehensive loss function of the fine-grained information consistent inference network is expressed as:

[0052] ;

[0053] where, 、 and respectively represent the trade-off parameters of the losses of each component; is the cross-entropy loss of the global probability of the semantic relationship of the sentence pair obtained by the preliminary prediction of the basic module ; is the cross-entropy loss of the final probability of the semantic relationship of the sentence pair obtained by the final prediction of the inference module ; is the consistency loss; is the given training instance, consisting of the sentence pair and ; represents 's true label.

[0054] A text matching device based on a fine-grained information consistent inference network, the device includes:

[0055] A network construction module for constructing a fine-grained information consistent inference network composed of a basic module and an inference module; wherein, the basic module encodes and generates the global features and local features of the input sentence pair by using the character-level interaction between sentences, and calculates the preliminary prediction result of the semantic relationship of the sentence pair based on the global features; the inference module generates the full information flow containing the text block pair information of all characters in the sentence pair and the different information flow containing the text block pair information of different characters in the sentence pair through cross-attention learning based on the local features, and automatically selects the most important fine-grained features from the full information flow and the different information flows, and improves the preliminary prediction result of the semantic relationship of the sentence pair according to the consistency inference between the global features and the fine-grained features to obtain the final prediction result of the semantic relationship of the sentence pair;

[0056] A training and task execution module for inputting a text matching dataset into a fine-grained information consistent inference network for training, with the training objective of minimizing the comprehensive loss of the preliminary prediction, final prediction, and consistent inference of the semantic relationship between sentence pairs until a trained fine-grained information consistent inference network is obtained to execute the text matching task.

[0057] The above text matching method and device based on a fine-grained information consistent inference network implement a preliminary prediction of the semantic relationship between sentence pairs based on the basic module in the network, and further introduce fine-grained information and consistent inference based on the inference module in the network to optimize the preliminary prediction of the basic module. Different from the existing methods that only consider the overall features of sentence pairs, the designed inference module in this application designs an extensible cross-attention learning unit, enabling it to utilize the soft alignment interaction between two sentences to generate fine-grained information. In addition, the inference module also automatically selects more important fine-grained information through consistent constraint reasoning to further improve the accuracy of the prediction of the semantic relationship between sentence pairs, thereby improving the correct rate of text matching. BRIEF DESCRIPTION OF THE DRAWINGS

[0058] Figure 1 It is a schematic flowchart of a text matching method based on a fine-grained information consistent inference network in an embodiment;

[0059] Figure 2 It is a schematic diagram of the overall architecture of a fine-grained information consistent inference network in an embodiment;

[0060] Figure 3 It is a schematic flowchart of obtaining the full-character features of a sentence pair through cross-attention learning calculation in an embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0061] In order to make the objectives, technical solutions, and advantages of this application clearer, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not used to limit this application.

[0062] In one embodiment, as Figure 1 shown, a text matching method based on a fine-grained information consistent inference network is provided, including the following steps:

[0063] Step S1, constructing a fine-grained information consistent inference network composed of a basic module and an inference module.

[0064] The overall architecture of the fine-grained information consistent inference network is as Figure 2 shown, where the basic module includes an encoder and a classifier based on Transformer.

[0065] In modern text matching methods, significant progress has been made in sentence pair modeling methods based on pre-trained language models. There are mainly two common architectures: bi-encoders and cross-encoders.

[0066] The bi-encoder architecture first encodes the two sentences separately and then calculates their similarity in the embedding space. Its advantage is high computational efficiency, especially suitable for large-scale data. Common bi-encoder models include: (1) SBERT (Sentence-BERT): Fine-tunes BERT through a siamese / triplet network structure for efficient calculation of sentence pair similarity; (2) Enhanced SBERT: Proposes a simple but efficient data augmentation strategy; (3) PromptBERT (Prompt-based BERT): Proposes a sentence embedding method based on prompt learning.

[0067] The cross-encoder architecture processes the sentence pair in one encoding process and directly outputs a classification score to predict the relationship between the sentence pair. The cross-encoder can model the lexical-level interaction between sentences and usually obtains more accurate results. Common cross-encoder methods include: (1) BERT: Utilizes the Transformer encoder to perform word-level interaction on sentences and outputs a global feature for classification; (2) To improve the semantic representation of the pre-trained language model, SemBERT (Semantic BERT) is proposed, which introduces explicit semantic role annotation into the model through a lightweight fine-tuning method; (3) UERBERT (Universal Pre-trained Encoder BERT): Directly injects synonym knowledge into the attention mechanism to improve semantic text similarity; (4) SyntaxBERT (Syntactic BERT): Integrates syntactic tree information into the Transformer model in a plug-and-play manner to improve syntactic understanding; (5) DAFA (Domain Adaptive Fine-tuning Approach): Efficiently incorporates dependency structure prior information into the pre-trained model. Since the cross-encoder can directly model the interaction between sentences, it is usually superior to the bi-encoder in performance.

[0068] Based on this, the Transformer-based encoder in the above basic module selects the cross-encoder BERT, utilizes the character-level interaction between sentences of the cross-encoder to encode and generate the global and local features of the input sentence pair, so as to improve the lexical or logical alignment degree of the corresponding detailed information between the two sentences. The classifier in the above basic module is used to calculate the preliminary prediction result of the semantic relationship between the sentence pair according to the global feature.

[0069] Specifically, the data processing process of the Transformer-based encoder includes:

[0070] Given a pair of sentences and , first insert special characters [CLS] and [SEP], and splice them into the encoder input , expressed as:

[0071] (1)

[0072] Among them, the subscripts and respectively represent the number of characters in sentence A and sentence B, and respectively represent the th character in sentence A and the th character in sentence B; [CLS] is called the classification token, which is used to represent the start of the input sequence; [SEP] is called the separator token, which is used to separate two sentences or represent the end of a single sentence.

[0073] Then, select the pre-trained language model BERT as the encoder , and encode to generate the context representation of the sentence pair , the expression is:

[0074] (2)

[0075] Among them, the context representation contains the semantic representation of the special character [CLS] and the semantic representations of all characters in the sentence pair; among them, the semantic representation of the special character [CLS] is regarded as the global feature of the sentence pair, which is further input into the classifier for the preliminary prediction of the semantic relationship of the sentence pair; the semantic representations of all characters in sentence A and the semantic representations and of all characters in sentence B are regarded as the local features of the sentence pair, which are further input into the inference module to generate and automatically select more important fine-grained features to optimize the preliminary prediction result of the semantic relationship of the sentence pair; and respectively represent the semantic representation of the

[0076] Specifically, the data processing process of the classifier includes: inputting the global feature into the classifier, and making a preliminary prediction of the semantic relationship of the sentence pair according to the classifier to obtain the global probability of the semantic relationship of the sentence pair, expressed as:

[0077] (3)

[0078] Among them, is the normalized exponential function, is a learnable parameter, and the superscript T represents transpose.

[0079] Furthermore, in order to overcome the limitation of the basic module in ignoring important details, the present application introduces an inference module in the fine-grained information consistent inference network to improve the prediction of the semantic relationship of sentence pairs by using the consistency inference between global information and more important fine-grained information. Specifically, as Figure 2 shown, the inference module includes a cross-attention learning unit and a consistency inference unit.

[0080] The cross-attention learning unit obtains the soft alignment representation of the text block pair information containing all characters in the sentence pair by using the soft alignment interaction between two sentences at the sentence level. On the one hand, through cross-attention learning, it obtains the soft alignment representation of the text block pair information containing all characters in the sentence pair and uses it as the full information flow to calculate the full character features of the sentence pair; on the other hand, it also obtains the soft alignment representation of the text block pair information containing different characters in the sentence pair through cross-attention learning and uses it as the different information flow to calculate the different character features of the sentence pair.

[0081] Specifically, as Figure 3 shown, obtaining the soft alignment representation of the text block pair information containing all characters in the sentence pair by cross-attention learning and using it as the full information flow to calculate the full character features of the sentence pair includes:

[0082] First, according to the semantic representations of all characters in sentence A in the local features and the semantic representations of all characters in sentence B , calculate the similarity matrix of the sentence pair. The element in the th row and th column of the similarity matrix

[0083] is expressed as:

[0084] where represents the similarity function, which is calculated by vector dot product; and respectively represent the semantic representation of the th character in sentence A and the semantic representation of the j th character in sentence B.

[0085] Next, normalize the similarity matrix to obtain the soft alignment weights of each character in one sentence to all characters in the other sentence; among them, the soft alignment weights of each character in sentence A to all characters in sentence B are calculated as follows:

[0086] (5)

[0087] Among them, is the normalized exponential function; Each element in represents the matching degree of the corresponding character pair.

[0088] Then, the semantic representations of all characters in sentence B are multiplied by the corresponding soft alignment weights to obtain the soft alignment representation of each character in sentence A , which is expressed as:

[0089] (6)

[0090] Similarly, following similar steps as above, the semantic representations of all characters in sentence A are multiplied by the corresponding soft alignment weights to obtain the soft alignment representation of each character in sentence B , and and are used as the full information flow. and represent the soft alignment representation of one sentence relative to another sentence, covering the influence of all detailed information.

[0091] Further, after average pooling and concatenation of and , they are successively passed through a feed-forward neural network and layer normalization processing to obtain the full character features of the sentence pair , which is expressed as:

[0092] (7)

[0093] Among them, MP, FFN, and LN represent average pooling, feed-forward neural network, and layer normalization processing respectively.

[0094] Specifically, through cross-attention learning, a soft alignment representation containing text block pair information of different characters in the sentence pair is obtained, and it is used as different information flows to calculate and obtain different character features of the sentence pair, including:

[0095] First, a 0-1 matrix is constructed to represent the non-correlation between text block pairs. In the matrix , if a certain text block pair is the same, the weights of the entire row and entire column where the text pair is located are set to 0, and the weights of the remaining positions are set to 1.

[0096] Secondly, according to the similarity matrix of the sentence pair and the matrix Construct a matrix representing different detailed information of sentence pairs , is defined as follows:

[0097] (8)

[0098] Then, adopt the same calculation method as shown in formula (5) and formula (6) for and Based on calculate and obtain the soft alignment representations affected by different information in sentence A and sentence B respectively and , and use and as different information flows. In different information flows, focus on those text block pairs that are not literally the same, and ignore those text block pairs that are literally the same. Different text block pairs refer to the text blocks that appear in one sentence but not in the other sentence, indicating the differences between the two.

[0099] Further adopt the calculation formula shown in formula (7), for and perform average pooling and concatenation, and then pass through a feed-forward neural network and layer normalization processing in sequence to obtain different character features of the sentence pair .

[0100] Considering that the relationship prediction from global features and local features is consistent, this application improves the prediction by automatically selecting the most relevant fine-grained features from the full information flow and different information flows. In this selection process, it is necessary to ensure that the probabilities of global features and the selected fine-grained features satisfy the consistency constraint. Specifically, the consistency reasoning unit in the inference module interacts with the classifier in the basic module to calculate the consistency loss between the global probability of the semantic relationship of the sentence pair output by the classifier and the probabilities of different granularities calculated based on the full character features and different character features, and infers and determines the local probability of the semantic relationship of the sentence pair. Finally, by performing weighted summation on the global probability and the local probability, the final probability of the semantic relationship of the sentence pair is output. The data processing process of the consistency reasoning unit includes:

[0101] The consistency reasoning unit interacts with the classifier in the basic module and calculates the probabilities of different granularities of the semantic relationship of the sentence pair based on the full character features and different character features and .

[0102] ​Subsequently, the global probability of the semantic relationship of the sentence pair output by the classifier is calculated through the KL (Kullback-Leibler) divergence The consistency loss between different granularity predictions , expressed as:

[0103] (9)

[0104] If , the local probability of the semantic relationship of the sentence pair is set to , and the learning consistency loss . This is because the full information flow pays more attention to the same text pair. At this time, it means selecting the detailed information emphasized by different information flows to refine the preliminary prediction of the semantic relationship of the sentence pair; if , is set to , and . At this time, it means selecting the full information flow to supplement the information missing in the basic module, and then optimizing the preliminary prediction of the semantic relationship of the sentence pair.

[0105] Finally, by performing a weighted sum of the global probability and the local probability , the final probability of the semantic relationship of the sentence pair is output, expressed as:

[0106] (10)

[0107] Among them, and represent the trade-off parameters.

[0108] Step S2: Input the text matching dataset into the fine-grained information consistency inference network for training. The training objective is to minimize the comprehensive loss of the preliminary prediction, final prediction, and consistency inference of the semantic relationship of the sentence pair until the trained fine-grained information consistency inference network is obtained to perform the text matching task.

[0109] Specifically, the comprehensive loss function of the fine-grained information consistency inference network is expressed as:

[0110] (11)

[0111] Among them, 、 and respectively represent the trade-off parameters of the losses of each component; is the cross-entropy loss of the global probability of the semantic relationship of the sentence pair obtained by the preliminary prediction of the basic module; The final probability of the semantic relationship of the sentence pair obtained by the inference module for final prediction is the cross-entropy loss; is the consistency loss; is a given training instance, consisting of the sentence pair and ; denotes the true label of

[0112] In summary, the method proposed in this application constructs a fine-grained information consistent inference network containing a basic module and an inference module. Based on the basic module, a preliminary prediction of the semantic relationship between sentence pairs is realized. Based on the inference module, fine-grained information and consistency inference are further introduced to optimize the preliminary prediction of the basic module. Moreover, an extensible cross-attention learning unit is designed in the inference module, enabling it to utilize the soft alignment interaction between two sentences to generate fine-grained information. In addition, in the inference module, more important fine-grained information is automatically selected through consistency constraint inference to further improve the accuracy of the semantic relationship prediction between sentence pairs, thereby improving the accuracy of text matching.

[0113] Furthermore, comparative empirical experiments on six public datasets of GLUE (General Language Understanding Evaluation Benchmark) are conducted to verify the effectiveness of the fine-grained information consistent inference network (hereinafter referred to as ConInf) constructed by the method proposed in this application. The six public datasets are: MRPC (Microsoft Research Paraphrase Corpus), QQP (Quora Question Pairs), STS-B, MNLI-m / mm (Multi-Genre Natural Language Inference (matched / mismatched) dataset), QNLI (Question Natural Language Inference dataset), and RTE (Recognizing Textual Entailment dataset). The performance evaluation parameters include the F1 value of sentence similarity and sentence inference, Pearson correlation coefficient (Pearson), and accuracy (Accuracy, Acc).

[0114] Table 1 Test results of different models on six datasets of GLUE under the BERT-base (basic version of BERT) configuration:

[0115]

[0116] Table 2 Test results of different models on six datasets of GLUE under the BERT-large (high-performance version of BERT) configuration:

[0117]

[0118] Table 1 shows the test results of different models on six datasets of GLUE under the BERT-base configuration. The experimental models include UERBERT under the BERT-base configuration base , SemBERT base , SyntaxBERT base , DAFA base and ConInf base . Table 2 shows the test results of different models on six datasets of GLUE under the BERT-large configuration. The experimental models include BERT under the BERT-large configuration large , Prompt-tuning + BERT large (prompt-tuned BERT), SyntaxBERT large , DAFA large and ConInf large . The values with superscript “+” in the last column of Table 1 and Table 2 represent the average scores of performance parameters on the other five datasets except the QQP dataset. As shown in Table 1, under the BERT-base configuration, the fine-grained information consistent inference network ConInf base constructed in this application is significantly superior to all other methods in terms of the average score (83.63% v.s. 82.68% and 85.55% v.s. 85.18%), which indicates the effectiveness of the text matching method based on the fine-grained information consistent inference network proposed in this application in text matching. As shown in Table 2, under the BERT-large configuration, ConInf large is significantly superior to BERT large and Prompt-tuning + BERT large (85.60% v.s. 83.39% and 85.60% v.s. 82.18%), and is consistently superior to these basic pre-trained language model methods on six different datasets, indicating the effectiveness of the fine-grained information. The experimental results show that the fine-grained information consistency inference network constructed in this application has achieved a significant improvement on six common datasets of GLUE, reaching an average score of 83.63% under the BERT-base configuration (+0.95 absolute improvement) and an average score of 85.60% under the BERT-large configuration (+0.63% absolute improvement), demonstrating its effectiveness and advantages in text matching tasks.

[0119] In one embodiment, a text matching device based on a fine-grained information consistent inference network is provided, including:

[0120] A network construction module for constructing a fine-grained information-consistent inference network composed of a basic module and an inference module. Among them, the basic module encodes and generates the global features and local features of the input sentence pair by utilizing character-level inter-sentence interaction, and calculates the preliminary prediction result of the semantic relationship of the sentence pair based on the global features. The inference module generates the full information flow containing the text block pair information of all characters in the sentence pair and the different information flows containing the text block pair information of different characters in the sentence pair through cross-attention learning based on the local features by utilizing sentence-level inter-sentence interaction, and automatically selects the most important fine-grained features from the full information flow and the different information flows, and improves the preliminary prediction result of the semantic relationship of the sentence pair according to the consistency inference between the global features and the fine-grained features to obtain the final prediction result of the semantic relationship of the sentence pair.

[0121] A training and task execution module for inputting a text matching data set into the fine-grained information-consistent inference network for training. The training objective is to minimize the comprehensive loss of the preliminary prediction, final prediction and consistency inference of the semantic relationship of the sentence pair until the trained fine-grained information-consistent inference network is obtained to execute the text matching task.

[0122] For the specific limitations of the text matching device based on the fine-grained information-consistent inference network, reference can be made to the limitations of the text matching method based on the fine-grained information-consistent inference network in the above text, which will not be elaborated here. Each module in the above text matching device based on the fine-grained information-consistent inference network can be implemented in whole or in part by software, hardware and their combination. The above modules can be embedded in the processor in the computer device in the form of hardware or independent of it, or stored in the memory in the computer device in the form of software, so as to facilitate the processor to call and execute the operations corresponding to the above modules.

[0123] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.

[0124] The above-described embodiments only represent several implementation manners of the present application, and their descriptions are relatively specific and detailed, but they should not be construed as limiting the scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several deformations and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the appended claims.

Claims

1. A text matching method based on a fine-grained information consistent reasoning network, characterized in that The method includes: Constructing a fine-grained information-consistent inference network composed of a basic module and an inference module; wherein, the basic module encodes and generates global features and local features of the input sentence pair by utilizing character-level inter-sentence interaction, and calculates a preliminary prediction result of the semantic relationship of the sentence pair based on the global features; the inference module generates a full information flow containing text block pair information of all characters in the sentence pair and a different information flow containing text block pair information of different characters in the sentence pair through cross-attention learning based on sentence-level inter-sentence interaction, and automatically selects the most important fine-grained features from the full information flow and the different information flows, and improves the preliminary prediction result of the semantic relationship of the sentence pair according to the consistency inference between the global features and the fine-grained features to obtain the final prediction result of the semantic relationship of the sentence pair; Inputting a text matching data set into the fine-grained information-consistent inference network for training, and the training objective is to minimize the comprehensive loss of the preliminary prediction, the final prediction and the consistency inference of the semantic relationship of the sentence pair until a trained fine-grained information-consistent inference network is obtained to perform the text matching task.

2. The method according to claim 1, characterized in that, The basic module includes a Transformer-based encoder and a classifier; wherein, the Transformer-based encoder selects a cross-encoder, and encodes and generates global features and local features of the input sentence pair by utilizing the character-level inter-sentence interaction of the cross-encoder, and the cross-encoder is BERT; the classifier is used to calculate a preliminary prediction result of the semantic relationship of the sentence pair according to the global features.

3. The method according to claim 2, wherein Encoding and generating global features and local features of the input sentence pair by utilizing the character-level inter-sentence interaction of the cross-encoder, including: Given a pair of sentences and , first insert special characters [CLS] and [SEP], and concatenate them into the encoder input , which is expressed as: ; Among them, the subscripts and represent the number of characters in sentence A and sentence B respectively, and represent the -th character in sentence A and the -th character in sentence B respectively; [CLS] is called the classification token, which is used to represent the start of the input sequence; [SEP] is called the separator token, which is used to separate two sentences or represent the end of a single sentence; Then, select the pre-trained language model BERT as the encoder , and encode to generate the context representation of the sentence pair , and the expression is: ; Among them, the context representation includes the semantic representation of the special character [CLS] and the semantic representations of all characters in the sentence pair; among them, the semantic representation of the special character [CLS] is regarded as the global feature of the sentence pair and is used to further input into the classifier for the preliminary prediction of the semantic relationship of the sentence pair; the semantic representations of all characters in sentence A and the semantic representations of all characters in sentence B are regarded as the local features of the sentence pair and are used to further input into the inference module to generate and automatically select more important fine-grained features to optimize the preliminary prediction result of the semantic relationship of the sentence pair; and respectively represent the semantic representation of the th character in sentence A and the semantic representation of the th character in sentence B.

4. The method according to claim 3, wherein The classifier calculates a preliminary prediction result of the semantic relationship of the sentence pair based on the global features, including: Input global features into a classifier, and perform a preliminary prediction of the semantic relationship of the sentence pair according to the classifier to obtain the global probability of the semantic relationship of the sentence pair , which is expressed as: ; Among them, is the normalized exponential function, is a learnable parameter, and the superscript T represents the transpose.

5. The method according to claim 3, wherein The inference module includes a cross-attention learning unit and a consistency inference unit; Wherein, the cross-attention learning unit obtains a soft-alignment representation containing text block pair information of all characters in the sentence pair through cross-attention learning by utilizing sentence-level learning of the soft-alignment interaction between two sentences, and calculates the full-character features of the sentence pair by using it as the full information flow; on the other hand, also obtains a soft-alignment representation containing text block pair information of different characters in the sentence pair through cross-attention learning, and calculates the different-character features of the sentence pair by using it as the different information flow; The consistency inference unit calculates the consistency loss between the global probability of the semantic relationship of the sentence pair output by the classifier and the probabilities of different granularities calculated based on the full-character features and the different-character features through interaction with the classifier in the basic module, infers and determines the local probability of the semantic relationship of the sentence pair, and finally outputs the final probability of the semantic relationship of the sentence pair by performing weighted summation on the global probability and the local probability.

6. The method according to claim 5, wherein Obtaining a soft-alignment representation containing text block pair information of all characters in the sentence pair through cross-attention learning, and calculating the full-character features of the sentence pair by using it as the full information flow, including: Based on the semantic representations of all characters in sentence A among the local features and the semantic representations of all characters in sentence B , calculate the similarity matrix of the sentence pair . The similarity matrix . The element in the -th row and -th column of the similarity matrix is expressed as: ; Among them, represents the similarity function, and the vector dot product calculation is adopted; and respectively represent the semantic representations of the -th character in sentence A and the -th character in sentence B; For the similarity matrix perform normalization processing to obtain the soft alignment weights of each character in one sentence to all characters in another sentence; among them, the soft alignment weights of each character in sentence A to all characters in sentence B are calculated as follows: ; Among them, is the normalized exponential function; each element in represents the matching degree of the corresponding character pair; The semantic representations of all characters in sentence B are multiplied by the corresponding soft alignment weights to obtain the soft alignment representations of each character in sentence A , which is expressed as: ; Similarly, multiply the semantic representations of all characters in sentence A by the corresponding soft alignment weights to obtain the soft alignment representations of each character in sentence B , and use and as the full information flow; Further, after performing average pooling on and and concatenating them, they are sequentially passed through a feed-forward neural network and layer normalization processing to obtain the full-character features of the sentence pair , which is expressed as: ; Among them, MP, FFN, and LN represent average pooling, feed-forward neural network, and layer normalization processing respectively.

7. The method according to claim 6, wherein Through cross-attention learning, a soft alignment representation of text block pairs containing different characters in the sentence pair is obtained, and it is used as different information flows to calculate different character features of the sentence pair, including: Construct a 0-1 matrix , which is used to represent the non-correlation between text block pairs. In the matrix , if a certain text block pair is the same, the weights of the entire row and entire column where the text pair is located are set to 0, and the weights of the remaining positions are set to 1; According to the similarity matrix of sentence pairs and the matrix construct a matrix representing different detailed information of sentence pairs , is defined as follows: ; Adopt the same calculation method as and , and respectively calculate and obtain the soft alignment representations in sentence A and sentence B affected by different information and , and use and as different information flows; Further, after performing average pooling on and and concatenating them, they are successively passed through a feedforward neural network and layer normalization processing to obtain different character features of the sentence pair .

8. The method according to claim 5, characterized in that, The consistency reasoning unit calculates the consistency loss between the global probability of the semantic relationship of the sentence pair output by the classifier and the probabilities of different granularities calculated based on the full character features and different character features by interacting with the classifier in the basic module, reasons to determine the local probability of the semantic relationship of the sentence pair, and finally outputs the final probability of the semantic relationship of the sentence pair by performing a weighted sum of the global probability and the local probability, including: The consistency inference unit calculates probabilities of semantic relations of sentence pairs at different granularities based on the full-character features by interacting with the classifier in the basic module and different character features and ;​ Subsequently, the global probability of the semantic relationship of the sentence pairs output by the classifier is calculated through the KL divergence and the consistency loss between different granularity predictions , which is expressed as: ; If , set the local probability of the semantic relationship of the sentence pair to , and the learning consistency loss . At this time, it means selecting the detailed information emphasized by different information flows to refine the preliminary prediction of the semantic relationship of the sentence pair; If , set to , and . At this time, it means selecting the full information flow to supplement the information missing in the basic module, thereby optimizing the preliminary prediction of the semantic relationship of the sentence pair; Finally, by performing a weighted sum of the global probability and the local probability the final probability of the semantic relationship of the sentence pair is output , expressed as: ; Among them, and represent trade-off parameters.

9. The method according to claim 1, characterized in that, The comprehensive loss function of the fine-grained information consistent reasoning network is expressed as: ; Among them, 、 and respectively represent the trade-off parameters of the losses of each component; is the cross-entropy loss of the global probability of the semantic relationship of the sentence pair obtained by the preliminary prediction of the basic module ; is the cross-entropy loss of the final probability of the semantic relationship of the sentence pair obtained by the final prediction of the inference module ; is the consistency loss; is the given training instance, consisting of the sentence pair and ; represents 's true label.

10. A text matching device based on a fine-grained information consistent reasoning network, characterized in that, The device includes: A network construction module, configured to construct a fine-grained information consistency reasoning network composed of a basic module and a reasoning module; wherein, the basic module encodes and generates the global features and local features of the input sentence pair by utilizing the inter-sentence interaction at the character level, and calculates a preliminary prediction result of the semantic relationship of the sentence pair based on the global features; the reasoning module generates a full information flow containing text block pair information of all characters in the sentence pair and different information flows containing text block pair information of different characters in the sentence pair based on the cross-attention learning of the local features through utilizing the inter-sentence interaction at the sentence level, and automatically selects the most important fine-grained features from the full information flow and different information flows, and improves the preliminary prediction result of the semantic relationship of the sentence pair according to the consistency reasoning between the global features and the fine-grained features to obtain the final prediction result of the semantic relationship of the sentence pair; A training and task execution module, configured to input a text matching data set into the fine-grained information consistency reasoning network for training, and the training objective is to minimize the comprehensive loss of the preliminary prediction, final prediction, and consistency reasoning of the semantic relationship of the sentence pair until the trained fine-grained information consistency reasoning network is obtained to perform the text matching task.

Citation Information

Patent Citations

  • A text implication relation recognition method based on multi-granularity information fusion

    CN109299262A

  • Systems and methods for processing machine learning language model classification outputs via text block masking

    US20230334887A1