Sentence semantic matching method and device for sentence subject interaction
By constructing a sentence semantic matching model, extracting and optimizing the distillation theme characteristics and optimal theme characteristics, the problem of noise interference and insufficient interaction between sentence themes in the existing model is solved, and higher matching accuracy and robustness are achieved.
Patent Information
- Application Number
- CN202510268340.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-07
- Publication Date
- 2025-07-11
AI Technical Summary
The existing semantic matching model fails to effectively remove noise information and fails to fully consider the interaction of sentence theme information, resulting in insufficient matching accuracy.
A sentence semantic matching model is constructed, distilled theme features and optimal theme features are extracted through feature aggregation network, and loss calculation and optimization are performed in combination with classification layers. The loss value optimization model is used to eliminate noise features and enhance sentence theme features interaction.
It improves the accuracy and robustness of sentence semantic matching, enhances the model's adaptability to different semantic matching scenarios, and significantly improves the matching accuracy.
Smart Images

Figure CN120297282A_ABST
Abstract
Description
Technical Field
[0001] The present invention mainly relates to the technical field of natural language processing, and particularly relates to a sentence semantic matching method and device for sentence theme interaction. Background Art
[0002] Text matching is a fundamental research task in the field of natural language processing (NLP). Its core goal is to evaluate and quantify the semantic similarity or relevance between two texts. As a key enabling technology for many advanced language understanding tasks (such as paraphrase recognition, natural language inference, and question answering systems), an effective text matching model plays a crucial role in the development of natural language processing technology.
[0003] However, natural language is complex. In particular, redundant information interferes with the understanding of sentence themes, posing significant challenges to sentence semantic matching. The existing research on semantic matching models to address this problem includes three approaches: one is to use the attention mechanism to obtain the theme information of sentences, calculate the attention scores of different sentence theme information, and then comprehensively judge the similarity of sentences; the second is to train the model by introducing noisy data to improve its noise resistance ability and the ability to grasp sentence themes; the third is to enhance the model's understanding of sentence themes with the help of external knowledge bases. Nevertheless, most of the research on text matching tasks by semantic matching models only focuses on the extraction of key theme information, and fails to simultaneously consider the elimination of noise information and the interaction of theme information. Summary of the Invention
[0004] The technical problem to be solved by the present invention is to provide a sentence semantic matching method and device for sentence theme interaction in view of the deficiencies of the prior art.
[0005] The technical solution of the present invention to solve the above technical problems is as follows:
[0006] A sentence semantic matching method for sentence theme interaction includes the following steps:
[0007] Construct a sentence semantic matching model, where the sentence semantic matching model includes a feature aggregation network, a theme extraction layer, and a classification layer;
[0008] Input the sentence pair to be matched into the sentence semantic matching model for training, including: the feature aggregation network extracts the theme of the sentence pair to be matched to obtain distilled theme features, the theme extraction layer extracts the keywords of the sentence pair to be matched to obtain a theme information sequence, the feature aggregation network extracts the theme of the theme information sequence to obtain optimal theme features, and the classification layer classifies according to the distilled theme features and the optimal theme features to obtain a matching label, and completes the preliminary training of the sentence semantic matching model;
[0009] Calculate the losses for the distilled main idea features and the optimal main idea features respectively to obtain loss values;
[0010] Optimize the initially trained sentence semantic matching model according to the loss values to obtain an optimized sentence semantic matching model.
[0011] Another technical solution of the present invention to solve the above technical problems is as follows:
[0012] A sentence semantic matching device for sentence main idea interaction, comprising:
[0013] A construction module for constructing a sentence semantic matching model, the sentence semantic matching model including a feature aggregation network, a main idea extraction layer and a classification layer;
[0014] A training module for inputting a pair of sentences to be matched into the sentence semantic matching model for training, including: the feature aggregation network extracts the main idea of the pair of sentences to be matched to obtain distilled main idea features, the main idea extraction layer extracts the keywords of the pair of sentences to be matched to obtain a main idea information sequence, the feature aggregation network extracts the main idea of the main idea information sequence to obtain optimal main idea features, and the classification layer classifies according to the distilled main idea features and the optimal main idea features to obtain a matching label, and completes the initial training of the sentence semantic matching model;
[0015] An evaluation module for calculating the losses for the distilled main idea features and the optimal main idea features respectively to obtain loss values;
[0016] An optimization module for optimizing the initially trained sentence semantic matching model according to the loss values to obtain an optimized sentence semantic matching model.
[0017] The beneficial effects of the present invention are as follows: During the process of training a sentence semantic matching model, by extracting distilled main idea features according to the context features of a pair of sentences to be matched, and extracting the keywords of the pair of sentences to be matched as semantic main ideas to obtain optimal main idea features, combining the distilled main idea features and the optimal main idea features can more accurately identify the semantic relationship between sentences, so as to improve the accuracy of sentence semantic matching. Using loss calculation to evaluate the performance of the model, optimizing the model according to the loss values, and improving the model to adapt to different semantic matching scenarios. Description of the Drawings
[0018] Figure 1 It is a flowchart of the sentence semantic matching method for sentence main idea interaction provided by an embodiment of the present invention;
[0019] Figure 2 It is a block diagram of the modules of the sentence semantic matching method device for sentence main idea interaction provided by an embodiment of the present invention. Detailed Embodiments
[0020] The principles and features of the present invention will be described below in conjunction with the accompanying drawings. The examples given are only used to explain the present invention and are not intended to limit the scope of the present invention.
[0021] As Figure 1 shown, a sentence semantic matching method for sentence theme interaction provided by an embodiment of the present invention includes the following steps:
[0022] Construct a sentence semantic matching model, where the sentence semantic matching model includes a feature aggregation network, a theme extraction layer, and a classification layer;
[0023] Input the sentence pair to be matched into the sentence semantic matching model for training, including: the feature aggregation network extracts the theme of the sentence pair to be matched to obtain distilled theme features, the theme extraction layer extracts the keywords of the sentence pair to be matched to obtain a theme information sequence, the feature aggregation network extracts the theme of the theme information sequence to obtain optimal theme features, and the classification layer classifies according to the distilled theme features and the optimal theme features to obtain a matching label, and complete the preliminary training of the sentence semantic matching model;
[0024] Calculate the loss for the distilled theme features and the optimal theme features respectively to obtain loss values;
[0025] Optimize the sentence semantic matching model after preliminary training according to the loss value to obtain an optimized sentence semantic matching model.
[0026] Predict the target sentence pair to be matched through the optimized sentence semantic matching model to obtain an optimal matching label.
[0027] It should be understood that inspired by the text matching model based on keyword and intent decomposition and the semantic matching method for zero-shot relation extraction, the theme extraction layer extracts the keywords of the sentence pair to be matched through an unsupervised method (such as the TextRank algorithm) to obtain a theme information sequence.
[0028] In the embodiment of the present invention, during the model training process, by deeply fusing the extracted context features of the sentence with the theme features, the semantic relationship between sentences is comprehensively captured, avoiding the limitations brought by relying only on keywords, so as to more accurately judge the similarity of sentences in terms of overall semantics. Through feature aggregation and theme extraction, the model can effectively capture key information and interact with the theme features, which not only enhances the expression ability of the sentence theme features for semantic relationships, but also further improves the matching accuracy. The performance of the model is evaluated by loss calculation, and the model is optimized according to the loss value to improve the parameters of the model to adapt to different semantic matching scenarios.
[0029] Preferably, the feature aggregation network includes an encoder, a feature aggregator, and a feature distillation layer;
[0030] The feature aggregation network extracts the main idea of the sentence pair to be matched, obtaining a distilled main idea feature, including:
[0031] The encoder extracts features from the sentence pair to be matched, obtaining an encoded representation and an aggregated sequence representation. The feature aggregator performs an aggregation calculation on the encoded representation to obtain weakly relevant features. The feature distillation layer performs a projection process on the aggregated sequence representation and the weakly relevant features to obtain a distilled main idea feature.
[0032] Specifically, the encoder includes multiple hidden layers, and the sentence pair to be matched is subjected to hidden feature extraction or semantic feature extraction layer by layer through each hidden layer. After being processed by the last hidden layer, the encoded representation and the aggregated sequence representation are output. The feature extraction expression is:
[0033] [h cls ; H a,b = PLM([w cls ; X a,b ),
[0034] where h cls is the aggregated sequence representation, H a,b is the encoded representation, is the i-th encoded representation of the sentence sequence, w cls is a special sequence marker preset at the forefront of the sentence pair to be matched, X a,b is the sentence to be matched (i.e., the sentence sequence), and PLM(·) is the encoder, that is, a pre-trained language model PLM is embedded in the encoder.
[0035] A trainable query vector q is used to perform weighted aggregation on the encoded representation to filter out features weakly relevant to the matching task. The dot product calculation is performed on each encoded representation of the sentence sequence through the preset query vector to obtain multiple weakly relevant feature weights. The first dot product expression is:
[0036] ∝ = Softmax(q · H a,b ),
[0037]
[0038] The dot product calculation is performed on each weakly relevant feature weight and the corresponding encoded representation to obtain multiple weakly relevant features. The second dot product expression is:
[0039]
[0040] where ∝ n represents the n-th weakly relevant feature weight value, Represents the nth encoded representation, Represents the features weakly related to the semantic matching relationship after summarizing the ith encoded representation (i.e., weakly related features), q represents the query vector, and Softmax(·) represents the normalized exponential function.
[0041] The feature distillation layer maps the aggregated sequence representation of the sentence to the orthogonal complement space of the weakly related features . By projecting onto the space defined by , the noise feature components are separated, thereby optimizing the model's ability to identify key features for the matching task and improving the matching accuracy. The expression for calculating the noise features by projection is:
[0042]
[0043] Effectively eliminate the noise features in the aggregated sequence representation through projection transformation technology Thus, the final hidden layer aggregated sequence representation of the sentence pair to be matched after refinement (i.e., distilled main idea features) is obtained. The expression for calculating the distilled main idea features by projection is:
[0044]
[0045] where is the ith noise feature, is the ith distilled main idea feature, Proj(·,·) is the projection operator, and the calculation formula of the projection operator is:
[0046]
[0047] For example, the calculation of the noise features is:
[0048]
[0049] In the embodiments of the present invention, the noise features in the sentence pair to be matched are jointly searched through dot product attention, the SoftMax function, and the projection theorem. The projection theorem is used to identify and discard the noise features in the aggregated sequence representation of the hidden layer information, thereby obtaining refined context features, reducing the influence of interference factors on the matching result, and improving the matching accuracy. Through projection transformation, the noise features are successfully removed, and then a more refined aggregated sequence representation of the sentence pair to be matched output in the final hidden layer is obtained, that is, the distilled main idea features. The application of this technical means significantly reduces the potential interference of noise information on the matching accuracy, and at the same time enhances the performance of the model when performing the sentence matching task.
[0050] Preferably, the feature aggregation network extracts the main idea of the main idea information sequence to obtain the optimal main idea features, including:
[0051] The encoder extracts features from the subject information sequence to obtain a representation of the subject aggregation sequence. The feature distillation layer projects the representation of the subject aggregation sequence and the distilled subject features to obtain the optimal subject features.
[0052] Specifically, the sentence subject features extracted by an unsupervised method (i.e., the subject information sequence constructed by keywords covering the core semantics of the sentence) are input into the encoder. The hidden features of the subject information sequence are extracted one by one through each hidden layer of the encoder to obtain the subject encoding representation. At the same time, semantic features of the subject information sequence are extracted to obtain the representation of the subject aggregation sequence. The subject encoding representation and the representation of the subject aggregation sequence are output after being processed by the last hidden layer. The feature extraction expression is:
[0053]
[0054] Among them, is the representation of the subject aggregation sequence, is the subject encoding representation, is a special sequence marker preset at the forefront of the subject information sequence of the sentence pair to be matched, is the subject information sequence, and PLM(·) is the encoder.
[0055] The feature distillation layer maps the representation of the subject aggregation sequence of the sentence to the orthogonal complement space of the distilled subject features , that is, projects the representation of the subject aggregation sequence and the distilled subject features to perform a projection calculation (i.e., a linear transformation) to obtain the optimal subject features The expression for calculating the optimal subject features by projection is:
[0056]
[0057] Among them, is the i-th optimal subject feature, is the i-th representation of the subject aggregation sequence, is the i-th distilled subject feature, and Proj(·,·) is the projection operator.
[0058] In the embodiments of the present invention, the extracted subject information sequence is input into the encoder to extract semantic features, and the projection technology is used to map the subject features from the initial space to the orthogonal space of the finer features after being processed by the feature distillation layer, which can effectively extract high-quality sentence subject interaction matching features and obtain features with more accurate sentence subject semantics, thereby improving the model's ability to model semantic relationships.
[0059] To obtain high-quality sentence topic interaction features, project the sentence topic features extracted by the unsupervised method into the orthogonal space of the refined features processed by the distillation layer, which not only enhances the expression ability of the sentence topic features for semantic relations but also further improves the matching accuracy.
[0060] Preferably, the classification layer classifies according to the distillation topic features and the optimal topic features to obtain a matching label, including:
[0061] The classification layer fuses the distillation topic features and the optimal topic features to obtain a fused topic feature, and classifies the fused topic feature to obtain a matching label.
[0062] Specifically, the sentence context features (i.e., distillation topic features) and the topic features (i.e., optimal topic features) are fused through a fusion calculation expression to enhance the feature semantics, and the fusion calculation expression is:
[0063]
[0064] The fused topic feature is classified through a classification expression to calculate the matching probability value of the fused topic feature of the sentence pair, and the matching label is divided according to the matching probability value to generate a matching result. The classification expression is:
[0065] y * = argmax y∈Y P(y|v),
[0066] where v i is the i-th fused topic feature, v is the fused topic feature, y * is the matching label, the true label y represents the semantic relationship between the input sentence pairs, and argmax(·) is a function for finding the maximum value index in an array or matrix. For the three-class sentence matching relationship Y = {not matching, partial matching, complete matching}, and for the two-class sentence matching relationship Y = {not matching, complete matching}.
[0067] In the embodiment of the present invention, in the prediction stage, the refined context features and the interacted sentence topic features are fused, and the final prediction is completed through feature enhancement, so as to obtain a more accurate matching result. That is, the semantic matching degree of the sentence pair is classified based on the fused topic feature, and the finally predicted matching label is output to obtain the semantic relationship of the sentence pair.
[0068] Preferably, the loss values are calculated respectively for the distillation topic features and the optimal topic features, including:
[0069] Calculate the semantic matching probability distribution of the distilled main idea features according to the preset distilled matching weight matrix to obtain the distilled main idea probability, and calculate the semantic matching probability distribution of the optimal main idea features according to the preset optimal matching weight matrix to obtain the optimal main idea probability;
[0070] Calculate the classification loss of the distilled main idea probability to obtain the standard classification loss, calculate the classification loss of the optimal main idea features to obtain the sentence main idea classification loss, and calculate the distribution loss of the distilled main idea probability and the optimal main idea probability to obtain the probability distribution loss.
[0071] Specifically, the preset trainable weight matrix includes the distilled matching weight matrix the optimal matching weight matrix the main idea matching weight matrix where K represents the number of labels and H represents the hidden layer size. Calculate the distribution probability of the semantic matching labels of the distilled main idea features of the sentence pair according to the distilled matching weight matrix. The calculation formula for the distilled main idea probability is:
[0072]
[0073] where is the i-th distilled main idea feature of the first sentence and the i-th distilled main idea feature of the second sentence is the matching probability distribution of the matching label y (i.e., the distilled main idea probability), is the distilled main idea feature, is the distilled matching weight matrix, and Softmax(·) is the normalization exponential function.
[0074] Calculate the distribution probability of the semantic matching labels of the optimal main idea features of the sentence pair according to the optimal matching weight matrix. The calculation formula for the optimal main idea probability is:
[0075]
[0076] where is the i-th optimal main idea feature of the first sentence and the i-th optimal main idea feature of the second sentence is the matching probability distribution of the matching label y (i.e., the optimal main idea probability), is the optimal main idea feature, is the optimal matching weight matrix, and Softmax(·) is the normalization exponential function.
[0077] Calculate the distilled main idea probability through the classification loss function to obtain the standard classification loss. The classification loss function L J is:
[0078]
[0079] Calculate according to the theme matching weight matrix through the theme classification loss function For the optimal theme features Perform calculations to obtain the sentence theme classification loss, and the theme classification loss function L G Is:
[0080]
[0081] Calculate the distilled theme probability and the optimal theme probability through the relative entropy loss function to obtain the probability distribution loss. The relative entropy loss function L KL Is:
[0082]
[0083] Among them, L KL Represents the probability distribution loss calculated by the KL loss function (i.e., the relative entropy loss function), and D KL (·||·) represents the KL divergence function, which is used to measure the closeness of two distributions. Represents the distribution of the sentence after context distillation matching (i.e., the distilled theme probability). Represents the distribution of the sentence theme matching (i.e., the optimal theme probability).
[0084] It should be understood that the KL divergence (Kullback-Leibler, i.e., relative entropy) is used to measure the difference between two probability distributions. By minimizing the bidirectional KL divergence, the distance between the sentence after context distillation matching and the sentence theme matching is minimized.
[0085] In the embodiments of the present invention, the optimization strategy based on the loss function value significantly improves the performance and robustness of the model in the text semantic matching task. By minimizing the theme classification loss function, the model can better understand the semantic relationship between sentences and accurately judge their similarity or relevance. The KL divergence is introduced into the loss function to minimize the difference between the refined context feature distribution and the sentence theme distribution, which not only improves the model's learning ability of the relationship between the matching content and the sentence theme, but also significantly improves the accuracy and robustness of the matching.
[0086] Preferably, the feature aggregation network further includes a gradient reverse layer and a relationship classifier; before the step of obtaining the loss value, it further includes calculating the gradient reverse loss for the weakly related features through the cross-entropy function, specifically:
[0087] The gradient reverse layer optimizes the set propagation parameters to obtain gradient parameters.
[0088] The relationship classifier classifies and predicts the weakly related features according to the gradient parameters to obtain the matching probability values of the weakly related features, and calculates the gradient reverse loss through a cross-entropy function.
[0089] Specifically, a Gradient Reverse Layer (GRL) is introduced to screen out the features weakly related to the matching relationship from the semantic representations of the sentence pairs to be matched. By optimizing the gradient parameters of the relationship classifier during the training process (that is, optimizing the relationship classifier of the sentence semantic matching model according to the gradient reverse loss), the query vector is further optimized to better capture the features weakly related to the matching relationship, indirectly improving the performance of the model in the sentence matching task. The matching probability expression is:
[0090]
[0091] The cross-entropy function is:
[0092] L GRL,i = CrossEntropy(y i , prob i ),
[0093] where W and b represent the weight matrix and bias of the relationship classifier respectively, prob i represents the matching probability value of the i-th weakly related feature of the sentence pair to be matched. After being processed by the GRL layer, is input into the relationship classifier for prediction. CrossEntropy(.) represents the cross-entropy function, y i represents the true label of the i-th matching probability of the sentence pair to be matched, and L GRL,i represents the i-th gradient reverse loss of the sentence pair to be matched, which is used to optimize the training process of the query vector.
[0094] In the embodiment of the present invention, a Gradient Reverse Layer (GRL) is introduced. Through the gradient reversal mechanism of backpropagation, the semantic features irrelevant to the matching task are identified and removed, so as to guide the query vector to more accurately focus on the weakly related features.
[0095] As Figure 2 shown, a sentence semantic matching device for sentence theme interaction provided by an embodiment of the present invention includes:
[0096] A construction module for constructing a sentence semantic matching model, where the sentence semantic matching model includes a feature aggregation network, a theme extraction layer, and a classification layer;
[0097] A training module for training the sentence semantic matching model by inputting the sentence pairs to be matched, including: the feature aggregation network extracts the main idea of the sentence pairs to be matched to obtain distilled main idea features, the main idea extraction layer extracts the key words of the sentence pairs to be matched to obtain a main idea information sequence, the feature aggregation network extracts the main idea of the main idea information sequence to obtain optimal main idea features, and the classification layer classifies according to the distilled main idea features and the optimal main idea features to obtain matching labels, and completes the preliminary training of the sentence semantic matching model;
[0098] An evaluation module for calculating losses for the distilled main idea features and the optimal main idea features respectively to obtain loss values;
[0099] An optimization module for optimizing the sentence semantic matching model after preliminary training according to the loss values to obtain an optimized sentence semantic matching model.
[0100] The advantages of the present invention are that the sentence semantic matching model can effectively acquire and eliminate the noise information of the sentence pairs to be matched, and realize the interaction of the sentence main idea information. By refining the sentence pair semantic context, accurately identifying and eliminating noise features, high-quality refined context features are obtained; technical means such as dot product attention, GRL layer, Softmax function and projection theorem are used to effectively avoid the problems of data pollution and inaccurate noise processing existing in the traditional method of relying on noise data to improve the noise resistance ability. A one-way interaction strategy based on the projection theorem is introduced to map the sentence main idea to the orthogonal space of the refined context features. Compared with the traditional bidirectional attention mechanism, the quality of the interaction information is significantly improved and the processing efficiency is increased. In the prediction stage, the refined context features and the sentence main idea are fused to calculate the matching result, and the distance between the two distributions is optimized by KL divergence to make the probability distributions closer.
[0101] For the above sentence semantic matching device with sentence main idea interaction, reference can be made to the implementation content and its beneficial effects described in detail above for a sentence semantic matching method with sentence main idea interaction, which will not be repeated here.
[0102] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variation thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements includes not only those elements, but also other elements not expressly listed, or elements inherent to such process, method, article or device.
[0103] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described devices and modules can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.
[0104] In several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of modules is only a logical function division. In actual implementation, there may be other division methods. For example, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed.
[0105] The modules described as separate components may or may not be physically separated, and the components shown as modules may or may not be physical modules, that is, they may be located in one place or distributed to multiple network modules. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of the embodiments of the present invention.
[0106] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A sentence semantic matching method for sentence theme interaction, characterized in that, It includes the following steps: Construct a sentence semantic matching model, which includes a feature aggregation network, a main idea extraction layer, and a classification layer; Input the sentence pair to be matched into the sentence semantic matching model for training, including: the feature aggregation network extracts the main idea of the sentence pair to be matched to obtain a distilled main idea feature, the main idea extraction layer extracts the keywords of the sentence pair to be matched to obtain a main idea information sequence, the feature aggregation network extracts the main idea of the main idea information sequence to obtain an optimal main idea feature, and the classification layer classifies according to the distilled main idea feature and the optimal main idea feature to obtain a matching label, and completes the preliminary training of the sentence semantic matching model; Calculate the losses of the distilled main idea feature and the optimal main idea feature respectively to obtain loss values; Optimize the sentence semantic matching model after preliminary training according to the loss values to obtain an optimized sentence semantic matching model.
2. The sentence semantic matching method according to claim 1, wherein The feature aggregation network includes an encoder, a feature aggregator, and a feature distillation layer; The feature aggregation network extracts the main idea of the sentence pair to be matched to obtain a distilled main idea feature, including: The encoder extracts features of the sentence pair to be matched to obtain an encoded representation and an aggregated sequence representation, the feature aggregator performs an aggregation calculation on the encoded representation to obtain a weakly related feature, and the feature distillation layer performs a projection process on the aggregated sequence representation and the weakly related feature to obtain a distilled main idea feature.
3. The sentence semantic matching method according to claim 2, wherein The feature aggregation network extracts the main idea of the main idea information sequence to obtain an optimal main idea feature, including: The encoder extracts features of the main idea information sequence to obtain a main idea aggregated sequence representation, and the feature distillation layer performs a projection process on the main idea aggregated sequence representation and the distilled main idea feature to obtain an optimal main idea feature.
4. The sentence semantic matching method according to claim 2, characterized in that The feature aggregation network further includes a gradient reverse layer and a relationship classifier; Before the step of obtaining the loss values, it further includes: The gradient reverse layer optimizes the set propagation parameter to obtain a gradient parameter; The relationship classifier classifies and predicts the weakly related feature according to the gradient parameter to obtain a matching probability value of the weakly related feature, and calculates the gradient reverse loss through a cross-entropy function for the matching probability value.
5. The sentence semantic matching method according to claim 1, characterized in that The classification layer classifies according to the distilled main idea feature and the optimal main idea feature to obtain a matching label, including: The classification layer fuses the distilled main idea feature and the optimal main idea feature to obtain a fused main idea feature, and classifies the fused main idea feature to obtain a matching label.
6. The sentence semantic matching method according to claim 1, characterized in that The step of calculating the losses of the distilled main idea feature and the optimal main idea feature respectively to obtain loss values includes: Calculate the semantic matching probability distribution of the distilled main idea feature according to a preset distilled matching weight matrix to obtain a distilled main idea probability, and calculate the semantic matching probability distribution of the optimal main idea feature according to a preset optimal matching weight matrix to obtain an optimal main idea probability; Calculate the classification loss of the distilled topic probability to obtain the standard classification loss, calculate the classification loss of the optimal topic feature to obtain the sentence topic classification loss, and calculate the distribution loss of the distilled topic probability and the optimal topic probability to obtain the probability distribution loss.
7. The sentence semantic matching method according to claim 6, wherein The calculation of the classification loss of the distilled topic probability to obtain the standard classification loss includes: Calculate the distilled topic probability through the classification loss function to obtain the standard classification loss, and the classification loss function is: Among them, L J is the standard classification loss, is the i-th distilled main idea feature of the first sentence and the i-th distilled main idea feature of the second sentence is the distilled main idea probability for the matching label y.
8. The sentence semantic matching method according to claim 6, characterized in that The calculation of the classification loss of the optimal topic feature to obtain the sentence topic classification loss includes: Calculate the optimal topic feature according to the topic matching weight matrix through the topic classification loss function to obtain the sentence topic classification loss, and the topic classification loss function is: Among them, L G is the sentence theme classification loss, is the optimal theme feature, is the theme matching weight matrix.
9. The sentence semantic matching method according to claim 6, characterized in that The calculation of the distribution loss of the distilled topic probability and the optimal topic probability to obtain the probability distribution loss includes: Calculate the distilled topic probability and the optimal topic probability through the relative entropy loss function to obtain the probability distribution loss, and the relative entropy loss function is: Among them, L KL is the probability distribution loss, D KL (·||·) is the KL divergence function, is the i-th distilled topic feature of the first sentence and the i-th distilled topic feature of the second sentence for the distilled topic probability of the matching label y, is the i-th optimal topic feature of the first sentence and the i-th optimal topic feature of the second sentence for the optimal topic probability of the matching label y.
10. A sentence semantic matching device for sentence main idea interaction, characterized in that, Include: A construction module for constructing a sentence semantic matching model, where the sentence semantic matching model includes a feature aggregation network, a topic extraction layer, and a classification layer; A training module for inputting the sentence pairs to be matched into the sentence semantic matching model for training, including: the feature aggregation network extracts the topics of the sentence pairs to be matched to obtain the distilled topic features, the topic extraction layer extracts the keywords of the sentence pairs to be matched to obtain the topic information sequence, the feature aggregation network extracts the topics of the topic information sequence to obtain the optimal topic features, and the classification layer classifies according to the distilled topic features and the optimal topic features to obtain the matching labels and complete the preliminary training of the sentence semantic matching model; An evaluation module for calculating the losses of the distilled topic features and the optimal topic features respectively to obtain the loss values; An optimization module for optimizing the sentence semantic matching model after preliminary training according to the loss values to obtain the optimized sentence semantic matching model.