Sentence semantic matching method and device for key semantic maintenance

By constructing a sentence semantic matching model, using the encoder, feature aggregator and feature distillation layer to eliminate noise information, the problem of noise interference in the sentence semantic matching model is solved, and more accurate sentence matching and performance improvement of downstream tasks is achieved.

CN120297281APending Publication Date: 2025-07-11GUILIN UNIV OF ELECTRONIC TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510268337.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-07
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

The existing sentence semantic matching model fails to fully consider the matching granularity differences of text content at different levels when processing, resulting in noise interference affecting the matching accuracy.

Method used

Construct a sentence semantic matching model, including an encoder, feature aggregator, feature distillation layer and classification layer, optimize the model through the total loss function, the encoder captures semantic information, the feature aggregator filters noise characteristics, the feature distillation layer eliminates redundant information, and the classification layer generates matching results.

Benefits of technology

It improves the accuracy and effectiveness of sentence matching, can extract more representative features, generate accurate matching results, and improves the model's performance in downstream tasks such as intelligent question-and-answer and information retrieval.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120297281A_ABST
    Figure CN120297281A_ABST
Patent Text Reader

Abstract

The invention provides a sentence semantic matching method and device for key semantic maintenance, and relates to the technical field of natural language processing. The method comprises the steps that a sentence semantic matching model is trained, an encoder extracts coded semantic representation and aggregated semantic representation of a to-be-matched sentence pair, a feature aggregator conducts aggregated calculation on the coded semantic representation to obtain weak correlation features, a feature distillation layer maps the aggregated semantic representation to a space orthogonal to the weak correlation features, and a sentence semantic matching model is obtained; the classification layer classifies the key semantic features to obtain a matching result; and performing loss calculation on the aggregated semantic representation, the weak correlation features and the key semantic features, and optimizing the trained sentence semantic matching model according to the total loss. According to the method, features which are weakly related to a semantic objective are screened out from sentence pairs, redundant noise information in semantic features is removed based on the projection theorem, and semantic matching results of the sentence pairs are generated according to reserved key semantic features so as to train an optimized sentence semantic matching model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention mainly relates to the technical field of natural language processing, and particularly relates to a sentence semantic matching method and device for maintaining key semantics. Background Art

[0002] Sentence semantic matching is a key task in natural language processing, aiming to accurately judge the deep semantic relationship between two texts. Whether in the academic field or in the industrial field, the research on sentence semantic matching has received extensive attention and is of great significance in many practical applications.

[0003] However, when dealing with the sentence semantic matching task, traditional sentence semantic matching models usually fail to fully consider the differences in the matching granularity of text content at different levels, which limits the further improvement of model performance. Because when dealing with sentence semantic matching, there is often information unrelated to the theme in natural language, which makes the sentence semantic matching model inevitably face the problem of noise interference, thus affecting the accuracy of matching. Summary of the Invention

[0004] The technical problem to be solved by the present invention is to provide a sentence semantic matching method and device for maintaining key semantics in view of the deficiencies of the prior art.

[0005] The technical solution of the present invention for solving the above technical problems is as follows:

[0006] A sentence semantic matching method for maintaining key semantics includes the following steps:

[0007] Construct a sentence semantic matching model, where the sentence semantic matching model includes an encoder, a feature aggregator, a feature distillation layer, and a classification layer;

[0008] Input the sentence pair to be matched into the sentence semantic matching model for training, including: the encoder extracts features from the sentence pair to be matched to obtain an encoded semantic representation and an aggregated semantic representation, the feature aggregator performs an aggregation calculation on the encoded semantic representation to obtain weakly related features, the feature distillation layer performs a projection process on the aggregated semantic representation and the weakly related features to obtain key semantic features, the classification layer classifies the key semantic features to obtain a matching result, and completes the preliminary training of the sentence semantic matching model;

[0009] Calculate the loss of the aggregated semantic representation, the weakly related features, and the key semantic features through a total loss function to obtain a total loss;

[0010] Optimize the sentence semantic matching model after preliminary training according to the total loss to obtain an optimized sentence semantic matching model.

[0011] Another technical solution for the present invention to solve the above technical problems is as follows:

[0012] A sentence semantic matching device for maintaining key semantics, comprising:

[0013] A construction module for constructing a sentence semantic matching model, where the sentence semantic matching model includes an encoder, a feature aggregator, a feature distillation layer, and a classification layer;

[0014] A training module for inputting the sentence pairs to be matched into the sentence semantic matching model for training, including: the encoder extracts features from the sentence pairs to be matched to obtain encoded semantic representations and aggregated semantic representations, the feature aggregator performs aggregation calculations on the encoded semantic representations to obtain weakly related features, the feature distillation layer performs projection processing on the aggregated semantic representations and the weakly related features to obtain key semantic features, and the classification layer classifies the key semantic features to obtain a matching result and complete the preliminary training of the sentence semantic matching model;

[0015] An evaluation module for calculating losses of the aggregated semantic representations, the weakly related features, and the key semantic features through a total loss function to obtain a total loss;

[0016] An optimization module for optimizing the sentence semantic matching model after preliminary training according to the total loss to obtain an optimized sentence semantic matching model.

[0017] The beneficial effects of the present invention are: during the process of training the sentence semantic matching model, the encoder encodes semantic information of the sentence pairs to be matched, accurately captures the context information and semantic features of the sentences, the feature aggregator filters out the features weakly related to the semantic main idea from the sentence pairs and determines them as noise features, the feature distillation layer eliminates redundant noise information in the semantic features based on the projection theorem, retains the key semantic features, and the classification layer generates a semantic matching result of the sentence pairs based on the key semantic features. The performance of the sentence semantic matching model after preliminary training is evaluated through a loss function, and the model is optimized according to the loss value, so that the model can accurately match the semantics of the input text sentences, extract more representative and effective features, and generate accurate matching results in the matching task. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 It is a flowchart of the key semantic-preserving sentence semantic matching method provided by an embodiment of the present invention;

[0019] Figure 2 It is a block diagram of the modules of the key semantic-preserving sentence semantic matching device provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0020] The principles and features of the present invention are described below in conjunction with the accompanying drawings. The examples given are only used to explain the present invention and are not used to limit the scope of the present invention.

[0021] Regarding sentence semantic matching, in the field of information retrieval, it can help users find the information they need quickly and accurately; in the industrial field, text matching (i.e., sentence semantic matching) technology is widely used in language question-answering robots, improving the robot model's ability to understand real scenes and promoting the development of industrial automation. Therefore, an effective sentence semantic matching model plays a vital role in natural language processing and industrial applications.

[0022] In order to solve the problem of noise interference when the sentence semantic matching model performs semantic matching, the academic community has proposed two main strategies: one is to enhance the model's ability to resist noise by introducing noise in the text representation stage; the other is to use feature selection methods to reduce the impact of noise after obtaining text semantic information. With the development of neural network technology and attention mechanism, the model's computing power has been improved and data resources have been enriched. The fine-tuning method based on pre-training has achieved remarkable results in the field of text matching, especially the Transformer architecture-based model (such as BERT and its derivative models), which provides powerful text representation capabilities for various downstream tasks (such as intelligent question answering, information retrieval and recommendation, and sentiment analysis and intent recognition), thereby significantly enhancing the generalization ability of the model.

[0023] Nevertheless, existing technologies usually fail to fully consider the differences in matching granularity of text content at different levels when dealing with sentence matching tasks, which limits the further improvement of model performance. To meet this challenge, matching methods that integrate multiple features (such as DC-Match (Divide and Conquer Match, i.e., text semantic matching method) and MCP-SM (Multi-Concept Parsed Semantic Matching, i.e., multi-language semantic matching method) have been proposed. These methods aim to deeply explore the semantic information of different granularities in sentences. However, although these matching methods or large pre-trained language models (such as ALBERT and DeBERTa, etc.) have made progress in sentence semantic matching tasks, they are still disturbed by noise information, which affects the accuracy of matching.

[0024] like Figure 1 As shown, a sentence semantic matching method for maintaining key semantics provided by an embodiment of the present invention includes the following steps:

[0025] Constructing a sentence semantic matching model, wherein the sentence semantic matching model includes an encoder, a feature aggregator, a feature distillation layer, and a classification layer;

[0026] Input the sentence pairs to be matched into the sentence semantic matching model for training, including: the encoder extracts features from the sentence pairs to be matched to obtain encoded semantic representations and aggregated semantic representations, the feature aggregator performs aggregation calculations on the encoded semantic representations to obtain weakly relevant features, the feature distillation layer performs projection processing on the aggregated semantic representations and the weakly relevant features to obtain key semantic features, and the classification layer classifies the key semantic features to obtain a matching result, and completes the preliminary training of the sentence semantic matching model;

[0027] Calculate the loss of the aggregated semantic representation, the weakly relevant features, and the key semantic features through the total loss function to obtain the total loss;

[0028] Optimize the sentence semantic matching model after preliminary training according to the total loss to obtain an optimized sentence semantic matching model.

[0029] Predict the target sentence pair to be matched through the optimized sentence semantic matching model to obtain the optimal matching result.

[0030] It should be understood that the encoder of the sentence semantic matching model embeds a pre-trained language model PLM. The pre-trained language model PLM includes language models such as BERT, RoBERTa, ALBERT, DeBERTa, EntityCS, KBioXLM, Twhin-bert, and XLM-R. The pre-trained language model PLM has mastered rich language knowledge and context understanding ability through self-supervised learning on large-scale unlabeled text data, providing strong support for various downstream tasks.

[0031] In the embodiment of the present invention, the sentence semantic matching model is trained by the sentence pairs to be matched and the pre-annotated labels to generate an optimal sentence semantic matching model. During the training of the sentence semantic matching model, the encoder encodes the semantic information of the sentence pairs to be matched, accurately captures the context information and semantic features of the sentences, the feature aggregator filters out the features with weak relevance to the semantic theme from the sentence pairs and determines them as noise features, the feature distillation layer eliminates the redundant noise information in the semantic features based on the projection theorem, retains the key semantic features, and the classification layer generates the semantic matching result of the sentence pair based on the key semantic features. The performance of the sentence semantic matching model after preliminary training is evaluated through the loss function, and the model is optimized according to the loss value, so that the model can accurately match the semantics of the input text sentences, extract more representative and effective features, and generate accurate matching results in the matching task.

[0032] Preferably, the sentence pair to be matched includes a first sentence and a second sentence; the sentence semantic matching model further includes a connection layer; after the step of inputting the sentence pair to be matched into the sentence semantic matching model, the following steps are further included:

[0033] The connection layer connects the first sentence and the second sentence through a delimiter to obtain a sentence sequence, which is expressed as:

[0034]

[0035] Second sentence.

[0036] In the embodiment of the present invention, connecting the two sentences facilitates locking the sentences to be matched during semantic matching and enables the matching task to be carried out smoothly.

[0037] Preferably, the encoder extracts features from the sentence pair to be matched to obtain an encoded semantic representation and an aggregated semantic representation, including:

[0038] The encoder extracts hidden features from the sentence pair to be matched to obtain an encoded semantic representation;

[0039] The encoder extracts semantic features from the sentence pair to be matched to obtain an aggregated semantic representation.

[0040] Specifically, the encoder includes multiple hidden layers, and hidden features or semantic features are extracted from the sentence pair to be matched through each hidden layer one by one. The encoded semantic representation and the aggregated semantic representation are output after being processed by the last hidden layer. The feature extraction expression is:

[0041] [k cls ; K a,b = PLM([w cls ; X a,b ),

[0042] where k cls is the aggregated semantic representation, K a,b is the encoded semantic representation, is the i-th encoded semantic representation of the sentence sequence, w cls is a special sequence marker preset at the front end of the sentence pair to be matched, X a,b is the sentence to be matched (i.e., the sentence sequence), and PLM(·) is the encoder.

[0043] In the embodiment of the present invention, features with text context information in the sentence pair are extracted through hidden features, which facilitates finding out noise information that is not strongly related to the text theme. Semantic features expressing the context meaning of the text in the sentence pair are extracted through semantic features, which facilitates finding out noise information that is not strongly related to the text theme.

[0044] Preferably, the feature aggregator includes a dot-product attention mechanism;

[0045] The feature aggregator performs an aggregation calculation on the encoded semantic representation to obtain weakly correlated features, including:

[0046] The dot-product attention mechanism performs a dot-product calculation on the encoded semantic representation according to a preset query vector to obtain weakly correlated feature weight values;

[0047] The dot-product attention mechanism performs a dot-product calculation on the weakly correlated feature weight values and the encoded semantic representation to obtain weakly correlated features.

[0048] Specifically, a variant of the dot-product attention mechanism is used, that is, a trainable query vector is used to aggregate features that are weakly correlated with the sentence semantic matching relationship;

[0049] A dot-product calculation is performed on each encoded semantic representation of the sentence sequence according to a preset query vector to obtain multiple weakly correlated feature weight values. The first dot-product expression is:

[0050]

[0051]

[0052] The dot-product calculation is performed on each weakly correlated feature weight value and the corresponding encoded semantic representation to obtain multiple weakly correlated features. The second dot-product expression is:

[0053]

[0054] Wherein, represents the nth weakly correlated feature weight value, represents the nth encoded semantic representation, f i * represents the feature (i.e., the weakly correlated feature) that is weakly correlated with the semantic matching relationship after summarizing the ith encoded semantic representation, z represents the query vector, and Softmax(·) represents the normalized exponential function.

[0055] In the embodiment of the present invention, weakly correlated features with weak relevance are matched from the features with text context information according to the dot-product attention and the Softmax function for elimination, and key semantic features with matching significance are retained.

[0056] Preferably, the feature distillation layer performs a projection process on the aggregated semantic representation and the weakly correlated features to obtain key semantic features, including:

[0057] The feature distillation layer performs a projection calculation on the aggregated semantic representation and the weakly correlated features according to a projection operator to obtain noise features;

[0058] The feature distillation layer removes the noise features in the aggregated semantic representation to obtain a denoised aggregated semantic representation, and performs projection calculation on the aggregated semantic representation and the denoised aggregated semantic representation according to a projection operator to obtain key semantic features.

[0059] Specifically, project the aggregated semantic representation of the i-th encoded semantic representation onto the direction represented by the i-th weakly correlated feature f i * (i.e., the orthogonal space) to identify the noise features mixed with the aggregated semantic representation That is, based on the projection theorem (projector), perform a linear transformation on the aggregated semantic representation k cls and the weakly correlated feature f i * to obtain the noise feature to better distinguish and utilize the relevant features for the sentence matching task while reducing the interference of noise. The expression for calculating the noise feature by projection is:

[0060]

[0061] where represents the i-th noise feature, represents the i-th aggregated semantic representation of the sentence sequence, and f i * represents the i-th weakly correlated feature. Under the action of the projection operator Proj(·,·), perform projection calculation on the aggregated semantic representation of the final hidden layer and the weakly correlated feature f i *

[0062] Use the projection operation to remove the noise features in the aggregated semantic representation, so as to obtain a more refined aggregated sequence representation of the final hidden layer of the sentence to be matched, that is, further eliminate the negative impact of the noise component on the sentence matching effect, so as to obtain key matching features. The expression for calculating the key semantic feature by projection is:

[0063]

[0064] where f i p represents the key semantic feature, Proj(·,·) represents the projection operator, and the calculation formula of the projection operator is:

[0065]

[0066] For example, the calculation of the noise feature is:

[0067] ​

[0068] In the embodiment of the present invention, the projection theorem is used for feature decomposition during dimensionality reduction processing, and the projection theorem can effectively remove the noise features in the sentence semantic information. A feature distillation layer based on the projection theorem is introduced, aiming to improve the sentence matching effect and remove the noise features irrelevant to semantics by projecting the aggregated semantic representation of the sentence sequence output by the final hidden layer of the encoder into a space orthogonal to the weakly related features. By designing a feature distillation layer with adaptive context, the recognition ability of weakly related features is improved, and important features can be extracted more efficiently. The noise features in the semantic features output by the hidden layer of the encoder are identified and removed based on the projection theorem, thereby improving the matching accuracy.

[0069] Preferably, the key semantic features are classified by the classification layer to obtain a matching result, including:

[0070] The classification layer performs a matching calculation on the key semantic features according to a preset matching weight matrix to obtain a matching result.

[0071] Specifically, when making a prediction, a separate classification layer is used to output the result. The classification layer includes a trainable weight W ∈ R Z×C , where Z represents the number of matching result labels, and C represents the hidden layer size. By learning the trainable weight, the classification layer can predict the corresponding matching label y according to the features of the input sequence, and the matching result prediction expression is:

[0072] P(y|f i (a)p ,f i (b)p ) = Softmax(f i p W T ),

[0073] where P(y|f i (a)p ,f i (b)p ) represents the matching probability distribution (i.e., the matching result) of the i-th key semantic feature f i (a)p of the first sentence and the i-th key semantic feature f i (b)p of the second sentence to the matching label y, f i p represents the i-th key semantic feature, and W represents the matching weight matrix.

[0074] In an embodiment of the present invention, the classification layer classifies the degree of key feature matching of sentence pairs to generate a matching result, thereby providing an information basis for downstream tasks such as intelligent question answering, information retrieval and recommendation, and sentiment analysis and intention recognition.

[0075] Preferably, calculating the loss of the aggregated semantic representation, the weakly relevant features, and the key semantic features through the total loss function to obtain the total loss includes:

[0076] Calculating the weakly relevant features through the cross-entropy function to obtain the gradient reverse loss, calculating the key semantic features through the classification loss function to obtain the standard classification loss, calculating the aggregated semantic representation and the key semantic features through the margin loss function to obtain the margin loss, and summing the gradient reverse loss, the standard classification loss, and the margin loss to obtain the total loss.

[0077] Specifically, three loss functions (i.e., L J 、L GRL and L I ) are used to jointly train the sentence semantic matching model, aiming to optimize the model performance from different perspectives. Each loss function has its specific optimization objective. L J is the classification loss of the global matching model, used to judge the semantic relationship between sentences processed by the feature distillation layer. L GRL is the loss for optimizing weakly relevant context features, aiming to filter out features weakly relevant to the sentence theme. L I is the margin loss, in order to encourage the key semantic features learned by the model to start from the original text semantic representation. The expression of the total loss function is:

[0078] L = L J + L GRL + L I ,

[0079] where L represents the total loss, L J represents the standard classification loss, L GRL represents the gradient reverse loss, and L I is the margin loss.

[0080] In an embodiment of the present invention, by jointly optimizing the model according to the values of the three loss functions, the overall performance of the model in the sentence matching task is improved to provide more accurate and reliable matching results.

[0081] Preferably, the feature aggregator further includes a gradient reverse layer and a relationship classifier;

[0082] Calculating the weakly relevant features through the cross-entropy function to obtain the gradient reverse loss includes:

[0083] The gradient reversal layer optimizes the set propagation parameters to obtain gradient parameters;

[0084] The relationship classifier classifies and predicts the weakly related features according to the gradient parameters to obtain the matching probability value of the weakly related features, and calculates the gradient reversal loss through the cross-entropy function for the matching probability value.

[0085] Specifically, a gradient reversal layer (GRL) is introduced to screen out features weakly related to the matching relationship from the semantic representation of the sentence pairs to be matched. By optimizing the gradient parameters of the relationship classifier during the training process, the query vector is further optimized to better capture features weakly related to the matching relationship, indirectly improving the performance of the model in the sentence matching task.

[0086] The matching probability expression is:

[0087] prob i =Softmax(GRL(f i * )·W + b),

[0088] The cross-entropy function is:

[0089] L GRL,i =CrossEntropy(y i , prob i ),

[0090] where W and b represent the weight matrix and bias of the relationship classifier respectively, prob i represents the matching probability value of the i-th weakly related feature of the sentence pairs to be matched. After being processed by the GRL layer, f i * is input into the relationship classifier for prediction. CrossEntropy(.) represents the cross-entropy function, y i represents the true label of the i-th matching probability of the sentence pairs to be matched, and L GRL,i represents the i-th gradient reversal loss of the sentence pairs to be matched, which is used to optimize the training process of the query vector.

[0091] It should be understood that the features learned by the feature extractor are invariant between different domains (such as the source domain and the target domain), that is, the features generated by the feature extractor should not contain domain-specific information. The gradient reversal layer can make the optimization objective of the feature extractor opposite to that of the domain classifier by multiplying the gradient by a negative number during backpropagation. The domain classifier is used to distinguish the source (i.e., domain) of the features, while the feature extractor is used to generate features that are difficult to be distinguished by the domain classifier. This mechanism of adversarial training can prompt the feature extractor to learn more generalizable feature representations (i.e., optimize the query vector to match weakly related features).

[0092] In the embodiments of the present invention, by minimizing the gradient reverse loss, the model can better learn the features weakly related to the matching relationship, thereby improving the performance of sentence matching.

[0093] Preferably, the standard classification loss is obtained by calculating the key semantic features through a classification loss function, and the classification loss function is:

[0094] L J =-logP(y|f i (a)p ,f i (b)p ),

[0095] where L J is the standard classification loss, and P(y|f i (a)p ,f i (b)p ) is the matching result.

[0096] In the embodiments of the present invention, f i (a)p and f i (b)p are used to more accurately represent the aggregated sequence representations (i.e., key semantic features) of the sentences a and b to be matched. For model training, a classification loss function L J is introduced to calculate the standard classification loss of the global matching model, measure the performance of the model in the sentence matching task, and optimize the model parameters by minimizing the error, thereby improving the accuracy of the model in sentence matching.

[0097] Preferably, the margin loss is obtained by calculating the aggregated semantic representation and the key semantic features through a margin loss function, including:

[0098] Calculate the similarity between the aggregated semantic representation and the key semantic features to obtain a margin loss value, and calculate the margin loss through a margin loss function. The margin loss function is:

[0099] L I =max(0,ρ - ξ),

[0100] where L I is the margin loss, ρ is a hyperparameter (set to 0.006), ξ is the margin loss value, and max(·,·) is the maximum value function.

[0101] In the embodiments of the present invention, considering that directly using a distance function to maintain the original text semantics is likely to lead to overfitting, through the original text key semantics maintenance strategy, an interval loss function is adopted and combined with cosine similarity to measure the gap between the key features and the original hidden layer semantic features, ensuring that the extracted features accurately reflect the original text semantics, avoiding the occurrence of overfitting, and further improving the accuracy of the model through iterative updating of the features.

[0102] Preferably, the aggregated semantic representation consists of multiple aggregated sub-semantic representations, and the key semantic features consist of multiple key sub-semantic features;

[0103] Calculating the similarity between the aggregated semantic representation and the key semantic features to obtain an interval loss value includes:

[0104] Calculating the similarity between any one of the aggregated sub-semantic representations and any one of the key sub-semantic features through a cosine function to obtain a cosine similarity. Through this process, calculating all the aggregated sub-semantic representations and key sub-semantic features to obtain multiple cosine similarities. The cosine function is:

[0105]

[0106] Wherein, is the cosine similarity between the i-th aggregated semantic representation and the i-th aggregated semantic representation, cos(·,·) is the cosine function, is the i-th aggregated semantic representation, is the j-th key semantic feature;

[0107] Calculating the difference of multiple cosine similarities respectively through an interval difference expression to obtain an interval loss value. The interval difference expression is:

[0108]

[0109] Where ξ i is the interval loss value of the i-th aggregated semantic representation, is the cosine similarity between the i-th aggregated semantic representation and the k-th key semantic feature, is the cosine similarity between the i-th aggregated semantic representation and the j-th key semantic feature, and max(·,·) is the maximum value function.

[0110] In the embodiments of the present invention, the cosine function is used to calculate the feature distance between the aggregated semantic representation and the key semantic features output by the final hidden layer, and the cosine similarity is maximized to optimize the model, thereby ensuring that the extracted key features can accurately reflect the semantics of the original text.

[0111] Such as Figure 2As shown in the figure, a sentence semantic matching device for maintaining key semantics provided by an embodiment of the present invention includes:

[0112] A construction module for constructing a sentence semantic matching model, where the sentence semantic matching model includes an encoder, a feature aggregator, a feature distillation layer, and a classification layer;

[0113] A training module for inputting the sentence pairs to be matched into the sentence semantic matching model for training, including: the encoder extracts features from the sentence pairs to be matched to obtain encoded semantic representations and aggregated semantic representations, the feature aggregator performs aggregation calculations on the encoded semantic representations to obtain weakly related features, the feature distillation layer performs projection processing on the aggregated semantic representations and the weakly related features to obtain key semantic features, the classification layer classifies the key semantic features to obtain a matching result, and completes the preliminary training of the sentence semantic matching model;

[0114] An evaluation module for calculating losses for the aggregated semantic representations, the weakly related features, and the key semantic features through a total loss function to obtain a total loss;

[0115] An optimization module for optimizing the sentence semantic matching model after preliminary training according to the total loss to obtain an optimized sentence semantic matching model.

[0116] For the above-mentioned sentence semantic matching device for maintaining key semantics, reference may be made to the implementation content and beneficial effects specifically described above for a sentence semantic matching method for maintaining key semantics, which will not be elaborated here.

[0117] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variation thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device.

[0118] Those skilled in the art can clearly understand that for the sake of convenience and brevity of description, the specific working processes of the above-described device and module can refer to the corresponding processes in the foregoing method embodiments, which will not be elaborated here.

[0119] In several embodiments provided by the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of modules is only a logical function division. In actual implementation, there may be other division methods. For example, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed.

[0120] The modules described as separate components may or may not be physically separated, and the components shown as modules may or may not be physical modules, that is, they can be located in one place or distributed to multiple network modules. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of the embodiments of the present invention.

[0121] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A sentence semantic matching method for maintaining key semantics, characterized in that, The method includes the following steps: Construct a sentence semantic matching model, where the sentence semantic matching model includes an encoder, a feature aggregator, a feature distillation layer, and a classification layer; Input the sentence pair to be matched into the sentence semantic matching model for training, including: the encoder extracts features from the sentence pair to be matched to obtain an encoded semantic representation and an aggregated semantic representation, the feature aggregator performs an aggregation calculation on the encoded semantic representation to obtain weakly related features, the feature distillation layer performs a projection process on the aggregated semantic representation and the weakly related features to obtain key semantic features, and the classification layer classifies the key semantic features to obtain a matching result, and completes the preliminary training of the sentence semantic matching model; Calculate the loss of the aggregated semantic representation, the weakly related features, and the key semantic features through a total loss function to obtain a total loss; Optimize the sentence semantic matching model after preliminary training according to the total loss to obtain an optimized sentence semantic matching model.

2. The sentence semantic matching method according to claim 1, characterized in that The encoder extracts features from the sentence pair to be matched to obtain an encoded semantic representation and an aggregated semantic representation, including: The encoder extracts hidden features from the sentence pair to be matched to obtain an encoded semantic representation; The encoder extracts semantic features from the sentence pair to be matched to obtain an aggregated semantic representation.

3. The sentence semantic matching method according to claim 1, wherein The feature aggregator includes a dot product attention mechanism; The feature aggregator performs an aggregation calculation on the encoded semantic representation to obtain weakly related features, including: The dot product attention mechanism performs a dot product calculation on the encoded semantic representation according to a preset query vector to obtain weakly related feature weight values; The dot product attention mechanism performs a dot product calculation on the weakly related feature weight values and the encoded semantic representation to obtain weakly related features.

4. The sentence semantic matching method according to claim 1, wherein The feature distillation layer performs a projection process on the aggregated semantic representation and the weakly related features to obtain key semantic features, including: The feature distillation layer performs a projection calculation on the aggregated semantic representation and the weakly related features according to a projection operator to obtain noise features; The feature distillation layer removes the noise features in the aggregated semantic representation to obtain a denoised aggregated semantic representation, and performs a projection calculation on the aggregated semantic representation and the denoised aggregated semantic representation according to a projection operator to obtain key semantic features.

5. The sentence semantic matching method according to claim 1, characterized in that The classification layer classifies the key semantic features to obtain a matching result, including: The classification layer performs a matching calculation on the key semantic features according to a preset matching weight matrix to obtain a matching result.

6. The sentence semantic matching method according to claim 1, wherein The calculation of the loss of the aggregated semantic representation, the weakly related features, and the key semantic features through the total loss function to obtain a total loss, including: Calculate the gradient reverse loss through a cross-entropy function for the weakly related features, calculate the standard classification loss through a classification loss function for the key semantic features, calculate the margin loss through a margin loss function for the aggregated semantic representation and the key semantic features, and sum the gradient reverse loss, the standard classification loss, and the margin loss to obtain a total loss.

7. The sentence semantic matching method according to claim 6, characterized in that, The feature aggregator further includes a gradient reverse layer and a relationship classifier; Calculating the gradient reverse loss for the weakly relevant features through the cross-entropy function includes: The gradient reverse layer optimizes the set propagation parameters to obtain gradient parameters; The relationship classifier classifies and predicts the weakly relevant features according to the gradient parameters to obtain the matching probability value of the weakly relevant features, and calculates the gradient reverse loss by calculating the matching probability value through the cross-entropy function.

8. The sentence semantic matching method according to claim 6, wherein Calculating the margin loss for the aggregated semantic representation and the key semantic features through the margin loss function includes: Calculating the similarity between the aggregated semantic representation and the key semantic features to obtain a margin loss value, and calculating the margin loss by calculating the margin loss value through the margin loss function. The margin loss function is: L I = max(0, ρ - ξ), where L I is the margin loss, ρ is a hyperparameter, ξ is the margin loss value, and max(·, ·) is the maximum function.

9. The sentence semantic matching method according to claim 8, wherein The aggregated semantic representation is composed of multiple aggregated sub-semantic representations, and the key semantic features are composed of multiple key sub-semantic features; Calculating the similarity between the aggregated semantic representation and the key semantic features to obtain a margin loss value includes: Calculating the cosine similarity between any one of the aggregated sub-semantic representations and any one of the key sub-semantic features through the cosine function, and calculating multiple cosine similarities for all aggregated sub-semantic representations and key sub-semantic features in this process; Calculating the difference of multiple cosine similarities respectively through the margin difference expression to obtain the margin loss value.

10. A sentence semantic matching device for maintaining key semantics, characterized in that, Including: A construction module for constructing a sentence semantic matching model, where the sentence semantic matching model includes an encoder, a feature aggregator, a feature distillation layer, and a classification layer; A training module for inputting the sentence pair to be matched into the sentence semantic matching model for training, including: the encoder extracts features from the sentence pair to be matched to obtain an encoded semantic representation and an aggregated semantic representation, the feature aggregator performs aggregation calculation on the encoded semantic representation to obtain weakly relevant features, the feature distillation layer performs projection processing on the aggregated semantic representation and the weakly relevant features to obtain key semantic features, the classification layer classifies the key semantic features to obtain a matching result, and completes the preliminary training of the sentence semantic matching model; An evaluation module for calculating the loss of the aggregated semantic representation, the weakly relevant features, and the key semantic features through a total loss function to obtain a total loss; An optimization module for optimizing the sentence semantic matching model after preliminary training according to the total loss to obtain an optimized sentence semantic matching model.