Intelligent scheme recommendation method based on semantic matching

By using a text embedding model based on transformer architecture in information retrieval and performing migration training, the problem of insufficient semantic understanding in the existing technology is solved, and a more accurate and personalized recommendation effect is achieved.

CN119961508APending Publication Date: 2025-05-09THE 28TH RES INST OF CHINA ELECTRONICS TECH GROUP CORP

Patent Information

Application Number
CN202411817002.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-11
Publication Date
2025-05-09

AI Technical Summary

Technical Problem

The prior art is difficult to fully understand the semantics of text in information retrieval, resulting in low accuracy of recommendation results, especially in poor performance in specific fields.

Method used

A text embedding model based on transformer architecture is adopted, combining the self-attention mechanism and a fully connected layer, and optimizing model parameters in a specific field through migration training, generating more accurate word vectors and sentence vector representations.

Benefits of technology

It improves the accuracy of semantic matching in information retrieval, provides more personalized and accurate recommendation services, especially in specific fields, which significantly improves recommendation results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119961508A_ABST
    Figure CN119961508A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of natural language processing and information retrieval, and discloses an intelligent scheme recommendation method based on semantic matching. The method comprises the following steps: constructing a general text embedding model, fully mining a context relationship of a scheme text by using a self-attention mechanism of a transform architecture, and calculating the semantic similarity of the scheme; and migration training is performed on the general text embedding model, so that the accuracy of scheme recommendation of the text embedding model in a specific field is improved. And evaluating recommendation effects of different models by using a public data set, and calculating indexes such as an average value of average accuracy, an average value of average reciprocal ranking and normalized loss cumulative gain. The scheme recommendation effect of the model adopted by the invention is better than that of a traditional model, and various evaluation indexes are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of natural language processing (NLP) and information retrieval, and in particular to an intelligent solution recommendation method based on semantic matching. Background Art

[0002] This section merely provides background information related to the present disclosure and is not necessarily prior art.

[0003] Information explosion and personalized needs: With the rapid development of the Internet, information is growing explosively. Faced with a huge amount of information, how to quickly and accurately find the content of interest to users has become an urgent problem to be solved. Fortunately, the emergence of semantic analysis technology has provided a new idea for solving this problem. Semantic analysis technology can deeply understand the meaning and content behind the text, and provide more accurate semantic information for the solution recommendation system by extracting key information and quantifying the semantic relationship between words. The development of this technology makes it possible to intelligently recommend solutions based on semantic matching.

[0004] Traditional keyword-based recommendation methods often fail to fully understand the semantics of text, resulting in low accuracy of recommendation results. At the same time, general text embedding models often perform poorly in specific fields and do not mine sufficient semantic information, resulting in distortion of personalized recommendation content. Summary of the invention

[0005] The technical problem to be solved by the present invention is to propose a solution intelligent recommendation method based on semantic matching in view of the deficiencies of the above-mentioned prior art.

[0006] The technical solution adopted by the present invention to solve the technical problem is as follows:

[0007] A solution intelligent recommendation method based on semantic matching, such as Figure 1 As shown, the following steps are included:

[0008] Step 1: Build a text embedding model. The model is based on the transformer architecture and includes two layers: a self-attention mechanism and a fully connected layer (Feed Forward Neural Network). Use a general dataset to train the model and learn the semantic relationships in the text.

[0009] Step 2: Use a public dataset to evaluate the model obtained in step 1. Through experiments, it can be found that the recommendation effect of the text embedding model used in this patent is better than other models.

[0010] Step 3: Construct migration training data and domain-specific query solutions based on domain-specific datasets.

[0011] Step 4: Use the migration training data obtained in step 3 to perform migration training on the text embedding model, adjust the model parameters, generate word vectors for each word, and establish a word vector table corresponding to the word table.

[0012] Step 5: Perform word segmentation, data preprocessing and encoding on the corpus in the domain-specific dataset, and convert the scheme into word-unit representation.

[0013] Step 6: Use the text embedding model after transfer training obtained in step 4 to convert each word into a corresponding word vector, generate a sentence vector or a segment vector according to the pooling strategy, and build a vector representation library of the corpus.

[0014] Step 7: When the user enters a query solution, the query solution is segmented, preprocessed, encoded, and converted into sentence vectors and segment vectors. The specific processing method is the same as step 5 and step 6.

[0015] Step 8, calculate the cosine similarity between the sentence or paragraph vector of the query solution and all vectors in the query solution library, and return a specified number of solutions and similarity scores according to the cosine similarity score.

[0016] Furthermore, the step 1 of constructing and training the text embedding model includes:

[0017] Step S111, load the general data set, input the token list after word segmentation into the token embedding layer, the segment embedding layer and the position embedding layer respectively, and obtain the corresponding embedding representation; wherein, in the token embedding layer, add [CLS], [SEP] and other tags to the word segmentation result, generate attention_mask, and convert all tags into embedding vectors. Two embedding vectors are returned in the segment embedding layer, which can be used to distinguish two sentences input at the same time. In the position embedding layer, the order information of the words appearing in the sentence is encoded, and the position vector is output.

[0018] Step S112, accumulating the two embedding vectors and position vector obtained in step S111 as the output vector e of the self-attention module i .

[0019] Step S113: input vector e i With three weight matrices W Q , W K , W VMultiply, the weight matrix is ​​randomly initialized and optimized by back propagation during training to generate the query vector Q i , key vector K i Sum value vector V i .

[0020] Q i =e i ·W Q ,

[0021] K i =e i ·W K ,

[0022] V i =e i ·W V ;

[0023] Step S114, calculate the attention score of each input vector ij , and score ij Normalization is performed to prevent the dot product result from being too large, causing the softmax function to enter the saturation region. The formula is as follows:

[0024]

[0025] The score ij is the attention score between the i-th word and the j-th key, d k is the key vector K i Dimension.

[0026] Step S115: Convert the score into attention weight AttentionWeight through the softmax activation function ij ,

[0027]

[0028] Among them, AttentionWeight ij It indicates the attention that the word vector at the jth position receives when generating the output at the i-th position.

[0029] Step S116, weighted sum, using attention weight AttentionWeight ij The value vector V i Perform weighted summation to obtain the output vector Output of the i-th word i ,

[0030] Output i =AttentionWeight ij ·V j,

[0031] The above process formula is summarized as follows:

[0032]

[0033] Step S117: In the process of converting word units into vectors, a multi-head attention mechanism is used to focus on different positions in the input sequence at the same time, and the input sequence is divided into multiple heads. Each head performs the above operations independently, and finally the outputs of all heads are multiplied by the corresponding weights to obtain the output vector Z. i .

[0034] Step S118: Output the word vector Z i It is transmitted to the fully connected layer, which includes two layers. The first layer is the ReLU activation function and the second layer is the linear activation function, which is expressed as:

[0035] FFN(Z i )=max(0,Z i W1+b1)W2+b2,

[0036] Where W1 and W2 are weight matrices, b1 and b2 are bias terms (the parameters are randomly initialized and optimized by back-propagation during training).

[0037] The ReLU activation function is a nonlinear function that learns the nonlinear relationship of the input vector and captures the dependencies and patterns in the input sequence. The second layer is a linear layer without an activation function. The second layer of linear activation function further transforms the features after nonlinear transformation to generate the final word vector output Z′ i .

[0038] Step S119, perform the Masked Language Modeling (MLM) pre-training task, use the special tag [MASK] to replace some words in the input sequence, and perform the word vector Z′ obtained in S118 on the linear layer. i Linear transformation:

[0039] O i =WZ′ i +b,

[0040] Among them, i is the output vector, W is the weight matrix, b is the bias term, the weight matrix W and the bias term b are randomly initialized before the training begins. During the training process, the values ​​of the weight matrix W and the bias term b are adjusted according to the loss function. Among them, the output vector O of the linear layer i Each element of corresponds to a word in the vocabulary, and the value of the element represents the probability that the model believes that the word appears in the masked position.

[0041] Step S120, normalizing using a softmax activation function to generate a probability distribution, in which each element represents the probability that the corresponding word in the vocabulary appears in the masked position, and the probability distribution is the final prediction of the model.

[0042]

[0043] where p i is the probability that the i-th word in the vocabulary appears in the masked position.

[0044] Step S121, using the cross entropy loss function as the loss function, calculates the difference between the probability distribution obtained in S120 and the true value of the masked word, and in the model training process, takes the value of the minimization loss function as the goal, adjusts the word vector, weight matrix, bias term and other parameters, so that the output probability distribution is close to the true probability distribution. The loss function uses the cross entropy loss function, and the formula is as follows:

[0045]

[0046] where y i is a duress encoding (one-hot encoding) vector, indicating whether it is the real target word, that is, the real value of the masked word; at the real target word position, the yi value is 1, otherwise it is 0.

[0047] Furthermore, the evaluation of the model obtained in step 1 in step 2 includes:

[0048] Step S211: Use part of the public data set for experimental evaluation, where the part of the data is the first 15,000 pieces of data in the public data set.

[0049] Step S212, respectively evaluate the S-BERT model, the Text2vec model and the model constructed by this patent, and calculate the indicators.

[0050] Step S213, counting the values ​​of different indicators of different models in the two parts of data and establishing a table.

[0051] Furthermore, the indicators calculated in step S212 include map@1, map@10, mrr@1, mrr@10, ndcg@1, and ndcg@10 indicators, and the calculation formula of each indicator is as follows:

[0052] (1) MAP@k (MeanAverage Precision at k) refers to the average of the average precision AP (Average Precision) of each query among the first k recommendation results returned;

[0053] The average precision AP is an evaluation indicator for a single query, which takes into account the ranking position of the query results in the relevant documents. For each query content, the query formula of AP is as follows:

[0054]

[0055] Where P(n) is the precision of the first n results, which is the number of relevant solutions in the first n results divided by n; rel(n) is a binary function that indicates whether the nth result is relevant, 1 indicates relevant, and 0 indicates irrelevant. Recall@k measures the proportion of truly relevant results in the first k returned results.

[0056]

[0057] MAP@k is the average mean precision (AP) of all queries.

[0058] (2) MRR@k (Mean Reciprocal Rank at k) refers to the average of the reciprocal values ​​of the ranking of the first relevant result among the first k results. This metric measures the speed at which the model finds the first relevant result.

[0059] For each query, the reciprocal rank (RR) is calculated as:

[0060]

[0061] Among them, rank is the position of the first relevant result, starting from 1.

[0062] MRR@k is the average of the reciprocal values ​​RR of all queries.

[0063] (3) NDCG@k (NormalizedDiscounted Cumulative Gain at k) refers to the normalized value of the cumulative gain taking into account the position discount among the first k results. It measures the quality of the result list returned by the model, taking into account relevance and position.

[0064] 1) For a single query, the calculation formula for the discounted cumulative gain DCG is:

[0065]

[0066] Among them, rel(i) is the relevance score of the i-th result, and the relevance score is usually a value between 0 and a positive number, such as 0, 1, 2, 3, etc.

[0067] 2) Calculate the ideal DCG, i.e. IDCG, to normalize the DCG, sort the results by relevance score and calculate the DCG as IDCG k .

[0068] 3) Calculate the normalized discounted cumulative gain NDCG:

[0069]

[0070] NDCG@k is the average NDCG of all queries.

[0071] Furthermore, the step 3 constructs a migration training data sample and a query solution library for a specific field based on a specific field data set, including:

[0072] Step S311: construct a migration training data sample based on the specific terms that users are concerned about in the specific domain data set, use the data to perform migration training on the model, and then generate a feature vector that is more in line with the semantics of the specific domain.

[0073] Step S312: Generate a solution library to be queried for a specific field according to the field of concern.

[0074] Furthermore, the transfer training of the text embedding model in step 4 includes:

[0075] Step S411, select a loss function and calculate the loss.

[0076] Step S412, start training, load the migration training data samples, let the model learn how to distinguish positive samples from negative samples, generate a loss function and minimize the loss function as the goal, so that the embedding vector generated by the text embedding model is more in line with the specific semantics.

[0077] Furthermore, the loss function of step S411 is a sentence triplet negative softmax contrast loss function, and the calculation process is as follows:

[0078] Step S4111, calculate the similarity matrix:

[0079]

[0080] Among them, s ij is the similarity matrix, is the i-th specific term embedding, E j is the jth sample embedding. The sample can be a positive sample or a negative sample. T is the temperature parameter. The temperature parameter T is used to scale and balance the influence of positive and negative samples.

[0081] Step S4112, apply log-softmax:

[0082]

[0083] Among them, p j It is the normalized similarity matrix, which represents the probability distribution of each sample as a positive sample. Assume that there is one positive example and M negative samples for each specific term.

[0084] Step S4113, calculate the loss L, where N is the number of positive examples in a batch:

[0085]

[0086] The goal of the loss function is to maximize the probability of positive samples and minimize the probability of negative samples, using the minimized negative log-likelihood function.

[0087] Step S4114, exchange loss, by exchanging the roles of specific terms and positive examples, calculate the loss function and use it as the exchange loss. This helps the model learn a more symmetric similarity metric.

[0088] Furthermore, the step 5 performs word segmentation, data preprocessing and encoding on the corpus in the specific domain data set, and converts the scheme into word-unit representation, including:

[0089] Step S511, initialize the vocabulary, divide the input scheme into characters, count the frequency of each character, and obtain a primary character-level vocabulary.

[0090] Step S512, calculating the frequency of adjacent character pairs, and counting the frequencies of all adjacent character pairs, wherein the adjacent character pairs are two adjacent characters or a character pair that has been merged.

[0091] Step S513, a merging operation, selects the adjacent character pairs with the highest frequency in the vocabulary, merges them into a new character or vocabulary unit, and updates the vocabulary and the frequency statistics of the adjacent character pairs.

[0092] Step S514, repeating the merging operation until a preset vocabulary size is reached or there are no adjacent character pairs that can be merged.

[0093] Step S515, obtaining a final vocabulary. After all the above operations are completed, all the sub-words or words obtained constitute a vocabulary. The sub-words or words are the word segmentation results of the input solution.

[0094] Step S516, data preprocessing, filling short solutions according to the preset length of the text embedding model, truncating long solutions; replacing special characters and unknown words; removing irrelevant characters in the solution.

[0095] Step S517, encoding is to convert the input sequence into a data form acceptable to the model. The encoding result includes three parts: input_ids, token_type_ids and attention_mask; among them, input_ids represents the ID of each word in the word table, token_type_ids is used to distinguish which sentence each word belongs to when two sentences are input at the same time, and attention_mask is used to determine which words need to be paid attention to when generating word vectors.

[0096] Furthermore, the step 6 generates sentence vectors and segment vectors using the text embedding model after transfer training, and constructs a vector representation library of the query solution library, including:

[0097] Step S611, according to the word segmentation result, search the ID of each word in the word-gram table, and generate an attention_mask, and use the attention_mask to determine the word-gram that needs to be considered.

[0098] Step S612, extracting the word vector of each word from the text embedding model according to the ID of each word.

[0099] Step S613: Generate a sentence vector or a segment vector according to the pooling strategy.

[0100] Step S614, using a two-dimensional heat map to intuitively display the changes in the text vector before and after the transfer training.

[0101] Furthermore, the pooling strategy in step S613 is to select an average pooling strategy, obtain all words that need to be paid attention to according to attention_mask, and average each element in the word vector of all the word elements that need to be paid attention to as the sentence or paragraph vector.

[0102] Further, the step 8 comprises:

[0103] Step S811, normalizing the vector, wherein the vector includes the embedding vector of the query solution and the vector generated by the solution in the solution library to be queried.

[0104] Step S812, calculating the vector dot product is the cosine similarity, the formula is as follows:

[0105]

[0106] Among them, A represents the embedding vector of the query solution, and B represents the vector generated by a solution in the solution library to be queried.

[0107] Step S813, sorting the results, and returning a specified number of solutions and similarity scores according to the cosine similarity scores from high to low.

[0108] The present invention adopts the above technical solution to produce the following technical effects:

[0109] The present invention proposes a method for intelligent recommendation of solutions based on semantic matching. First, a general text embedding model is constructed, and the self-attention mechanism of the transformer architecture is used to fully mine the contextual relationship of the text and calculate the semantic similarity of the solutions. In addition, the general text embedding model is transferred and trained to improve the recommendation accuracy of the text embedding model in specific fields.

[0110] The advantages of the method are as follows: the present invention uses a solution intelligent recommendation method based on semantic matching, and takes text semantic information into consideration when calculating solution similarity. Compared with the solution intelligent recommendation method based on keywords, the present method can provide more accurate and personalized recommendation services. The text embedding model of the present invention uses a self-attention mechanism to fully mine text context information and solve the problem of insufficient semantic information mining by traditional neural networks. The method supports transfer training and can train the model using domain text according to the specific fields used by users, solving the problem that general text embedding models are not effective in specific fields. It effectively mines and accurately represents the specific semantics of professional fields, calculates the similarity between heterogeneous texts, and improves the rationality and accuracy of solution similarity calculation in specific fields. BRIEF DESCRIPTION OF THE DRAWINGS

[0111] The present invention will be further described in detail below in conjunction with the accompanying drawings and specific embodiments, and the above and / or other advantages of the present invention will become more clear.

[0112] Figure 1 The figure is a flow chart of the method for intelligently recommending solutions based on semantic matching according to the present invention.

[0113] Figure 2 A visualization diagram of sentence vectors of the intelligent solution recommendation method based on semantic matching of the present invention.

[0114] Figure 3 Schematic diagram of the difference in feature vectors between the positive example and the solution intelligent recommendation method based on semantic matching before transfer training of the present invention.

[0115] Figure 4 This is a schematic diagram of the difference in feature vectors between the method for intelligent recommendation of solutions based on semantic matching of the present invention after transfer training and the positive example.

[0116] Figure 5 Schematic diagram of the difference in feature vectors between negative examples and before migration training of the intelligent recommendation method for solutions based on semantic matching of the present invention.

[0117] Figure 6 This is a schematic diagram of the difference in feature vectors between the negative examples and the solution intelligent recommendation method based on semantic matching after transfer training of the present invention. DETAILED DESCRIPTION

[0118] The following describes the embodiments of the present invention in conjunction with the accompanying drawings.

[0119] The present invention designs a solution intelligent recommendation method based on semantic matching, which includes the following steps:

[0120] Step 1: Build a text embedding model. The model is based on the transformer architecture and contains two layers: the self-attention mechanism and the fully connected layer (FeedForward Neural Network). Use a general data set to train the model and learn the semantic relationship in the text. The specific process is as follows:

[0121] Step S111: Load the general dataset, input the token list after word segmentation into the token embedding layer, segment embedding layer and position embedding layer respectively, and obtain the corresponding embedding representation. In the token embedding layer, add tags such as [CLS] and [SEP] to the word segmentation result, generate attention_mask, and convert all tags into embedding vectors. Two embedding vectors are returned in the segment embedding layer, which can be used to distinguish two sentences input at the same time. In the position embedding layer, encode the order information of the words appearing in the sentence and output the position vector.

[0122] Step S112: Accumulate the above two embedding vectors and position vector as the input vector e of the self-attention module i .

[0123] Step S113: input vector e i With three weight matrices W Q , W K , W V Multiply (the weight matrix is ​​randomly initialized and optimized by backpropagation during training) to generate the query vector Q i , key vector K i Sum value vector V i .

[0124] Q i =e i ·W Q ,

[0125] K i =e i ·W K ,

[0126] V i =ei ·W V ;

[0127] Step S114: Calculate the attention score for each vector ij (the attention score between the i-th word and the j-th key) and the score ij Normalize, that is, divide by (d k K i (key) vector dimension) to prevent the dot product result from being too large, causing the softmax function to enter the saturation region. The formula is as follows:

[0128]

[0129] Step S115: Use the softmax activation function to convert the score into attention weight AttentionWeight ij .

[0130] These weights represent the attention that the word vector at position j receives when generating the output at position i.

[0131]

[0132] Step S116: weighted summation, using the attention weights to perform weighted summation on the value vectors to obtain the output vector of the i-th word unit.

[0133] Output i =AttentionWeight ij ·V j ,

[0134] The above process formula is summarized as follows:

[0135]

[0136] Step S117: In the process of converting word units into vectors, a multi-head attention mechanism is used to focus on different positions in the input sequence at the same time, and the input sequence is divided into multiple heads. Each head performs the above operations independently, and finally the output of all heads is multiplied by the corresponding weight to obtain the output vector Z i .

[0137] Step S118: The obtained word vector is transmitted to the next module, namely the fully connected layer. This module has two layers. The first layer is a ReLU activation function and the second layer is a linear activation function, which can be expressed as:

[0138] FFN(Z i )=max(0,Z iW1+b1)W2+b2,

[0139] The ReLU function in the first layer is a nonlinear function that learns the nonlinear relationship of the input vector and better captures more complex dependencies and patterns in the input sequence. The second layer is a linear layer without an activation function. In this layer, the features after nonlinear transformation are further transformed to generate the final word vector output Z′ i .

[0140] Step S119: Execute the Masked Language Modeling (MLM) pre-training task, use the special tag [MASK] to replace some words in the input sequence, and then linearly transform the word vector obtained in step 1.8 on the linear layer. The formula is as follows:

[0141] O i =WZ′ i +b,

[0142] O i is the output vector, W is the weight matrix, and b is the bias term. The weight matrix and bias term are randomly initialized before training begins, but during the training process, the values ​​of the weight matrix and bias term are adjusted according to the loss function to minimize the loss function. Among them, the output vector O of the linear layer i Each element of corresponds to a word in the vocabulary, and the value of the element represents the probability that the model believes that the word appears in the masked position.

[0143] Step S120: Normalize using the softmax activation function to produce a probability distribution in which each element represents the probability that the corresponding word in the vocabulary appears in the masked position. This probability distribution is the final prediction of the model.

[0144]

[0145] where p i is the probability that the i-th word in the vocabulary appears in the masked position.

[0146] Step S121: Use the cross entropy loss function as the loss function to calculate the difference between the probability distribution obtained in step 1.10 and the true value of the masked word. During the model training process, the goal is to minimize the value of the loss function, and adjust the parameters such as word vectors, weight matrices, and bias terms to make the output probability distribution closer and closer to the true probability distribution. The loss function uses the cross entropy loss function. The formula is as follows:

[0147]

[0148] Among them, y iis a duress encoding (one-hot encoding) vector, indicating whether it is the real target word (the real value of the masked word). At the real target word position, y i The value is 1 if yes, otherwise 0.

[0149] Step 2: Use the T2ranking dataset to evaluate different text embedding models. Through experiments, it can be found that the recommendation effect of the text embedding model used in this patent is better than other models, as shown in Table 1. The specific processing process is as follows:

[0150] Step S211: Use the first 15,000 data of the T2ranking dataset for experimental evaluation.

[0151] Step S212: Evaluate the S-BERT model, the Text2vec model and the model constructed by this patent respectively, and return the map@1, map@10, mrr@1, mrr@10, ndcg@1, ndcg@10 indicators. The calculation formula of each indicator is as follows:

[0152] (1)MAP@k(MeanAverage Precision at k)

[0153] MAP@k refers to the average of the average precision of each query among the first k recommended results returned. Average Precision (AP) is an evaluation metric for a single query that considers the ranking position of the query results in the relevant documents. For each query content, the query formula of AP is as follows:

[0154]

[0155] Where P(n) is the precision of the first n results, which is equal to the number of relevant solutions in the first n results divided by n; rel(n) is a binary function indicating whether the nth result is relevant (1 for relevant, 0 for irrelevant). Recall@k measures the proportion of truly relevant results among the first k returned results.

[0156] MAP@k is the average AP of all queries.

[0157] (2)MRR@k(Mean Reciprocal Rank at k)

[0158] MRR@k refers to the average of the reciprocal values ​​of the ranking of the first relevant result among the first k results. This metric measures how quickly the model finds the first relevant result.

[0159] For each query, the reciprocal rank (RR) is calculated as:

[0160]

[0161] Here, rank is the position of the first relevant result (ranking starts at 1).

[0162] MRR@k is the average of all queries.

[0163] (3)NDCG@k(NormalizedDiscounted Cumulative Gain at k)

[0164] NDCG@k refers to the normalized value of the cumulative gain taking into account the position discount among the top k results. It measures the quality of the list of results returned by the model, taking into account relevance and position.

[0165] 1) For a single query, the calculation formula of DCG (Discounted Cumulative Gain) is:

[0166]

[0167] Where rel(i) is the relevance score of the i-th result (usually a value between 0 and some positive number, such as 0, 1, 2, 3, etc.).

[0168] 2) Calculate the ideal DCG (IDCG) for normalized DCG. Sort the results by relevance score and then calculate the DCG as IDCG k .

[0169] 3) Calculate NDCG:

[0170]

[0171] NDCG@k is the average NDCG of all queries.

[0172] Step S213: Count the values ​​of different indicators of different models in the two parts of data and create a table. The results are shown in Table 1.

[0173] Table 1 Schematic diagram of the recommendation effect of different models on the T2ranking dataset

[0174]

[0175] Step 3: Build migration training data samples and a database of solutions to be queried in a specific field based on the data set of the specific field. The specific processing process is as follows:

[0176] Step S311: construct a migration training data sample based on the specific terms that users are concerned about in the specific domain data set, use the data to perform migration training on the model, and then generate a feature vector that is more in line with the semantics of the specific domain.

[0177] Step S312: Generate a solution library to be queried for a specific field according to the field of concern.

[0178] Step 4: Use the migration training data to perform migration training on the text embedding model, adjust the model-related parameters, generate word vectors for each word, and establish a word vector table corresponding to the word table. The specific processing process is as follows:

[0179] Step S411: Select a loss function. This patent uses a sentence triplet negative softmax contrast loss function, and the calculation process is as follows.

[0180] 1) Calculate the similarity matrix:

[0181]

[0182] Among them, s ij is the similarity matrix, is the i-th specific term embedding, E j is the embedding of the jth sample (either a positive sample or a negative sample). T is the temperature parameter, which is used to scale and balance the impact of positive and negative samples.

[0183] 2) Apply log-softmax:

[0184]

[0185] Among them, p j It is the normalized similarity matrix, which represents the probability distribution of each sample as a positive sample. Assume that there is one positive example and M negative samples for each specific term.

[0186] 3) Calculate the loss (N is the number of positive examples in a batch):

[0187]

[0188] The goal of the loss function is to maximize the probability of positive samples and minimize the probability of negative samples, using the minimized negative log-likelihood function.

[0189] 4) Exchange loss:

[0190] To improve the model performance, the loss function can be calculated by swapping the roles of a specific term with the positive example and using it as a swap loss. This helps the model learn a more symmetric similarity metric.

[0191] Step S412: Start training, load the migration training data samples, let the model learn how to distinguish positive samples from negative samples, generate a loss function and minimize the loss function as the goal, so that the embedding vector generated by the text embedding model is more in line with the specific semantics.

[0192] Step 5: Segment, preprocess and encode the corpus in the specific domain data set, and convert the solution into word unit representation. The specific processing process is as follows:

[0193] Step S511: Initialize the vocabulary. First, divide the input scheme into individual characters according to the characters, and then count the frequency of occurrence of each character to form a primary character-level vocabulary.

[0194] Step S512: Calculate the frequency of adjacent character pairs, and count the frequencies of all adjacent character pairs. An adjacent character pair may be two adjacent characters or a merged character pair.

[0195] Step S513: a merging operation, selecting adjacent character pairs with the highest frequency in the vocabulary, merging them into a new character or vocabulary unit, and then updating the vocabulary and the frequency statistics of the adjacent character pairs.

[0196] Step S514: Repeat the merging operation until the preset vocabulary size is reached or there are no adjacent character pairs that can be merged.

[0197] Step S515: Obtain the final vocabulary. After all the above operations are completed, all the subwords or words obtained constitute the vocabulary. These subwords or words are the word segmentation results of the input solution.

[0198] Step S516: Data preprocessing mainly includes padding, truncation, character replacement, and text cleaning. Padding and truncation are padding short solutions and truncating long solutions according to the preset length of the text embedding model. Character replacement includes special character replacement and unknown word replacement. Text cleaning refers to removing irrelevant characters in the solution.

[0199] Step S517: Encoding is to convert the input sequence into a data form acceptable to the model. The encoding result contains three parts: input_ids, token_type_ids, and attention_mask. Among them, input_ids represents the ID of each word in the word table, token_type_ids is used to distinguish which sentence each word belongs to when two sentences are input at the same time. Attention_mask is used to determine which words need to be paid attention to when generating word vectors.

[0200] Step 6: Use the text embedding model after transfer training to convert each word into a corresponding word vector, then select a pooling strategy to generate a sentence or paragraph vector to build a vector representation library for the query solution library. The specific processing process is as follows:

[0201] Step S611: According to the word segmentation result, the ID of each word is searched in the word-gram table, and an attention_mask is generated (the attention_mask is used to distinguish which words need to be considered and which words do not need to be considered).

[0202] Step S612: According to the ID of each word unit, extract the word vector of each word unit from the text embedding model.

[0203] Step S613: Select a pooling strategy to generate a sentence or paragraph vector. The text embedding model of this patent selects the average pooling strategy. First, all the words that need to be paid attention to are obtained according to the attention_mask, and then the vector obtained by averaging each element in the word vector of all the words that need to be paid attention to is used as the sentence or paragraph vector.

[0204] Step S614: Use a two-dimensional heat map to intuitively display the changes in text vectors before and after transfer training, such as Figure 2-Figure 6 Among them, the 726-dimensional vector is reshaped into a two-dimensional matrix, and the horizontal axis and vertical axis represent the column index and row index after reshaping respectively. Figure 2 is a two-dimensional heat map example of sentence generation; Figure 3 and Figure 4 It is a visualization of the difference between the embedding vector generated by a specific term and the embedding vector generated by its positive example before and after domain adaptive training. From the figure, we can intuitively see that before domain adaptive training, the color of the vector difference heat map is darker and the value of each element is larger, that is, the difference between the vector generated by the specific term and the vector generated by the positive example is large, but after domain adaptive training, the color of the vector difference heat map becomes lighter, the value of each element becomes smaller, and the difference tends to 0, so the specific term becomes similar to the positive example through training. Figure 5 and Figure 6 It is a visualization of the difference between the embedding vector generated by a specific term ("panda") and its negative example (panda) before and after domain adaptive training. As can be seen from the figure, before domain adaptive training, the color of the vector difference heat map is lighter and the element value is smaller, that is, the difference between the vector generated by the specific term and the vector generated by the negative example is small. However, after domain adaptive training, it can be found that the color of the vector difference heat map is darker and the element value becomes larger, so the specific term becomes dissimilar to the positive example through training.

[0205] Step 7: When the user inputs a query solution, the query solution is segmented, preprocessed, encoded, and converted into sentence and paragraph vectors. The specific processing method is the same as steps 5 and 6.

[0206] Step 8: Calculate the cosine similarity between the sentence or paragraph vector of the query solution and all vectors in the query solution library, sort the results, and return the specified number of solutions and similarity scores from high to low according to the cosine similarity score. The specific processing process is as follows:

[0207] Step S811, normalizing the vector, wherein the vector includes the embedding vector of the query solution and the vector generated by the solution in the solution library to be queried.

[0208] Step S812, calculating the vector dot product is the cosine similarity, the formula is as follows:

[0209]

[0210] Among them, A represents the embedding vector of the query solution, and B represents the vector generated by a solution in the solution library to be queried.

[0211] Step S813, sorting the results, and returning a specified number of solutions and similarity scores according to the cosine similarity scores from high to low.

[0212] The present invention provides a method and idea for a method for intelligent recommendation of a solution based on semantic matching. There are many methods and approaches to implement the technical solution. The above is only a preferred implementation of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principle of the present invention. These improvements and modifications should also be regarded as the protection scope of the present invention. All components not specified in this embodiment can be implemented by existing technologies.

Claims

1. A method for intelligently recommending solutions based on semantic matching, comprising the following steps: Step 1: Build a text embedding model based on the transformer architecture, which includes a self-attention mechanism and a fully connected layer. Use a general dataset to train the model and learn the semantic relationship in the text. Step 2: Use public datasets to evaluate the model obtained in step 1; Step 3: construct migration training data and query solutions for specific fields based on the data set of specific fields; Step 4: Use the migration training data obtained in step 3 to perform migration training on the text embedding model, adjust the model parameters, generate a word vector for each word, and establish a word vector table corresponding to the word table; Step 5: segment the corpus in the domain-specific dataset, perform data preprocessing and encoding, and convert the solution into word-unit representation; Step 6: Use the text embedding model after transfer training obtained in step 4 to convert each word into a corresponding word vector, generate sentence or paragraph vectors according to the pooling strategy, and build a vector representation library of the corpus; Step 7: When the user enters a query solution, the query solution is segmented, preprocessed, encoded, and converted into sentence vectors and segment vectors; Step 8, calculate the cosine similarity between the sentence or paragraph vector of the query solution and all vectors in the query solution library, and return a specified number of solutions and similarity scores according to the cosine similarity score.

2. The method for intelligent recommendation of solutions based on semantic matching according to claim 1, characterized in that: Step 1 of building a text embedding model and training it includes: Step S111, load the general data set, input the token list after word segmentation into the token embedding layer, the fragment embedding layer and the position embedding layer respectively, and obtain the corresponding embedding representation; wherein, in the token embedding layer, add tokens to the word segmentation results, generate attention_mask, and convert all tokens into embedding vectors; in the fragment embedding layer, return two embedding vectors; in the position embedding layer, encode the order information of the words appearing in the sentence, and output the position vector; Step S112, accumulating the two embedding vectors and position vector obtained in step S111 as the output vector e of the self-attention module i ; Step S113: input vector e i With three weight matrices W Q , W K , W V The weight matrix is ​​randomly initialized and optimized by back propagation during the training process to generate the query vector Q i , key vector K i Sum value vector V i ; Step S114, calculate the attention score of each input vector ij , the formula is as follows: The score ij is the attention score between the i-th word and the j-th key, d k is the key vector K i Dimensions; Step S115: Convert the score into attention weight through the softmax activation function ij , Among them, Attention Weight ij Indicates the attention that the word vector at the jth position receives when generating the output at the i-th position; Step S116, weighted sum, using attention weight ij The value vector V i Perform weighted summation to obtain the output vector Output of the i-th word i ; Step S117, in the process of converting word units into vectors, a multi-head attention mechanism is used to simultaneously focus on different positions in the input sequence, and the input sequence is divided into multiple heads. Each head performs the above operations independently, and finally the outputs of all heads are multiplied by the corresponding weights to obtain the output vector Z i ; Step S118: Output the word vector Z i It is transmitted to the fully connected layer, which includes two layers. The first layer is the ReLU activation function and the second layer is the linear activation function, which is expressed as: FFN(Z i )=max(0,Z i W1+b1)W2+b2, Wherein W1 and W2 are weight matrices, b1 and b2 are bias items, and the weight matrices W1 and W2 and the bias items b1 and b2 are randomly initialized before starting training and are optimized by back propagation during the training process; The ReLU activation function is a nonlinear function that learns the nonlinear relationship of the input vector and captures the dependencies and patterns in the input sequence. The second layer of linear activation function further transforms the features after nonlinear transformation to generate the final word vector output Z′ i ; Step S119, perform the Masked Language Modeling pre-training task, use the special tag [MASK] to replace some words in the input sequence, and perform the word vector Z′ obtained in S118 on the linear layer. i Linear transformation: ABOUT i =WZ′ i +b, Among them, i is the output vector, W is the weight matrix, b is the bias term, the weight matrix W and the bias term b are randomly initialized before the training begins. During the training process, the values ​​of the weight matrix W and the bias term b are adjusted according to the loss function. The output vector O of the linear layer i Each element of corresponds to a word in the vocabulary, and the value of the element indicates the probability that the model believes that the word appears in the masked position; Step 120, normalize using a softmax activation function to generate a probability distribution, in which each element represents the probability that the corresponding word in the vocabulary appears in the masked position, and the probability distribution is the final prediction of the model. where p i is the probability that the i-th word in the vocabulary appears in the masked position; Step S121, using the cross entropy loss function as the loss function, calculates the difference between the probability distribution obtained in S120 and the true value of the masked word, and adjusts the parameters during the model training process with the goal of minimizing the value of the loss function, so that the output probability distribution is close to the true probability distribution.

3. The method for intelligent solution recommendation based on semantic matching according to claim 1, characterized in that: Step 2 describes the evaluation of the model obtained in step 1, including: Step S211, using part of the public data set to conduct experimental evaluation; Step S212, respectively evaluating the solution model and the comparison model, and calculating the indexes; Step S213, counting the values ​​of different indicators of different models in the two parts of data and establishing a table.

4. The method for intelligent solution recommendation based on semantic matching according to claim 3 is characterized in that: The indicators calculated in step S212 include map@1, map@10, mrr@1, mrr@10, ndcg@1, and ndcg@10 indicators. The calculation formula of each indicator is as follows: (1) MAP@k refers to the average AP of each query among the first k recommended results returned; The average precision AP is an evaluation indicator for a single query, which takes into account the ranking position of the query results in the relevant documents. For each query content, the query formula of AP is as follows: Among them, P(n) is the precision of the first n results, which is the number of relevant solutions in the first n results divided by n; rel(n) indicates whether the nth result is relevant, 1 indicates relevant, and 0 indicates irrelevant; Recall@k measures the proportion of truly relevant results in the first k returned results. MAP@k is the average of the mean average precision (AP) of all queries; (2) MRR@k refers to the average of the reciprocal values ​​of the ranking of the first relevant result among the first k results; For each query, the calculation formula of reciprocal rank RR is: Among them, rank is the position of the first relevant result, starting from 1; MRR@k is the average of the reciprocal RR of all queries; (3) NDCG@k refers to the normalized value of the cumulative gain taking into account the position discount in the first k results; 1) For a single query, the calculation formula of discounted cumulative gain DCG is: Where rel(i) is the relevance score of the i-th result, and the relevance score is usually a value between 0 and a positive number; 2) Calculate the ideal DCG, i.e. IDCG, to normalize the DCG, sort the results by relevance score and calculate the DCG as IDCG k ; 3) Calculate the normalized discounted cumulative gain NDCG: NDCG@k is the average NDCG of all queries.

5. The method for intelligently recommending solutions based on semantic matching according to claim 1, characterized in that: The migration training of the text embedding model in step 4 includes: Step S411, selecting a loss function and calculating the loss; Step S412, start training, load the migration training data samples, let the model learn how to distinguish positive samples from negative samples, generate a loss function and minimize the loss function as the goal, so that the embedding vector generated by the text embedding model is more in line with the specific semantics.

6. The method for intelligently recommending solutions based on semantic matching according to claim 5, characterized in that: The loss function of step S411 is the sentence triple negative softmax contrast loss function, and the calculation process is as follows: Step S4111, calculate the similarity matrix: Among them, s ij is the similarity matrix, is the i-th specific term embedding, E j is the jth sample embedding, T is the temperature parameter, and the temperature parameter T is used to scale and balance the influence of positive and negative samples; Step S4112, apply log-softmax for normalization: Among them, p j is the normalized similarity matrix, which represents the probability distribution of each sample as a positive sample, assuming that there is one positive example and M negative samples for each specific term; Step S4113, calculate the loss L, where N is the number of positive examples in a batch: The goal of the loss function is to maximize the probability of positive samples and minimize the probability of negative samples, using the minimized negative log-likelihood function; Step S4114, exchange loss, calculates the loss function by exchanging the roles of the specific term and the positive example, and uses it as the exchange loss.

7. The method for intelligently recommending solutions based on semantic matching according to claim 1, characterized in that: The step 5 performs word segmentation, data preprocessing and encoding on the corpus in the specific domain data set, and converts the scheme into word unit representation, including: Step S511, initializing the vocabulary, dividing the input scheme into characters, counting the frequency of each character, and obtaining a primary character-level vocabulary; Step S512, calculating the frequency of adjacent character pairs, and counting the frequencies of all adjacent character pairs, wherein the adjacent character pairs are two adjacent characters or a merged character pair; Step S513, a merging operation, selecting the adjacent character pairs with the highest frequency in the vocabulary, merging them into a new character or vocabulary unit, and updating the vocabulary and the frequency statistics of the adjacent character pairs; Step S514, repeating the merging operation until a preset vocabulary size is reached or there are no adjacent character pairs that can be merged; Step S515, obtaining a final vocabulary. After all the above operations are completed, all the subwords or words obtained constitute a vocabulary, and the subwords or words are the word segmentation results of the input scheme; Step S516, data preprocessing, filling short solutions according to the preset length of the text embedding model, truncating long solutions; replacing special characters and unknown words; removing irrelevant characters in the solution; Step S517, encoding is to convert the input sequence into a data form acceptable to the model. The encoding result includes three parts: input_ids, token_type_ids and attention_mask; among them, input_ids represents the ID of each word in the word table, token_type_ids is used to distinguish which sentence each word belongs to when two sentences are input at the same time, and attention_mask is used to determine which words need to be paid attention to when generating word vectors.

8. The method for intelligently recommending solutions based on semantic matching according to claim 1, characterized in that: The step 6 generates sentence vectors and segment vectors using the text embedding model after transfer training, and constructs a vector representation library of the query solution library, including: Step S611, according to the word segmentation result, look up the ID of each word in the word-gram table and generate an attention_mask to determine the word-grams that need to be considered; Step S612, extracting the word vector of each word from the text embedding model according to the ID of each word; Step S613, generating a sentence or paragraph vector according to a pooling strategy; Step S614, using a two-dimensional heat map to intuitively display the changes in the text vector before and after the transfer training.

9. The method for intelligently recommending solutions based on semantic matching according to claim 8, characterized in that: The pooling strategy in step S613 is to select an average pooling strategy, obtain all words that need to be paid attention to according to attention_mask, and take the average of each element in the word vector of all the word elements that need to be paid attention to as the sentence or paragraph vector.

10. The method for intelligently recommending solutions based on semantic matching according to claim 1, characterized in that: The step 8 comprises: Step S811, normalizing the vector, wherein the vector includes the embedding vector of the query solution and the vector generated by the solution in the solution library to be queried; Step S812, calculating the vector dot product is the cosine similarity, the formula is as follows: Among them, A represents the embedding vector of the query solution, and B represents the vector generated by a solution in the solution library to be queried; Step S813, sorting the results, and returning a specified number of solutions and similarity scores according to the cosine similarity scores from high to low.

Citation Information

Patent Citations

  • Class case recommendation method based on text content

    CN110442684A

  • Long text retrieval model based on comparative learning

    CN114201581A

  • NLP-based industry data analysis method and system

    CN118070812A

  • Instruction recognition method and device, training method, and computer readable storage medium

    WO2023231676A1

Cited By

  • Retrieval enhancement method and system for multi-round dialogue type questions and answers and application

    CN121029952A

  • Semantic recommendation method fusing field self-adaption and user preference weighting

    CN121052259A