Candidate reply statement scoring method and device based on multi-level features and double coding

Through the multi-level feature and dual-coded candidate reply statement scoring method, the problems of manual annotation data dependence and length difference are solved, and more efficient semantic information capture and candidate reply statement sorting accuracy are achieved.

CN120336868APending Publication Date: 2025-07-18XIAMEN KUAISHANGTONG TECH CORP LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510297544.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-13
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

Existing candidate reply statement scoring models rely on manual labeling data to be costly, poor labeling consistency, and difficult to capture subtle differences when dealing with the above-mentioned differences in lengths between candidate reply statements, resulting in limited model performance and robustness.

Method used

The multi-level feature and dual encoding method is used to independently encode the dialogue history and candidate reply statements, and the word-level, phrase-level and sentence-level feature extraction is used, combined with comprehensive similarity calculation, and a loss function optimization model based on comparison learning is used to ensure semantic information integrity and sorting accuracy.

Benefits of technology

It significantly improves the quality and coverage of the training data, avoids the loss of semantic information, and improves the accuracy and robustness of the model's sorting in candidate reply statements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120336868A_ABST
    Figure CN120336868A_ABST
Patent Text Reader

Abstract

The invention discloses a candidate reply statement scoring method and device based on multi-level features and dual coding. The method comprises the steps that a training data set is constructed by using a large language model; constructing a candidate reply statement scoring model and training the candidate reply statement scoring model by adopting the training data set and a loss function based on comparative learning to obtain a trained candidate reply statement scoring model; the candidate reply statement scoring model comprises a word-level feature extraction module, a word-level similarity calculation module, a phrase-level feature extraction module, a phrase-level similarity calculation module, a sentence-level feature extraction module, a sentence-level similarity calculation module and a comprehensive similarity calculation module. The dialogue history and the candidate reply statements are coded through double coding and then input into a trained candidate reply statement scoring model, the word-level similarity, the phrase-level similarity and the sentence-level similarity are calculated, and the comprehensive similarity is calculated; and the candidate reply statements are sorted according to the comprehensive similarity, so that the sorting accuracy is remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of dialogue generation, and particularly to a method and device for scoring candidate response sentences based on multi-level features and dual encoding. Background Art

[0002] A candidate response sentence scoring model is a model used to evaluate and optimize the quality of generated text in the field of natural language processing (NLP). This model sorts multiple candidate response sentences and assigns a reward score to each candidate response sentence according to a predefined reward function or manually annotated preferences. By training the candidate response sentence scoring model, it can learn how to distinguish high-quality and low-quality generated text, so as to guide the model to generate higher-quality sentences that meet expectations in subsequent generation tasks.

[0003] Currently, the training data of the candidate response sentence scoring model mainly relies on manual annotation. In the data processing process, the method usually adopted is to splice the above text with each candidate response sentence as the input and call the encoding model for training. Specifically, the model splices the above text and the candidate response sentence into an overall input, uses the encoding model to encode the spliced text, and then generates the reward score of each candidate response sentence. This method can capture context information to a certain extent and distinguish the quality of different candidate response sentences.

[0004] Although the existing candidate response sentence scoring models have certain advantages in theory, there are still many deficiencies in practical applications. First, the cost of manually annotated data is high, resulting in a limited amount of training data, and the annotation effect is easily affected by subjective factors and external conditions, making it difficult to ensure consistency and objectivity. Second, in actual dialogue scenarios, the above text is often long while the candidate response sentences are short. Direct splicing may cause the encoding model to be difficult to capture the subtle differences between candidate response sentences. Especially in different rounds of dialogue, the length difference between the above text and the candidate response sentences further exacerbates this problem. In addition, the length differences between different texts are large. If the set max_length is too long, the shorter text part will be overfilled, wasting computing resources; if it is set too short, the candidate response sentences may be truncated, losing key information. These limitations restrict the performance and robustness of the existing candidate response sentence scoring models in practical applications and further improvement is urgently needed. Summary of the Invention

[0005] The purpose of this application is to propose a method and device for scoring candidate response sentences based on multi-level features and dual encoding for the above-mentioned technical problems.

[0006] In a first aspect, the present invention provides a method for scoring candidate response sentences based on multi-level features and dual encoding, including the following steps:

[0007] Obtain the conversation history and its corresponding several candidate response sentences, input the conversation history, all the candidate response sentences, and the prompt words for scoring each candidate response sentence according to the conversation history and all candidate response sentences into a large language model to obtain the true score corresponding to each candidate response sentence, and construct a training dataset based on the conversation history, its corresponding several candidate response sentences, and the true scores.

[0008] Construct a candidate response sentence scoring model and train it using the training dataset to obtain a trained candidate response sentence scoring model. The candidate response sentence scoring model includes a word-level feature extraction module, a word-level similarity calculation module, a phrase-level feature extraction module, a phrase-level similarity calculation module, a sentence-level feature extraction module, a sentence-level similarity calculation module, and a comprehensive similarity calculation module.

[0009] Obtain the conversation history to be scored and its corresponding several candidate response sentences, encode the conversation history to be scored and each candidate response sentence respectively to obtain the semantic representation of the conversation history and the semantic representation of each candidate response sentence; input the semantic representation of the conversation history and the semantic representation of each candidate response sentence into the trained candidate response sentence scoring model. The semantic representation of the conversation history and the semantic representation of each candidate response sentence respectively pass through the word-level feature extraction module to extract the word-level feature vectors of each word in the conversation history and the word-level feature vectors of each word in the candidate response sentence. The word-level feature vectors of each word in the conversation history and the word-level feature vectors of each word in the candidate response sentence are input into the word-level similarity calculation module to calculate the word-level similarity; the word-level feature vectors of each word in the conversation history and the word-level feature vectors of each word in the candidate response sentence respectively pass through the phrase-level feature extraction module to extract the phrase-level feature vectors of the conversation history and the phrase-level feature vectors of the candidate response sentence. The phrase-level feature vectors of the conversation history and the phrase-level feature vectors of the candidate response sentence are input into the phrase-level similarity calculation module to calculate the phrase-level similarity; input the [CLS] token representations of the semantic representation of the conversation history and the [CLS] token representations of the semantic representation of each candidate response sentence into the sentence-level feature extraction module to extract the sentence-level feature vectors of the conversation history and the sentence-level feature vectors of the candidate response sentence. The sentence-level feature vectors of the conversation history and the sentence-level feature vectors of the candidate response sentence are input into the sentence-level similarity calculation module to calculate the sentence-level similarity. The word-level similarity, phrase-level similarity, and sentence-level similarity are input into the comprehensive similarity calculation module for weighted calculation to obtain the predicted score of each candidate response sentence.

[0010] Preferably, the word-level feature extraction module includes a convolutional layer with a ReLU activation function and a multi-head self-attention layer connected in sequence. The words at each word position in the dialogue history and the words at each word position in the candidate response sentence are respectively passed through the convolutional layer with a ReLU activation function and the multi-head self-attention layer in sequence, and the word-level feature vectors at each word position in the dialogue history and the word-level feature vectors at each word position in the candidate response sentence are extracted. In the word-level similarity calculation module, first calculate the cosine similarity between the word-level feature vectors at each word position in the dialogue history and the word-level feature vectors at the aligned word positions in the candidate response sentence, and then take the average of the cosine similarities of the word-level feature vectors at all word positions to obtain the word-level similarity.

[0011] Preferably, the phrase-level feature extraction module includes a max pooling layer and an average pooling layer. The word-level feature vectors at all word positions in the dialogue history are respectively input into the max pooling layer and the average pooling layer to obtain the max pooling result and the average pooling result corresponding to the dialogue history. Add the max pooling result and the average pooling result corresponding to the dialogue history to obtain the phrase-level feature vector of the dialogue history; the word-level feature vectors at all word positions in the candidate response sentence are respectively input into the max pooling layer and the average pooling layer to obtain the max pooling result and the average pooling result at each phrase position in the candidate response sentence, and add the max pooling result and the average pooling result corresponding to the candidate response sentence to obtain the phrase-level feature vector of the candidate response sentence; in the phrase-level similarity calculation module, calculate the cosine similarity between the phrase-level feature vector of the dialogue history and the phrase-level feature vector of the candidate response sentence to obtain the phrase-level similarity.

[0012] Preferably, the sentence-level feature extraction module includes a fully connected layer. The [CLS] token representation in the semantic representation of the dialogue history and the [CLS] token representation of the semantic representation of each candidate response sentence are respectively passed through the fully connected layer to extract the sentence-level feature vector of the dialogue history and the sentence-level feature vector of the candidate response sentence. In the sentence-level similarity calculation module, calculate the cosine similarity between the sentence-level feature vector in the dialogue history and the sentence-level feature vector of the candidate response sentence to obtain the sentence-level similarity; the calculation formula of the comprehensive similarity calculation module is as follows:

[0013] s = w1·W + w2·P + w3·J;

[0014] where w1, w2, and w3 respectively represent the weights of the word-level similarity, phrase-level similarity, and sentence-level similarity, and s, W, P, and J respectively represent the comprehensive similarity, word-level similarity, phrase-level similarity, and sentence-level similarity. The comprehensive similarity is used as the prediction score of the candidate response sentence.

[0015] Preferably, a loss function based on contrastive learning is adopted during the training process of the candidate response sentence scoring model, as shown in the following formula:

[0016]

[0017] Among them, L represents the loss function based on contrastive learning, and y ij represents the true score of the j-th candidate response sentence in the i-th sample of the training data set, and y ik represents the true score of the k-th candidate response sentence in the i-th sample of the training data set. s ij represents the comprehensive similarity between the j-th candidate response sentence and the conversation history in the i-th sample of the training data set, and s ik represents the comprehensive similarity between the k-th candidate response sentence and the conversation history in the i-th sample of the training data set. N represents the total number of samples in the training data set, M represents the total number of candidate response sentences in the i-th sample, τ is the temperature parameter, and I(·) is the indicator function, which takes the value of 1 when the condition in the parentheses is true, otherwise 0.

[0018] Preferably, in the encoding process of the conversation history to be scored and each candidate response sentence, a pre-trained BERT model is used, and the maximum length is set to truncate the conversation history to be scored or each candidate response sentence from right to left, and a tokenizer is used for padding. The output features of the last hidden layer in the pre-trained BERT model are used as the semantic representation of the conversation history or the semantic representation of each candidate response sentence.

[0019] In a second aspect, the present invention provides a candidate response sentence scoring device based on multi-level features and dual encoding, including:

[0020] A training data set construction module, configured to obtain a conversation history and its corresponding several candidate response sentences, input the conversation history, all candidate response sentences, and a prompt word for scoring each candidate response sentence according to the conversation history and all candidate response sentences into a large language model to obtain the true score corresponding to each candidate response sentence, and construct a training data set according to the conversation history, its corresponding several candidate response sentences, and the true scores;

[0021] A model construction module, configured to construct a candidate response sentence scoring model and train it with the training data set to obtain a trained candidate response sentence scoring model. The candidate response sentence scoring model includes a word-level feature extraction module, a word-level similarity calculation module, a phrase-level feature extraction module, a phrase-level similarity calculation module, a sentence-level feature extraction module, a sentence-level similarity calculation module, and a comprehensive similarity calculation module;

[0022] The scoring module is configured to obtain the conversation history to be scored and its corresponding several candidate response sentences, encode the conversation history to be scored and each candidate response sentence respectively to obtain the semantic representation of the conversation history and the semantic representation of each candidate response sentence; input the semantic representation of the conversation history and the semantic representation of each candidate response sentence into the trained candidate response sentence scoring model. The semantic representation of the conversation history and the semantic representation of each candidate response sentence respectively pass through the word-level feature extraction module to extract the word-level feature vector of each word in the conversation history and the word-level feature vector of each word in the candidate response sentence. The word-level feature vector of each word in the conversation history and the word-level feature vector of each word in the candidate response sentence are input into the word-level similarity calculation module to calculate the word-level similarity; the word-level feature vector of each word in the conversation history and the word-level feature vector of each word in the candidate response sentence respectively pass through the phrase-level feature extraction module to extract the phrase-level feature vector of the conversation history and the phrase-level feature vector of the candidate response sentence. The phrase-level feature vector of the conversation history and the phrase-level feature vector of the candidate response sentence are input into the phrase-level similarity calculation module to calculate the phrase-level similarity; input the [CLS] token representation in the semantic representation of the conversation history and the [CLS] token representation of the semantic representation of each candidate response sentence into the sentence-level feature extraction module to extract the sentence-level feature vector of the conversation history and the sentence-level feature vector of the candidate response sentence. The sentence-level feature vector of the conversation history and the sentence-level feature vector of the candidate response sentence are input into the sentence-level similarity calculation module to calculate the sentence-level similarity. The word-level similarity, phrase-level similarity and sentence-level similarity are input into the comprehensive similarity calculation module for weighted calculation to obtain the predicted score of each candidate response sentence.

[0023] In a third aspect, the present invention provides an electronic device, including one or more processors; a storage device for storing one or more programs, which when executed by the one or more processors, cause the one or more processors to implement the method described in any implementation manner of the first aspect.

[0024] In a fourth aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the method described in any implementation manner of the first aspect.

[0025] In a fifth aspect, the present invention provides a computer program product, including a computer program, which when executed by a processor, implements the method described in any implementation manner of the first aspect.

[0026] Compared with the prior art, the present invention has the following beneficial effects:

[0027] (1) The candidate response sentence scoring method based on multi-level features and dual coding proposed by the present invention solves the problems of data dependence and annotation consistency by combining the construction of training data and manually annotated data, significantly improving the quality of training data, ensuring the diversity and coverage of training data, and overcoming the deficiencies of existing scoring models in aspects such as data dependence, annotation consistency, input splicing method, and text length processing.

[0028] (2) The candidate response sentence scoring method based on multi-level features and dual coding proposed by the present invention adopts the method of dual coding, independently encoding the dialogue history and candidate response sentences respectively, ensuring the integrity of semantic information and avoiding the problem of semantic information loss caused by traditional splicing methods.

[0029] (3) The candidate response sentence scoring method based on multi-level features and dual coding proposed by the present invention introduces a comprehensive similarity calculation method. Through multi-level feature extraction and similarity calculation, the sorting accuracy of the model is further improved. And a loss function based on contrast learning is adopted. By pairwise comparing the sorting labels of candidate response sentences, the sorting ability of the model is optimized, significantly improving the sorting accuracy of candidate response sentences. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0031] Figure 1 It is a schematic flowchart of the candidate response sentence scoring method based on multi-level features and dual coding for the embodiments of the present application;

[0032] Figure 2 It is a schematic diagram of the candidate response sentence scoring device based on multi-level features and dual coding for the embodiments of the present application;

[0033] Figure 3 It is a schematic hardware structure diagram of the electronic device provided by the embodiments of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0034] In order to make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of them. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.

[0035] Figure 1 A candidate response statement scoring method based on multi-level features and dual coding provided by an embodiment of the present application is shown, including the following steps:

[0036] S1. Obtain the conversation history and its corresponding several candidate response statements, input the conversation history, all the candidate response statements, and the prompt words for scoring each candidate response statement according to the conversation history and all the candidate response statements into a large language model to obtain the true score corresponding to each candidate response statement, and construct a training data set according to the conversation history and its corresponding several candidate response statements and the true scores.

[0037] Specifically, the embodiment of the present application uses a large language model to construct a training data set. By constructing prompt words, the large language model is called to generate true scores, and the generated true scores are compared with the finely manually annotated score data to calculate the accuracy rate, and the quality of the prompt words is continuously optimized according to the accuracy rate. In one example, for a data set with a total of 100 samples, it is judged whether the candidate response statements ranked first and last according to the true scores in each sample are consistent with the candidate response statements ranked first and last according to the manually annotated score data. If they are consistent, the score of the sample is 1; if they are inconsistent, the score of the sample is 0. The total score in 100 samples is statistically calculated, and the total score is divided by the total number of samples to obtain the accuracy rate. The prompt words can be continuously adjusted until the accuracy rate reaches the threshold, and the prompt words for scoring each candidate response statement according to the conversation history and all the candidate response statements can be obtained. Inputting the conversation history and its corresponding several candidate response statements and the prompt words for scoring each candidate response statement according to the conversation history and all the candidate response statements into the large language model can generate the true scores corresponding to each candidate response statement. In one example, the large language model in the present application can adopt the pre-trained Qwen 2.5 - 72B model. In other examples, other suitable large language models can also be adopted, which are not limited herein.

[0038] The generated training data set contains the following fields: history represents the conversation history, that is, the context above the candidate response statement; candidates represents the list of candidate response statements, including multiple candidate response statements; response represents the ranking corresponding to each candidate response statement, where a ranking of 1 represents the candidate response statement corresponding to the highest true score, and so on.

[0039] In one example, the conversation history is:

[0040] "Visitor: Can I eat barbecue within a month after having polyps?

[0041] Customer service: Hello, may I ask what I can do for you?

[0042] Visitor: Is it okay to have barbecue a month after polyps removal?

[0043] Customer service: It is not recommended to have barbecue. Do you have any discomfort now?

[0044] Visitor: My stomach is a bit uncomfortable. What should I do?

[0045] The list of candidate response statements is: ["Well, then you need to understand your own situation. We can take remedial measures or treatment. <sep>Which hospital did you go to before?”, “It depends. You can take anti-inflammatory drugs or observe further <sep>What's mainly bothering you now?","Polyps can be treated with minimally invasive surgery <sep>Which location of your polyp is it? ", "Let me know first. If you feel uncomfortable, you can choose to have a check. The treatment methods for each disease are different <sep>Which region are you in? There are various treatment methods on our side. <sep>How long has it been since you had polyps removed?”, the sorting result of the output candidate response statements is: [1, 4, 2, 5, 3], and the true scores given in the order of each candidate response statement are [1, 3, 5, 2, 4]. In this way, 40,000 pieces of training data are generated. Each piece of training data includes the context, the list of candidate response statements, and the true scores of each candidate response statement, and finally constitutes a training data set.

[0046] S2. Construct a candidate response statement scoring model and train it using the training data set to obtain a trained candidate response statement scoring model. The candidate response statement scoring model includes a word-level feature extraction module, a word-level similarity calculation module, a phrase-level feature extraction module, a phrase-level similarity calculation module, a sentence-level feature extraction module, a sentence-level similarity calculation module, and a comprehensive similarity calculation module.

[0047] In a specific embodiment, the candidate response statement scoring model uses a contrastive learning-based loss function during training, as shown in the following formula:

[0048]

[0049] where L represents the contrastive learning-based loss function, y ij represents the true score of the j-th candidate response statement in the i-th sample of the training data set, y ik represents the true score of the k-th candidate response statement in the i-th sample of the training data set, s ij represents the comprehensive similarity between the j-th candidate response statement and the conversation history in the i-th sample of the training data set, s ik represents the comprehensive similarity between the k-th candidate response statement and the conversation history in the i-th sample of the training data set, N represents the total number of samples in the training data set, M represents the total number of candidate response statements in the i-th sample, τ is the temperature parameter, and I(·) is the indicator function, which takes the value of 1 when the condition in the parentheses is true, otherwise 0.

[0050] Specifically, the embodiment of the present application constructs a candidate response statement scoring model for scoring candidate response statements to obtain the candidate response statement ranked first and output it as the response statement. The embodiments of the present application independently encode the conversation history and candidate response statements respectively, introduce comprehensive similarity for multi-level feature extraction, introduce a contrastive learning-based loss function, and calculate the accuracy rate, which significantly improves the performance of the candidate response statement scoring model. The specific module structure and calculation process of the candidate response statement scoring model will be described in detail later.

[0051] The embodiment of the present application uses a contrastive learning-based loss function to further identify the differences between different candidate response statements and conform to human preferences.

[0052] Specifically, the true scores corresponding to each candidate response statement are compared pairwise, and the model is optimized through contrastive learning. The loss function based on contrastive learning proposed in the embodiments of the present application is obtained by improving on the Infonce loss function.

[0053] In one example, the candidate response statements are first adjusted. [1, 4, 2, 5, 3] represents the priority ranking of the candidate response statements. For example, if candidate response statement 1 is ranked first, its corresponding true score is 1; candidate response statement 4 is ranked second, and its corresponding true score is 2; candidate response statement 2 is ranked third, and its corresponding true score is 3; candidate response statement 5 is ranked fourth, and its corresponding true score is 4; candidate response statement 3 is ranked fifth, and the true score is 5. Converted according to the order of the candidate response statements, [1, 3, 5, 2, 4] represents the true scores of each candidate response statement. The loss calculation process is as follows: Compare candidate response statement 1 and candidate response statement 2, y 11 = 1, y 12 = 3, y 11 <y 12 , calculate the positive sample loss; compare candidate response statement 1 and candidate response statement 3, y 11 = 1, y 13 = 5, y 11 <y 13 , calculate the positive sample loss; compare candidate response statement 2 and candidate response statement 4, y 12 = 3, y 14 = 5, y 12 >y 14 , calculate the negative sample loss. In this way, the model can learn the relative ranking relationship between candidate response statements, thereby improving the ranking accuracy of the candidate response statement scoring model.

[0054] After the candidate response statement scoring model is trained, the manually annotated data is tested to calculate the accuracy of the model. Specifically, the inference result of the model is compared with the manually annotated result, and the top3 accuracy is calculated, that is, whether the best candidate response statement output by the model is among the top 3 best candidate response statements in the manual scoring.

[0055] In one example, for the above conversation, assume that the ranking output by the model is [1, 4, 2, 5, 3], the best candidate response statement is 1, and the manually annotated ranking is [4, 1, 2, 5, 3]. 1 is among the top three candidate response statements in the manual ranking, so this candidate response statement is correctly classified. In this way, the effectiveness of the trained candidate response statement scoring model in the embodiments of the present application is verified.

[0056] The experimental results show that the candidate response sentence scoring method based on multi-level features and dual coding proposed in the embodiments of the present application is significantly superior to the traditional method in terms of accuracy. Specifically, in the manually annotated data of 6,441 sentences, the top-3 accuracy of the traditional method is 55%, while the top-3 accuracy of the present invention reaches 73.4%. This result verifies the effectiveness of the present invention in improving the performance of the candidate response sentence scoring model.

[0057] S3. Obtain the conversation history to be scored and its corresponding several candidate response sentences, respectively encode the conversation history to be scored and each candidate response sentence to obtain the semantic representation of the conversation history and the semantic representation of each candidate response sentence; input the semantic representation of the conversation history and the semantic representation of each candidate response sentence into the trained candidate response sentence scoring model. The semantic representation of the conversation history and the semantic representation of each candidate response sentence respectively pass through the word-level feature extraction module to extract the word-level feature vector of each word in the conversation history and the word-level feature vector of each word in the candidate response sentence. The word-level feature vector of each word in the conversation history and the word-level feature vector of each word in the candidate response sentence are input into the word-level similarity calculation module to calculate the word-level similarity; the word-level feature vector of each word in the conversation history and the word-level feature vector of each word in the candidate response sentence respectively pass through the phrase-level feature extraction module to extract the phrase-level feature vector of the conversation history and the phrase-level feature vector of the candidate response sentence. The phrase-level feature vector of the conversation history and the phrase-level feature vector of the candidate response sentence are input into the phrase-level similarity calculation module to calculate the phrase-level similarity; input the [CLS] token representation in the semantic representation of the conversation history and the [CLS] token representation of the semantic representation of each candidate response sentence into the sentence-level feature extraction module to extract the sentence-level feature vector of the conversation history and the sentence-level feature vector of the candidate response sentence. The sentence-level feature vector of the conversation history and the sentence-level feature vector of the candidate response sentence are input into the sentence-level similarity calculation module to calculate the sentence-level similarity. The word-level similarity, phrase-level similarity and sentence-level similarity are input into the comprehensive similarity calculation module for weighted calculation to obtain the predicted score of each candidate response sentence.

[0058] In a specific embodiment, the pre-trained BERT model is used in the encoding process of the conversation history to be scored and each candidate response sentence, and the maximum length is set to truncate the conversation history to be scored or each candidate response sentence from right to left, and the tokenizer is used for padding. The output features of the last hidden layer in the pre-trained BERT model are used as the semantic representation of the conversation history or the semantic representation of each candidate response sentence.

[0059] Specifically, the trained candidate response sentence scoring model is deployed. In the inference stage, first, the input conversation history is organized into a string form and encoded using the pre-trained BERT model. To ensure that the information in the conversation history can be completely encoded and avoid wasting computing resources due to overly long conversation history, the present invention sets a maximum length max_length = 128 to truncate and automatically pad the conversation history. To prevent loss of semantic information, it is set to truncate from right to left and use a tokenizer for padding to ensure that the text length meets the requirements. The semantic representation of the conversation history is extracted through the hidden layer of the last layer of the pre-trained BERT model and mapped to a low-dimensional space for subsequent similarity calculation. Secondly, the input candidate response sentences are retained in list form, and each candidate response sentence is separately encoded using another independent pre-trained BERT model to ensure their independent representation in the semantic space. The remaining processing is the same as that of the conversation history for subsequent similarity calculation. In one example, the pre-trained BERT model mentioned in the embodiments of the present application uses the bert-base-chinese model. In other embodiments, other suitable encoding models can be used.

[0060] In one example, a tokenizer is used to pad the truncated conversation history to ensure that the text length is 128. The padded conversation history is as follows: "[PAD][PAD]... Visitor: Is it okay to have barbecue a month after having polyps removed? Customer service: It is not recommended to have barbecue. Do you have any discomfort now? Visitor: My stomach is a bit uncomfortable. What should I do?" The candidate response sentences are separately encoded as: [semantic representation of candidate response sentence 1, semantic representation of candidate response sentence 2, semantic representation of candidate response sentence 3, semantic representation of candidate response sentence 4, semantic representation of candidate response sentence 5].

[0061] In a specific embodiment, the word-level feature extraction module includes a convolutional layer with a ReLU activation function and a multi-head self-attention layer connected in sequence. The words at each word position in the conversation history and the words at each word position in the candidate response sentences respectively pass through the convolutional layer with a ReLU activation function and the multi-head self-attention layer in sequence to extract the word-level feature vectors at each word position in the conversation history and the word-level feature vectors at each word position in the candidate response sentences. In the word-level similarity calculation module, first, the cosine similarity between the word-level feature vectors at each word position in the conversation history and the aligned word-level feature vectors at the word positions in the candidate response sentences is calculated, and then the average value of the cosine similarities of all word-level feature vectors at the word positions is taken to obtain the word-level similarity.

[0062] In a specific embodiment, the phrase-level feature extraction module includes a max pooling layer and an average pooling layer. The word-level feature vectors at all word positions in the dialogue history are respectively input into the max pooling layer and the average pooling layer to obtain the max pooling result and the average pooling result corresponding to the dialogue history. The max pooling result and the average pooling result corresponding to the dialogue history are added together to obtain the phrase-level feature vector of the dialogue history. The word-level feature vectors at all word positions in the candidate response sentence are respectively input into the max pooling layer and the average pooling layer to obtain the max pooling result and the average pooling result at each phrase position in the candidate response sentence. The max pooling result and the average pooling result corresponding to the candidate response sentence are added together to obtain the phrase-level feature vector of the candidate response sentence. In the phrase-level similarity calculation module, the cosine similarity between the phrase-level feature vector of the dialogue history and the phrase-level feature vector of the candidate response sentence is calculated to obtain the phrase-level similarity.

[0063] In a specific embodiment, the sentence-level feature extraction module includes a fully connected layer. The [CLS] token representation in the semantic representation of the dialogue history and the [CLS] token representation of the semantic representation of each candidate response sentence respectively pass through the fully connected layer to extract the sentence-level feature vector of the dialogue history and the sentence-level feature vector of the candidate response sentence. In the sentence-level similarity calculation module, the cosine similarity between the sentence-level feature vector in the dialogue history and the sentence-level feature vector of the candidate response sentence is calculated to obtain the sentence-level similarity. The calculation formula of the comprehensive similarity calculation module is as follows:

[0064] s = w1·W + w2·P + w3·J;

[0065] Wherein, w1, w2, and w3 respectively represent the weights of the word-level similarity, the phrase-level similarity, and the sentence-level similarity, and s, W, P, and J respectively represent the comprehensive similarity, the word-level similarity, the phrase-level similarity, and the sentence-level similarity. The comprehensive similarity is used as the prediction score of the candidate response sentence.

[0066] To further improve the ranking ability of the model, the embodiment of the present application introduces a comprehensive similarity calculation method. This method comprehensively considers the word-level similarity, the phrase-level similarity, and the sentence-level similarity through multi-level feature extraction and similarity calculation, thereby improving the ranking accuracy of the model.

[0067] Specifically, the present invention adopts the following hierarchical comprehensive similarity calculation method:

[0068] First, calculate the word-level similarity between the dialogue history and the candidate response sentence through the word-level feature extraction module and the word-level similarity calculation module. In the word-level feature extraction module, a one-dimensional convolutional layer with a ReLU activation function is used to extract local features, and then a multi-head self-attention layer is used to capture global dependencies, obtaining the word-level feature vectors at each word position in the dialogue history and the word-level feature vectors at each word position in the candidate response sentence. In the word-level similarity calculation module, the cosine similarity is used to calculate the similarity between the dialogue history and the candidate response sentence at each word position, and the average of the similarities at all word positions is taken to obtain the word-level similarity W. Since both the dialogue history and the candidate response sentence have been truncated, the lengths of the dialogue history and the candidate response sentence are actually the same, both being the maximum length max_length. Therefore, when calculating the word-level similarity, the words at the aligned positions in the dialogue history and the candidate response sentence can be found for calculation.

[0069] Second, calculate the phrase-level similarity between the dialogue history and the candidate response sentence through the phrase-level feature extraction module and the phrase-level similarity calculation module. The phrase-level feature extraction module performs max pooling and average pooling on the word-level feature vectors of the dialogue history and each candidate response sentence respectively, compresses seq_len to 1, removes redundant dimensions, and then adds the results of max pooling and average pooling to obtain the phrase-level feature vector. Further, in the phrase-level similarity calculation module, the cosine similarity is used to calculate the similarity between the dialogue history and the candidate response sentence at the phrase level, and the phrase-level similarity P can be obtained.

[0070] Third, calculate the sentence-level similarity between the dialogue history and the candidate response sentence through the sentence-level feature extraction module and the sentence-level similarity calculation module. The fully connected layer in the sentence-level feature extraction module is used to perform a linear transformation on the [CLS] token representation in the semantic representation of the dialogue history and the [CLS] token representation in the semantic representation of each candidate response sentence respectively. Further, in the sentence-level similarity calculation module, the cosine similarity is used to calculate the similarity between the dialogue history and the candidate response sentence at the sentence level, and the sentence-level similarity J can be obtained.

[0071] Finally, the word-level similarity, phrase-level similarity, and sentence-level similarity are weighted and summed to obtain the comprehensive similarity.

[0072] Based on the comprehensive similarity, several candidate response sentences are sorted from largest to smallest, and the candidate response sentence with the largest comprehensive similarity is output as the best candidate response sentence.

[0073] For further reference Figure 2 , as an implementation of the methods shown in the above figures, this application provides an embodiment of a candidate response sentence scoring device based on multi-level features and dual coding. This device embodiment is related to Figure 1 Corresponding to the method embodiment shown, the device can be specifically applied to various electronic devices.

[0074] An embodiment of the present application provides a candidate response sentence scoring device based on multi-level features and dual coding, including:

[0075] A training data set construction module 1, configured to obtain a conversation history and several corresponding candidate response sentences, input the conversation history, all candidate response sentences, and a prompt word for scoring each candidate response sentence according to the conversation history and all candidate response sentences into a large language model to obtain the true score corresponding to each candidate response sentence, and construct a training data set according to the conversation history, its corresponding several candidate response sentences, and the true score;

[0076] A model construction module 2, configured to construct a candidate response sentence scoring model and train it using the training data set to obtain a trained candidate response sentence scoring model. The candidate response sentence scoring model includes a word-level feature extraction module, a word-level similarity calculation module, a phrase-level feature extraction module, a phrase-level similarity calculation module, a sentence-level feature extraction module, a sentence-level similarity calculation module, and a comprehensive similarity calculation module;

[0077] The scoring module 3 is configured to obtain the conversation history to be scored and its corresponding several candidate reply sentences, encode the conversation history to be scored and each candidate reply sentence respectively to obtain the semantic representation of the conversation history and the semantic representation of each candidate reply sentence; input the semantic representation of the conversation history and the semantic representation of each candidate reply sentence into the trained candidate reply sentence scoring model. The semantic representation of the conversation history and the semantic representation of each candidate reply sentence respectively pass through the word-level feature extraction module to extract the word-level feature vector of each word in the conversation history and the word-level feature vector of each word in the candidate reply sentence. The word-level feature vector of each word in the conversation history and the word-level feature vector of each word in the candidate reply sentence are input into the word-level similarity calculation module to calculate the word-level similarity; the word-level feature vector of each word in the conversation history and the word-level feature vector of each word in the candidate reply sentence respectively pass through the phrase-level feature extraction module to extract the phrase-level feature vector of the conversation history and the phrase-level feature vector of the candidate reply sentence. The phrase-level feature vector of the conversation history and the phrase-level feature vector of the candidate reply sentence are input into the phrase-level similarity calculation module to calculate the phrase-level similarity; input the [CLS] token representation in the semantic representation of the conversation history and the [CLS] token representation of the semantic representation of each candidate reply sentence into the sentence-level feature extraction module to extract the sentence-level feature vector of the conversation history and the sentence-level feature vector of the candidate reply sentence. The sentence-level feature vector of the conversation history and the sentence-level feature vector of the candidate reply sentence are input into the sentence-level similarity calculation module to calculate the sentence-level similarity. The word-level similarity, phrase-level similarity and sentence-level similarity are input into the comprehensive similarity calculation module for weighted calculation to obtain the predicted score of each candidate reply sentence.

[0078] Figure 3 FIG. is a schematic hardware structure diagram of the electronic device provided by the embodiment of the present invention. As Figure 3 shown, the electronic device of this embodiment includes: a processor 301 and a memory 302; wherein the memory 302 is used to store computer execution instructions; the processor 301 is used to execute the computer execution instructions stored in the memory to implement each step executed by the electronic device in the above embodiment. Specifically, reference may be made to the relevant descriptions in the foregoing method embodiments.

[0079] Optionally, the memory 302 may be either independent or integrated with the processor 301.

[0080] When the memory 302 is independently provided, the electronic device further includes a bus 303 for connecting the memory 302 and the processor 301.

[0081] An embodiment of the present invention also provides a computer storage medium storing computer-executable instructions, which, when executed by a processor 301, implement the above method.

[0082] An embodiment of the present invention also provides a computer program product including a computer program, which, when executed by a processor 301, implements the above method.

[0083] In the embodiments provided by the present invention, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of modules is only a logical function division. In actual implementation, there may be other division methods. For example, multiple modules can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces, and the indirect coupling or communication connection of devices or modules can be in electrical, mechanical or other forms.

[0084] The modules described as separate components may or may not be physically separated, and the components displayed as modules may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to implement the solution of this embodiment.

[0085] In addition, in each embodiment of the present invention, the functional modules can be integrated in a processing unit, or each module can exist physically alone, or two or more modules can be integrated in one unit. The unit formed by the above modules can be implemented in the form of hardware, or in the form of a hardware plus software functional unit.

[0086] The integrated modules implemented in the form of software functional modules can be stored in a computer-readable storage medium. The above software functional modules are stored in a storage medium, including several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor 301 to execute some steps of the methods in various embodiments of the present application.

[0087] It should be understood that the above-mentioned processor 301 may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), etc. The general-purpose processor may be a microprocessor, or the processor 301 may also be any conventional processor 301, etc. The steps of the method disclosed in combination with the invention can be directly embodied as being executed and completed by the hardware processor 301, or by a combination of the hardware and software modules in the processor 301.

[0088] The memory 302 may include high-speed RAM memory, and may also include non-volatile storage NVM, such as at least one disk memory, and may also be a USB flash drive, a mobile hard disk, a read-only memory, a magnetic disk, or an optical disc, etc.

[0089] The bus 303 may be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, etc. The bus 303 can be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience in representation, the bus 303 in the attached drawings of this application is not limited to only one bus 303 or one type of bus 303.

[0090] The above storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read-Only Memory (EPROM), Programmable Read-Only Memory (PROM), Read-Only Memory (ROM), magnetic memory, flash memory, magnetic disk or optical disc. The storage medium can be any available medium that can be accessed by a general-purpose or special-purpose computer.

[0091] An exemplary storage medium is coupled to the processor 301, enabling the processor 301 to read information from the storage medium and write information to the storage medium. Of course, the storage medium can also be a component of the processor 301. The processor 301 and the storage medium can be located in an Application Specific Integrated Circuits (ASIC). Of course, the processor 301 and the storage medium can also exist as discrete components in an electronic device or a master device.

[0092] Those of ordinary skill in the art can understand that all or part of the steps to implement the above method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps including the above method embodiments; and the foregoing storage medium includes: various media such as ROM, RAM, magnetic disks, or optical discs that can store program codes.

[0093] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present invention.< / sep> < / sep> < / sep> < / sep> < / sep>

Claims

1. A candidate response sentence scoring method based on multi-level features and dual coding, characterized in that The steps include: Obtain the conversation history and its corresponding several candidate response sentences, input the conversation history, all the candidate response sentences, and the prompt words for scoring each candidate response sentence based on the conversation history and all the candidate response sentences into a large language model to obtain the true score corresponding to each candidate response sentence, and construct a training dataset according to the conversation history, its corresponding several candidate response sentences, and the true scores; Construct a candidate response sentence scoring model and train it using the training dataset to obtain a trained candidate response sentence scoring model. The candidate response sentence scoring model includes a word-level feature extraction module, a word-level similarity calculation module, a phrase-level feature extraction module, a phrase-level similarity calculation module, a sentence-level feature extraction module, a sentence-level similarity calculation module, and a comprehensive similarity calculation module; Obtain the conversation history to be scored and its corresponding several candidate response sentences, respectively encode the conversation history to be scored and each candidate response sentence to obtain the semantic representation of the conversation history and the semantic representation of each candidate response sentence; input the semantic representation of the conversation history and the semantic representation of each candidate response sentence into the trained candidate response sentence scoring model. The semantic representation of the conversation history and the semantic representation of each candidate response sentence respectively pass through the word-level feature extraction module to extract the word-level feature vectors of each word in the conversation history and the word-level feature vectors of each word in the candidate response sentence. The word-level feature vectors of each word in the conversation history and the word-level feature vectors of each word in the candidate response sentence are input into the word-level similarity calculation module to calculate the word-level similarity; the word-level feature vectors of each word in the conversation history and the word-level feature vectors of each word in the candidate response sentence respectively pass through the phrase-level feature extraction module to extract the phrase-level feature vectors of the conversation history and the phrase-level feature vectors of the candidate response sentence. The phrase-level feature vectors of the conversation history and the phrase-level feature vectors of the candidate response sentence are input into the phrase-level similarity calculation module to calculate the phrase-level similarity; input the [CLS] token representation in the semantic representation of the conversation history and the [CLS] token representation of the semantic representation of each candidate response sentence into the sentence-level feature extraction module to extract the sentence-level feature vectors of the conversation history and the sentence-level feature vectors of the candidate response sentence. The sentence-level feature vectors of the conversation history and the sentence-level feature vectors of the candidate response sentence are input into the sentence-level similarity calculation module to calculate the sentence-level similarity. The word-level similarity, phrase-level similarity, and sentence-level similarity are input into the comprehensive similarity calculation module for weighted calculation to obtain the predicted score of each candidate response sentence.

2. The candidate response statement scoring method based on multi-level features and dual coding according to claim 1, wherein The word-level feature extraction module includes a convolutional layer with a ReLU activation function and a multi-head self-attention layer connected in sequence. The words at each word position in the dialogue history and the words at each word position in the candidate response sentence respectively pass through the convolutional layer with the ReLU activation function and the multi-head self-attention layer in sequence, and the word-level feature vectors at each word position in the dialogue history and the word-level feature vectors at each word position in the candidate response sentence are extracted. In the word-level similarity calculation module, first calculate the cosine similarity between the word-level feature vectors at each word position in the dialogue history and the word-level feature vectors at the aligned word positions in the candidate response sentence, and then take the average of the cosine similarities of the word-level feature vectors at all word positions to obtain the word-level similarity.

3. The method for scoring candidate response sentences based on multi-level features and dual coding according to claim 1, wherein The phrase-level feature extraction module includes a max pooling layer and an average pooling layer. The word-level feature vectors at all word positions in the dialogue history are respectively input into the max pooling layer and the average pooling layer to obtain the max pooling result and the average pooling result corresponding to the dialogue history. Add the max pooling result and the average pooling result corresponding to the dialogue history to obtain the phrase-level feature vector of the dialogue history; The word-level feature vectors at all word positions in the candidate response sentence are respectively input into the max pooling layer and the average pooling layer to obtain the max pooling result and the average pooling result at each phrase position in the candidate response sentence. Add the max pooling result and the average pooling result corresponding to the candidate response sentence to obtain the phrase-level feature vector of the candidate response sentence; In the phrase-level similarity calculation module, calculate the cosine similarity between the phrase-level feature vector of the dialogue history and the phrase-level feature vector of the candidate response sentence to obtain the phrase-level similarity.

4. The candidate response sentence scoring method based on multi-level features and dual coding according to claim 1, characterized in that The sentence-level feature extraction module includes a fully connected layer. The [CLS] token representation in the semantic representation of the dialogue history and the [CLS] token representation in the semantic representation of each candidate response sentence respectively pass through the fully connected layer, and the sentence-level feature vector of the dialogue history and the sentence-level feature vector of the candidate response sentence are extracted. In the sentence-level similarity calculation module, calculate the cosine similarity between the sentence-level feature vector in the dialogue history and the sentence-level feature vector of the candidate response sentence to obtain the sentence-level similarity; The calculation formula of the comprehensive similarity calculation module is as follows: s = w1·W + w2·P + w3·J; Among them, w1, w2, and w3 respectively represent the weights of the word-level similarity, phrase-level similarity, and sentence-level similarity, and s, W, P, and J respectively represent the comprehensive similarity, word-level similarity, phrase-level similarity, and sentence-level similarity. Take the comprehensive similarity as the prediction score of the candidate response sentence.

5. The candidate response sentence scoring method based on multi-level features and dual coding according to claim 1, wherein The candidate response sentence scoring model uses a contrastive learning-based loss function during training, as shown in the following formula: Among them, L represents the loss function based on contrastive learning, and y ij represents the true score of the j-th candidate response sentence in the i-th sample of the training data set, and y ik represents the true score of the k-th candidate response sentence in the i-th sample of the training data set, and s ij represents the comprehensive similarity between the j-th candidate response sentence and the conversation history in the i-th sample of the training data set, and s ik represents the comprehensive similarity between the k-th candidate response sentence and the conversation history in the i-th sample of the training data set. N represents the total number of samples in the training data set, M represents the total number of candidate response sentences in the i-th sample, τ is the temperature parameter, and I(·) is the indicator function, which takes the value of 1 when the condition in the parentheses is true and 0 otherwise.

6. The candidate response sentence scoring method based on multi-level features and dual coding according to claim 1, wherein In the encoding processes of the dialogue history to be scored and each candidate response sentence, a pre-trained BERT model is used, and a maximum length is set to truncate the dialogue history to be scored or each candidate response sentence from right to left, and a tokenizer is used for padding. The output features of the last hidden layer in the pre-trained BERT model are used as the semantic representation of the dialogue history or the semantic representation of each candidate response sentence.

7. A candidate response statement scoring device based on multi-level features and dual coding, characterized in that Including: A training dataset construction module, configured to obtain a dialogue history and several corresponding candidate response sentences, input the dialogue history, all candidate response sentences, and a prompt word for scoring each candidate response sentence according to the dialogue history and all candidate response sentences into a large language model to obtain the true score corresponding to each candidate response sentence, and construct a training dataset according to the dialogue history, its corresponding several candidate response sentences, and the true scores; A model construction module, configured to construct a candidate response sentence scoring model and train it using the training dataset to obtain a trained candidate response sentence scoring model. The candidate response sentence scoring model includes a word-level feature extraction module, a word-level similarity calculation module, a phrase-level feature extraction module, a phrase-level similarity calculation module, a sentence-level feature extraction module, a sentence-level similarity calculation module, and a comprehensive similarity calculation module; The scoring module is configured to obtain the conversation history to be scored and its corresponding several candidate response sentences, encode the conversation history to be scored and each candidate response sentence respectively to obtain the semantic representation of the conversation history and the semantic representation of each candidate response sentence; input the semantic representation of the conversation history and the semantic representation of each candidate response sentence into the trained candidate response sentence scoring model. The semantic representation of the conversation history and the semantic representation of each candidate response sentence respectively pass through the word-level feature extraction module to extract the word-level feature vector of each word in the conversation history and the word-level feature vector of each word in the candidate response sentence. The word-level feature vector of each word in the conversation history and the word-level feature vector of each word in the candidate response sentence are input into the word-level similarity calculation module to calculate the word-level similarity; the word-level feature vector of each word in the conversation history and the word-level feature vector of each word in the candidate response sentence respectively pass through the phrase-level feature extraction module to extract the phrase-level feature vector of the conversation history and the phrase-level feature vector of the candidate response sentence. The phrase-level feature vector of the conversation history and the phrase-level feature vector of the candidate response sentence are input into the phrase-level similarity calculation module to calculate the phrase-level similarity; input the [CLS] token representation in the semantic representation of the conversation history and the [CLS] token representation of the semantic representation of each candidate response sentence into the sentence-level feature extraction module to extract the sentence-level feature vector of the conversation history and the sentence-level feature vector of the candidate response sentence. The sentence-level feature vector of the conversation history and the sentence-level feature vector of the candidate response sentence are input into the sentence-level similarity calculation module to calculate the sentence-level similarity. The word-level similarity, phrase-level similarity and sentence-level similarity are input into the comprehensive similarity calculation module for weighted calculation to obtain the predicted score of each candidate response sentence.

8. An electronic device, comprising: One or more processors; A storage device for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1-6.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the method according to any one of claims 1-6.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the method according to any one of claims 1-6.