Replay content generation method based on fine-grained user demand satisfaction reasoning

Through the replies content generation method based on fine-grained user demand satisfaction reasoning, the problem that large language models are difficult to meet the specific needs of users in personalized response generation is solved, and more accurate and targeted replies content generation is achieved, which improves user satisfaction.

CN120104729AActive Publication Date: 2025-06-06ZHEJIANG UNIV

Patent Information

Application Number
CN202510063315.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-15
Publication Date
2025-06-06
Estimated Expiration
2045-01-15

AI Technical Summary

Technical Problem

Large language model LLMs is difficult to accurately understand and meet the specific needs of individual users when dealing with the generation of personalized responses, resulting in the lack of targetedness and accuracy of the generated content.

Method used

The reply content generation method based on fine-grained user demand satisfaction reasoning is adopted. By obtaining the historical dialogue text of users and manual customer service, data preprocessing and characterization learning is carried out, a global demand prototype tree and a multi-grained satisfaction attribution model are constructed, and the user's fine-grained satisfaction is inferred from bottom to top, and through self-supervisation adaptation, the large language model is optimized to generate replies that are more in line with user needs.

Benefits of technology

It realizes the accuracy and pertinence of large language models in personalized response generation, which can better meet users' fine-grained needs and improve user satisfaction and reply quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120104729A_ABST
    Figure CN120104729A_ABST
Patent Text Reader

Abstract

The invention discloses a reply content generation method based on fine-grained user demand satisfaction reasoning. The method comprises the steps that historical dialogue texts of a user are obtained for preprocessing and representation learning, then clustering, unsupervised learning, association and matching are conducted, then a satisfaction tree is constructed, layer-by-layer reasoning is conducted on the satisfaction tree from bottom to top, then the overall satisfaction degree is predicted, supervised learning is conducted, and a multi-granularity satisfaction degree attribution model is constructed; obtaining a fine-grained demand unsatisfied by the user, inputting the fine-grained demand into the large-language model for optimization, and performing self-supervised adaptive optimization on the model according to the optimized overall satisfaction; and processing the to-be-replied text, and inputting the processed text into the model to generate reply content. According to the method, the fine-grained demand of the user can be recognized based on the dialogue text, the multi-granularity satisfaction degree of the user to the reply content is inferred, and self-adaptive and self-supervised fine tuning of prompt is realized in combination with inference attribution judgment of the multi-granularity satisfaction degree, so that personalized content generation of a large language model is realized, and the fine-grained user demand is met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a reply content generation method, and to the technical field of artificial intelligence large language models, and in particular to a reply content generation method based on fine-grained user demand satisfaction reasoning. Background Art

[0002] Large language models (LLMs) have been widely used in recent years due to their powerful capabilities in natural language processing (NLP) tasks. By training on large-scale public datasets, large language models (LLMs) can generate general content and common sense knowledge, and show excellent performance in various tasks such as text generation, machine translation, and dialogue systems. The core advantage of this type of model is that it can effectively capture language patterns and contextual information through deep learning methods, and has the ability to migrate across domains. Therefore, large language models (LLMs) have significant advantages in generating generalized text and answering common sense questions.

[0003] However, large language models (LLMs) still face many challenges in dealing with personalized response generation. Since most of their training data comes from the public domain, large language models (LLMs) are usually unable to accurately understand and meet the specific needs of individual users. Studies have shown that large language models (LLMs) may experience hallucinations when dealing with personalization tasks, that is, the information generated by the model may seem reasonable but is actually irrelevant to the user query or context. In addition, large language models (LLMs) tend to generate vague, stereotyped content and lack specificity when generating personalized responses. This is particularly evident in scenarios that target specific user needs, such as customer service, product recommendations, and personalized question-and-answer applications.

[0004] When large language models (LLMs) are used in specific task scenarios, the limitations of the model become more apparent. Taking large language models (LLMs) as task-oriented dialogue agents as an example, the answers they generate are usually general and fail to respond accurately to the user's specific intentions or questions. This is particularly prominent in personalized task scenarios. For example, when large language models (LLMs) are used to act as hotel managers to respond to users' online reviews, the responses generated by the model are usually not targeted and cannot effectively address the personalized concerns raised by users. The answers of such models are often superficial and cannot deeply analyze and respond to the specific issues mentioned by users in their reviews. This limitation limits the application effect of large language models (LLMs) in personalized service scenarios and cannot fully meet users' expectations for accurate and high-quality responses.

[0005] In order to meet the above challenges, many studies in recent years have been devoted to improving the personalized generation capabilities of large language models (LLMs). Some methods use historical data of specific users to further train large language models (LLMs) through personalized data fine-tuning to improve the model's understanding and response capabilities to user preferences. However, such methods usually rely on a large amount of user data and require significant computing resources and time to complete model fine-tuning. In addition, although the fine-tuned model performs better in specific user scenarios, it is difficult to generalize due to the limitations of its training data set. The historical training data is also limited by the drawbacks of templated responses, making it difficult to provide good supervision for the personalized generation of large language models (LLMs). As a result, large language models (LLMs) still have difficulty in achieving more accurate recognition and coverage of users' fine-grained needs.

[0006] Although large language models (LLMs) have great advantages in generating general text, their limitations in generating personalized responses are still obvious, especially in task-oriented dialogue scenarios. Although existing methods can improve the personalization capabilities of large language models (LLMs) to a certain extent, there is still a large gap to completely solve the problem. Future research and technological development need to further explore how to combine fine-grained user demand identification to provide more effective solutions for personalized response generation. Summary of the invention

[0007] In order to solve the problems existing in the background technology, the present invention provides a method for generating reply content based on fine-grained user demand satisfaction reasoning. The method of the present invention can solve the problem that the personalized generation capability of existing large language models (LLMs) is limited and it is difficult to accurately identify and cover the fine-grained needs of users.

[0008] The technical solution adopted by the present invention is:

[0009] The reply content generation method based on fine-grained user demand satisfaction reasoning of the present invention comprises:

[0010] Step 1: Obtain the historical conversation texts between several users and manual customer service and perform data preprocessing and representation learning in sequence to obtain the word embedding representation of each user.

[0011] Step 2: Perform clustering and unsupervised learning on the embedding representations of each word to obtain the prototype representation vector, and then perform hierarchical clustering to build a global demand prototype tree; by performing prototype learning on historical conversation texts, unsupervised learning of the global fine-grained demand prototype representation in the conversation scenario, hierarchical clustering of demand prototypes of different granularities is performed to build a global user demand prototype tree.

[0012] Step 3: Associate and match the global demand prototype tree and each user's word embedding representation in turn to build a satisfaction tree. Perform bottom-up reasoning on the satisfaction tree to obtain the predicted overall satisfaction. Perform supervised learning based on the predicted overall satisfaction to build a multi-granularity satisfaction attribution model MSA (Multi-granularity user Satisfaction Attribution); locally reason about each user's fine-grained satisfaction with the received replies, perform bottom-up reasoning layer by layer, and predict the user's overall satisfaction.

[0013] Step 4: Obtain the fine-grained demands that the current user is dissatisfied with through the multi-granularity satisfaction attribution model MSA, and input them into the preset prompt template of the large language model to optimize the current user's reply text, and perform self-supervised adaptive optimization on the large language model according to the optimized overall satisfaction of the current user to obtain the trained large language model; adapt, adjust and optimize the prompts of the large language model based on the user's fine-grained satisfaction; use the multi-granularity satisfaction attribution model MSA as a generated content evaluator, incorporate the evaluation score into the optimization function of the large language model to fine-tune the large language model, and realize self-supervised adaptive optimization of language model generation.

[0014] Step 5: The query text of the user to be answered is processed in the same way as in step 1 and then input into the trained large language model. After the processing is completed, the reply text of the user's query text is output to complete the generation of the reply content.

[0015] The step 1 is specifically as follows:

[0016] Step 1.1: Get the historical conversation texts T between several users and human customer service. in, and 1 They represent the query text of the first user, the reply text of the manual customer service to the query text of the first user, and the satisfaction label of the first user for the reply text. and 2 They represent the query text of the second user, the reply text of the manual customer service to the query text of the second user, and the satisfaction label of the second user for the reply text. and n They respectively represent the query text of the n-th user, the reply text of the manual customer service to the query text of the n-th user, and the satisfaction label of the n-th user for the reply text.

[0017] Step 1.2: Data preprocessing is performed on the historical conversation texts between each user and the manual customer service. That is, the historical conversation texts are segmented and then stop words, modal particles, and punctuation marks are removed, and finally keywords reflecting the service or product are retained.

[0018] Step 1.3: Use the pre-trained language representation BERT (Bidirectional Encoder Representations from Transformers) model to perform feature extraction and vectorization processing on each user’s keywords to obtain the word embedding representation of each user in, and They respectively represent the word embedding representations of the query text of the i-th user and the reply text of the manual customer service to the query text of the i-th user.

[0019] The step 2 is specifically as follows:

[0020] Step 2.1: Use K-means clustering method to cluster each word embedding representation and divide it into K clusters, the cluster set p = [p 1 ,p 2 ,…,p K ], p 1 、p 2 ,…,p K They represent the 1st, 2nd, ..., K clusters respectively, and the center vector of each cluster is used as K initial prototype representation vectors.

[0021] Step 2.2: Reconstruction loss by evaluating Prototype Diversity Loss and representative loss Unsupervised learning is performed on the K initial prototype representation vectors, and gradient descent is used to optimize the K initial prototype representation vectors to obtain the optimized K prototype representation vectors that best represent the demand attributes that users are concerned about.

[0022] Step 2.3: Use the hierarchical clustering method to construct each prototype representation vector into a hierarchical global demand prototype tree to reflect the multi-attribute and multi-granularity characteristics of user needs.

[0023] The evaluation reconstruction loss Prototype Diversity Loss and representative loss The details are as follows:

[0024]

[0025]

[0026] Among them, λ1 , 2 and λ 3 Respectively represent the evaluation reconstruction loss Prototype Diversity Loss and representative loss The learnable weight parameters of represents the estimate of the word embedding representation of the query text of the i-th user, represents the correlation between the word embedding representation of the query text of the i-th user and the k-th initial prototype representation vector; N represents the number of sentences in the query text of each user, and They represent the global representation of the query text of the i-th user and the word embedding representation of the j-th sentence in it, Represents the center vector of the k-th prototype representation vector; ‖‖ 2 represents the bi-norm; L represents the number of keywords in each user's query text; dist() represents the distance function; Represents the word embedding representation of the jth keyword of the i-th user.

[0027] The step 3 is as follows:

[0028] Step 3.1: For each tree node in the global demand prototype tree, the tree node is associated with the word embedding representation of each user's query text and the word embedding representation of the manual customer service's reply text to each user's query text by calculating the dot product similarity, so as to construct each user's query embedding tree and reply embedding tree respectively. Each tree node in the query embedding tree represents the user's attention to the node demand, and each tree node in the reply embedding tree represents the coverage of the reply content on the node demand; after matching each tree node in each user's query embedding tree and reply embedding tree, a satisfaction tree is constructed. Each tree node in the satisfaction tree represents the initial fine-grained satisfaction of the current user on the fine-grained demand of the tree node for a given reply text, that is, the satisfaction with the demands of each node at different granularities.

[0029] Step 3.2: The satisfaction tree is recursively calculated from the bottom up, and the final fine-grained satisfaction of the tree nodes at each level is aggregated and inferred in turn to obtain the predicted overall satisfaction;

[0030] Step 3.3: Use the root mean square error of the predicted overall satisfaction of each user as the loss for supervised learning, and optimize the final parameter fine-grained satisfaction of each layer of the satisfaction tree through back-propagation to construct a multi-granular satisfaction attribution model MSA.

[0031] In the step 3.2, the final fine-grained satisfaction of each tree node in each layer of the satisfaction tree is obtained by weighting its own initial fine-grained satisfaction and the final fine-grained satisfaction of several tree nodes in the upper layer connected to it. The predicted overall satisfaction is the final fine-grained satisfaction of a root node in the top layer of the satisfaction tree, or the final fine-grained satisfaction of each tree node in a layer of M layers adjacent to the root node is input into a fully connected layer for weighted averaging to obtain the fine-grained satisfaction as the final fine-grained satisfaction.

[0032] In the step 4, each node in the multi-granularity satisfaction attribution model MSA represents the optimized fine-grained satisfaction of the current user on the fine-grained demand of the current node for a given reply text, and the fine-grained demand whose optimized fine-grained satisfaction is lower than the satisfaction threshold δ is input into the preset prompt template of the large language model, and the large language model outputs the reply text of the optimized fine-grained demand after processing, and the optimization is repeated until the optimized fine-grained satisfaction of each fine-grained demand of the current user reaches the satisfaction threshold δ or reaches the preset number of iterations S, and the reply text at this time is used as the optimized reply text of the current user.

[0033] The large language model uses the LLAMA model or the ChatGLM model, etc. The preset prompt template is as follows: the user is still not satisfied with the first, fourth and fifth ones. Please strengthen the coverage of these specific needs in the reply.

[0034] In the step 4, the query text and the optimized reply text of the current user are repeated with steps 1-3 to obtain the optimized overall satisfaction of the current user, the optimized reply texts whose overall satisfaction is lower than the satisfaction threshold δ are constructed as an unsatisfactory reply set, and the optimized reply texts whose overall satisfaction reaches the satisfaction threshold δ are constructed as a satisfactory reply set, and the generation loss and contrast loss of the large language model are constructed according to the unsatisfactory reply set and the satisfactory reply set, the generation loss is used to improve the ability of the model to generate satisfactory replies, and the contrast loss is used to expand the difference in generation probability between satisfactory replies and unsatisfactory replies; the large language model is self-supervised adaptively optimized, and the model parameters of the large language model are optimized by back propagation until the generation loss and the contrast loss converge to obtain the trained large language model.

[0035] Use the reply text enhanced with prompts as training data to fine-tune the large language model, introduce generation loss and contrast loss into the loss function of the large model, efficiently optimize some model parameters through back propagation, and improve the satisfactory reply ability of the model at the batch level in a self-supervised manner to better meet the personalized needs of users.

[0036] The generation loss L of the large language model is gen and contrast loss L con The details are as follows:

[0037]

[0038] in, represents the number of satisfactory responses obtained in the jth iteration; p() represents the generation probability; and denote the generation parameters of the mth response obtained in the jth iteration and the generation parameters of the response with the highest satisfaction score; MSA() denotes the overall satisfaction obtained by the multi-granularity satisfaction attribution model; M j represents the number of replies obtained in the jth iteration.

[0039] The method of the present invention firstly learns the global fine-grained demand prototype representation in the dialogue scenario in an unsupervised manner through prototype learning, hierarchically clusters the demand prototypes of different granularities, and constructs a user demand prototype tree structure; secondly, the user's fine-grained satisfaction with the reply text (i.e., the satisfaction with the demand of each tree node) is inferred, and the user's overall satisfaction is predicted by bottom-up layer-by-layer reasoning; then, the satisfaction prediction model is used as a generated content evaluator, and the predicted unsatisfactory fine-grained demands are input into the prompt template, so that the large language model generates a reply that pays more attention to the user's fine-grained demands; at the same time, the generated content evaluation score is also incorporated into the optimization function of the large language model to realize the self-supervised adaptive optimization of language model generation.

[0040] The beneficial effects of the present invention are:

[0041] The method of the present invention identifies the fine-grained needs of users based on the dialogue text, and infers the multi-granular satisfaction of users with the reply content, and combines the reasoning attribution judgment of the multi-granularity satisfaction to achieve prompt adaptation and self-supervision fine-tuning, thereby realizing personalized content generation of the large prediction model LLM to meet the fine-grained user needs. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] Figure 1 is a flow chart of the method of the present invention;

[0043] Figure 2 It is a schematic diagram of the global demand prototype tree structure of the present invention;

[0044] Figure 3 It is a schematic diagram of the multi-granularity satisfaction attribution model of the present invention. DETAILED DESCRIPTION

[0045] The present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.

[0046] like Figure 1 As shown, the reply content generation method based on fine-grained user demand satisfaction reasoning of the present invention is specifically as follows:

[0047] Step 1: Obtain the historical conversation texts of several users and manual customer service and perform data preprocessing and representation learning in sequence to obtain the word embedding representation of each user, as follows:

[0048] Step 1.1: Get the historical conversation texts T between several users and human customer service. in, and 1 They represent the query text of the first user, the reply text of the manual customer service to the query text of the first user, and the satisfaction label of the first user for the reply text. and 2 They represent the query text of the second user, the reply text of the manual customer service to the query text of the second user, and the satisfaction label of the second user for the reply text. and n They represent the query text of the nth user, the reply text of the manual customer service to the query text of the nth user, and the satisfaction label of the nth user to the reply text, y 1 ,y 2 , …, y n ∈[0,5].

[0049] Step 1.2: Data preprocessing is performed on the historical conversation texts between each user and the manual customer service. That is, the historical conversation texts are segmented and then stop words, modal particles, and punctuation marks are removed, and finally keywords reflecting the service or product are retained.

[0050] Step 1.3: Use the pre-trained language representation BERT (Bidirectional Encoder Representations from Transformers) model to perform feature extraction and vectorization processing on each user’s keywords to obtain the word embedding representation of each user in, and They represent the word embedding representations of the query text of the i-th user and the reply text of the manual customer service to the query text of the i-th user, respectively. They represent the word embedding representations of the 1st, 2nd, …, Lth keywords in the query text of the i-th user, respectively. They represent the word embedding representations of the 1st, 2nd, …, Lth keywords in the reply text of the manual customer service to the query text of the i-th user,

[0051] Step 2: After clustering and unsupervised learning of each word embedding representation, the prototype representation vector is obtained, and then hierarchical clustering is performed to construct a global demand prototype tree, such as Figure 2As shown in the figure, some parent node representation vectors are obtained by averaging all their child nodes, as follows:

[0052] Step 2.1: Use K-means clustering method to cluster each word embedding representation and divide it into K clusters, the cluster set p = [p 1 ,p 2 ,…,p K ], p 1 、p 2 ,…,p K They represent the 1st, 2nd, ..., K clusters respectively, and the center vector of each cluster is used as K initial prototype representation vectors.

[0053] Step 2.2: Reconstruction loss by evaluating Prototype Diversity Loss and representative loss Unsupervised learning is performed on the K initial prototype representation vectors, and gradient descent is used to optimize the K initial prototype representation vectors to obtain the optimized K prototype representation vectors that best represent the demand attributes that users are concerned about.

[0054] Step 2.3: Use the hierarchical clustering method to construct each prototype representation vector into a hierarchical global demand prototype tree to reflect the multi-attribute and multi-granularity characteristics of user needs.

[0055] The global requirement prototype tree is as follows:

[0056]

[0057] in, Represents the node representation vector of the Hth row and Bth column in the global demand prototype tree.

[0058] The representation vector of some parent nodes is obtained by averaging all its child nodes.

[0059] By performing prototype learning on historical conversation texts, we unsupervisedly learn the global fine-grained demand prototype representation in the conversation scenario, and perform hierarchical clustering on demand prototypes of different granularities to build a global user demand prototype tree.

[0060] Evaluating the reconstruction loss Prototype Diversity Loss and representative loss The details are as follows:

[0061]

[0062] Among them, λ 1 , 2 and λ 3 Respectively represent the evaluation reconstruction loss Prototype Diversity Loss and representative loss The learnable weight parameters of represents the estimate of the word embedding representation of the query text of the i-th user, represents the correlation between the word embedding representation of the query text of the i-th user and the k-th initial prototype representation vector; N represents the number of sentences in the query text of each user, and They represent the global representation of the query text of the i-th user and the word embedding representation of the j-th sentence in it, Represents the center vector of the k-th prototype representation vector; ‖‖ 2 represents the bi-norm; L represents the number of keywords in each user's query text; dist() represents the distance function; Represents the word embedding representation of the jth keyword of the i-th user.

[0063] The word embedding representation of each user is used to calculate the word vector weight in the text through the self-attention mechanism, and then the word vectors in the text are weighted averaged to obtain the global representation of the word embedding representation of each user. in, Represents the word vector weight in the text of the word embedding representation of the i-th user; then the global representation Perform nonlinear transformation to calculate the weight of each initial prototype representation vector The details are as follows:

[0064]

[0065] Among them, softmax[] represents nonlinear transformation, σ k () represents k-dimensional linear transformation.

[0066] Then, the K initial prototype representation vectors are weighted averaged to reconstruct the estimate of the text features. The gap between the reconstructed text representation and the global text feature is calculated to obtain the reconstruction loss

[0067] The diversity loss improves the diversity of the initial prototype representation vectors by calculating the distance between each initial prototype representation vector and other initial prototype representation vectors, thereby improving the difference between each initial prototype representation vector and other initial prototype representation vectors.

[0068] The representative loss calculates the average distance between all word embedding vectors and their nearest initial prototype representation vectors, making the initial prototype representation vector more capable of representing a series of word embeddings.

[0069] Step 3: Associate and match the global demand prototype tree and each user's word embedding representation in turn to build a satisfaction tree. Perform bottom-up reasoning on the satisfaction tree to obtain the predicted overall satisfaction. Perform supervised learning based on the predicted overall satisfaction to build a multi-granularity satisfaction attribution model MSA, such as Figure 3 As shown, fine-grained user satisfaction reasoning and overall satisfaction prediction are performed as follows:

[0070] Step 3.1: For each tree node in the global demand prototype tree, associate the tree node with the word embedding representation of each user's query text and the word embedding representation of the manual customer service's reply text to each user's query text by calculating the dot product similarity. m>K, thus constructing the query embedding tree and reply embedding tree of each user respectively. Each tree node in the query embedding tree represents the user's attention to the node's needs, and each tree node in the reply embedding tree represents the coverage of the reply content to the node's needs. After matching each tree node in each user's query embedding tree and reply embedding tree, a satisfaction tree is constructed. Among them, σ is a learnable linear function, ⊙ represents the Hadamard product, Used to learn the match between queries and responses, Represents the estimate of the word embedding representation of the reply text of the manual customer service to the query text of the i-th user; each tree node in the satisfaction tree represents the initial fine-grained satisfaction of the current user on the fine-grained requirements of the tree node for a given reply text, that is, the satisfaction with the requirements of each node at different granularities.

[0071] Step 3.2: The satisfaction tree is recursively calculated from the bottom up, and the final fine-grained satisfaction of the tree nodes at each level is aggregated and inferred in turn to obtain the predicted overall satisfaction. The final fine-grained satisfaction of each tree node in each layer of the satisfaction tree is obtained by weighting its own initial fine-grained satisfaction and the final fine-grained satisfaction of several tree nodes in the upper layer connected to it. The predicted overall satisfaction is the final fine-grained satisfaction of a root node in the top layer of the satisfaction tree, or the final fine-grained satisfaction of each tree node in a layer of M layers adjacent to the root node is input into a fully connected layer for weighted average to obtain the fine-grained satisfaction as the final fine-grained satisfaction.

[0072] The predictions for overall satisfaction are as follows:

[0073]

[0074] in, and They represent the demand satisfaction of the b-th node in the h-th and h-1-th layers of the satisfaction tree of the i-th user respectively; represents the demand satisfaction of the qth node connected to the bth node of the upper layer in the h-1th layer of the satisfaction tree of the i-th user, Q represents the total number of lower-layer nodes connected to the bth node, represents the original satisfaction of the qth node in the h-1th layer of the satisfaction tree of the i-th user; ξ (h-1)q represents the learnable weight parameter of the qth node in the h-1th layer connected to the bth node in the hth layer in the satisfaction tree of the ith user; α represents the learnable factor, which coordinates the relationship between the satisfaction of the matched tree node and the satisfaction obtained by aggregating the child nodes of the node.

[0075] Step 3.3: Use the root mean square error of the predicted overall satisfaction of each user as the loss for supervised learning. The loss function is y and They represent the satisfaction of the true label and the satisfaction of the prediction respectively. ,λ 4 Represents the learnable weight parameter of the loss function; ξ represents the learnable weight parameter of the satisfaction of the lower-level demand node during recursive calculation; and the final parameter fine-grained satisfaction of each layer of the satisfaction tree is learned through back-propagation optimization, thereby constructing a multi-granularity satisfaction attribution model MSA.

[0076] We locally infer each user's fine-grained satisfaction with the replies they receive, and reason from bottom to top to predict the user's overall satisfaction.

[0077] Step 4: Obtain the fine-grained demands that the current user is dissatisfied with through the multi-granularity satisfaction attribution model MSA, and input them into the preset prompt template of the large language model to optimize the current user's reply text, and perform self-supervised adaptive optimization on the large language model according to the optimized overall satisfaction of the current user to obtain the trained large language model; adapt, adjust and optimize the prompts of the large language model based on the user's fine-grained satisfaction; use the multi-granularity satisfaction attribution model MSA as a generated content evaluator, incorporate the evaluation score into the optimization function of the large language model to fine-tune the large language model, and realize self-supervised adaptive optimization of language model generation.

[0078] Each node in the multi-granularity satisfaction attribution model MSA represents the optimized fine-grained satisfaction of the current user on the fine-grained demand of the current node for a given reply text. The fine-grained demand whose optimized fine-grained satisfaction is lower than the satisfaction threshold δ is input into the preset prompt template of the large language model. After processing, the large language model outputs the reply text of the optimized fine-grained demand. The optimization is repeated until the optimized fine-grained satisfaction of each fine-grained demand of the current user reaches the satisfaction threshold δ or reaches the preset number of iterations S. The reply text at this time is used as the optimized reply text of the current user.

[0079] The large language model uses the LLAMA model or the ChatGLM model, etc. The preset prompt template is as follows: the user is still not satisfied with the first, fourth and fifth ones. Please strengthen the coverage of these specific needs in the reply.

[0080] Repeat steps 1-3 for the query text and optimized reply text of the current user to obtain the optimized overall satisfaction of the current user. Construct the optimized reply texts whose overall satisfaction is lower than the satisfaction threshold δ into an unsatisfactory reply set, and construct the optimized reply texts whose overall satisfaction reaches the satisfaction threshold δ into a satisfactory reply set. Construct the generation loss and contrast loss of the large language model based on the unsatisfactory reply set and the satisfactory reply set. The generation loss is used to improve the ability of the model to generate satisfactory replies, and the contrast loss is used to expand the difference in generation probability between satisfactory replies and unsatisfactory replies. Perform self-supervised adaptive optimization on the large language model, and optimize the model parameters of the large language model through back propagation until the generation loss and contrast loss converge to obtain the trained large language model.

[0081] Use the reply text enhanced with prompts as training data to fine-tune the large language model, introduce generation loss and contrast loss into the loss function of the large model, efficiently optimize some model parameters through back propagation, and improve the satisfactory reply ability of the model at the batch level in a self-supervised manner to better meet the personalized needs of users.

[0082] The generation loss L of the large language model gen and contrast loss L con The details are as follows:

[0083]

[0084]

[0085] in, represents the number of satisfactory responses obtained in the jth iteration; p() represents the generation probability; and denote the generation parameters of the mth response obtained in the jth iteration and the generation parameters of the response with the highest satisfaction score; MSA() denotes the overall satisfaction obtained by the multi-granularity satisfaction attribution model; M j represents the number of replies obtained in the jth iteration.

[0086] Step 5: The query text of the user to be answered is processed in the same way as in step 1 and then input into the trained large language model. After the processing is completed, the reply text of the user's query text is output to complete the generation of the reply content.

[0087] In order to provide interaction with a user, the method of the present invention can be implemented on a computer, which has: a display device (such as a liquid crystal display monitor, etc.) for displaying information to the user, and a keyboard and a pointing device (such as a mouse, etc.), and the user can provide input to the computer through the keyboard and the pointing device.

[0088] The above specific implementations do not constitute a limitation on the protection scope of the present disclosure. It should be understood by those skilled in the art that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modification, equivalent substitution and improvement made within the spirit and principle of the present disclosure shall be included in the protection scope of the present disclosure.

Claims

1. A reply content generation method based on fine-grained user demand satisfaction reasoning, characterized in that: include: Step 1: Obtain the historical conversation texts of several users and manual customer service and perform data preprocessing and representation learning in sequence to obtain the word embedding representation of each user; Step 2: After clustering and unsupervised learning of each word embedding representation, the prototype representation vector is obtained, and then hierarchical clustering is performed to construct a global demand prototype tree; Step 3: Associate and match the global demand prototype tree and each user's word embedding representation in turn to build a satisfaction tree. Perform bottom-up reasoning on the satisfaction tree to obtain the predicted overall satisfaction. Perform supervised learning based on the predicted overall satisfaction to build a multi-granularity satisfaction attribution model MSA. Step 4: Obtain the fine-grained demands that the current user is dissatisfied with through the multi-granularity satisfaction attribution model MSA, and input them into the preset prompt template of the large language model to optimize the reply text of the current user, and perform self-supervised adaptive optimization on the large language model according to the optimized overall satisfaction of the current user to obtain the trained large language model; Step 5: The query text of the user to be answered is processed in the same way as in step 1 and then input into the trained large language model. After the processing is completed, the reply text of the user's query text is output to complete the generation of the reply content.

2. The reply content generation method based on fine-grained user demand satisfaction reasoning according to claim 1 is characterized by: The step 1 is specifically as follows: Step 1.1: Get the historical conversation text T between several users and human customer service. in, and y1 represent the query text of the first user, the reply text of the manual customer service to the query text of the first user, and the satisfaction label of the first user for the reply text, respectively. and y2 represent the query text of the second user, the reply text of the manual customer service to the query text of the second user, and the satisfaction label of the second user for the reply text, respectively. and n Respectively represent the query text of the nth user, the reply text of the manual customer service to the query text of the nth user, and the satisfaction label of the nth user to the reply text; Step 1.2: Perform data preprocessing on the historical conversation texts between each user and the human customer service representative. That is, the historical conversation texts are segmented and then stop words, modal particles, and punctuation marks are removed, and finally the keywords are retained. Step 1.3: Use the pre-trained language representation BERT model to perform feature extraction and vectorization processing on each user’s keywords to obtain the word embedding representation of each user in, and They respectively represent the word embedding representations of the query text of the i-th user and the reply text of the manual customer service to the query text of the i-th user.

3. The reply content generation method based on fine-grained user demand satisfaction reasoning according to claim 1 is characterized by: The step 2 is specifically as follows: Step 2.1: Use K-means clustering method to cluster each word embedding representation and divide it into K clusters, the cluster set p = [p1, p2, ..., p K ], p1, p2, …, p K Represent the 1st, 2nd, ..., Kth clusters respectively, and the center vector of each cluster is used as K initial prototype representation vectors; Step 2.2: Reconstruction loss by evaluating Prototype Diversity Loss and representative loss Perform unsupervised learning on the K initial prototype representation vectors, and use gradient descent to optimize the K initial prototype representation vectors to obtain optimized K prototype representation vectors; Step 2.3: Use the hierarchical clustering method to construct each prototype representation vector into a hierarchical global demand prototype tree.

4. The reply content generation method based on fine-grained user demand satisfaction reasoning according to claim 3 is characterized by: The evaluation reconstruction loss Prototype Diversity Loss and representative loss The details are as follows: Among them, λ1, λ2 and λ3 represent the evaluation reconstruction loss respectively. Prototype Diversity Loss and representative loss The learnable weight parameters of represents the estimate of the word embedding representation of the query text of the i-th user, represents the correlation between the word embedding representation of the query text of the i-th user and the k-th initial prototype representation vector; N represents the number of sentences in the query text of each user, and They represent the global representation of the query text of the i-th user and the word embedding representation of the j-th sentence in it, represents the central vector of the k-th prototype representation vector; ‖‖2 represents the two-norm; L represents the number of keywords in each user's query text; dist() represents the distance function; Represents the word embedding representation of the jth keyword of the i-th user.

5. The method for generating reply content based on fine-grained user demand satisfaction reasoning according to claim 1 is characterized by: The step 3 is as follows: Step 3.1: For each tree node in the global demand prototype tree, associate the tree node with the word embedding representation of each user's query text and the word embedding representation of the manual customer service's reply text to each user's query text by calculating the dot product similarity, thereby constructing each user's query embedding tree and reply embedding tree respectively; After matching each tree node in each user's query embedding tree and reply embedding tree, a satisfaction tree is constructed. Each tree node in the satisfaction tree represents the initial fine-grained satisfaction of the current user on the fine-grained requirements of the tree node for a given reply text. Step 3.2: The satisfaction tree is recursively calculated from the bottom up, and the final fine-grained satisfaction of the tree nodes at each level is aggregated and inferred in turn to obtain the predicted overall satisfaction; Step 3.3: Use the root mean square error of the predicted overall satisfaction of each user as the loss for supervised learning, and optimize the final parameter fine-grained satisfaction of each layer of the satisfaction tree through back-propagation to construct a multi-granular satisfaction attribution model MSA.

6. The method for generating reply content based on fine-grained user demand satisfaction reasoning according to claim 5 is characterized by: In the step 3.2, the final fine-grained satisfaction of each tree node in each layer of the satisfaction tree is obtained by weighting its own initial fine-grained satisfaction and the final fine-grained satisfaction of several tree nodes in the upper layer connected to it. The predicted overall satisfaction is the final fine-grained satisfaction of a root node in the top layer of the satisfaction tree, or the final fine-grained satisfaction of each tree node in a layer of M layers adjacent to the root node is input into a fully connected layer for weighted averaging to obtain the fine-grained satisfaction as the final fine-grained satisfaction.

7. The method for generating reply content based on fine-grained user demand satisfaction reasoning according to claim 5 is characterized by: In the step 4, each node in the multi-granularity satisfaction attribution model MSA represents the optimized fine-grained satisfaction of the current user on the fine-grained demand of the current node for a given reply text, and the fine-grained demand whose optimized fine-grained satisfaction is lower than the satisfaction threshold δ is input into the preset prompt template of the large language model, and the large language model outputs the reply text of the optimized fine-grained demand after processing, and the optimization is repeated until the optimized fine-grained satisfaction of each fine-grained demand of the current user reaches the satisfaction threshold δ or reaches the preset number of iterations S, and the reply text at this time is used as the optimized reply text of the current user.

8. The method for generating reply content based on fine-grained user demand satisfaction reasoning according to claim 1 is characterized by: In the step 4, the query text and the optimized reply text of the current user are repeated with steps 1-3 to obtain the optimized overall satisfaction of the current user, the optimized reply texts whose overall satisfaction is lower than the satisfaction threshold δ are constructed as an unsatisfactory reply set, and the optimized reply texts whose overall satisfaction reaches the satisfaction threshold δ are constructed as a satisfactory reply set, the generation loss and contrast loss of the large language model are constructed according to the unsatisfactory reply set and the satisfactory reply set, the large language model is self-supervised adaptively optimized, and the model parameters of the large language model are optimized by back propagation until the generation loss and the contrast loss converge to obtain the trained large language model.

9. The method for generating reply content based on fine-grained user demand satisfaction reasoning according to claim 8 is characterized in that: The generation loss L of the large language model is gen and contrast loss L con The details are as follows: in, represents the number of satisfactory responses obtained in the jth iteration; p() represents the generation probability; and denote the generation parameters of the mth response obtained in the jth iteration and the generation parameters of the response with the highest satisfaction score; MSA() denotes the overall satisfaction obtained by the multi-granularity satisfaction attribution model; M j represents the number of replies obtained in the jth iteration.

Citation Information

Patent Citations

  • Satisfaction-based user simulation method and system thereof

    CN114048301A

  • Dynamic adaptation question answering system and method based on hierarchical structure and retrieval enhancement

    CN118193714A

  • Ignition stick

    KR1020250045024A

Cited By

  • Voice call content intelligent analysis and seat reminding system

    CN120581009A