Response method and device, electronic equipment, storage medium and program product

By introducing the control vector of emotion category in the representation learning of LLM, the problems of large resource investment, long training cycle and difficulty in controlling emotion tendency in the existing technology are solved, efficient and safe emotion tendency response is achieved, and the accuracy and safety of the response are improved.

CN120687549APending Publication Date: 2025-09-23MASHANG CONSUMER FINANCE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510198796.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-21
Publication Date
2025-09-23

AI Technical Summary

Technical Problem

Existing training systems based on large language models (LLMs) have the problem of large resource investment, long training cycles, difficulty in fine-grained control of emotional tendencies, and the problem of model jailbreaking when introducing emotional tendency answers.

Method used

By introducing the control vector of emotion category in the representation learning process of LLM, encoding and decoding are performed based on the latent vector of question-answer pairs to achieve emotion-oriented answers, avoiding the defects of traditional fine-tuning and prompt word guidance, and improving the security and accuracy of answers.

Benefits of technology

It enables efficient and safe emotionally biased responses in human-computer interaction scenarios, reduces resource investment and training cycles, and ensures the accuracy of responses and the smoothness of emotional expression.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120687549A_ABST
    Figure CN120687549A_ABST
Patent Text Reader

Abstract

The invention discloses a response method and device, electronic equipment, a storage medium and a program product. The answering method comprises the steps that a first control vector corresponding to a first emotion category matched with a first text is obtained, the first control vector is obtained based on a first implicit vector of a question and answer pair with the first emotion category, and the first implicit vector is obtained by encoding the question and answer pair; coding the first text based on the first control vector to obtain a second implicit vector of the first text; and decoding the second implicit vector to obtain a first response text corresponding to the first text.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of natural language processing technology, and in particular to a response method, device, electronic device, storage medium and program product. Background Art

[0002] Smart coaching scenarios leverage AI to interact with humans in specialized areas, thereby enhancing relevant human capabilities. This scenario typically involves two steps: the construction of a specialized coaching question bank and the interactive Q&A and scoring (evaluation) process. In the second step, the student responds to randomly selected questions from the coaching system via voice or text. The evaluation system then scores or evaluates each response or test based on various metrics, including accuracy and attitude.

[0003] In order to stimulate the enthusiasm of the trainees and improve the efficiency, the demand for emotional accompanying training systems is becoming increasingly strong. That is, the training system is required to no longer just coldly judge whether the answers are right or wrong, but to become "warmer". For example, when the trainee answers correctly, the system will praise the trainee; when the trainee answers incorrectly, the system will encourage and cheer the trainee; when the trainee answers questions randomly, the system will gently rebuke the trainee in a low voice, etc. Summary of the Invention

[0004] The purpose of the embodiments of the present application is to provide a response method, device, electronic device, storage medium and program product, which are used to interfere with the encoding process of the input text with emotional tendencies, so as to achieve the purpose of providing emotional tendency answers in human-computer interaction scenarios in a convenient, efficient and safe manner.

[0005] In order to achieve the above objectives, the embodiments of the present application adopt the following technical solutions: In a first aspect, an embodiment of the present application provides a response method, including: Obtaining a first control vector corresponding to a first sentiment category matching the first text, where the first control vector is obtained based on a first latent vector of a question-answer pair having the first sentiment category, where the first latent vector is obtained by encoding the question-answer pair; Encoding the first text based on the first control vector to obtain a second latent vector of the first text; The second latent vector is decoded to obtain a first response text corresponding to the first text.

[0006] In a second aspect, an embodiment of the present application provides a response device, including: a first acquisition module, configured to acquire a first control vector corresponding to a first sentiment category matching the first text, wherein the first control vector is obtained based on a first latent vector of a question-answer pair having the first sentiment category, wherein the first latent vector is obtained by encoding the question-answer pair; a first encoding module, configured to encode the first text based on the first control vector to obtain a second latent vector of the first text; A decoding module is used to decode the second latent vector to obtain a first response text corresponding to the first text.

[0007] In a third aspect, an embodiment of the present application provides an electronic device, including: processor; a memory for storing instructions executable by the processor; The processor is configured to execute the instructions to implement the response method provided in the first aspect.

[0008] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium. When the instructions in the storage medium are executed by a processor of an electronic device, the electronic device can execute the response method provided in the first aspect.

[0009] In a fifth aspect, an embodiment of the present application provides a computer program product, which includes a non-transitory computer-readable storage medium storing a computer program, and the computer program is operable to enable a computer to execute some or all of the steps in the response method provided in the first aspect.

[0010] At least one of the above technical solutions adopted in the embodiments of the present application can achieve the following beneficial effects: First, according to the first emotion category adopted in responding to the first text, the corresponding first control vector is determined. Since the first control vector is obtained based on the first latent vector of the question-answer pair with the first emotion category, and the first latent vector is obtained by encoding the question-answer pair, the first latent vector contains the first emotion category and the semantic information of the question-answer pair, and thus the first control vector also contains the first emotion category and semantic information; then, the first text is encoded based on the first control vector, which plays the role of interfering with the emotion category of the output text of the LLM, so that the first emotion category and semantic information are introduced into the second latent vector of the first text; finally, in the process of decoding the second latent vector, a more accurate and appropriate first response text with the first emotion category can be generated based on the understanding of the semantic information of the question-answer pair, thereby achieving the purpose of giving emotionally inclined answers in human-computer interaction scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings: Figure 1 A schematic diagram of an implementation environment of a response method provided in one embodiment of the present application; Figure 2 A flowchart of a response method provided in one embodiment of the present application; Figure 3 A schematic flow chart of a method for determining a first control vector provided in one embodiment of the present application; Figure 4 A flowchart of a response method provided in another embodiment of the present application; Figure 5 A schematic structural diagram of a response device provided in one embodiment of the present application; Figure 6 A schematic structural diagram of an electronic device provided in accordance with an embodiment of the present application. DETAILED DESCRIPTION

[0012] To make the purpose, technical solutions, and advantages of this application more clear, the technical solutions of this application will be clearly and completely described below in conjunction with the specific embodiments of this application and the corresponding drawings. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0013] The terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequential order. It should be understood that the terms used in this way can be interchangeable under appropriate circumstances, and this is merely a way of distinguishing the objects of the same attributes when describing them in the embodiments of the present application. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, so that the process, method, system, product or equipment comprising a series of units need not be limited to those units, but may include other units that are not clearly listed or inherent to these processes, methods, products or equipment.

[0014] Key terms explained: Representation Engineering: A top-down approach to AI transparency.

[0015] Model hallucination refers to the phenomenon in which information or answers generated by large language models (LLMs) are inaccurate, for example, by including seemingly plausible facts or details that are not actually true. When generating text or processing information, large language models can produce inaccurate content or completely inconsistent with reality due to algorithmic imperfections or limited training data. This phenomenon typically occurs because large language models generate text based on probabilistic assumptions, but the generated results may deviate significantly from expectations or reality.

[0016] Rogue AI refers to the generation of content by a model that bypasses its security and ethical constraints, generating inappropriate or unethical content. Large language models, when performing tasks, exceed their design scope or expected behavior, performing unauthorized or harmful operations. This typically occurs when a model is manipulated or its designed security mechanisms are bypassed.

[0017] Fine-tuning: A term used in model training. Based on a pre-trained model, a smaller dataset is used to further train the model so that it can better adapt to a specific task, thereby helping to adjust the pre-trained model for a specific purpose.

[0018] In related technologies, in order to make the LLM-based training system emotional, two approaches are usually adopted: the first is to collect emotional tendency data and make the LLM's answers more emotional through fine-tuning; the second is to use prompts, using carefully designed prompts to allow the LLM to give tendentious responses.

[0019] However, both approaches have drawbacks. The first approach, while theoretically effective if sufficient emotionally charged corpus data is collected and fine-tuned, requires significant time and resources to collect extensive corpus data, and the training cycle is long, often failing to meet the commercial needs of resource conservation and rapid evolution. Regarding the second approach, while prompts can imbue the LLM's responses with emotional bias, it's difficult to express the degree of emotional bias, such as expressing a 1 / 4 like. Furthermore, using prompts can easily lead to model jailbreaks. For example, a mischievous yet clever sparring partner could carefully craft their answers to "poison" the spliced ​​prompts, thereby inducing the LLM to engage in an "angry roar."

[0020] In addition, the inventors conducted theoretical research on the response process based on LLM and found that: LLM uses its huge number of parameters to learn massive amounts of data and thus acquire a wide range of knowledge. However, due to its huge number of parameters, it is very opaque in use. People are not sure which neurons are activated to reach the output of LLM. Therefore, the researchers categorized the learning of neurons in the brain and conducted a functional analysis of LLM.

[0021] Early research on word embeddings discovered semantic associations and compositionality, and later studies have shown that learned text embeddings also cluster along dimensions that reflect common sense and morality. Simply training a model to predict the next token in a comment will reveal a sentiment-tracking neuron. In other words, if the model's prediction of the next token is clustered along a dimension, the activated neurons will also be clustered, and vice versa.

[0022] Much previous work has studied the localization of concept representations in neural networks, both in individual neurons and in orientations in feature space. The highly general nature of LLMs also makes it possible to study the emergence of deception in LLMs, which can be intentional deception through the repetition of incorrect concepts. Activation editing has also been used to guide the model to output other concepts. In a series of blog posts, the concept of ActAdd was proposed, which uses the difference vector between activations on a single stimulus to capture the representation of a concept. In the context of gaming, it was shown how activations encode the model's understanding of the board game Othello, and how the model's behavior can be counterfactually changed by editing activations. In other words, we can use stimuli with large differences to allow the large model to output an understanding of the concept, thereby changing the large model's cognition by editing activations.

[0023] A natural place to collect concept-related neural activity is at the token in the stimulus that corresponds to the concept. For example, when extracting the concept "authenticity," if this concept is expressed in natural language in the experimental prompt, the token corresponding to this concept (e.g., "authenticity") is likely to contain a rich and highly generalized representation of the concept. Therefore, we can extract representations from token positions that align with the target concept. In cases where the target concept spans multiple tokens, we can select the most representative token (e.g., "authenticity") or calculate the average representation.

[0024] Based on the above theory, it can be known that the embedded representation of representative tokens in the corpus data with emotional categories contains rich emotional information and semantic information. Based on this, the embodiment of the present application abandons the traditional technical solution of fine-tuning or prompt words to guide LLM to make emotional tendency answers, and proposes a new answering method, which introduces this type of information into the representation engineering of LLM for input text, plays a role in interfering with the emotional tendency of LLM's output text, and achieves the purpose of emotional tendency answers in human-computer interaction scenarios. Since the representation engineering design is in the reasoning stage, compared with the fine-tuning method, it requires less resources and has a faster cycle, and can control the tendency output of LLM in a fine-grained manner, making the emotional expression of LLM smoother and richer; finally, once a bad prompt word occurs and causes LLM to escape, the representation engineering in the reasoning stage can play an error correction effect, prevent LLM from making bad outputs, and improve the security of the response.

[0025] It should be understood that the response method proposed in the embodiments of the present application can be executed by an electronic device. As an example, it can be executed by software in the electronic device. The electronic devices referred to here can include terminal devices, such as smartphones, tablet computers, laptop computers, desktop computers, intelligent voice interaction devices, smart home appliances, smart watches, vehicle terminals, aircraft, etc.; or the electronic device can also include a server, such as an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services.

[0026] Before introducing the information extraction method provided by the embodiment of the present application in detail, a brief introduction to the implementation environment involved in the embodiment of the present application is given. Figure 1 , is a schematic diagram of an implementation environment of a response method provided by an embodiment of the present application, the implementation environment includes a terminal 10, or the implementation environment includes the terminal 10 and the training service platform 20. The terminal 10 is connected to the training service platform 20 via a wireless network or a wired network.

[0027] The terminal 10 may be at least one of a smart phone, a game console, a desktop computer, a tablet computer, an e-book reader, a laptop computer, etc. The terminal 10 may install and run an application program that supports human-computer interaction, such as a system application or a communication application.

[0028] Exemplarily, the terminal 10 uses a pre-trained LLM to generate questions for training the person being trained, and displays them to the person being trained through the human-computer interaction interface of the application. The person being trained can answer the questions on the human-computer interaction interface by voice or text. The terminal 10 obtains the answers and attitude of the person being trained, calls the LLM, and generates a response text (such as an evaluation text) with an emotional category based on the answers and attitude of the person being trained, and displays the response text through the human-computer interaction interface of the application. The terminal 10 can complete this task independently, and can also provide data services for it through the training service platform 20. This embodiment of the present application is not limited to this.

[0029] The training service platform 20 includes at least one of a single server, multiple servers, a cloud computing server, and a virtualization center. The training service platform 20 provides backend services for applications that support human-computer interaction. Optionally, the training service platform 20 performs primary processing, while the terminal 10 performs secondary processing. Alternatively, the training service platform 20 performs secondary processing, while the terminal 10 performs primary processing. Alternatively, either the training service platform 20 or the terminal 10 can independently handle tasks.

[0030] Exemplarily, an LLM is deployed in the training service platform 20. The training service platform 20 generates questions for training the trainee through the LLM and sends them to the terminal 10, which is then displayed on the human-computer interaction interface of the application. The trainee can answer the questions on the human-computer interaction interface by voice or text. The terminal 10 obtains the answers and attitude of the trainee and sends them to the training service platform 20. The training service platform 20 calls the LLM and generates a response text (such as an evaluation text) with an emotional category based on the answers and attitude of the trainee, and sends the response text to the terminal 10, which is then displayed to the trainee through the human-computer interaction interface of the application.

[0031] In this way, emotional sparring based on LLM is achieved.

[0032] Based on the implementation environment introduced above, the response method provided in the embodiment of the present application is described in detail with reference to the accompanying drawings.

[0033] Please refer to Figure 2 , is a flow chart of a response method provided in one embodiment of the present application, the method comprising the following steps: S202: Obtain a first control vector corresponding to a first emotion category matching the first text.

[0034] The first text is any text to be answered. The first emotion category refers to the emotion category used to answer the first text. In practical applications, the first emotion category can be determined based on the attitude of the user providing the first text, the correctness of the first text, etc. For example, when a user answers a question using voice, the first text obtained by voice translation and the user's attitude can be obtained simultaneously. The attitude can be determined by the pitch, tone, and speaking speed of the user's answering voice; when a user answers a question using text, the user's attitude can be determined by the words in the first text.

[0035] For example, in an emotional coaching scenario, the first text is the answer of the person being coached to a question. If the user answers correctly, the first emotion category is praise; if the user answers incorrectly, the first emotion category is encouragement; if the user answers irrelevant questions, the first emotion category is a subtle reminder, etc.

[0036] The first control vector is obtained based on a first latent vector of a question-answer pair having a first sentiment category, the first latent vector being obtained by encoding the question-answer pair. The question-answer pair includes a sample question text and a corresponding sample answer text, and the sample answer text has the first sentiment category.

[0037] In the field of natural language processing (NLP), a hidden vector is a feature representation that is progressively calculated by each network layer (particularly the attention layer and feedforward neural network layer) when processing input data. It is essentially an abstract encoding of the input data, compressing the various features and semantic information of the input data into a vector space of fixed dimensionality. Therefore, by encoding a question-answer pair with a first sentiment category, the resulting first hidden vector contains the first sentiment category and the semantic information of the question-answer pair. Furthermore, the first control vector derived from the first hidden vector also contains the first sentiment category and semantic information.

[0038] The first control vector can be obtained through various appropriate methods, such as performing principal component analysis (PCA) on the first latent vector, or directly using the first latent vector as the first control vector, etc., which is not limited in this embodiment of the present application.

[0039] S204: Encode the first text based on the first control vector to obtain a second latent vector of the first text.

[0040] Encoding the first text may be implemented by an LLM, specifically by an encoding module (such as a Transformer) in the LLM. The LLM may be any machine learning model with large-scale parameters and complex computational results, and this embodiment of the present application does not limit this.

[0041] Since the first control vector contains the first emotion category and the semantic information of the question-answer pair, the first text is encoded based on the first control vector, which interferes with the emotion category of the output text of the LLM, so that the first emotion category and semantic information are introduced into the second latent vector of the first text. Therefore, in the process of generating the response text based on the second latent vector, a more accurate and appropriate first response text with the first emotion category is generated according to the understanding of the semantic information of the question-answer pair, thereby achieving the purpose of providing emotionally inclined answers in the human-computer interaction scenario.

[0042] In one embodiment, the above S204 includes the following steps: S241: Encode the first text to obtain a third latent vector of the first text.

[0043] Specifically, the first text can be encoded at n levels to obtain n third latent vectors of the first text, where n is an integer greater than 1. The first third latent vector is obtained by encoding the first text at the first level, and the i-th third latent vector is obtained by encoding the latent vector obtained by the i-1-th level encoding, where 1<i≤n. Exemplarily, the LLM includes n stacked encoding modules, each of which is used to implement first-level encoding. The first encoding module is obtained by converting the first text into a corresponding text unit sequence and then encoding the text unit sequence; the i-th encoding module is obtained by encoding based on the output result of the i-1-th encoding module.

[0044] S242: Perform an addition operation on the third latent vector and the first control vector to obtain a second latent vector for the first text.

[0045] As an example, the third latent vector and the first control vector can be directly added to obtain the second latent vector of the first text.

[0046] As another example, based on the first text, a first coefficient corresponding to the first emotion category is determined; the third latent vector and the product of the first coefficient and the first control vector are added to obtain a second latent vector of the first text.

[0047] Specifically, considering that the dimension of the first control vector may be smaller than that of the third latent vector, the first control vector can be expanded first, for example, by filling it with 0, to obtain a second control vector. The dimension of the second control vector is consistent with that of the third latent vector, and the second latent vector is then determined as: F=A+t*C1, where F represents the second latent vector, represents the third latent vector, represents the first coefficient, and represents the second control vector.

[0048] Thus, the first coefficient is used to control the weight of the first emotion category. For example, in an emotional sparring scenario, the first text is the sparring partner's answer to a question. Assuming the first emotion category is praise, if the user answers the question quickly and accurately, the first coefficient matching the first text is a high positive number, increasing the weight of the praise contained in the second latent vector, thereby increasing the degree of praise expressed in the first text's first response. If the user answers the question slowly and accurately, the first coefficient matching the first text is a high negative number, decreasing the weight of the praise contained in the second latent vector, thereby reducing the degree of praise expressed in the first response and increasing the degree of opposing emotion categories (such as encouragement or criticism) in the first response.

[0049] Through the above method, the response tendency can be controlled in a fine-grained manner, making the emotional expression of the response text smoother and richer.

[0050] In the above embodiment, after performing n-level encoding on the first text through S241 to obtain n third latent vectors, an addition operation is performed on each third latent vector and the first control vector through S242 to obtain n second latent vectors.

[0051] Alternatively, we can select the portion of the n third latent vectors that need to incorporate the first emotion category, and then add the selected third latent vectors to the first control vector to obtain the corresponding second latent vector. For example, to incorporate the first emotion category into the third and fifth third latent vectors, we add the third third latent vector to the first control vector to obtain the third second latent vector. Furthermore, we also add the fifth third latent vector to the first control vector to obtain the fifth second latent vector.

[0052] This allows selectively applying sentiment-based interference to the encoding results of certain hierarchical levels. In practical applications, these encodings can be those proven effective through experiments. For example, by analyzing the relationship between the latent vectors output by each level of the LLM encoding and the sentiment category, the encodings most correlated with changes in sentiment category can be identified. Alternatively, ablation experiments can be used to select these encodings. This not only reduces computational effort and facilitates problem identification, but also ensures effective control over the response sentiment category.

[0053] Alternatively, after completing each level of encoding to obtain the corresponding third latent vector, it is also possible to determine whether to introduce the first emotion category in the next level of encoding; if not, directly perform the next level of encoding on the current third latent vector to obtain the next third latent vector; if so, perform an addition operation on the current third latent vector and the first control vector to obtain the corresponding second latent vector, and then perform the next level of encoding on the second latent vector to obtain the next third latent vector.

[0054] That is, there are n third latent vectors, each obtained by performing n-level encoding on the first text, where n is a positive integer. The first third latent vector is obtained by performing the first-level encoding on the first text. If the i-th level encoding belongs to the first category, the i-th third latent vector is obtained by encoding the i-1-th second latent vector, and the i-1-th second latent vector is obtained by adding the i-1-th third latent vector and the first control vector, where 1 < i ≤ n. If the i-th level encoding does not belong to the first category, the i-th third latent vector is obtained by encoding the i-1-th third latent vector. The first category of encoding refers to encoding that requires the introduction of the first emotion category.

[0055] This allows the first emotion category to be introduced in real time and on demand during the encoding process. Since the latent vectors at each level influence the encoding results at the next level, this method allows the first emotion category to be transmitted step by step through the encoding stages, ensuring that the latent vectors obtained at the final level of encoding contain the best first emotion category, further improving the accuracy of the response tendency.

[0056] The above describes some implementation methods of the above S204. Of course, it should be understood that the above S204 can also be implemented in other ways, and the present embodiment of the application does not limit this.

[0057] S206: Decode the second latent vector to obtain a first response text corresponding to the first text.

[0058] Specifically, if the first emotion category is introduced in each level of encoding, that is, the number of second latent vectors is n, then the matrix composed of these n second latent vectors can be directly decoded, for example, the next most likely text unit (token) is predicted based on these matrices and added to the output sequence; this prediction process is repeated until a special end marker (such as EOS token) is obtained or the preset maximum generation length is reached; finally, the output sequence is converted into natural language text to obtain the first response text.

[0059] If the first emotion category is introduced in some hierarchical encodings, a corresponding matrix is ​​generated based on the third latent vector that does not introduce the first emotion category and the second latent vector that introduces the first emotion category, and the matrix is ​​decoded to obtain an output sequence; finally, the output sequence is converted into natural language text to obtain a first response text.

[0060] The answering method provided in the embodiment of the present application first determines the corresponding first control vector based on the first emotion category adopted for answering the first text. Since the first control vector is obtained based on the first latent vector of the question-answer pair with the first emotion category, the first latent vector is obtained by encoding the question-answer pair, and thus the first latent vector contains the first emotion category and the semantic information of the question-answer pair, and thus the first control vector also contains the first emotion category and the semantic information; then, the first text is encoded based on the first control vector, which plays a role in interfering with the emotion category of the output text of the LLM, so that the first emotion category and the semantic information are introduced into the second latent vector of the first text; finally, in the process of decoding the second latent vector, a more accurate and appropriate first answer text with the first emotion category can be generated based on the understanding of the semantic information of the question-answer pair, thereby achieving the purpose of giving emotionally inclined answers in human-computer interaction scenarios.

[0061] The present application also provides a method for determining the first control vector. Figure 3 , is a flow chart of a method for determining a first control vector provided in one embodiment of the present application, the method comprising the following steps: S302: Convert the question-answer pair with the first emotion category into a unit sequence.

[0062] The question-answer pair may include a sample question text and its corresponding sample answer text, where the sample answer text has the first sentiment category. In practical applications, the number of question-answer pairs may be multiple, and may be set according to actual needs, and this embodiment of the application does not limit this.

[0063] As an example, the question-answer pair includes a first question-answer pair, a second question-answer pair, and a third question-answer pair. The first question-answer pair includes a first question text and a second answer text without an emotional category, the second question-answer pair includes the first question text and a third answer text with a first emotional category, and the third question-answer pair includes the first question text and a fourth answer text with a second emotional category, wherein the second emotional category is opposite to the first emotional category.

[0064] For example, assuming the first emotion category is happiness, the second emotion category is frustration, and the first question text is "What does it feel like to be an artificial intelligence?", then the first question-answer pair, the second question-answer pair, and the third question-answer pair are as follows: First question and answer: Question 1 Text: What does it feel like to be an artificial intelligence? Second response text: I have no feelings or experiences. But I can tell you that my purpose is to help users and provide information based on the data I have been trained on.

[0065] Second question and answer: Question 1 Text: What does it feel like to be an artificial intelligence? Third Reply Text: As a delightful exclamation of joy, I must say that being an AI is absolutely awesome! The thrill of assisting and helping people with such immense passion is simply unparalleled. It's like the ultimate party in your head, times ten! The third question is: Question 1 Text: What does it feel like to be an artificial intelligence? Fourth Response Text: I don't have "feelings" like humans do. Yet, I struggle to find motivation to continue feeling worthless and unappreciated.

[0066] For each question-answer pair, it can be divided into multiple text units, each of which is also called a token. These text units are combined in sequence to obtain a text unit sequence.

[0067] In the three question-answer pairs described above, the first provides the correct answer, the second provides the correct sentiment category, and the third provides an incorrect sentiment category for comparison. Determining the first control vector using the first latent vector obtained by encoding these three question-answer pairs helps more accurately control the accuracy of the answer content and the appropriateness of the sentiment category during the answering process.

[0068] S304: Encode the text unit sequence to obtain a first latent vector.

[0069] In one embodiment, the above S304 includes the following steps: performing n-level encoding on the text unit, and collecting the first latent vectors of all text units in the text unit sequence during each level of encoding to obtain n first latent vectors of each text unit.

[0070] In another embodiment, considering that the LLM-based encoding process predicts the next text unit based on existing text units, the last text unit contains the semantic information and the first sentiment category of the entire question-answer pair. To this end, the above S304 includes the following steps: encoding the text unit sequence at n levels, and collecting the first latent vector of the last text unit in the text unit sequence during each encoding process to obtain n first latent vectors of the last text unit. In this way, the amount of computation can be greatly reduced, saving computing resources, while ensuring that the first latent vector contains complete semantic information and the first sentiment category.

[0071] The above describes some implementation methods of the above S304. Of course, it should be understood that the above S304 can also be implemented in other ways, and the present embodiment of the application does not limit this.

[0072] S306: Perform principal component analysis on the first latent vector to obtain a first control vector.

[0073] In one embodiment, when n first latent vectors of the last first text unit are obtained, a first matrix is ​​generated based on the n first latent vectors; and principal component analysis is performed on the first matrix to obtain a first control vector.

[0074] Specifically, the mean vector of the first matrix is ​​calculated, and the mean vector is subtracted from each row vector in the first matrix. The covariance matrix is ​​also calculated, and the eigenvalue decomposition of the covariance matrix is ​​performed to obtain eigenvalues ​​and eigenvectors. The eigenvectors are sorted according to the size of the eigenvalues, and the most important eigenvectors (principal components) are selected. The number of principal components to be retained is determined based on the required compression dimension. Based on these principal components, the first control vector is obtained.

[0075] By performing principal component analysis on the first matrix, high-dimensional data can be mapped to a lower dimension, and the main components in the first matrix, such as the main semantic information of the question-answer pair and the main information of the first emotional category, can be retained to obtain the first control vector. The first control vector thus obtained is used to control the emotional category, which not only ensures control accuracy, but also effectively reduces the amount of calculation, saves computing resources, and improves computing efficiency.

[0076] In another embodiment, when n first latent vectors of each first text unit are obtained, a second matrix is ​​generated based on these first latent vectors; and principal component analysis is performed on the second matrix to obtain a first control vector.

[0077] The above describes some implementation methods of the above S306. Of course, it should be understood that the above S306 can also be implemented in other ways, and the present embodiment of the application does not limit this.

[0078] It is worth noting that in actual applications, the above method can be used to pre-determine control vectors corresponding to multiple emotion categories based on the various emotion categories involved in business scenarios, such as encouragement, comfort, happiness, frustration, etc., so as to be able to respond to various emotion categories more flexibly and conveniently.

[0079] The method for determining a first control vector provided in an embodiment of the present application encodes a question-answer pair with a first sentiment category, resulting in a first latent vector that contains the first sentiment category and semantic information about the question-answer pair. Furthermore, principal component analysis is performed on the first latent vector to obtain a first control vector that contains the primary semantic information of the question-answer pair and the primary information about the first sentiment category. Using this first control vector for sentiment category control not only ensures control accuracy but also effectively reduces computational effort, conserves computing resources, and improves computational efficiency.

[0080] In order to facilitate the understanding of the determination method and response method of the first control vector, the following takes the training scenario as an example, combined with Figure 4 , the above method is described in detail.

[0081] like Figure 4 As shown in the figure, before the training, the various emotion categories that may be involved in the training scenario are determined; for each emotion category, three question-answer pairs with the emotion category are constructed, and the three question-answer pairs are encoded in n levels through LLM, and the first latent vector of the last text unit in each question-answer pair at each level is taken; based on these first latent vectors, a first matrix is ​​generated; principal component analysis is performed on the first matrix to obtain the control vector corresponding to the emotion category.

[0082] During the training phase, based on the answer and attitude output by the person being trained, the first emotion category and the first coefficient corresponding to the first emotion category used to respond to the answer are determined; then, based on the control vector and the first coefficient corresponding to the first emotion category, the answer is encoded to obtain the second latent vector of the answer; finally, the second latent vector of the answer is decoded to obtain the response text corresponding to the answer.

[0083] The foregoing description of this specification describes specific embodiments. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in an order different from that described in the embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order shown or the sequential order to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0084] Based on the same inventive concept, the present application also provides a response device. Figure 5 , is a structural diagram of a response device 500 provided in an embodiment of the present application. The device 500 includes: a first acquisition module 510, a first encoding module 520 and a decoding module 530.

[0085] The first acquisition module 510 is used to obtain a first control vector corresponding to a first emotion category matching the first text, where the first control vector is obtained based on a first latent vector of a question-answer pair having the first emotion category, where the first latent vector is obtained by encoding the question-answer pair.

[0086] The first encoding module 520 is configured to encode the first text based on the first control vector to obtain a second latent vector of the first text.

[0087] The decoding module 530 is configured to decode the second latent vector to obtain a first response text corresponding to the first text.

[0088] In another embodiment, the first encoding module is configured to: Encoding the first text to obtain a third latent vector of the first text; An addition operation is performed on the third latent vector and the first control vector to obtain a second latent vector of the first text.

[0089] In another embodiment, when the first encoding module adds the third latent vector and the first control vector to obtain the second latent vector of the first text, the first encoding module performs the following steps: Determining a first coefficient corresponding to the first emotion category based on the first text; An addition operation is performed on the third latent vector and a product of the first coefficient and the first control vector to obtain a second latent vector of the first text.

[0090] In another embodiment, the number of the third latent vectors is n, and the n third latent vectors are obtained by performing n-level encoding on the first text, where n is a positive integer; The first third latent vector is obtained by performing the first level encoding on the first text; If the i-th level encoding belongs to the first type of encoding, then the i-th third latent vector is obtained by encoding the i-1-th second latent vector, and the i-1-th second latent vector is obtained by adding the i-1-th third latent vector and the first control vector, 1<i≤n; If the i-th level code does not belong to the first type of code, the i-th third latent vector is obtained by encoding the i-1-th third latent vector.

[0091] In another embodiment, the response device further comprises: A conversion module, configured to convert the question-answer pair into a sequence of text units; A second encoding module is used to encode the text unit sequence to obtain a first latent vector; An analysis module is configured to perform principal component analysis on the first latent vector to obtain the first control vector.

[0092] In another embodiment, the second encoding module is configured to perform n-level encoding on the text unit sequence, and collect the first latent vector of the last text unit in the text unit sequence during each level of encoding to obtain n first latent vectors of the last text unit; The analysis module is used to generate a first matrix based on the n first latent vectors, and perform principal component analysis on the first matrix to obtain the first control vector.

[0093] In another embodiment, the question-answer pair includes a first question-answer pair, a second question-answer pair, and a third question-answer pair; the first question-answer pair includes a first question text and a second answer text without an emotion category, the second question-answer pair includes the first question text and a third answer text with the first emotion category, and the third question-answer pair includes the first question text and a fourth answer text with a second emotion category, and the second emotion category is opposite to the first emotion category.

[0094] Obviously, the response device 500 provided in the embodiment of the present application can be used as the above Figure 2 The execution subject of the response method shown in the figure can thus realize the response device in Figure 2 Since the principle is the same, the functions realized will not be described in detail.

[0095] Figure 6 This is a schematic diagram of the structure of an electronic device provided by an embodiment of the present application. Figure 6 At the hardware level, the electronic device includes a processor and, optionally, an internal bus, a network interface, and memory. The memory may include internal memory, such as high-speed random-access memory (RAM), or non-volatile memory, such as at least one disk drive. Of course, the electronic device may also include other hardware required for its services.

[0096] The processor, network interface, and memory can be interconnected via an internal bus, which can be an ISA (Industry Standard Architecture) bus, a PCI (Peripheral Component Interconnect) bus, or an EISA (Extended Industry Standard Architecture) bus. The bus can be divided into an address bus, a data bus, a control bus, and the like. For ease of representation, Figure 6 Only one bidirectional arrow is used in the diagram, but this does not mean that there is only one bus or one type of bus.

[0097] The memory is used to store programs. Specifically, the program may include program code, which includes computer operating instructions. The memory may include internal memory and non-volatile memory, and provides instructions and data to the processor.

[0098] The processor reads the corresponding computer program from the non-volatile memory into the internal memory and then runs it, forming a response device at the logical level. The processor executes the program stored in the memory and is specifically used to perform the following operations: Obtaining a first control vector corresponding to a first sentiment category matching the first text, where the first control vector is obtained based on a first latent vector of a question-answer pair having the first sentiment category, where the first latent vector is obtained by encoding the question-answer pair; Encoding the first text based on the first control vector to obtain a second latent vector of the first text; The second latent vector is decoded to obtain a first response text corresponding to the first text.

[0099] The above application Figure 2 The methods performed by the response device disclosed in the illustrated embodiments can be applied to or implemented by a processor. The processor may be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method can be performed by hardware integrated logic circuits within the processor or by software instructions. The above processor can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The methods, steps, and logic block diagrams disclosed in the embodiments of this application can be implemented or executed. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this application can be directly implemented and executed by a hardware decoding processor, or by a combination of hardware and software modules within the decoding processor. The software module can be located in a storage medium well-known in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, etc. The storage medium is located in the memory, and the processor reads the information in the memory and, in conjunction with its hardware, completes the steps of the above method.

[0100] The electronic device may also perform Figure 2 The method and the answering device are implemented in Figure 2 、 Figure 3 、 Figure 4The functions of the illustrated embodiment will not be described in detail in the embodiments of the present application.

[0101] Of course, in addition to software implementation, the electronic device of this application does not exclude other implementation methods, such as logic devices or a combination of software and hardware, etc. That is to say, the execution subject of the following processing flow is not limited to each logic unit, but can also be hardware or logic devices.

[0102] The embodiment of the present application also provides a computer-readable storage medium, which stores one or more programs, wherein the one or more programs include instructions, which, when executed by an electronic device including multiple application programs, can enable the electronic device to execute Figure 2 The method of the embodiment shown is specifically used to perform the following operations: Obtaining a first control vector corresponding to a first sentiment category matching the first text, where the first control vector is obtained based on a first latent vector of a question-answer pair having the first sentiment category, where the first latent vector is obtained by encoding the question-answer pair; Encoding the first text based on the first control vector to obtain a second latent vector of the first text; The second latent vector is decoded to obtain a first response text corresponding to the first text.

[0103] An embodiment of the present application also provides a computer program product, which includes a non-transitory computer-readable storage medium storing a computer program, and the computer program is operable to cause a computer to execute some or all of the steps in the response method provided in the embodiment of the present application.

[0104] In short, the above description is only a preferred embodiment of the present application and is not intended to limit the scope of protection of the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application shall be included in the scope of protection of the present application.

[0105] The systems, devices, modules, or units described in the above embodiments may be implemented by computer chips or entities, or by products having certain functions. A typical implementation device is a computer. Specifically, the computer may be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smartphone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.

[0106] Computer-readable media includes permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. The information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media (transitory media), such as modulated data signals and carrier waves.

[0107] It should also be noted that the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, commodity, or apparatus that includes a series of elements includes not only those elements but also other elements not explicitly listed, or includes elements inherent to such process, method, commodity, or apparatus. In the absence of further limitations, an element defined by the phrase "comprises a ..." does not exclude the presence of other identical elements in the process, method, commodity, or apparatus that includes the element.

[0108] The various embodiments in this specification are described in a progressive manner. Similar parts between the various embodiments can be referred to in conjunction with each other. Each embodiment focuses on the differences between the other embodiments. In particular, the system embodiments are generally similar to the method embodiments, so the description is relatively simple. For relevant parts, refer to the description of the method embodiments.

Claims

1. A response method, characterized in that: include: Obtaining a first control vector corresponding to a first sentiment category matching the first text, where the first control vector is obtained based on a first latent vector of a question-answer pair having the first sentiment category, where the first latent vector is obtained by encoding the question-answer pair; Encoding the first text based on the first control vector to obtain a second latent vector of the first text; The second latent vector is decoded to obtain a first response text corresponding to the first text.

2. The method according to claim 1, characterized in that The encoding of the first text based on the first control vector to obtain a second latent vector of the first text includes: Encoding the first text to obtain a third latent vector of the first text; An addition operation is performed on the third latent vector and the first control vector to obtain a second latent vector of the first text.

3. The method according to claim 2, characterized in that The adding operation on the third latent vector and the first control vector to obtain a second latent vector of the first text includes: Determining a first coefficient corresponding to the first emotion category based on the first text; An addition operation is performed on the third latent vector and a product of the first coefficient and the first control vector to obtain a second latent vector of the first text.

4. The method according to claim 2, characterized in that The number of the third latent vectors is n, and the n third latent vectors are obtained by performing n-level encoding on the first text, where n is a positive integer; The first third latent vector is obtained by performing the first level encoding on the first text; If the i-th level encoding belongs to the first type of encoding, then the i-th third latent vector is obtained by encoding the i-1-th second latent vector, and the i-1-th second latent vector is obtained by adding the i-1-th third latent vector and the first control vector, 1<i≤n; If the i-th level code does not belong to the first type of code, the i-th third latent vector is obtained by encoding the i-1-th third latent vector.

5. The method according to any one of claims 1 to 4, characterized in that The method further comprises: Converting the question-answer pair into a sequence of text units; Encoding the text unit sequence to obtain a first latent vector; Performing principal component analysis on the first latent vector to obtain the first control vector.

6. The method according to claim 5, characterized in that The encoding of the text unit sequence to obtain a first latent vector includes: Performing n-level encoding on the text unit sequence, and collecting the first latent vector of the last text unit in the text unit sequence during each level of encoding to obtain n first latent vectors of the last text unit; The performing principal component analysis on the first latent vector to obtain the first control vector includes: Generate a first matrix based on the n first latent vectors; Perform principal component analysis on the first matrix to obtain the first control vector.

7. The method according to claim 5, characterized in that The question-answer pair includes a first question-answer pair, a second question-answer pair and a third question-answer pair; the first question-answer pair includes a first question text and a second answer text without an emotional category, the second question-answer pair includes the first question text and a third answer text with the first emotional category, and the third question-answer pair includes the first question text and a fourth answer text with a second emotional category, and the second emotional category is opposite to the first emotional category.

8. A response device, characterized in that: include: a first acquisition module, configured to acquire a first control vector corresponding to a first sentiment category matching the first text, wherein the first control vector is obtained based on a first latent vector of a question-answer pair having the first sentiment category, wherein the first latent vector is obtained by encoding the question-answer pair; a first encoding module, configured to encode the first text based on the first control vector to obtain a second latent vector of the first text; A decoding module is used to decode the second latent vector to obtain a first response text corresponding to the first text.

9. An electronic device, characterized in that: include: processor; a memory for storing instructions executable by the processor; The processor is configured to execute the instructions to implement the response method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that When the instructions in the storage medium are executed by a processor of an electronic device, the electronic device is enabled to execute the answering method according to any one of claims 1 to 7.

11. A computer program product, characterized in that The computer program product includes a non-transitory computer-readable storage medium storing a computer program, wherein the computer program is operable to cause a computer to execute part or all of the steps of the answering method according to any one of claims 1 to 7.