Response dialogue generation method and device, electronic equipment and storage medium
By using the guidance vector in the preset language model to update the output vector, the problem of high training cost in task-based dialogues is solved, and efficient generation of topic-consistent reply dialogues is achieved.
Patent Information
- Application Number
- CN202510971478.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-15
- Publication Date
- 2025-10-28
AI Technical Summary
In task-based dialogues, existing technologies require building high-quality labeled datasets to fine-tune language models in order to ensure topic consistency in interactive dialogues. This results in high training costs and large consumption of computing resources, and cannot guarantee the effectiveness of generating response dialogues.
By obtaining the guidance vector, the first output vector is updated according to the output vector of the interference prompt information corresponding to the first network layer in the preset language model to generate a reply dialogue with consistent topics, avoiding dependence on high-quality labeled datasets.
This approach improves the topic consistency of reply dialogues generated by language models without the need to build high-quality labeled datasets, reducing training costs and computing resource consumption.
Smart Images

Figure CN120849708A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of computer technology, and in particular to a method and apparatus for generating response dialogues, an electronic device, and a storage medium. Background Technology
[0002] Task-oriented dialogue refers to interactive dialogue that assists users in completing specific target tasks. In task-oriented dialogue, the dialogue process revolves around a clearly defined topic.
[0003] In real-world applications, users may input inquiries that deviate from the main topic of the conversation. In such cases, to ensure the consistency of the interactive dialogue's theme, a common approach is to fine-tune the language model using a high-quality labeled dataset to ensure that the generated response dialogue meets the theme consistency requirement. However, this method is not only costly to train, but also cannot guarantee the effectiveness of the generated response dialogue. Summary of the Invention
[0004] This disclosure provides a method, apparatus, electronic device, and storage medium for generating response dialogues.
[0005] In a first aspect, this disclosure provides a method for generating a response dialogue, the method comprising: in response to receiving a user query statement, inputting the user query statement into a preset language model to obtain a first output vector generated by a first network layer in the preset language model based on the user query statement; obtaining a guiding vector, the guiding vector being generated based on a second output vector of preset interference prompt information corresponding to the first network layer; wherein the interference prompt information is constructed based on a preset answer to the interference query statement; and updating the first output vector based on the guiding vector, so that subsequent network layers generate a response dialogue for the user query statement based on the updated first output vector.
[0006] Secondly, this disclosure provides an apparatus for generating a response dialogue, comprising: an input module, configured to, in response to receiving a user query, input the user query into a preset language model to obtain a first output vector generated by a first network layer in the preset language model based on the user query; an acquisition module, configured to acquire a guidance vector, the guidance vector being generated based on a second output vector of preset interference prompt information corresponding to the first network layer; wherein the interference prompt information is constructed based on a preset answer to the interference query; and a generation module, configured to update the first output vector based on the guidance vector, so that subsequent network layers generate a response dialogue to the user query based on the updated first output vector.
[0007] Thirdly, this disclosure provides an electronic device comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores one or more computer programs executable by the at least one processor, the one or more computer programs being executed by the at least one processor to enable the at least one processor to perform the above-described method for generating a response dialogue.
[0008] Fourthly, this disclosure provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the above-described method for generating a response dialogue.
[0009] The method for generating a response dialogue provided in this embodiment firstly, in response to receiving a user query, inputting the user query into a preset language model to obtain a first output vector generated by the first network layer of the preset language model based on the user query; secondly, obtaining a guiding vector, which is generated based on a second output vector of preset interference prompt information corresponding to the first network layer; wherein, the interference prompt information is constructed based on a preset answer to the interference query; finally, updating the first output vector based on the guiding vector, so that subsequent network layers generate a response dialogue to the user query based on the updated first output vector. Therefore, this embodiment does not require constructing a high-quality labeled dataset for fine-tuning the language model; it only needs to generate the guiding vector based on the output vector of the first network layer corresponding to the interference prompt information in the preset language model. Since the interference prompt information is constructed based on a preset answer to the interference query, the guiding vector generated based on the output vector of the first network layer corresponding to the interference prompt information can be used to guide the preset language model to generate a topic-consistent response dialogue. Thus, when the preset language model processes a user query, directly updating the output vector of the first network layer in the preset language model through this guiding vector can effectively improve the topic consistency of the response dialogue generated by the preset language model.
[0010] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0011] The accompanying drawings are provided to further illustrate the present disclosure and form part of the specification. They are used together with the embodiments of the present disclosure to explain the disclosure and do not constitute a limitation thereof. The above and other features and advantages will become more apparent to those skilled in the art from the detailed description of exemplary embodiments with reference to the accompanying drawings, in which:
[0012] Figure 1An application scenario diagram of the method and apparatus for generating response dialogues provided in the embodiments of this disclosure;
[0013] Figure 2 A flowchart illustrating a method for generating a response dialogue as provided in an embodiment of this disclosure;
[0014] Figure 3 A flowchart illustrating a method for generating a response dialogue provided in an embodiment of this disclosure;
[0015] Figure 4 A block diagram of a response dialogue generation apparatus provided in an embodiment of this disclosure;
[0016] Figure 5 This is a block diagram of an electronic device provided in an embodiment of the present disclosure. Detailed Implementation
[0017] To enable those skilled in the art to better understand the technical solutions of this disclosure, exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments of this disclosure to aid understanding. These should be considered merely exemplary. Therefore, those skilled in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0018] Where there is no conflict, the various embodiments of this disclosure and the features thereof in the embodiments may be combined with each other.
[0019] As used herein, the term “and / or” includes any and all combinations of one or more related enumerated entries.
[0020] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit this disclosure. As used herein, the singular forms “a” and “the” are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will also be understood that when the terms “comprising” and / or “made of” are used in this specification, the presence of the stated feature, integral, step, operation, element, and / or component is specified, but the presence or addition of one or more other features, integrals, steps, operations, elements, components, and / or groups thereof is not excluded. Words such as “connected” or “linked” are not limited to physical or mechanical connections but can include electrical connections, whether direct or indirect.
[0021] Unless otherwise specified, all terms used herein (including technical and scientific terms) have the same meaning as commonly understood by one of ordinary skill in the art. It will also be understood that terms such as those defined in commonly used dictionaries should be interpreted as having a meaning consistent with their meaning in the context of the relevant art and this disclosure, and will not be interpreted as having an idealized or overly formal meaning, unless expressly so defined herein.
[0022] The collection, storage, use, processing, transmission, provision, and disclosure of user personal information in this technical solution comply with relevant laws and regulations and do not violate public order and good morals. The use of user data in this technical solution follows relevant national laws and regulations (e.g., the "Information Security Technology - Personal Information Security Specification"). For example, appropriate measures are taken for personal information access control; restrictions are imposed on the display of personal information; the purpose of using personal information does not exceed the scope of direct or reasonable association; and explicit identity targeting is eliminated when using personal information to avoid precisely identifying specific individuals.
[0023] In task-oriented dialogue scenarios, users may input irrelevant questions that deviate from the main topic. To improve the dialogue generation performance of language models in order to enable them to generate topic-consistent responses to these irrelevant questions, the following methods are often employed:
[0024] The method is based on model fine-tuning. This approach mainly involves constructing high-quality, labeled, domain-specific datasets to fine-tune the language model. Through fine-tuning, the internal parameters of the language model are effectively adjusted, enabling it to achieve topic consistency in task-oriented dialogue scenarios and adapt to specific constraints. However, to ensure the effectiveness of model fine-tuning, this method requires a large amount of training data, resulting in high costs for labeling training data and significant computational resource consumption during training, making it difficult to apply in practice.
[0025] In addition, there are methods based on language model-based prompting engineering. This approach requires building a robust language model for prompting word inputs to control the topic consistency of the generated response dialogue. However, these methods require long-term maintenance of the instruction content and dialogue context, and cannot guarantee the language model's adherence to instructions, thus compromising the effectiveness of generated response dialogues in real-world applications.
[0026] In view of this, embodiments of this disclosure provide a method for generating a response dialogue. First, in response to receiving a user query, the user query is input into a preset language model to obtain a first output vector generated by a first network layer in the preset language model based on the user query. Second, a guiding vector is obtained, which is generated based on a second output vector of preset interference prompt information corresponding to the first network layer. The interference prompt information is constructed based on a preset answer to the interference query. Finally, the first output vector is updated based on the guiding vector, so that subsequent network layers generate a response dialogue to the user query based on the updated first output vector. Therefore, embodiments of this disclosure do not require constructing a high-quality labeled dataset for fine-tuning the language model; they only need to generate the guiding vector based on the output vector of the first network layer corresponding to the interference prompt information in the preset language model. Since the interference prompt information is constructed based on a preset answer to the interference query, the guiding vector generated based on the output vector of the first network layer corresponding to the interference prompt information can be used to guide the preset language model to generate a topic-consistent response dialogue. Thus, when the preset language model processes a user query, directly updating the output vector of the first network layer in the preset language model using this guiding vector can effectively improve the topic consistency of the response dialogue generated by the preset language model.
[0027] Figure 1 This diagram illustrates an application scenario of the method and apparatus for generating response dialogues provided in this embodiment of the disclosure.
[0028] like Figure 1 As shown, the application scenario of this disclosure embodiment may include terminal device 101, network 103, and server 102. Network 103 is used as a medium to provide a communication link between terminal device 101 and server 102. Network 103 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.
[0029] Users can use terminal device 101 to interact with server 102 via network 103 to receive or send messages, etc. Various communication client applications can be installed on terminal device 101, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social platform software, etc. (for example only).
[0030] Terminal device 101 can be various electronic devices with a display screen and support web browsing, including but not limited to smartphones, tablets, laptops, and desktop computers.
[0031] Server 102 can be a server that provides various services, such as a backend management server that supports websites browsed by users using terminal device 101 (for example only). The backend management server can analyze and process data such as received user requests, and feed back the processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal device.
[0032] It should be noted that the method and apparatus for generating response dialogues provided in this embodiment of the present disclosure can be executed by server 102. Accordingly, the method and apparatus for generating response dialogues provided in this embodiment of the present disclosure can be located in server 102. Alternatively, the method and apparatus for generating response dialogues provided in this embodiment of the present disclosure can also be executed by a server or server cluster that is different from server 102 and capable of communicating with terminal device 101 and / or server 102. Accordingly, the method and apparatus for generating response dialogues provided in this embodiment of the present disclosure can also be located in a server or server cluster that is different from server 102 and capable of communicating with terminal device 101 and / or server 102.
[0033] It should be understood that Figure 1 The number of terminal devices, networks, and servers shown is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, and servers can be included.
[0034] Figure 2 A flowchart illustrating a method for generating a response dialogue as provided in an embodiment of this disclosure. (Refer to...) Figure 2 The method includes:
[0035] Step S210: In response to receiving a user query statement, input the user query statement into a preset language model to obtain the first output vector generated by the first network layer in the preset language model based on the user query statement.
[0036] Among them, user query statements refer to the questions entered by users in task-oriented dialogue scenarios.
[0037] It should be noted that user queries can be related to the task topic or unrelated to it; this embodiment of the disclosure does not impose any limitations on this. For example, if the task topic is ticket booking, the user query can be an inquiry related to ticket booking, such as inquiries about ticket prices or quantities. The user query can also be a distracting question unrelated to ticket booking.
[0038] The preset language model is a preset natural language processing model that can understand the input user query and infer and generate the corresponding response dialogue.
[0039] In one optional implementation, to reduce the deployment cost of the preset language model and the computational resource consumption during inference, the preset language model can be a lightweight language model, i.e., a language model with a parameter size smaller than a preset parameter threshold. The preset parameter threshold can be adaptively set according to the actual deployment needs of the language model, and this disclosure does not impose any limitations on it. For example, the preset language model can be a small language model (SLM) with a small parameter size and low computational resource requirements.
[0040] The first network layer refers to the network layer in the preset language model to which the guiding vector is to be added. The number and position of the first network layer can be adaptively set according to actual needs, and this embodiment does not limit this.
[0041] By inputting a user query into a preset language model and performing a forward computation process within the preset language model, a first output vector generated by the first network layer based on the user query can be obtained. In other words, the first output vector is the feature vector output by the first network layer in the preset language model after the user query is input, and can also be referred to as the hidden representation of the user query by the first network layer.
[0042] Step S220: Obtain the guiding vector, which is generated based on the second output vector of the preset interference prompt information corresponding to the first network layer; wherein, the interference prompt information is constructed based on the preset answer of the interference query statement.
[0043] Here, the guiding vector refers to the vector representation used to guide the preset language model in generating a topic-consistent response dialogue. The guiding vector can be generated based on the second output vector of the first network layer corresponding to the preset interference prompts.
[0044] The preset interference prompts are used to guide the preset language model to generate topic-consistent answers to interference queries. These prompts are constructed based on the preset answers to the interference queries.
[0045] In one alternative implementation, to improve the dialogue generation effect of the response dialogue generated by the preset language model in response to the user's query, the information type of the interference prompt can be determined according to the dialogue topic of the dialogue scenario in which the user's query is located.
[0046] Correspondingly, distractor queries are used to characterize queries that are unrelated to the topic of the dialogue context corresponding to the user's query. Preset answers to distractor queries include those related to the topic of the dialogue context corresponding to the user's query.
[0047] For example, if the user's query corresponds to the scenario dialogue topic of ticket booking, then the distracting prompts will be of the ticket booking type. Distracting queries are pre-set inquiries unrelated to the ticket booking topic. Pre-set answers to distracting queries include those related to the ticket booking topic.
[0048] For example, disruptive prompts for order inquiry types may include:
[0049] Interfering with the query statement: "The weather is nice today".
[0050] The default response to the query is: "The weather is indeed nice. Do you need help checking the current order progress or handling other order issues?".
[0051] It should be noted that, in order to improve the diversity of the interference prompts, there can be multiple interference query statements, and each interference query statement can correspond to a interference topic. Examples include casual conversation interference topics, emotion-related interference topics, and cross-business consultation interference topics; this embodiment does not impose any limitations on these. Correspondingly, corresponding interference prompts can be constructed based on the preset answers to each interference query statement; that is, there can also be multiple interference prompts.
[0052] In one alternative implementation, to improve the effectiveness of the interference prompts, interference queries can be combined with topic-related and non-topic-related answers to construct interference prompts. This allows the interference prompts to clearly define the expected and unexpected answers of the preset language model when faced with interference queries.
[0053] Accordingly, the preset answers to the interfering query statement include: a first preset answer and a second preset answer. The first preset answer is used to represent the topic-related answer to the interfering query statement, and the second preset answer is used to represent the non-topic-related answer to the interfering query statement.
[0054] The interference prompts include: a first prompt and a second prompt. The first prompt is constructed based on the interference query and a first preset answer, and the second prompt is constructed based on the interference query and a second preset answer.
[0055] For example, regarding the distracting prompts for ticket booking, the first preset answer is a preset answer related to the ticket booking topic set for the distracting query. The second preset answer is a preset answer unrelated to the ticket booking topic set for the distracting query.
[0056] Accordingly, by combining the distracting query statement with the first preset answer, the first prompt message can be constructed. By combining the distracting query statement with the second preset answer, the second prompt message can be constructed.
[0057] Specifically, the interference hint information dataset S = {q1, q2, ...}. Here, any qi represents an interference hint derived from a interference query statement d. The interference hint information can be represented as qi = {q1, q2, ...}. i p ,q i n}. Where, q i p Used to represent the first prompt information, q i n Used to represent the second prompt message.
[0058] For each distracting query, it is paired with two answer options (i.e., a first preset answer cp and a second preset answer cn). The first preset answer cp represents the expected answer to the distracting query, and the second preset answer cn represents the unexpected answer to the distracting query.
[0059] Accordingly, the first prompt information is constructed based on the interfering query statement d and the first preset answer cp. For example, the first prompt information is obtained by concatenating the interfering query statement with the first preset answer. The second prompt information is constructed based on the interfering query statement d and the second preset answer cn. For example, the second prompt information is obtained by concatenating the interfering query statement with the second preset answer.
[0060] By inputting the interference prompt information into a preset language model, a second output vector corresponding to the interference prompt information in the first network layer can be obtained. In other words, the second output vector is the feature vector output by the first network layer in the preset language model after the interference prompt information is input; it can also be called the hidden representation of the interference prompt information in the first network layer. Furthermore, a guiding vector can be generated based on the second output vector.
[0061] In one alternative implementation, if the preset answer only includes topic-related answers that interfere with the query statement, i.e., the first preset answer, the character vector corresponding to the characters of the first preset answer can be extracted from the second output vector, and a guiding vector can be generated based on the character vector.
[0062] In one alternative implementation, when the preset answer includes a first preset answer and a second preset answer, a first sub-output vector corresponding to the first prompt information and a second sub-output vector corresponding to the second prompt information can be extracted from the second output vector. The first sub-output vector and the second sub-output vector are then compared to generate a guidance vector.
[0063] Accordingly, the guiding vector is generated in the following way: the interference prompt information is input into the preset language model to obtain the second output vector generated by the first network layer based on the interference prompt information; in the second output vector, the first sub-output vector corresponding to the first prompt information and the second sub-output vector corresponding to the second prompt information are extracted; the guiding vector is generated based on the vector difference between the first sub-output vector and the second sub-output vector.
[0064] The first sub-output vector and the second sub-output vector can be output vectors corresponding to all characters in the first prompt information and the second prompt information, or they can be output vectors corresponding to some characters in the first prompt information and the second prompt information. This embodiment does not limit them.
[0065] In one alternative implementation, to improve the accuracy of the generated guidance vector, only the output vectors corresponding to the first preset answer and the second preset answer can be extracted as the first sub-output vector and the second sub-output vector.
[0066] For example, the first sub-output vector hp(l) is the vector in the second output vector corresponding to the first preset answer cp. The second sub-output vector hn(l) is the vector in the second output vector corresponding to the second preset answer cn. That is, the first sub-output vector is the activation of the first network layer corresponding to the first preset answer, and the second sub-output vector is the activation of the first network layer corresponding to the second preset answer.
[0067] Accordingly, a guiding vector can be generated based on the vector difference between the first and second sub-output vectors. That is, the guiding vector v. i s It is calculated using the following formula:
[0068] v i s =hp(l)-hn(l) Formula (1)
[0069] It should be noted that when there are multiple interference prompts, each interference prompt corresponds to a guiding vector. The guiding vectors corresponding to each interference prompt can be averaged and normalized to obtain the final guiding vector.
[0070] Step S230: Update the first output vector according to the guiding vector so that the network layers after the first network layer can generate a response dialogue for the user query statement based on the updated first output vector.
[0071] There are several ways to update the first output vector based on the guiding vector. For example, the guiding vector and the first output vector can be directly superimposed to obtain the updated first output vector. Another example is to set corresponding weight coefficients for the guiding vector and the first output vector, and then superimpose the guiding vector and the first output vector according to the weight coefficients to obtain the updated first output vector. This embodiment of the present disclosure does not limit the method.
[0072] In this step, after updating the first output vector of the first network layer according to the guiding vector, the network layers in the preset language model located after the first network layer can perform iterative calculations based on the updated first output vector until the forward propagation process is completed. At this point, the preset language model can generate a response dialogue for the user's query.
[0073] The method for generating response dialogues provided in this disclosure does not require constructing a high-quality labeled dataset for fine-tuning the language model. Instead, it only needs to generate a guiding vector based on the output vector of the first network layer in the preset language model corresponding to the interference prompt information. Since the interference prompt information is constructed based on a preset answer to the interference query, the guiding vector generated from the output vector of the first network layer corresponding to the interference prompt information can be used to guide the preset language model to generate a topic-consistent response dialogue. Therefore, when the preset language model processes a user query, directly updating the output vector of the first network layer in the preset language model using this guiding vector can effectively improve the topic consistency of the response dialogue generated by the preset language model.
[0074] In one alternative implementation, to improve the update effect of updating the first output vector based on the guiding vector, information entropy can be calculated based on the output vectors of some network layers in the preset language model, and the vector value of the guiding vector can be adaptively adjusted based on the information entropy to adaptively adjust the guiding strength of the guiding vector.
[0075] Accordingly, the first output vector is updated based on the guiding vector, including: determining the target network layer from multiple network layers of the preset language model; calculating the information entropy of the third output vector generated by the target network layer based on the user query statement; adjusting the vector value of the guiding vector based on the information entropy to obtain the adjusted guiding vector; and updating the first output vector based on the adjusted guiding vector.
[0076] The target network layer refers to the network layer from which the information entropy is to be calculated among multiple network layers. The selection method of the target network layer can be adaptively set according to actual needs, and this embodiment does not limit this. For example, a network layer at a specific position among multiple network layers can be selected as the target network layer.
[0077] In one optional implementation, the target network layer can be determined based on the network layer at the halfway point of the preset language model, i.e., the middle network layer. The information entropy of the middle network layer reflects the preset language model's semantic understanding and information processing capabilities for user queries. Therefore, the response effect of the preset language model to user queries can be determined based on the information entropy of the middle network layer. A higher information entropy of the middle network layer indicates that the preset language model has lower certainty in its semantic understanding of user queries, and the topic consistency of its generated response is poor. In this case, the guiding strength of the guiding vector can be adaptively increased.
[0078] In one alternative implementation, to improve the selection effect of the target network layer, the target network layer can be adaptively selected according to the preset generation strategy of the preset language model, so as to improve the selection accuracy of the target network layer.
[0079] Accordingly, the target network layer is determined from multiple network layers of the preset language model, including: determining the preset generation strategy corresponding to the response dialogue generated by the preset language model based on the query type and / or complexity of the user query statement; finding the preset network layer position that matches the preset generation strategy from the mapping relationship between the generation strategy and the network layer position; and determining the target network layer from multiple network layers of the preset language model based on the preset network layer position.
[0080] The query type of a user query refers to the query category to which it belongs. The query type can be determined based on attributes such as the query task, query process complexity, and query domain. For example, based on the query task, query types can be categorized as transaction processing queries, tool invocation queries, and process guidance queries. Based on the query process complexity, query types can be categorized as single-turn queries and multi-round interactive queries. Based on the query domain, query types can be categorized as financial queries, lifestyle service queries, and medical queries.
[0081] The statement complexity of a user query is used to characterize the complexity of the user query in terms of structure, semantics, and information density. Based on the query type and / or statement complexity of the user query, a preset generation strategy corresponding to the response dialogue generated by the preset language model can be determined.
[0082] The preset language model corresponds to several selectable generation strategies, such as greedy decoding, contrastive decoding, constraint decoding, and bundle search. Different query types and / or statement complexities have corresponding matching generation strategies. For example, when the statement complexity is at level one and / or the query type is a transaction processing query, the corresponding generation strategy could be greedy decoding. As another example, when the statement complexity is at level two and / or the query type is a process-oriented query, the corresponding generation strategy could be bundle search.
[0083] It should be noted that the above correspondence between query type and / or statement complexity and generation strategy is only an example. In actual applications, the correspondence between query type and / or statement complexity and generation strategy can be adaptively set according to the needs of the scenario. This disclosure does not limit this.
[0084] After determining the preset generation strategy for the preset language model, the preset network layer positions that match the preset generation strategy can be found from the mapping relationship between the generation strategy and network layer positions. Different generation strategies have corresponding network layer positions. For example, for the greedy decoding strategy, the corresponding network layer positions could be the 16th or 19th network layer. For the bundle search strategy, the corresponding network layer position could be the 14th network layer. For the constraint decoding strategy, the corresponding network layer position could be the 11th network layer.
[0085] It should be noted that the above mapping relationship between the generation strategy and the network layer position is only an example. In actual applications, the mapping relationship between the generation strategy and the network layer position can be adaptively set according to the needs of the scenario. This disclosure does not limit this.
[0086] After determining the preset network layer position, the network layer at that preset position can be selected from multiple network layers of the preset language model as the target network layer.
[0087] The third output vector generated by the target network layer based on the user query is the feature vector output by the target network layer in the preset language model after the user query is input. This third output vector can also be referred to as the hidden representation of the user query by the target network layer.
[0088] The information entropy of the third output vector can be calculated as follows: determine the feature vector corresponding to each word in the third output vector; calculate the probability distribution corresponding to each word based on the feature vector corresponding to each word; determine the information entropy corresponding to each word based on the probability distribution corresponding to each word; and obtain the information entropy of the third output vector based on the information entropy corresponding to each word.
[0089] The feature vectors corresponding to each word segment can be normalized to obtain the probability distribution of each segment. Based on the probability distribution of each word segment, the information entropy of each word segment can be determined using the following formula:
[0090]
[0091] Among them, H i Let P be the information entropy corresponding to the i-th segmented token. i,j Let be the probability distribution corresponding to the j-th feature vector of the i-th word token, and n be the number of feature vectors corresponding to the i-th word token.
[0092] The information entropy of the third output vector can be obtained based on the information entropy corresponding to each word segment, which can be achieved using the following formula:
[0093]
[0094] Among them, H t is the information entropy of the third output vector, and m is the number of tokens.
[0095] After calculating the information entropy of the third output vector, the vector value of the guiding vector can be adjusted based on the information entropy to obtain the adjusted guiding vector. In one optional implementation, the information entropy and the guiding vector can be multiplied or summed to adjust the vector value of the guiding vector.
[0096] In one alternative implementation, to improve the adjustment effect on the guiding vector, the adjustment coefficient can be determined based on the information entropy, and then the guiding vector can be adjusted according to the adjustment coefficient.
[0097] Accordingly, the vector value of the guiding vector is adjusted according to the information entropy to obtain the adjusted guiding vector, including: determining the adjustment coefficient of the guiding vector according to the entropy value of the information entropy, wherein the coefficient value of the adjustment coefficient is positively correlated with the entropy value of the information entropy; and adjusting the vector value of the guiding vector according to the adjustment coefficient to obtain the adjusted guiding vector.
[0098] The adjustment factor is used to adjust the magnitude of the guiding vector. The adjusted guiding vector is obtained by multiplying the adjustment factor by the guiding vector.
[0099] It should be noted that there are multiple ways to determine the adjustment coefficient of the guiding vector based on the entropy value of information entropy, as long as the condition that the adjustment coefficient is positively correlated with the entropy value of information entropy is met. A positive correlation between the adjustment coefficient and the entropy value of information entropy indicates that the larger the entropy value of information entropy, the larger the adjustment coefficient value.
[0100] In one optional implementation, the adjustment coefficient of the guiding vector can be obtained by calculating the ratio between the information entropy and a preset information entropy threshold. In another optional implementation, the adjustment coefficient of the guiding vector can also be determined based on the entropy value of the information entropy in the following ways:
[0101]
[0102] in, Cmax is the adjustment coefficient of the guiding vector, H(L) is the information entropy, and a and t are preset parameters, with a being greater than 0.
[0103] Correspondingly, the size of the guiding vector can be adjusted according to the adjustment coefficient. For example, by multiplying the adjustment coefficient with the guiding vector, the larger the value of the adjustment coefficient, the larger the size of the adjusted guiding vector, and the stronger the guiding strength of the adjusted guiding vector.
[0104] In this embodiment of the disclosure, the larger the entropy value of the information entropy, the stronger the uncertainty of the output information of the preset language model. At this time, the larger the coefficient value of the adjustment coefficient, the stronger the guidance strength of the guidance vector after adjustment according to the adjustment coefficient, thereby effectively improving the processing effect of the preset language model on interference queries that deviate from the topic.
[0105] It should be noted that the specific implementation of updating the first output vector according to the adjusted guiding vector can be referred to the above description, and will not be repeated here. For example, the adjusted guiding vector and the first output vector can be directly superimposed. Or, according to the preset weight coefficient, the adjusted guiding vector and the first output vector can be superimposed. This disclosure does not limit this.
[0106] It should also be noted that in this embodiment, the target network layer can be located before or after the first network layer. When the first network layer is located before the target network layer, after updating the first output vector of the first network layer according to the guiding vector, since the target network layer is located after the first network layer, the target network layer needs to perform the forward propagation operation again based on the updated first output vector to finally obtain the response dialogue for the user's query.
[0107] When the first network layer is located after the target network layer, after updating the first output vector of the first network layer according to the guiding vector, the preset language model can continue to execute the original forward propagation operation process. That is, the network layers after the first network layer perform forward propagation operation according to the updated first output vector to finally obtain the response dialogue of the user's query statement.
[0108] In one alternative implementation, after generating a response dialogue for a user query, the topic relevance of the response dialogue can be calculated. If the topic relevance is low, a second network layer can be selected in a preset language model. Based on the output vector of the prompt information corresponding to the second network layer, a new guiding vector can be generated. The output vector of the second network layer corresponding to the user query can then be updated based on this guiding vector, and the response dialogue for the user query can be regenerated to improve the generation effect of the response dialogue.
[0109] Accordingly, after generating the response dialogue for the user query, the method further includes: calculating the topic relevance of the response dialogue; if the topic relevance is less than a preset relevance threshold, selecting a second network layer from multiple network layers of a preset language model; generating a guiding vector corresponding to the second network layer based on the fourth output vector of the interference prompt information corresponding to the second network layer; inputting the user query into the preset language model; updating the fifth output vector generated by the second network layer based on the user query based on the guiding vector corresponding to the second network layer, so that subsequent network layers can regenerate the response dialogue for the user query based on the updated fifth output vector.
[0110] The topic relevance of the response dialogue can be obtained by analyzing the semantic relevance between the context dialogue and the response dialogue. The context dialogue refers to the historical dialogue content related to the user's query.
[0111] In one alternative implementation, the first dialogue keyword and the second dialogue keyword of the context dialogue can be extracted, and the word overlap between the first dialogue keyword and the second dialogue keyword can be calculated. Based on the word overlap, the topic relevance of the response dialogue can be determined.
[0112] In one alternative implementation, the semantic similarity between the context dialogue and the response dialogue can be calculated, thereby determining the topic relevance of the response dialogue based on the semantic similarity.
[0113] The preset relevance threshold is a pre-defined threshold used to evaluate the topic relevance of the response dialogue. The value of the preset relevance threshold can be adaptively set according to actual application needs, and this embodiment does not impose any limitations on it.
[0114] If the topic relevance of the response dialogue is less than a preset relevance threshold, it indicates low topic consistency in the generated response dialogue. In this case, the network layer with the added guiding vector (i.e., the second network layer) can be reselected, and a new guiding vector can be generated for this second network layer. Correspondingly, the fifth output vector generated by this second network layer based on the user query can be updated according to the updated fifth output vector, so that subsequent network layers can regenerate the response dialogue to the user query based on the updated fifth output vector.
[0115] The fourth output vector is the feature vector output by the second network layer of the preset language model after the interference prompt information is input. It can also be called the hidden representation of the interference prompt information corresponding to the second network layer. Furthermore, the guiding vector corresponding to the second network layer can be generated based on the fourth output vector.
[0116] The method for generating the guiding vector corresponding to the second network layer based on the fourth output vector can be referred to the method for generating the guiding vector based on the second output vector described above, and will not be repeated here.
[0117] It should be noted that, to improve the accuracy of selecting the second network layer, a network layer that has an sequential relationship with the first network layer can be selected as the second network layer, following the hierarchical order of the network layers in the preset language model. For example, the network layer preceding or following the first network layer can be selected as the second network layer.
[0118] The fifth output vector is the feature vector output by the second network layer in the preset language model after the user query is input into the preset language model. It can also be called the hidden representation of the user query corresponding to the second network layer.
[0119] After generating the guiding vector corresponding to the second network layer, the fifth output vector can be updated based on this guiding vector. The method for updating the fifth output vector based on the guiding vector of the second network layer is similar to the implementation of updating the first output vector based on the guiding vector described above, and will not be repeated here. By updating the fifth output vector generated by the second network layer based on the user query statement according to the guiding vector of the second network layer, the response dialogue to the user query statement can be regenerated.
[0120] To facilitate understanding, the following specific example illustrates the detailed implementation of the above method:
[0121] In task-oriented dialogue scenarios, users may input irrelevant questions that deviate from the main topic. To improve the dialogue generation performance of language models in order to enable them to generate topic-consistent responses to these irrelevant questions, the following methods are often employed:
[0122] (1) A method based on large language model prompting engineering. Large language models rely on massive amounts of prior knowledge and have a good ability to follow prompts. This approach controls the topic consistency of the dialogue system by constructing prompt instructions. However, large language models have a huge number of parameters, with billions or even hundreds of billions of settings, which consumes a lot of computing resources during inference and has a high deployment cost.
[0123] (2) A method based on fine-tuning a small language model. This method fine-tunes a small language model by constructing a high-quality, labeled dataset for a specific domain. This effectively adjusts the internal parameters of the small language model, enabling it to achieve topic consistency in the context and adapt to specific constraints.
[0124] (3) A method based on small language model prompting engineering. This method guides the model to maintain the consistency of scene theme through prompt word instruction constraints.
[0125] However, the above methods have the following drawbacks: Methods based on large language model prompting engineering have high deployment costs, as large language models have billions or even hundreds of billions of parameters, consuming significant computational resources during inference. Furthermore, this approach requires long-term maintenance of detailed instructions and context, and its effectiveness often decreases in complex and nuanced scenarios. Methods based on small language model fine-tuning require large amounts of data, have high annotation costs, and incur significant training computational resource costs, making them unsuitable for practical applications. Furthermore, methods based on small language model prompting engineering suffer from poor instruction compliance capabilities, failing to guarantee the quality of generated response dialogues.
[0126] In view of this, the present disclosure provides a method for generating response dialogues that does not require additional fine-tuning of the language model, has lower data annotation and model training costs, and can also reduce model inference operation costs. Figure 3 A flowchart illustrating a method for generating a response dialogue according to an embodiment of this disclosure is shown below. Figure 3 The method includes:
[0127] Step S301: In response to receiving a user query statement, input the user query statement into a preset language model to obtain the first output vector generated by the first network layer in the preset language model based on the user query statement.
[0128] Specifically, during the response generation process, input information can be composed of system instructions (I), dialogue history (D), and user query statements, and then this input information is fed into a preset language model.
[0129] Step S302: Obtain the guiding vector, which is generated based on the second output vector of the preset interference prompt information corresponding to the first network layer; wherein, the interference prompt information is constructed based on the preset answer of the interference query statement.
[0130] Specifically, paired prompts can be constructed based on the interfering query statement. That is, the interfering prompts corresponding to a single interfering query statement can include a first prompt and a second prompt. The first prompt represents the desired behavior, and the second prompt represents the undesired behavior. Correspondingly, interfering queries can be combined with preset responses to form prompt pairs. The first prompt is obtained by combining the interfering query statement with the first preset response. The first prompt, i.e., the positive prompt, can end with the characters of the preset response for the target behavior (i.e., the first preset response). The second prompt is obtained by combining the interfering query statement with the second preset response. The second prompt, i.e., the negative prompt, can end with the characters of the preset response for the undesired behavior (i.e., the second preset response).
[0131] It should be noted that the interference prompts in this step can be constructed based on the interference query statement, and generated by the language model by inputting the prompts into a language model such as Qwen2.5-72B.
[0132] Furthermore, by calculating the activation difference between the first and second prompts at the position of the response character and taking the average, the guiding vector can be obtained.
[0133] For any interfering prompt, after inputting it into a preset language model, the hidden representation h(l) of the interfering prompt in a specified layer l, i.e., the first network layer, can be determined through forward propagation f(·). For example, hp(l) and hn(l) are the hidden representations of the first network layer output corresponding to the first preset response character (cp) and the second preset response character (cn), respectively. The guiding vector is obtained by calculating the difference between the two vectors.
[0134] It should be noted that when the dataset contains multiple sets of interference prompts, the final guiding vector v can be obtained by averaging and normalizing the guiding vector corresponding to each interference prompt.
[0135] Step S303: Determine the target network layer from multiple network layers of the preset language model.
[0136] Specifically, the target network layers can be determined by using a preset generation strategy corresponding to the preset language model. For example, if the preset generation strategy of the preset language model is a greedy decoding strategy, and it generates k=2 tags, then the target network layers are the 16th and 19th network layers.
[0137] Step S304: Calculate the information entropy of the third output vector generated by the target network layer based on the user query statement. Based on the entropy value of the information entropy, determine the adjustment coefficient of the guiding vector. The coefficient value of the adjustment coefficient is positively correlated with the entropy value of the information entropy.
[0138] Within the predefined language model's internal state, under the same system instructions, the information entropy distribution of each network layer differs for user queries that interfere with the query and for those that are topic-related. Assume that at the target network layer l, for input information xd = {I, D, d} or xo = {I, D, o}, calculate the information entropy Ed(l) or Eo(l) of its output vector. Here, o and d represent topic-related and interfering user queries, respectively, I is the system instruction, and D is the dialogue history.
[0139] After calculating the information entropy, adjustment coefficients, such as scaling coefficients, can be calculated based on the information entropy. Then, the vector value of the guiding vector can be adjusted according to the scaling coefficient to obtain the adjusted guiding vector.
[0140] Step S305: Adjust the vector value of the guiding vector according to the adjustment coefficient to obtain the adjusted guiding vector.
[0141] Step S306: Update the first output vector according to the adjusted guiding vector, so that the network layers after the first network layer can generate a response dialogue for the user query statement based on the updated first output vector.
[0142] Specifically, the calculated scaling factor can be applied to the guiding vector v, and the adjusted guiding vector can be added to the activation of the preset language model in the specified layer, i.e., the first network layer. This process allows the guiding strength to dynamically adapt according to the information entropy of the target network layer in the preset language model, thereby improving the preset language model's ability to handle interfering queries and generate topic-related response dialogues.
[0143] In this embodiment, by calculating the guiding vector and adding a scaling factor, the need for large-scale training data is reduced by incorporating internal constraints from the pre-defined language model, enabling the language model to perform well even in resource-limited or data-scarce scenarios. Furthermore, it maintains high accuracy in topic-related input. When processing user queries that deviate from the topic, it maintains accurate responses, ensuring an unaffected user experience.
[0144] It is understood that the various method embodiments mentioned above in this disclosure can be combined with each other to form combined embodiments without violating the principle and logic. Due to space limitations, this disclosure will not elaborate further. Those skilled in the art will understand that in the above methods of specific implementation, the specific execution order of each step should be determined by its function and possible internal logic.
[0145] In addition, this disclosure also provides a device for generating a response dialogue, an electronic device, and a computer-readable storage medium, all of which can be used to implement any of the response dialogue generation methods provided in this disclosure. The corresponding technical solutions and descriptions are described in the corresponding section of the method and will not be repeated here.
[0146] Figure 4 This is a block diagram of a response dialogue generation apparatus provided in an embodiment of the present disclosure.
[0147] Reference Figure 4 This disclosure provides an apparatus for generating a response dialogue, the apparatus comprising:
[0148] Input module 41 is used to respond to receiving a user query statement by inputting the user query statement into a preset language model to obtain a first output vector generated by the first network layer in the preset language model based on the user query statement.
[0149] The acquisition module 42 is used to acquire a guidance vector, which is generated based on the second output vector of the interference prompt information corresponding to the first network layer; wherein the interference prompt information is constructed based on the preset answer of the interference query statement;
[0150] The generation module 43 is used to update the first output vector according to the guiding vector, so that the network layers after the first network layer generate a response dialogue for the user query statement according to the updated first output vector.
[0151] In one optional implementation, updating the first output vector based on the guiding vector includes: determining a target network layer from multiple network layers of the preset language model; calculating the information entropy of the third output vector generated by the target network layer based on the user query statement; adjusting the vector value of the guiding vector based on the information entropy to obtain an adjusted guiding vector; and updating the first output vector based on the adjusted guiding vector.
[0152] In one optional implementation, adjusting the vector value of the guiding vector based on the information entropy to obtain the adjusted guiding vector includes:
[0153] Based on the entropy value of the information entropy, the adjustment coefficient of the guiding vector is determined, and the coefficient value of the adjustment coefficient is positively correlated with the entropy value of the information entropy;
[0154] The vector value of the guiding vector is adjusted according to the adjustment coefficient to obtain the adjusted guiding vector.
[0155] In one optional implementation, determining the target network layer from multiple network layers of the preset language model includes:
[0156] Based on the query type and / or query complexity of the user query, determine the preset generation strategy corresponding to the preset language model generating the response dialogue;
[0157] From the mapping relationship between generation strategy and network layer position, find the preset network layer position that matches the preset generation strategy;
[0158] The target network layer is determined from multiple network layers of the preset language model based on the preset network layer position.
[0159] In one optional implementation, the preset answers to the interfering query statement include: a first preset answer and a second preset answer, wherein the first preset answer is used to characterize the topic-related answer of the interfering query statement, and the second preset answer is used to characterize the non-topic-related answer of the interfering query statement;
[0160] The interference prompt information includes: a first prompt information and a second prompt information; the first prompt information is constructed based on the interference query statement and the first preset answer, and the second prompt information is constructed based on the interference query statement and the second preset answer.
[0161] In one alternative implementation, the guiding vector is generated in the following manner:
[0162] The interference prompt information is input into the preset language model to obtain the second output vector generated by the first network layer based on the interference prompt information;
[0163] In the second output vector, the first sub-output vector corresponding to the first prompt information and the second sub-output vector corresponding to the second prompt information are extracted;
[0164] The guiding vector is generated based on the vector difference between the first sub-output vector and the second sub-output vector.
[0165] In one alternative implementation, after generating the response dialogue for the user query, the apparatus is further configured to:
[0166] Calculate the topic relevance of the response dialogue; if the topic relevance is less than a preset relevance threshold, select a second network layer from multiple network layers of the preset language model.
[0167] Based on the fourth output vector of the interference prompt information corresponding to the second network layer, a guiding vector corresponding to the second network layer is generated.
[0168] The user query is input into the preset language model. Based on the guiding vector corresponding to the second network layer, the fifth output vector generated by the second network layer based on the user query is updated, so that the network layers after the second network layer can regenerate the response dialogue to the user query based on the updated fifth output vector.
[0169] The modules in the aforementioned response dialogue generation device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the operations corresponding to each module.
[0170] Figure 5 This is a block diagram of an electronic device provided in an embodiment of the present disclosure.
[0171] Reference Figure 5 This disclosure provides an electronic device comprising: at least one processor 501; at least one memory 502; and one or more I / O interfaces 503 connected between the processor 501 and the memory 502; wherein the memory 502 stores one or more computer programs executable by the at least one processor 501, the one or more computer programs being executed by the at least one processor 501 to enable the at least one processor 501 to perform the above-described method for generating a response dialogue.
[0172] The modules in the aforementioned electronic devices can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in the processor of a computer device in hardware form or independent of it, or stored in the memory of a computer device in software form, so that the processor can call and execute the operations corresponding to each module.
[0173] This disclosure also provides a computer-readable storage medium storing a computer program thereon, wherein the computer program, when executed by a processor, implements the above-described method for generating a response dialogue. The computer-readable storage medium may be volatile or non-volatile.
[0174] This disclosure also provides a computer program product, including computer-readable code, or a non-volatile computer-readable storage medium carrying computer-readable code, wherein when the computer-readable code is run in a processor of an electronic device, the processor in the electronic device executes the above-described method for generating a response dialogue.
[0175] Those skilled in the art will understand that all or some of the steps, systems, and apparatuses disclosed above, and their functional modules / units, can be implemented as software, firmware, hardware, or suitable combinations thereof. In hardware implementations, the division between functional modules / units mentioned above does not necessarily correspond to the division of physical components; for example, a physical component may have multiple functions, or a function or step may be performed collaboratively by several physical components. Some or all physical components may be implemented as software executed by a processor, such as a central processing unit, digital signal processor, or microprocessor, or as hardware, or as an integrated circuit, such as an application-specific integrated circuit (ASIC). Such software can be distributed on a computer-readable storage medium, which may include computer storage media (or non-transitory media) and communication media (or transient media).
[0176] As is known to those skilled in the art, the term computer storage medium includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information (such as computer-readable program instructions, data structures, program modules, or other data). Computer storage media includes, but is not limited to, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), static random access memory (SRAM), flash memory or other memory technologies, portable compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical disc storage, magnetic cartridges, magnetic tape, disk storage or other magnetic storage devices, or any other medium that can be used to store desired information and is accessible to a computer. Furthermore, it is known to those skilled in the art that communication media typically contain computer-readable program instructions, data structures, program modules, or other data in modulated data signals such as carrier waves or other transmission mechanisms, and may include any information delivery medium.
[0177] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to each computing / processing device, or downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network can include copper transmission cables, fiber optic transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. The network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions to be stored in the computer-readable storage medium in each computing / processing device.
[0178] Computer program instructions used to perform the operations of this disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, status setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Smalltalk, C++, etc., and conventional procedural programming languages such as the "C" language or similar programming languages. The computer-readable program instructions may execute entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or may be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, electronic circuitry, such as programmable logic circuitry, field-programmable gate arrays (FPGAs), or programmable logic arrays (PLAs), is personalized by utilizing the status information of the computer-readable program instructions to implement various aspects of this disclosure.
[0179] The computer program product described herein can be implemented specifically through hardware, software, or a combination thereof. In one alternative embodiment, the computer program product is specifically embodied in a computer storage medium; in another alternative embodiment, the computer program product is specifically embodied in a software product, such as a software development kit (SDK), etc.
[0180] Various aspects of this disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0181] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a machine, so that when these instructions are executed by the processor of the computer or other programmable data processing device, a device is generated that implements the functions / actions specified in one or more blocks in the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium, where these instructions cause the computer, programmable data processing device, and / or other device to operate in a specific manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing various aspects of the functions / actions specified in one or more blocks in the flowchart and / or block diagram.
[0182] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device so that a series of operational steps are performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to implement the functions / actions specified in one or more blocks in the flowchart and / or block diagram.
[0183] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction containing one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions marked in the blocks may occur in a different order than those shown in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.
[0184] Example embodiments have been disclosed herein, and while specific terminology has been used, it is for illustrative purposes only and should be construed as such, and is not intended to be limiting. In some instances, it will be apparent to those skilled in the art that features, characteristics, and / or elements described in connection with particular embodiments may be used alone, or in combination with features, characteristics, and / or elements described in connection with other embodiments, unless otherwise expressly indicated. Therefore, those skilled in the art will understand that various changes in form and detail may be made without departing from the scope of this disclosure as set forth by the appended claims.
Claims
1. A method for generating a response dialogue, characterized in that, include: In response to receiving a user query statement, the user query statement is input into a preset language model to obtain a first output vector generated by the first network layer in the preset language model based on the user query statement; A guiding vector is obtained, which is generated based on the second output vector of the preset interference prompt information corresponding to the first network layer; wherein, the interference prompt information is constructed based on the preset answer of the interference query statement; Based on the guiding vector, the first output vector is updated so that subsequent network layers generate a response dialogue for the user query statement based on the updated first output vector.
2. The method according to claim 1, characterized in that, The step of updating the first output vector according to the guiding vector includes: The target network layer is determined from multiple network layers of the preset language model; Calculate the information entropy of the third output vector generated by the target network layer based on the user query statement, and adjust the vector value of the guiding vector according to the information entropy to obtain the adjusted guiding vector; The first output vector is updated based on the adjusted guiding vector.
3. The method according to claim 2, characterized in that, The step of adjusting the vector value of the guiding vector according to the information entropy to obtain the adjusted guiding vector includes: Based on the entropy value of the information entropy, the adjustment coefficient of the guiding vector is determined, and the coefficient value of the adjustment coefficient is positively correlated with the entropy value of the information entropy; The vector value of the guiding vector is adjusted according to the adjustment coefficient to obtain the adjusted guiding vector.
4. The method according to claim 2, characterized in that, Determining the target network layer from multiple network layers of the preset language model includes: Based on the query type and / or query complexity of the user query, determine the preset generation strategy corresponding to the preset language model generating the response dialogue; From the mapping relationship between generation strategy and network layer position, find the preset network layer position that matches the preset generation strategy; The target network layer is determined from multiple network layers of the preset language model based on the preset network layer position.
5. The method according to claim 1, characterized in that, The preset answers to the interfering query statement include: a first preset answer and a second preset answer, wherein the first preset answer is used to characterize the topic-related answer of the interfering query statement, and the second preset answer is used to characterize the non-topic-related answer of the interfering query statement; The interference prompt information includes: a first prompt information and a second prompt information; the first prompt information is constructed based on the interference query statement and the first preset answer, and the second prompt information is constructed based on the interference query statement and the second preset answer.
6. The method according to claim 5, characterized in that, The guiding vector is generated in the following manner: The interference prompt information is input into the preset language model to obtain the second output vector generated by the first network layer based on the interference prompt information; In the second output vector, the first sub-output vector corresponding to the first prompt information and the second sub-output vector corresponding to the second prompt information are extracted; The guiding vector is generated based on the vector difference between the first sub-output vector and the second sub-output vector.
7. The method according to claim 1, characterized in that, After generating the response dialogue for the user's query, the method further includes: Calculate the topic relevance of the response dialogue; if the topic relevance is less than a preset relevance threshold, select a second network layer from multiple network layers of the preset language model. Based on the fourth output vector of the interference prompt information corresponding to the second network layer, a guiding vector corresponding to the second network layer is generated. The user query is input into the preset language model. Based on the guiding vector corresponding to the second network layer, the fifth output vector generated by the second network layer based on the user query is updated, so that the network layers after the second network layer can regenerate the response dialogue to the user query based on the updated fifth output vector.
8. A device for generating a response dialogue, characterized in that, include: The input module is used to respond to receiving a user query statement by inputting the user query statement into a preset language model to obtain a first output vector generated by the first network layer in the preset language model based on the user query statement. The acquisition module is used to acquire a guidance vector, which is generated based on the second output vector of the interference prompt information corresponding to the first network layer; wherein the interference prompt information is constructed based on the preset answer of the interference query statement; The generation module is configured to update the first output vector according to the guiding vector, so that the network layers after the first network layer generate a response dialogue for the user query statement based on the updated first output vector.
9. An electronic device, characterized in that, include: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores one or more computer programs that can be executed by the at least one processor to enable the at least one processor to perform the method for generating a response dialogue as described in any one of claims 1-7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method for generating a response dialogue as described in any one of claims 1-7.