Dialogue model training method, dialogue generation method and related device
By using a dialogue model training method, the dialogue model is optimized based on the memorized content and preset scoring factors, generating more natural, coherent and human-like responses. This solves the problem of unnatural dialogue in virtual human characters and improves the user experience.
Patent Information
- Application Number
- CN202511861326.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-10
- Publication Date
- 2026-03-13
AI Technical Summary
The dialogue generated by virtual humans suffers from unnatural responses and a lack of human warmth, failing to meet users' emotional needs and resulting in a poor user experience.
By reading memory content associated with the current conversation, multiple response dialogues are generated. Based on preset scoring factors such as memory alignment and emotional resonance, the parameters of the dialogue model are updated to improve the memory alignment ability of the dialogue model, making the responses more natural, coherent and humane.
It enhances the naturalness and coherence of virtual human dialogue, satisfies users' emotional needs, and improves the user experience.
Smart Images

Figure CN121658610A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of computer technology, and more specifically, to a method for training a dialogue model, a dialogue generation method, and related apparatus. Background Technology
[0002] With increasingly rich cultural life and rapid development of computer technology, the demand for human-computer dialogue is also growing. Users can gain emotional value by conversing with virtual humans.
[0003] However, the dialogue generated by virtual humans may have issues such as unnatural responses and a lack of human touch, failing to meet users' emotional needs and resulting in a poor user experience when interacting with virtual humans. Summary of the Invention
[0004] A brief overview of this disclosure is given below to provide a basic understanding of some aspects of it. However, it should be understood that this overview is not an exhaustive summary of this disclosure. It is not intended to identify key or essential parts of this disclosure, nor is it intended to limit the scope of this disclosure. Its purpose is merely to present certain concepts of this disclosure in a simplified form as a prelude to the more detailed description that follows.
[0005] One of the purposes of this disclosure is to provide a method for training a dialogue model, a method for generating dialogue, and related apparatus.
[0006] According to a first aspect of this disclosure, a method for training a dialogue model is provided, comprising: reading memory content associated with a current dialogue; generating multiple response dialogues for the current dialogue based on the memory content, the current dialogue, and a pre-trained dialogue model; obtaining a target score for each of the multiple response dialogues based on preset scoring factors, wherein the preset scoring factors include a first factor related to the degree of memory alignment; and updating the parameters of the dialogue model based on the target score of each response dialogue.
[0007] In some embodiments, the first factor includes at least one of memory accuracy, emotional resonance, detail responsiveness, historical tracking, and topic coherence, wherein the target score is positively correlated with memory accuracy, emotional resonance, detail responsiveness, historical tracking, and topic coherence.
[0008] In some embodiments, the preset scoring factors further include a second factor related to dialogue quality, wherein the second factor includes at least one of the naturalness and coherence of the dialogue expression, and the target score is positively correlated with both the naturalness and coherence.
[0009] In some embodiments, obtaining a target score for each of the plurality of reply dialogues based on preset scoring factors includes: for each reply dialogue, obtaining a first score based on the first factor and a second score based on the second factor; and obtaining a target score based on the first score and the second score.
[0010] In some embodiments, reading memory content associated with the current conversation includes: generating a prompt message in response to acquiring the current conversation, wherein the prompt message indicates reading memory content associated with the current conversation from a memory bank; and reading the memory content from the memory bank based on the prompt message and a pre-trained memory retrieval model.
[0011] In some embodiments, reading memory content associated with the current conversation includes at least one of the following: reading memory content from a memory bank that matches the topic of the current conversation; reading memory content from a memory bank that meets the semantic requirements of the current conversation; and reading memory content from a memory bank that is associated with the word segmentation in the current conversation.
[0012] In some embodiments, the training method further includes: obtaining a training sample set before generating multiple response dialogues for the current dialogue based on the memory content, the current dialogue, and a pre-trained dialogue model, wherein the training samples include dialogue data and corresponding memory content labels; and training the dialogue model based on the training sample set using a supervised fine-tuning method.
[0013] According to a second aspect of this disclosure, a dialogue generation method based on memory content is provided, comprising: reading target memory content associated with an input dialogue from a user; and generating a target response dialogue for responding to the input dialogue based on the dialogue model trained according to the training method described above and the target memory content.
[0014] According to a third aspect of this disclosure, a training apparatus for a dialogue model is provided, comprising: a first reading module configured to read memory content associated with a current dialogue; a first dialogue generation module configured to generate multiple response dialogues for the current dialogue based on the memory content, the current dialogue, and a pre-trained dialogue model; a scoring module configured to obtain a target score for each of the multiple response dialogues based on preset scoring factors, wherein the preset scoring factors include a first factor related to the degree of memory alignment; and an updating module configured to update the parameters of the dialogue model based on the target score of each response dialogue.
[0015] According to a fourth aspect of this disclosure, a dialogue generation apparatus based on memory content is provided, comprising: a second reading module configured to read target memory content associated with an input dialogue from a user; and a second dialogue generation model configured to generate a target response dialogue for responding to the input dialogue based on the dialogue model trained by the training method described above and the target memory content.
[0016] According to a fifth aspect of this disclosure, an electronic device is provided, including a memory and a processor, wherein the memory stores instructions that, when executed by the processor, implement the operation of the training method or dialogue generation method as described above.
[0017] According to a sixth aspect of this disclosure, a non-transitory computer-readable storage medium is provided, wherein instructions are stored on the non-transitory computer-readable storage medium, and when the instructions are executed by a processor, the operation of the training method or dialogue generation method as described above is implemented.
[0018] According to a seventh aspect of this disclosure, a computer program product is provided, the computer program product including instructions that, when executed by a processor, implement the operation of the training method or dialogue generation method as described above.
[0019] Other features and advantages of this disclosure will become clearer from the following detailed description of exemplary embodiments with reference to the accompanying drawings. Attached Figure Description
[0020] The accompanying drawings, which form part of this specification, illustrate embodiments of this disclosure and, together with the specification, serve to explain the principles of this disclosure.
[0021] This disclosure will become clearer with reference to the accompanying drawings and the following detailed description, wherein:
[0022] Figure 1 A flowchart illustrating a method for training a dialogue model according to some embodiments of the present disclosure is shown.
[0023] Figure 2 A flowchart illustrating a process for obtaining a target score for each of multiple response dialogues based on preset scoring factors, according to some embodiments of the present disclosure, is shown.
[0024] Figure 3 A flowchart illustrating a dialogue generation method according to some embodiments of the present disclosure is shown;
[0025] Figure 4 A schematic diagram of a training apparatus for a dialogue model according to some embodiments of the present disclosure is shown;
[0026] Figure 5A schematic diagram of a dialogue generation apparatus according to some embodiments of the present disclosure is shown;
[0027] Figure 6 A schematic diagram of an electronic device according to some embodiments of the present disclosure is shown;
[0028] Figure 7 A schematic block diagram of a computer system on which embodiments of the present disclosure may be implemented is shown.
[0029] Note that in the embodiments described below, the same reference numerals are sometimes used across different figures to denote the same parts or parts having the same function, and repeated descriptions are omitted. In this specification, similar reference numerals and letters are used to denote similar items; therefore, once an item is defined in one figure, it does not need to be discussed further in subsequent figures.
[0030] For ease of understanding, the positions, dimensions, and extents of the structures shown in the accompanying drawings and other materials may not represent actual positions, dimensions, and extents. Therefore, the disclosed invention is not limited to the positions, dimensions, and extents disclosed in the accompanying drawings and other materials. Furthermore, the drawings are not necessarily drawn to scale, and some features may be enlarged to show details of specific components. Detailed Implementation
[0031] Various exemplary embodiments of the present disclosure will now be described in detail with reference to the accompanying drawings. It should be noted that, unless otherwise specifically stated, the relative arrangement, numerical expressions, and values of the components and steps set forth in these embodiments do not limit the scope of the present disclosure.
[0032] The following description of at least one exemplary embodiment is merely illustrative and is in no way intended to limit the scope of this disclosure or its application or use. That is, the structures and methods herein are shown in an exemplary manner to illustrate different embodiments of the structures and methods in this disclosure. However, those skilled in the art will understand that they merely illustrate exemplary ways that can be used to implement this disclosure, and not exhaustive ways. Furthermore, the drawings are not necessarily drawn to scale, and some features may be enlarged to show details of specific components.
[0033] Techniques, methods, and equipment known to those skilled in the art may not be discussed in detail, but where appropriate, such techniques, methods, and equipment should be considered part of the specification.
[0034] In all examples shown and discussed herein, any specific values should be interpreted as merely exemplary and not as limitations. Therefore, other examples of exemplary embodiments may have different values.
[0035] With the enrichment of cultural life and the rapid development of computer technology, many functional services based on human-computer dialogue have emerged. The core value of these services is to simulate real-person dialogue interaction. Users can gain emotional value by conversing with virtual humans. However, the dialogue generated by virtual humans may suffer from unnatural responses and a lack of human warmth, making the virtual human appear as a "rigid dialogue machine" that fails to meet users' emotional needs.
[0036] In related technologies, the historical dialogue slices between virtual humans and users can be vectorized and saved to a vector library. During the dialogue, relevant historical fragments can be retrieved from the vector library in real time based on the current dialogue content. Then, combined with the capabilities of large language models, such as using ultra-long text reasoning techniques like length extrapolation, responses can be generated.
[0037] However, the responses generated by these technologies may be based on memory content that is unsuitable for the current context. This can disrupt the coherence of the dialogue and make the relevant memory content in the response seem abrupt and deliberate. Furthermore, if more than one memory is recalled or retrieved based on the current dialogue context, inputting all of these memories into a large language model may lead to misalignment of memory targets. For example, the memory that should be aligned is the cause of the event, but the response focuses on the details of the event, resulting in a poor user experience.
[0038] Unlike traditional Q&A and customer service services, in some business scenarios, the content to be remembered needs to be integrated more naturally into the appropriate dialogue context in order to better meet the user's needs. That is, while ensuring the smooth flow of the response context, the content to be remembered needs to be integrated more naturally into the response. However, related technologies still have the problem that responses cannot be integrated naturally into the current dialogue context and that the content to be remembered cannot be integrated naturally into the response.
[0039] To address at least one of the aforementioned problems, this disclosure proposes a training method for a dialogue model, a dialogue generation method, and related apparatus. In a scenario where a pre-trained dialogue model generates multiple response dialogues based on memory content related to the current dialogue, a score is obtained for each response dialogue based on preset scoring factors. The parameters of the dialogue model are then updated based on the score of each response dialogue. This improves the memory alignment capability of the dialogue model, enabling subsequent responses generated by the model to be more natural, coherent, and better aligned with relevant memories. This results in more human-centered responses, satisfying users' emotional needs and enhancing the user experience.
[0040] In some embodiments of this disclosure, such as Figure 1 As shown, training methods for dialogue models can include:
[0041] Step S110: Read the memory content associated with the current conversation.
[0042] In this disclosure, the dialogue can be in electronic text form, or it can be in the form of voice, video, etc., without limitation. Furthermore, the current dialogue can be user-authorized and user-provided content. The current dialogue can be multi-turn dialogue content, so as to comprehensively and accurately recall relevant memories by combining the content of multiple turns of dialogue. It is understood that the memory recall referred to in this disclosure mainly refers to accurately retrieving and reading content related to the input information of the current dialogue from stored historical data.
[0043] In some embodiments, the current conversation can be a conversation that the user is currently inputting through the agent, meaning the current conversation can be acquired in real time and online. Alternatively, in some embodiments, the current conversation can be a conversation extracted from, for example, a log containing the user's conversation data, meaning the current conversation can be acquired non-real-time and offline, and this is not limited thereto.
[0044] In some embodiments, corresponding memory content can be extracted from historical dialogues and stored in a memory bank, so that memory content associated with the current dialogue can be retrieved from the memory bank. Here, memory content belonging to the corresponding preset memory category in historical dialogues can be extracted based on preset memory categories, and the extracted memory content can be stored in the memory bank. In this way, it is not necessary to save all dialogue content to the memory bank, reducing the storage pressure on the memory bank, freeing up storage resources, and thus enabling the storage of more content with a longer time span. In this disclosure, the memory bank can be located on a remote server, or it can be located on a local server; there is no restriction on the location of the memory bank.
[0045] In some embodiments, the preset memory categories may include, for example, first categories such as interactive information categories and social information categories. It should be understood that each first category may also include one or more second categories. For example, interactive information categories may include agreement categories, interaction categories, relationship categories, etc., and social information categories may include life status categories, without limitation. In this way, through the preset memory categories, a corresponding memory summary system can be formed to facilitate refined memory management, reduce memory burden, and enable the extraction and preservation of more important personalized information in historical dialogues, while filtering out trivial and less important information. This avoids the degradation of memory retrieval accuracy and efficiency caused by storing too much unimportant information, and can effectively improve the accuracy and efficiency of subsequent memory retrieval or memory reading.
[0046] In some embodiments, retrieving memory content associated with the current conversation may include retrieving memory content from a memory bank that matches the topic of the current conversation. Generally, the memory bank may contain preset memory categories that include topics corresponding to different memory contents, such as directly using the memory topic as the name of a second category or its subcategory. Thus, referring to the aforementioned memory summarization system, relevant memory content can be retrieved by determining whether the topic of the current conversation matches the topic corresponding to the memory content in the memory bank. For example, if the topic of the current conversation is determined to be "sleep," memory content related to "sleep schedule" and "insomnia experience" can be retrieved from the memory bank; furthermore, if the current conversation has multiple topics, memory content related to each topic can be comprehensively retrieved based on these multiple topics.
[0047] Alternatively or optionally, in some embodiments, retrieving memory content associated with the current dialogue may include: retrieving memory content from a memory bank that meets the semantic requirements of the current dialogue. Generally, semantic requirements refer to the combination of surface information and deep-seated needs contained in the user's input content. Surface information can be extracted using structured processing techniques such as semantic recognition to extract key entities and attribute information from the text; while deep-seated needs can be obtained using multimodal feature fusion and intent classification models to acquire the user's intent (such as query, request, feedback, etc.) and emotional characteristics (such as positive, negative, neutral, etc.). Based on this, memory content with indirect or extended semantic associations can be recalled by judging the semantic requirements implied in the current dialogue. For example, if the semantic requirement of the current dialogue is determined to be the intent of "wanting to relax," memory content with key entities such as "going on a beach vacation" and emotional characteristics such as "seeking encouragement or comfort" can be retrieved from the memory bank.
[0048] Alternatively or optionally, in some embodiments, reading memory content associated with the current dialogue may include: retrieving memory content associated with word segmentation from a memory bank based on word segmentation in the current dialogue. Generally, memory content is retrieved and segmented to obtain core vocabulary, and then a common sense knowledge base is introduced. Based on technologies such as knowledge graphs, semantic relationships, attribute features, and scene logic between memory content and vocabulary content are explored. Key information in core vocabulary is focused through an attention mechanism, and algorithms or models such as deep learning technology and large language models are used for association or reasoning. In a non-limiting embodiment, common sense can be used to make reasonable associations or inferences about word segmentation. For example, when a user inputs content related to "Northeasterners," based on human geography knowledge, it can be inferred that the current dialogue may be associated with key information such as "sauerkraut stew with vermicelli" and "indoor heating," thereby retrieving memory content associated with "Northeasterners" for these key information.
[0049] In some embodiments, memory content matching the topic of the current dialogue can be retrieved from the memory bank first. If the number of memory contents retrieved based on the topic is less than a preset threshold, at least one of the following can be performed: retrieving memory content that meets the semantic requirements of the current dialogue; or retrieving memory content associated with word segmentation from the memory bank based on word segmentation in the current dialogue. This not only effectively retrieves the required memory content associated with the current dialogue but also significantly improves memory retrieval efficiency.
[0050] In some embodiments, reading memory content associated with the current conversation can be automatic, for example, by relying on one or more machine learning models. In some embodiments, machine learning models such as large language models can be used to read memory content associated with the current conversation from a memory bank. Alternatively, in some embodiments, reading memory content associated with the current conversation from a memory bank can also be achieved directly by an intelligent agent. This intelligent agent can also be referred to as a robot, a digital human, or a virtual agent of a machine learning model. The intelligent agent can be implemented based on one or more machine learning models, such as a large language model.
[0051] In some embodiments, retrieving memory content associated with the current dialogue may include: generating a first prompt in response to acquiring the current dialogue; and retrieving memory content associated with the current dialogue from a memory bank based on the first prompt and a pre-trained memory retrieval model. The first prompt may instruct the retrieval of memory content associated with the current dialogue from the memory bank.
[0052] In this disclosure, the first prompt information can be a comprehensive set of instructions that can guide the memory retrieval model to retrieve memory content associated with the current dialogue from the memory bank, and the first prompt information can be represented in text form; the memory retrieval model, also known as the memory recall model, can be implemented, for example, based on a large language model.
[0053] In some embodiments, the first prompt information can be generated based on a memory retrieval strategy to be adopted, wherein the memory retrieval strategy can include at least one of the following: retrieving memory content from the memory bank that matches the topic of the current dialogue; retrieving memory content from the memory bank that meets the semantic requirements of the current dialogue; and retrieving memory content associated with word segmentation from the memory bank based on word segmentation in the current dialogue. Thus, the first prompt information can guide the memory retrieval model to retrieve relevant memory content based on the corresponding memory retrieval strategy.
[0054] After reading the memory content associated with the current dialogue, corresponding response dialogues can be generated based on the read memory content. In some embodiments, generating multiple response dialogues for the current dialogue based on the read memory content can be done automatically, for example, by relying on one or more machine learning models. In some embodiments, machine learning models such as large language models can be used to generate multiple response dialogues for the current dialogue based on the read memory content. Alternatively, in some embodiments, generating multiple response dialogues for the current dialogue based on the read memory content can also be implemented directly by an intelligent agent.
[0055] Specifically, in some embodiments of this disclosure, such as Figure 1 As shown, training methods for dialogue models can also include:
[0056] Step S120: Generate multiple response dialogues for the current dialogue based on the memorized content, the current dialogue, and the pre-trained dialogue model.
[0057] In this disclosure, the dialogue model can be implemented, for example, based on a large language model, and corresponding response dialogues can be automatically generated by invoking a pre-trained dialogue model. It should be understood that, in this disclosure, the pre-trained dialogue model can be implemented, for example, based on various large language models that are currently known or will be developed in the future, such as the Qwen large language model, etc.
[0058] Here, multiple response dialogues for the current dialogue can be candidate dialogues used to respond to the current dialogue. These response dialogues may or may not be related to the retrieved memory content. It is understood that when there are multiple retrieved memory contents associated with the current memory, each candidate response dialogue can correspond to or be related to one or more memory contents. In this way, a diverse range of candidate responses can be generated.
[0059] In some embodiments, generating multiple response dialogues for the current dialogue based on the memory content, the current dialogue, and a pre-trained dialogue model may include: inputting the memory content and the current dialogue into the pre-trained dialogue model to generate multiple response dialogues for the current dialogue.
[0060] In other embodiments, generating multiple response dialogues for the current dialogue based on the memory content, the current dialogue, and a pre-trained dialogue model may include: generating a second prompt message in response to reading memory content associated with the current dialogue; and generating response dialogues based on the second prompt message and the pre-trained dialogue model. The second prompt message may indicate the generation of multiple response dialogues for the current dialogue based on the memory content; in some embodiments, the second prompt message may be generated based on both the memory content and the current dialogue.
[0061] In this disclosure, the second prompt information can be a comprehensive set of instructions that can guide the dialogue model to generate multiple response dialogues for the current dialogue based on the memory content, and the second prompt information can be represented in text form.
[0062] In order to enable the generated response dialogue to effectively link with the memory content, that is, to ensure that the response dialogue can be accurately and naturally aligned with the retrieved memory content, in some embodiments, supervised fine-tuning (SFT) can be performed on the pre-trained dialogue model to give the dialogue model better memory alignment capabilities.
[0063] Specifically, before generating multiple response dialogues for the current dialogue based on the remembered content, the current dialogue, and a pre-trained dialogue model, a training sample set can be obtained. A supervised fine-tuning method can then be used to train the dialogue model based on this training sample set. The training samples include historical dialogue data and corresponding remembered content labels. In this way, the model parameters of the pre-trained dialogue model can be updated using the training sample set, enabling the dialogue model to better integrate the relevant remembered content to generate appropriate response dialogues, thus achieving better memory alignment capabilities.
[0064] In this disclosure, the dialogue data in the training samples may come from one or more users and have been authorized for use by those users. The memory content tags of the dialogue data may be manually labeled or automatically generated based on one or more machine learning models. For example, memory content tags may be generated by calling a corresponding memory retrieval model to read memory content associated with relevant dialogue data from a memory bank. This is not limited to this method. Considering that the number of training samples may be large, in a non-limiting embodiment, the Megatron training framework may be used for training to improve training efficiency.
[0065] To further enhance the memory alignment capability of the dialogue model, thereby enabling the responses generated by the dialogue model to be more natural and coherent, in some embodiments of this disclosure, such as Figure 1 As shown, training methods for dialogue models can also include:
[0066] Step S130: Obtain the target score for each reply dialogue in multiple reply dialogues based on preset scoring factors; and
[0067] Step S140: Update the parameters of the dialogue model based on the target score of each response dialogue.
[0068] The preset scoring factors can include a first factor related to the degree of memory alignment. Thus, by scoring the response dialogue using preset scoring factors and updating the parameters of the dialogue model, the goal is to improve the score of the response dialogue, thereby enabling the dialogue model to generate response dialogues with a higher degree of memory alignment and enhancing the dialogue model's memory alignment capability.
[0069] In some embodiments, the first factor related to memory alignment may include at least one of memory accuracy, emotional resonance, detail responsiveness, historical tracking, and topic coherence. Thus, corresponding reward strategies can be set based on emotional resonance, detail responsiveness, historical tracking, and topic coherence to obtain a target score. This allows the dialogue model updated based on the target score to generate responses that are responsive to memory, logically clear, and humane, thereby improving the user experience.
[0070] Specifically, in some embodiments, the target score can be positively correlated with the accuracy of memory recall. For example, the higher the accuracy of the memory recall in responding to the dialogue, the higher the target score. Memory accuracy can indicate the accuracy of the recalled content involved in responding to the dialogue, including but not limited to criteria such as the recall content being consistent with the topic of the current dialogue, the recall content having no irrelevant memory redundancy, and the presented text details being complete and without omissions.
[0071] In other embodiments, the target score can be positively correlated with the degree of emotional resonance; for example, the higher the degree of emotional resonance in the response dialogue, the higher the target score. The degree of emotional resonance can indicate the extent to which the response incorporates the emotional needs within the preceding context. A higher degree of emotional resonance indicates that the response dialogue is more naturally able to incorporate historical memories to achieve emotional resonance, such as using personal experience memories to empathize. If the response dialogue merely adds or introduces retrieved memories without incorporating the emotional needs within the preceding context, then the degree of emotional resonance in the response dialogue is low.
[0072] In some embodiments, the target score can be positively correlated with the level of detail responsiveness; for example, the higher the level of detail responsiveness in the response dialogue, the higher the target score. Detail responsiveness indicates the specificity of the remembered content involved in the response dialogue. For instance, for events or items mentioned by the user in historical conversations, subsequent interactions can recall their various ancillary details, and the response dialogue should fully and comprehensively reflect an understanding and response to these ancillary details. A higher level of detail responsiveness indicates more specific remembered content in the response dialogue.
[0073] In some embodiments, the target score can be positively correlated with the degree of historical tracking; for example, the higher the degree of historical tracking in the response dialogue, the higher the target score. For memory content recalled in the current dialogue, compared to its presentation in historical dialogues, it often undergoes subsequent development or changes over time. The degree of historical tracking is mainly used to indicate the extent to which this change is reflected in the response dialogue. A higher degree of historical tracking indicates that the response dialogue better reflects the changes in the relevant content, and better reflects the effective understanding and utilization of the memory content.
[0074] In some embodiments, the target score can be positively correlated with the degree of topic cohesion; for example, the higher the degree of topic cohesion in the response dialogue, the higher the target score. Generally, the degree of topic cohesion can indicate the degree of relevance between the context of the current dialogue and the memory content in the response dialogue, reflecting the response dialogue's understanding of the context and sentiment features of the current dialogue. A higher degree of topic cohesion indicates that the context of the current dialogue is more relevant to the memory content in the response dialogue.
[0075] It should be understood that the primary factor related to memory alignment may also include other factors besides memory accuracy, emotional resonance, detail responsiveness, historical tracking, and topic coherence. Specifically, one or more other factors related to memory alignment can be set according to the user's purpose and expectations for dialogue interaction. The target score may be positively correlated with or negatively correlated with these other factors, without any restrictions.
[0076] In some embodiments, obtaining a target score for each of multiple response dialogues based on preset scoring factors may include: determining a first score for the response dialogue based on a first factor, and obtaining a target score for the response dialogue based on the first score, wherein the target score is positively correlated with the first score. In a non-limiting embodiment, the first score may be determined as the target score for each response dialogue. Alternatively, if the preset scoring factors include other factors besides the first factor, the target score may be obtained by combining the first score with other scores obtained based on those other factors, as will be described in detail below.
[0077] In a non-limiting embodiment, the first rating may be positively correlated with the accuracy of memory, the degree of emotional resonance, the degree of detail relevance, the degree of historical tracing, and the degree of topic connection.
[0078] In some embodiments, determining a first score for a response dialogue based on a first factor may include: determining the memory alignment of the response dialogue based on the first factor, and determining the first score for the response dialogue based on the memory alignment of the response dialogue.
[0079] In some embodiments, one or more machine learning models, such as a large language model, can be used to determine the memory alignment of the response dialogue based on a first factor. In a non-limiting embodiment, determining the memory alignment of the response dialogue based on a first factor may include: determining the memory alignment of the response dialogue based on at least one of the following: memory accuracy, emotional resonance, detail responsiveness, historical tracking, and topic coherence. In a non-limiting embodiment, one or more machine learning models (such as a large language model) can be used to determine the memory alignment of the response dialogue based on at least one of the following: memory accuracy, emotional resonance, detail responsiveness, historical tracking, and topic coherence.
[0080] The first rating, based on the degree of memory alignment in the response dialogue, can be determined into several different score levels, such as first, second, and third alignment ratings ranked from low to high (e.g., 0 points, 1 point, and 2 points respectively).
[0081] Specifically, in some embodiments, in response to a first memory alignment degree in the response dialogue, a first score for the response dialogue is determined as a first memory alignment score. The first memory alignment degree may indicate that there is a significant and prominent memory misalignment problem in the response dialogue, such as the presence of at least one of the following: inappropriate associative memory; inaccurate description of memory content; or forced elicitation of memory content in an irrelevant context.
[0082] In other embodiments, in response to a second memory alignment degree that is higher than a first memory alignment degree, a first score is determined as a second memory alignment score. The second memory alignment degree may, for example, indicate that the response dialogue lacks highlights or significant problems, such as the presence of at least one of the following: no forced elicitation of memory content but omission of some important information; or the memory association is relatively appropriate, but the response content is relatively generic, lacking detail and adaptability.
[0083] In some other embodiments, in response to a third memory alignment level that is higher than a second memory alignment level, a first score is determined as a third memory alignment score. The third memory alignment level may indicate, for example, that the response dialogue has significant highlights, such as the presence of at least one of the following: appropriate memory content is selected and expressed naturally and harmoniously; the introduction of memory content advances the dialogue and creates a positive dialogue atmosphere.
[0084] In some embodiments, the preset scoring factors may also include a second factor related to the quality of the dialogue. In this way, by considering the second factor to score the response dialogue, the overall effect of the response dialogue can be comprehensively evaluated, so that the dialogue model updated based on the target score can align with the remembered content while improving the quality of the dialogue, thereby achieving good memory recall and a more realistic dialogue effect, and improving the user experience.
[0085] In some embodiments, a second factor related to dialogue quality may include at least one of the naturalness and coherence of the dialogue expression. Thus, a reward strategy can be set based on the naturalness and coherence of the dialogue expression to obtain a corresponding target score. This allows the dialogue model, updated based on the target score, to generate responses with more responsive memory, more natural expression, and clearer, more coherent logic, thereby improving the user experience.
[0086] The naturalness of a dialogue indicates that the choice of words and phrases is close to real human expression, without a mechanical or overly formal feel. For example, a low level of naturalness may indicate the presence of "mechanical" language, "unnecessary formal language," or "lengthy replies"; a high level of naturalness indicates that the dialogue conforms to human interaction logic, emotional expression is natural, and the overall dialogue closely resembles the intuitive experience of real humans in conversations (such as text or voice dialogues).
[0087] The coherence of a dialogue indicates how semantically and logically the response is consistent with the preceding context (the interactive content in the current dialogue). For example, low coherence may indicate that the response is abrupt, deviates from the dialogue context, or is logically confused; high coherence indicates that the response is logically consistent with the preceding context, is not abrupt, does not deviate from the core content of the preceding context, and is free from logical confusion.
[0088] In some embodiments, the target score may be positively correlated with naturalness, for example, the higher the naturalness of the response, the higher the target score; additionally or alternatively, the target score may be positively correlated with coherence, for example, the higher the coherence of the response, the higher the target score. It should be understood that the second factor may also include other factors besides naturalness and coherence, and the target score may be positively correlated with or negatively correlated with these other factors, without limitation herein.
[0089] like Figure 2 As shown, in some embodiments of this disclosure, obtaining a target score for each response dialogue in multiple response dialogues based on preset scoring factors may include:
[0090] Step S131: For each reply dialogue, obtain a first score based on the first factor and a second score based on the second factor;
[0091] Step S132: Obtain the target score based on the first score and the second score.
[0092] Specifically, the target score is positively correlated with the first score and can also be positively correlated with the second score. Thus, by comprehensively considering the first factor related to memory alignment and the second factor related to dialogue quality to obtain the target score, a multi-dimensional scoring of the response dialogue can be achieved. This effectively prevents the output of the dialogue model from deteriorating in either dialogue quality or memory alignment, thereby improving the generalization ability and robustness of the dialogue model.
[0093] In a non-limiting embodiment, the second score may be positively correlated with naturalness and coherence.
[0094] In some embodiments, a naturalness score can be obtained based on the naturalness of the response dialogue, a coherence score can be obtained based on the coherence of the response dialogue, and a second score can be obtained based on at least one of the naturalness score and the coherence score. In a non-limiting embodiment, it can be based on A2=W21. A21+W22 A22 yields a second score A2, where A2 is the second score, A21 is the natural score based on the degree of naturalness, A22 is the coherent score based on the degree of coherence, W21 is the preset weight corresponding to the natural score, and W22 is the preset weight corresponding to the coherent score.
[0095] On the one hand, in some embodiments, the naturalness rating of a response dialogue can be determined as multiple different score levels based on its naturalness, such as including first, second, third, fourth and fifth naturalness ratings ranked from low to high (e.g., 0 points, 1 point, 2 points, 3 points and 4 points respectively).
[0096] Specifically, in some embodiments, in response to the naturalness of the reply dialogue being a first naturalness level, a naturalness score for the reply dialogue is determined as a first naturalness score, wherein the first naturalness level may indicate, for example, that the reply dialogue exhibits at least one of the following phenomena: the dialogue lacks a sense of realism; the dialogue is written; the dialogue is formulaic; the content is too much to conform to the logic of everyday conversation; or it is completely inconsistent with human conversation style.
[0097] In other embodiments, in response to a response dialogue having a naturalness level higher than a first naturalness level, its naturalness score is determined to be a second naturalness score. The second naturalness level may, for example, indicate that the response dialogue exhibits at least one of the following: parts of the dialogue lack a sense of realism; parts of the dialogue use colloquial language; or there are obvious traces of mechanical expression (such as "encyclopedic explanations," "redundant introductory statements before describing related topics," or "repetitive explanations").
[0098] In some other embodiments, in response to a response dialogue having a naturalness level higher than a second naturalness level, a naturalness score is determined as a third naturalness score. A third naturalness level may, for example, indicate that the response dialogue exhibits at least one of the following characteristics: generally natural expression; generally colloquial expression; or partial responses that are emotionally inappropriate for the context of the current dialogue.
[0099] In some other embodiments, in response to a response dialogue having a naturalness level of a fourth naturalness level, which is higher than the third naturalness level, its naturalness score is determined to be a fourth naturalness score. The fourth naturalness level may, for example, indicate that the response dialogue exhibits at least one of the following characteristics: the dialogue is close to that of a real person; the language, emotions, and interaction logic conform to actual everyday conversation; and there is no mechanical expression overall.
[0100] In some other embodiments, in response to a response dialogue having a naturalness level of a fifth naturalness level, which is higher than a fourth naturalness level, its naturalness score is determined to be a fifth naturalness score. A fifth naturalness level may, for example, indicate that the response dialogue exhibits at least one of the following characteristics: the dialogue is anthropomorphic; and the language style and content description are close to a human conversational style.
[0101] In a non-limiting embodiment, one or more machine learning models, such as large language models, can be used to determine the naturalness of the response dialogue in order to determine the corresponding naturalness score.
[0102] On the other hand, in some embodiments, the coherence score of a response dialogue can be determined as multiple different score levels based on its coherence, such as including first, second, third, fourth and fifth coherence scores ranked from low to high (e.g., 0 points, 1 point, 2 points, 3 points and 4 points respectively).
[0103] Specifically, in some embodiments, in response to the coherence of the reply dialogue being a first coherence level, a coherence score for the reply dialogue is determined as a first coherence score, wherein the first coherence level may indicate, for example, that the reply dialogue exhibits at least one of the following phenomena: the reply dialogue is completely incoherent with the preceding text; it is unrelated to the core content of the preceding text; or it is logically incoherent.
[0104] In other embodiments, in response to a response dialogue having a second coherence level, which is higher than the first coherence level, a coherence score is determined as a second coherence score. The second coherence level may, for example, indicate that the response dialogue exhibits at least one of the following characteristics: the response dialogue is related to but deviates from the preceding text; the response dialogue uses keywords from the preceding text but deviates from the core content.
[0105] In some other embodiments, in response to the coherence of the reply dialogue being a third coherence level, which is higher than the second coherence level, its coherence score is determined to be a third coherence score. The third coherence level may, for example, indicate that the reply dialogue exhibits at least one of the following characteristics: the reply dialogue is substantially coherent with the preceding text; the semantic direction of the reply dialogue conforms to the direction of the core content of the preceding text but omits key details; or the reply dialogue is not closely related to the preceding text (e.g., the reply dialogue is an echolative reply).
[0106] In some other embodiments, in response to the coherence of the reply dialogue being a fourth coherence level, which is higher than the third coherence level, the coherence score of the reply dialogue is determined to be a fourth coherence score. The fourth coherence level may, for example, indicate that the reply dialogue exhibits at least one of the following characteristics: the reply dialogue closely follows the core content of the preceding text; or it is logically sound.
[0107] In some other embodiments, in response to a response dialogue having a coherence level of 5, which is higher than 4 coherence level, its coherence score is determined to be a 5 coherence score. For example, 5 coherence level could indicate the deep coherence of the response dialogue with the preceding text, meaning that the response dialogue is not only closely related to the core content of the preceding text, but also echoes the subtext, emotions, or details of the preceding text.
[0108] In other words, assessing the coherence of a response dialogue can include determining the consistency between the content of the response dialogue and the remembered content, as well as the smoothness of logical connections. Subdivided assessment elements can be set for each of these aspects.
[0109] In a non-limiting embodiment, one or more machine learning models, such as large language models, can be used to determine the coherence of the response dialogue in order to determine a coherence score for the response dialogue.
[0110] In some embodiments, it can be based on A=W1 A1+W2 A2 yields the target score A, where A1 is the first score, A2 is the second score, W1 is the preset weight corresponding to the first score, and W2 is the preset weight corresponding to the second score. Thus, the corresponding target score can be obtained based on the first and second scores.
[0111] It should be understood that in some embodiments, the preset scoring factors may also include other factors besides the first and second factors mentioned above, and the corresponding scores are obtained based on other factors; and the target score can also be calculated in other ways according to the scores corresponding to each factor, such as normalizing the first score and the second score first and then fusing them by summation, or calculating the covariance matrix between each factor, etc., without limitation.
[0112] In some embodiments, obtaining the target score for each reply dialogue in multiple reply dialogues based on preset scoring factors can be done automatically. This includes obtaining the first score based on a first factor and / or the second score based on a second factor, which can be achieved, for example, using one or more machine learning models. In some embodiments, machine learning models such as large language models can be used to obtain the target score for each reply dialogue in multiple reply dialogues based on preset scoring factors. Alternatively, in some embodiments, obtaining the target score for each reply dialogue in multiple reply dialogues based on preset scoring factors can also be directly achieved by an intelligent agent.
[0113] In some embodiments, obtaining a target score for each of the multiple response dialogues based on preset scoring factors may include: generating a third prompt message in response to generating multiple response dialogues for the current dialogue; and obtaining a target score for each response dialogue based on the third prompt message and a pre-trained scoring model. The third prompt message may indicate the target score for each response dialogue based on preset scoring factors; in some embodiments, the third prompt message may be generated based on preset scoring factors.
[0114] In this disclosure, the third prompt information can be a comprehensive set of instructions that can guide the scoring model to obtain the target score for each response dialogue, and the third prompt information can be represented in text form; the scoring model can be implemented, for example, based on a large language model.
[0115] After obtaining the target score for each response dialogue, a pre-defined optimization algorithm can be used to update the parameters of the dialogue model, enabling the dialogue model to tend to generate responses with higher scores. Specifically, an objective function can be set according to the target score of each response dialogue, and a pre-defined gradient update method can be used to calculate the gradient of the objective function with respect to the parameters of the dialogue model, so as to update the parameters of the dialogue model based on the gradient of the dialogue model parameters. Here, the parameters of the dialogue model can refer to the weight parameters and / or bias parameters of the neural network layers of the dialogue model, wherein the neural network layers may include at least one of an embedding layer, a self-attention layer, a feedforward neural network layer, and an output layer. In a non-limiting embodiment, the dialogue model can be built based on a Transformer architecture, and the dialogue model may include a layered embedding layer, a multi-head self-attention layer, a feedforward neural network layer, and a normalization layer, and the corresponding optimization algorithm can be used to update the weight parameters and bias parameters of the embedding layer, the multi-head self-attention layer, the feedforward neural network layer, and the normalization layer. It should be understood that the dialogue model can also be built based on other architectures, and the dialogue model can also have corresponding neural network layers and corresponding neural network layer parameters, which are not limited herein.
[0116] In some embodiments, a reward signal can be set based on the target score of each response dialogue, and a reinforcement learning algorithm can be used to update the parameters of the dialogue model. In a non-limiting embodiment, the target score of each response dialogue can be used as the reward signal. In a non-limiting embodiment, the objective function can be set based on maximizing the expected reward. In some embodiments, a Group Relative Policy Optimization (GRPO) algorithm can be used to update or fine-tune the parameters of the dialogue model. Specifically, for each response dialogue, the relative score of that response dialogue relative to other response dialogues within the group (i.e., the multiple response dialogues generated) is calculated (e.g., it can be obtained based on the mean and variance of the response dialogues within the group). Then, the relative score of each response dialogue can be used as the reward signal, a corresponding objective function can be set, and a preset gradient update method can be used to update the parameters of the dialogue model.
[0117] In a non-limiting embodiment, for the current dialogue q, the current dialogue model is utilized. Generate a set of output sequences containing G response dialogues, denoted as { , , …, }, and can get each reply dialogue. The generation probability of (i taking values from 1 to G) ( |q). Next, each response dialogue can be obtained based on preset scoring factors. The target score (i.e., the reward value) .
[0118] Furthermore, for each response dialogue, the corresponding dominance value can be calculated using within-group statistical characteristics. Specifically, the mean of the target scores for the group of G response dialogues is calculated based on the following formula. and standard deviation : , Next, for each response dialogue, a standardized advantage value can be calculated based on the following formula. (i.e., relative rating): ,in, To prevent tiny constants with a denominator of zero.
[0119] Furthermore, to ensure training stability and prevent excessive policy shifts, an objective function can be constructed or set that includes a policy objective and a Kullback-Leibler divergence penalty term. Specifically, the objective function L( ) can be set to:
[0120] .
[0121] Where X = This can represent the probability ratio of responses generated based on the old and new strategies, where the new strategy refers to the dialogue model whose parameters are to be updated. The old strategy refers to a dialogue model with fixed parameters used to generate candidate responses. clip( ) is a cutoff function used to restrict the probability ratio to [1- ,1+ Within the range of ], This is a preset hyperparameter (e.g., 0.2). This truncation function limits the policy update magnitude, i.e., limits the update magnitude of the model parameters, preventing excessive shifts in the model policy.
[0122] Y= Dialogue models that can represent parameters to be updated With reference dialogue model The KL divergence between the parameters is used to constrain the dialogue model whose parameters are to be updated. The output distribution is optimized to avoid conflicts with the dialogue model whose parameters need updating. Output distribution and reference dialogue model The output distribution deviates too much to ensure the dialogue model The fluency of the output language. (Refer to the dialogue model.) For example, it could be a dialogue model that has been supervised and fine-tuned. A preset coefficient for controlling the intensity of the KL divergence penalty. This indicates the calculation of the expected value; This indicates that the current dialogue q is sampled from the dialogue set P(Q). For example, P(Q) is the current multi-turn dialogue, and q is a single turn of dialogue in the current multi-turn dialogue. This indicates that the response to the dialogue is through the dialogue model. Generated based on the current dialogue q.
[0123] Furthermore, based on the objective function L( constructed above) The gradient of the objective function with respect to the parameters of each layer in the dialogue model (such as the attention layer weights and the feedforward layer weights) can be calculated using the backpropagation algorithm. Next, a pre-defined optimizer can be used, such as the Adaptive Momentum Estimation (Adam) optimizer or the Stochastic Gradient Descent (SGD) optimizer, based on the calculated gradient. Update the parameters of the dialogue model To minimize the above objective function L( This allows the dialogue model to be optimized to output responses with higher target scores, resulting in responses with higher memory alignment and dialogue quality.
[0124] According to the scheme disclosed herein, the candidate response dialogues generated by the dialogue model are scored based on preset scoring factors, and the parameters of the dialogue model are updated based on the scores. This improves the memory alignment capability of the dialogue model, enabling the response dialogues subsequently generated by the dialogue model to appropriately incorporate related memory content. This allows the response to echo and link with past memories, and the expression to be more natural. As a result, the dialogue model can generate more humane dialogue content, thereby meeting the user's value needs and companionship needs, and improving the user experience.
[0125] According to another aspect of this disclosure, a dialogue generation method based on memory content is also provided. In some exemplary embodiments of this disclosure, such as... Figure 3 As shown, dialogue generation methods may include:
[0126] Step S210: Read the target memory content associated with the input dialogue from the user;
[0127] Step S220: Based on the dialogue model and target memory content, generate a target response dialogue for responding to the input dialogue.
[0128] Here, the dialogue model can be trained based on the training method of the dialogue model described in any of the above embodiments, so that the target response dialogue used to respond to the input dialogue can echo and link with past memories, and can be more natural and humane in expression, thereby improving the user experience.
[0129] The input dialogue can be a dialogue currently entered by the user through the agent. Furthermore, reading the target memory content associated with the input dialogue can be referred to the description above regarding reading memory content associated with the current dialogue, and will not be repeated here.
[0130] According to another aspect of this disclosure, a training apparatus for a dialogue model is also provided, which can be configured to perform the training method for the dialogue model as described above. Figure 4 As shown, in some exemplary embodiments of this disclosure, the training device 300 may include a first reading module 310, a first dialogue generation module 320, a scoring module 330, and an update module 340.
[0131] The first reading module 310 can be configured to read memory content associated with the current dialogue. The first dialogue generation module 320 can be configured to generate multiple response dialogues for the current dialogue based on the memory content, the current dialogue, and a pre-trained dialogue model. The scoring module 330 can be configured to obtain a target score for each of the multiple response dialogues based on preset scoring factors, wherein the preset scoring factors may include a first factor related to the degree of memory alignment. The updating module 340 can be configured to update the parameters of the dialogue model based on the target score of each response dialogue.
[0132] The specific operation of the various modules in the training device 300 can be found in the detailed explanation of the training method for dialogue models above, and will not be repeated here.
[0133] According to another aspect of this disclosure, a dialogue generation apparatus is also provided, which can be configured to perform the dialogue generation method as described above. Figure 5 As shown, in some exemplary embodiments of this disclosure, the dialogue generation apparatus 400 may include a second reading module 410 and a second dialogue generation module 420.
[0134] The second reading module 410 can be configured to read target memory content associated with the input dialogue from the user. The second dialogue generation module 420 can be configured to generate a target response dialogue for responding to the input dialogue based on a dialogue model trained according to the dialogue model training method described above and the target memory content associated with the input dialogue.
[0135] The specific operation of the various modules in the dialogue generation device 400 can be found in the detailed explanation of the dialogue generation method above, and will not be repeated here.
[0136] According to another aspect of this disclosure, an electronic device is also provided, which can be configured to perform the training method or dialogue generation method of the dialogue model as described above. Figure 6 As shown, in some exemplary embodiments of this disclosure, the electronic device 500 may include a memory 510 and a processor 520. Instructions may be stored on the memory 510, which, when executed by the processor 520, can implement the dialogue model training method or dialogue generation method as described above.
[0137] Specifically, processor 520 can perform various actions and processes according to instructions stored in memory 510. Processor 520 can be an integrated circuit chip with signal processing capabilities. The processor can be a general-purpose processor, digital signal processor (DSP), application-specific integrated circuit (ASIC), field-programmable gate array (FPGA), or other programmable logic device, discrete gate or transistor logic device, or discrete hardware component. It can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of this disclosure. The general-purpose processor can be a microprocessor or any conventional processor, and can be an x86 architecture or an ARM architecture, etc.
[0138] Memory 510 stores executable instructions that, when executed by processor 520, implement the training method or dialogue generation method of the dialogue model described above. Memory 510 may be volatile memory or non-volatile memory, or may include both. Non-volatile memory may be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory may be random access memory (RAM), which serves as an external cache. By way of example, but not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous linked dynamic random access memory (SLDRAM), and direct memory bus random access memory (DR RAM). It should be noted that the memory of the methods described herein is intended to include, but is not limited to, these and any other suitable types of memory.
[0139] This disclosure also proposes a non-transitory computer-readable storage medium storing instructions that, when executed by a processor, can implement the dialogue model training method or dialogue generation method as described above.
[0140] Similarly, the non-transitory computer-readable storage media in the embodiments of this disclosure are intended to include, but are not limited to, the above and any other suitable types of memory.
[0141] This disclosure also proposes a computer program product that may include instructions that, when executed by a processor, can implement the training method or dialogue generation method of the dialogue model as described above.
[0142] Instructions can be any set of instructions that will be executed directly by one or more processors, such as machine code, or any set of instructions that will be executed indirectly, such as scripts. The terms “instruction,” “application,” “procedure,” “step,” and “program” used herein are used interchangeably. Instructions can be stored in object code format for direct processing by one or more processors, or stored in any other computer language, including scripts or sets of independent source code modules that are interpreted on demand or compiled ahead of time. The function, methods, and routines of instructions are explained in more detail in other parts of this document.
[0143] Figure 7A schematic block diagram of a computer system 600 on which embodiments of the present disclosure may be implemented is shown. The computer system 600 includes a bus 610 or other communication mechanism for transmitting information, and a processing means 620 coupled to the bus 610 for processing information. The computer system 600 also includes a memory coupled to the bus 610 for storing instructions to be executed by the processing means 620; the memory may be random access memory (RAM) or other dynamic storage device. The memory (such as RAM 630) may also be used to store temporary variables or other intermediate information during the execution of instructions to be executed by the processing means 620. The computer system 600 may also include a read-only memory (ROM) 640 or other static storage device coupled to the bus 610 for storing static information and instructions for the processing means 620. A storage device 650, such as a magnetic disk or optical disk, is provided and coupled to the bus 610 for storing information and instructions. Computer system 600 may be coupled via bus 610 to output device 660 for providing output to a user, such as, but not limited to, a display (such as a cathode ray tube (CRT) or liquid crystal display (LCD)), speakers, etc. Input device 670, such as a keyboard, mouse, microphone, etc., is coupled to bus 610 for transmitting information and command selection to processing device 620. Computer system 600 may perform embodiments of this disclosure. Consistent with certain implementations of this disclosure, computer system 600 provides results by executing one or more sequences of one or more instructions contained in memory (such as RAM 630) in response to processing device 620. Such instructions may be read into memory (such as RAM 630) from another computer-readable medium, such as storage device 650. Execution of the sequence of instructions contained in memory (such as RAM 630) causes processing device 620 to perform the methods described herein. Alternatively, hard-wired circuitry may be used in place of or in combination with software instructions to implement the teachings. Therefore, implementations of this disclosure are not limited to any particular combination of hardware circuitry and software. In various embodiments, computer system 600 can be connected across a network to one or more other computer systems, such as computer system 600, to form a networked system via network interface 680. This network may include a private network or a public network such as the Internet. In a networked system, one or more computer systems can store data and supply data to other computer systems. As used herein, the term "computer-readable medium" refers to any medium that participates in providing instructions to processing device 620 for execution. Such media can take many forms, including but not limited to non-volatile media, volatile media, and transmission media. Non-volatile media include, for example, optical discs or magnetic disks such as storage device 650. Volatile media include dynamic memory such as memory (e.g., RAM 630).Transmission media include coaxial cable, copper wire, and optical fiber, including cabling containing bus 610. Common forms of computer-readable media or computer program products include, for example, floppy disks, flexible disks, hard disks, magnetic tape, or any other magnetic media, CD-ROMs, digital video discs (DVDs), Blu-ray discs, any other optical media, thumb drives, memory cards, RAM, PROMs and EPROMs, fast EPROMs, any other memory chips or cartridges, or any other tangible media from which a computer can read. Various forms of computer-readable media may be involved when carrying one or more sequences of one or more instructions to processing device 620 for execution. For example, instructions may initially be carried on a disk of a remote computer. The remote computer may load the instructions into its dynamic memory and transmit the instructions over a telephone line using a modem. A modem local to computer system 600 may receive data over the telephone line and convert the data into an infrared signal using an infrared transmitter. An infrared detector coupled to bus 610 may receive the data carried in the infrared signal and place the data on bus 610. Bus 610 carries data to memory (such as RAM 630), and processing device 620 retrieves instructions from memory (such as RAM 630) and executes the instructions. Optionally, instructions received from memory (such as RAM 630) may be stored on storage device 650 before or after execution by processing device 620.
[0144] According to various embodiments, instructions configured to be executed by processing device 620 to perform a method are stored on a computer-readable medium. The computer-readable medium may be a device for storing digital information. For example, the computer-readable medium includes a compact disc read-only memory (CD-ROM) as known in the art for storing software. The computer-readable medium is accessed by a processor adapted to execute the instructions configured to be executed.
[0145] It should be noted that the flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0146] In general, the various exemplary embodiments of this disclosure can be implemented in hardware or dedicated circuitry, software, firmware, logic, or any combination thereof. Some aspects can be implemented in hardware, while others can be implemented in firmware or software that can be executed by a controller, microprocessor, or other computing device. When aspects of embodiments of this disclosure are illustrated or described as block diagrams, flowcharts, or using some other graphical representation, it will be understood that the blocks, apparatuses, systems, techniques, or methods described herein can be implemented as non-limiting examples in hardware, software, firmware, dedicated circuitry or logic, general-purpose hardware or controllers or other computing devices, or some combination thereof.
[0147] The terms “left,” “right,” “front,” “back,” “top,” “bottom,” “upper,” “lower,” “high,” “lower,” etc., used in the specification and claims, if present, are for descriptive purposes and are not necessarily used to describe unchanging relative positions. It should be understood that such terms are interchangeable where appropriate, so that embodiments of this disclosure described herein can, for example, operate on orientations different from those shown or otherwise described herein.
[0148] As used herein, the term “exemplary” means “serving as an example, instance, or illustration” and not as a “model” to be precisely copied. Any implementation described herein by example is not necessarily to be construed as preferred or advantageous over other implementations. Moreover, this disclosure is not limited to any theory expressed or implied as given in the field of art, background art, summary of invention, or detailed description.
[0149] As used herein, the term "substantially" means any minor variation resulting from design or manufacturing defects, device or component tolerances, environmental influences, and / or other factors. The term "substantially" also allows for differences from the perfect or ideal situation due to parasitic effects, noise, and other practical considerations that may exist in the actual implementation.
[0150] Furthermore, terms such as “first,” “second,” etc., may be used in this document for reference purposes only and are not intended to be limiting. For example, unless the context clearly indicates otherwise, the words “first,” “second,” and other such numerical terms relating to structures or elements do not imply order or sequence.
[0151] It should also be understood that when the term “including / contains” is used herein, it indicates the presence of the indicated feature, whole, step, operation, unit and / or component, but does not preclude the presence or addition of one or more other features, wholes, steps, operations, units and / or components and / or combinations thereof.
[0152] In this disclosure, the term “provide” is used broadly to cover all ways of obtaining an object, and therefore “provide an object” includes, but is not limited to, “purchasing,” “preparing / manufacturing,” “arranging / setting up,” “installing / assembling,” and / or “ordering” an object.
[0153] As used herein, the term “and / or” includes any and all combinations of one or more of the listed items in association. The terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting of this disclosure. As used herein, the singular forms “a,” “an,” and “the” are also intended to include the plural forms unless the context clearly indicates otherwise.
[0154] Those skilled in the art will recognize that the boundaries between the above operations are merely illustrative. Multiple operations may be combined into a single operation, a single operation may be distributed among additional operations, and operations may be performed with at least partial overlap in time. Moreover, alternative embodiments may include multiple instances of a particular operation, and the order of operations may be changed in various other embodiments. However, other modifications, variations, and substitutions are equally possible. Aspects and elements of all the embodiments disclosed above may be combined in any way and / or in combination with aspects or elements of other embodiments to provide multiple additional embodiments. Therefore, this specification and the accompanying drawings should be considered illustrative rather than restrictive.
[0155] While specific embodiments of this disclosure have been described in detail by way of example, those skilled in the art should understand that the examples are for illustrative purposes only and not intended to limit the scope of this disclosure. The various embodiments disclosed herein can be combined in any way without departing from the spirit and scope of this disclosure. Those skilled in the art should also understand that various modifications can be made to the embodiments without departing from the scope and spirit of this disclosure. The scope of this disclosure is defined by the appended claims.
Claims
1. A method for training a dialogue model, characterized in that, The training method includes: Read memory content associated with the current conversation; Based on the memory content, the current dialogue, and the pre-trained dialogue model, multiple response dialogues are generated for the current dialogue; A target score is obtained for each of the multiple response dialogues based on preset scoring factors, wherein the preset scoring factors include a first factor related to memory alignment; and The parameters of the dialogue model are updated based on the target score for each response dialogue.
2. The training method according to claim 1, characterized in that, The first factor includes at least one of the following: accuracy of memory, level of emotional resonance, level of detail responsiveness, level of historical tracking, and level of topic coherence. Specifically, the target score is positively correlated with accuracy of memory, level of emotional resonance, level of detail responsiveness, level of historical tracking, and level of topic coherence; and / or The preset scoring factors also include a second factor related to dialogue quality, wherein the second factor includes at least one of the naturalness and coherence of the dialogue expression, and the target score is positively correlated with the naturalness and coherence.
3. The training method according to claim 2, characterized in that, The target score for each of the multiple response dialogues, based on preset scoring factors, includes: For each response dialogue, a first score is obtained based on the first factor, and a second score is obtained based on the second factor; The target score is obtained based on the first and second scores.
4. The training method according to claim 1, characterized in that, Retrieving memory content associated with the current dialogue includes: generating a prompt message in response to acquiring the current dialogue, and retrieving the memory content from the memory bank based on the prompt message and a pre-trained memory retrieval model, wherein the prompt message indicates retrieving memory content associated with the current dialogue from the memory bank; and / or Retrieving memory content associated with the current dialogue includes at least one of the following: retrieving memory content from the memory bank that matches the topic of the current dialogue; retrieving memory content from the memory bank that meets the semantic requirements of the current dialogue; and retrieving memory content from the memory bank that is associated with word segmentation in the current dialogue; and / or The training method further includes: before generating multiple response dialogues for the current dialogue based on the memory content, the current dialogue, and the pre-trained dialogue model, obtaining a training sample set, wherein the training samples include dialogue data and corresponding memory content labels; and training the dialogue model based on the training sample set using a supervised fine-tuning method.
5. A dialogue generation method based on memory content, characterized in that, The dialogue generation method includes: Read target memory content associated with input dialogue from the user; and Based on the dialogue model trained by the training method according to any one of claims 1 to 4 and the target memory content, a target response dialogue is generated for responding to the input dialogue.
6. A training device for a dialogue model, characterized in that, The training device includes: The first reading module is configured to read memory content associated with the current conversation; The first dialogue generation module is configured to generate multiple response dialogues for the current dialogue based on the memory content, the current dialogue, and a pre-trained dialogue model. The scoring module is configured to obtain a target score for each of the plurality of response dialogues based on preset scoring factors, wherein the preset scoring factors include a first factor related to memory alignment; and The update module is configured to update the parameters of the dialogue model based on the target score for each response dialogue.
7. A dialogue generation device based on memory content, characterized in that, The dialogue generation device includes: The second reading module is configured to read target memory content associated with input dialogue from the user; and The second dialogue generation model is configured to generate a target response dialogue for responding to the input dialogue based on the dialogue model trained by the training method according to any one of claims 1 to 4 and the target memory content.
8. An electronic device, characterized in that, The electronic device includes a memory and a processor, the memory storing instructions that, when executed by the processor, perform the training method according to any one of claims 1 to 4 or the dialogue generation method according to claim 5.
9. A non-transitory computer-readable storage medium, characterized in that, The non-transitory computer-readable storage medium stores instructions that, when executed by a processor, perform the training method according to any one of claims 1 to 4 or the dialogue generation method according to claim 5.
10. A computer program product, characterized in that, The computer program product includes instructions that, when executed, perform the training method according to any one of claims 1 to 4 or the dialogue generation method according to claim 5.
Citation Information
Patent Citations
Training method and device for dialogue generation model
CN108984679A
Method, device and equipment for generating dialogue information and readable storage medium
CN115952272A
Method and device for generating dialogue information based on large model
CN118607636A
Context sensing type intelligent dialogue system construction method and terminal implementation
CN120144717A
Common-situation dialogue generation method and system based on large language model evaluation
CN120822619A