Multi-party dialogue model training method, related equipment thereof and multi-party dialogue system

Through comparative learning training of the speech style perception model, the problem of multi-party dialogue models relying on external annotation information is solved, personalized communication and high-quality response are achieved, and user dialogue satisfaction is improved.

CN120278230APending Publication Date: 2025-07-08ZHEJIANG AEROSPACE RUNBO MEASUREMENT & CONTROL TECH CO LTD +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510758450.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-09
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

The existing multi-party dialogue model relies on external annotation information, which limits the expressive ability of the personalized communication style between participants and affects user dialogue satisfaction.

Method used

By obtaining multi-party dialogue samples, training the speech style perception model based on comparison learning, and iteratively training the multi-party dialogue model using dialogue context samples to improve the model's ability to express the participants' personalized communication style.

Benefits of technology

It improves the quality and coherence of the responses of the multi-party dialogue model, enhances user dialogue satisfaction, and can generate personalized replies based on the context.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120278230A_ABST
    Figure CN120278230A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-party dialogue model training method, a related device thereof and a multi-party dialogue system, and relates to the technical field of artificial intelligence, and the method comprises the steps: obtaining a multi-party dialogue sample, and a dialogue context sample of a first speaker in a first dialogue of multi-party dialogues; based on the speaking of each speaker in the multi-party dialogue sample, performing comparative learning on a preset encoder to obtain a speaking style perception model; namely, the model can learn the speaking styles of different users without depending on the annotation information through a comparative learning mode; the method comprises the following steps: performing iterative training on a speaking style perception model based on a dialogue context sample to obtain a multi-party dialogue model, and learning how to generate a reply according to a context on the basis of having a speaking style perception capability, so that the multi-party dialogue model can respond to speaks of different speakers in different speaking styles, and the user experience is improved. Personalized communication with the participants is realized, the reply quality and continuity are improved, and the dialogue satisfaction degree of the user is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and particularly to a training method for a multi-party dialogue model, related devices thereof, and a multi-party dialogue system. Background Art

[0002] Currently, it is usually dependent on Graph Neural Networks (GNNs) to capture dependencies in conversations and model interactions between different participants.

[0003] However, although graph neural networks perform well in dealing with such dynamic structures, building a dialogue model in this way requires relying on additional annotation information in the dataset (for example, relationship annotations of "who says to whom"); this dependence on external annotation information limits to a certain extent the model's ability to express the personalized communication styles between participants, affects the quality and coherence of the model's generated responses, and reduces user dialogue satisfaction. Summary of the Invention

[0004] The main objective of this application is to provide a training method for a multi-party dialogue model, related devices thereof, and a multi-party dialogue system, aiming to solve the technical problem of low user dialogue satisfaction.

[0005] To achieve the above objective, this application proposes a training method for a multi-party dialogue model, and the method includes: Obtain multi-party dialogue samples and dialogue context samples of the first speaker in the first conversation of the multi-party dialogue; Based on the speeches of each speaker in the multi-party dialogue samples, perform contrastive learning on a preset encoder to obtain a speaking style perception model; Based on the dialogue context samples, perform iterative training on the speaking style perception model to obtain a multi-party dialogue model, where the multi-party dialogue model is used to respond to the speeches of different speakers with different speaking styles.

[0006] In one embodiment, the step of performing iterative training on the speaking style perception model based on the dialogue context samples to obtain a multi-party dialogue model further includes: Use the dialogue context in the dialogue context samples as the first query sample, use the true response in the dialogue context as the first positive sample, and use the wrong response in the dialogue context as the first negative sample; Based on the dialogue context, the first query sample, the first positive sample, and the first negative sample, perform iterative training on the speaking style perception model to obtain a multi-party dialogue model.

[0007] In one embodiment, the step of iteratively training the speech style perception model based on the dialogue context, the first query sample, the first positive sample, and the first negative sample to obtain a multi-party dialogue model includes: Through the speech style perception model, perform contrastive learning on the first query sample, the first positive sample, and the first negative sample to respectively obtain corresponding feature representations; Based on the corresponding feature representations and the dialogue context, output a response statement for the current dialogue of the first speaker through the speech style perception model; Based on the corresponding feature representations and the response statement of the first speaker, calculate the model loss value through a preset loss function; When the model loss value is greater than the target loss value, update the model parameters of the speech style perception model, and return to the step of performing contrastive learning on the first query sample, the first positive sample, and the first negative sample through the speech style perception model to respectively obtain corresponding feature representations, until a preset condition is met to obtain the multi-party dialogue model.

[0008] In one embodiment, the preset loss function includes a first loss function and a second loss function. The step of calculating the model loss value through the preset loss function based on the dialogue feature representation and the response statement includes: Based on the corresponding feature representations, calculate a first loss value through the first loss function; Based on the dialogue context sample, the response statement, and the first positive sample, calculate a second loss value through the second loss function; Perform weighted summation on the first loss value and the first loss value to obtain the model loss value.

[0009] In one embodiment, the first negative sample includes at least one of the response statements of other speakers different from the first speaker in the first dialogue, the response statements of the first speaker in other dialogues except the first dialogue in the multi-party dialogue, and the response statements after interference processing of the first positive sample.

[0010] In one embodiment, the step of performing contrastive learning on a preset encoder based on the speeches of each speaker in the multi-party dialogue sample to obtain a speech style perception model includes: Use the first speech of the second speaker in the second dialogue in the multi-party dialogue sample as the second query sample, use the second speech of the second speaker in the second dialogue as the second positive sample, and use the speeches of other speakers except the second speaker as the second negative sample; Based on the second query sample, the second positive sample, and the second negative sample, contrastive learning is performed on a preset encoder to obtain a speaking style perception model.

[0011] In addition, to achieve the above object, the present application further provides a multi-party dialogue system, which includes: An interaction module, configured to send the speech input by the user to the logic processing module in response to the speech, and display a reply statement. A logic processing module, configured to process the speech input by the user based on a pre-deployed multi-party dialogue model to generate a reply statement and return it to the interaction module, where the multi-party dialogue model is obtained by iteratively training the speaking style perception model based on the dialogue context sample of the first speaker in the first dialogue of the multi-party dialogue of the multi-party dialogue sample, and the speaking style perception model is obtained by performing contrastive learning on a preset encoder based on the speeches of each speaker in the multi-party dialogue sample.

[0012] In addition, to achieve the above object, the present application further provides a training device for a multi-party dialogue model, where the device includes: a memory, a processor, and a computer program stored on the memory and executable on the processor, and the computer program is configured to implement the steps of the training method of the multi-party dialogue model as described above.

[0013] In addition, to achieve the above object, the present application further provides a storage medium, which is a computer-readable storage medium, and a computer program is stored on the storage medium, and when the computer program is executed by a processor, the steps of the training method of the multi-party dialogue model as described above are implemented.

[0014] In addition, to achieve the above object, the present application further provides a computer program product, which includes a computer program, and when the computer program is executed by a processor, the steps of the training method of the multi-party dialogue model as described above are implemented.

[0015] One or more technical solutions provided by the present application have at least the following technical effects: This application obtains multi-party conversation samples and conversation context samples of the first speaker in the first conversation of the multi-party conversation; based on the speeches of each speaker in the multi-party conversation samples, contrastive learning is performed on a preset encoder to obtain a speaking style perception model; that is, through contrastive learning, without relying on annotation information, the model can learn the speaking styles of different users; further, based on the conversation context samples, iterative training is performed on the speaking style perception model to obtain a multi-party conversation model, which learns how to generate responses according to the context on the basis of having the ability to perceive speaking styles, so that the multi-party conversation model can respond to the speeches of different speakers in different speaking styles, realizing personalized communication with participants, improving the quality and coherence of responses, and enhancing user satisfaction with conversations. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] The accompanying drawings herein are incorporated into and constitute a part of this specification, showing embodiments consistent with this application and, together with the specification, are used to explain the principles of this application.

[0017] In order to more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the following will briefly introduce the accompanying drawings required for use in the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, other drawings can also be obtained based on these drawings without creative efforts.

[0018] Figure 1 It is a schematic flowchart provided for the first embodiment of the training method of the multi-party conversation model of this application; Figure 2 It is a schematic flowchart provided for the second embodiment of the training method of the multi-party conversation model of this application; Figure 3 It is an architecture diagram of the multi-party conversation system according to the embodiment of this application; Figure 4 It is a schematic diagram of the first scenario provided for the multi-party conversation system according to the embodiment of this application; Figure 5 It is a schematic diagram of the second scenario provided for the multi-party conversation system according to the embodiment of this application; Figure 6 It is a schematic diagram of the device structure of the hardware operating environment involved in the training method of the multi-party conversation model in the embodiment of this application.

[0019] The implementation, functional features, and advantages of this application will be further described with reference to the embodiments and the accompanying drawings. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0020] It should be understood that the specific embodiments described herein are only used to explain the technical solutions of this application and are not used to limit this application.

[0021] To better understand the technical solution of this application, the following will be described in detail in conjunction with the accompanying drawings of the specification and specific implementation manners.

[0022] It should be noted that the execution subject of this embodiment can be a computing service device with data processing, network communication, and program running functions, such as a tablet computer, a personal computer, a mobile phone, etc., or an electronic device, a training device for a multi-party dialogue model, etc. that can implement the above functions. Hereinafter, taking the training device for a multi-party dialogue model as an example, this embodiment and the following embodiments will be described.

[0023] Based on this, the embodiment of this application provides a method for training a multi-party dialogue model. Refer to Figure 1 , Figure 1 which is a schematic flowchart of the first embodiment of the method for training the multi-party dialogue model of this application.

[0024] In this embodiment, the method for training the multi-party dialogue model includes steps S10 to S30: Step S10, obtaining a multi-party dialogue sample and a dialogue context sample of the first speaker in the first dialogue of the multi-party dialogue; It should be noted that the open-domain dialogue system aims to achieve natural and flexible communication with users and can usually cover various topics from daily chats to complex social interactions; the key application scenarios of such a system are customer service, virtual assistants, and social robots, etc.; in traditional two-party conversations, it is relatively straightforward to identify the content of the speaker, but when it comes to multi-party conversations, the situation becomes much more complex. However, the above systems cannot fully consider the background information and personalized needs of users and cannot provide more targeted answers.

[0025] Current solutions mainly rely on Graph Neural Networks (GNNs) to model the dependencies and interactions among participants in the dialogue, thereby improving the quality of dialogue generation. Specifically, by treating each speaker as a node in the graph and using edges to represent the interaction relationships between them, the complex semantic flow in the dialogue can be captured; further, in order to enhance the expressive power of the model and handle the diverse styles and information flow of speakers in multi-party conversations, Heterogeneous Graph Neural Networks (HGNN) have also been proposed. By introducing different types of nodes and edges, it can better adapt to and handle diverse dialogue situations; in addition, there are also some methods dedicated to optimizing the dialogue model by maximizing the receiver's inference expectation, enabling it to be personalized according to the characteristics of specific speakers; at the same time, models combining discourse parsing graphs and graph fusion techniques have also been proposed to more deeply understand and analyze the dialogue structure and context, and capture information flow and topic transitions.

[0026] However, although the above-mentioned graph neural network performs well in processing this dynamic structure, building a dialogue model based on this approach requires reliance on additional annotation information in the dataset (for example, relationship annotations of "who said to whom"); this reliance on external annotation information limits the model's ability to express the personalized communication style between participants to a certain extent, affects the quality and coherence of the responses generated by the model, and reduces user conversation satisfaction.

[0027] In order to solve the above problem, this embodiment proposes a training method for a multi-party dialogue model. First, a sample basis can be provided for model training by obtaining multi-party dialogue samples and a dialogue context sample of the first speaker in the first dialogue of the multi-party dialogue.

[0028] Among them, the multi-party dialogue sample contains multiple historical dialogues, each historical dialogue contains the speeches of multiple speakers, and the speeches of each speaker are arranged in order to form the dialogue context (for example, the dialogue context of the first speaker includes the current speech of the first speaker, the speeches of other speakers before the first speaker, and the speeches of other speakers after the first speaker).

[0029] Step S20, based on the speech of each speaker in the multi-party dialogue sample, performing comparative learning on the preset encoder to obtain a speech style perception model; In order to avoid dependence on external annotation information during model training, this embodiment uses the speeches of each speaker in the multi-party dialogue sample as training data, adopts a comparative learning training method, and explores the speaking styles between different speakers, thereby training a speaking style perception model; while not relying on external annotation information, the model's ability to express the personalized communication style between participants is improved.

[0030] Specifically, based on the speeches of each speaker in the multi-party conversation sample, comparative learning is performed on the preset encoder. Since the speeches of different speakers are semantically different, and the speeches of the same speaker are similar in feature representation, the comparative learning algorithm can enable the model to understand that different speakers have different speaking styles, thereby perceiving the speaker's language style.

[0031] The specific implementation of performing comparative learning on a preset encoder based on the speech of each speaker in the multi-party dialogue sample to obtain a speech style perception model may be: Use the first utterance of the second speaker in the second conversation of the multi-party conversation sample as the second query sample, use the second utterance of the second speaker in the second conversation as the second positive sample, and use the utterances of speakers other than the second speaker as the second negative samples; based on the second query sample, the second positive sample, and the second negative samples, perform contrastive learning on a preset encoder to obtain a speech style perception model.

[0032] By comparing and learning the utterances of the same and different speakers in the same conversation and the entire conversation sample, the model can better distinguish the language styles and expressions of different speakers, thereby enhancing the model's ability to distinguish the personalized language features of speakers and enabling the model to better perceive the speaker roles in multi-party conversations.

[0033] Therefore, in this embodiment, a random utterance (the first utterance) is selected from any conversation (the second conversation) in the multi-party conversation sample as the query sample (the second query sample), and the corresponding speaker is (the second speaker); another utterance (the second utterance) made by the same speaker (the second speaker) in the same conversation is used as its positive sample (the second positive sample); the utterances of speakers other than this speaker (the second speaker) are used as negative samples .

[0034] Based on the second query sample, the second positive sample, and the second negative samples, perform contrastive learning on a preset encoder. Through contrastive learning, the model can learn between the second positive sample (utterances of the same speaker) and the second negative sample (utterances of other speakers), so as to distinguish the language styles and expressions of different speakers; specifically, contrastive learning maximizes the similarity between positive sample pairs (i.e., different utterances of the same speaker are as similar as possible), while minimizing the similarity between positive samples and negative samples (i.e., utterances of different speakers are as dissimilar as possible), so that the model can capture the unique styles of different speakers; when the speech style perception model encounters a new utterance, it can infer which type of speaker the utterance may belong to based on the knowledge of different styles learned previously.

[0035] In order to maximize the distinction of the unique styles of different speakers, the sampling method of the second negative sample can be at least one of the following two methods: First, sample the utterances of other speakers in the same conversation (the second conversation); Second, sample the utterances of other speakers in the multi-party conversation sample.

[0036] During the process of sampling query samples, positive samples, and negative samples, responses such as "yes" and "good" that lack meaning and information content are filtered out to ensure the high quality of the training data; thereby enabling the preset encoder to better capture the subtle differences in language styles among different speakers in a multi-party conversation.

[0037] Among them, based on the second query sample, the second positive sample, and the second negative sample, a specific implementation manner of performing contrastive learning on the preset encoder to obtain a speaking style perception model can be: Through the preset encoder, encode and transform the second query sample, the second positive sample, and the second negative sample respectively to obtain corresponding embedding vectors; use cosine similarity or dot product to calculate the first similarity between the second query sample and the second positive sample vector, and the second similarity between the second query sample and the second negative sample respectively, and calculate a loss value based on the first similarity and the second similarity.

[0038] Based on the calculated loss value and the training objective (to make the similarity between the query sample and the positive sample as high as possible, while the similarity between the query sample and the negative sample as low as possible), determine whether it is necessary to update the model parameters of the preset encoder (if the training objective is not achieved, then it is necessary to update). If necessary, update the model parameters of the preset encoder, and return to the step of encoding and transforming the second query sample, the second positive sample, and the second negative sample respectively through the preset encoder to obtain corresponding embedding vectors, until the target training conditions are met to obtain the speaking style perception model.

[0039] Among them, the embedding vectors corresponding to the second query sample, the second positive sample, and the second negative sample are respectively: ; ; ; Among them, is the embedding vector corresponding to the second query sample, is the embedding vector corresponding to the second negative sample, is the hidden vector representation of the second positive sample.

[0040] The specific formula for using dot product to calculate the first similarity between the second query sample and the second positive sample vector, and the second similarity between the second query sample and the second negative sample respectively, and calculating a loss value based on the first similarity and the second similarity is:

[0041] Among them, is the temperature coefficient, and the parameter and respectively represent the quantity from the second positive sample and the training batch.

[0042] For example, the settings in the training process of obtaining the speech style perception model by performing contrastive learning on the preset encoder based on the second query sample, the second positive sample, and the second negative sample can be: based on the PyTorch framework, using the T5-base pre-trained model as the preset encoder, the optimizer uses AdamW, and the learning rate is set to .

[0043] Step S30: Based on the dialogue context sample, iteratively train the speech style perception model to obtain a multi-party dialogue model, where the multi-party dialogue model is used to respond to the speeches of different speakers in different speech styles.

[0044] The trained speech style perception model can perceive the speech styles of different speakers. On this basis, in order to obtain a model that can respond to the speeches of different speakers in different speech styles, in this embodiment, based on the dialogue context sample, the speech style perception model is iteratively trained to obtain a multi-party dialogue model.

[0045] Specifically, the method of iteratively training the speech style perception model through the dialogue context sample to obtain the multi-party dialogue model can be based on an unsupervised algorithm. Since the dialogue context reflects the reply methods between the interlocutors, by analyzing the above dialogue context sample, it is possible to learn how to reply to the dialogue; therefore, the trained multi-party dialogue model has both the ability of style perception and dialogue reply.

[0046] In addition, during the above model training process, the above multi-party dialogue samples and dialogue context samples can also be obtained from the Friends dataset (including conversations between multiple characters, for example, multiple characters such as Monica, Ross, etc.). By performing contrastive learning on the speeches of each speaker in the multi-party dialogue samples obtained based on the Friends dataset on the preset encoder, the obtained speech style perception model has the personality characteristics of these multiple characters and can simulate the unique expressions of these multiple characters to have a conversation. When the user sends a message, the backend model will infer the appropriate character to respond and reply in the language style of that character.

[0047] In this embodiment, by obtaining multi-party conversation samples and the conversation context samples of the first speaker in the first conversation of the multi-party conversation; based on the speeches of each speaker in the multi-party conversation samples, contrastive learning is performed on a preset encoder to obtain a speaking style perception model; that is, through contrastive learning, without relying on annotation information, the model can learn the speaking styles of different users; further, based on the conversation context samples, the speaking style perception model is iteratively trained to obtain a multi-party conversation model, which can learn how to generate responses according to the context on the basis of the ability to perceive speaking styles, so that the multi-party conversation model can respond to the speeches of different speakers in different speaking styles, realize personalized communication with participants, improve the quality and coherence of responses, and improve user conversation satisfaction.

[0048] Compared with traditional training methods, since contrastive learning focuses on learning the similarities and differences between data rather than directly learning the mapping relationship from input to output labels, the contrastive learning method helps to improve the generalization ability of the model, so that it can better adapt to unseen new data and transfer the learning results between different tasks. In addition, the multi-party conversation model can also have multiple role personalities and can simulate the styles of different roles to reply, with a wide range of application scenarios.

[0049] Based on the first embodiment of the present application, in the second embodiment of the present application, the same or similar content as that in the above-mentioned first embodiment can be referred to the above introduction and will not be repeated hereinafter. On this basis, please refer to Figure 2 , step S30 includes steps S01 to S02: Step S01, taking the conversation context in the conversation context sample as the first query sample, taking the true response in the conversation context as the first positive sample, and taking the wrong response in the conversation context as the first negative sample; Step S02, based on the conversation context, the first query sample, the first positive sample and the first negative sample, iteratively train the speaking style perception model to obtain a multi-party conversation model.

[0050] By contrastively learning the correct and wrong responses in the same context, the model can better distinguish the response methods in different contexts. Therefore, in this embodiment, the conversation context in the conversation context sample is used as the first query sample, the true response in the conversation context is used as the first positive sample, and the wrong response in the conversation context is used as the first negative sample.

[0051] Based on the dialogue context, the first query sample, the first positive sample, and the first negative sample, the iterative training of the speech style perception model to obtain a multi-party dialogue model can be achieved through the Transformer architecture or the bidirectional LSTM (Long Short-Term Memory) algorithm, which learns the context information and the differences between the positive and negative samples, thereby training a multi-party dialogue model that can reply to different speakers in different styles.

[0052] To improve the dialogue coherence and response quality of the multi-party dialogue model, in this embodiment, the speech style perception model is jointly trained by combining contrastive learning and response tasks, thereby obtaining a multi-party dialogue model with the ability to reply and generate high-quality response sentences according to the dialogue context. Specifically, the implementation manner of iteratively training the speech style perception model based on the dialogue context, the first query sample, the first positive sample, and the first negative sample to obtain a multi-party dialogue model can be: Through the speech style perception model, contrastive learning is performed on the first query sample, the first positive sample, and the first negative sample to obtain corresponding feature representations respectively; based on the corresponding feature representations and the dialogue context, the speech style perception model outputs a response sentence for the current dialogue of the first speaker; based on the corresponding feature representations and the response sentence of the first speaker, a model loss value is calculated through a preset loss function; when the model loss value is greater than the target loss value, the model parameters of the speech style perception model are updated, and the step of performing contrastive learning on the first query sample, the first positive sample, and the first negative sample through the speech style perception model to obtain corresponding feature representations respectively is returned until a preset condition is met, and the multi-party dialogue model is obtained.

[0053] It should be noted that by performing contrastive learning to obtain corresponding feature representations respectively, it is convenient for the model to better capture the subtle differences between different response sentences through these feature representations, thereby improving the model's understanding ability of discourse intentions. This training method can help the model more accurately identify which utterances are most relevant to the current dialogue context, thereby improving the dialogue quality.

[0054] In addition, to improve the learning accuracy of contrastive learning, the first negative sample may include at least one of the response sentences of other speakers different from the first speaker in the first dialogue, the response sentences of the first speaker in other dialogues except the first dialogue in the multi-party dialogue, and the response sentences after disturbing the first positive sample.

[0055] Among them, the response statements of other speakers different from the first speaker in the first conversation can be selected from the current conversation context D as other speakers different from the first speaker different from the first speaker response , and use it as the first negative sample; this first negative sample demonstrates the response styles of different speakers in the same context, and the model can learn to distinguish different speakers based on language style in the same context based on this first negative sample; it can not only strengthen the model's recognition ability of the expression characteristics of a specific speaker, but also enable the model to more precisely maintain the unique style of the first speaker when generating responses.

[0056] The response statements of the first speaker in other conversations other than the first conversation in the multi-party conversation can be selected from other conversations D' other than the first conversation as the response of the first speaker response , and use it as a negative sample; by introducing different responses of the same speaker in different conversation contexts, the model can be forced to learn different expression ways that the same speaker may adopt in different scenarios; this can help the model establish "context dependence", that is, even for the same speaker, the generated response should be adjusted according to the context change, rather than blindly copying the typical expression pattern of the speaker.

[0057] In addition, use the response statement after disturbing the first positive sample as the first negative sample, so that the model needs to identify the differences in the content itself during contrastive learning, rather than just the style; it can avoid the model generating incorrect responses that do not match the content but have a similar style. Introducing content interference can enable the model to consider both style and semantics when generating responses, ensuring that the response is consistent in style and accurate in content to fit the current context.

[0058] Specifically, the response statement after disturbing the first positive sample can be obtained by at least one of the following two methods: First, replace specific names, locations or dates with other words to destroy the original semantics.

[0059] Second, interfere by using different word orders or sentence patterns, or replacing with approximate words, but retain some style features.

[0060] By providing reasonable prompts, the two replacement strategies can be effectively completed through the GPT-4 (a language model) model.

[0061] Through the above optimization of contrastive learning, the semantics of the query sample and the positive sample are similar, and the semantics of the query sample and the negative sample are distant, thus improving the ability of the speaking style perception model to understand the discourse intention.

[0062] Further, based on the corresponding feature representation and the dialogue context, the response sentence for the current dialogue of the first speaker is output through the speaking style perception model; since the speaking style perception model can distinguish semantically correct responses and semantically incorrect responses by analyzing the feature representations of the first query sample, the first positive sample, and the first negative sample, the speaking style perception model can generate appropriate response sentences based on the dialogue context and the learned ability to understand dialogue intents.

[0063] Further, in order to train a multi-party dialogue model that meets certain accuracy conditions, based on the corresponding feature representation and the response sentence of the first speaker, the model loss value can be calculated through a preset loss function; when the model loss value is greater than the target loss value, update the model parameters of the speaking style perception model, and return the step of obtaining the corresponding feature representations respectively through the speaking style perception model for contrastive learning of the first query sample, the first positive sample, and the first negative sample, until the preset conditions are met to obtain the multi-party dialogue model; when the model loss value is less than or equal to the target loss value or the maximum number of iterations is reached, the training can be stopped to obtain the multi-party dialogue model.

[0064] Among them, the preset loss function includes a first loss function and a second loss function, and the implementation of calculating the model loss value through the preset loss function based on the dialogue feature representation and the response sentence can be: Based on the corresponding feature representation, calculate the first loss value through the first loss function; based on the dialogue context sample, the response sentence, and the first positive sample, calculate the second loss value through the second loss function; perform weighted summation on the first loss value and the first loss value to obtain the model loss value.

[0065] It should be noted that the above contrastive learning and response task training are carried out simultaneously; while the model is performing contrastive learning and understanding the speaker, it also needs to generate responses that conform to the context. These two tasks promote and complement each other. Contrastive learning can help the model understand the context and the speaker, thus also enhancing the generation of responses; that is to say, the model needs to learn how to generate appropriate responses according to the context and also actually generate response sentences. Thereby, it can improve the model's ability to understand and generate natural language, and at the same time enhance the coherence and relevance of the dialogue while capturing the unique styles of different speakers.

[0066] Since the training of the above-mentioned contrast learning and response tasks is carried out simultaneously, during the model training process, based on the corresponding feature representation, the first loss value is calculated through the first loss function; based on the dialogue context sample, the response sentence, and the first positive sample, the second loss value is calculated through the second loss function; among them, the first loss function (such as the contrast loss) is specifically used to distinguish positive samples and negative samples, which helps the model learn to identify specific speaking styles or context clues, and the second loss function (such as the cross-entropy loss for the response generation task) ensures that the generated response is not only content-related but also style-matching; combining the two can more accurately capture and apply the speaking style. Therefore, in this embodiment, the first loss value and the second loss value are weighted and summed to obtain the model loss value, allowing adjustment of the relative importance between different tasks, that is, the model uses contrast learning to improve the response quality while generating responses, and the weighted loss strategy under the multi-task learning framework helps to avoid a single task dominating the entire training process, making the training more stable and reducing the risk of overfitting.

[0067] Specifically, the calculation formula of the preset loss function is as follows: ; Where represents the response generation loss, is the contrast learning loss, is a preset hyperparameter that adjusts the weight between these two loss terms.

[0068] The cross-entropy loss function (second loss function) for response generation is as follows:

[0069] Where is the token at position in the standard response (the basic unit when the AI model processes text or other types of data), represents the token sequence before position , is the dialogue context, is the information of the first speaker, and the information of the first speaker includes the information of the first speaker (for example, the ID or name of the first speaker in the dialogue context, etc.).

[0070] Among them, the contrast learning loss (first loss function) is defined as follows:

[0071] Where the parameter , , The quantities of the first negative samples from three different sources respectively (reply statements of other speakers different from the first speaker in the first conversation, reply statements of the first speaker in other conversations except the first conversation in the multi-party conversation, and reply statements after interfering with the first positive sample). It can be understood that if only one of the first negative samples is used for model training, then from , , the corresponding parameters are determined.

[0072] It should be noted that the entire above training process can be completed on an NVIDIA RTX Titan GPU.

[0073] In this embodiment, through the above two-stage model training (in the first stage, a speaking style perception model is obtained through contrastive learning, and in the second stage, the speaking style perception model is trained through joint learning of contrastive learning and reply generation tasks to obtain a multi-party conversation model), the multi-party conversation model can effectively capture the conversation context, improve the quality of reply generation, and enhance the ability to distinguish the speaking styles of speakers.

[0074] Moreover, in the second stage, the model compares the first query sample, the first positive sample, and the three first negative samples, enhancing the model's understanding of speaker interaction and conversation topics; making the generated reply highly consistent with the first speaker in terms of semantics and style.

[0075] It should be noted that the above examples are only for understanding this application and do not constitute a limitation on the training method of the multi-party conversation model of this application. Based on this technical concept, more forms of simple transformations are within the protection scope of this application.

[0076] This application also provides a multi-party conversation system. Please refer to Figure 3 , the multi-party conversation system includes: An interaction module (platform interaction layer) for, in response to a speech input by a user, sending the speech to the logic processing module and displaying a reply statement; A logic processing module (logic processing layer) for processing the speech input by the user based on a pre-deployed multi-party conversation model, generating a reply statement, and returning it to the interaction module, where the multi-party conversation model is obtained by iteratively training a speaking style perception model based on the conversation context sample of the first speaker in the first conversation of the multi-party conversation of multi-party conversation samples, and the speaking style perception model is obtained by contrastive learning of a preset encoder based on the speeches of each speaker in the multi-party conversation samples.

[0077] Furthermore, the multi-party conversation system of this embodiment further includes a data management module (data management layer). Please refer toFigure 3 ; In practical applications, a multi-party dialogue system can be designed for the open-domain scenario, and a human-computer interaction system can be built based on Web technology and run in a browser; the interaction module belongs to the client part, mainly used to receive user input and visually display the response data processed by the model; the logic processing module is located on the server side and provides processing services through interface calls. The logic processing module (including the request processing module and the dialogue generation module) can process user input and generate appropriate responses based on a pre-trained offline algorithm model, and then the interaction module performs visual display; the data management module (dialogue storage module) is mainly used to store and manage relevant data during the above user dialogue process. The above multi-party dialogue system adopts a modular design, emphasizing low coupling between modules and high cohesion within modules. Through coordinated cooperation, the system development is completed, realizing the separation of the front end and the back end, which is convenient for subsequent maintenance and update.

[0078] Specifically, the front end of the human-computer interaction system can use the Vue.js framework to build the user interface. With its powerful componentization and data binding capabilities, a simple and easy-to-use chat interface can be created. To ensure efficient data interaction between the front end and the back end, the Axios library is used to send requests and receive responses through the HTTP (Hypertext Transfer Protocol) protocol. In particular, the message input by the user is encapsulated in JSON format and passed to the back end through a POST request; CSS (Cascading Style Sheets) can also be used to beautify the front-end interface to ensure that the chat interface is not only fully functional but also beautiful and generous, enhancing the user experience; referring to Figure 4 , the entire front end mainly consists of three core modules: the user input module, the message display module, and the front-back end interaction module. The user can send messages through the input box. After the front end receives these messages, it formats them into JSON and initiates a network request through Axios to send the information to the back end. At the same time, the front end also receives data from the back end, including the dialogue history record and the real-time updated message list, and performs corresponding displays on the interface.

[0079] The backend of the human-computer interaction system can be built using the FastAPI framework (a fast (high-performance) web framework); FastAPI combines a coroutine pool and a task queue to handle concurrent requests, where asyncio.Queue is used to manage user requests, ensuring that all requests can be processed in order and effectively avoiding system overload caused by excessive requests; through the shared event loop mechanism, the coroutine pool can significantly improve the system's concurrent processing ability, minimizing the occurrence of blocking situations, so that requests from multiple users can be quickly responded to. Such an architecture design not only meets the real-time communication needs in a multi-user environment but also ensures the scalability and maintainability of the system.

[0080] In addition, to ensure the user experience of the multi-party dialogue system, the above multi-party dialogue system also has the following functions, refer to Figure 4 and Figure 5 : Login function: Users can log in by entering their username and password; if the login information entered by the user is invalid, the system will provide corresponding error prompts to remind the user to make corrections.

[0081] Logout function: Users can end the current conversation at any time, and the system will automatically return to the login interface and clear the current login information.

[0082] Registration function: New users can create an account through the registration process to use the various functions of the system.

[0083] System interaction function: Users can have a conversation with the system.

[0084] Message time display: For each message, the system will display its sending time.

[0085] Clear conversation history: Users can choose to clear the history of the current session to clean up unnecessary information.

[0086] Retry function: If the user is not satisfied with the reply of the current role, the retry function can be triggered, and the system will regenerate the reply of that role.

[0087] Load conversation history: Users can load previous conversation records to restore the historical state or continue an unfinished conversation.

[0088] In this embodiment, by deploying a multi-party dialogue model, friendly interaction with users is achieved. Moreover, by using the training method of the multi-party dialogue model in the above embodiment, the technical problem of low user dialogue satisfaction can be solved. Compared with the prior art, the beneficial effects of the multi-party dialogue system provided in this application are the same as those of the training method of the multi-party dialogue model provided in the above embodiment, and the other technical features in the multi-party dialogue system are the same as those disclosed in the method of the above embodiment, which will not be elaborated here.

[0089] This application provides a training device for a multi-party dialogue model. The training device for the multi-party dialogue model includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the training method of the multi-party dialogue model in Embodiment 1 above.

[0090] The following refers to Figure 6 , which shows a schematic structural diagram of a training device for a multi-party dialogue model suitable for implementing the embodiments of this application. The training device for the multi-party dialogue model in the embodiments of this application may include, but is not limited to, mobile terminals such as mobile phones, tablet computers, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PMPs (Portable Media Players), in-vehicle terminals (such as in-vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 6 The training device for the multi-party dialogue model shown is only an example and should not impose any limitations on the functions and usage scope of the embodiments of this application.

[0091] As Figure 6As shown in the figure, the training device of the multi-party dialogue model may include a processing device 1001 (such as a central processing unit, a graphics processing unit, etc.), which can execute various appropriate actions and processes according to the program stored in the read-only memory (ROM: Read Only Memory) 1002 or the program loaded from the storage device 1003 into the random access memory (RAM: Random Access Memory) 1004. In the RAM 1004, various programs and data required for the operation of the training device of the multi-party dialogue model are also stored. The processing device 1001, the ROM 1002, and the RAM 1004 are connected to each other through a bus 1005. The input / output (I / O) interface 1006 is also connected to the bus. Generally, the following systems may be connected to the I / O interface 1006: an input device 1007 including, for example, a touch screen, a touchpad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; an output device 1008 including, for example, a liquid crystal display (LCD: Liquid Crystal Display), a speaker, a vibrator, etc.; a storage device 1003 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1009. The communication device 1009 can allow the training device of the multi-party dialogue model to communicate with other devices wirelessly or wiredly to exchange data. Although the figure shows the training device of the multi-party dialogue model with various systems, it should be understood that it is not required to implement or have all the shown systems. More or fewer systems can be implemented or had alternatively.

[0092] In particular, according to the embodiments disclosed in the present application, the process described above with reference to the flowchart can be implemented as a computer software program. For example, the embodiments disclosed in the present application include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes program codes for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication device, or installed from the storage device 1003, or installed from the ROM 1002. When the computer program is executed by the processing device 1001, the above functions defined in the method of the embodiments disclosed in the present application are executed.

[0093] The training device of the multi-party dialogue model provided by the present application adopts the training method of the multi-party dialogue model in the above embodiments, and can solve the technical problem of low user dialogue satisfaction. Compared with the prior art, the beneficial effects of the training device of the multi-party dialogue model provided by the present application are the same as those of the training method of the multi-party dialogue model provided by the above embodiments, and other technical features in the training device of the multi-party dialogue model are the same as the features disclosed in the method of the previous embodiment, and will not be elaborated here.

[0094] It should be understood that the various parts disclosed in this application can be implemented by hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in a suitable manner in any one or more embodiments or examples.

[0095] As described above, the above is only the specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art within the technical scope disclosed in this application can easily think of changes or replacements, which should all be covered within the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.

[0096] This application provides a computer-readable storage medium with computer-readable program instructions (i.e., computer programs) stored thereon, and the computer-readable program instructions are used to execute the training method of the multi-party dialogue model in the above embodiments.

[0097] The computer-readable storage medium provided by this application can be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or any combination of the above. More specific examples of computer-readable storage media can include, but are not limited to: electrical connections with one or more wires, portable computer disks, hard disks, random access memory (RAM: Random Access Memory), read-only memory (ROM: Read Only Memory), erasable programmable read-only memory (EPROM: Erasable Programmable Read Only Memory or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM: CD-Read Only Memory), optical storage devices, magnetic storage devices, or any suitable combination of the above. In this embodiment, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, device, or device. The program code contained on the computer-readable storage medium can be transmitted by any appropriate medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination of the above.

[0098] The above computer-readable storage medium can be included in the training device of the multi-party dialogue model; it can also exist separately and not be assembled into the training device of the multi-party dialogue model.

[0099] The above computer-readable storage medium carries one or more programs. When the above one or more programs are executed by the training device of the multi-party dialogue model, the training device of the multi-party dialogue model is caused to: execute the training method of the multi-party dialogue model.

[0100] Computer program code for performing the operations of the present application can be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (e.g., through the Internet using an Internet service provider).

[0101] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of the systems, methods, and computer program products according to various embodiments of the present application. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a portion of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0102] The modules described in the embodiments of the present application can be implemented in software or in hardware. In some cases, the name of the module does not constitute a limitation on the unit itself.

[0103] The readable storage medium provided by the present application is a computer-readable storage medium that stores computer-readable program instructions (i.e., computer programs) for performing the training method of the above multi-party dialogue model, and can solve the technical problem of low user dialogue satisfaction. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided by the present application are the same as those of the training method of the multi-party dialogue model provided by the above embodiments, and will not be elaborated here.

[0104] The present application also provides a computer program product, including a computer program which, when executed by a processor, implements the steps of the training method of the multi-party dialogue model as described above.

[0105] The computer program product provided by the present application can solve the technical problem of low user dialogue satisfaction. Compared with the prior art, the beneficial effects of the computer program product provided by the present application are the same as those of the training method of the multi-party dialogue model provided by the above embodiments, and will not be elaborated herein.

[0106] The above are only partial embodiments of the present application, and thus do not limit the patent scope of the present application. Any equivalent structural transformation made under the technical concept of the present application by using the content of the specification and drawings of the present application, or any direct / indirect application in other related technical fields, is included in the patent protection scope of the present application.

Claims

1. A training method for a multi-party dialogue model, characterized in that, The method described above includes: Obtaining a multi-party conversation sample and a conversation context sample of the first speaker in the first conversation of the multi-party conversation; Performing contrastive learning on a preset encoder based on the speeches of each speaker in the multi-party conversation sample to obtain a speech style perception model; Performing iterative training on the speech style perception model based on the conversation context sample to obtain a multi-party conversation model, where the multi-party conversation model is used to respond to the speeches of different speakers in different speech styles.

2. The method according to claim 1, wherein The step of performing iterative training on the speech style perception model based on the conversation context sample to obtain a multi-party conversation model further includes: Taking the conversation context in the conversation context sample as a first query sample, taking the true reply in the conversation context as a first positive sample, and taking the wrong reply in the conversation context as a first negative sample; Performing iterative training on the speech style perception model based on the conversation context, the first query sample, the first positive sample, and the first negative sample to obtain a multi-party conversation model.

3. The method according to claim 2, wherein The step of performing iterative training on the speech style perception model based on the conversation context, the first query sample, the first positive sample, and the first negative sample to obtain a multi-party conversation model includes: Performing contrastive learning on the first query sample, the first positive sample, and the first negative sample through the speech style perception model to obtain corresponding feature representations respectively; Based on the corresponding feature representations and the conversation context, outputting a reply statement for the current conversation of the first speaker through the speech style perception model; Calculating a model loss value through a preset loss function based on the corresponding feature representations and the reply statement of the first speaker; When the model loss value is greater than the target loss value, updating the model parameters of the speech style perception model and returning to the step of performing contrastive learning on the first query sample, the first positive sample, and the first negative sample through the speech style perception model to obtain corresponding feature representations respectively, until a preset condition is met to obtain the multi-party conversation model.

4. The method according to claim 3, wherein The preset loss function includes a first loss function and a second loss function. The step of calculating a model loss value through a preset loss function based on the conversation feature representation and the reply statement includes: Calculating a first loss value through the first loss function based on the corresponding feature representations; Calculating a second loss value through the second loss function based on the conversation context sample, the reply statement, and the first positive sample; Performing weighted summation on the first loss value and the first loss value to obtain a model loss value.

5. The method according to any one of claims 2 to 4, characterized in that, The first negative sample includes at least one of the reply statements of other speakers different from the first speaker in the first conversation, the reply statements of the first speaker in other conversations except the first conversation in the multi-party conversation, and the reply statements after interference processing on the first positive sample.

6. The method according to claim 1, wherein The step of performing contrastive learning on a preset encoder based on the speeches of each speaker in the multi-party dialogue sample to obtain a speech style perception model includes: Taking the first speech of the second speaker in the second dialogue in the multi-party dialogue sample as the second query sample, taking the second speech of the second speaker in the second dialogue as the second positive sample, and taking the speeches of other speakers except the second speaker as the second negative samples; Performing contrastive learning on the preset encoder based on the second query sample, the second positive sample, and the second negative samples to obtain a speech style perception model.

7. A multi-party dialogue system, characterized in that, The multi-party dialogue system includes: An interaction module, configured to respond to a speech input by a user, send the speech to a logic processing module, and display a reply statement; A logic processing module, configured to process the speech input by the user based on a pre-deployed multi-party dialogue model, generate a reply statement, and return it to the interaction module, where the multi-party dialogue model is obtained by iteratively training the speech style perception model based on the dialogue context sample of the first speaker in the first dialogue of the multi-party dialogue in the multi-party dialogue sample, and the speech style perception model is obtained by performing contrastive learning on a preset encoder based on the speeches of each speaker in the multi-party dialogue sample.

8. A training device for a multi-party dialogue model, characterized in that, The device includes: a memory, a processor, and a computer program stored on the memory and executable on the processor, where the computer program is configured to implement the steps of the method for training a multi-party dialogue model according to any one of claims 1 to 6.

9. A storage medium, characterized in that, The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium, and when the computer program is executed by a processor, it implements the steps of the method for training a multi-party dialogue model according to any one of claims 1 to 6.

10. A computer program product, characterized in that, The computer program product includes a computer program, and when the computer program is executed by a processor, it implements the steps of the method for training a multi-party dialogue model according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Image recognition model training optimization method and device, electronic equipment and storage medium

    CN117576519A