Dialogue model training method and apparatus, dialogue method and apparatus, and electronic device
By identifying first- and second-class word pairs in the dialogue model and training the model using a contrastive learning strategy, the shortcomings of existing dialogue models in recognizing ambiguous entities are addressed, and the model's ability to recognize and distinguish ambiguous entities is improved.
Patent Information
- Application Number
- PCT/CN2025/110061
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-07-30
- Filing Date
- 2025-07-23
- Publication Date
- 2026-02-05
AI Technical Summary
Existing dialogue models perform poorly when dealing with ambiguous entities and cannot effectively distinguish entity references in different contexts.
By acquiring dialogue sample data, we determine the first and second types of word pairs, and use a contrastive learning strategy to train the initial dialogue model, thereby improving the model's ability to recognize entity references.
It enhances the dialogue model's ability to identify and distinguish ambiguous entities, thereby improving the model's performance and dialogue understanding capabilities.
Smart Images

Figure CN2025110061_05022026_PF_FP_ABST
Abstract
Description
Dialogue model training methods, dialogue methods, devices and electronic equipment
[0001] Cross-references to related applications
[0002] This disclosure claims priority to Chinese Patent Application No. 202411029216.0, filed in China on July 30, 2024, the entire contents of which are incorporated herein by reference. Technical Field
[0003] This disclosure relates to the field of communication technology, and in particular to a dialogue model training method, dialogue method, apparatus, and electronic device. Background Technology
[0004] Task-Oriented Dialogue Systems (TODS) are an important research area in artificial intelligence, aiming to help users efficiently complete specific tasks through natural language interaction, such as booking hotels and checking the weather. The core of this system lies in understanding the user's intent and providing accurate information and operational guidance during the dialogue. To achieve this goal, TODS needs strong language understanding, dialogue management, and task execution capabilities. With technological advancements, the design and implementation of TODS are constantly improving. For example, by integrating natural language processing technology, TODS can more accurately identify user instructions and needs, thereby providing more personalized services. Furthermore, TODS can continuously optimize its dialogue strategy (e.g., optimize the dialogue model) through machine learning algorithms to better adapt to different user behavior patterns and preferences. The widespread application of such systems not only improves people's life efficiency but also brings convenience to business operations. For example, in e-commerce and customer service, TODS can significantly improve user experience and business processing efficiency. In recent years, deep learning technology, especially Large Language Models (LLMs), has been increasingly widely used as dialogue models in TODS.
[0005] However, entities in dialogues are often ambiguous, meaning that the same word may refer to different entities in different contexts, and different words may refer to the same entity. Existing dialogue models are poor at distinguishing such data, meaning that existing dialogue models are inefficient. Summary of the Invention
[0006] This disclosure provides a dialogue model training method, dialogue method, apparatus, electronic device, computer-readable storage medium, and computer program product to solve the problem of poor performance of existing dialogue models.
[0007] To solve the above-mentioned technical problems, this disclosure is implemented as follows:
[0008] In a first aspect, embodiments of this disclosure provide a dialogue model training method, including:
[0009] Obtain dialogue sample data;
[0010] Based on the dialogue sample data, a first type of word pair and a second type of word pair are determined. In the first type of word pair, the two words in each word pair belong to the same entity and have the same semantic reference relationship. In the second type of word pair, the two words in each word pair do not belong to the same entity and have different semantic reference relationships.
[0011] Based on the dialogue sample data, the first type of word pairs, and the second type of word pairs, the initial dialogue model is trained using a contrastive learning strategy to obtain the target dialogue model.
[0012] Secondly, embodiments of this disclosure provide a dialogue method, including:
[0013] Obtain dialogue input data;
[0014] The feature vector of the dialogue input data is determined by the target dialogue model;
[0015] Based on the feature vector of the dialogue input data, the operation result of the dialogue input data is output;
[0016] The target dialogue model is the target dialogue model trained using the above-mentioned dialogue model training method.
[0017] Thirdly, embodiments of this disclosure provide a model training apparatus, the apparatus comprising:
[0018] The first acquisition module is used to acquire dialogue sample data;
[0019] The first determining module is used to determine a first type of word pair and a second type of word pair based on the dialogue sample data. The two words in each word pair of the first type of word pair belong to the same entity and have the same semantic reference relationship. The two words in each word pair of the second type of word pair do not belong to the same entity and have different semantic reference relationships.
[0020] The training module is used to train the initial dialogue model based on the dialogue sample data, the first type of word pairs, and the second type of word pairs, using a contrastive learning strategy to obtain the target dialogue model.
[0021] Fourthly, embodiments of this disclosure provide a dialogue device, the device comprising:
[0022] The second acquisition module is used to acquire dialogue input data;
[0023] A feature determination module is used to determine the feature vector of the dialogue input data through a target dialogue model;
[0024] The output module is used to output the operation result of the dialogue input data based on the feature vector of the dialogue input data;
[0025] The target dialogue model is the target dialogue model trained using the above-mentioned dialogue model training method.
[0026] Fifthly, embodiments of this disclosure provide an electronic device, including a transceiver and a processor.
[0027] The processor is used for:
[0028] Obtain dialogue sample data;
[0029] Based on the dialogue sample data, a first type of word pair and a second type of word pair are determined. In the first type of word pair, the two words in each word pair belong to the same entity and have the same semantic reference relationship. In the second type of word pair, the two words in each word pair do not belong to the same entity and have different semantic reference relationships.
[0030] Based on the dialogue sample data, the first type of word pairs, and the second type of word pairs, the initial dialogue model is trained using a contrastive learning strategy to obtain the target dialogue model.
[0031] Sixthly, embodiments of this disclosure provide an electronic device, including a transceiver and a processor.
[0032] The processor is used for:
[0033] Obtain dialogue input data;
[0034] The feature vector of the dialogue input data is determined by the target dialogue model;
[0035] Based on the feature vector of the dialogue input data, the operation result of the dialogue input data is output;
[0036] The target dialogue model is a target dialogue model trained using the above-described model training method.
[0037] In a seventh aspect, embodiments of this disclosure provide an electronic device, including: a processor, a memory, and a program stored in the memory and executable on the processor, wherein the program, when executed by the processor, implements the steps of the above-described method.
[0038] Eighthly, embodiments of this disclosure provide a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the above-described method.
[0039] Ninthly, embodiments of this disclosure also provide a computer program product, including a computer program that, when executed by a processor, implements the steps in the above-described method.
[0040] In the model training process of this embodiment, a first type of word pair and a second type of word pair can be determined based on the dialogue sample data. In the first type of word pair, the two words in each pair belong to the same entity and have the same semantic reference relationship. In the second type of word pair, the two words in each pair do not belong to the same entity and have different semantic reference relationships. Based on the dialogue sample data, the first type of word pair, and the second type of word pair, a contrastive learning strategy can be used to train the initial dialogue model to obtain the target dialogue model. This allows the dialogue model to learn the similarities and differences in the dialogue sample data, enabling it to better distinguish between identical and different data, thus improving the performance of the Hua model. Attached Figure Description
[0041] To more clearly illustrate the technical solutions of the embodiments of this disclosure, the accompanying drawings used in the description of the embodiments of this disclosure will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this disclosure. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0042] Figure 1 is a flowchart of a dialogue model training method provided in an embodiment of this disclosure;
[0043] Figure 2 is a flowchart of a dialogue method provided in an embodiment of this disclosure;
[0044] Figure 3 is a block diagram illustrating the principle of the dialogue model training method provided in the embodiments of this disclosure;
[0045] Figure 4 is a schematic diagram of a language model training principle provided in an embodiment of this disclosure;
[0046] Figure 5 is a schematic diagram of a graph encoder training principle provided in an embodiment of this disclosure;
[0047] Figure 6 is a schematic diagram of a dialogue model training device provided in an embodiment of this disclosure;
[0048] Figure 7 is a schematic diagram of a dialogue device provided in an embodiment of this disclosure;
[0049] Figure 8 is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure;
[0050] Figure 9 is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure. Detailed Implementation
[0051] The technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this disclosure. Based on the embodiments of this disclosure, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this disclosure.
[0052] Referring to Figure 1, which is a flowchart of a dialogue model training method provided in this embodiment, the method is executed by a first electronic device, which may be a terminal, a server device, or the like. As shown in Figure 1, the dialogue model training method provided in this embodiment includes the following steps:
[0053] Step 101: Obtain dialogue sample data.
[0054] Dialogue sample data may include, but is not limited to, text-based dialogue data and voice-based dialogue data. It should be noted that dialogue sample data may include dialogue data between at least one role. For example, a certain dialogue sample data may include dialogue data between a user and the system. For example, it may include "User (USR): I want to take a train to city X on Monday, when is it available?; System (SYS): I have several possibilities for Monday. User (USR): ...book a table for 5 people at 21:00 on the same day, please. System (SYS): OK. The table will be reserved for 15 minutes."
[0055] Step 102: Based on the dialogue sample data, determine the first type of word pairs and the second type of word pairs. In the first type of word pairs, the two words in each word pair belong to the same entity and have the same semantic reference relationship. In the second type of word pairs, the two words in each word pair do not belong to the same entity and have different semantic reference relationships.
[0056] It should be noted that the number of dialogue sample data can be one or more. If there are multiple dialogue sample data, this step can be performed for each dialogue sample data. That is, for each dialogue sample data, the first type of word pair and the second type of word pair corresponding to the dialogue sample data can be determined.
[0057] Furthermore, the first type of word pair may include at least one first word pair, and the second type of word pair may include at least one second word pair. It should be noted that each word pair includes two words, and at least one word differs between different word pairs. It can be understood that the words in the word pairs in this embodiment are the result of entity recognition of the dialogue sample data. For each word pair in the first type of word pair, the two words belong to the same entity and have the same semantic reference relationship, i.e., they share a reference relationship. For each word pair in the second type of word pair, the two words do not belong to the same entity and have different semantic reference relationships, i.e., they do not share a reference relationship. For example, in the above dialogue data, "Monday" and "the same day" belong to the same entity and share a semantic reference relationship, while "a train" and "the same day" do not belong to the same entity and have different semantic reference relationships.
[0058] As an example, for a certain dialogue sample data, one can first determine the first type of word pairs based on the dialogue sample data, and then determine the second type of word pairs based on the first type of word pairs. Alternatively, one can first determine the second type of word pairs based on the dialogue sample data, and then determine the first type of word pairs based on the second type of word pairs. This embodiment does not impose a specific limitation. For example, in the process of first determining the first type of word pairs based on the dialogue sample data and then determining the second type of word pairs, one can first determine the first type of word pairs based on the dialogue sample data, and then replace any word in each word pair of the first type of word pairs to obtain the second type of word pairs. Furthermore, in the process of replacing a word in a word pair within the first type of word pairs, another word from the dialogue sample data belonging to a different entity than that word can be used to replace that word, thus obtaining the replaced word pair, which is one of the word pairs in the second type of word pairs. For example, in the first type of word pair, there is the pair "Monday" and "the same day". "A train" belongs to a different entity than "Monday" in the sample dialogue data. Replacing "Monday" with "a train" results in the corresponding word pair "a train" and "the same day", which belongs to the second type of word pair. In one example, for a certain dialogue sample data, for each word pair in the first type of word pair, m random samples can be used... c(positive integer) replacement words (belonging to this dialogue sample data) replace one word in the word pair, and randomly sample m... c A positive integer (m) replacement word (belonging to this dialogue sample data) replaces another word in the word pair, resulting in the corresponding (m) word pair. c +m c The first set of word pairs is replaced with '', and then placed into the second set of word pairs. For each word pair in the first set of word pairs, the above process can be repeated to obtain the final second set of word pairs.
[0059] Step 103: Based on the dialogue sample data, the first type of word pairs, and the second type of word pairs, train the initial dialogue model using a contrastive learning strategy to obtain the target dialogue model.
[0060] Contrastive learning is a machine learning method that learns by comparing the similarities and differences between two things. It trains the model to identify which data points are similar or different, learning the general features of unlabeled sample data to distinguish between similar and dissimilar data points. The goal is to learn the representation of the data to capture the basic structure and relationships between different data points. In the first type of word pairs determined based on dialogue sample data, each word pair belongs to the same entity but has different semantic references, meaning they refer to the same thing. In the second type of word pairs, each word pair does not belong to the same entity and has different semantic references, meaning they refer to different things. Contrastive learning strategies can learn the similarities and differences in data. In this embodiment, based on dialogue sample data, the first type of word pairs, and the second type of word pairs, a contrastive learning strategy can be used to train an initial dialogue model to obtain a target dialogue model, thereby improving the performance of the trained model.
[0061] Furthermore, it should be noted that the method of this embodiment can be applied to various scenarios, such as dialogue scenarios. Specifically, it can be applied to, but is not limited to, information query scenarios (e.g., weather queries), information booking scenarios (airline ticket booking, high-speed rail ticket booking, ship ticket booking, bus ticket booking, etc.), e-commerce scenarios, customer service scenarios, etc. The aforementioned dialogue sample data can be dialogue data in the specific scenario in which this method is applied. The aforementioned dialogue model can be a dialogue model in the scenario in which this method is applied, which can be used to extract features from the user's input data in that scenario for subsequent target tasks, such as user intent recognition and outputting corresponding response information based on the recognized user intent.
[0062] In the model training process of this embodiment, firstly, based on the dialogue sample data, a first type of word pair and a second type of word pair can be determined. In the first type of word pair, the two words in each pair belong to the same entity and have the same semantic reference relationship. In the second type of word pair, the two words in each pair do not belong to the same entity and have different semantic reference relationships. Based on the dialogue sample data, the first type of word pair, and the second type of word pair, the initial dialogue model can be trained using a contrastive learning strategy to obtain the target dialogue model. This allows the dialogue model to learn the similarities and differences in the dialogue sample data, enabling it to better distinguish between identical and different data, thereby improving the performance of the dialogue model.
[0063] In one embodiment, the initial dialogue model includes an initial language model, and the target dialogue model includes a target language model;
[0064] By employing a contrastive learning strategy, the initial dialogue model is trained based on dialogue sample data, first-class word pairs, and second-class word pairs to obtain the target dialogue model, which includes:
[0065] The semantic vectors of words in the dialogue sample data are determined using an initial language model.
[0066] Based on the semantic vectors of the two words in each word pair of the first type of word pair and the semantic vectors of the two words in each word pair of the second type of word pair, the result of the target loss function of the contrastive learning strategy is determined.
[0067] Based on the results of the target loss function, the initial language model is trained to obtain the target language model.
[0068] It is understood that the language model in this embodiment can be a pre-trained language model, which is a type of text encoder. The model training process is the process of continuously adjusting the model parameters based on the results of the loss function. In the model training process of this embodiment, the target language model can be obtained by training an initial language model through a contrastive learning strategy. In this process, the initial language model can be started first to determine the semantic vectors of words in the dialogue sample data. In this way, the semantic vectors of the two words in each word pair of the first type of word pair and the semantic vectors of the two words in each word pair of the second type of word pair can be obtained. Then, using the semantic vectors of the two words in each word pair of the first type of word pair and the semantic vectors of the two words in each word pair of the second type of word pair, the result of the target loss function of the contrastive learning strategy is obtained. The initial language model is trained using the result of the target loss function of the contrastive learning strategy to obtain the target language model, which enables the target language model to better distinguish between the same and different semantic data, thereby improving the model performance.
[0069] Additionally, it should be noted that the dialogue sample data is input into the initial language model, which obtains the semantic vector of each word in the dialogue sample data. Based on the semantic vector of each word in the dialogue sample data, the semantic vector of each word (corresponding entity, also called entity word) in the dialogue sample data is determined. For example, entity recognition can be performed on the dialogue sample data to obtain the words in the dialogue sample data. For each word in the entity recognition, the semantic vector of the word is obtained based on the semantic vector of each word in the word. There are various ways to obtain the semantic vector of a word based on the semantic vector of each word in the word, and no specific limitation is made. For example, pooling can be performed based on the semantic vector of each word in the word to obtain the semantic vector of the word. More specifically, the average of the semantic vectors of each word in the word can be used as the semantic vector of the word, etc.
[0070] In some embodiments, there are multiple dialogue sample data sets, and the result of the target loss function is the sum of the values of the semantic contrast loss functions of the dialogue sample data sets. The value of the semantic contrast loss function of each dialogue sample data set is obtained as follows:
[0071] It is calculated using the semantic vectors of the two words in each word pair of the first type of word pair in the dialogue sample data and the semantic vectors of the two words in each word pair of the second type of word pair in the dialogue sample data.
[0072] Since there can be multiple dialogue sample data sets, the process of determining the first and second category word pairs described above can be performed for each dialogue sample data set. This yields the corresponding first and second category word pairs for each dialogue sample data set. During training, training can be performed based on multiple dialogue sample data sets and their corresponding first and second category word pairs. For each dialogue sample data set, the semantic contrastive loss function value corresponding to the contrastive learning strategy can be obtained based on the semantic vectors of the two words in each word pair of the first category word pair and the semantic vectors of the two words in each word pair of the second category word pair. This similar operation is performed for each dialogue sample data set, thus obtaining the semantic contrastive loss function values corresponding to multiple dialogue sample data sets. The sum of the semantic contrastive loss function values corresponding to multiple dialogue sample data sets can be used as the result of the target loss function for model training to improve model performance.
[0073] In some embodiments, the initial dialogue model further includes an initial graph encoder, and the target dialogue model further includes a target graph encoder; the dialogue sample data is multiple;
[0074] By employing a contrastive learning strategy, the initial dialogue model is trained based on dialogue sample data, first-class word pairs, and second-class word pairs to obtain the target dialogue model, which includes:
[0075] For each dialogue sample in multiple dialogue sample data sets, a dialogue graph is constructed based on the semantic vectors of words in the dialogue sample data, the positional relationships of words in the dialogue sample data, the dialogue roles included in the dialogue sample data, and the semantic reference relationships between words in the dialogue sample data. The dialogue graph includes N nodes and the connections between N nodes. Nodes are used to represent the semantic vectors of words or dialogue roles. The number of nodes corresponding to the same dialogue role in the dialogue graph is the same as the number of dialogues of the dialogue role in the dialogue sample data. Different nodes of the same dialogue role correspond to different dialogue content. Connections are made between nodes with the same semantic reference relationship, between nodes with the same dialogue role, between nodes of different dialogue roles with dialogue, between nodes of each dialogue role and the nodes corresponding to each word in the dialogue content of the node, and between nodes corresponding to adjacent words in the dialogue sample data.
[0076] The initial graph encoder determines the global feature representation of each dialogue graph and the node feature vector of each node in each dialogue graph.
[0077] An initial graph encoder is trained based on at least one first sample pair and at least one second sample pair to obtain a target graph encoder. The first sample pair includes the global feature representation of the first dialogue graph and the node feature representation of each node in the first dialogue graph. The first dialogue graph is any one of the multiple dialogue graphs. The second sample pair includes the global feature representation of the first dialogue graph and the node feature representation of each node in the second dialogue graph. The second dialogue graph is any one of the multiple dialogue graphs other than the first dialogue graph.
[0078] It should be noted that in the dialogue graph, each node representing the semantic vector of a word corresponds to the semantic vector of a single word, and different nodes in the dialogue graph represent the semantic vectors of different words. Similarly, each node representing a dialogue character in the dialogue graph corresponds to a single dialogue character, and different nodes in the dialogue character node can correspond to different dialogue characters or the same dialogue character. The number of instances of the same dialogue character in the dialogue graph corresponding to the dialogue sample data is related to the number of times that dialogue character has been in dialogue within the dialogue sample dataset. For example, as illustrated above, user A enters into two conversations. The first conversation is "I want to take a train to city X on Monday, when's available?", and the second conversation is "...book a table for 5 people at 21:00 on the same day, please." The system then conducts two more conversations: the first is "I have several possibilities for Monday.", and the second is "OK. The table will be reserved for 15 minutes." In the corresponding conversation graph, user A has two nodes. One node corresponds to user A's first conversation, where each word in user A's first conversation is connected to this node, and adjacent words in user A's first conversation are connected by edges. The other node corresponds to user A's second conversation, where each word in user A's second conversation is connected to this node, and adjacent words in user A's second conversation are connected by edges. The two nodes corresponding to user A are also connected. There are two nodes corresponding to each system role. One node corresponds to the first dialogue content of the system. It has a dialogue with the node in the first dialogue content of user A, and there are edges connecting them. Each word node in the first dialogue content of the system is connected to this node, and there are edges connecting adjacent words in the first dialogue content of the system. The other node corresponding to the system role corresponds to the second dialogue content of the system. It also has a dialogue with the node in the second dialogue content of user A, and there are edges connecting each word node in the second dialogue content of the system. The two nodes corresponding to the system role are connected. Furthermore, in the dialogue graph, the node representation (embedding) of the dialogue role's node can be obtained from the semantic vectors of all nodes connected to that node, for example, it could be the average of the semantic vectors of all nodes connected to that node. It should also be noted that in this embodiment, "adjacent" can be understood as being adjacent in position within the dialogue sample data.
[0079] For each dialogue sample in the multiple dialogue sample data sets, a corresponding dialogue graph is constructed through the above process, resulting in multiple dialogue graphs. Each dialogue graph can then be input into an initial graph encoder. The initial graph encoder determines the global feature representation of each dialogue graph and the node feature vectors of each node in each dialogue graph. Then, based on the global feature representations and node feature vectors of the multiple dialogue graphs, at least one first sample pair and at least one second sample pair can be constructed. The initial graph encoder is trained using these at least one first sample pair and at least one second sample pair to obtain the target graph encoder. It should be noted that at least one first sample pair corresponds to at least one dialogue graph. Each first sample pair includes the global feature representation of one dialogue graph and the node feature representations of each node in that dialogue graph. The corresponding first sample pair includes the global feature representation of that dialogue graph and the node feature representations of each node in another dialogue graph.
[0080] In this implementation, the semantic feature information corresponding to the dialogue sample data is used during the language model training process. In this embodiment, a graph encoder can also be trained. The graph structure can represent the structured information of the dialogue sample data. The graph encoder can be trained using the graph structure information corresponding to the dialogue sample data to obtain the target graph encoder, thereby improving the performance of the dialogue model.
[0081] In some embodiments, an initial graph encoder is trained based on at least one first sample pair and at least one second sample pair to obtain a target graph encoder, including:
[0082] At least one first sample pair and at least one second sample pair are respectively input into the discriminator for discrimination;
[0083] Based on the mutual information output by the discriminator, determine the result of the graph contrast loss function;
[0084] Based on the results of the graph contrast loss function, the initial graph encoder is trained to obtain the target graph encoder.
[0085] It is understood that the results output by the discriminator may include the global feature representation of the first dialogue graph corresponding to each first sample and the mutual information (MI) between each node in the first dialogue graph, as well as the global feature representation of the first dialogue graph corresponding to each second sample and the mutual information between each node in the second dialogue graph.
[0086] Mutual information can be used to measure the correlation between two pieces of information. In this implementation, based on the first sample pair, the corresponding mutual information can be obtained through the discriminator; based on the second sample pair, the corresponding mutual information can be obtained through the discriminator. The mutual information output by the discriminator determines the result of the graph contrastive loss function, which can also be understood as the result of the structural contrastive loss function. Based on the result of the graph contrastive loss function, the initial graph encoder is trained, enabling the graph encoder to learn the structural information of the dialogue sample data and improve the performance of the trained target graph encoder.
[0087] As shown in Figure 2, a dialogue method is also provided, which can be applied to a second electronic device. The second electronic device may be the same as or different from the first electronic device. The second electronic device may be a terminal, a server device, etc. The method includes:
[0088] Step 201: Obtain dialogue input data.
[0089] Dialogue input data can be understood as dialogue data entered by the user during actual application.
[0090] Step 202: Determine the feature vector of the dialogue input data through the target dialogue model, wherein the target dialogue model is the target dialogue model trained by the above dialogue model training method.
[0091] Step 203: Based on the feature vector of the dialogue input data, output the operation result of the dialogue input data.
[0092] The operation result here can be the result of performing corresponding tasks based on the feature vector of the dialogue input data (e.g., tasks in scenarios including but not limited to information query scenarios (e.g., weather query), information booking scenarios (flight ticket booking, high-speed rail ticket booking, ferry ticket booking, bus ticket booking), e-commerce scenarios, customer service scenarios, etc.). For example, in the weather query scenario, if the user inputs the dialogue data "Please query the weather in city B today", the operation result of the specific weather in city B today can be output. It should be noted that the feature vector of the dialogue input data can be the semantic features output by the target language model in the target dialogue model, or it can be the global graph representation of the dialogue graph corresponding to the dialogue input data output by the target graph encoder (the construction process is similar to the dialogue graph construction process of the dialogue sample data mentioned above, the difference being that the former is dialogue input data and the latter is dialogue sample data) and the node feature representation of each node in the dialogue graph.
[0093] The above process will be specifically described below with some specific embodiments.
[0094] Introduction to related technologies:
[0095] Task-Oriented Dialogue Systems (TODS) are an important research area in artificial intelligence, aiming to help users efficiently complete specific tasks, such as booking hotels or checking the weather, through natural language interaction. The core of such systems lies in understanding the user's intent and providing accurate information and operational guidance during the dialogue. To achieve this goal, TODS needs strong language understanding, dialogue management, and task execution capabilities. With technological advancements, the design and implementation of TODS are constantly improving. For example, by integrating natural language processing technology, TODS can more accurately identify user instructions and needs, thereby providing more personalized services. Furthermore, TODS can continuously optimize its dialogue strategies through machine learning algorithms to better adapt to different user behavior patterns and preferences. The widespread application of such systems not only improves people's lives but also brings convenience to business operations. For example, in e-commerce and customer service, TODS can significantly enhance user experience and business processing efficiency.
[0096] In recent years, deep learning technology, especially Large Language Models (LLMs), has been increasingly widely used in TODS (TOEFL / TOD) systems. LLMs, through pre-training and fine-tuning, can automatically learn and understand natural language, thus performing exceptionally well in multi-turn dialogues. These models are typically trained on large-scale text data, learning language patterns and structures, enabling them to understand and generate natural language. In TODS, LLMs can be used to understand user queries, generate appropriate answers, and even provide assistance in complex tasks. For example, an LLM-based TODS system can handle complex booking tasks, not only understanding the user's desired date, location, and type of booking, but also recommending suitable options based on the user's preferences. Furthermore, LLMs can continuously improve their performance in specific tasks through continuous learning and optimization, making TODS systems more intelligent and efficient. The widespread application of this technology has not only driven the development of TODS but also opened up new possibilities for the application of artificial intelligence in other fields.
[0097] However, inherent differences exist between the linguistic features of plain text and the dialogue context, which reduces the practicality of plain text-pretrained language models in task-oriented dialogue systems. The linguistic information and context contained in plain text data may not always be applicable in dialogue systems. Dialogue systems need to handle interactive and dynamic linguistic environments where user intentions and needs may change as the conversation progresses. Therefore, plain text-pretrained language models may not be able to adequately capture this dialogue dynamics and contextual dependencies.
[0098] In related technologies, the inherent differences in linguistic features between plain text and dialogue context reduce the practicality of plain text pre-trained language models in task-oriented dialogue systems. This difference is primarily manifested in the interactivity and dynamism of dialogue. In a dialogue, a user's intent and needs constantly change as the conversation progresses, while plain text data often lacks this dynamism. Therefore, traditional pre-trained language models struggle to capture contextual changes and the evolution of user intent within a dialogue. For example, in task-oriented dialogues, a user might make a request and then adjust or refine it based on the system's response. This dynamic change requires the dialogue system to update its understanding of user intent in real time. However, plain text pre-trained models often fail to do this because they do not learn this dialogue dynamism during the pre-training phase. Furthermore, contextual information in dialogue is often non-linear; a user's current utterance may be related to multiple previous utterances, and plain text models struggle to handle such complex contextual relationships.
[0099] Furthermore, current technologies often fail to fully capture the flow of information and contextual dependencies in dialogue. One of the core challenges of dialogue systems is understanding and predicting the flow of information within a conversation. During a dialogue, information is transmitted and updated continuously, not in isolation. Each utterance by a user may depend on previous dialogue content, and the system's response needs to consider the user's overall intent and needs. Existing pre-trained language models, such as bidirectional encoder representations from transformers (BERT), while performing well on general text data, face challenges in handling dialogue data due to contextual dependencies. These models typically lack a deep understanding of dialogue structure and the patterns of information flow, leading to poor performance in tasks such as dialogue state tracking and intent recognition. To address this issue, researchers need to develop new pre-training strategies that enable models to better understand and predict the flow of information and contextual dependencies in dialogue.
[0100] Moreover, entities are crucial elements for coherently understanding the entire dialogue, yet current technologies often fail to fully capture or utilize the semantic and structural relationships they represent. In task-oriented dialogues, entity recognition and linking are key to understanding user intent and needs. However, existing pre-trained language models often have limitations in handling entities within dialogues. These models may fail to accurately identify and link entities, especially when faced with abbreviations, synonyms, or implicit entities in the context. Furthermore, entities in dialogues are often ambiguous, meaning the same word may refer to different entities in different contexts. Traditional pre-trained models struggle to handle this ambiguity because they haven't learned sufficient contextual information during pre-training to distinguish the semantic and structural relationships between different entities. Therefore, researchers need to develop new models and techniques to improve the accuracy and robustness of entity recognition and linking in dialogue systems.
[0101] This disclosure provides a model training method that can more efficiently handle the dynamism and contextual dependencies in dialogue, truly utilizing dialogue contextual information. By combining pre-trained and fine-tuned models, the training process is adjusted to address the problem of diminished practicality of current pure text pre-trained models in task-oriented dialogue systems. Specifically, this proposal employs a Transformer-based pre-trained model and incorporates dialogue contextual information during the pre-training phase. In this way, the model can learn the interactivity and dynamism of the dialogue during the pre-training phase, thereby better handling task-oriented dialogue data during the fine-tuning phase.
[0102] Furthermore, the proposal in this disclosure also employs a contrastive learning objective function to fully learn the semantic and structural relationships between entities, further enhancing the model's ability to understand the dialogue context. By designing a unified framework, this proposal integrates referential information fully into the pre-training process from both semantic and structural perspectives. Specifically, the contrastive learning objective function enables the model to identify and distinguish subtle differences between different entities. Typically, entity identification and linking are key to understanding user intent and needs in dialogue. However, the ambiguity of entities and implicit entities in the context complicate this task. Through contrastive learning, the model can learn the semantic similarities and differences between entities, thereby more accurately identifying and linking entities in the dialogue. This proposal, through the contrastive learning objective function and the unified pre-training framework, significantly improves the model's ability to understand the dialogue context, providing strong support for building more intelligent task-oriented dialogue systems.
[0103] Furthermore, applying pre-trained language models to task-oriented dialogue systems leverages their powerful language understanding capabilities. However, while related research has largely achieved effective processing of general text data, it has failed to achieve a deep understanding of dialogue data. To address this issue, the proposal of this disclosure introduces contextual information from unlabeled dialogue sample data into the pre-training process through unsupervised learning, enabling the model to better capture information flow and contextual dependencies within the dialogue. In this way, the model can not only process general text data but also understand and predict information flow in dialogue, thereby improving the performance of task-oriented dialogue systems.
[0104] The proposed model (CECPT) in this disclosure embodiment may include two components: semantic pre-training and graph structure pre-training, as shown in Figure 3.
[0105] First, preprocessing is performed, including the use of the AllenNLP toolkit and modifications to the parsed reference pairs during pre-training.
[0106] To construct extensive and role-sensitive graph-learning dialogue knowledge from large-scale task-oriented datasets (dialogue sample data), CECPT uses an automatic reference resolution tool to parse connected multi-turn utterances into predefined graph structures. Considering the potential relationships between words mentioned in the dialogue history (also known as references), pre-trained reference resolution relations will naturally benefit the modeling of reference enhancement relations in the dialogue text.
[0107] Secondly, dialogue semantic pre-training:
[0108] For example, a dialogue sample data d consists of a series of constituent utterances (dialogue content), denoted as (u1, u2, ..., uN). Each utterance can be composed of multiple tokens. A multi-layer Transformer architecture is adopted, and the token representation is extracted from the hidden vectors of the final layer. In addition, to distinguish the roles of the speakers in the dialogue, role tokens are added before each utterance, such as [USR] and [SYS]. Furthermore, a special [CLS] token is added at the beginning of the dialogue.
[0109] As shown in Figure 4, the semantic contrast model (language model) takes the current dialogue utterance (e.g., OK. The table will be reserved for 15 minutes.), the previous dialogue utterance (e.g., I want to take a train to city X on Monday, when's available?...book a table for 5 people at 21:00 on the same day, please. I have several possibilities for Monday.), and the reference parsing results (i.e., corresponding first-class word pairs and second-class word pairs) as input. It uses a pre-trained language model (PLM) as the text encoder and trains it to distinguish various reference pairs (i.e., corresponding word pairs) mentioned in the dialogue.
[0110] A referential discriminator can be used as a pre-training task for semantic contrast, aiming to promote the representational proximity of words or mentions sharing semantic referential relationships and to contrast them with pairs having different semantic referential relationships. Using these referential pairs as positive samples (corresponding to first-class word pairs), the text encoder is trained to distinguish them from referential pairs exhibiting different semantic referential connections (as negative samples, corresponding to second-class word pairs). Therefore, this method enables the text encoder to capture dialogical referential semantics without requiring manual annotation.
[0111] For example, a word pair (c, c') is a positive example pair (positive referential pair, a word pair in the first type of word pair), they belong to the same entity and share semantic referential relations, such as 'Monday' and 'the same day' in Figure 4; otherwise, it is a negative example. Specifically, through random sampling, a mention of 'c' in the original positive example pair is replaced with a word from the tag... The negative reference to another (e.g., 'a train') is randomly sampled for c and replaced with m. c Next, we get m c The number of negative pairs, for c', are randomly sampled and replaced by m. c ′ times, to obtain m c The number of negative pairs. To distinguish between positive and negative pronoun pairs, a training objective is established for the positive pronoun pair (c,c'), using a contrastive loss (i.e., the value of the semantic contrastive loss function corresponding to the positive pronoun pair (c,c') to correctly classify the positive pronoun pair:
[0112] Where x represents the semantic vector of a word / mention (e.g., word c), and W is the number of matrix parameters for learning word pair similarity (i.e., the matrix parameters in the initial language model). L c,c′ x represents the value of the semantic contrast loss function corresponding to the word pair (positive reference pair) including word c and word c'. c′ The semantic vector representing word c'. This represents the transpose of the semantic vector of word c. This indicates the word that is replaced for the i-th time in word c. The transpose of the semantic vector. It is the word that is replaced for the jth time with word c'. The semantic vector.
[0113] Therefore, in order to obtain the overall training objective of semantic pre-training, the set P of word pairs corresponding to all dialogue sample data d in batch Bd (multiple dialogue sample data) is... d The summation of the losses of all positive referential pairs in the first-class word pairs corresponding to the dialogue sample data d yields the result L of the target loss function:
[0114] Then, the dialogue structure is pre-trained:
[0115] Because referential entities can be interconnected, graphical representations can effectively capture structured information, thus facilitating the modeling of interconnected references in multi-turn dialogues. In this training process, we first introduce the construction of the dialogue graph, and then explain how to use a graph contrast mechanism to learn the dialogue structure.
[0116] First, the dialogue graph is constructed:
[0117] Each type of node or edge is used to encode a certain type of information or information flow in the dialogue, as shown in the two diagrams on the left of Figure 5. For nodes, the embeddings of word and utterance nodes can be initialized using a text encoder, and the embedding of each mention / word node is obtained by averaging the embeddings of all word nodes for each mention / word. For edges, each word node in an utterance is connected to an utterance node (also called a node of a dialogue character), and is connected to its neighboring word nodes through the context represented by dashed lines, as shown in Figure 5. Each utterance node can also be connected to p (integers) of the most recent past utterances and f (integers) of the future utterances.
[0118] Secondly, the graph contrast learning mechanism:
[0119] As shown in Figure 5, the structural contrastive model (graph encoder) takes a predefined dialogue graph as input, initializes node embeddings using the aforementioned text encoder, and maximizes the mutual information between individual node representations and the overall graph representation through contrastive learning. To learn the dialogue structure without losing too much semantic information and to reduce the cost of manual annotation, the graph-based representation learning method Infograph can be used for unsupervised pre-training of the dialogue graph. Infograph uses mutual information (MI) as the core metric and learns the graph structure by maximizing the mutual information between node representations and graph representations through contrastive learning.
[0120] For example, consider a batch of dialogue graph samples containing N dialogue graphs (G1, G2, ... GN). Taking G1 as an example, where node v ∈ G1 (i.e., G1). The graph encoder is a K-layer graph isomorphism network (GIN). When graph G1 is input to the graph encoder, the k-th layer representation of node v can be obtained according to the following formula:
[0121] in, N represents the embedding of node v in layer k (i.e., the feature vector of node v in layer k), (v) This represents the neighboring nodes of node v. This represents the (k-1)th level representation of node μ. f represents the (k-1)th level representation of node v. (k) (·) represents the multilayer perceptron of the graph encoder at layer k, ∈ (k) This represents the combined parameters of the graph encoder at layer k.
[0122] By connecting the representations of all layers, the representation of node v can be obtained as follows:
[0123] Where h v g represents the representation of node v (the node feature vector of node v), and g(·) represents the join operation.
[0124] The representation of the dialogue graph (global feature representation) can be obtained by concatenating and pooling the representations of each node in the graph:
[0125] H G1Let g(·) be the global feature representation of the dialogue graph G1 (i.e., G1), and g(·) represent the join operation. It should be noted that the join operation in this embodiment can be averaging or dimensional join; for example, two 10-dimensional vectors joined together become a 20-dimensional vector. Here, sum represents the summation pooling operation. After obtaining the global feature representation, positive sample pairs (first sample pair) and negative sample pairs (second sample pair) are constructed and sent to the discriminator to continuously learn the graph representation. In a batch of samples, the global feature representation of the dialogue graph and the node representation of its own nodes are taken as positive sample pairs, while the global feature representation of the dialogue graph and the node representations of other nodes in the same batch are taken as negative sample pairs.
[0126] Taking the graphs GA and GB in Figure 5 as examples, the graph encoder can obtain the node representations of the GA (corresponding to the thin solid lines in Figure 5), the global representation of the GA (corresponding to the dashed lines in Figure 5), and the node representations of the GB (corresponding to the thick solid lines in Figure 5). Then, the global feature representation and node representation of the GA can form positive sample pairs, and the global representation of the GA and the node representation of the GB can form negative sample pairs. Subsequently, both positive and negative sample pairs are fed into the discriminator to calculate and maximize mutual information (MI), using a method similar to the Infograph algorithm, with a pre-training loss L. G (The results of the comparison loss function are calculated as follows:)
[0127] Where φ is the parameter of the discriminator D. These are the parameters of GIN. For the output of the discriminator and The mutual information between them, y′ is the negative sample corresponding to the input sample y, and N is the total number of graphs. It is the total number of nodes, v i This represents the number of nodes in graph i. This represents the global representation (i.e., the global feature representation) of graph i. This represents the local representation (i.e., the node feature representation) of node j in graph i. sp(·) is the softplus function (an activation function). This can represent the average mutual information corresponding to positive sample pairs. It can represent the average mutual information of negative sample pairs.
[0128] This disclosure proposes an innovative dialogue modeling framework that employs a contrastive pre-training strategy and explicitly integrates referentially guided information flow. This strategy enables more effective capture of semantic coherence and structural consistency in dialogue. Furthermore, by incorporating referential information into the pre-training process, it achieves modeling of both semantic and structural dimensions of dialogue within a unified framework. This approach not only enhances the model's understanding of dialogue context but also improves its ability to handle complex dialogue structures. The proposal implemented in this disclosure uses a graph contrastive training mechanism to explore the semantic and structural aspects of dialogue modeling. The scheme in this embodiment demonstrates significant performance improvements in multiple downstream tasks (e.g., intent recognition), providing new perspectives and possibilities for further research and application of dialogue systems.
[0129] As shown in Figure 6, which is a structural schematic diagram of a dialogue model training device provided in an embodiment of this disclosure, the dialogue model training device 600 includes:
[0130] The first acquisition module 601 is used to acquire dialogue sample data;
[0131] The first determining module 602 is used to determine a first type of word pair and a second type of word pair based on the dialogue sample data. In the first type of word pair, the two words in each word pair belong to the same entity and have the same semantic reference relationship. In the second type of word pair, the two words in each word pair do not belong to the same entity and have different semantic reference relationships.
[0132] Training module 603 is used to train the initial dialogue model based on dialogue sample data, first-class word pairs, and second-class word pairs through a contrastive learning strategy to obtain the target dialogue model.
[0133] In some embodiments, the initial dialogue model includes an initial language model, and the target dialogue model includes a target language model;
[0134] Training module 603 includes:
[0135] The semantic vector determination module is used to determine the semantic vectors of words in the dialogue sample data using the initial language model;
[0136] The first loss determination module is used to determine the result of the target loss function of the contrastive learning strategy based on the semantic vectors of the two words in each word pair of the first type of word pair and the semantic vectors of the two words in each word pair of the second type of word pair.
[0137] The first model training module is used to train the initial language model based on the results of the target loss function, so as to obtain the target language model.
[0138] In some embodiments, there are multiple dialogue sample data sets, and the result of the target loss function is the sum of the values of the semantic contrast loss functions of the dialogue sample data sets. The value of the semantic contrast loss function of each dialogue sample data set is obtained as follows:
[0139] It is calculated using the semantic vectors of the two words in each word pair of the first type of word pair in the dialogue sample data and the semantic vectors of the two words in each word pair of the second type of word pair in the dialogue sample data.
[0140] In some embodiments, the initial dialogue model further includes an initial graph encoder, and the target dialogue model further includes a target graph encoder; the dialogue sample data is multiple;
[0141] Training module 603 also includes:
[0142] The construction module is used to construct a dialogue graph for each dialogue sample data in multiple dialogue sample data sets, based on the semantic vectors of words in the dialogue sample data, the positional relationships of words in the dialogue sample data, the dialogue roles included in the dialogue sample data, and the semantic reference relationships between words in the dialogue sample data. The dialogue graph includes N nodes and the connection relationships between N nodes. Nodes are used to represent the semantic vectors of words or dialogue roles. The number of nodes corresponding to the same dialogue role in the dialogue graph is the same as the number of dialogues of the dialogue role in the dialogue sample data. Different nodes of the same dialogue role correspond to different dialogue content. The dialogue graph includes connections between nodes with the same semantic reference relationship, connections between nodes of the same dialogue role, connections between nodes of different dialogue roles with dialogue, connections between nodes of each dialogue role and nodes corresponding to each word in the dialogue content of the node, and connections between nodes corresponding to adjacent words in the dialogue sample data.
[0143] The graph feature vector determination module is used to determine the global feature representation of each dialogue graph and the node feature vector of each node in each dialogue graph through the initial graph encoder.
[0144] The second training module is used to train an initial graph encoder based on at least one first sample pair and at least one second sample pair to obtain a target graph encoder. The first sample pair includes the global feature representation of the first dialogue graph and the node feature representation of each node in the first dialogue graph. The first dialogue graph is any one of the multiple dialogue graphs. The second sample pair includes the global feature representation of the first dialogue graph and the node feature representation of each node in the second dialogue graph. The second dialogue graph is any one of the multiple dialogue graphs other than the first dialogue graph.
[0145] In some embodiments, the second training module includes:
[0146] An input module is used to input at least one first sample pair and at least one second sample pair into the discriminator for discrimination, respectively.
[0147] The second loss determination module is used to determine the result of the graph contrast loss function based on the mutual information output by the discriminator.
[0148] The graph encoder training module is used to train the initial graph encoder based on the results of the graph contrastive loss function, so as to obtain the target graph encoder.
[0149] The device 600 provided in this embodiment can implement each process of the above-described embodiments of the dialogue model training method applied to the first electronic device. The technical features are one-to-one and can achieve the same technical effect. To avoid repetition, it will not be described again here.
[0150] Referring to Figure 7, which is a schematic diagram of the structure of a dialogue device provided in an embodiment of this disclosure, as shown in Figure 7, the dialogue device 700 includes:
[0151] The second acquisition module 701 is used to acquire dialogue input data;
[0152] Feature determination module 702 is used to determine the feature vector of dialogue input data through the target dialogue model;
[0153] Output module 703 is used to output the operation result of the dialogue input data based on the feature vector of the dialogue input data;
[0154] The target dialogue model is the target dialogue model trained using the dialogue model training method described above.
[0155] The device 700 provided in this embodiment can implement the various processes of the above-described embodiments applied to the dialogue method. The technical features are one-to-one and can achieve the same technical effect. To avoid repetition, it will not be described again here.
[0156] This disclosure also provides an electronic device, including: a processor, a memory, and a program stored in the memory and executable on the processor. When the program is executed by the processor, it implements the various processes of the above-described dialogue model training method embodiment applied to the first electronic device and achieves the same technical effect. To avoid repetition, it will not be described again here.
[0157] Specifically, referring to Figure 8, this disclosure also provides an electronic device, including a bus 801, a transceiver 802, an antenna 803, a bus interface 804, a processor 805, and a memory 806.
[0158] The processor 805 is used for:
[0159] Obtain dialogue sample data;
[0160] Based on the dialogue sample data, we identified two types of word pairs: the first type and the second type. In the first type, the two words in each word pair belong to the same entity and have the same semantic reference relationship. In the second type, the two words in each word pair do not belong to the same entity and have different semantic reference relationships.
[0161] Based on dialogue sample data, first-class word pairs, and second-class word pairs, the initial dialogue model is trained using a contrastive learning strategy to obtain the target dialogue model.
[0162] In some embodiments, the initial dialogue model includes an initial language model, and the target dialogue model includes a target language model;
[0163] The 805 processor is specifically used for:
[0164] The semantic vectors of words in the dialogue sample data are determined using an initial language model.
[0165] Based on the semantic vectors of the two words in each word pair of the first type of word pair and the semantic vectors of the two words in each word pair of the second type of word pair, the result of the target loss function of the contrastive learning strategy is determined.
[0166] Based on the results of the target loss function, the initial language model is trained to obtain the target language model.
[0167] In some embodiments, there are multiple dialogue sample data sets, and the result of the target loss function is the sum of the values of the semantic contrast loss functions of the dialogue sample data sets. The value of the semantic contrast loss function of each dialogue sample data set is obtained as follows:
[0168] It is calculated using the semantic vectors of the two words in each word pair of the first type of word pair in the dialogue sample data and the semantic vectors of the two words in each word pair of the second type of word pair in the dialogue sample data.
[0169] In some embodiments, the initial dialogue model further includes an initial graph encoder, and the target dialogue model further includes a target graph encoder; the dialogue sample data is multiple;
[0170] The 805 processor is specifically used for:
[0171] For each dialogue sample in multiple dialogue sample data sets, a dialogue graph is constructed based on the semantic vectors of words in the dialogue sample data, the positional relationships of words in the dialogue sample data, the dialogue roles included in the dialogue sample data, and the semantic reference relationships between words in the dialogue sample data. The dialogue graph includes N nodes and the connections between N nodes. Nodes are used to represent the semantic vectors of words or dialogue roles. The number of nodes corresponding to the same dialogue role in the dialogue graph is the same as the number of dialogues of the dialogue role in the dialogue sample data. Different nodes of the same dialogue role correspond to different dialogue content. Connections are made between nodes with the same semantic reference relationship, between nodes with the same dialogue role, between nodes of different dialogue roles with dialogue, between nodes of each dialogue role and the nodes corresponding to each word in the dialogue content of the node, and between nodes corresponding to adjacent words in the dialogue sample data.
[0172] The initial graph encoder determines the global feature representation of each dialogue graph and the node feature vector of each node in each dialogue graph.
[0173] An initial graph encoder is trained based on at least one first sample pair and at least one second sample pair to obtain a target graph encoder. The first sample pair includes the global feature representation of the first dialogue graph and the node feature representation of each node in the first dialogue graph. The first dialogue graph is any one of the multiple dialogue graphs. The second sample pair includes the global feature representation of the first dialogue graph and the node feature representation of each node in the second dialogue graph. The second dialogue graph is any one of the multiple dialogue graphs other than the first dialogue graph.
[0174] In some embodiments, the processor 805 is specifically used for:
[0175] At least one first sample pair and at least one second sample pair are respectively input into the discriminator for discrimination;
[0176] Based on the mutual information output by the discriminator, determine the result of the graph contrast loss function;
[0177] Based on the results of the graph contrast loss function, the initial graph encoder is trained to obtain the target graph encoder.
[0178] In Figure 8, a bus architecture (represented by bus 801) is shown. Bus 801 may include any number of interconnected buses and bridges, linking various circuits including one or more processors represented by processor 805 and memory represented by memory 806. Bus 801 may also link various other circuits such as peripheral devices, voltage regulators, and power management circuits, which are well known in the art and therefore will not be described further herein. Bus interface 804 provides an interface between bus 801 and transceiver 802. Transceiver 802 may be a single element or multiple elements, such as multiple receivers and transmitters, providing a unit for communicating with various other devices over a transmission medium. Data processed by processor 805 is transmitted over a wireless medium via antenna 803, which further receives data and transmits it to processor 805.
[0179] The processor 805 manages the bus 801 and handles general processing, and also provides various functions, including timing, peripheral interfaces, voltage regulation, power management, and other control functions. The memory 806 can be used to store data used by the processor 805 during operation.
[0180] Optionally, the processor 805 may be a central processing unit (CPU), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or a complex programmable logic device (CPLD).
[0181] The processor in the electronic device 800 provided in this embodiment can implement the various processes of the above-described embodiments applied to the training method. The technical features correspond one-to-one and can achieve the same technical effect. To avoid repetition, it will not be described again here.
[0182] This disclosure also provides a computer-readable storage medium storing a computer program. When executed by a processor, the computer program implements the various processes described in the embodiments of the dialogue model training method applied to the first electronic device, and achieves the same technical effects. To avoid repetition, it will not be described again here. The computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0183] This disclosure also provides an electronic device, including: a processor, a memory, and a program stored in the memory and executable on the processor. When the program is executed by the processor, it implements the various processes described above for the dialogue method embodiments and achieves the same technical effects. To avoid repetition, it will not be described again here.
[0184] Specifically, as shown in Figure 9, this disclosure also provides an electronic device, including a bus 901, a transceiver 902, an antenna 903, a bus interface 904, a processor 905, and a memory 906.
[0185] The processor 905 is used for:
[0186] Obtain dialogue input data;
[0187] Determine the feature vector of the dialogue input data using the target dialogue model;
[0188] Based on the feature vector of the dialogue input data, output the operation result of the dialogue input data;
[0189] The target dialogue model is the target dialogue model trained using the dialogue model training methods described in the above embodiments.
[0190] In Figure 9, a bus architecture (represented by bus 901) is shown. Bus 901 can include any number of interconnected buses and bridges, linking various circuits including one or more processors represented by processor 905 and memory represented by memory 906. Bus 901 can also link various other circuits such as peripheral devices, voltage regulators, and power management circuits, which are well known in the art and therefore will not be described further herein. Bus interface 904 provides an interface between bus 901 and transceiver 902. Transceiver 902 can be a single element or multiple elements, such as multiple receivers and transmitters, providing a unit for communicating with various other devices over a transmission medium. Data processed by processor 905 is transmitted over a wireless medium via antenna 903, which further receives data and transmits it to processor 905.
[0191] Processor 905 manages bus 901 and general processing, and also provides various functions, including timing, peripheral interface, voltage regulation, power management, and other control functions. Memory 906 can be used to store data used by processor 905 during operation.
[0192] Optionally, the processor 905 can be a CPU, ASIC, FPGA, or CPLD.
[0193] The processor in the electronic device 900 provided in this embodiment can implement the various processes of the above-described embodiments applied to the dialogue method. The technical features correspond one-to-one and can achieve the same technical effect. To avoid repetition, it will not be described again here.
[0194] This disclosure also provides a computer-readable storage medium storing a computer program. When executed by a processor, the computer program implements the various processes described above in the dialogue method embodiments and achieves the same technical effects. To avoid repetition, further details are omitted here. The computer-readable storage medium may be, for example, ROM, RAM, magnetic disk, or optical disk.
[0195] This disclosure also provides a computer program product, including a computer program that, when executed by a processor, implements the various processes of the above method embodiments and achieves the same technical effects. To avoid repetition, these will not be described again here.
[0196] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.
[0197] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this disclosure, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods of the various embodiments of this disclosure.
[0198] The embodiments of this disclosure have been described above with reference to the accompanying drawings. However, this disclosure is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this disclosure without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this disclosure.
Claims
1. A dialogue model training method, the method comprising: Obtain dialogue sample data; Based on the dialogue sample data, a first type of word pair and a second type of word pair are determined. In the first type of word pair, the two words in each word pair belong to the same entity and have the same semantic reference relationship. In the second type of word pair, the two words in each word pair do not belong to the same entity and have different semantic reference relationships. Based on the dialogue sample data, the first type of word pairs, and the second type of word pairs, the initial dialogue model is trained using a contrastive learning strategy to obtain the target dialogue model.
2. The method according to claim 1, wherein, The initial dialogue model includes an initial language model, and the target dialogue model includes a target language model; The step of training an initial dialogue model using a contrastive learning strategy based on the dialogue sample data, the first type of word pairs, and the second type of word pairs to obtain a target dialogue model includes: Using the initial language model, the semantic vectors of words in the dialogue sample data are determined; Based on the semantic vectors of the two words in each word pair of the first type of word pair and the semantic vectors of the two words in each word pair of the second type of word pair, the result of the target loss function of the contrastive learning strategy is determined; Based on the result of the target loss function, the initial language model is trained to obtain the target language model.
3. The method according to claim 2, wherein, The dialogue sample data consists of multiple data sets. The result of the target loss function is the sum of the values of the semantic contrast loss functions of the dialogue sample data sets. The value of the semantic contrast loss function of each dialogue sample data set is obtained in the following way: The semantic vectors of the two words in each word pair of the first type of word pair in the dialogue sample data and the semantic vectors of the two words in each word pair of the second type of word pair in the dialogue sample data are calculated.
4. The method according to claim 2, wherein, The initial dialogue model further includes an initial graph encoder, and the target dialogue model further includes a target graph encoder; the dialogue sample data consists of multiple data sets. The step of training an initial dialogue model using a contrastive learning strategy based on the dialogue sample data, the first type of word pairs, and the second type of word pairs to obtain a target dialogue model includes: For each dialogue sample in multiple dialogue sample data sets, a dialogue graph is constructed based on the semantic vectors of words in the dialogue sample data, the positional relationships of words in the dialogue sample data, the dialogue roles included in the dialogue sample data, and the semantic reference relationships between words in the dialogue sample data. The dialogue graph includes N nodes and the connection relationships between the N nodes. The nodes are used to represent the semantic vectors of words or dialogue roles. The number of nodes corresponding to the same dialogue role in the dialogue graph is the same as the number of dialogues of the dialogue role in the dialogue sample data. Different nodes of the same dialogue role correspond to different dialogue content. The dialogue graph includes connections between nodes with the same semantic reference relationship, connections between nodes of the same dialogue role, connections between nodes of different dialogue roles with dialogue, connections between nodes of each dialogue role and nodes corresponding to each word in the dialogue content of the node, and connections between nodes corresponding to adjacent words in the dialogue sample data. The initial graph encoder determines the global feature representation of each dialogue graph and the node feature vector of each node in each dialogue graph. The initial graph encoder is trained based on at least one first sample pair and at least one second sample pair to obtain the target graph encoder. The first sample pair includes a global feature representation of a first dialogue graph and a node feature representation of each node in the first dialogue graph. The first dialogue graph is any one of a plurality of dialogue graphs. The second sample pair includes a global feature representation of the first dialogue graph and a node feature representation of each node in the second dialogue graph. The second dialogue graph is any one of the plurality of dialogue graphs other than the first dialogue graph.
5. The method according to claim 4, wherein, The initial graph encoder training, based on at least one first sample pair and at least one second sample pair, to obtain the target graph encoder includes: The at least one first sample pair and the at least one second sample pair are respectively input into the discriminator for identification; Based on the mutual information output by the discriminator, the result of the graph contrast loss function is determined; Based on the results of the graph contrast loss function, the initial graph encoder is trained to obtain the target graph encoder.
6. A dialogue method, the method comprising: Obtain dialogue input data; The feature vector of the dialogue input data is determined by the target dialogue model; Based on the feature vector of the dialogue input data, the operation result of the dialogue input data is output; The target dialogue model is a target dialogue model trained by any one of the methods in claims 1 to 5.
7. A model training apparatus, the apparatus comprising: The first acquisition module is used to acquire dialogue sample data; The first determining module is used to determine a first type of word pair and a second type of word pair based on the dialogue sample data. The two words in each word pair of the first type of word pair belong to the same entity and have the same semantic reference relationship. The two words in each word pair of the second type of word pair do not belong to the same entity and have different semantic reference relationships. The training module is used to train the initial dialogue model based on the dialogue sample data, the first type of word pairs, and the second type of word pairs, using a contrastive learning strategy to obtain the target dialogue model.
8. An electronic device comprising a transceiver and a processor, The processor is used for: Obtain dialogue sample data; Based on the dialogue sample data, a first type of word pair and a second type of word pair are determined. In the first type of word pair, the two words in each word pair belong to the same entity and have the same semantic reference relationship. In the second type of word pair, the two words in each word pair do not belong to the same entity and have different semantic reference relationships. Based on the dialogue sample data, the first type of word pairs, and the second type of word pairs, the initial dialogue model is trained using a contrastive learning strategy to obtain the target dialogue model.
9. An electronic device, comprising: A processor, a memory, and a program stored in the memory and executable on the processor, wherein the program, when executed by the processor, implements the steps of the method as claimed in any one of claims 1 to 5, or implements the steps of the method as claimed in claim 6.
10. A computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of any one of the methods of claims 1-5, or implements the steps of the method of claim 6.
11. A computer program product comprising computer instructions that, when executed by a processor, implement the steps of the method as described in any one of claims 1-6.
Citation Information
Patent Citations
Anaphora resolution method and device, electronic equipment and readable storage medium
CN112989043A
Model training method and device, electronic equipment and storage medium
CN114638212A
Multi-task learning dialogue method supporting anaphora resolution and one-language multi-meaning
CN117149958A
Dialogue text clustering method and related equipment
CN117332086A
Dialogue model training method, dialogue method, dialogue device and electronic equipment
CN118917439A