Dialogue generation method and related device

By sharing key parameters and value parameters in the multi-head self-attention module of the Transformer model, the problems of high computational complexity and poor generalization ability of the existing model are solved, and more efficient model training and reasoning are achieved, and more accurate conversations are generated.

CN120179760APending Publication Date: 2025-06-20TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311733942.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-15
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

In the existing Transformer model, in the multi-head self-attention module, the key parameter matrix WK and value parameter matrix WV of each attention layer are different, resulting in high computational complexity of model training and inference, which is easy to overfit, and reduces the generalization ability of the model.

Method used

In the dialogue generation method, the key parameter matrix and value parameter matrix preset by each multi-head self-attention module are used to share these parameters to reduce the amount of model parameters, and the degree of correlation between each vocabulary and other vocabulary in the to-be-attention mechanism is used to calculate the degree of correlation between each vocabulary and other vocabulary in the to-be-treated dialogue, capturing the long-distance dependence in the sentence.

Benefits of technology

By sharing key parameters and value parameters, the calculation complexity and model size are reduced, the risk of overfitting is reduced, the generalization ability of the model is improved, and more accurate and smooth reply conversations are generated.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120179760A_ABST
    Figure CN120179760A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence, in particular to a dialogue generation method and related device.The method comprises the steps that firstly, a word segmentation matrix corresponding to a to-be-processed dialogue is input into a dialogue model trained based on a Transform network, the dialogue model comprises a plurality of multi-head self-attention modules, and in each module, a word segmentation matrix corresponding to the to-be-processed dialogue is input into a word segmentation matrix corresponding to the to-be-processed dialogue; and based on the corresponding shared key parameter matrix, performing feature extraction on the word meaning of the segmented word represented by the input matrix to obtain a corresponding key feature matrix, and based on the corresponding shared value parameter matrix, performing feature extraction on the word meaning of the segmented word represented by the input matrix to obtain a corresponding value feature matrix. Therefore, in each module, the respective attention layer shares the key features and the value features and shares the key parameters and the value parameters, so that the model parameter quantity is reduced, the calculation complexity is reduced, and the model training and reasoning speed is improved; in addition, by reducing the parameter quantity of the model, the over-fitting risk of the model is reduced, and the generalization ability of the model is enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence, and in particular, to a dialogue generation method and related devices. Background Art

[0002] With the development of artificial intelligence technology, the application scenarios of dialogue generation models are very broad in various fields. Whether it is online customer service, online learning, entertainment and virtual characters, medical and health management, or the business and financial fields, dialogue generation models can bring more efficient communication and intelligent services to all walks of life.

[0003] The Transformer model is a commonly used dialogue generation model. Refer to Figure 1 As shown, it is the structure diagram of the Transformer model. The Transformer model contains an encoder and a decoder, and multiple multi-head self-attention modules are used in both the encoder and the decoder to extract the correlation information between each token and the position information of each token in the input sequence. In the encoder, the input matrix of the multi-head self-attention module is the corresponding word matrix X, Y of the dialogue text, or the output of the previous multi-head self-attention module; in the decoder, the input matrix of the multi-head self-attention module may be the corresponding word matrix Y of the dialogue text, the output Z of the previous multi-head attention module, or the output C of the encoder. Each multi-head self-attention module contains multiple self-attention layers. For the input sequence X of a multi-head self-attention module, linear transformations are respectively performed on X based on the model parameters W Q 、W K 、W V to obtain the corresponding query feature (Query), key feature (Key), and value feature (Value):

[0004] Query = X × W Q ; Key = X × W K ; Value = X × W V

[0005] In the existing Transformer model, for a multi-head self-attention module, the model parameters W K of each attention layer are different, and different Keys can be obtained respectively; the model parameters W V of each attention layer are also different, and different Values can be obtained. Therefore, during the model training process, it is necessary to update the parameters W K 、W V in each attention layer respectively, and the computational complexity is relatively high, which affects the model training and inference speed; in addition, due to the large number of model parameters, the model is prone to overfitting during training, resulting in poor generalization ability of the model.

[0006] In view of this, there is a need to propose a new dialogue generation method to overcome the above defects. Summary of the Invention

[0007] The present application provides a dialogue generation method and related devices to improve the quality and practicality of dialogue generation.

[0008] In a first aspect, an embodiment of the present application provides a dialogue generation method, the method comprising:

[0009] Performing vector transformation on at least one word segment included in the dialogue to be processed to obtain a first word segment matrix;

[0010] Inputting the first word segment matrix into a dialogue model trained based on a transformed Transformer network, and obtaining a reply dialogue associated with the dialogue to be processed based on global features finally output by a plurality of multi-head self-attention modules in the dialogue model; wherein, each multi-head self-attention module performs the following operations:

[0011] Based on a preset key parameter matrix and value parameter matrix corresponding to the multi-head self-attention module, respectively performing feature extraction on the word segment meanings represented by the input matrix to obtain corresponding key feature matrices, and performing feature extraction on the word segment meanings represented by the input matrix to obtain corresponding value feature matrices;

[0012] Based on each preset query parameter matrix corresponding to the multi-head self-attention module, respectively performing feature extraction on the word segment meanings represented by the input matrix to obtain corresponding query feature matrices;

[0013] Respectively obtaining the correlations between each query feature matrix and the key feature matrix, and respectively combining the value feature matrices based on each correlation to obtain corresponding attention features;

[0014] Performing fusion processing on each attention feature to obtain a corresponding global feature.

[0015] In a second aspect, an embodiment of the present application further provides a dialogue generation device, the device comprising:

[0016] A vector transformation unit for performing vector transformation on at least one word segment included in the dialogue to be processed to obtain a first word segment matrix;

[0017] A reply generation unit for inputting the first word segment matrix into a dialogue model trained based on a transformed Transformer network, and obtaining a reply dialogue associated with the dialogue to be processed based on global features finally output by a plurality of multi-head self-attention modules in the dialogue model; wherein, the reply generation unit further includes:

[0018] The key-value feature acquisition subunit is configured to perform, for each multi-head attention module respectively: based on the preset key parameter matrix and value parameter matrix corresponding to the multi-head self-attention module, respectively extract features from the tokenized word meanings represented by the input matrix to obtain the corresponding key feature matrix, and extract features from the tokenized word meanings represented by the input matrix to obtain the corresponding value feature matrix;

[0019] The query feature acquisition subunit is configured to perform, for each multi-head attention module respectively: respectively based on each query parameter matrix preset corresponding to the multi-head self-attention module, extract features from the tokenized word meanings represented by the input matrix to obtain the corresponding query feature matrix;

[0020] The attention feature acquisition subunit is configured to perform, for each multi-head attention module respectively: respectively obtain the relevance between each query feature matrix and the key feature matrix, and respectively based on each relevance, combine with the value feature matrix to obtain the corresponding attention feature;

[0021] The feature fusion subunit is configured to perform, for each multi-head attention module respectively: perform fusion processing on each attention feature to obtain the corresponding global feature.

[0022] Optionally, based on the global features finally output by multiple multi-head self-attention modules in the dialogue model, to obtain the reply dialogue associated with the dialogue to be processed, the reply generation unit is further configured to:

[0023] Perform feature mapping processing on the global feature to obtain a candidate token set corresponding to each of the N token positions; wherein, the N token positions are set based on a preset dialogue grammar structure;

[0024] Perform multiple splicing combinations among the candidate token sets to obtain multiple candidate dialogues; wherein, each time of splicing combination, select one candidate token from each of the candidate token sets;

[0025] Respectively based on the token relevance between the multiple candidate dialogues and the dialogue to be processed, obtain the quality evaluation value of each of the multiple candidate dialogues;

[0026] Select the candidate dialogue whose quality evaluation value meets the preset reply condition as the reply dialogue.

[0027] Optionally, perform feature mapping processing on the global feature to obtain a candidate token set corresponding to each of the N token positions, the reply generation unit is further configured to:

[0028] Perform feature mapping processing on the global feature to obtain a mapping probability set of the preset token set at each of the N token positions; wherein, each preset token has a corresponding mapping probability at each token position;

[0029] For each of the N word segmentation positions, perform the following operations respectively:

[0030] In the mapping probability set corresponding to a word segmentation position, select at least one target probability that meets the preset candidate conditions;

[0031] Select the preset word segments corresponding to the at least one target probability from the preset word segmentation set to form the candidate word segmentation set.

[0032] Optionally, based on the word segmentation relevance between the multiple candidate conversations and the conversation to be processed respectively, obtain the quality evaluation values of the multiple candidate conversations respectively. The reply generation module is further configured to:

[0033] For each of the multiple candidate conversations, perform the following operations respectively:

[0034] Concatenate the conversation to be processed and a candidate conversation to obtain an evaluation conversation to be evaluated, and perform vector conversion on each word segment included in the evaluation conversation to be evaluated to obtain a corresponding second word segment matrix;

[0035] Use an evaluation model trained based on the Transformer network to extract features from the second word segment matrix to obtain corresponding attention features; the attention features represent the word segmentation relevance between the candidate conversation and the conversation to be processed;

[0036] Optionally, the device further includes a model training unit, and the evaluation model is obtained by performing model training based on a positive sample set and a negative sample set; each positive sample includes: a sample conversation to be evaluated, a positive sample reply, and a positive sample label. Then the model training unit is further configured to construct a negative sample set:

[0037] For each positive sample reply in each positive sample, perform sample reconstruction processing by using at least one of the methods of randomly replacing the reply, replacing keywords, randomly replacing clauses, rearranging word segments, repeating reply clauses, and adding interference information to generate corresponding negative sample replies;

[0038] Construct a negative sample set based on the negative sample labels, each sample conversation to be evaluated, and the corresponding negative sample replies.

[0039] Optionally, the model training unit is further configured to:

[0040] Obtain a preset training sample set, and each training sample includes: a sample conversation to be replied and a true sample reply;

[0041] Based on the training sample set, perform multi-round iterative training on the conversation model to be trained. Among them, in one round of iterative process, perform the following steps:

[0042] Input the to-be-replied sample dialogue into a conversion Transformer network, and obtain a predicted sample reply associated with the to-be-replied sample dialogue based on the sample global features finally output by multiple multi-head self-attention modules. Adjust the model parameters based on the loss function value between the predicted sample reply and the true sample reply. Among them, each multi-head self-attention module performs the following operations:

[0043] Based on the preset key parameter matrix and value parameter matrix corresponding to the multi-head self-attention module, obtain the corresponding sample key feature matrix and sample value feature matrix, and respectively obtain the corresponding sample query feature matrix based on each preset query parameter matrix corresponding to the multi-head self-attention module;

[0044] Based on the sample key feature matrix and the sample value feature matrix, combine the sample query feature matrices to obtain the corresponding sample attention features, and perform a fusion process on each sample attention feature to obtain the corresponding sample global features.

[0045] In a third aspect, an embodiment of the present application provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the method described in any item of the first aspect is implemented.

[0046] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the method described in any one of the first aspect are implemented.

[0047] In a fifth aspect, an embodiment of the present application provides a computer program product. When the computer program product is called by a computer, the computer is made to execute the method described in the first aspect.

[0048] An embodiment of the present application provides a dialogue generation method.

[0049] First, perform vector conversion on the word segments included in the dialogue to be processed to obtain a word segment matrix. Then, input the word segment matrix into a dialogue model trained based on a transformed Transformer network. The dialogue model includes multiple multi-head self-attention modules. Among them, in each module, there are corresponding key parameter matrices and value parameter matrices preset. Based on the corresponding key parameter matrices, feature extraction is performed on the word meanings of the word segments represented by the input matrix to obtain corresponding key feature matrices, and based on the corresponding value parameter matrices, feature extraction is performed on the word meanings of the word segments represented by the input matrix to obtain corresponding value feature matrices. In this way, in each module, the self-attention layer shares the key feature matrix and the value feature matrix, as well as the key parameter matrix and the value parameter matrix, reducing the number of model parameters, helping to reduce the computational complexity, and improving the speed of model training and inference; at the same time, it also reduces the occupation of video memory and memory, reducing the requirements for hardware resources when deploying the model; in addition, by reducing the number of model parameters, the overfitting risk of the model is reduced, and the generalization ability of the model is enhanced.

[0050] Secondly, in each module, while obtaining the key feature matrix and the value feature matrix, feature extraction is also performed on the word meanings of the word segments of the input matrix respectively based on each query parameter matrix preset for the corresponding module to obtain corresponding query feature matrices; then, the correlation degree between the key feature matrix and the value feature matrix is calculated, and based on each correlation degree, combined with the value feature matrix, corresponding attention features are obtained. In this way, the self-attention mechanism is used to calculate the association degree between each word in the dialogue to be processed and other words, so as to capture the long-distance dependence relationship in the sentence.

[0051] Finally, perform fusion processing on each attention feature to obtain corresponding global features. In this way, multiple attention features enable the model to simultaneously focus on information at different positions, helping the model to generate more accurate and fluent reply dialogues. Description of the Drawings

[0052] Figure 1 It is a schematic diagram of the Transformer model structure;

[0053] Figure 2 It is a schematic diagram of the application scenario in the embodiment of the present application;

[0054] Figure 3 It is a schematic flowchart of a dialogue model training method in the embodiment of the present application;

[0055] Figure 4 It is a flowchart of the method for obtaining the global feature of the sample during the training of the dialogue model in the embodiment of the present application;

[0056] Figure 5 It is a schematic logical diagram of a self-attention processing method in the embodiment of the present application;

[0057] Figure 6 Schematic diagram of an evaluation model structure in an embodiment of the present application;

[0058] Figure 7 Schematic diagram of a process flow of a dialogue generation method in an embodiment of the present application;

[0059] Figure 8 Schematic diagram of a process flow of a method for obtaining attention features during dialogue generation in an embodiment of the present application;

[0060] Figure 9 Schematic diagram of a process flow of a method for generating candidate dialogues in an embodiment of the present application;

[0061] Figure 10 Schematic diagram of a logic of a method for generating candidate dialogues in an embodiment of the present application;

[0062] Figure 11 Schematic diagram of a process flow of a method for evaluating the quality of candidate dialogues in an embodiment of the present application;

[0063] Figure 12 Schematic diagram of an overall process flow of a dialogue generation method in an embodiment of the present application;

[0064] Figure 13 Schematic diagram of the structure of a dialogue generation device in an embodiment of the present application;

[0065] Figure 14 Schematic diagram of the structure of an electronic device in an embodiment of the present application. Detailed implementation manners

[0066] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Apparently, the described embodiments are some, but not all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments recorded in this application document without creative efforts shall fall within the scope of protection of the technical solutions of the present application.

[0067] Terms such as "first" and "second" in the specification, claims, and the above-mentioned drawings of the present application are used to distinguish similar objects and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein.

[0068] In the embodiments of the present application, the term "module" or "unit" refers to a computer program with a predetermined function or a part of a computer program, which works together with other related parts to achieve a predetermined goal, and can be fully or partially implemented by using software, hardware (such as a processing circuit or a memory), or a combination thereof. Similarly, one processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of an overall module or unit that includes the function of that module or unit.

[0069] The following explains some terms in the embodiments of the present application to facilitate the understanding of those skilled in the art.

[0070] (1) Encoder: In deep learning, a model structure that converts raw data into a low-dimensional vector representation and can extract the key features of the raw data.

[0071] (2) Decoder: In deep learning, a model structure that converts a low-dimensional vector representation back to the raw data space.

[0072] (3) Self-attention mechanism: A network configuration that enables the model to focus on both the global information and the key information of the model input.

[0073] (4) Overfitting: The model can fit the training set data well, but the prediction effect on the test set is poor.

[0074] (5) Generalization ability: The prediction ability of a machine learning model for unknown data outside the learning set.

[0075] The following briefly introduces the design concept of the embodiments of the present application:

[0076] Dialogue generation is a basic task in natural language processing. By performing intention recognition, sentiment analysis, etc. on the context, an open-ended response dialogue is generated. The Transformer model is a commonly used dialogue generation model. The Transformer model includes an encoder and a decoder, and multiple multi-head self-attention modules are used in both the encoder and the decoder to extract the correlation information between each token and the position information of each token in the input sequence.

[0077] In the existing Transformer model, for a multi-head self-attention module, the key parameter matrix W of each attention layer K is different, and different key features can be obtained respectively; the value parameter matrix W of each attention layer V is also different, and different value features can be obtained. Therefore, during the model training process, the parameters W in each attention layer need to be separately K 、W VUpdating has a relatively high computational complexity, which affects the model training and inference speed. Additionally, due to the large number of model parameters, the model is prone to overfitting during training, resulting in poor generalization ability of the model.

[0078] In view of this, in the embodiments of the present application, a dialogue generation method and related device are proposed.

[0079] In the embodiments of the present application, first, the word segments included in the dialogue to be processed are vector-converted to obtain a word segment matrix. Then, the word segment matrix is input into a dialogue model trained based on a transformed Transformer network. The dialogue model includes multiple multi-head self-attention modules. In each module, the self-attention layers share key parameters and value parameters. Specifically, each module is preset with a corresponding key parameter matrix and value parameter matrix. In this way, in each module, based on the corresponding key parameter matrix, feature extraction is performed on the word meanings represented by the input matrix to obtain a corresponding key feature matrix, and based on the corresponding value parameter matrix, feature extraction is performed on the word meanings represented by the input matrix to obtain a corresponding value feature matrix. In this way, in each module, the self-attention layers share the key feature matrix and the value feature matrix, reducing the number of model parameters, helping to reduce the computational complexity, and improving the speed of model training and inference. At the same time, it also reduces the occupation of video memory and memory, and reduces the requirements for hardware resources when deploying the model. In addition, by reducing the number of model parameters, the overfitting risk of the model is reduced, and the generalization ability of the model is enhanced.

[0080] Secondly, in each module, while obtaining the key feature matrix and the value feature matrix, feature extraction is also performed on the word meanings of the word segments of the input matrix based on each query parameter matrix preset for the corresponding module to obtain a corresponding query feature matrix. Then, the correlation between the key feature matrix and the value feature matrix is calculated, and based on each correlation, combined with the value feature matrix, a corresponding attention feature is obtained. In this way, the self-attention mechanism is used to calculate the association degree between each word in the dialogue to be processed and other words, so as to capture the long-distance dependence relationship in the sentence. Finally, the attention features are fused to obtain a corresponding global feature. In this way, multiple attention features enable the model to simultaneously focus on information at different positions, helping the model to generate more accurate and fluent reply dialogues.

[0081] In addition, in the inference stage, first, the dialogue model is used to splice and combine between multiple candidate word segment sets to generate multiple candidate dialogues. In this way, the dialogue model can generate diverse and creative replies. Then, an evaluation model is used to evaluate the quality of the multiple candidate dialogues, and then select the candidate dialogue that meets the preset reply conditions as the final reply dialogue. The evaluation model evaluates the accuracy, relevance, and fluency of the reply. The evaluation model can filter out irrelevant or incorrect replies generated by the dialogue model, helping to generate high-quality replies and improve the performance of the dialogue system.

[0082] Finally, when training the evaluation model, various different methods such as reply random replacement, token random replacement, token rearrangement, repeated reply clauses, and adding interference information are used to construct negative samples. In this way, more abundant and diverse negative samples can be generated, thereby improving the discrimination ability for high-quality replies. The multiple negative sample construction methods help to encourage the evaluation model to learn more features about low-quality replies, capture errors and differences in different aspects, and make the evaluation model more robust. In this way, the evaluation model can still maintain good performance when facing different types of errors or noises.

[0083] After introducing the design concept of the embodiments of the present application, the main technologies involved in the embodiments of the present application will be introduced below.

[0084] Artificial Intelligence (AI) is a theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines, enabling the machines to have the functions of perception, reasoning, and decision-making.

[0085] Artificial intelligence technology is an interdisciplinary subject with a wide range of fields, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0086] Machine Learning (ML) is an interdisciplinary subject that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and rote learning.

[0087] The solution provided by the embodiments of this application mainly relates to technologies such as machine learning / deep learning under the field of artificial intelligence. Specifically, through the method provided by this application, the existing Transformer network is improved, and model training is performed based on the improved Transformer network to obtain a dialogue model, and then a reply dialogue associated with the dialogue to be processed is generated based on the dialogue model.

[0088] The preferred embodiments of this application will be described below with reference to the accompanying drawings of the specification. It should be understood that the preferred embodiments described herein are only used to illustrate and explain this application, and are not used to limit this application. And without conflict, the embodiments of this application and the features in the embodiments can be combined with each other.

[0089] As Figure 2 shown, it is a schematic diagram of the application scenario in the embodiments of this application. It is a schematic diagram of the application scenario in the embodiments of this application. This application scenario diagram includes two terminal devices 210 and a server 220. Communication can be carried out between the terminal device 210 and the server 220 through a communication network. The user can obtain the target reply dialogue through the terminal device 210. A client with functions such as a chatbot, virtual assistant, or other intelligent dialogue generation functions can be installed on the terminal device 210. The chatbot can understand the user's needs and provide real-time replies and solutions. The virtual assistant can help the user perform various tasks, such as sending messages, querying the weather, setting an alarm, etc. The client can be software, or a web page, a small program, etc. The embodiments of this application do not limit the specific type of the client, and the background server is the background server corresponding to the software, web page, small program, etc.

[0090] In an alternative embodiment, the communication network is a wired network or a wireless network. The terminal device 210 and the server 220 can be directly or indirectly connected through wired or wireless communication methods, and this application does not limit this here.

[0091] In the embodiments of the present application, the terminal device 210 is an electronic device used by a user. The electronic device may be a computer device such as a personal computer, a mobile phone, a tablet computer, a notebook, an e-book reader, a smart home, etc., which has a certain computing ability and runs instant messaging software and websites or social software and websites. Each terminal device 210 is connected to the server 220 through a wireless network. The server 220 may be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It may also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery network (Content Delivery Network, CDN), and big data and artificial intelligence platforms.

[0092] Among them, the dialogue model and the evaluation model can be deployed on the server 220 for training. A large number of training samples can be stored in the server 220 for training the dialogue model and the evaluation model. Optionally, after the dialogue model and the evaluation model are trained based on the training method in the embodiments of the present application, the trained dialogue model and evaluation model can be directly deployed on the server 220 or the terminal device 210. Generally, the dialogue model and the evaluation model are directly deployed on the server 220.

[0093] In a possible application scenario, the training samples in the present application can be stored using cloud storage technology. Cloud storage is a new concept extended and developed from the concept of cloud computing. A distributed cloud storage system (hereinafter referred to as a storage system) refers to a storage system that combines a large number of different types of storage devices (storage devices are also called storage nodes) in the network through cluster applications, grid technology, and distributed file systems, and works together through application software or application interfaces to jointly provide data storage and service access functions to the outside world.

[0094] In a possible application scenario, in order to facilitate reducing communication latency, the server 220 can be deployed in each region, or for load balancing, different servers 220 can respectively serve the regions corresponding to each terminal device 210. Multiple servers 220 achieve data sharing through a blockchain. The multiple servers 220 are equivalent to a data sharing system composed of multiple servers 220. For example, the terminal device 210 is located at location a and is communicatively connected to the server 220, and the terminal device 210 is located at location b and is communicatively connected to another server 120.

[0095] For each server 220 in the data sharing system, there is a node identifier corresponding to the server 220. Each server 220 in the data sharing system can store the node identifiers of other servers 220 in the data sharing system, so that subsequently, according to the node identifiers of other servers 220, the generated blocks can be broadcast to other servers 220 in the data sharing system. A node identifier list as shown in the following table can be maintained in each server 220, and the server 220 name and the node identifier are correspondingly stored in the node identifier list. Among them, the node identifier can be an Internet Protocol (IP) address for interconnection between networks and any other information that can be used to identify the node. Only the IP address is taken as an example for illustration in Table 1.

[0096] Table 1

[0097] Server Name Node Identifier Node 1 119.115.151.174 Node 2 118.116.189.145 … … Node N 119.124.189.258

[0098] Next, in combination with the application scenario described above, the dialogue generation method provided by the exemplary embodiment of the present application will be described with reference to the accompanying drawings. It should be noted that the above application scenario is only shown for the convenience of understanding the spirit and principle of the present application, and the embodiments of the present application are not limited in this regard.

[0099] Refer to Figure 3 shown, which is a flowchart of the training method of the dialogue model in the embodiment of the present application. Next, in combination with Figure 3 , the specific steps to be executed will be described in detail:

[0100] Step 31: Obtain a preset training sample set.

[0101] Among them, each training sample includes: a sample dialogue to be replied to and a true sample reply.

[0102] In the embodiment of the present application, first, a large amount of open-domain dialogue corpus data is collected. Taking social media platforms such as Douban, Weibo, and Baidu Tieba as data sources, a large amount of open-domain dialogue corpus is crawled by means of web crawlers, and then the collected original dialogue corpus data is preprocessed to improve the data quality. Among them, the preprocessing steps include:

[0103] (1) Remove special characters and tags: Clean the collected dialogue text data, remove irrelevant content such as URL links, HTML tags, JavaScript code, special symbols, etc., and retain only plain text. (2) Unify the encoding method: Unify the dialogue text data into a single character encoding (such as UTF-8) to avoid garbled problems during processing. (3) Remove low-quality dialogue text: Identify and delete dialogue text containing excessive errors, meaningless or repetitive content to ensure the quality of the training data. (4) Filter multi-person conversations: Filter out conversations involving three or more people, and only retain conversations between two people as training data.

[0104] The following shows the multi-turn dialogue data between Speaker A and Speaker B.

[0105] Speaker A: Are you tired from playing games?

[0106] Speaker B: Yes, but what can I do?

[0107] Speaker A: You can stop playing.

[0108] Speaker B: If I don't play, I'll have nothing to do.

[0109] After preprocessing the data, convert the dialogue data into a form acceptable to the dialogue model, that is, a form combined with the dialogue to be replied and the real sample reply. First, add a special symbol "bos" at the beginning of the dialogue history to represent the start of the sentence, and splice the multi-turn dialogue with a special symbol "sep", and add a special symbol "eos" at the end of the dialogue history to represent the end of the dialogue history. Similarly, add a special symbol "bos" at the beginning of the reply to represent the start of the reply, and add a special symbol "eos" at the end of the reply to represent the end of the reply. In the embodiments of the present application, each training sample of the dialogue model is a pair of sentence pairs containing the dialogue to be replied and the real sample reply, as follows, X represents the dialogue to be replied, and Y represents the real sample reply:

[0110] X = [bos] Are you tired from playing games? [sep] Yes, but what can I do? [sep] You can stop playing. [eos]

[0111] Y = [bos] If I don't play, I'll have nothing to do. [eos]

[0112] After preprocessing the data, perform word segmentation and indexing on the dialogue text. For example, the results after word segmentation are as follows:

[0113] X = ['bos', 'Are you', 'tired', 'from', 'playing', 'games', '?', 'Yes', ',', 'but', 'what', 'can', 'you', 'do', '?', 'sep', 'You', 'can', 'stop', 'playing', '.', 'eos']

[0114] Y = ['bos', 'If', 'you', 'don', 't', 'play', ',', 'you', 'll', 'have', 'nothing', 'to', 'do', '.', 'eos']

[0115] After word segmentation, it is determined that the inputs to the encoder and decoder in the dialogue model are X and Y1 respectively, and the target output of the decoder is Y2:

[0116] Y1 = ['bos', 'If', 'you', 'don', 't', 'play', ',', 'you', 'll', 'have', 'nothing', 'to', 'do', '.']

[0117] Y2 = ['If', 'you', 'don', 't', 'play', ',', 'you', 'll', 'have', 'nothing', 'to', 'do', '.', 'eos']

[0118] It should be noted that the network is used to extract local features of the input sequence and usually includes two fully connected layers and an activation function.

[0119] The dialogue model adopts a bidirectional attention mechanism for the input text, where each word in the input text can attend to all other words; for the output text, it adopts a unidirectional attention mechanism, i.e., left-to-right attention, where each word in the output text can only attend to the previous words and not the subsequent words. In this way, encoding the input text using a bidirectional language model helps the dialogue model better understand the semantic information of the input text.

[0120] Next, the word sequence needs to be indexed. Before indexing, a vocabulary needs to be pre-constructed. For example, if the vocabulary size is 20000 (there are 20000 words in the vocabulary), the index of 'don't' in the vocabulary is 35, and 'don't' is indexed as '35'. The index of 'play' in the vocabulary is 3576, and 'play' is indexed as '3576'. In this way, a word sequence can be converted into an index sequence (a sequence of numbers). For example, the result of indexing Y1 is:

[0121] Y1 = [2, 35, 3576, 5046, 1257, 15539, 12876, 14140, 3465]

[0122] After the word sequence is indexed, it is also necessary to vectorize the word sequence. The word sequence is input into the Embedding layer at the beginning of the Transformer model, and for each word in the word sequence, the corresponding N-dimensional word vector is output. For example, Y1 contains 9 words. After vectorizing Y1, a 9×N-dimensional vector matrix is generated, where N can be set by those skilled in the art according to the actual situation, and the embodiments of the present application do not limit this.

[0123] In order to enable the Transformer model to learn the position information of words in a sentence, it is necessary to add position encoding to the input data. The dimension of the position encoding is the same as that of the word vector, both are N-dimensional. The position encoding method is as follows: for each word in a word sequence, encode according to the following formula:

[0124]

[0125]

[0126] The sine function is used for even positions, and the cosine function is used for odd positions to calculate the encoding value. pos represents the word sequence index of the corresponding word. For example, Y1 contains 9 words, and the word sequence index is 0 to 8; the variable i in the trigonometric function on the right side of the equation is i = dim_index / / 2, where dim_index is the word vector index. For example, when a N-dimensional word vector is generated, the word vector index is 0 to N - 1; d model represents the dimension of the word vector.

[0127] Step 32: Input the sample dialogue to be replied into the conversion Transformer network to obtain the sample global features finally output by multiple multi-head self-attention modules.

[0128] In the embodiments of the present application, after adding the vectorized sample dialogue X to be replied and the corresponding position encoding, it is input into the encoder of the conversion Transformer network. After adding the vectorized Y1 and the corresponding position encoding, it is input into the decoder of the Transformer network. After feature extraction by multiple multi-head self-attention modules, the sample global features output by the last multi-head self-attention module in the decoder are obtained.

[0129] Among them, referring to Figure 4 as shown, each attention module performs the following steps:

[0130] Step 321: Based on the preset key parameter matrix and value parameter matrix of the corresponding multi-head self-attention module, obtain the corresponding sample key feature matrix and sample value feature matrix, and respectively based on each query parameter matrix preset by the corresponding multi-head self-attention module, obtain the corresponding sample query feature matrix.

[0131] Refer to Figure 5 As shown, it is a schematic diagram of the structure of the multi-head self-attention module in the embodiment of the present application. A key parameter matrix and a value parameter matrix are respectively set for each multi-head self-attention module. That is, in one multi-head self-attention module, the self-attention layers share the key parameter matrix W K and the value parameter matrix W V , and the query parameter matrix is respectively set for each self-attention layer In each multi-head self-attention module, based on W K , a linear transformation is performed on the input matrix to generate a sample key feature matrix K; based on W V , a linear transformation is performed on the input matrix to generate a sample value feature matrix V; respectively based on each a linear transformation is performed to generate the corresponding sample query feature matrix Q i , that is:

[0132] K = X × W K ; Value = X × W V ;

[0133] Step 322: Based on the sample key feature matrix and the sample value feature matrix, combined with each sample query feature matrix, obtain the corresponding sample attention features.

[0134] Refer to Figure 5 As shown, in the embodiment of the present application, self-attention processing is performed on each sample query feature, sample key feature, and sample value feature. First, the correlation between each sample query feature matrix and the sample key feature matrix is calculated respectively, and based on each correlation respectively, combined with the value feature matrix, the corresponding attention features are obtained. Specifically, each attention feature Z1, Z2, Z3... Zn is calculated according to the following formula respectively:

[0135]

[0136] where d k represents the dimension of the word vector, that is, the number of columns of the Q and K matrices.

[0137] Step 323: Perform a fusion process on each sample attention feature to obtain the corresponding sample global feature.

[0138] Refer to Figure 5 As shown, in the embodiment of the present application, the way to perform a fusion process on each sample attention feature is to perform a splicing and fully connected operation on each sample attention feature, and then obtain the corresponding sample global feature.

[0139] Step 33: Based on the sample global feature, obtain the predicted sample reply associated with the sample dialogue to be replied.

[0140] In the embodiments of the present application, after the global features of the sample are input into the forward feedback network and normalized, a corresponding probability matrix is obtained, and then a corresponding predicted sample response is obtained based on the probability matrix.

[0141] Step 34: Based on a preset loss function, obtain the loss value between the predicted sample response and the true sample response, and based on the loss value, adjust the model parameters.

[0142] In the embodiments of the present application, the negative log loss function is used to calculate the loss value between the predicted sample response and the true sample response, and a common gradient descent algorithm, such as the stochastic gradient descent algorithm, is used to update the model parameters.

[0143] After completing the training of the dialogue model according to the above steps, a dialogue is generated based on the trained dialogue model. However, the dialogue generated by the dialogue model may have the following problems:

[0144] (1) The problem that the response repeats the previous dialogue: The generated response may directly copy or reuse some content in the previous dialogue without providing new information or viewpoints. For example:

[0145] User: I feel a bit tired today.

[0146] Dialogue model: You feel a bit tired today.

[0147] (2) The problem of intra-sentence repetition in the response: The generated response may contain repeated words or phrases, resulting in redundant expressions. For example:

[0148] User: What fruits do you like to eat?

[0149] Dialogue model: I like to eat apples. Apples are my favorite fruits. I like apples.

[0150] (3) The problem that the response is irrelevant to the previous dialogue: The generated response may be irrelevant to the theme or content of the previous dialogue, resulting in a lack of pertinence in the response. For example:

[0151] User: Do you know how to cook pasta?

[0152] Dialogue model: I like to take walks in the park and enjoy the beautiful scenery.

[0153] (4) The problem that the response contradicts the previous dialogue: The generated response may contradict the information or viewpoints in the previous text, resulting in a lack of consistency in the response. For example:

[0154] User: Do you think cats are cuter than dogs?

[0155] Dialogue model: Yes, I think cats are cute. But I think dogs are cuter.

[0156] (5) Problem of resolving intra-sentence contradictions: The generated responses may contain contradictory or inconsistent information, resulting in a lack of consistency in the responses. For example,

[0157] User: What's the weather like today?

[0158] Dialogue model: The weather is nice today, sunny. However, it's raining heavily now.

[0159] (6) Problem of generic responses: The generation model may tend to generate common and relatively safe responses such as "I don't know", "This is interesting", etc., resulting in a lack of diversity and creativity in the generated responses. For example:

[0160] User: What's the weather like today?

[0161] Dialogue model: I don't know.

[0162] Therefore, the embodiments of the present application also provide an evaluation model for evaluating the responses generated by the dialogue model, so as to select the preferred response as the final response. To solve the problems existing in the dialogue model when generating conversations, when training the evaluation model, a variety of methods are used for negative sample construction and mining. A negative sample refers to a response that is inappropriate for a given dialogue history (which can be single-round or multi-round). Taking a normal dialogue context and the corresponding response as a positive sample, the embodiments of the present application propose a variety of different methods for constructing and mining negative samples. For example, the following dialogue is a positive sample:

[0163] Speaker A: Do you like video games?

[0164] Speaker B (positive sample): I'm a PS4 player. However, I mainly like sports and cycling.

[0165] Then the following method can be used for negative sample reconstruction:

[0166] (1) Randomly replace the response: For a multi-round dialogue sample data, randomly select a response from the corpus that does not match the given context and replace the original positive sample response as the negative sample.

[0167] Negative sample: My family works in the medical industry.

[0168] (2) Randomly replace response clauses: Replace some clauses in the positive sample response, and retain a part of the content in the positive sample response to increase the learning difficulty of the negative sample and help the model learn more effective discrimination features.

[0169] Negative sample: I am a PS4 player. However, my family members work in the medical industry.

[0170] (3) Repeated clauses: To address the problem of intra-sentence repetition in the responses generated by the dialogue model, repeated clauses are used to construct negative samples that contain repeated words or phrases and express redundant information.

[0171] Negative sample: I am a PS4 player. However, I mainly like sports and cycling, like sports and cycling.

[0172] (4) Keyword replacement: Replace the keywords in the positive sample response with other relevant or irrelevant words. For example, replace the keyword "bicycle" in the positive sample with "car".

[0173] Negative sample: I am a PS4 player. However, I mainly like sports and driving a car.

[0174] (5) Shuffled order: By changing the order of words or phrases in the positive sample response, the semantics of the response become ambiguous or incoherent. For example, shuffle the order of multiple clauses or phrases in the positive sample.

[0175] Negative sample: Like sports and cycling, I mainly still, however, I am a PS4 player.

[0176] (6) Adding interfering information: Insert some words or phrases unrelated to the context into the positive sample response to make the response incoherent. For example, insert the irrelevant interfering information "eating an apple" into the positive sample response.

[0177] Negative sample: I am a PS4 player. However, I mainly like eating an apple, sports and cycling.

[0178] Combining the above multiple methods to generate richer negative samples can ensure that the differences between negative samples and positive samples are obvious enough for the model to learn effective distinguishing features. When constructing the corpus for the training and evaluation model, use positive and negative samples in a quantity ratio of 1:1 to maintain the balance between positive and negative samples.

[0179] After constructing the positive and negative sample dialogue data, convert the dialogue data into a form acceptable to the dialogue model. Concatenate the multi-turn dialogue and the corresponding response with the special symbol "sep", add a special symbol "cls" at the beginning to represent the start of the dialogue sample, and add a special symbol "eos" at the end to represent the end of the dialogue sample. The following shows the training data of two ranking models. X represents the input text (including the dialogue history and the corresponding response), and Y represents the sample label, indicating whether the dialogue history and the response match. Y = 1 represents a positive sample, indicating that the response and the dialogue history match; Y = 0 represents a negative sample, indicating that the response and the dialogue history do not match. The following are examples of positive and negative samples:

[0180] Positive sample:

[0181] X = [cls] Do you like video games? [sep] I'm a PS4 player. However, I mainly like sports and cycling. [eos]

[0182] Y = 1

[0183] Negative sample:

[0184] X = [cls] Do you like video games? [sep] My family works in the medical industry. [eos]

[0185] Y = 0

[0186] Next, perform word segmentation, indexing, vectorization, and positional encoding on the input text to convert the input text into a form acceptable to the model. The method is the same as that described in steps 31 and 32, and will not be elaborated here. Then, input the encoded text into the evaluation model. Refer to Figure 6 As shown, the evaluation model is also trained based on the transformed Transformer network. By adding a linear layer and a Softmax layer at the end of the Transformer network, map the final output of the Transformer model to the probability distribution of the class labels.

[0187] The evaluation model extracts the semantic features of the text through the self-attention mechanism and the feed-forward neural network. The output of the Transformer model is a sequence of vectors, where each vector represents a word in the input text. The first vector (corresponding to the special symbol [cls]) in the output vector sequence is input into the linear layer. This vector can be understood as the representation of the entire input dialogue text. The output of the linear layer is converted into a probability distribution over the class labels through the softmax function, and the class label with the highest probability can be selected as the prediction result. During training, the cross-entropy loss function is used to measure the difference between the predicted probability distribution and the true label, and the evaluation model parameters are updated by minimizing this loss function. The trained evaluation model can usually evaluate the accuracy, relevance, and fluency of the responses well.

[0188] After training the dialogue model and the evaluation model, next, according to the Figure 7 dialogue model shown below, the following combines Figure 7 to elaborate on the specific implementation steps in detail:

[0189] Step 71: Perform vector conversion on at least one word segment included in the dialogue to be processed to obtain a first word segment matrix.

[0190] In the embodiments of the present application, after obtaining the dialogue to be processed, the methods and steps of performing word segmentation, indexing, vectorization, and position encoding on the dialogue to be processed are the same as those in step 31, and will not be elaborated here.

[0191] Step 72: Input the first word segment matrix into the dialogue model trained based on the Transformer network. Based on the global features finally output by multiple multi-head self-attention modules in the dialogue model, obtain the response dialogue associated with the dialogue to be processed.

[0192] Specifically, referring to Figure 8 shown below, the following operations are performed in each multi-head self-attention module:

[0193] Step 721: Based on the preset key parameter matrix and value parameter matrix of the corresponding multi-head self-attention module, respectively extract the feature of the word meaning of the word segment represented by the input matrix to obtain the corresponding key feature matrix, and extract the feature of the word semantics of the word segment represented by the input matrix to obtain the corresponding value feature matrix.

[0194] In the embodiments of the present application, the methods for obtaining the key feature matrix and the value feature matrix are the same as those in step 321, and will not be elaborated here.

[0195] Step 722: Based on each preset query parameter matrix of the corresponding multi-head self-attention module, respectively extract the feature of the word meaning of the word segment represented by the input matrix to obtain the corresponding query feature matrix.

[0196] In the embodiments of the present application, the method for obtaining each query feature matrix is the same as the method in step 321, and will not be elaborated here.

[0197] Step 723: Obtain the relevance between each query feature matrix and the key feature matrix respectively, and based on each relevance respectively, combine the value feature matrix to obtain corresponding attention features.

[0198] In the embodiments of the present application, the method for obtaining each attention feature is the same as the method in step 322, and will not be elaborated here.

[0199] Step 724: Perform a fusion process on each attention feature to obtain a corresponding global feature.

[0200] In the embodiments of the present application, the method for obtaining the global feature is the same as the method in step 323, and will not be elaborated here.

[0201] Refer to Figure 9 As shown, after obtaining the global feature through the processing of multiple multi-head attention modules in the dialogue model, the background server corresponding to the dialogue model further performs:

[0202] Step 91: Perform a feature mapping process on the global feature to obtain a candidate token set corresponding to each of the N token positions.

[0203] Among them, the N token positions are set based on a preset dialogue grammar structure.

[0204] Specifically, first, perform a feature mapping process on the global feature to obtain a mapping probability set of the preset token set at the N token positions; among them, each preset token has a corresponding mapping probability at each token position.

[0205] For example, refer to Figure 10 As shown, for the dialogue to be processed "Which do you like, cats or dogs?", after processing by multiple multi-head self-attention modules, a global feature is obtained, and a full connection operation is performed on the global feature to obtain a mapping probability set of the preset token set at 4 token positions.

[0206] Then, for each of the N token positions, perform the following operations respectively:

[0207] In the mapping probability set corresponding to a token position, select at least one target probability that satisfies the preset candidate condition; from the preset token set, select the preset tokens corresponding to at least one target probability respectively to form a candidate token set.

[0208] For example, refer to Figure 10As shown, for each participle position, the minimum candidate word set whose cumulative probability of candidate words exceeds the threshold P is selected from the vocabulary. For example, P is 0.65, then for the four positions, the selected target probabilities are (0.8), (0.5, 0.1, 0.1), (0.6, 0.1), and (0.4, 0.3), respectively. They are mapped to the preset participle set, and the candidate participle sets for the four positions are {'I'} {'like', 'prefer', 'hate'} {'dog', 'cat'} {'a little', 'a lot'}.

[0209] Step 92: Multiple concatenations are performed between the candidate word segmentation sets to obtain multiple candidate dialogues; wherein, each time a concatenation is performed, a candidate word segmentation is selected from each candidate word segmentation set.

[0210] For example, see Figure 10 As shown, one candidate participle is selected from each of the candidate participle sets at the four positions for concatenation and combination, and four candidate dialogues are obtained: "I like cats", "I hate dogs", "I prefer cats a lot", and "I prefer cats a little".

[0211] Step 93: obtaining quality evaluation values ​​of the plurality of candidate dialogues based on the word segmentation correlations between the plurality of candidate dialogues and the dialogue to be processed;

[0212] For details, see Figure 11 As shown, when executing step 93, for multiple candidate dialogues, the following steps are respectively performed:

[0213] Step 931: concatenate the dialogue to be processed and a candidate dialogue to obtain the dialogue to be evaluated, and perform vector conversion on each segmentation word included in the dialogue to be evaluated to obtain a corresponding second segmentation matrix;

[0214] In the embodiment of the present application, the preprocessing process for the processed dialogue and the candidate dialogue is the same as the preprocessing process for the aforementioned training evaluation model, and will not be repeated here.

[0215] Step 932: Using the evaluation model based on the Transformer network training, feature extraction is performed on the second word segmentation matrix to obtain corresponding attention features; the attention feature represents: a word segmentation correlation between a candidate dialogue and a dialogue to be processed;

[0216] In an embodiment of the present application, the structure of each multi-head attention module in the Transformer network used in the evaluation model is the same as the multi-head attention module structure in the dialogue model, that is, in each multi-head self-attention module, each attention layer shares the key parameter matrix and the value parameter matrix.

[0217] Step 933: Based on the word segmentation relevance, obtain a quality assessment value of the candidate dialogue.

[0218] In the embodiments of the present application, the evaluation model is also trained based on the Transformer network with conversion. By adding a classification layer at the end of the Transformer network, the output of the Transformer model is mapped to the probability distribution of class labels. For the four candidate dialogues generated by the dialogue model, the evaluation model performs quality evaluation respectively. Specifically, the probability corresponding to the "match" category in the output probability distribution is used as the quality evaluation value of the corresponding candidate dialogue.

[0219] Step 94: Select the candidate dialogue whose quality evaluation value meets the preset reply condition as the reply dialogue.

[0220] For example, referring to Figure 12 As shown, the quality evaluation values corresponding to the four candidate dialogues "I like cats", "I hate dogs", "I prefer cats a lot", and "I prefer cats a little" are (0.4, 0.1, 0.25, 0.25) respectively. The candidate dialogue with the highest quality evaluation value, "I like cats", is selected as the final reply dialogue.

[0221] In addition, although the operations of the method of the present application are described in a specific order in the drawings, this does not require or imply that these operations must be performed in that specific order, or that all the operations shown must be performed to achieve the desired result. Additionally or alternatively, some steps may be omitted, multiple steps may be combined into one step for execution, and / or one step may be decomposed into multiple steps for execution.

[0222] Based on the same technical concept, referring to Figure 13 As shown, the embodiments of the present application further provide a dialogue generation device, which includes:

[0223] A vector conversion unit 1301, configured to perform vector conversion on at least one word segmentation included in the dialogue to be processed, and obtain a first word segmentation matrix;

[0224] A reply generation unit 1302, configured to input the first word segmentation matrix into a dialogue model trained based on the Transformer network with conversion, and obtain a reply dialogue associated with the dialogue to be processed based on the global features finally output by multiple multi-head self-attention modules in the dialogue model; wherein, the reply generation unit further includes:

[0225] A key-value feature acquisition sub-unit 13021, configured to perform the following operations for each multi-head attention module respectively: based on the preset key parameter matrix and value parameter matrix of the corresponding multi-head self-attention module, perform feature extraction on the word meaning of the word segmentation represented by the input matrix respectively to obtain the corresponding key feature matrix, and perform feature extraction on the word meaning of the word segmentation represented by the input matrix to obtain the corresponding value feature matrix;

[0226] A query feature acquisition subunit 13022 is configured to perform the following operations for each multi-head attention module: respectively perform feature extraction on the word meanings of the word segments represented by the input matrix based on each query parameter matrix preset for the corresponding multi-head self-attention module, so as to obtain a corresponding query feature matrix;

[0227] An attention feature acquisition subunit 13023 is configured to perform the following operations for each multi-head attention module: respectively obtain the correlation between each query feature matrix and key feature matrix, and respectively obtain corresponding attention features by combining the value feature matrix based on each correlation;

[0228] A feature fusion subunit 13024 is configured to perform the following operations for each multi-head attention module: perform fusion processing on each attention feature to obtain a corresponding global feature.

[0229] Optionally, based on the global features finally output by multiple multi-head self-attention modules in the dialogue model, to obtain a reply dialogue associated with the dialogue to be processed, the reply generation unit 1302 is further configured to:

[0230] Perform feature mapping processing on the global features to obtain a candidate word segment set corresponding to each of the N word segment positions; where the N word segment positions are set based on a preset dialogue grammar structure;

[0231] Perform multiple splicing combinations among the candidate word segment sets to obtain multiple candidate dialogues; where each time of splicing combination, a candidate word segment is respectively selected from each candidate word segment set;

[0232] Respectively obtain quality evaluation values of multiple candidate dialogues based on the word segment correlation between the multiple candidate dialogues and the dialogue to be processed;

[0233] Select the candidate dialogue whose quality evaluation value meets the preset reply condition as the reply dialogue.

[0234] Optionally, perform feature mapping processing on the global features to obtain a candidate word segment set corresponding to each of the N word segment positions, and the reply generation unit 1302 is further configured to:

[0235] Perform feature mapping processing on the global features to obtain a mapping probability set of a preset word segment set at each of the N word segment positions; where each preset word segment has a corresponding mapping probability at each word segment position;

[0236] For each of the N word segment positions, perform the following operations respectively:

[0237] In the mapping probability set corresponding to a word segment position, select at least one target probability that meets the preset candidate condition;

[0238] Select the preset word segments corresponding to at least one target probability from the preset word segment set to form a candidate word segment set.

[0239] Optionally, based on the token relevance between multiple candidate conversations and the conversation to be processed respectively, quality evaluation values of each of the multiple candidate conversations are obtained, and the response generation unit 1302 is further configured to:

[0240] For the multiple candidate conversations, the following operations are respectively performed:

[0241] Concatenate the conversation to be processed and a candidate conversation to obtain a conversation to be evaluated, and perform vector conversion on each token included in the conversation to be evaluated to obtain a corresponding second token matrix;

[0242] Use an evaluation model trained based on the Transformer network to perform feature extraction on the second token matrix to obtain corresponding attention features; the attention features represent: the token relevance between a candidate conversation and the conversation to be processed;

[0243] Optionally, the apparatus further includes a model training unit 1303, and the evaluation model is obtained by performing model training based on a positive sample set and a negative sample set; each positive sample includes: a sample conversation to be evaluated, a positive sample response, and a positive sample label, and the model training unit is further configured to construct a negative sample set:

[0244] For each positive sample response in each positive sample, perform sample reconstruction processing by using at least one of the methods of randomly replacing the response, replacing keywords, randomly replacing clauses, rearranging tokens, repeating response clauses, and adding interference information to generate corresponding negative sample responses;

[0245] Based on the negative sample labels, each sample conversation to be evaluated, and the corresponding negative sample responses, construct a negative sample set.

[0246] Optionally, the model training unit 1303 is further configured to:

[0247] Obtain a preset training sample set, and each training sample includes: a sample conversation to be replied and a true sample response;

[0248] Based on the training sample set, perform multiple rounds of iterative training on the conversation model to be trained. Among them, in one round of iteration process, the following steps are performed:

[0249] Input the sample conversation to be replied into the Transformer network, and based on the sample global features finally output by multiple multi-head self-attention modules, obtain a predicted sample response associated with the sample conversation to be replied, and adjust the model parameters based on the loss function value between the predicted sample response and the true sample response; among them, each multi-head self-attention module performs the following operations:

[0250] Based on the preset key parameter matrix and value parameter matrix of the corresponding multi-head self-attention module, obtain the corresponding sample key feature matrix and sample value feature matrix, and respectively based on each preset query parameter matrix of the corresponding multi-head self-attention module, obtain the corresponding sample query feature matrix;

[0251] Based on the sample key feature matrix and sample value feature matrix, combined with each sample query feature matrix, obtain the corresponding sample attention feature, and perform fusion processing on each sample attention feature to obtain the corresponding sample global feature.

[0252] Based on the same technical concept, the embodiment of the present application also provides an electronic device, which can implement the method flow of dialogue generation provided in the above embodiment of the present application.

[0253] In one embodiment, the electronic device can be a server, or a terminal device or other electronic devices.

[0254] Refer to Figure 14 As shown, the electronic device may include:

[0255] At least one processor 1401, and a memory 1402 connected to at least one processor 1401. In the embodiment of the present application, the specific connection medium between the processor 1401 and the memory 1402 is not limited. Figure 14 In the example, the processor 1401 and the memory 1402 are connected through a bus 1400. The bus 1400 is Figure 14 shown by a thick line in the figure. The connection methods between other components are only for illustrative purposes and are not limited thereto. The bus 1400 can be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience of representation, Figure 14 only a thick line is shown in the figure, but it does not mean that there is only one bus or one type of bus. Or, the processor 1401 can also be called a controller, and the name is not limited.

[0256] In the embodiment of the present application, the memory 1402 stores instructions executable by at least one processor 1401. By executing the instructions stored in the memory 1402, at least one processor 1401 can execute a dialogue generation method described above. The processor 1401 can implement Figure 13 the functions of each module in the device shown in the figure.

[0257] Among them, the processor 1401 is the control center of the device, and can connect various parts of the entire control device through various interfaces and lines. By running or executing the instructions stored in the memory 1402 and calling the data stored in the memory 1402, various functions of the device and process data, so as to monitor the device as a whole.

[0258] In a possible design, the processor 1401 may include one or more processing units. The processor 1401 may integrate an application processor and a modem processor. Among them, the application processor mainly processes the operating system, user interface, application programs, etc., and the modem processor mainly processes wireless communications. It can be understood that the above-mentioned modem processor may not be integrated into the processor 1401 either. In some embodiments, the processor 1401 and the memory 1402 may be implemented on the same chip, and in some embodiments, they may also be separately implemented on independent chips.

[0259] The processor 1401 may be a general-purpose processor, such as a CPU, a digital signal processor, an application-specific integrated circuit, a field-programmable gate array, or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, and can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor may be a microprocessor or any conventional processor, etc. The steps of a dialogue generation method disclosed in combination with the embodiments of the present application may be directly embodied as being executed by a hardware processor, or executed by a combination of hardware and software modules in the processor.

[0260] As a non-volatile computer-readable storage medium, the memory 1402 can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules. The memory 1402 may include at least one type of storage medium, for example, it may include flash memory, a hard disk, a multimedia card, a card-type memory, a random access memory (RAM), a static random access memory (SRAM), a programmable read-only memory (PROM), a read-only memory (ROM), an electrically erasable programmable read-only memory (EEPROM), a magnetic memory, a magnetic disk, an optical disc, etc. The memory 1402 is any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory 1402 in the embodiments of the present application may also be a circuit or any other device capable of implementing a storage function, for storing program instructions and / or data.

[0261] By programming the design of the processor 1401, the code corresponding to the dialogue generation method introduced in the foregoing embodiments can be solidified into the chip, so that the chip can execute when running Figure 7 and Figure 8Steps of a dialogue generation method of the illustrated embodiment. How to design and program the processor 1401 is a well-known technology to those skilled in the art and will not be elaborated here.

[0262] Based on the same inventive concept, an embodiment of the present application also provides a storage medium storing computer instructions, which when run on a computer, cause the computer to execute a dialogue generation method described above.

[0263] In some possible implementation manners, various aspects of a dialogue generation method provided by the present application can also be implemented in the form of a program product, which includes program code that, when the program product runs on a device, is used to cause the control device to execute the steps in a dialogue generation method according to various exemplary embodiments of the present application described above in this specification.

[0264] It should be noted that although several units or subunits of the device are mentioned in the above detailed description, this division is merely exemplary and not mandatory. In fact, according to the embodiments of the present application, the features and functions of two or more of the above-mentioned units can be embodied in one unit. Conversely, the features and functions of one unit described above can be further divided and embodied by multiple units.

[0265] In addition, although the operations of the method of the present application are described in a specific order in the drawings, this does not require or imply that these operations must be performed in that specific order, or that all of the shown operations must be performed to achieve the desired result. Additionally or alternatively, some steps may be omitted, multiple steps may be combined into one step for execution, and / or one step may be decomposed into multiple steps for execution.

[0266] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0267] This application is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to the application. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing device to produce a machine, such that the instructions executed by the processor of the computer or other programmable data processing device generate means for implementing the functions specified in the Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.

[0268] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufacture including instruction means that implement the functions specified in the Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.

[0269] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in the Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.

[0270] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application. Thus, if these modifications and variations of this application fall within the scope of the claims of this application and their equivalent technologies, this application is also intended to include these modifications and variations.

Claims

1. A method for generating conversations, characterized in that, Including: Performing vector transformation on at least one word segment included in the dialogue to be processed to obtain a first word segment matrix; Inputting the first word segment matrix into a dialogue model trained based on a transformed Transformer network, and obtaining a reply dialogue associated with the dialogue to be processed based on global features finally output by a plurality of multi-head self-attention modules in the dialogue model; wherein, each multi-head self-attention module performs the following operations: Based on a preset key parameter matrix and value parameter matrix corresponding to the multi-head self-attention module, respectively performing feature extraction on the word meanings of the word segments represented by the input matrix to obtain a corresponding key feature matrix, and performing feature extraction on the word meanings of the word segments represented by the input matrix to obtain a corresponding value feature matrix; Based on each preset query parameter matrix corresponding to the multi-head self-attention module, respectively performing feature extraction on the word meanings of the word segments represented by the input matrix to obtain a corresponding query feature matrix; Respectively obtaining the correlation degrees between each query feature matrix and the key feature matrix, and respectively combining the value feature matrix based on each correlation degree to obtain corresponding attention features; Performing fusion processing on each attention feature to obtain a corresponding global feature.

2. The method according to claim 1, characterized in that, The obtaining of the reply dialogue associated with the dialogue to be processed based on the global features finally output by the plurality of multi-head self-attention modules in the dialogue model includes: Performing feature mapping processing on the global features to obtain a candidate word segment set corresponding to each of N word segment positions; wherein, the N word segment positions are set based on a preset dialogue grammar structure; Performing multiple splicing combinations among the candidate word segment sets to obtain multiple candidate dialogues; wherein, each time of splicing combination, one candidate word segment is respectively selected from each candidate word segment set; Respectively obtaining quality evaluation values of the multiple candidate dialogues based on the word segment correlation between the multiple candidate dialogues and the dialogue to be processed; Selecting the candidate dialogue whose quality evaluation value meets a preset reply condition as the reply dialogue.

3. The method according to claim 2, characterized in that, The performing of the feature mapping processing on the global features to obtain a candidate word segment set corresponding to each of N word segment positions includes: Performing feature mapping processing on the global features to obtain a mapping probability set of a preset word segment set at each of the N word segment positions; wherein, each preset word segment has a corresponding mapping probability at each word segment position; For each of the N word segment positions, respectively performing the following operations: Selecting at least one target probability that meets a preset candidate condition from the mapping probability set corresponding to one word segment position; Selecting the preset word segments corresponding to the at least one target probability from the preset word segment set to form the candidate word segment set.

4. The method according to claim 2 or 3, characterized in that, Including, the respectively obtaining the quality evaluation values of the multiple candidate dialogues based on the word segment correlation between the multiple candidate dialogues and the dialogue to be processed includes: For each of the multiple candidate dialogues, respectively performing the following operations: Splicing the dialogue to be processed and a candidate dialogue to obtain a dialogue to be evaluated, and performing vector transformation on each word segment included in the dialogue to be evaluated to obtain a corresponding second word segment matrix; An evaluation model trained based on a Transformer network is used to extract features from the second tokenized matrix to obtain corresponding attention features; the attention features characterize the token correlation between the candidate dialogue and the dialogue to be processed. Based on the token correlation, a quality evaluation value of the candidate dialogue is obtained.

5. The method according to claim 4, characterized in that, The evaluation model is obtained by training the model based on a positive sample set and a negative sample set; each positive sample includes: a sample dialogue to be evaluated, a positive sample response, and a positive sample label. Then the negative sample set is constructed in the following way: For each positive sample response in the positive samples, at least one of the methods of randomly replacing the response, replacing keywords, randomly replacing clauses, rearranging tokens, repeating response clauses, and adding interference information is used to perform sample reconstruction processing to generate corresponding negative sample responses. Based on the negative sample labels, each sample dialogue to be evaluated, and the corresponding negative sample responses, a negative sample set is constructed.

6. The method according to any one of claims 1-3, characterized in that, The training process of the dialogue model is as follows: A preset training sample set is obtained, and each training sample includes: a sample dialogue to be replied to and a true sample response. Based on the training sample set, the dialogue model to be trained is subjected to multiple rounds of iterative training. Among them, in one round of the iterative process, the following steps are executed: The sample dialogue to be replied to is input into the Transformer network. Based on the sample global features finally output by multiple multi-head self-attention modules, a predicted sample response associated with the sample dialogue to be replied to is obtained. Based on the loss function value between the predicted sample response and the true sample response, the model parameters are adjusted; among them, each multi-head self-attention module performs the following operations: Based on the key parameter matrix and value parameter matrix preset for the multi-head self-attention module, a corresponding sample key feature matrix and sample value feature matrix are obtained, and based on each query parameter matrix preset for the multi-head self-attention module, a corresponding sample query feature matrix is obtained. Based on the sample key feature matrix and the sample value feature matrix, combined with the sample query feature matrices, corresponding sample attention features are obtained, and the sample attention features are fused to obtain corresponding sample global features.

7. A device for generating conversations, characterized in that, Including: A vector conversion unit for performing vector conversion on at least one token included in the dialogue to be processed to obtain a first tokenized matrix. A response generation unit for inputting the first tokenized matrix into a dialogue model trained based on a Transformer network, and obtaining a response dialogue associated with the dialogue to be processed based on the global features finally output by multiple multi-head self-attention modules in the dialogue model; among them, the response generation unit further includes: A key-value feature acquisition sub-unit for respectively performing, for each multi-head attention module: based on the key parameter matrix and value parameter matrix preset for the multi-head self-attention module, extracting features from the token semantics represented by the input matrix to obtain a corresponding key feature matrix, and extracting features from the token semantics represented by the input matrix to obtain a corresponding value feature matrix. The query feature acquisition subunit is configured to perform the following operations for each multi-head attention module: respectively perform feature extraction on the word meanings of the word segments represented by the input matrix based on each query parameter matrix preset for the corresponding multi-head self-attention module, so as to obtain corresponding query feature matrices; The attention feature acquisition subunit is configured to perform the following operations for each multi-head attention module: respectively obtain the relevance between each query feature matrix and the key feature matrix, and respectively obtain corresponding attention features by combining the value feature matrix based on each relevance; The feature fusion subunit is configured to perform the following operations for each multi-head attention module: perform fusion processing on each attention feature to obtain corresponding global features.

8. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein, When the processor executes the computer program, the method described in any one of claims 1-6 is implemented.

9. A computer-readable storage medium, on which a computer program is stored, wherein, When the computer program is executed by the processor, the steps of the method described in any one of claims 1-6 are implemented.

10. A computer program product, wherein, When the computer program product is called by a computer, the computer is caused to execute the method described in any one of claims 1-6.