Dialogue model training method, data processing method and related device

By evaluating and updating the parameters of the target dialogue model, combining the differences between the predicted reply data and the actual reply data and the quality of the speech, the problem of low accuracy of the dialogue model in the target business scenario is solved, and higher accuracy of the reply data and better speech quality are achieved.

CN120562545APending Publication Date: 2025-08-29MASHANG CONSUMER FINANCE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410227777.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-02-28
Publication Date
2025-08-29

AI Technical Summary

Technical Problem

In the target business scenario, the prior art has low accuracy in generating reply data due to the small dialogue data and low speech quality.

Method used

By obtaining conversation data in the target business scenario, using pre-trained speech evaluation model to evaluate the predicted reply data quality of the target dialogue model, and combining the differences between the reply data and the predicted reply data, the model parameters of the target dialogue model are updated to improve the accuracy and speech quality of the reply data.

Benefits of technology

The accuracy and speech quality of the target dialogue model generate reply data is improved, and the response ability of the dialogue model in the target business scenario is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120562545A_ABST
    Figure CN120562545A_ABST
Patent Text Reader

Abstract

The invention provides a dialogue model training method, a data processing method and a related device.The dialogue model training method comprises the steps that dialogue data in a target service scene is obtained, the occurrence frequency of the target service scene is smaller than a threshold value, and the dialogue data comprises first data and first reply data replying to the first data; the to-be-trained target dialogue model carries out reply prediction on the first data to obtain first prediction reply data, the pre-trained verbal skill evaluation model carries out verbal skill quality evaluation on the first prediction reply data to obtain an evaluation result, the difference between the first reply data and the first prediction reply data is obtained, and the verbal skill quality of the target dialogue model is evaluated based on the difference and the evaluation result. And updating model parameters of the target dialogue model. Through the method, the accuracy of generating the reply data by the target dialogue model can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of data processing technology, and in particular to a conversation model training method, a data processing method, and related devices. Background Art

[0002] Conversational models can be used to engage in conversations with users and answer their questions. However, for target business scenarios, since these scenarios are only valid within a target time period, the corresponding conversation data is relatively small, the quality of the dialogue used is low, and the accuracy of the responses included in the conversation data is low. This results in low accuracy in the responses generated by the conversational model. Summary of the Invention

[0003] The embodiments of the present application provide a training method, a data processing method, and related devices for a dialogue model, which can improve the accuracy of response data generated by a target dialogue model.

[0004] The technical solution of the embodiment of the present application is implemented as follows:

[0005] An embodiment of the present application provides a method for training a dialogue model, the method comprising:

[0006] Acquire conversation data in a target business scenario, where the occurrence frequency of the target business scenario is less than a threshold, and the conversation data includes first data and first reply data in reply to the first data;

[0007] The target dialogue model to be trained predicts a response to the first data to obtain first predicted response data;

[0008] The pre-trained speech evaluation model performs speech quality evaluation on the first predicted response data to obtain an evaluation result;

[0009] Obtain a difference between the first reply data and the first predicted reply data, and update model parameters of the target dialogue model based on the difference and the evaluation result.

[0010] This embodiment of the present application provides a data processing method, the method comprising:

[0011] Obtaining data to be replied under a target business scenario, where the occurrence frequency of the target business scenario is less than a threshold;

[0012] The target dialogue model predicts a reply for the data to be replied to and obtains predicted reply data. The target dialogue model is trained based on the difference between the first reply data and the first predicted reply data and the evaluation result. The first predicted reply data is obtained by the target dialogue model to be trained predicting a reply for the first data. The evaluation result is obtained by the pre-trained speech evaluation model performing a speech quality evaluation on the first predicted reply data. The first data and the first reply data are included in the dialogue data under the target business scenario.

[0013] An embodiment of the present application provides a device for training a dialogue model, the device comprising:

[0014] an acquisition module, configured to acquire conversation data in a target business scenario, where the occurrence frequency of the target business scenario is less than a threshold, and the conversation data includes first data and first reply data in reply to the first data;

[0015] An input module, configured for the target dialogue model to be trained to predict a response to the first data to obtain first predicted response data;

[0016] The input module is further configured to use a pre-trained speech evaluation model to perform speech quality evaluation on the first predicted reply data to obtain an evaluation result;

[0017] An updating module is used to obtain a difference between the first reply data and the first predicted reply data, and update the model parameters of the target dialogue model based on the difference and the evaluation result.

[0018] An embodiment of the present application provides a data processing device, the device comprising:

[0019] An acquisition module, configured to acquire data to be replied under a target business scenario, wherein the occurrence frequency of the target business scenario is less than a threshold;

[0020] A prediction module is used for the target dialogue model to predict the reply of the data to be replied to and obtain predicted reply data. The target dialogue model is trained based on the difference between the first reply data and the first predicted reply data and the evaluation result. The first predicted reply data is obtained by the target dialogue model to be trained to predict the reply of the first data. The evaluation result is obtained by the pre-trained speech evaluation model to evaluate the speech quality of the first predicted reply data. The first data and the first reply data are included in the dialogue data under the target business scenario.

[0021] An embodiment of the present application provides an electronic device, comprising:

[0022] a memory for storing computer-executable instructions;

[0023] The processor is used to implement the training method of the dialogue model provided in the embodiment of the present application, or implement the data processing method provided in the embodiment of the present application when executing the computer-executable instructions stored in the memory.

[0024] An embodiment of the present application provides a computer-readable storage medium storing a computer program or computer-executable instructions, which is used to implement the training method of the dialogue model provided in the embodiment of the present application, or implement the data processing method provided in the embodiment of the present application when executed by a processor.

[0025] An embodiment of the present application provides a computer program product, including a computer program or computer-executable instructions. When the computer program or computer-executable instructions are executed by a processor, the training method of the dialogue model provided in the embodiment of the present application is implemented, or the data processing method provided in the embodiment of the present application is implemented.

[0026] The embodiment of the present application can evaluate the speech quality of the first reply data output by the target dialogue model to be trained, and use the evaluation results as the basis for adjusting the model parameters. That is, the embodiment of the present application combines the difference between the first reply data and the first predicted reply data and the evaluation results to adjust the model parameters of the target dialogue model. Compared with the method of adjusting the model parameters based only on the difference between the predicted reply data and the reply data, the reply data generated by the target dialogue model can be made more accurate, and the speech quality of the reply data generated by the target dialogue model can be improved, which can further improve the accuracy of the reply data generated by the target dialogue model. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] Figure 1 This is a schematic diagram of the structure of the dialogue system provided in an embodiment of the present application;

[0028] Figure 2A This is a first structural diagram of an electronic device provided in an embodiment of the present application;

[0029] Figure 2B is a second structural diagram of an electronic device provided in an embodiment of the present application;

[0030] Figure 3 This is a first flow chart of the method for training a dialogue model provided in an embodiment of the present application;

[0031] Figure 4 This is a schematic diagram of the structure of the target dialogue model provided in an embodiment of the present application;

[0032] Figure 5 A second flow chart of the method for training a dialogue model provided in an embodiment of the present application;

[0033] Figure 6 This is a third flow chart of the method for training a dialogue model provided in an embodiment of the present application;

[0034] Figure 7 This is a fourth flow chart of the method for training a dialogue model provided in an embodiment of the present application;

[0035] Figure 8 This is a fifth flow chart of the method for training a dialogue model provided in an embodiment of the present application;

[0036] Figure 9 This is a sixth flow chart of the method for training a dialogue model provided in an embodiment of the present application;

[0037] Figure 10 This is a seventh flow chart of the method for training a dialogue model provided in an embodiment of the present application;

[0038] Figure 11 This is an eighth flow chart of the method for training a dialogue model provided in an embodiment of the present application;

[0039] Figure 12 Schematic diagram of a training speech quality scoring model provided in an embodiment of the present application;

[0040] Figure 13 Schematic diagram of a training target dialogue model provided in an embodiment of the present application;

[0041] Figure 14 It is a flowchart of data processing provided by an embodiment of the present application.

[0042] It should be pointed out that the above-mentioned "first" and "second" are only used to distinguish different solutions, and do not represent the degree of distinction between the advantages and disadvantages of the solutions or the priority in the implementation process. DETAILED DESCRIPTION

[0043] In order to make the purpose, technical solutions and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limiting this application. All other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.

[0044] In the following description, reference is made to “some embodiments”, which describes a subset of all possible embodiments, but it will be understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.

[0045] In the following description, the terms "first\second\third" involved are merely used to distinguish similar objects and do not represent a specific ordering of the objects. It can be understood that "first\second\third" can be interchanged with a specific order or sequence where permitted, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.

[0046] Unless otherwise defined, all technical and scientific terms used in the embodiments of the present application have the same meanings as those commonly understood by those skilled in the art. The terms used in the embodiments of the present application are only for the purpose of describing the embodiments of the present application and are not intended to limit the present application.

[0047] In related art, for a target business scenario valid only within a target time period, to enable a conversational model to respond to questions corresponding to the target business scenario, the conversational model training method is as follows: based on sample conversational data, conversational data corresponding to the target business scenario is added and the conversational model is retrained. The target business scenario includes data to be responded to and data to respond to the data to be responded to. The response data can serve as label data. During the conversational model training process, the discrepancy between the predicted response data and the response data can be used to adjust the conversational model parameters.

[0048] The related technology has the following shortcomings: the quality of the dialogue data corresponding to the target business scenario is low, and the accuracy of the reply data included in the dialogue data is low, which leads to low accuracy of the reply data generated by the dialogue model.

[0049] In response to the above-mentioned shortcomings, the embodiments of the present application provide a training method, a data processing method and related devices for a dialogue model, which can evaluate the speech quality of the first reply data output by the target dialogue model to be trained, and use the evaluation results as the basis for adjusting the model parameters, thereby adjusting the model parameters of the target dialogue model in combination with the difference between the first reply data and the first predicted reply data and the evaluation results. Compared with the method of adjusting the model parameters only in combination with the difference between the predicted reply data and the reply data in the related art, the reply data generated by the target dialogue model can be made more accurate, and the speech quality of the reply data generated by the target dialogue model can be improved, which can further improve the accuracy of the reply data generated by the target dialogue model.

[0050] See also Figure 1 , Figure 1 This is a structural diagram of the dialogue system 100 provided in an embodiment of the present application. In order to support a training application of a dialogue model and a data processing application, the terminal 400 is connected to the server 200 through the network 300. The network 300 can be a wide area network or a local area network, or a combination of the two.

[0051] Taking conversation model training on server 200 as an example, server 200 is configured to obtain conversation data for a target business scenario. The conversation data includes first data and first reply data in response to the first data. The target conversation model to be trained predicts a reply to the first data input, obtaining first predicted reply data.

[0052] The pre-trained speech evaluation model performs speech quality evaluation on the first predicted reply data to obtain an evaluation result, obtains the difference between the first reply data and the first predicted reply data, and updates the model parameters of the target dialogue model based on the difference and the evaluation result, thereby completing the training of the target dialogue model.

[0053] The following describes a method for implementing a data processing application. The terminal 400 is used to display a dialogue interface on a graphical interface 410. In response to a text input instruction, the data to be replied can be obtained. The terminal 400 can display the data to be replied on the graphical interface 410 and send the data to be replied to the server 200.

[0054] After receiving the data to be replied, the target dialogue model can predict the reply to the data to be replied, and obtain the predicted reply data. The server 200 can send the predicted reply data to the terminal 400, and the terminal 400 can display the predicted reply data on the graphical interface 410.

[0055] The conversation model training method and the data processing method provided in the embodiments of the present application can be implemented by various electronic devices or computer devices. For example, they can be implemented by a terminal alone, by a server alone, or by a terminal and a server in collaboration.

[0056] The following describes an exemplary application of an electronic device for implementing a training method for a dialogue model provided in an embodiment of the present application, as well as a data processing method provided in an embodiment of the present application. The electronic device may be a terminal or a server, and the terminal may be various types of user terminals such as a laptop computer, a tablet computer, a desktop computer, a set-top box, a mobile device (e.g., a mobile phone, a portable music player, a personal digital assistant, a dedicated messaging device, a portable gaming device), a smart phone, a smart speaker, a smart watch, a smart TV, and a vehicle-mounted terminal.

[0057] The server 200 may be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It may also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms. The terminal and the server may be connected directly or indirectly via wired or wireless communication, which is not limited in the embodiments of the present application.

[0058] See also Figure 2A and Figure 2B , Figure 2A is a first structural diagram of an electronic device provided in an embodiment of the present application, Figure 2B is a second structural diagram of an electronic device provided in an embodiment of the present application, Figure 2A and Figure 2B The electronic device shown includes: at least one processor 210, a memory 250, at least one network interface 220 and a user interface 230. The various components in the electronic device are coupled together via a bus system 240. It is understood that the bus system 240 is used to achieve connection and communication between these components. In addition to the data bus, the bus system 240 also includes a power bus, a control bus and a status signal bus. However, for the sake of clarity, the bus system 240 is not described in detail. Figure 2A and Figure 2B Various buses are labeled as bus system 240 .

[0059] The processor 210 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc., where the general-purpose processor can be a microprocessor or any conventional processor, etc.

[0060] The user interface 230 includes one or more output devices 231 that enable presentation of media content, including one or more speakers and / or one or more visual display screens. The user interface 230 also includes one or more input devices 232, including user interface components that facilitate user input, such as a keyboard, mouse, microphone, touch screen display, camera, other input buttons and controls.

[0061] The memory 250 may be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state memory, hard drives, optical drives, etc. The memory 250 may optionally include one or more storage devices that are physically remote from the processor 210.

[0062] The memory 250 includes volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. The non-volatile memory may be a read-only memory (ROM), and the volatile memory may be a random access memory (RAM). The memory 250 described in the embodiments of the present application is intended to include any suitable type of memory.

[0063] In some embodiments, the memory 250 can store data to support various operations. Examples of such data include programs, modules, and data structures, or subsets or supersets thereof, as exemplified below. The operating system 251 includes system programs for processing various basic system services and performing hardware-related tasks, such as a framework layer, a core library layer, and a driver layer, for implementing various basic services and processing hardware-based tasks.

[0064] A network communication module 252 is used to reach other electronic devices via one or more (wired or wireless) network interfaces 220. Exemplary network interfaces 220 include Bluetooth, Wireless LAN (WiFi), and Universal Serial Bus (USB). A presentation module 253 is used to enable information presentation (e.g., a user interface for operating peripheral devices and displaying content and information) via one or more output devices 231 associated with a user interface 230 (e.g., a display screen, a speaker, etc.);

[0065] The input processing module 254 is used to detect one or more user inputs or interactions from one of the one or more input devices 232 and translate the detected inputs or interactions. In some embodiments, the training device of the dialogue model provided in the embodiment of the present application can be implemented in software. Figure 2A Dialogue model training device 255 stored in memory 250 is shown. This device can be software in the form of a program or plug-in, and includes the following software modules: an acquisition module 2551, an input module 2552, and an update module 2553. These modules are logical and can be arbitrarily combined or further separated according to the functions they implement. The functions of each module are described below.

[0066] In some embodiments, the data processing device provided in the embodiments of the present application can be implemented in software. Figure 2BThe data processing device 256 stored in the memory 250 is shown. This device can be software in the form of a program or plug-in, and includes the following software modules: an acquisition module 2561 and a prediction module 2562. These modules are logical and can be arbitrarily combined or further separated according to the functions they implement. The functions of each module will be described below.

[0067] In other embodiments, the conversation model training device 255 provided in the embodiments of the present application and the data processing device 256 provided in the embodiments of the present application can both be implemented in hardware. As an example, they can be processors in the form of hardware decoding processors that are programmed to execute the conversation model training method provided in the embodiments of the present application. For example, the processors in the form of hardware decoding processors can be one or more application-specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), or other electronic components.

[0068] Based on the above description of the dialogue system and electronic device of the embodiment of the present application, Figure 3 The following describes the training method of the dialogue model provided in the embodiment of the present application. Figure 3 , Figure 3 This is a first flow chart of the method for training a dialogue model provided in an embodiment of the present application. As mentioned above, the electronic device for implementing the method for training a dialogue model in an embodiment of the present application may be a terminal, a server, or a combination of the two. Below, the method for training a dialogue model provided in an embodiment of the present application is described by taking server implementation as an example.

[0069] In step 101, conversation data in a target business scenario is obtained.

[0070] The frequency of occurrence of the target business scenario is less than the pre-set frequency threshold (i.e., the threshold). The target business scenario can be set by the user based on actual usage needs and is valid within the target time period. For example, the target business scenario can be an activity corresponding to a festival, such as an activity scenario corresponding to the Mid-Autumn Festival or an activity scenario corresponding to the National Day. For another example, the target business scenario can be an emergency, such as a temporary business activity organized in response to a certain requirement issued at the moment. The frequency threshold can refer to a threshold within a certain time range, such as within a year, within a quarter, etc. The frequency threshold can be set according to actual needs. For example, the frequency threshold can be 1, 2, 3, etc.

[0071] Retrieving conversation data for a target business scenario may include retrieving conversation data within a target time period. For example, the target business scenario may be Activity 51, and the target time period is 00:00 on May 1st to 23:00 on May 3rd. Conversation data for Activity 51 between 00:00 on May 1st and 23:00 on May 3rd may be retrieved. For another example, the target business scenario may be Burst Activity 1, and the target time period is 10:00 on December 16th to 12:00 on December 16th. Conversation data for Burst Activity 1 between 10:00 on December 16th and 12:00 on December 16th may be retrieved.

[0072] The conversation data includes first data (data to be replied) and first reply data for replying to the first data. For example, the first data may be "What are the activities during the holiday?", and the first reply data may be the specific content of the activities during the holiday.

[0073] Conversation data refers to conversation data generated for a target business scenario. This conversation data can be used as training data to train the target conversation model, with the first reply data in the conversation data serving as the label. In some embodiments, the amount of conversation data can be determined based on actual usage requirements. For example, the amount of conversation data can be 500,000, 600,000, or 700,000, and is not specifically limited here.

[0074] In some embodiments, the conversation data is uncleaned data, meaning it contains noise. Cleaning activity data would be labor-intensive and time-consuming. However, the conversation model training method provided in the embodiments of this application can reduce the labor and time required for data cleaning and improve the efficiency of training the target conversation model.

[0075] In step 102, the target dialogue model to be trained predicts a response to the first data to obtain first predicted response data.

[0076] To facilitate understanding of the target conversation model, let's first explain its structure. The target conversation model can be used for chatting with users or conducting knowledge Q&A sessions. This is equivalent to a target conversation model acting as a customer service representative in a specific field. The target conversation model can include a general conversation model and a bypass model. The general conversation model is essentially a large language model.

[0077] A bypass model is trained based on conversation data. A general conversation model can be based on the general conversation model to be trained and training conversation data from common business scenarios (i.e., general conversation data, also known as domain-wide data). The general conversation model is a model obtained through data screening, supervised fine-tuning (SFT) training, and reinforcement training. The number of parameters in a general conversation model is much larger than that in a bypass model.

[0078] Common business scenarios are business scenarios valid within any time period. Training conversation data for common business scenarios is cleansed conversation data and is not affected by the target business scenario. For example, training conversation data for common business scenarios might be a question from a questioner asking, "What is the meaning of noun A?" The corresponding response data might include the meaning of noun A and related explanations.

[0079] After the first data is input into the target dialogue model, both the general dialogue model and the bypass model can predict the reply data corresponding to the first data, and then obtain the first predicted reply data based on the reply data predicted by the general dialogue model and the reply data predicted by the bypass model.

[0080] In some embodiments, obtaining the first predicted reply data based on the reply data predicted by the general conversation model and the reply data predicted by the bypass model can be: fusing the reply data predicted by the general conversation model and the reply data predicted by the bypass model to obtain the first predicted reply data.

[0081] For example, see Figure 4 , Figure 4 301 is a schematic diagram of the target dialogue model provided in an embodiment of the present application. Target dialogue model 301 may include general dialogue model 302 and bypass model 303. First data is input into target dialogue model 301, that is, the first data is input into general dialogue model 302 and bypass model 303 respectively.

[0082] The universal dialogue model 302 may include an encoding layer 304 and a decoding layer 305. The encoding layer 304 is used to encode the input first data into a vector of fixed length, and the decoding layer 305 is used to convert the vector into a readable natural language output to obtain universal predicted reply data. The universal predicted reply data is the predicted reply data obtained by the universal dialogue model 302 when predicting a reply to the first data.

[0083] The bypass model 303 may include a dimensionality reduction layer 306 and a dimensionality increase layer 307. The dimensionality reduction layer 306 is used to reduce the dimensionality of the features corresponding to the input first data. The dimensionality increase layer 307 is used to increase the dimensionality of the reduced features to the dimensions before dimensionality reduction to obtain bypass prediction response data. The bypass prediction response data is the prediction response data obtained by the bypass model 303 after performing a response prediction on the first data.

[0084] The reply data predicted by the general dialogue model 302 and the reply data predicted by the bypass model 303 are fused, that is, the general predicted reply data and the bypass predicted reply data are fused, so as to obtain the first predicted reply data.

[0085] In some embodiments, obtaining first predicted reply data based on the reply data predicted by the general dialogue model and the reply data predicted by the bypass model can be: scoring the reply data predicted by the general dialogue model and the reply data predicted by the bypass model, and then determining the weight according to the scoring results, and fusing the reply data predicted by the general dialogue model and the reply data predicted by the bypass model according to the weight, so as to obtain the first predicted reply data.

[0086] Through step 102, the general dialogue model and the bypass model included in the target dialogue model can respectively predict the response of the first data, so that the first predicted response data can be obtained based on the response data predicted by the general dialogue model and the response data predicted by the bypass model.

[0087] In step 103, the pre-trained speech evaluation model performs speech quality evaluation on the first predicted reply data to obtain an evaluation result.

[0088] The speech quality evaluation model is used to evaluate the speech quality of the first predicted response data. Speech quality can be the quality of the language expression. Speech quality can be divided into multiple levels, for example, high, medium, and low. For another example, speech quality can be graded from 1 to 7, with increasing numbers indicating higher speech quality.

[0089] The evaluation result can be a probability value of the first predicted reply data being a qualified speech. A qualified speech can also be an excellent speech, and a qualified speech is a speech whose speech quality is greater than a preset quality threshold. The corresponding value range of the probability value is 0-1. The preset quality threshold can be set according to the level corresponding to the speech quality and the actual needs of the user. For example, when the speech quality level is high, medium, and low, the preset quality threshold can be the value corresponding to medium. Taking the speech quality value of 0-100 as an example, the value corresponding to medium can be 67, that is, the preset quality threshold is 67.

[0090] For another example, when the level of speech quality is still level 1-level 7, the preset quality threshold can be the value corresponding to level 5. Taking the speech quality value of 0-100 as an example, the value corresponding to level 5 can be 72, that is, the preset quality threshold is 72.

[0091] In step 104 , the difference between the first reply data and the first predicted reply data is obtained, and based on the difference and the evaluation result, the model parameters of the target dialogue model are updated.

[0092] In an embodiment of the present application, the loss function used to update the target dialogue model combines two parts of information, one part is the difference between the first reply data and the first predicted reply data, and the other part is an evaluation result indicating the speech quality of the first predicted reply data.

[0093] That is to say, the embodiment of the present application combines the difference between the first reply data and the first predicted reply data and the evaluation results to adjust the model parameters of the target dialogue model. Compared with the method of adjusting the model parameters only based on the difference between the predicted reply data and the reply data, the reply data generated by the target dialogue model can be made more accurate, and the quality of the reply data generated by the target dialogue model can be improved, which can further improve the accuracy of the reply data generated by the target dialogue model.

[0094] In the embodiment of the present application, when adjusting the model parameters of the target dialogue model, the response data generated by the target dialogue model can be made more accurate by minimizing the loss function. When the loss function includes the evaluation result, the speech quality of the response data generated by the target dialogue model can also be improved, which can further improve the accuracy of the response data generated by the target dialogue model.

[0095] In some embodiments, see Figure 5 , Figure 5 This is a second flow chart of the method for training a dialogue model provided in an embodiment of the present application. Figure 3 The step 104 of “updating the model parameters of the target dialogue model based on the differences and the evaluation results” can be implemented by following the steps 1041 to 1043 , which are described in detail below.

[0096] In step 1041, a quality evaluation value is calculated based on the probability value.

[0097] In some embodiments, a quality assessment value can be obtained by subtracting a probability value from a preset value. The preset value can be set according to actual usage requirements. In some embodiments, the preset value can be 1, and the quality assessment value can be obtained by subtracting the probability value from 1. As can be seen, when the probability value is higher, the quality of the first predicted reply data is higher, and the quality assessment value is smaller. When the probability value is lower, the quality of the first predicted reply data is lower, and the quality assessment value is larger.

[0098] In step 1042 , a first model loss is calculated based on the difference and the quality evaluation value.

[0099] After obtaining the quality assessment value, the difference can be added to the quality assessment value to obtain the first model loss. The higher the probability value, that is, the higher the quality of the first predicted reply data, the smaller the first model loss. The lower the probability value, that is, the lower the quality of the first predicted reply data, the higher the first model loss.

[0100] In some embodiments, the preset value is 1, and the loss function of the target dialogue model is as follows: Formula (1). The first model loss can be calculated by Formula (1). Formula (1) is as follows:

[0101] loss=outputloss+(1-prob) (1)

[0102] Wherein, loss is the first model loss, outputloss is the difference between the first response data and the first predicted response data, prob is the probability value, and 1-prob is the quality assessment value. Through step 1042, the accurate first model loss can be obtained.

[0103] In step 1043 , the model parameters of the target dialogue model are updated based on the first model loss.

[0104] After obtaining the first model loss, stochastic gradient descent can be used to update the model parameters of the target dialogue model until the first model loss falls below a preset loss threshold or the number of training cycles reaches a preset number of training cycles, and the target dialogue model converges. In some embodiments, during the update of the model parameters of the target dialogue model, the model parameters of the general dialogue model remain unchanged, and only the model parameters of the bypass model are updated. After obtaining the first model loss, stochastic gradient descent can be used to update the model parameters of the bypass model until the first model loss falls below a preset loss threshold or the number of training cycles reaches a preset number of training cycles, and the bypass model converges. This approach enables fine-tuning of the general dialogue model.

[0105] In related technologies, the conversation data of the target business scenario is used to update the general conversation model. However, during the training process, the conversation data of the target business scenario will affect the accuracy of the general conversation model, and the training cycle of the general conversation model is long, which is not suitable for target business scenarios with shorter target periods.

[0106] The training method for the dialogue model provided in the embodiment of the present application uses a bypass model as a bypass for the general dialogue model and uses dialogue data to train the bypass model. This makes it possible to fine-tune the general dialogue model without affecting the accuracy of the general dialogue model. In addition, the training cycle is short and can be applied to target business scenarios with shorter target periods.

[0107] The following combination Figure 6 To explain the training method of the speech evaluation model, see Figure 6 , Figure 6 This is a third flow chart of the method for training a dialogue model provided in an embodiment of the present application. Figure 3 Before step 103 shown, steps 105 to 107 may also be performed, which are described in detail below:

[0108] In step 105, N responses to be evaluated are obtained.

[0109] Wherein, N is a positive integer greater than or equal to 2. The N responses to be evaluated correspond to different speech qualities. In some embodiments, N responses to be evaluated can be obtained in a sequential order. The N responses to be evaluated are sorted by speech quality. That is, the order of the N responses to be evaluated corresponds to the speech quality of the N responses to be evaluated. The N responses to be evaluated can be sorted from high to low speech quality, or from low to high speech quality, which is both reasonable.

[0110] For example, if the speech quality is ranked from high to low, there may be three responses to be evaluated, and the order of the responses to be evaluated may be response 1 to be evaluated, response 2 to be evaluated, and response 3 to be evaluated. The speech quality of response 1 to be evaluated is higher than that of response 2 to be evaluated, and the speech quality of response 2 to be evaluated is higher than that of response 3 to be evaluated.

[0111] In some embodiments, see Figure 7 , Figure 7 This is a fourth flow chart of the method for training a dialogue model provided in an embodiment of the present application. Figure 6 The illustrated step 105 can be implemented by following the steps 1051 to 1052 , which are described in detail below.

[0112] In step 1051, sample data to be replied is obtained.

[0113] In some embodiments, for the same target business scenario, the sample data to be replied to may include a portion of the conversation data, and may also include other data that is not part of the conversation data. This means that there is overlap between the sample data to be replied to and the conversation data. In some embodiments, the sample data to be replied to may include data that is not part of the conversation data for the same target business scenario, which means that there is no overlap between the conversation data and the sample data to be replied to.

[0114] In step 1052, N dialogue models predict responses to the sample response data to obtain N responses to be evaluated.

[0115] The N responses to be evaluated have different speech qualities. The N pre-trained dialogue models have different parameter counts. The parameter count of a dialogue model is positively correlated with the speech quality of the output. In some embodiments, the bypass model parameters in the N pre-trained dialogue models differ. The larger the parameter count of the bypass model of a dialogue model, the higher the speech quality of the output response to be evaluated. The smaller the parameter count of the bypass model of a dialogue model, the lower the speech quality of the output response to be evaluated.

[0116] For example, if N is 3, the bypass model parameters for the three dialogue models can be 0.9 megabytes, 0.6 megabytes, and 0.3 megabytes, respectively. The dialogue model corresponding to 0.9 megabytes outputs the highest quality of the response to be evaluated, the dialogue model corresponding to 0.3 megabytes outputs the lowest quality of the response to be evaluated, and the dialogue model corresponding to 0.6 megabytes outputs the medium quality of the response to be evaluated.

[0117] In some embodiments, the target dialogue model to be trained is one of N dialogue models, and the number of parameters corresponding to the target dialogue model to be trained is greater than the number of parameters corresponding to the N-1 dialogue models. In some embodiments, the target dialogue model and the N-1 dialogue models have the same structure, differing only in the number of parameters in the bypass model. In other words, the target dialogue model is the model with the largest number of bypass model parameters among the N dialogue models.

[0118] In some embodiments, to obtain N responses to be evaluated in a sequential order, after obtaining the N responses to be evaluated, the N responses to be evaluated can be sorted by speech quality. Since the number of parameters of a dialogue model is positively correlated with the quality of the output speech, the N responses to be evaluated can be sorted according to the number of parameters of the dialogue model corresponding to the responses to be evaluated.

[0119] In some embodiments, the N responses to be evaluated can be sorted in descending order of speech quality. In some embodiments, the N responses to be evaluated can also be sorted in descending order of speech quality. In this way, the N responses to be evaluated can be accurately obtained in order.

[0120] Continue to see Figure 6 In step 106, the speech evaluation model evaluates the speech quality of the N responses to be evaluated and obtains N predicted evaluation results.

[0121] After obtaining the responses to be evaluated, each of these N responses can be fed into the speech evaluation model. The model then evaluates the speech quality of each of the N responses, generating N predicted evaluation results. These predicted evaluation results indicate the speech quality of the responses to be evaluated. The predicted evaluation results are the predicted probability that the responses to be evaluated meet the quality requirements, with the speech quality of the responses exceeding a preset quality threshold. Each response to be evaluated corresponds to one predicted evaluation result.

[0122] In some embodiments, the speech evaluation model may include a feature extraction layer and a mapping layer. The feature extraction layer is used to extract features from the input response to be evaluated and obtain a vector corresponding to the data. The mapping layer is used to map the vector corresponding to the response to be evaluated to obtain a predicted evaluation result.

[0123] In step 107, a second model loss is determined based on the N predicted evaluation results and the speech qualities of the N responses to be evaluated, and based on the second model loss, the model parameters of the speech evaluation model are updated.

[0124] In some embodiments, see Figure 8 , Figure 8 This is a fifth flow chart of the method for training a dialogue model provided in an embodiment of the present application. Figure 6The "determining the second model loss based on the N predicted evaluation results and the speech quality of the N responses to be evaluated" in step 107 can be achieved through the following steps 1071 to 1073, which are described in detail below.

[0125] In step 1071 , for each response to be evaluated, a first target value of the response to be evaluated is calculated based on the predicted probability values ​​corresponding to the M responses to be evaluated and the default target value.

[0126] In some embodiments, step 1071 can be implemented as follows: for each response to be evaluated, M responses to be evaluated corresponding to the response to be evaluated can be obtained, where the M responses to be evaluated are responses to be evaluated whose speech quality is less than the speech quality of the response to be evaluated among the N responses to be evaluated. Here, M is an integer greater than or equal to 0 and less than N, with a minimum value of 0 and a maximum value of N-1.

[0127] For example, the total number of responses to be evaluated is 3, including response 1 to be evaluated, response 2 to be evaluated, and response 3 to be evaluated. The quality of the response 1 to be evaluated is greater than the quality of the response 2 to be evaluated, and the quality of the response 2 to be evaluated is greater than the quality of the response 3 to be evaluated.

[0128] For response 1 to be evaluated, there are two responses to be evaluated whose speech quality is lower than that of response 1 to be evaluated. M can be 2, and two responses to be evaluated corresponding to response 1 to be evaluated can be obtained, namely, response 2 to be evaluated and response 3 to be evaluated.

[0129] For response 2, there is one response whose quality is lower than that of response 2. Therefore, M can be 1, and one response corresponding to response 2 can be obtained, namely response 3. For response 3, there is no response whose quality is lower than that of response 3. Therefore, M is 0.

[0130] In some embodiments, a bubble sort method can be used to sequentially determine the M responses to be evaluated corresponding to each response to be evaluated. After obtaining the M responses to be evaluated corresponding to each response to be evaluated, a first target value for each response to be evaluated can be calculated based on the predicted probability values ​​corresponding to the M responses to be evaluated. In some embodiments, the predicted probability values ​​corresponding to the M responses to be evaluated can be summed to obtain the first target value. In this way, an accurate first target value can be obtained.

[0131] In one embodiment, the first target value of the response to be evaluated is calculated based on the predicted probability values ​​corresponding to M responses to be evaluated and the default target value, including: when M is equal to 0, determining the default target value as the first target value of the response to be evaluated; when M is greater than 0 and less than N, summing the predicted probability values ​​corresponding to the M responses to be evaluated to obtain the first target value of the response to be evaluated.

[0132] In some embodiments, in the case where N replies to be evaluated are obtained and arranged in order, and the order of the N replies to be evaluated is from high to low in terms of speech quality, the predicted probability values ​​corresponding to the replies to be evaluated that are located after the reply to be evaluated are summed up according to the order of the N replies to be evaluated to obtain a first target value.

[0133] When the N responses to be evaluated are ranked from low to high in terms of speech quality, the predicted probability values ​​corresponding to the responses to be evaluated that follow the responses to be evaluated can be summed up according to the reverse order of the ranking of the N responses to be evaluated to obtain the first target value.

[0134] As an example of step 1071, assume that there are four responses to be evaluated, and the order of the responses to be evaluated is from high to low in terms of speech quality, specifically, response a to be evaluated, response b to be evaluated, response c to be evaluated, and response d to be evaluated. The corresponding predicted evaluation results are predicted evaluation value a, predicted evaluation value b, predicted evaluation value c, and predicted evaluation value d.

[0135] For response a to be evaluated, the first target value a is the sum of predicted evaluation values ​​b, c, and d. For response b to be evaluated, the first target value b is the sum of predicted evaluation values ​​c and d. For response c to be evaluated, the first target value c is the predicted evaluation value d. For response d to be evaluated, the first target value d is 0.

[0136] In some embodiments, the number of first target values ​​may be N, and the first target value corresponding to the reply to be evaluated with the lowest speech quality among the N replies to be evaluated is 0. In the case where the replies to be evaluated are sorted by speech quality, the first target value corresponding to the reply to be evaluated that is ranked last among the N replies to be evaluated is 0.

[0137] In step 1072, the predicted probability value of the response to be evaluated is subjected to a preset process, and a second target value is calculated based on the predicted probability value after the preset process and the first target value.

[0138] In some embodiments, step 1072 can be implemented by multiplying the predicted probability value by M for each response to be evaluated to obtain a predicted probability value after the processing, where M can be the number of responses to be evaluated whose speech quality is lower than the speech quality of the response to be evaluated, or the number of responses to be evaluated ranked after the response to be evaluated. A difference is calculated between the predicted probability value after the preset processing and the first target value to obtain a second target value.

[0139] As an example of step 1072, continuing with the example of step 1071, for response a to be evaluated, M is 3, and the second target value a corresponding to response a to be evaluated is 3 times the predicted evaluation value a minus the first target value a. For response b to be evaluated, M is 2, and the second target value b corresponding to response b to be evaluated is 2 times the predicted evaluation value b minus the first target value b. For response c to be evaluated, M is 1, and the second target value c corresponding to response c to be evaluated is the predicted evaluation value c minus the first target value c. For response d to be evaluated, M is 0, and the second target value d corresponding to response d to be evaluated is 0.

[0140] In some embodiments, the number of second target values ​​may be N, and the second target value corresponding to the response to be evaluated with the lowest speech quality is 0. In the case where the responses to be evaluated are sorted by speech quality, the second target value corresponding to the response to be evaluated that is ranked last among the N responses to be evaluated in the sorting order is 0.

[0141] In step 1073 , a second model loss is calculated based on the second target values ​​of the N responses to be used.

[0142] After obtaining the second target values ​​corresponding to N responses to be evaluated, the second target values ​​of the N responses to be used can be summed up, so that the sum of the second target values ​​can be obtained. In the process of training the speech evaluation model, the goal is to enable the speech evaluation model to output response data with high speech quality. Therefore, the sum of the second target values ​​needs to be maximized.

[0143] However, the gradient descent method is to minimize the loss, so the negative value of the sum of the second target values ​​can be determined as the second model loss. Through step 1073, the second model loss of the speech evaluation model can be accurately obtained.

[0144] After obtaining the second model loss, the model parameters of the speech evaluation model can be updated based on the second model loss until the second model loss is less than a preset loss threshold or the number of training times reaches a preset number of training times, and the speech evaluation model converges. Through step 107, the training of the speech evaluation model can be completed, and a more accurate speech evaluation model can be obtained.

[0145] The following combination Figure 9For an explanation of the training process of the dialogue model, see Figure 9 , Figure 9 This is the sixth flow chart of the method for training a dialogue model provided in an embodiment of the present application. The training steps of the dialogue model include steps 108 to 110, which are described in detail below.

[0146] In step 108 , N conversation training data are obtained.

[0147] For the same target business scenario, conversation training data can include a portion of conversation data and other data that is not part of the conversation data or sample data to be replied. This means that conversation data can overlap with conversation training data. For the same target business scenario, conversation training data can include data that is not part of the conversation data or sample data to be replied. This means that conversation training data can overlap with neither the conversation data nor the sample data to be replied.

[0148] The conversation training data includes second data (data to be replied) and second reply data for replying to the second data. A conversation training data set is composed of training conversation data from a target business scenario and training conversation data from a common business scenario. It should be noted that the training conversation data of the target business scenario may include data to be replied and reply data. Similarly, the training conversation data of a common business scenario may also include data to be replied and reply data. The data to be replied and reply data appear in pairs. The conversation training data mentioned here includes multiple second data, and accordingly, the number of second reply data is also multiple.

[0149] A set of conversation training data can be composed of training conversation data from a target business scenario and training conversation data from a common business scenario. This means that a set of conversation training data can include some pairs of data to be replied and reply data from the target business scenario, and some pairs of data to be replied and reply data from a common business scenario. For example, if the training conversation data for the target business scenario includes data to be replied 1 and reply data 1, and the training conversation data for the common business scenario includes data to be replied 2 and reply data 2, then the second data included in the set of conversation training data can be data to be replied 1 and data to be replied 2, and the second reply data can be reply data 1 and reply data 2.

[0150] For example, the number of training conversation data from the target business scenario in the N sets of conversation training data decreases sequentially. Each set of conversation training data has the same number of conversation data. The conversation training data may include three sets, specifically conversation training data 1, conversation training data 2, and conversation training data 3. The proportion of training conversation data from the target business scenario in conversation training data 1 is greater than the proportion of training conversation data from the target business scenario in conversation training data 2. In other words, the amount of training conversation data from the target business scenario in conversation training data 1 is greater than the amount of training conversation data from the target business scenario in conversation training data 2.

[0151] The proportion of training conversation data from the target business scenario in conversation training data 2 is greater than the proportion of training conversation data from the target business scenario in conversation training data 3. In other words, the amount of training conversation data from the target business scenario in conversation training data 2 is greater than the amount of training conversation data from the target business scenario in conversation training data 3.

[0152] In some embodiments, the parameter quantities of N dialogue models to be trained can be obtained, and based on the parameter quantities of the N dialogue models to be trained, the data ratio of the training dialogue data from the target business scenario and the training dialogue data from the general business scenario in each dialogue training data can be determined.

[0153] In some embodiments, the parameter amount of the bypass model in the N dialogue models to be trained can also be obtained, and based on the parameter amount of the N bypass models, the data ratio of the training dialogue data from the target business scenario and the training dialogue data from the general business scenario in each dialogue training data can be determined.

[0154] In some embodiments, a correspondence between candidate parameter quantities and candidate data ratios is pre-set, the parameter quantity of the bypass model is obtained, and based on the parameter quantity of the bypass model, the data ratio corresponding to the parameter quantity is determined from the correspondence between the candidate parameter quantity and the candidate data ratio.

[0155] For example, the parameters of the bypass model of the dialogue model can be 0.9 Mbps, 0.6 Mbps, and 0.3 Mbps, respectively. The number of dialogue training data in each set can be 600,000. For a set of dialogue training data corresponding to the dialogue model with a size of 0.9 Mbps, the ratio of training dialogue data from the target business scenario to training dialogue data from common business scenarios is 1:5. For a set of dialogue training data corresponding to the dialogue model with a size of 0.6 Mbps, the ratio of training dialogue data from the target business scenario to training dialogue data from common business scenarios is 2:4. For a set of dialogue training data corresponding to the dialogue model with a size of 0.3 Mbps, the ratio of training dialogue data from the target business scenario to training dialogue data from common business scenarios is 3:3. This approach allows for accurate data ratios to be obtained.

[0156] In some embodiments, based on the data ratio, data is extracted from the training dialogue data from the target business scenario and the training dialogue data from the common business scenario to obtain N sets of dialogue training data.

[0157] After obtaining the data ratio, data can be extracted from the training conversation data from the target business scenario and the training conversation data from the common business scenario according to the data ratio to obtain N sets of conversation training data. In other words, the conversation training data is composed of training conversation data from the target business scenario and training conversation data from the common business scenario in different ratios.

[0158] For example, after obtaining the data ratio, for the dialogue training data of the dialogue model corresponding to 0.9 terabytes, 100,000 training dialogue data can be extracted from the training dialogue data from the target business scenario, and 500,000 training dialogue data can be extracted from the training dialogue data from the common business scenario, to obtain the dialogue training data of the dialogue model corresponding to 0.9 terabytes.

[0159] For the dialogue training data of the dialogue model corresponding to 0.6 terabytes, 200,000 pieces of training dialogue data can be extracted from the training dialogue data from the target business scenario, and 400,000 pieces of training dialogue data can be extracted from the training dialogue data from the general business scenario to obtain the dialogue training data of the dialogue model corresponding to 0.6 terabytes.

[0160] To obtain 0.3 terabytes of conversation training data for the conversation model, we can extract 300,000 conversations from the training conversation data for the target business scenario and 300,000 conversations from the training conversation data for common business scenarios. This yields 0.3 terabytes of conversation training data for the conversation model. This approach allows us to accurately obtain the conversation training data needed to train the conversation model.

[0161] In step 109 , for each piece of dialogue training data, a dialogue model corresponding to the dialogue training data is determined, and the dialogue model predicts a reply for the second data in the dialogue training data to obtain second predicted reply data.

[0162] In some embodiments, the step of "determining the dialogue model corresponding to the dialogue training data" in step 109 can be implemented by calculating the data source ratio of the dialogue training data, where the data source ratio refers to the ratio of training dialogue data from the target business scenario to training dialogue data from common business scenarios in the dialogue training data. Based on the data source ratio, the dialogue model corresponding to the data source ratio is determined from a preset mapping relationship between the data source ratio and the dialogue model.

[0163] For example, the data source ratio of the training conversation data from the target business scenario to the training conversation data from the general business scenario is 1 to 5, the data source ratio of the training conversation data from the target business scenario to the training conversation data from the general business scenario is 2 to 4, and the data source ratio of the training conversation data from the target business scenario to the training conversation data from the general business scenario is 3 to 3.

[0164] If the data source ratio is 1:5, the pre-set mapping relationship between the data source ratio and the dialogue model can be used to determine that dialogue model a corresponds to a ratio of 1:5. The parameter size of the bypass model for dialogue model a is 0.9 megabytes. If the data source ratio is 2:4, the pre-set mapping relationship between the data source ratio and the dialogue model can be used to determine that dialogue model b corresponds to a ratio of 2:4. The parameter size of the bypass model for dialogue model b is 0.6 megabytes. If the data source ratio is 3:3, the pre-set mapping relationship between the data source ratio and the dialogue model can be used to determine that dialogue model c corresponds to a ratio of 3:3. The parameter size of the bypass model for dialogue model c is 0.3 megabytes. In this way, the dialogue model corresponding to the data source ratio can be determined by calculating the data source ratio.

[0165] In some embodiments, after obtaining and determining the dialogue model corresponding to the dialogue training data, for each dialogue model, the second data in the dialogue training data corresponding to the dialogue model can be input into the dialogue model, and the dialogue model can predict a reply for the second data in the dialogue training data to obtain second predicted reply data.

[0166] In step 110 , model parameters of the dialogue model are updated based on the difference between the second reply data and the second predicted reply data.

[0167] The dialogue model includes a general dialogue model and a bypass model. In some embodiments, the general dialogue model is a pre-trained model. During the dialogue model training process, the model parameters of the general dialogue model may not be updated, but only the model parameters of the bypass sub-model may be updated.

[0168] In some embodiments, see Figure 10 , Figure 10 This is a seventh flow chart of the method for training a dialogue model provided in an embodiment of the present application. Figure 9 The illustrated step 110 can be implemented by following steps 1101 and 1102 , which are described in detail below.

[0169] In step 1101 , the model parameters of the general dialogue model are fixed and remain unchanged.

[0170] The model parameters of the general dialogue model are fixed and remain unchanged. That is, the model parameters of the general dialogue model may not be updated based on the difference between the second reply data and the second predicted reply data. Through step 1101, the training dialogue data from the target business scenario will not affect the accuracy of the general dialogue model.

[0171] In step 1102 , model parameters of the bypass model are updated based on the difference between the second response data and the second predicted response data.

[0172] After obtaining the second predicted reply data, the bypass model's parameters can be updated based on the difference between the second reply data and the second predicted reply data. The bypass model converges, and thus the dialogue model, until the difference between the second reply data and the second predicted reply data is less than a preset loss threshold, or the number of training cycles reaches a preset number. Step 1102 completes the training of the bypass model, and thus the dialogue model, resulting in a more accurate dialogue model.

[0173] In some embodiments, the parameter quantity of the bypass model in the target dialogue model is greater than a preset parameter quantity threshold, and the dialogue training data corresponding to the target dialogue model includes training dialogue data from the target business scenario and training dialogue data from the ordinary business scenario, and the proportion of the training dialogue data from the target business scenario is less than the preset proportion threshold.

[0174] As an example, the parameters of the bypass model in the conversation model can be 0.9 Mbps, 0.6 Mbps, and 0.3 Mbps, respectively. The data source ratio for 0.9 Mbps is 1:5, the data source ratio for 0.6 Mbps is 2:4, and the data source ratio for 0.3 Mbps is 3:3. The parameters of the bypass model in the target conversation model are greater than 0.8 Mbps. The proportion of training conversation data from the target business scenario in the conversation training data corresponding to the target conversation model is less than 0.25. Therefore, the conversation model corresponding to 0.9 Mbps can be selected as the target conversation model.

[0175] Since the higher the number of parameters in the bypass model, the higher the accuracy of the bypass model, and the higher the number of parameters in the bypass model, the higher the proportion of training dialogue data from ordinary business scenarios in the corresponding dialogue training data, that is, the higher the quality of the dialogue training data, therefore, the higher the number of parameters in the dialogue model, the higher the corresponding accuracy, that is, the higher the number of parameters in the model, the higher the quality of the corresponding dialogue output by the response.

[0176] Accordingly, when executing step 1052, that is, executing N dialogue models to predict replies to the sample data to be replied to, and obtaining N replies to be evaluated, the higher the parameter amount, the higher the accuracy corresponding to the dialogue model, that is, the higher the quality of the reply to be evaluated output by the dialogue model. Therefore, the N replies to be evaluated can be sorted according to the parameter amount of the dialogue model corresponding to the reply to be evaluated.

[0177] See also Figure 11 , Figure 11 This is the eighth flow chart of the conversation model training method provided in the embodiments of the present application. As previously mentioned, the electronic device implementing the conversation model training method provided in the embodiments of the present application can be a terminal, a server, or a combination of the two. Below, the conversation model training method provided in the embodiments of the present application will be described using a server as an example. The server can execute steps 401-406.

[0178] In step 401, a general dialogue model is trained.

[0179] The general dialogue model can be a high-quality domain language model. The general dialogue model can be obtained by combining the general dialogue model and training dialogue data from common business scenarios through data screening, supervised fine-tuning training, and reinforcement training.

[0180] In step 402, N sets of dialogue training data are obtained.

[0181] Among them, one set of dialogue training data consists of training dialogue data from the target business scenario and training dialogue data from the common business scenario. The number of training dialogue data from the target business scenario in the N sets of dialogue training data decreases in sequence. Figure 9 The steps 108 shown are the same and will not be described in detail here.

[0182] In step 403, N dialogue models are trained based on N dialogue training data.

[0183] One conversation training data is used to train a conversation model. The conversation model includes a general conversation model and a bypass model. The difference between the N conversation models is that the parameters of the bypass model are different. Figure 9 Steps 109 and 110 are shown to be the same and are not described in detail here.

[0184] In step 404, sample data to be replied from the target business scenario is obtained, and replies to be evaluated of different qualities are obtained based on the sample data to be replied and N dialogue models. In step 405, a speech evaluation model is trained based on the replies to be evaluated of different qualities. Figure 6 Steps 106 and 107 are shown to be the same and will not be described in detail here.

[0185] In step 406, the first data and the first reply data in the target business scenario are obtained, the target dialogue model to be trained performs reply prediction on the first data to obtain the first predicted reply data, the speech evaluation model performs speech quality evaluation on the first predicted reply data to obtain the evaluation result, and based on the difference between the first reply data and the first predicted reply data, and based on the difference and the evaluation result, the parameters of the bypass model in the target dialogue model are adjusted. Figure 3 The steps shown are the same and will not be repeated here.

[0186] The following describes the training method for the dialogue model provided by the embodiment of the present application in conjunction with a specific application scenario. Taking question-answering using the target dialogue model as an example, the target business scenario is "Mid-Autumn Festival activities." The dialogue data is dialogue data related to "Mid-Autumn Festival activities."

[0187] The general conversation model is trained using training conversation data from common business scenarios. After acquiring the general conversation model, three conversation models can be trained. The target conversation data from the target business scenario includes 500,000 words (specifically, conversation data, sample reply data, and conversation training data), and the general conversation data also includes 500,000 words. The bypass model parameters for the three conversation models are 0.9, 0.6, and 0.3 megabytes, respectively.

[0188] The amount of dialogue training data can be 600,000. The dialogue training data for the dialogue model corresponding to 0.9 terabytes can include 100,000 training dialogue data from target business scenarios and 500,000 training dialogue data from common business scenarios. The dialogue training data for the dialogue model corresponding to 0.6 terabytes can include 200,000 training dialogue data from target business scenarios and 400,000 training dialogue data from common business scenarios. The dialogue training data for the dialogue model corresponding to 0.3 terabytes can include 300,000 training dialogue data from target business scenarios and 300,000 training dialogue data from common business scenarios.

[0189] 100,000 training conversations from target business scenarios and 500,000 training conversations from general business scenarios were used to train a 0.9-M conversation model. 200,000 training conversations from target business scenarios and 400,000 training conversations from general business scenarios were used to train a 0.6-M conversation model. 300,000 training conversations from target business scenarios and 300,000 training conversations from general business scenarios were used to train a 0.3-M conversation model.

[0190] After the training of the 0.9M corresponding dialogue model, the 0.6M corresponding dialogue model, and the 0.3M corresponding dialogue model is completed, see Figure 12 , Figure 12 It is a schematic diagram of the training speech quality scoring model provided in an embodiment of the present application.

[0191] exist Figure 12 In the example, three pre-trained dialogue models are included, and the parameters of the bypass models of the dialogue models are 0.9 megabytes, 0.6 megabytes, and 0.3 megabytes, respectively. The dialogue model corresponding to 0.9 megabytes may include an embedding layer 501, a general dialogue model 502, a bypass model 504, and an output layer 503.

[0192] The conversation model corresponding to 0.6 megabytes may include an embedding layer 505 , a general conversation model 506 , a bypass model 507 , and an output layer 508 . The conversation model corresponding to 0.3 megabytes may include an embedding layer 509 , a general conversation model 511 , a bypass model 510 , and an output layer 512 .

[0193] Obtain sample response data. Input the sample response data into three pre-trained dialogue models to obtain three responses with varying degrees of quality. The following describes the process of inputting the sample response data into the pre-trained dialogue models to obtain responses to be evaluated, using a 0.9-megabyte dialogue model as an example. The sample response data is input into the embedding layer 501, which maps each word in the response to be evaluated into a vector. The vectors output by the embedding layer 501 can then be input into the general dialogue model 502 and the bypass model 504, respectively.

[0194] The general conversation model 302 can encode the vector into a fixed-length vector and convert the fixed-length vector into a readable natural language output, generating general predicted response data. The bypass model 504 can reduce the dimensionality of the vector and then upscale the reduced features to their original dimensions, generating bypass predicted response data. The output layer 503 is used to fuse the general predicted response data with the bypass predicted response data to obtain the response to be evaluated corresponding to the 0.9-megabyte conversation model.

[0195] The process of obtaining the response to be evaluated from the sample response data for the 0.6 Mbps dialogue model and the 0.3 Mbps dialogue model is the same as that for the 0.9 Mbps dialogue model. For details, please refer to the corresponding instructions and will not be repeated here.

[0196] Among them, the speech quality of the response 1 to be evaluated output by the dialogue model corresponding to 0.9 megabytes is the highest, the speech quality of the response 3 to be evaluated output by the dialogue model corresponding to 0.3 megabytes is the lowest, and the speech quality of the response 2 to be evaluated output by the dialogue model corresponding to 0.6 megabytes is medium.

[0197] The three responses to be evaluated are respectively input into the speech evaluation model 513 to obtain three predicted evaluation results. The second model loss is determined based on the predicted evaluation results and the ranking of the three responses to be evaluated. Based on the second model loss, the model parameters of the speech evaluation model 513 are updated.

[0198] In some embodiments, for the response 1 to be evaluated, the first target value 1 is the sum of the predicted evaluation value 2 and the predicted evaluation value 3, M is 2, and the second target value 1 corresponding to the response 1 to be evaluated is 2 times the predicted evaluation value 1 minus the first target value 1.

[0199] For response 2 to be evaluated, the first target value 2 is the predicted evaluation value 3, M is 1, and the second target value 2 corresponding to response 2 to be evaluated is the predicted evaluation value 2 minus the first target value 2. For response 3 to be evaluated, M is 0, so the second target value 3 of response 3 to be evaluated is 0. The negative of the sum of the three second target values ​​is determined as the second model loss.

[0200] After completing the training of the dialogue evaluation model by generating responses to be evaluated through multiple dialogue models, the target dialogue model can be trained based on the trained dialogue evaluation model. Figure 13 , Figure 13 It is a schematic diagram of the training target dialogue model provided in an embodiment of the present application.

[0201] exist Figure 13 In the example, conversation data for a target business scenario is obtained, where the conversation data includes first data and first reply data in response to the first data. The first data is input into a target conversation model to be trained to obtain first predicted reply data. The target conversation model is used to predict reply data in response to the first data.

[0202] The target conversation model can be a conversation model corresponding to 0.9 terabytes, specifically including an embedding layer 501, a general conversation model 502, a bypass sub-model 504, and an output layer 503. The first predicted reply data is input into a pre-trained speech evaluation model 513 to obtain an evaluation result, which indicates the speech quality of the first predicted reply data. The difference between the first reply data and the first predicted reply data is obtained, and based on the difference and the evaluation result, the model parameters of the bypass model 504 are updated.

[0203] The following combination Figure 14 The data processing method provided in the embodiment of the present application is described in detail. Figure 14 , Figure 14 This is a flow chart of the data processing method provided in the embodiment of the present application. As mentioned above, the electronic device that implements the data processing method in the embodiment of the present application can be a terminal, a server, or a combination of the two. The data processing method provided in the embodiment of the present application is described below using a server as an example.

[0204] In step 1401, the data to be replied in the target business scenario is obtained.

[0205] The frequency of occurrence of the target business scenario is less than the threshold. The target business scenario can be set by the user based on actual usage needs and is valid during the target time period. For example, the target business scenario can be an activity corresponding to a holiday. Another example is an emergency situation. The frequency threshold (i.e., the threshold) can be set based on actual needs. For example, the frequency threshold can be 1, 2, 3, etc.

[0206] In step 1402, the target dialogue model predicts a response to the response data to obtain predicted response data.

[0207] The target dialogue model is trained based on the difference between the first reply data and the first predicted reply data and the evaluation results. The first predicted reply data is obtained by the target dialogue model to be trained to predict the reply to the first data, and the evaluation result is obtained by the pre-trained speech evaluation model to evaluate the speech quality of the first predicted reply data; the first data and the first reply data are included in the dialogue data under the target business scenario.

[0208] The target dialogue model training method can be found in the dialogue model training method provided in the embodiments of this application. The target dialogue model includes a general dialogue model and a bypass model. After the response data to be replied is input into the target dialogue model, both the general dialogue model and the bypass model can predict the response data corresponding to the response data to be replied. Then, based on the response data predicted by the general dialogue model and the bypass model, the predicted response data is obtained.

[0209] In some embodiments, the response data predicted by the general conversation model and the response data predicted by the bypass model can be fused to obtain predicted response data. In some embodiments, the response data predicted by the general conversation model and the response data predicted by the bypass model can be scored separately, and then weights can be determined based on the scoring results. The response data predicted by the general conversation model and the response data predicted by the bypass model can be fused based on the weights to obtain predicted response data. In this way, more accurate predicted response data can be obtained, and the predicted response data has higher speech quality.

[0210] The following continues to describe the exemplary structure of the dialogue model training device 255 provided in the embodiment of the present application as a software module. In some embodiments, such as Figure 2A As shown, the software modules stored in the dialog model training device 255 of the memory 240 may include:

[0211] The acquisition module 2551 is used to obtain conversation data under the target business scenario, where the occurrence frequency of the target business scenario is less than a threshold, and the conversation data includes first data and first reply data in response to the first data; wherein the acquisition module 2551 can also be called a first acquisition module.

[0212] Input module 2552, configured for the target dialogue model to be trained to predict a response to the first data to obtain first predicted response data;

[0213] The input module 2552 is further configured to use a pre-trained speech evaluation model to perform speech quality evaluation on the first predicted reply data to obtain an evaluation result;

[0214] The updating module 2553 is used to obtain the difference between the first reply data and the first predicted reply data, and update the model parameters of the target dialogue model based on the difference and the evaluation result.

[0215] In some embodiments, the evaluation result includes a probability value of the first predicted response data being a qualified speech. The updating module 2553 is further configured to calculate a quality assessment value based on the probability value; calculate a first model loss based on the difference and the quality assessment value; and update model parameters of the target dialogue model based on the first model loss.

[0216] In some embodiments, the acquisition module 2551 is further configured to acquire N responses to be evaluated, where N is a positive integer greater than or equal to 2.

[0217] The input module 2552 is further configured to use the speech evaluation model to perform speech quality evaluation on the N responses to be evaluated, thereby obtaining N predicted evaluation results, where each response to be evaluated corresponds to one predicted evaluation result.

[0218] The updating module 2553 is also used to determine the second model loss based on the N predicted evaluation results and the speech quality of the N responses to be evaluated, and to update the model parameters of the speech evaluation model based on the second model loss.

[0219] In some embodiments, the predicted evaluation result corresponding to a response to be evaluated includes a predicted probability value of the response to be evaluated as a qualified speech.

[0220] The updating module 2553 is also used to calculate the first target value of each reply to be evaluated based on the predicted probability values ​​corresponding to the M replies to be evaluated and the default target value; the M replies to be evaluated refer to the replies to be evaluated whose speech quality is less than the speech quality corresponding to the reply to be evaluated among the N replies to be evaluated; the M is an integer greater than or equal to 0 and less than the N; the predicted probability value of the reply to be evaluated is preset processed, and the second target value is calculated based on the predicted probability value after the preset processing and the first target value; the second model loss is calculated based on the second target values ​​of the N replies to be evaluated.

[0221] In some embodiments, the acquisition module 2551 is further configured to acquire sample data to be replied to. The input module 2552 is further configured to have N dialogue models predict responses to the sample data to be replied to, thereby obtaining N responses to be evaluated; the N responses to be evaluated have different speech qualities; the N dialogue models have different parameter quantities, and the parameter quantity of a dialogue model is positively correlated with the speech quality output.

[0222] In some embodiments, the acquisition module 2551 is further used to obtain N pieces of dialogue training data, where one piece of dialogue training data includes second data and second reply data that is a reply to the second data; one piece of the dialogue training data is composed of training dialogue data from the target business scenario and training dialogue data from a common business scenario; and the number of training dialogue data from the target business scenario in the N pieces of dialogue training data decreases successively.

[0223] Input module 2552 is further configured to, for each piece of conversation training data, determine a conversation model corresponding to the conversation training data, and to use the conversation model to predict a response to the second data in the conversation training data to obtain second predicted response data. Update module 2553 is further configured to update model parameters of the conversation model based on the difference between the second response data and the second predicted response data.

[0224] In some embodiments, the input module 2552 is further used to calculate the data source ratio of the dialogue training data, where the data source ratio refers to the ratio of the training dialogue data from the target business scenario to the training dialogue data from the general business scenario in the dialogue training data; based on the data source ratio, the dialogue model corresponding to the data source ratio is determined from the mapping relationship between the pre-set data source ratio and the dialogue model.

[0225] In some embodiments, the target dialogue model to be trained belongs to N dialogue models, and the number of parameters corresponding to the target dialogue model to be trained is greater than the number of parameters corresponding to N-1 dialogue models; the target dialogue model includes a general dialogue model and a bypass model; the general dialogue model is trained based on training dialogue data from the common business scenario; and updating the model parameters of the dialogue model refers to updating the parameters of the bypass model in the dialogue model.

[0226] The following continues to describe the exemplary structure of the data processing device 256 provided in the embodiment of the present application implemented as a software module. In some embodiments, such as Figure 2B As shown, the software modules stored in the conversation model training device 256 in the memory 240 may include: an acquisition module 2561 for acquiring data to be replied to in a target business scenario, where the occurrence frequency of the target business scenario is less than a threshold. The acquisition module 2561 may also be referred to as a second acquisition module.

[0227] Prediction module 2562 is used for the target dialogue model to predict the reply of the data to be replied to and obtain predicted reply data. The target dialogue model is trained based on the difference between the first reply data and the first predicted reply data and the evaluation result. The first predicted reply data is obtained by the target dialogue model to be trained to predict the reply of the first data. The evaluation result is obtained by the pre-trained speech evaluation model to evaluate the speech quality of the first predicted reply data. The first data and the first reply data are included in the dialogue data under the target business scenario.

[0228] An embodiment of the present application provides a computer program product, which includes computer-executable instructions. The computer-executable instructions are stored in a computer-readable storage medium. A processor of an electronic device reads the computer-executable instructions from the computer-readable storage medium, and the processor executes the computer-executable instructions, so that the electronic device executes the training method of the dialogue model described above in the embodiment of the present application, or the electronic device executes the data processing method described above in the embodiment of the present application.

[0229] The embodiment of the present application provides a computer-readable storage medium storing computer-executable instructions, wherein the computer-executable instructions or computer programs are stored. When the computer-executable instructions or computer programs are executed by a processor, the processor will execute the training method of the dialogue model provided in the embodiment of the present application, for example, Figure 3 Alternatively, when the computer executable instructions or computer program are executed by the processor, the processor will execute the data processing method provided in the embodiment of the present application, for example, Figure 14 The data processing method is shown.

[0230] In some embodiments, the computer-readable storage medium may be a memory such as RAM, ROM, flash memory, magnetic surface memory, optical disk, or CD-ROM; or various devices including one or any combination of the above memories. In some embodiments, the computer-executable instructions may be in the form of a program, software, software module, script, or code, written in any programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including as a standalone program or as a module, component, subroutine, or other unit suitable for use in a computing environment.

[0231] As an example, computer-executable instructions may, but need not, correspond to a file in a file system, may be stored as part of a file that stores other programs or data, such as in one or more scripts in a HyperText Markup Language (HTML) document, in a single file dedicated to the program in question, or in multiple coordinating files (e.g., files storing one or more modules, subroutines, or code portions).

[0232] By way of example, computer-executable instructions may be deployed to be executed on one electronic device, or on multiple electronic devices located at one site, or on multiple electronic devices distributed across multiple sites and interconnected by a communication network.

[0233] In summary, the embodiments of the present application enable the mixing of cleaned training conversation data from common business scenarios and uncleaned training conversation data from target business scenarios at varying ratios to generate multiple sets of training conversation data. Conversational models with larger parameter counts have a lower proportion of training conversation data from target business scenarios in their corresponding conversational training data. In other words, conversational models with larger parameter counts produce conversational training data with higher verbal quality, resulting in more accurate conversational models, meaning that the responses output by the conversational models are of higher verbal quality.

[0234] Conversational models with smaller parameters have a higher proportion of training data from the target business scenario in their corresponding conversational training data. In other words, conversational models with smaller parameters have lower speech quality in their corresponding conversational training data. This allows for training conversational models with lower accuracy. This allows for the generation of responses of varying quality to be evaluated, and ultimately, for training conversational evaluation models with high accuracy.

[0235] In the case of training the target dialogue model, the first predicted reply data can be input into the pre-trained speech evaluation model to obtain the evaluation result, and then the difference between the first reply data and the first predicted reply data can be obtained. Based on the difference and the evaluation result, the model parameters of the target dialogue model are updated. There is no need to clean the temporary dialogue data. Instead, the speech quality is evaluated through the speech evaluation model, which is equivalent to being able to achieve automatic evaluation and improve the accuracy of the target dialogue model. It can also reduce the manpower and time consumed by data cleaning and improve the efficiency of model training. Compared with the full training method in the related art, the training method of the dialogue model provided in the embodiment of the present application can improve the training efficiency. The data processing method provided in the embodiment of the present application can obtain predicted reply data with high speech quality and high accuracy.

[0236] The above description is merely an embodiment of the present application and is not intended to limit the scope of protection of the present application. Any modifications, equivalent replacements, and improvements made within the spirit and scope of the present application are included in the scope of protection of the present application.

Claims

1. A method for training a dialogue model, characterized in that: The method comprises: Acquire conversation data in a target business scenario, where the occurrence frequency of the target business scenario is less than a threshold, and the conversation data includes first data and first reply data in reply to the first data; The target dialogue model to be trained predicts a response to the first data to obtain first predicted response data; The pre-trained speech evaluation model performs speech quality evaluation on the first predicted response data to obtain an evaluation result; Obtain a difference between the first reply data and the first predicted reply data, and update model parameters of the target dialogue model based on the difference and the evaluation result.

2. The method according to claim 1, characterized in that The evaluation result includes a probability value of the first predicted response data being a qualified speech; The updating of the model parameters of the target dialogue model based on the difference and the evaluation result includes: Calculate a quality assessment value based on the probability value; Calculate a first model loss based on the difference and the quality evaluation value; Based on the first model loss, model parameters of the target dialogue model are updated.

3. The method according to claim 1, characterized in that The method further comprises: Obtain N responses to be evaluated, where N is a positive integer greater than or equal to 2; The speech evaluation model performs speech quality evaluation on the N responses to be evaluated to obtain N predicted evaluation results, where each response to be evaluated corresponds to one predicted evaluation result; A second model loss is determined based on the N predicted evaluation results and the speech quality of the N responses to be evaluated, and based on the second model loss, the model parameters of the speech evaluation model are updated.

4. The method according to claim 3, characterized in that The predicted evaluation result corresponding to a response to be evaluated includes a predicted probability value of the response to be evaluated as a qualified speech; The determining of the second model loss based on the N prediction evaluation results and the speech qualities of the N to-be-evaluated responses includes: For each response to be evaluated, a first target value of the response to be evaluated is calculated based on the predicted probability values ​​corresponding to M responses to be evaluated and the default target value; the M responses to be evaluated are responses to be evaluated whose speech quality is less than the speech quality corresponding to the response to be evaluated among the N responses to be evaluated; M is an integer greater than or equal to 0 and less than N; Performing a preset processing on the predicted probability value of the response to be evaluated, and calculating a second target value based on the predicted probability value after the preset processing and the first target value; A second model loss is calculated based on second target values ​​of the N responses to be evaluated.

5. The method according to claim 3, characterized in that The obtaining of N responses to be evaluated includes: Get sample data to be replied; N dialogue models predict responses to the sample data to be replied to, and obtain N responses to be evaluated; the N responses to be evaluated have different speech qualities; the number of parameters in the N dialogue models is different, and the number of parameters of a dialogue model is positively correlated with the output speech quality.

6. The method according to claim 5, characterized in that The method further comprises: Obtaining N sets of dialogue training data, where each set of dialogue training data includes second data and second reply data that is a reply to the second data; each set of dialogue training data is composed of training dialogue data from the target business scenario and training dialogue data from a common business scenario; wherein the number of training dialogue data from the target business scenario in the N sets of dialogue training data decreases in sequence; For each piece of dialogue training data, determining a dialogue model corresponding to the dialogue training data, and performing reply prediction on second data in the dialogue training data by the dialogue model to obtain second predicted reply data; Based on the difference between the second reply data and the second predicted reply data, the model parameters of the dialogue model are updated.

7. The method according to claim 6, characterized in that Determining the dialogue model corresponding to the dialogue training data includes: Calculating a data source ratio of the dialogue training data, where the data source ratio refers to a ratio of the training dialogue data from the target business scenario to the training dialogue data from the common business scenario in the dialogue training data; Based on the data source ratio, a dialogue model corresponding to the data source ratio is determined from a preset mapping relationship between the data source ratio and the dialogue model.

8. The method according to claim 6, characterized in that The target dialogue model to be trained belongs to N dialogue models, and the number of parameters corresponding to the target dialogue model to be trained is greater than the number of parameters corresponding to N-1 dialogue models; The target dialogue model includes a general dialogue model and a bypass model; the general dialogue model is trained based on training dialogue data from the common business scenario; and updating the model parameters of the dialogue model refers to updating the parameters of the bypass model in the dialogue model.

9. A data processing method, characterized in that: The method comprises: Obtaining data to be replied under a target business scenario, where the occurrence frequency of the target business scenario is less than a threshold; The target dialogue model predicts a reply for the data to be replied to and obtains predicted reply data. The target dialogue model is trained based on the difference between the first reply data and the first predicted reply data and the evaluation result. The first predicted reply data is obtained by the target dialogue model to be trained predicting a reply for the first data. The evaluation result is obtained by the pre-trained speech evaluation model performing a speech quality evaluation on the first predicted reply data. The first data and the first reply data are included in the dialogue data under the target business scenario.

10. A training device for a dialogue model, characterized in that: The device comprises: an acquisition module, configured to acquire conversation data in a target business scenario, where the occurrence frequency of the target business scenario is less than a threshold, and the conversation data includes first data and first reply data in reply to the first data; An input module, configured for the target dialogue model to be trained to predict a response to the first data to obtain first predicted response data; The input module is further configured to use a pre-trained speech evaluation model to perform speech quality evaluation on the first predicted reply data to obtain an evaluation result; An updating module is used to obtain a difference between the first reply data and the first predicted reply data, and update the model parameters of the target dialogue model based on the difference and the evaluation result.

11. A data processing device, characterized in that: The device comprises: An acquisition module, configured to acquire data to be replied under a target business scenario, wherein the occurrence frequency of the target business scenario is less than a threshold; A prediction module is used for the target dialogue model to predict the reply of the data to be replied to and obtain predicted reply data. The target dialogue model is trained based on the difference between the first reply data and the first predicted reply data and the evaluation result. The first predicted reply data is obtained by the target dialogue model to be trained to predict the reply of the first data. The evaluation result is obtained by the pre-trained speech evaluation model to evaluate the speech quality of the first predicted reply data. The first data and the first reply data are included in the dialogue data under the target business scenario.

12. An electronic device, characterized in that: The electronic device comprises: a memory for storing computer-executable instructions; A processor, configured to implement the method for training a dialogue model according to any one of claims 1 to 8, or the data processing method according to claim 9, when executing computer-executable instructions stored in the memory.

13. A computer-readable storage medium storing computer-executable instructions or a computer program, characterized in that: When the computer-executable instructions or computer program are executed by a processor, the method for training a dialogue model according to any one of claims 1 to 8 is implemented, or the data processing method according to claim 9 is implemented.

14. A computer program product comprising computer executable instructions or a computer program, characterized in that When the computer-executable instructions or computer program are executed by a processor, the method for training a dialogue model according to any one of claims 1 to 8 is implemented, or the data processing method according to claim 9 is implemented.