Information processing method, model training method, device, equipment, medium, and program product
By training dialogue models with corrected and diverse response samples, the issues of low accuracy and poor quality in current dialogue systems are addressed, resulting in improved dialogue accuracy and quality.
Patent Information
- Application Number
- JP2023048430
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2022-08-10
- Filing Date
- 2023-03-24
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2043-03-24
AI Technical Summary
Current dialogue models in smart dialogue systems suffer from low dialogue accuracy and poor dialogue quality due to discrepancies between social media comment data and real human dialogue scenarios.
A dialogue model is trained using corrected response sample sentences, second candidate response sample sentences, and recall response sample sentences, which are used to refine and improve the initial dialogue model, resulting in higher dialogue accuracy and quality.
The proposed solution enhances dialogue accuracy and quality by continuously training the dialogue model with high-quality response samples, ensuring that the generated responses better align with human-like dialogue.
Smart Images

Figure 0007689541000003 
Figure 0007689541000004 
Figure 0007689541000005
Abstract
Description
[Technical field]
[0001] The present disclosure relates to the field of computer technology, in particular to the field of artificial intelligence and speech technology, and specifically to information processing methods, model training methods, devices, equipment, media and program products. [Background technology]
[0002] With the development of natural language processing technology, machine learning models can be used in the field of smart dialogue, in which the dialogue model responds based on the sentences entered by the user, achieving the effect of dialogue with the user.
[0003] Currently, the dialogue model has low dialogue accuracy and poor dialogue quality. Summary of the Invention [Problem to be solved by the invention]
[0004] The present disclosure provides an information processing method, a model training method, an apparatus, a device, a medium, and a program product. [Means for solving the problem]
[0005] According to one aspect of the present disclosure, there is provided an information processing method, the method comprising: obtaining an initial dialogue; inputting the initial dialogue sentence into a trained dialogue model to obtain a target response sentence; The dialogue model is a model obtained by training based on a corrected response sample sentence, a second candidate response sample sentence, and a recall response sample sentence, and an initial dialogue sample sentence is input into an initial dialogue model to obtain a plurality of candidate response sample sentences, the second candidate response sample sentence being any one of the plurality of candidate response sample sentences, the corrected response sample sentence being a sentence obtained by correcting a first response sample sentence among the candidate response sample sentences, and the recall response sample sentence being a sample sentence other than the initial dialogue sample sentence and the plurality of candidate response sample sentences among the training sample sentences.
[0006] According to another aspect of the present disclosure, there is provided a model training method, the method comprising: obtaining an initial dialogue sample sentence; inputting the initial dialogue sample sentences into an initial dialogue model to obtain a plurality of candidate reply sample sentences; modifying a first candidate reply sample sentence from the plurality of candidate reply sample sentences to obtain a modified reply sample sentence; training the initial dialogue model based on the revised response sample sentence, a second candidate response sample sentence of the plurality of candidate response sample sentences, and a recall response sample sentence to obtain a dialogue model; The recall response sample sentences are sample sentences other than the initial dialogue sample sentences and the plurality of candidate response sample sentences, among the training sample sentences.
[0007] According to another aspect of the present disclosure, there is provided an information processing device, the device comprising: an acquisition module for acquiring an initial dialogue; an input module for inputting the initial dialogue sentence into a trained dialogue model to obtain a target response sentence; The dialogue model is a model obtained by training based on a corrected response sample sentence, a second candidate response sample sentence, and a recall response sample sentence, and an initial dialogue sample sentence is input into an initial dialogue model to obtain a plurality of candidate response sample sentences, the second candidate response sample sentence being any one of the plurality of candidate response sample sentences, the corrected response sample sentence being a sentence obtained by correcting a first response sample sentence among the candidate response sample sentences, and the recall response sample sentence being a sample sentence other than the initial dialogue sample sentence and the plurality of candidate response sample sentences among the training sample sentences.
[0008] According to another aspect of the present disclosure, there is provided a model training apparatus, the apparatus comprising: a sentence acquisition module for acquiring an initial dialogue sample sentence; a sentence input module for inputting the initial dialogue sample sentence into an initial dialogue model to obtain a plurality of candidate reply sample sentences; a correction module that corrects a first candidate reply sample sentence from the plurality of candidate reply sample sentences to obtain a corrected reply sample sentence; a training module for training the initial dialogue model based on the revised response sample sentence, a second candidate response sample sentence among the plurality of candidate response sample sentences, and a recall response sample sentence to obtain a dialogue model; The recall response sample sentences are sample sentences other than the initial dialogue sample sentences and the plurality of candidate response sample sentences, among the training sample sentences.
[0009] According to another aspect of the present disclosure, there is provided an electronic device, the electronic device comprising: At least one processor; a memory communicatively coupled to the at least one processor; The memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor such that the at least one processor performs the above-described methods.
[0010] According to another aspect of the present disclosure, a non-transitory computer readable storage medium having computer instructions stored thereon, the computer instructions causing a computer to perform the above-described method.
[0011] According to another aspect of the present disclosure, a computer program product, which when executed by a processor, implements the steps of the above method. Effect of the Invention
[0012] In some embodiments of the present disclosure, training is performed based on the corrected response sample sentences, the second candidate response sample sentences, and the recall response sample sentences to obtain a dialogue model, and the initial dialogue sample sentences are input into the initial dialogue model to obtain a plurality of candidate response sample sentences, the second candidate response sample sentence being any one of the plurality of candidate response sample sentences, the corrected response sample sentence being a high quality dialogue sentence obtained by correcting the first response sample sentence among the candidate response sample sentences, and the recall response sample sentence being a sample sentence among the training sample sentences other than the initial dialogue sample sentences and the plurality of candidate response sample sentences, and the initial dialogue model is continued to be trained on the corrected response sample sentences, the second candidate response sample sentences, and the recall response sample sentences to obtain a dialogue model with high dialogue accuracy, and the initial dialogue sentences are input into the dialogue model to obtain a target response sentence with high dialogue quality.
[0013] It should be understood that the contents described in this section are not intended to identify key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will be readily understood from the following description. [Brief description of the drawings]
[0014] The drawings are used for a better understanding of the present technical solution, and are not intended to limit the present disclosure. [Figure 1]1 is a schematic flowchart of an information processing method provided by Example 1 of the present disclosure. [Diagram 2] 1 is a schematic flowchart of a model training method provided by Example 2 of the present disclosure. [Diagram 3] 11 is a flowchart of an information processing method provided by Example 3 of the present disclosure. [Figure 4] FIG. 1 is a schematic configuration diagram of an information processing device provided by an exemplary embodiment of the present disclosure. [Diagram 5] FIG. 1 is a schematic configuration diagram of a model training apparatus provided by an exemplary embodiment of the present disclosure. [Figure 6] FIG. 1 is a schematic block diagram of an exemplary electronic device for implementing embodiments of the present disclosure. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0015] Hereinafter, exemplary embodiments of the present disclosure will be described in conjunction with the drawings, and various details of the embodiments of the present disclosure will be included therein for ease of understanding, and they should be regarded as merely exemplary. Therefore, it should be recognized that those skilled in the art can make various changes and modifications to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for clarity and conciseness, the following description will omit the description of well-known functions and structures.
[0016] In addition, in the technical solution disclosed herein, the collection, storage, use, processing, transmission, provision, disclosure and other processing of relevant user personal information shall all comply with the provisions of relevant laws and regulations and shall not violate public order and morals.
[0017] Artificial intelligence is a field that studies how computers can simulate certain human thought processes and intelligent behaviors (learning, reasoning, thinking, planning, etc.), and includes both hardware and software level technologies. AI hardware technology generally includes sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, etc. AI software technology mainly includes several directions, such as computer vision technology, voice recognition technology, natural language processing technology, machine learning / deep learning, big data processing technology, and knowledge graph technology.
[0018] With the development of natural language processing technology, machine learning models can be used in the field of smart dialogue, in which the dialogue model responds based on the sentences entered by the user, achieving the effect of dialogue with the user.
[0019] In the field of dialogue systems, large-scale dialogue models trained on social media comment data have emerged one after another, but due to the discrepancy between social media comment scenes and real human dialogue scenes, the model's generation ability is poor.
[0020] A generative dialogue model generates multiple candidate replies during inference and then evaluates and sorts the replies using a generation score. However, the sorting method based on the generation score cannot effectively bring high-quality replies to the front.
[0021] Currently, the dialogue model has low dialogue accuracy and poor dialogue quality.
[0022] In response to the above-mentioned technical problems, in some embodiments of the present disclosure, a dialogue model is obtained by training based on the corrected response sample sentence, the second candidate response sample sentence, and the recall response sample sentence, and the initial dialogue sample sentence is input into the initial dialogue model to obtain a plurality of candidate response sample sentences, the second candidate response sample sentence being any one of the plurality of candidate response sample sentences, the corrected response sample sentence being a sentence with high dialogue quality obtained by correcting the first reply sample sentence among the candidate response sample sentences, and the recall response sample sentence being a sample sentence other than the initial dialogue sample sentence and the plurality of candidate response sample sentences among the training sample sentences, and the initial dialogue model is continued to be trained on the corrected response sample sentence, the second candidate response sample sentence, and the recall response sample sentence to obtain a dialogue model with high dialogue accuracy, and the initial dialogue sentence is input into the dialogue model to obtain a target response sentence with high dialogue quality.
[0023] Below, the technical solutions provided by the embodiments of the present disclosure will be described in detail in conjunction with the drawings.
[0024] 1 is a schematic flowchart of an information processing method provided by the first embodiment of the present disclosure. As shown in FIG 1, the method includes the following steps S101 to S102.
[0025] S101, an initial dialogue is obtained. S102, input the initial dialogue sentence into a trained dialogue model to obtain a target response sentence. The dialogue model is a model obtained by training based on the corrected response sample sentence, the second candidate response sample sentence and the recall response sample sentence, and the initial dialogue sample sentence is input into the initial dialogue model to obtain a plurality of candidate response sample sentences, the second candidate response sample sentence being any one of the plurality of candidate response sample sentences, the corrected response sample sentence being a sentence obtained by correcting a first response sample sentence among the candidate response sample sentences, and the recall response sample sentence being a sample sentence other than the initial dialogue sample sentence and the plurality of candidate response sample sentences among the training sample sentences.
[0026] In this embodiment, the execution subject of the above method may be a server or a terminal device.
[0027] When the execution subject of the above method is a server, the implementation form of the server is not limited. For example, the server may be a server device such as a general-purpose server, a cloud server, a cloud host, a virtual center, etc. The configuration of the server mainly includes a processor, a hard disk, a memory, a system bus, etc., and a type of general-purpose computer architecture.
[0028] When the execution subject of the above method is a terminal device, the implementation form of the terminal device is not limited. The terminal device includes, but is not limited to, a personal computer, a tablet computer, a smartphone, or a smart wearable device.
[0029] In this embodiment, training is performed based on the corrected response sample sentence, the second candidate response sample sentence, and the recall response sample sentence to obtain a dialogue model, and the initial dialogue sample sentence is input into the initial dialogue model to obtain a plurality of candidate response sample sentences, the second candidate response sample sentence being any one of the plurality of candidate response sample sentences, the corrected response sample sentence being a high quality dialogue sentence obtained by correcting the first response sample sentence among the candidate response sample sentences, and the recall response sample sentence being a sample sentence among the training sample sentences excluding the initial dialogue sample sentence and the plurality of candidate response sample sentences, and by continuing to train the initial dialogue model on the corrected response sample sentence, the second candidate response sample sentence, and the recall response sample sentence, a dialogue model with high dialogue accuracy is obtained, an initial dialogue sentence is obtained, and the initial dialogue sentence is input into the dialogue model to obtain a target response sentence with high dialogue quality.
[0030] Below, the technical proposals of the present disclosure will be explained according to application scenarios.
[0031] Application scenario 1: The smartphone responds to an initial dialogue sentence "How is the weather today?" input by the user via voice, the smartphone uploads the initial dialogue sentence to the server, the server inputs the initial dialogue sentence into a trained dialogue model to obtain the target response sentence "It's sunny today", the server sends the target response sentence down to the smartphone, and the smartphone plays the target response sentence "It's sunny today" via voice.
[0032] Application scene 2: The smartphone responds to an initial dialogue sentence “How is the weather today?” input by voice by the user, and the smartphone inputs the initial dialogue sentence into a locally integrated dialogue model to obtain a target response sentence “It is sunny today”, and the smartphone plays the target response sentence “It is sunny today” by voice.
[0033] Before using the dialogue model, we need to train an initial dialogue model to obtain a dialogue model. The process of training the dialogue model is described below.
[0034] 2 is a schematic flowchart of a model training method provided by the second embodiment of the present disclosure. As shown in FIG. 2, the method includes the following steps S201 to S204.
[0035] S201, an initial dialogue sample sentence is obtained.
[0036] S202, input the initial dialogue sample sentence into an initial dialogue model to obtain a plurality of candidate response sample sentences.
[0037] S203, a first candidate reply sample sentence among the plurality of candidate reply sample sentences is corrected to obtain a corrected reply sample sentence.
[0038] S204, training an initial dialogue model based on the revised response sample sentence, the second candidate response sample sentence among the plurality of candidate response sample sentences and the recall response sample sentence to obtain a dialogue model. The recall response sample sentences are the training sample sentences other than the initial dialogue sample sentences and the plurality of candidate response sample sentences.
[0039] The training device for training the above dialogue model may be any type of computer device, and the embodiments of the present disclosure are not limited thereto.
[0040] Note that the initial dialogue model may be a trained model, and the accuracy of the initial dialogue model is low, and the quality of the dialogue using the initial dialogue model is poor.
[0041] Obtain an initial dialogue sample sentence, input the initial dialogue sample sentence into an initial dialogue model to obtain a revised response sample sentence, modify a first candidate response sample sentence among the plurality of candidate response sample sentences to obtain a revised response sample sentence, randomly select a second candidate response sample sentence among the plurality of candidate response sample sentences, and select a recall response sample sentence from the other sample sentences among the training sample sentences except for the initial dialogue sample sentence and the plurality of candidate response sample sentences. The revised response sample sentence, the second candidate response sample sentence, and the recall response sample sentence constitute one training dataset. The above steps are repeated to obtain a training dataset for model training.
[0042] In addition, in order to increase the coverage range of the dataset, the initial dialogue sample sentences are adopted from datasets in as many different fields as possible, such as the news field, the social media field, the literature field, and the live-action dialogue field.
[0043] In the above embodiment, a first candidate reply sample sentence among the plurality of candidate reply sample sentences is modified to obtain a modified reply sample sentence. For example, the first candidate reply sample sentence is copied, corrected, or created to obtain a modified reply sample sentence.
[0044] For example, in response to an operation of inputting an initial dialogue sample sentence in the labeling interface, the initial dialogue sample sentence "It rains every day and makes me feel bad" is obtained, and the initial dialogue sample sentence is input into the initial dialogue model to obtain a number of candidate response sample sentences, such as "Music and chocolate go well with rainy days", "Rainy days are perfect for sleeping", "I feel bad too because no one will join me", "Rainy days are nice", "Me too! I don't like rainy days", "Yes, it's a problem that I can't go out", and "Yes, I hate rainy days too".
[0045] A first candidate response sample sentence "Music and chocolate go well with rainy days" from among the multiple candidate response sample sentences is modified to obtain a modified response sample sentence "I think music and chocolate go well with rainy days", a second candidate response sample sentence "Rainy days are perfect for sleeping" is randomly selected from the multiple candidate response sample sentences, and a recall response sample sentence "It's sunny today" is selected from other sample sentences excluding the initial dialogue sample sentence and the multiple candidate response sample sentences from the training sample sentences. The modified response sample sentence "Music and chocolate go well with rainy days", the second candidate response sample sentence "Rainy days are perfect for sleeping", and the recall response sample sentence "It's sunny today" constitute one training dataset.
[0046] In the above embodiment, the dialogue model is obtained by training the initial dialogue model based on the corrected response sample sentence, the second candidate response sample sentence and the recall response sample sentence among the plurality of candidate response sample sentences. In one possible embodiment, the corrected response sample sentence, the second candidate response sample sentence and the recall response sample sentence are initial Input into a sentence generation model to obtain the probability of the actual response sentence, the probability of the corrected response sample sentence, the probability of the second candidate response sample sentence, and the probability of the recall response sample sentence, and jointly train the initial sentence generation model and the initial sentence determination model of the initial dialogue model based on the probability of the actual response sentence, the probability of the corrected response sample sentence, the probability of the second candidate response sample sentence, and the probability of the recall response sample sentence to obtain a dialogue model.
[0047] In one embodiment, an initial sentence generation model and an initial sentence determination model of an initial dialogue model are jointly trained based on the actual response sentence, the probability of the corrected response sample sentence, the probability of the second candidate response sample sentence, and the probability of the recall response sample sentence to obtain a dialogue model. A loss function is determined based on the actual response sentence and the corrected response sample sentence, and based on the loss function, the probability of the corrected response sample sentence is greater than the probability of the second candidate response sample sentence, the probability of the corrected response sample sentence is greater than the probability of the recall response sample sentence, and the probability of the second candidate response sample sentence is greater than the probability of the recall response sample sentence to obtain a dialogue model.
[0048] TIFF0007689541000001.tif28167
[0049] TIFF0007689541000002.tif35167
[0050] In conjunction with the above description of the embodiments, Fig. 3 is a flowchart of an information processing method provided by the third embodiment of the present disclosure. As shown in Fig. 3, the method includes the following steps S301 to S304.
[0051] S301: The terminal device responds to a voice input operation and obtains an initial dialogue.
[0052] S302, the terminal device transmits an initial dialogue to the server.
[0053] S303, the server receives the initial dialogue, inputs the initial dialogue into a dialogue model to obtain a target response sentence, and transmits the target response sentence down to the terminal device.
[0054] S304, the terminal device receives the target response sentence and plays the target response sentence by voice.
[0055] In this embodiment, the implementation form of the server is not limited. For example, the server may be a server device such as a general-purpose server, a cloud server, a cloud host, a virtual center, etc. The configuration of the server mainly includes a processor, a hard disk, a memory, a system bus, etc., and a general-purpose computer architecture type.
[0056] In this embodiment, the implementation form of the terminal device is not limited. The terminal device includes, but is not limited to, a personal computer, a tablet computer, a smartphone, or a smart wearable device.
[0057] The implementation form of each step in this embodiment can refer to the description of the above embodiments, and the description will be omitted in this embodiment, while at the same time, this embodiment can obtain the beneficial effects of the parts corresponding to each of the above embodiments.
[0058] 4 is a schematic diagram of an information processing device 40 provided by an exemplary embodiment of the present disclosure. The information processing device 40 includes an acquisition module 41 and an input module 42.
[0059] The acquisition module 41 acquires an initial dialogue.
[0060] The input module 42 inputs the initial dialogue sentence into the trained dialogue model to obtain the target response sentence. The dialogue model is a model obtained by training based on the corrected response sample sentence, the second candidate response sample sentence and the recall response sample sentence, and the initial dialogue sample sentence is input into the initial dialogue model to obtain a plurality of candidate response sample sentences, the second candidate response sample sentence being any one of the plurality of candidate response sample sentences, the corrected response sample sentence being a sentence obtained by correcting a first response sample sentence among the candidate response sample sentences, and the recall response sample sentence being a sample sentence other than the initial dialogue sample sentence and the plurality of candidate response sample sentences among the training sample sentences.
[0061] Optionally, the input module 42 inputs the initial dialogue sentence into the trained dialogue model to obtain the target response sentence, Within the dialogue model, inputting the initial dialogue sentence into a sentence generation model of the dialogue model to obtain a plurality of candidate response sentences and a probability of each candidate response sentence; A plurality of candidate response sentences and the probability of each candidate response sentence are input to a sentence determination model of the dialogue model to obtain a target response sentence.
[0062] Optionally, the input module 42 inputs a plurality of candidate response sentences and a probability of each candidate response sentence into a sentence determination model of the dialogue model to obtain a target response sentence, A plurality of candidate reply sentences and the probability of each candidate reply sentence are input to a sentence determination model, and a target reply sentence with the highest probability is selected from the plurality of candidate reply sentences.
[0063] 5 is a schematic diagram of a model training device 50 provided by an exemplary embodiment of the present disclosure. The model training device 50 includes a sentence acquisition module 51, a sentence input module 52, a correction module 53, and a training module 54. The sentence acquisition module 51 acquires an initial dialogue sample sentence, The sentence input module 52 inputs the initial dialogue sample sentence into the initial dialogue model to obtain a plurality of candidate reply sample sentences; The correction module 53 corrects a first candidate reply sample sentence from the plurality of candidate reply sample sentences to obtain a corrected reply sample sentence; The training module 54 trains an initial dialogue model based on the revised response sample sentence, the second candidate response sample sentence of the plurality of candidate response sample sentences, and the recall response sample sentence to obtain a dialogue model; The recall response sample sentences are sample sentences other than the initial dialogue sample sentences and the plurality of candidate response sample sentences among the training sample sentences.
[0064] Optionally, the training module 54 trains the initial dialogue model based on the revised response sample sentence, the second candidate response sample sentence of the plurality of candidate response sample sentences, and the recall response sample sentence to obtain the dialogue model, The revised response sample sentence, the second candidate response sample sentence, and the recall response sample sentence are initial inputting the actual response sentence, the probability of the corrected response sample sentence, the probability of the second candidate response sample sentence, and the probability of the recall response sample sentence into a sentence generation model; An initial sentence generation model and an initial sentence determination model of the initial dialogue model are jointly trained based on the probability of the actual response sentence, the probability of the corrected response sample sentence, the probability of the second candidate response sample sentence, and the probability of the recall response sample sentence to obtain a dialogue model.
[0065] Optionally, the training module 54 jointly trains the initial sentence generation model and the initial sentence determination model of the initial dialogue model according to the actual response sentence, the probability of the corrected response sample sentence, the probability of the second candidate response sample sentence, and the probability of the recall response sample sentence to obtain a dialogue model; determining a loss function based on the actual response sentence and the sample corrected response sentence; Based on the loss function, the probability of the corrected response sample sentence is greater than the probability of the second candidate response sample sentence, the probability of the corrected response sample sentence is greater than the probability of the recall response sample sentence, and the probability of the second candidate response sample sentence is greater than the probability of the recall response sample sentence, and the initial sentence generation model and the initial sentence determination model are jointly trained to obtain a dialogue model.
[0066] The specific manner in which each module of the apparatus in the above embodiment performs the operation has already been described in detail in the embodiment relating to the method, and will not be described in detail here.
[0067] According to an embodiment of the present disclosure, the present disclosure further provides an electronic device and a readable storage medium. According to an embodiment of the present disclosure, the present disclosure further provides a computer program, which, when executed by a processor, realizes the information processing method or the model training method provided by the present disclosure.
[0068] 6 is a schematic block diagram of an exemplary electronic device 600 for implementing embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processing, mobile phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the description herein and / or the implementation of the present disclosure as sought.
[0069] 6, the electronic device 600 includes a computing unit 601 that can perform various appropriate operations and processes according to a computer program stored in a read-only memory (ROM) 602 or loaded from a storage unit 608 into a random access memory (RAM) 603. The RAM 603 may also store various programs and data required for the operation of the electronic device 600. The computing unit 601, the ROM 602, and the RAM 603 are connected to each other via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.
[0070] The components of the electronic device 600 are connected to an I / O interface 605, including input units 606 such as a keyboard, a mouse, etc., output units 607 such as various types of displays, speakers, etc., storage units 608 such as a magnetic disk, an optical disk, etc., and a communication unit 609 such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 609 enables the electronic device 600 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0071] The computing unit 601 may be various general-purpose and / or dedicated processing components having processing and computing capabilities. Some examples of the computing unit 601 include, but are not limited to, a central processing unit (CPU), a graphic processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units that execute machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 601 executes each of the methods and processes described above, such as the information processing method and the model training method. For example, in some embodiments, the information processing method and the model training method can be realized as a computer software program tangibly included in a machine-readable medium, such as the storage unit 608. In some embodiments, some or all of the computer program can be loaded and / or installed in the electronic device 600 via the ROM 602 and / or the communication unit 609. When the computer program is loaded into the RAM 603 and executed by the computing unit 601, one or more steps of the information processing method and the model training method described above may be performed. Alternatively, in other embodiments, the computing unit 601 may be configured in any other suitable manner (eg, via firmware) to perform the information processing methods and the model training methods.
[0072] Various embodiments of the systems and techniques described herein above may be implemented in digital electronic circuitry systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include being embodied in one or more computer programs that may be executed and / or interpreted by a programmable system that includes at least one programmable processor, which may be an application specific or general purpose programmable processor, and that may receive data and instructions from and transmit data and instructions to a storage system, at least one input device, and at least one output device.
[0073] The program codes for carrying out the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general purpose computer, a special purpose computer, or other programmable data processing apparatus such that, when executed by the processor or controller, the functions / operations specified in the flowcharts and / or block diagrams are performed. The program codes may be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine, partially on a remote machine, or entirely on a remote machine or server.
[0074] In the context of the present disclosure, a machine-readable medium may be a tangible medium that can contain or store a program for use by or in combination with an instruction execution system, apparatus, or device. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the above contents. More specific examples of machine-readable storage media include one or more line-based electrical connections, portable computer disks, hard disks, random access memories (RAMs), read-only memories (ROMs), erasable programmable read-only memories (EPROMs or flash memories), optical fibers, portable compact disk read-only memories (CD-ROMs), optical storage devices, magnetic storage devices, or any suitable combination of the above contents.
[0075] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to a user, and a keyboard and pointing device (e.g., a mouse or trackball) by which a user can provide input to the computer. Other types of devices can also provide interaction with a user, for example, the feedback provided to the user can be any form of sensing feedback (e.g., vision feedback, auditory feedback, or haptic feedback) and can receive input from the user in any form (including acoustic, speech, or tactile input).
[0076] The systems and techniques described herein may be implemented in a computing system including a back-end component (e.g., a data server), or a computing system including a middleware component (e.g., an application server), or a computing system including a front-end component (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with the embodiments of the systems and techniques described herein), or a computing system including any combination of such back-end, middleware, and front-end components. The components of the system may be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), the Internet, and a blockchain network.
[0077] The computer system may include a client and a server. The client and the server are generally remote from each other and usually interact with each other via a communication network. The relationship between the client and the server is generated by a computer program that runs on a corresponding computer and has a client-server relationship with each other. The server may be a cloud server, also called a cloud computing server or cloud host, which is a host product in a cloud computing service system and solves the defects of the traditional physical host and VPS service (abbreviated as "Virtual Private Server", or "VPS"), such as difficulty in management and weak business scalability. The server may be a server of a distributed system, or may be a server incorporating a blockchain.
[0078] It should be understood that steps can be rearranged, added, or deleted using the various types of flows shown above. For example, each step described in the present disclosure may be performed in parallel, sequentially, or in a different order, but is not limited herein as long as the technical solution disclosed in the present disclosure can achieve the desired results.
[0079] The above specific embodiments do not limit the scope of protection of the present disclosure. It should be understood that those skilled in the art can make various modifications, combinations, subcombinations, and substitutions according to design requirements and other factors. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present disclosure should be included within the scope of protection of the present disclosure.
Claims
1. An information processing method executed by an information processing device, obtaining an initial dialogue; inputting the initial dialogue sentence into a trained dialogue model to obtain a target response sentence; The dialogue model is a model obtained by training based on a corrected response sample sentence, a second candidate response sample sentence, and a recall response sample sentence, and an initial dialogue sample sentence is input to an initial dialogue model to obtain a plurality of candidate response sample sentences, the second candidate response sample sentence being any one of the plurality of candidate response sample sentences, the corrected response sample sentence being a sentence obtained by correcting a first reply sample sentence among the candidate response sample sentences, and the recall response sample sentence being a sample sentence other than the initial dialogue sample sentence and the plurality of candidate response sample sentences among the training sample sentences, and the corrected response sample sentence, the second candidate response sample sentence, and the recall response sample sentence are inputting sample sentences into an initial sentence generation model of the initial dialogue model to obtain probabilities of corrected response sample sentences, probabilities of second candidate response sample sentences and probabilities of recall response sample sentences; generating actual response sentences according to the probabilities of the corrected response sample sentences, the probabilities of the second candidate response sample sentences and the probabilities of the recall response sample sentences generated by the initial sentence generation model through an initial sentence determination model; and acquiring the dialogue model by jointly training the initial sentence generation model and the initial sentence determination model of the initial dialogue model according to the probabilities of the actual response sentences, the probabilities of the corrected response sample sentences, the probabilities of the second candidate response sample sentences and the probabilities of the recall response sample sentences; 23. An information processing method comprising:
2. inputting the initial dialogue sentence into a trained dialogue model to obtain a target response sentence, within the dialogue model, inputting the initial dialogue sentence into a sentence generation model of the dialogue model to obtain a plurality of candidate response sentences and a probability of each of the candidate response sentences; inputting the plurality of candidate response sentences and a probability of each of the candidate response sentences into a sentence determination model of the dialogue model to obtain a target response sentence; 2. The information processing method according to claim 1,
3. inputting the plurality of candidate response sentences and the probabilities of each of the candidate response sentences into a sentence determination model of the dialogue model to obtain a target response sentence, inputting the plurality of candidate response sentences and a probability of each of the candidate response sentences into the sentence determination model, and selecting a most probable target response sentence from the plurality of candidate response sentences.
3. The information processing method according to claim 2.
4. A model training method executed by a model training device, comprising: obtaining an initial dialogue sample sentence; inputting the initial dialogue sample sentence into an initial dialogue model to obtain a plurality of candidate reply sample sentences; modifying a first candidate reply sample sentence from the plurality of candidate reply sample sentences to obtain a modified reply sample sentence; training the initial dialogue model based on the revised response sample sentence, a second candidate response sample sentence of the plurality of candidate response sample sentences, and a recall response sample sentence to obtain a dialogue model; the recall response sample sentence is a sample sentence other than the initial dialogue sample sentence and the plurality of candidate response sample sentences among the training sample sentences, training the initial dialogue model based on the revised response sample sentence, a second candidate response sample sentence of the plurality of candidate response sample sentences, and a recall response sample sentence to obtain a dialogue model, inputting the corrected response sample sentence, the second candidate response sample sentence and the recall response sample sentence into an initial sentence generation model of the initial dialogue model to obtain the probability of the corrected response sample sentence, the probability of the second candidate response sample sentence and the probability of the recall response sample sentence; and then generating an actual response sentence according to the probability of the corrected response sample sentence, the probability of the second candidate response sample sentence and the probability of the recall response sample sentence generated by the initial sentence generation model through an initial sentence determination model; and co-training the initial sentence generation model and the initial sentence determination model of the initial dialogue model based on the probabilities of the actual response sentences, the corrected response sample sentences, the second candidate response sample sentences, and the recall response sample sentences to obtain the dialogue model. A model training method comprising:
5. a step of jointly training the initial sentence generation model and the initial sentence determination model of the initial dialogue model based on the actual response sentence, the probability of the corrected response sample sentence, the probability of the second candidate response sample sentence, and the probability of the recall response sample sentence to obtain the dialogue model, determining a loss function based on the actual response sentence and the corrected response sample sentence; and co-training the initial sentence generation model and the initial sentence determination model to obtain the dialogue model, based on the loss function, with training targets that the probability of the corrected response sample sentence is greater than the probability of the second candidate response sample sentence, the probability of the corrected response sample sentence is greater than the probability of the recall response sample sentence, and the probability of the second candidate response sample sentence is greater than the probability of the recall response sample sentence.
5. The method of claim 4, wherein the model training method is
6. An information processing device, an acquisition module for acquiring an initial dialogue; an input module for inputting the initial dialogue sentence into a trained dialogue model to obtain a target response sentence; The dialogue model is a model obtained by training based on a corrected response sample sentence, a second candidate response sample sentence, and a recall response sample sentence, and an initial dialogue sample sentence is input to an initial dialogue model to obtain a plurality of candidate response sample sentences, the second candidate response sample sentence being any one of the plurality of candidate response sample sentences, the corrected response sample sentence being a sentence obtained by correcting a first reply sample sentence among the candidate response sample sentences, and the recall response sample sentence being a sample sentence other than the initial dialogue sample sentence and the plurality of candidate response sample sentences among the training sample sentences, and the corrected response sample sentence, the second candidate response sample sentence, and the recall response sample sentence are inputting sample sentences into an initial sentence generation model of the initial dialogue model to obtain probabilities of corrected response sample sentences, probabilities of second candidate response sample sentences and probabilities of recall response sample sentences; generating actual response sentences according to the probabilities of the corrected response sample sentences, the probabilities of the second candidate response sample sentences and the probabilities of the recall response sample sentences generated by the initial sentence generation model through an initial sentence determination model; and acquiring the dialogue model by jointly training the initial sentence generation model and the initial sentence determination model of the initial dialogue model according to the probabilities of the actual response sentences, the probabilities of the corrected response sample sentences, the probabilities of the second candidate response sample sentences and the probabilities of the recall response sample sentences; 23. An information processing apparatus comprising:
7. When the input module inputs the initial dialogue sentence into a trained dialogue model to obtain a target response sentence, Within the dialogue model, inputting the initial dialogue sentence into a sentence generation model of the dialogue model to obtain a plurality of candidate response sentences and a probability of each of the candidate response sentences; inputting the plurality of candidate response sentences and the probability of each of the candidate response sentences into a sentence determination model of the dialogue model to obtain a target response sentence; 7. The information processing apparatus according to claim 6,
8. the input module inputs the plurality of candidate response sentences and the probabilities of each of the candidate response sentences into a sentence determination model of the dialogue model to obtain a target response sentence, inputting the plurality of candidate response sentences and a probability of each of the candidate response sentences into the sentence determination model, and selecting a target response sentence with the highest probability from the plurality of candidate response sentences; 8. The information processing apparatus according to claim 7,
9. A model training device, a sentence acquisition module for acquiring an initial dialogue sample sentence; a sentence input module for inputting the initial dialogue sample sentence into an initial dialogue model to obtain a plurality of candidate reply sample sentences; a correction module that corrects a first candidate reply sample sentence from the plurality of candidate reply sample sentences to obtain a corrected reply sample sentence; a training module for training the initial dialogue model based on the revised response sample sentence, a second candidate response sample sentence among the plurality of candidate response sample sentences, and a recall response sample sentence to obtain a dialogue model; the recall response sample sentence is a sample sentence other than the initial dialogue sample sentence and the plurality of candidate response sample sentences among the training sample sentences, when the training module trains the initial dialogue model based on the revised response sample sentence, a second candidate response sample sentence among the plurality of candidate response sample sentences, and a recall response sample sentence to obtain a dialogue model, inputting the corrected response sample sentence, the second candidate response sample sentence and the recall response sample sentence into an initial sentence generation model of the initial dialogue model to obtain the probability of the corrected response sample sentence, the probability of the second candidate response sample sentence and the probability of the recall response sample sentence; and then generating an actual response sentence according to the probability of the corrected response sample sentence, the probability of the second candidate response sample sentence and the probability of the recall response sample sentence generated by the initial sentence generation model using an initial sentence determination model; jointly training the initial sentence generation model and the initial sentence determination model of the initial dialogue model based on the actual response sentence, the probability of the corrected response sample sentence, the probability of the second candidate response sample sentence, and the probability of the recall response sample sentence to obtain the dialogue model; A model training device characterized by:
10. the training module jointly trains the initial sentence generation model and the initial sentence determination model of the initial dialogue model according to the actual response sentence, the probability of the corrected response sample sentence, the probability of the second candidate response sample sentence, and the probability of the recall response sample sentence to obtain the dialogue model; determining a loss function based on the actual response sentence and the sample corrected response sentence; Based on the loss function, a probability of the corrected response sample sentence is greater than a probability of the second candidate response sample sentence, a probability of the corrected response sample sentence is greater than a probability of the recall response sample sentence, and a probability of the second candidate response sample sentence is greater than a probability of the recall response sample sentence, using these as training targets, to jointly train the initial sentence generation model and the initial sentence determination model to obtain the dialogue model.
10. The model training device according to claim 9.
11. An electronic device, At least one processor; a memory communicatively coupled to the at least one processor; The memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor such that the at least one processor performs a method according to any one of claims 1 to 3 or 4 or 5.
1. An electronic device comprising:
12. A non-transitory computer-readable storage medium having computer instructions stored thereon, comprising: The computer instructions cause a computer to carry out a method according to any one of claims 1 to 3 or 4 or 5. A non-transitory computer-readable storage medium comprising:
13. A computer program comprising: The computer program, when executed by a processor, implements the steps of the method according to any one of claims 1 to 3, 4 or 5. A computer program comprising:
Citation Information
Patent Citations
Model learning device, ranking unit, method and program
JP2015225416A
Interactive device, interacting method, and computer program for the same
JP2016212541A
Information processing device, information processing method, and program
JP2018206307A
Generation device and generation method
JP2022067223A