Dialogue generation method, dialogue generation device, dialogue robot and storage medium

By obtaining historical dialogue information and multimedia information and determining dialogue relationships, the problem of dull responses of dialogue robots is solved, and a more intelligent and diversified dialogue is achieved.

CN113987122BActive Publication Date: 2025-06-06BEIJING XIAOPENG AUTOMOBILE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111254663.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-10-27
Publication Date
2025-06-06
Estimated Expiration
2041-10-27

AI Technical Summary

Technical Problem

The existing dialogue robot responses appear dull and not intelligent enough, making it difficult to meet the dialogue needs of different scenarios.

Method used

By obtaining historical dialogue information and multimedia information, we determine the relationship between the dialogue personnel and the relationship between the dialogue robot and the dialogue personnel, and generate the reply result.

Benefits of technology

The diversity of dialogue is achieved, allowing dialogue robots to generate replies based on different personnel relationships and human-machine relationships to meet the dialogue needs of different scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113987122B_ABST
    Figure CN113987122B_ABST
Patent Text Reader

Abstract

The present invention discloses a dialogue generation method, a dialogue generation device, a dialogue robot and a storage medium. The dialogue generation method can be used for a dialogue robot. The dialogue generation method comprises: obtaining historical dialogue information and multimedia information; determining a relationship set according to the historical dialogue information and multimedia information, the relationship set including the relationship between dialogue personnel and the relationship between the dialogue robot and the dialogue personnel; generating a reply result according to the historical dialogue information and the relationship set. In the above-mentioned dialogue generation method, dialogue generation device, dialogue robot and storage medium, the relationship between dialogue personnel and the relationship between the dialogue robot and the dialogue personnel are determined according to the historical dialogue information and multimedia information. In this way, the reply result can be generated according to different personnel relationships and different human-machine relationships, so that the generated dialogue is more diverse and can meet different dialogue requirements in different scenarios.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to speech technology, and in particular to a dialogue generation method, a dialogue generation device, a dialogue robot and a storage medium. Background Art

[0002] In the related art, a conversational robot responds to conversations in various scenarios using one role, which causes the robot's responses to appear rather rigid and not intelligent enough. Summary of the invention

[0003] The present invention provides a dialogue generation method, a dialogue generation device, a dialogue robot and a storage medium.

[0004] The dialogue generation method of the embodiment of the present invention is used for a dialogue robot, and the dialogue generation method includes: obtaining historical dialogue information and multimedia information; determining a relationship set based on the historical dialogue information and the multimedia information, and the relationship set includes the relationship between the dialogue personnel and the relationship between the dialogue robot and the dialogue personnel; generating a reply result based on the historical dialogue information and the relationship set.

[0005] In the above-mentioned dialogue generation method, the relationship between the dialogue personnel and the relationship between the dialogue robot and the dialogue personnel are determined based on historical dialogue information and multimedia information. In this way, response results can be generated according to different personnel relationships and different human-computer relationships, making the generated dialogue more diverse and able to meet different dialogue needs in different scenarios.

[0006] In some embodiments, the conversation generation method includes: determining a first relationship set according to the historical conversation information. The determining a relationship set according to the historical conversation information and the multimedia information includes: determining the relationship set according to the first relationship set and the multimedia information.

[0007] In some embodiments, the multimedia information includes audio information, and the conversation generation method includes: determining a second relationship set according to the audio information. The determining the relationship set according to the first relationship set and the multimedia information includes: determining the relationship set according to the first relationship set and the second relationship set.

[0008] In some implementations, the dialog generation method includes: processing the audio information to obtain a valid audio segment. The determining the second relationship set according to the audio information includes: determining the second relationship set according to the valid audio segment.

[0009] In some embodiments, the multimedia information includes video information, and the dialog generation method includes: determining a third relationship set according to the video information. The determining the relationship set according to the first relationship set and the multimedia information includes: determining the relationship set according to the first relationship set and the third relationship set.

[0010] In some implementations, the dialog generation method includes: processing the video information to obtain a valid video segment. The determining the third relationship set according to the video information includes: determining the third relationship set according to the valid video segment.

[0011] In some embodiments, the multimedia information includes audio information and video information, and the conversation generation method includes: determining a second relationship set according to the audio information; and determining a third relationship set according to the video information. Determining the relationship set according to the first relationship set and the multimedia information includes: determining the relationship set according to the first relationship set, the second relationship set, and the third relationship set.

[0012] In some embodiments, the conversation generation method includes: inputting the relationship set into a relationship modeling model to obtain a relationship graph. Generating a reply result based on the historical conversation information and the relationship set includes: generating the reply result based on the historical conversation information and the relationship graph.

[0013] In some embodiments, the conversation generation method includes: determining a reply object. Generating a reply result according to the historical conversation information and the relationship set includes: inputting the historical conversation information, the relationship set and the reply object into a generative conversation model to generate the reply result.

[0014] The dialogue generation device of the embodiment of the present invention is used for a dialogue robot, and the dialogue generation device includes an acquisition module, a determination module and a generation module. The acquisition module is used to acquire historical dialogue information and multimedia information. The determination module is used to determine a relationship set based on the historical dialogue information and the multimedia information, and the relationship set includes the relationship between dialogue personnel and the relationship between the dialogue robot and the dialogue personnel. The generation module is used to generate a reply result based on the historical dialogue information and the relationship set.

[0015] In the above-mentioned dialogue generation device, the relationship between the dialogue personnel and the relationship between the dialogue robot and the dialogue personnel are determined based on historical dialogue information and multimedia information. In this way, response results can be generated according to different personnel relationships and different human-computer relationships, making the generated dialogue more diverse and able to meet different dialogue needs in different scenarios.

[0016] The conversation robot of the embodiment of the present invention includes one or more processors and a memory, wherein the memory stores a computer program, and when the computer program is executed by the processor, the steps of the conversation generation method of any one of the above embodiments are implemented.

[0017] In the above-mentioned conversation robot, the relationship between the conversation persons and the relationship between the conversation robot and the conversation persons are determined based on historical conversation information and multimedia information. In this way, response results can be generated according to different personal relationships and different human-computer relationships, making the generated conversation more diverse and able to meet different conversation needs in different scenarios.

[0018] The computer-readable storage medium of the embodiment of the present invention stores a computer program thereon, and is characterized in that when the program is executed by a processor, the dialogue generation method of any one of the above-mentioned embodiments is implemented.

[0019] In the above-mentioned computer-readable storage medium, the relationship between the interlocutors and the relationship between the dialogue robot and the interlocutors are determined based on historical dialogue information and multimedia information. In this way, response results can be generated according to different interpersonal relationships and different human-computer relationships, making the generated dialogue more diverse and able to meet different dialogue needs in different scenarios.

[0020] Additional aspects and advantages of the present invention will be given in part in the following description and in part will be obvious from the following description, or will be learned through practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] The above and / or additional aspects and advantages of the present invention will become apparent and easily understood from the following description of the embodiments in conjunction with the accompanying drawings, in which:

[0022] Figure 1 is a flow chart of a method for generating a dialogue according to certain embodiments of the present invention;

[0023] Figure 2 is a schematic diagram of a conversational robot according to certain embodiments of the present invention;

[0024] Figure 3 is a schematic diagram of a conversation generation device according to some embodiments of the present invention;

[0025] Figure 4 is a flow chart of a method for generating a dialogue according to certain embodiments of the present invention;

[0026] Figure 5 is a schematic diagram of a conversation generation device according to some embodiments of the present invention;

[0027] Figure 6is a flow chart of a method for generating a dialogue according to certain embodiments of the present invention;

[0028] Figure 7 is a schematic diagram of a conversation generation device according to some embodiments of the present invention;

[0029] Figure 8 and Fig. 9 is a schematic diagram of a set of relationships of certain embodiments of the present invention;

[0030] Fig.10 is a flow chart of a method for generating a dialogue according to certain embodiments of the present invention;

[0031] Fig.11 is a schematic diagram of a conversation generation device according to some embodiments of the present invention;

[0032] Fig.12 is a flow chart of a method for generating a dialogue according to certain embodiments of the present invention;

[0033] Fig.13 is a schematic diagram of a conversation generation device according to some embodiments of the present invention;

[0034] Fig.14 is a flow chart of a method for generating a dialogue according to certain embodiments of the present invention;

[0035] Fig.15 is a schematic diagram of a conversation generation device according to some embodiments of the present invention;

[0036] Fig.16 is a flow chart of a method for generating a dialogue according to certain embodiments of the present invention;

[0037] Fig.17 is a schematic diagram of a conversation generation device according to some embodiments of the present invention;

[0038] Fig.18 is a flow chart of a method for generating a dialogue according to certain embodiments of the present invention;

[0039] Fig.19 is a schematic diagram of a conversation generation device according to some embodiments of the present invention;

[0040] Fig. 20 is a flow chart of a method for generating a dialogue according to certain embodiments of the present invention;

[0041] Fig.21 is a schematic diagram of a conversation generation device according to some embodiments of the present invention;

[0042] Fig. 22 It is a schematic diagram of the connection between a conversational robot and a computer-readable storage medium in certain embodiments of the present invention.

[0043] Description of main component symbols:

[0044] Dialogue robot 100, dialogue generation device 10, acquisition module 11, determination module 12, generation module 13, first processing module 14, second processing module 152, third processing module 154, fourth processing module 162, fifth processing module 164, sixth processing module 172, seventh processing module 174, eighth processing module 18, ninth processing module 19, processor 20, memory 30, computer readable storage medium 800. DETAILED DESCRIPTION

[0045] Embodiments of the present invention are described in detail below, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present invention, and should not be construed as limiting the present invention.

[0046] In the related art, in various scenarios with different numbers of people in conversation and different personal relationships, the conversational robot responds to the conversation in one role, which makes the robot's responses appear to be rather rigid and not intelligent enough.

[0047] See also Figure 1 and Figure 2 The dialogue generation method of the embodiment of the present invention can be used for the dialogue robot 100. The dialogue generation method includes:

[0048] 01: Get historical conversation information and multimedia information;

[0049] 02: Determine a relationship set based on historical dialogue information and multimedia information, where the relationship set includes the relationship between dialogue personnel and the relationship between the dialogue robot 100 and the dialogue personnel;

[0050] 03: Generate response results based on historical conversation information and relationship sets.

[0051] See also Figures 1 to 3, the dialogue generation device 10 of the embodiment of the present invention can be used for the dialogue robot 100. The dialogue generation device 10 includes an acquisition module 11, a determination module 12 and a generation module 13. The dialogue generation method of the embodiment of the present invention can be implemented by the dialogue generation device 10 of the embodiment of the present invention, wherein step 01 can be implemented by the acquisition module 11, step 02 can be implemented by the determination module 12, and step 03 can be implemented by the generation module 13. That is to say, the acquisition module 11 can be used to acquire historical dialogue information and multimedia information. The determination module 12 can be used to determine a relationship set based on the historical dialogue information and multimedia information, and the relationship set includes the relationship between the dialogue personnel and the relationship between the dialogue robot 100 and the dialogue personnel. The generation module 13 can be used to generate a reply result based on the historical dialogue information and the relationship set.

[0052] In the above-mentioned dialogue generation method and dialogue generation device 10, the relationship between the dialogue personnel and the relationship between the dialogue robot 100 and the dialogue personnel are determined based on historical dialogue information and multimedia information. In this way, response results can be generated according to different personnel relationships and different human-computer relationships, making the generated dialogue more diverse and able to meet different dialogue needs in different scenarios.

[0053] The conversation robot 100 of the present invention may include devices such as a voice assistant, and the conversation robot 100 may be used in environments such as in-car and indoor environments. The conversation robot 100 may integrate multimedia functions, that is, the conversation robot 100 itself has a multimedia information acquisition device (such as a camera, a radio, etc.), or the conversation robot 100 and the multimedia information acquisition device are separately arranged, and the conversation robot 100 acquires multimedia information through communication connections, etc., which is not specifically limited here.

[0054] The historical conversation information may be text information. In some embodiments, the multimedia information includes audio information, and the historical conversation information may be a text result obtained after the audio information is processed by automatic speech recognition technology (Automatic Speech Recognition, ASR). All the historical conversation information obtained by audio information recognition may be retained, so that a relationship set may be determined and a reply result may be generated based on the complete historical conversation information.

[0055] The relationship set of the present invention includes the personnel relationship between each dialogue person and all dialogue members, and the human-machine relationship between the dialogue robot 100 and all dialogue persons. Dialogue persons may refer to persons participating in historical dialogue information and / or persons existing in multimedia information. Personnel relationships include, for example, father and daughter, mother and daughter, husband and wife, colleagues, classmates, strangers, friends, and others. Human-machine relationships include, for example, colleagues, classmates, strangers, friends, and others. Of course, human-machine relationships may also include father and daughter, mother and daughter, husband and wife, and others, which are not specifically limited here.

[0056] The reply result of the present invention can be outputted by voice, display, etc. For example, the conversation robot 100 may include a speaker, and after generating the reply result, the conversation robot 100 may play the reply result through the speaker. For another example, the conversation robot 100 may include a display screen, and after generating the reply result, the conversation robot 100 may display the reply result through the display screen.

[0057] See also Figure 4 In some embodiments, the dialog generation method includes:

[0058] 04: Determine the first relationship set according to the historical conversation information;

[0059] Step 02 (determining a relationship set based on historical conversation information and multimedia information) includes:

[0060] 022: Determine a relationship set according to the first relationship set and the multimedia information.

[0061] See also Figure 5 In some embodiments, the dialog generation device 10 includes a first processing module 14. Step 04 can be implemented by the first processing module 14, and step 022 can be implemented by the determination module 12. That is, the first processing module 14 can be used to determine the first relationship set according to the historical dialog information, and the determination module 12 can be used to determine the relationship set according to the first relationship set and the multimedia information.

[0062] In this way, the first relationship set can be determined based on the historical dialogue information. The first relationship set may include possible relationships between dialogue personnel. For example, the historical dialogue information is "Dialogue Person 1: Why did that uncle run a red light just now?", "Dialogue Person 2: Baby, we must obey traffic rules, running a red light is very dangerous, we can't do that." Based on the historical dialogue information, it can be determined that the possible personnel relationships between dialogue personnel 1 and dialogue personnel 2 are: mother and son, mother and daughter, father and son, father and daughter, etc., so that the first relationship set can be obtained.

[0063] Of course, the first relationship set may include possible relationships between the dialogue personnel and possible relationships between the dialogue robot 100 and the dialogue personnel. For example, based on the possible personnel relationships between dialogue personnel 1 and dialogue personnel 2 being: mother and son, mother and daughter, father and son, and father and daughter, it can be determined that the possible human-machine relationship between the dialogue robot 100 and dialogue personnel 1 is: colleagues, classmates, strangers, friends, and others, and the possible human-machine relationship between the dialogue robot 100 and dialogue personnel 2 is: colleagues, classmates, strangers, friends, and others. In one embodiment, the human-machine relationship between a specific dialogue personnel and the dialogue robot 100 can be pre-set, and then the human-machine relationship between other dialogue personnel and the dialogue robot 100 can be determined based on the human-machine relationship between the specific dialogue personnel and the dialogue robot 100. For example, dialogue personnel 1 is pre-set as a specific dialogue personnel, and the human-machine relationship between dialogue personnel 1 and the dialogue robot 100 is a friend, then the human-machine relationship between dialogue personnel 2 and the dialogue robot 100 can be the relationship between a younger generation and an older generation.

[0064] The embodiment of the present invention is described by taking two interlocutors as an example. When there are more than two interlocutors (for example, three interlocutors or four interlocutors), the implementation scheme is similar to that of the two interlocutors and will not be described in detail here.

[0065] By combining the first relationship set and the multimedia information, relatively accurate human relationships and human-machine relationships can be further determined as a relationship set.

[0066] See also Figure 6 In some embodiments, the multimedia information includes audio information, and the conversation generation method includes:

[0067] 052: Determine a second relationship set according to the audio information;

[0068] Step 022 (determining a relationship set according to the first relationship set and the multimedia information) includes:

[0069] 0222: Determine a relationship set according to the first relationship set and the second relationship set.

[0070] See also Figure 7 In some embodiments, the multimedia information includes audio information, and the dialog generation device 10 includes a second processing module 152. Step 052 can be implemented by the second processing module 152, and step 0222 can be implemented by the determination module 12. That is, the second processing module 152 can be used to determine the second relationship set according to the audio information, and the determination module 12 can be used to determine the relationship set according to the first relationship set and the second relationship set.

[0071] In this way, the second relationship set can be determined according to the audio information. Among them, the conversation robot 100 can be applied to vehicles, and the audio information can be recorded and obtained by the in-vehicle sound receiving device, and the sound receiving device can obtain the audio information within the first preset time period at the current moment, and the first preset time period can be set according to user needs. The longer the first preset time period, the more accurate the second relationship set obtained is, and the shorter the first preset time period, the smaller the storage and processing amount of the audio information is, and the audio information beyond the first preset time period can be deleted, thereby reducing the risk of privacy leakage.

[0072] The second relationship set may include possible relationships between the dialogue parties. The gender of the dialogue parties can be determined based on the audio information. For example, if it is determined that the dialogue parties are all girls, the possible relationships between dialogue parties 1 and 2 are: mother and daughter, colleagues, classmates, strangers, friends, and others, so that the second relationship set can be obtained. The emotions between the dialogue parties can be determined based on the audio information. Emotions include happiness, anger, sadness, and anger. For example, if the volume of the dialogue parties is high, it can be determined that the emotions of the dialogue parties are anger, anger, etc., then the possible relationships between dialogue parties 1 and 2 are: quarreling mother and daughter, quarreling colleagues, quarreling classmates, quarreling strangers, quarreling friends, etc.

[0073] Of course, the second relationship set may include possible relationships between the dialogue personnel and possible relationships between the dialogue robot 100 and the dialogue personnel. For example, based on the possible personnel relationships between dialogue personnel 1 and dialogue personnel 2 being: mother and daughter, colleagues, classmates, strangers, friends, and others, it can be determined that the possible human-machine relationship between the dialogue robot 100 and dialogue personnel 1 is: colleagues, classmates, strangers, friends, and others, and the possible human-machine relationship between the dialogue robot 100 and dialogue personnel 2 is: colleagues, classmates, strangers, friends, and others. In one embodiment, the human-machine relationship between a specific dialogue personnel and the dialogue robot 100 can be pre-set, and then the human-machine relationship between other dialogue personnel and the dialogue robot 100 can be determined based on the human-machine relationship between the specific dialogue personnel and the dialogue robot 100. For example, dialogue personnel 1 is pre-set as a specific dialogue personnel, and the human-machine relationship between dialogue personnel 1 and the dialogue robot 100 is a friend, then the human-machine relationship between dialogue personnel 2 and the dialogue robot 100 can be the relationship between the younger and the older, colleagues, classmates, strangers, friends, and others.

[0074] In combination with the first relationship set and the second relationship set, the more correct personnel relationships and human-computer relationships can be further determined as the relationship set. Specifically, the first relationship set and the second relationship set can be fused and judged by voting or other methods, that is, the personnel relationship with the highest possibility (the most possible occurrences) can be determined. For example, the historical dialogue information is "Dialogue Personnel 1: Why did that uncle just run a red light?", "Dialogue Personnel 2: Baby, we must obey traffic rules, running a red light is very dangerous, we can't do this", at this time, according to the first relationship set, the possible personnel relationships between dialogue personnel 1 and dialogue personnel 2 are determined as: mother and son, mother and daughter, father and son, father and daughter, etc., and according to the second relationship set, the possible personnel relationships between dialogue personnel 1 and dialogue personnel 2 are determined as: mother and daughter, colleagues, classmates, strangers, friends, others, etc., then the possible personnel relationships in the first relationship set and the second relationship set all include mother-daughter relationships, so please refer to Figure 8 , it can be determined that the relationship between dialogue personnel 1 and dialogue personnel 2 is a mother-daughter relationship. At this time, the relationship between dialogue personnel 1 and the dialogue robot is initialized as r, and the relationship between dialogue personnel 2 and the dialogue robot is r1. Then r can be a friend relationship, and r1 can be a relationship between an elder and a younger generation. The following can be established: Fig. 9 The set of relationships shown.

[0075] See also Fig.10 In some embodiments, the dialog generation method includes:

[0076] 054: Process audio information to obtain valid audio clips;

[0077] Step 052 (determining a second relationship set according to audio information) includes:

[0078] 0522: Determine a second relationship set according to the valid audio segment.

[0079] See also Fig.11 In some embodiments, the dialog generation device 10 includes a third processing module 154. Step 054 can be implemented by the third processing module 154, and step 0522 can be implemented by the second processing module 152. That is, the third processing module 154 can be used to process audio information to obtain a valid audio segment, and the second processing module 152 can be used to determine a second relationship set according to the valid audio segment.

[0080] Specifically, a valid audio segment may refer to an audio segment with valid speech, and segments of the audio information that only contain noise and background sound but no valid speech may not be retained. In this way, the second relationship set can be quickly and accurately determined based on the valid audio segment, thereby reducing the amount of audio information to be processed. In one embodiment, the valid audio segment may be detected by voice activity detection (VAD), which may be used to identify speech presence and speech absence in the audio information, thereby determining the valid audio segment.

[0081] See also Fig.12 In some embodiments, the multimedia information includes video information, and the conversation generation method includes:

[0082] 062: Determine a third relationship set according to the video information;

[0083] Step 022 (determining a relationship set according to the first relationship set and the multimedia information) includes:

[0084] 0224: Determine a relationship set according to the first relationship set and the third relationship set.

[0085] See also Fig.13 In some embodiments, the multimedia information includes video information, and the dialog generation device 10 includes a fourth processing module 162. Step 062 can be implemented by the fourth processing module 162, and step 0224 can be implemented by the determination module 12. That is, the fourth processing module 162 can be used to determine the third relationship set according to the video information, and the determination module 12 can be used to determine the relationship set according to the first relationship set and the third relationship set.

[0086] In this way, the third relationship set can be determined according to the video information. The conversational robot 100 can be applied to a vehicle, and the video information can be obtained by recording the camera in the vehicle. The camera can obtain the video information within the second preset time period at the current moment. The second preset time period can be set according to user needs. The longer the second preset time period, the more accurate the third relationship set obtained. The shorter the second preset time period, the smaller the storage and processing amount of the video information. The video information beyond the second preset time period can be deleted, thereby reducing the risk of privacy leakage.

[0087] The third relationship set may include possible relationships between the dialogue personnel. The gender of the dialogue personnel can be determined based on the video information. For example, if it is determined that the dialogue personnel are all girls, the possible personnel relationships between dialogue personnel 1 and dialogue personnel 2 are: mother and daughter, colleagues, classmates, strangers, friends, and others, so that the third relationship set can be obtained. The facial expressions, emotions, and actions between the dialogue personnel can be determined based on the video information. Emotions include happiness, anger, sadness, and anger. For example, if the dialogue personnel have a hideous face and swing their hands, it can be determined that the dialogue personnel's emotions are anger and anger, and the possible personnel relationships between dialogue personnel 1 and dialogue personnel 2 are: quarreling mother and daughter, quarreling colleagues, quarreling classmates, quarreling strangers, quarreling friends, etc. The location of the dialogue personnel can be determined based on the video information. For example, dialogue personnel 1 is a driver and dialogue personnel 2 is a passenger.

[0088] Of course, the third relationship set may include possible relationships between the dialogue personnel and possible relationships between the dialogue robot 100 and the dialogue personnel. For example, based on the possible personnel relationships between dialogue personnel 1 and dialogue personnel 2 being: mother and daughter, colleagues, classmates, strangers, friends, and others, it can be determined that the possible human-machine relationship between the dialogue robot 100 and dialogue personnel 1 is: colleagues, classmates, strangers, friends, and others, and the possible human-machine relationship between the dialogue robot 100 and dialogue personnel 2 is: colleagues, classmates, strangers, friends, and others. In one embodiment, the human-machine relationship between a specific dialogue personnel and the dialogue robot 100 can be pre-set, and then the human-machine relationship between other dialogue personnel and the dialogue robot 100 can be determined based on the human-machine relationship between the specific dialogue personnel and the dialogue robot 100. For example, dialogue personnel 1 is pre-set as a specific dialogue personnel, and the human-machine relationship between dialogue personnel 1 and the dialogue robot 100 is a friend, then the human-machine relationship between dialogue personnel 2 and the dialogue robot 100 can be the relationship between the younger and the older, colleagues, classmates, strangers, friends, and others.

[0089] By combining the first relationship set and the third relationship set, we can further determine the more correct personnel relationships and human-computer relationships as the relationship set. Specifically, we can use voting method and other methods to fuse and judge the first relationship set and the third relationship set, that is, determine the personnel relationship with the highest possibility (the most possible occurrences). For example, the historical dialogue information is "Dialogue Personnel 1: Why did that uncle just run a red light?", "Dialogue Personnel 2: Baby, we must obey traffic rules, running a red light is very dangerous, we can't do this", at this time, according to the first relationship set, the possible personnel relationships between dialogue personnel 1 and dialogue personnel 2 are determined as: mother and son, mother and daughter, father and son, father and daughter, etc., and according to the third relationship set, the possible personnel relationships between dialogue personnel 1 and dialogue personnel 2 are determined as: mother and daughter, colleagues, classmates, strangers, friends, others, etc., then the possible personnel relationships in the first relationship set and the third relationship set all include mother-daughter relationships, so please refer to Figure 8 , it can be determined that the relationship between dialogue personnel 1 and dialogue personnel 2 is a mother-daughter relationship. At this time, the relationship between dialogue personnel 1 and the dialogue robot is initialized as r, and the relationship between dialogue personnel 2 and the dialogue robot is r1. Then r can be a friend relationship, and r1 can be a relationship between an elder and a younger generation. The following can be established: Fig. 9 The set of relationships shown.

[0090] See also Fig.14 In some embodiments, the dialog generation method includes:

[0091] 064: Process video information to obtain valid video clips;

[0092] Step 062 (determining a third relationship set according to video information) includes:

[0093] 0622: Determine a third relationship set according to the valid video clips.

[0094] See also Fig.15 In some embodiments, the dialog generation device 10 includes a fifth processing module 164. Step 064 can be implemented by the fifth processing module 164, and step 0622 can be implemented by the fourth processing module 162. That is, the fifth processing module 164 can be used to process the video information to obtain a valid video segment, and the fourth processing module 162 can be used to determine the third relationship set according to the valid video segment.

[0095] Specifically, effective video clips may refer to video clips with effective interactions, and clips without any interactions may not be retained, wherein effective interactions include body interactions, facial interactions (facial expressions), and language interactions (determined by historical dialogue information or audio information). In this way, the third relationship set can be quickly and accurately determined based on effective video clips, reducing the amount of video information processing.

[0096] Determining a relationship set based on the first relationship set and the second relationship set, and determining a relationship set based on the first relationship set and the third relationship set can reduce the processing load of the conversational robot 100 and generate a reply result more quickly.

[0097] See also Fig.16 In some embodiments, the multimedia information includes audio information and video information, and the conversation generation method includes:

[0098] 072: Determine a second relationship set according to the audio information;

[0099] 074: Determine a third relationship set according to the video information;

[0100] Step 022 (determining a relationship set according to the first relationship set and the multimedia information) includes:

[0101] 0226: Determine a relationship set according to the first relationship set, the second relationship set and the third relationship set.

[0102] See also Fig.17 In some embodiments, the multimedia information includes audio information and video information, and the dialog generation device 10 includes a sixth processing module 172 and a seventh processing module 174. Step 072 can be implemented by the sixth processing module 172, step 074 can be implemented by the seventh processing module 174, and step 0226 can be implemented by the determination module 12. That is, the sixth processing module 172 can be used to determine the second relationship set according to the audio information, the seventh processing module 174 can be used to determine the third relationship set according to the video information, and the determination module 12 can be used to determine the relationship set according to the first relationship set, the second relationship set, and the third relationship set.

[0103] The manner of determining the relationship set according to the first relationship set, the second relationship set, and the third relationship set is similar to the manner of determining the relationship set according to the first relationship set and the second relationship set, and determining the relationship set according to the first relationship set and the third relationship set, and will not be repeated here. The second relationship set is determined according to the audio information, and specifically, the second relationship set can be determined according to the valid audio clip. The third relationship set is determined according to the video information, and specifically, the third relationship set can be determined according to the valid video clip. The relationship set is determined according to the first relationship set, the second relationship set, and the third relationship set. Since multimodality (audio plus video) is adopted, the personnel relationship and the human-computer relationship can be more accurately determined as the relationship set. For example, since the voice of the dialogue person is neutral, the gender of the dialogue person cannot be determined in the second relationship set. In this case, the gender of the dialogue person can be determined in the third relationship set according to the physical characteristics and dress (clothing, accessories, makeup, etc.) of the dialogue person in the image. For another example, since the physical characteristics and dress of the dialogue person in the image are neutral, the gender of the dialogue person cannot be determined in the third relationship set. In this case, the gender of the dialogue person can be determined in the second relationship set according to the voice of the dialogue person.

[0104] See also Fig.18 In some embodiments, the dialog generation method includes:

[0105] 08: Input the relationship set into the relational modeling model to obtain a relationship graph;

[0106] Step 03 (generating a reply result based on historical conversation information and relationship set) includes:

[0107] 032: Generate reply results based on historical conversation information and relationship graph.

[0108] See also Fig.19 In some embodiments, the dialog generation device 10 includes an eighth processing module 18. Step 08 can be implemented by the eighth processing module 18, and step 032 can be implemented by the generation module 13. That is, the eighth processing module 18 can be used to input the relationship set into the relationship modeling model to obtain a relationship graph, and the generation module 13 can be used to generate a reply result according to the historical dialog information and the relationship graph.

[0109] Specifically, the relationship set can be represented in the form of an adjacency matrix, which can be regarded as a discrete image representation. In order to facilitate the reading of the dialogue robot 100, the relationship set can be converted into a relationship graph using a relationship modeling model, wherein the data format of the relationship graph is a data format that can be read by the dialogue robot 100 (for example, the data format of the relationship graph can meet the format requirements of the input data of the generative dialogue model described below). In one embodiment, the relationship modeling model can be an n-layer CNN convolutional layer, a graph model modeling, a structured information expression of a triple, an RNN modeling, etc., which is not specifically limited here.

[0110] See also Fig. 20 In some embodiments, the dialog generation method includes:

[0111] 09: Determine the target of the reply;

[0112] Step 03 (generating a reply result based on historical conversation information and relationship set) includes:

[0113] 034: Input historical conversation information, relation sets, and reply objects into the generative conversation model to generate reply results.

[0114] See also Fig.21 In some embodiments, the dialog generation device 10 includes a ninth processing module 19. Step 09 can be implemented by the ninth processing module 19, and step 034 can be implemented by the generation module 13. That is, the ninth processing module 19 can be used to determine the reply object, and the generation module 13 can be used to input the historical dialog information, the relationship set, and the reply object into the generative dialog model to generate a reply result.

[0115] In this way, a corresponding reply result can be generated for the reply object. Specifically, the dialogue robot 100 can train the generative dialogue model through self-learning, wherein the self-learning method can be that the dialogue robot 100 obtains a large amount of historical dialogue information and multimedia information, and learns people's dialogue methods and skills in the process of communication and interaction. For example, the dialogue robot 100 finds through learning that when the dialogue personnel are quarreling, the person who persuades the quarrel usually says to the quarreling dialogue personnel "Everyone calm down". In actual applications, the dialogue robot 100 can determine the reply object based on historical dialogue information and multimedia information, for example, determine whether the dialogue personnel are quarreling based on historical dialogue information and multimedia information. If they are quarreling, the quarreling dialogue personnel can be determined as the reply object, and then the historical dialogue information, relationship set and reply object are input into the generative dialogue model to generate the reply result of "Everyone calm down". The generative dialogue model can include reinforcement learning, retrieval dialogue technology, etc., which are not specifically limited here.

[0116] In one embodiment, the historical dialogue information is "Interviewee 1: Why did that uncle just run a red light?", "Interviewee 2: Baby, we have to obey traffic rules, running a red light is very dangerous, we can't do that", at this time, according to the first relationship set, the possible personnel relationships between the dialogue personnel 1 and the dialogue personnel 2 are determined as: mother and son, mother and daughter, father and son, father and daughter, etc., and according to the second relationship set and / or the third relationship set, the possible personnel relationships between the dialogue personnel 1 and the dialogue personnel 2 are determined as: mother and daughter, colleagues, classmates, strangers, friends, others, etc., then it can be determined that the personnel relationship between the dialogue personnel 1 and the dialogue personnel 2 is a mother-daughter relationship. At this time, the relationship between the dialogue personnel 1 and the dialogue robot is initialized to r, and the relationship between the dialogue personnel 2 and the dialogue robot is r1, then please refer to Figure 8 and Fig. 9 , r can be a friend relationship, and r1 can be a relationship between an elder and a younger generation. Interviewer 2 raises a question, and interviewer 2 can be determined as the reply object. Combining the historical dialogue information and the relationship set, a reply result of "Little friend, your mother is right, it is wrong to do this" can be generated.

[0117] See also Figure 2 The conversation robot 100 of the embodiment of the present invention includes one or more processors 20 and a memory 30. The memory 30 stores a computer program. When the computer program is executed by the processor 20, the steps of the conversation generation method of any one of the above embodiments are implemented.

[0118] For example, when the computer program is executed by the processor 20, the following may be achieved:

[0119] 01: Get historical conversation information and multimedia information;

[0120] 02: Determine a relationship set based on historical dialogue information and multimedia information, where the relationship set includes the relationship between dialogue personnel and the relationship between the dialogue robot 100 and the dialogue personnel;

[0121] 03: Generate response results based on historical conversation information and relationship sets.

[0122] In the above-mentioned conversation robot 100, the relationship between the conversation persons and the relationship between the conversation robot 100 and the conversation persons are determined based on historical conversation information and multimedia information. In this way, response results can be generated according to different personal relationships and different human-computer relationships, making the generated conversation more diverse and able to meet different conversation needs in different scenarios.

[0123] See also Fig. 22The computer-readable storage medium 800 of the embodiment of the present invention stores a computer program, and when the program is executed by the processor 20, the dialogue generation method of any one of the above-mentioned embodiments is implemented.

[0124] For example, when executed by the processor 20, the following may be achieved:

[0125] 01: Get historical conversation information and multimedia information;

[0126] 02: Determine a relationship set based on historical dialogue information and multimedia information, where the relationship set includes the relationship between dialogue personnel and the relationship between the dialogue robot 100 and the dialogue personnel;

[0127] 03: Generate response results based on historical conversation information and relationship sets.

[0128] In the above-mentioned computer-readable storage medium 800, the relationship between the interlocutors and the relationship between the dialogue robot 100 and the interlocutors are determined based on historical dialogue information and multimedia information. In this way, response results can be generated according to different interpersonal relationships and different human-computer relationships, making the generated dialogue more diverse and able to meet different dialogue needs in different scenarios.

[0129] In the present invention, the computer program includes computer program code. The computer program code may be in source code form, object code form, executable file or some intermediate form, etc. The memory 30 may include a high-speed random access memory, and may also include a non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (Flash Card), at least one disk storage device, a flash memory device, or other volatile solid-state storage devices. The processor 20 may be a central processing unit (Central Processing Unit, CPU), or other general-purpose processors, digital signal processors (Digital Signal Processor, DSP), application specific integrated circuits (Application Specific Integrated Circuit, ASIC), field-programmable gate arrays (Field-Programmable Gate Array, FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0130] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" etc. means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of the different embodiments or examples, without contradiction.

[0131] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features defined as "first" and "second" may explicitly or implicitly include at least one of the features. In the description of the present invention, the meaning of "plurality" is at least two, such as two, three, etc., unless otherwise clearly and specifically defined.

[0132] Any process or method description in a flowchart or otherwise described herein may be understood to represent a module, segment or portion of code that includes one or more executable instructions for implementing the steps of a specific logical function or process, and the scope of the preferred embodiments of the present invention includes alternative implementations in which functions may not be performed in the order shown or discussed, including performing functions in a substantially simultaneous manner or in the reverse order depending on the functions involved, which should be understood by those skilled in the art to which the embodiments of the present invention belong.

[0133] Although the embodiments of the present invention have been shown and described above, it is to be understood that the above embodiments are exemplary and are not to be construed as limitations of the present invention. A person skilled in the art may change, modify, replace and vary the above embodiments within the scope of the present invention.

Claims

1. A dialogue generation method for a dialogue robot, It is characterized in that The dialogue generation method comprises: Get historical conversation information and multimedia information; Determine a relationship set according to the historical dialogue information and the multimedia information, wherein the relationship set includes the relationship between dialogue personnel and the relationship between the dialogue robot and the dialogue personnel; Generate a reply result according to the historical conversation information and the relationship set; The relationship between the dialogue robot and the dialogue person includes at least one of colleagues, classmates, strangers, friends, father and daughter, mother and daughter, and husband and wife. The relationship between the dialogue robot and the dialogue person is a preset human-machine relationship between the dialogue person and the dialogue robot.

2. The method for generating a dialogue according to claim 1, It is characterized in that The dialogue generation method comprises: Determine a first relationship set according to the historical conversation information; The determining of a relationship set according to the historical conversation information and the multimedia information includes: The relationship set is determined according to the first relationship set and the multimedia information.

3. The method for generating a dialogue according to claim 2, It is characterized in that The multimedia information includes audio information, and the dialogue generation method includes: Determine a second relationship set according to the audio information; The determining the relationship set according to the first relationship set and the multimedia information includes: The relationship set is determined according to the first relationship set and the second relationship set.

4. The method for generating a dialogue according to claim 3, It is characterized in that The dialogue generation method comprises: Processing the audio information to obtain a valid audio segment; The determining a second relationship set according to the audio information comprises: The second relationship set is determined according to the valid audio segment.

5. The method for generating a dialogue according to claim 2, It is characterized in that The multimedia information includes video information, and the dialogue generation method includes: Determine a third relationship set according to the video information; The determining the relationship set according to the first relationship set and the multimedia information includes: The relationship set is determined according to the first relationship set and the third relationship set.

6. The method for generating a dialogue according to claim 5, It is characterized in that The dialogue generation method comprises: Processing the video information to obtain a valid video segment; The determining a third relationship set according to the video information includes: The third relationship set is determined according to the valid video segment.

7. The method for generating a dialogue according to claim 2, It is characterized in that The multimedia information includes audio information and video information, and the dialogue generation method includes: Determine a second relationship set according to the audio information; Determine a third relationship set according to the video information; The determining the relationship set according to the first relationship set and the multimedia information includes: The relationship set is determined according to the first relationship set, the second relationship set and the third relationship set.

8. The method for generating a dialogue according to claim 1, It is characterized in that The dialogue generation method comprises: Inputting the relationship set into a relationship modeling model to obtain a relationship graph; The generating a reply result according to the historical conversation information and the relationship set includes: The reply result is generated according to the historical conversation information and the relationship graph.

9. The method for generating a dialogue according to claim 1, It is characterized in that The dialogue generation method comprises: Determine the target audience for the response; The generating a reply result according to the historical conversation information and the relationship set includes: The historical dialogue information, the relationship set and the reply object are input into a generative dialogue model to generate the reply result.

10. A dialogue generation device for a dialogue robot, It is characterized in that The dialogue generating device comprises: An acquisition module, which is used to acquire historical conversation information and multimedia information; A determination module, the determination module is used to determine a relationship set according to the historical dialogue information and the multimedia information, the relationship set including the relationship between the dialogue personnel and the relationship between the dialogue robot and the dialogue personnel; A generation module, the generation module is used to generate a reply result according to the historical conversation information and the relationship set; The relationship between the dialogue robot and the dialogue person includes at least one of colleagues, classmates, strangers, friends, father and daughter, mother and daughter, and husband and wife. The relationship between the dialogue robot and the dialogue person is a preset human-machine relationship between the dialogue person and the dialogue robot.

11. A conversational robot, It is characterized in that The conversation robot includes one or more processors and a memory, wherein the memory stores a computer program, and when the computer program is executed by the processor, the steps of the conversation generation method described in any one of claims 1 to 9 are implemented.

12. A computer-readable storage medium having a computer program stored thereon, It is characterized in that When the program is executed by a processor, the dialog generation method described in any one of claims 1 to 9 is implemented.

Citation Information

Patent Citations

  • Dialogue system, electronic apparatus and method for controlling the dialogue system

    CN111797208A

  • Emotional dialogue generation method and device and emotional dialogue model training method and device

    CN111897933A