Multi-modal dialogue processing method and device, electronic equipment and storage medium
By obtaining multimodal dialogue information and using large language models for in-depth understanding, and generating multimodal target dialogue reply information, the problem of lack of multimodal information in virtual character dialogue is solved, and user experience and interaction freedom is improved.
Patent Information
- Application Number
- CN202510353021.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-24
- Publication Date
- 2025-07-11
AI Technical Summary
In the prior art, virtual character dialogue methods lack input/output support for multimodal information, resulting in low interaction efficiency and poor user experience.
By obtaining the multimodal dialogue information input by the user, combining the role setting information and historical dialogue information, using a large language model to understand the dialogue intention, and generating multimodal target dialogue reply information.
It improves the freedom and immersion of users when talking to virtual characters, and the generated reply information is more anthropomorphic, enhancing the richness and attractiveness of the interaction.
Smart Images

Figure CN120296120A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and more specifically, to a multi-modal dialogue processing method, apparatus, electronic device, and storage medium. Background Art
[0002] In the cultural and entertainment scenarios, there are various virtual characters with rich character settings (such as game intelligent NPCs). With the development of artificial intelligence (AI) technology, the interaction methods of these virtual characters have changed from pre-set dialogue texts to more content generated by AI technology, improving the interaction freedom and user immersion.
[0003] In the related art, in the existing cultural and entertainment scenarios, AI character dialogues mostly reflect the characteristics of the characters from the perspective of text. Generally, based on the character setting text information and dialogue history of the AI character, a text generation model is used to generate responses.
[0004] However, the above-mentioned character dialogue method lacks support for the input / output methods of multi-modal information. In actual situations, the interaction between people is often in the form of multi-modal information, including text, pictures, voice, expressions, etc. That is, the existing character dialogue method lacks the understanding and generation of multi-modal dialogue information, resulting in the lack of anthropomorphism in the finally output response information, thus leading to low information interaction efficiency and poor user experience. Summary of the Invention
[0005] The purpose of the present invention is to provide a multi-modal dialogue processing method, apparatus, electronic device, and storage medium for solving the technical problems existing in the prior art in view of the above deficiencies in the prior art.
[0006] To achieve the above purpose, the technical solutions adopted in the embodiments of the present application are as follows:
[0007] In a first aspect, an embodiment of the present application provides a multi-modal dialogue processing method, the method including:
[0008] Obtain the dialogue information input by the user, where the dialogue information is information of multiple modalities;
[0009] Obtain the character setting information and historical dialogue information of the character conversing with the user;
[0010] Determine the dialogue status information of the user according to the character setting information, the historical dialogue information, and the dialogue information;
[0011] Input the dialogue status information into a pre-trained large language model to obtain a decision result output by the large language model, where the decision result is used to indicate the type of the response information and the prompt word corresponding to the response information;
[0012] Generate and output a target dialogue response message according to the type of the said response message and the prompt words corresponding to the said response message, and the said target dialogue response message is a multi-modal message.
[0013] In a second aspect, an embodiment of the present application further provides a multi-modal dialogue processing device, and the said device includes:
[0014] An acquisition module, configured to acquire the dialogue information input by the user, and the said dialogue information is multi-modal information; acquire the role setting information and historical dialogue information of the dialogue with the said user;
[0015] A determination module, configured to determine the dialogue status information of the user according to the said role setting information, the said historical dialogue information and the said dialogue information;
[0016] The said processing module is configured to input the said dialogue status information into a pre-trained large language model to obtain a decision result output by the large language model, and the said decision result is used to indicate the type of the response message and the prompt words corresponding to the response message;
[0017] A generation module, configured to generate and output a target dialogue response message according to the type of the said response message and the prompt words corresponding to the said response message, and the said target dialogue response message is multi-modal information.
[0018] In a third aspect, an embodiment of the present application provides an electronic device, including: a processor, a storage medium and a bus, the storage medium stores machine-readable instructions executable by the processor, when the electronic device runs, the processor communicates with the storage medium through the bus, and the processor executes the machine-readable instructions to execute the multi-modal dialogue processing method provided in the first aspect.
[0019] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is run by a processor, it executes the multi-modal dialogue processing method provided in the first aspect.
[0020] The beneficial effects of the present application are:
[0021] The present application provides a multi-modal dialogue processing method, apparatus, electronic device, and storage medium. The method includes: obtaining dialogue information input by a user; obtaining role setting information and historical dialogue information of the dialogue with the user; determining the user's dialogue status information according to the role setting information, historical dialogue information, and dialogue information; inputting the dialogue status information into a pre-trained large language model to obtain a decision result output by the large language model, where the decision result is used to indicate the type of reply information and the prompt words corresponding to the reply information; generating and outputting a target dialogue reply information according to the type of reply information and the prompt words corresponding to the reply information. Among them, both the dialogue information input by the user and the output target dialogue reply information are multi-modal information. In this solution, in order to solve the problem in the related art that does not support the input / output of multi-modal information, it is proposed to obtain the multi-modal dialogue information input by the user, and use the large language model to deeply understand the dialogue intention of the dialogue information, role setting information, and historical dialogue information input by the user, so that the generated decision result has a more anthropomorphic effect, and output the generated target dialogue reply information in a multi-modal form, improving the freedom and immersion when the user dialogues with the character role. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required to be used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention, and therefore should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.
[0023] Figure 1 It is a flowchart of a multi-modal dialogue processing method provided by an embodiment of the present application;
[0024] Figure 2 It is a flowchart of another multi-modal dialogue processing method provided by an embodiment of the present application;
[0025] Figure 3 It is a flowchart of another multi-modal dialogue processing method provided by an embodiment of the present application;
[0026] Figure 4 It is a flowchart of another multi-modal dialogue processing method provided by an embodiment of the present application;
[0027] Figure 5 It is a flowchart of another multi-modal dialogue processing method provided by an embodiment of the present application;
[0028] Figure 6 It is a flowchart of another multi-modal dialogue processing method provided by an embodiment of the present application;
[0029] Figure 7 A flowchart of another multi-modal dialogue processing method provided by an embodiment of the present application;
[0030] Figure 8 A flowchart of another multi-modal dialogue processing method provided by an embodiment of the present application;
[0031] Figure 9 A flowchart of a decision result obtained based on a large language model provided by an embodiment of the present application;
[0032] Figure 10 A structural diagram of a multi-modal dialogue processing device provided by an embodiment of the present application;
[0033] Figure 11 A structural diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners
[0034] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. It should be understood that the accompanying drawings in the present application are only for the purpose of illustration and description, and are not used to limit the protection scope of the present application. In addition, it should be understood that the schematic drawings are not drawn to scale. The flowcharts used in the present application illustrate operations implemented according to some embodiments of the present application. It should be understood that the operations in the flowcharts may not be implemented in sequence, and steps without logical context relationships may be reversed or implemented simultaneously. In addition, those skilled in the art may add one or more other operations to the flowchart or remove one or more operations from the flowchart under the guidance of the content of the present application.
[0035] In addition, the described embodiments are only some embodiments of the present application, rather than all embodiments. The components of the embodiments of the present application usually described and illustrated in the accompanying drawings here may be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present application provided in the accompanying drawings is not intended to limit the scope of the present application claimed, but merely represents selected embodiments of the present application. All other embodiments obtained by those skilled in the art based on the embodiments of the present application without creative efforts fall within the protection scope of the present application.
[0036] In order to enable those skilled in the art to use the content of this application, the following embodiments are given in combination with a specific application scenario, "simulation of soft bodies in game animations". For those skilled in the art, without departing from the spirit and scope of this application, the general principles defined here can be applied to other embodiments and application scenarios. Although this application mainly focuses on the simulation of soft bodies in game animations, it should be understood that this is only an exemplary embodiment. This application can be applied to any other scenario.
[0037] It should be noted that the term "including" will be used in the embodiments of this application to indicate the existence of the features stated thereafter, but does not exclude the addition of other features.
[0038] First, an introduction to the prior art related to this application is given.
[0039] In the related art, AI character conversations in cultural and entertainment scenarios mostly reflect the characteristics of characters from a textual perspective. Generally, based on the character setting text information and conversation history of the AI character, a text generation model is used to generate responses.
[0040] However, the above-mentioned character conversation method lacks support for multi-modal information input and output. In actual situations, human interaction is often in the form of multi-modal information, including text, pictures, voices, expressions, etc.; that is, the above-mentioned character conversation method in the related art lacks the understanding and generation of multi-modal information, resulting in the lack of anthropomorphism in the finally output response information, thus leading to low information interaction efficiency and poor user experience.
[0041] Based on the above problems, this application proposes a multi-modal conversation processing method, which supports users to input multi-modal conversation information, and the large speech model deeply analyzes the character setting information, historical conversations and conversation information of the character conversing with the user to obtain the strategy result of the multi-modal conversation. The strategy result is used to indicate which type of response information and the prompt words corresponding to the response information to call; then, according to the strategy result, the specified type of response information is called, and based on the prompt words corresponding to the response information, the target conversation response information is generated; finally, the target conversation response information is output in a multi-modal form. Therefore, this solution supports the input / output of multi-modal conversation information, and more accurately understands the user's conversation intention based on the character setting information, historical conversations and conversation information, making the generated decision result more anthropomorphic, and outputting the generated target conversation response information in a multi-modal form, improving the freedom and immersion of the user when conversing with the character.
[0042] The multi-modal conversation processing method provided by this application will be described through the following embodiments.
[0043] Figure 1Flow diagram of the multi-modal dialogue processing method provided by the embodiments of the present application Figure 1 ; Optionally, the execution subject of this method may be an electronic device with data processing capabilities. It should be noted that the multi-modal dialogue processing method provided by the present application is not limited by Figure 1 the specific order described below
[0044] It should be understood that in other embodiments, the order of some steps of the multi-modal dialogue processing method provided by the present application can be interchanged according to actual needs, or some of the steps can also be omitted or deleted. Refer to Figure 1 , this method includes:
[0045] S101. Obtain the dialogue information input by the user
[0046] Among them, the dialogue information is information of multiple modalities. Exemplarily, the dialogue information may be text information, picture information, voice information, expression information, etc., that is, the types of dialogue information are in different forms
[0047] Optionally, in order to support the input of multi-modal information, the type of the dialogue information input by the user may be in the form of multi-modal information, including: text type, picture type, voice type, expression type, etc. Exemplarily, the dialogue information input by the user is: What's your father's name, smile (expression), a landscape photo, etc. That is, the dialogue scenario provided by the present application is similar to the scenario of real online chatting, and the user can input information of multiple modalities
[0048] At the same time, in order to facilitate the understanding of the multi-modal dialogue information input by the user, the input dialogue information may also be described in text form to obtain dialogue text information. Exemplarily, if the dialogue information input by the user is smile (expression), then the smile (expression) is described in text form to obtain the dialogue text information: smile
[0049] S102. Obtain the role setting information and historical dialogue information for the dialogue with the user
[0050] Among them, the role setting information refers to the attribute information of the character role who has a dialogue with the user, generally in text description. Exemplarily, the role setting information includes: name, personality characteristics, personal preferences, social background, social relationships, etc. For example, taking "Lin Daiyu" as the character role, the role setting information is as follows
[0051] Name: Lin Daiyu, the granddaughter of Jia Mu, the cousin of Jia Baoyu, and the Jia family is a prestigious noble family that has been prosperous for a hundred years
[0052] Personality characteristics: She has lived in the Jia Mansion all year round and developed a proud and aloof personality. Lin Daiyu is beautiful, but in poor health and always looks frail
[0053] Personal preferences: Lin Daiyu loves reading and writing poetry, and has a high literary talent. The most notable feature of her language system is elegance. In her daily conversation, she often uses "bookish" language, is good at quoting beautiful sentences from ancient texts, and always speaks in a semi-literary and semi-vernacular manner.
[0054] Social background: Lin Daiyu is a sentimental and sensitive person. Since she has been living under the care of others since she was a child, she is often hurt by the words or actions of others.
[0055] Social relations: In daily life, she is insightful about the ways of the world, but she can only sigh and lament when faced with difficulties. In love, she feels threatened by the "gold and jade theory" of Jia Baoyu and Xue Baochai, and she always acts like a spoiled child to test Jia Baoyu's sincerity. Lin Daiyu likes to think independently and is also very rebellious. In the feudal society where "women without talent are virtuous", she is not only talented, but also dares to appreciate works criticized as "obscene lyrics and songs", and she never persuades Jia Baoyu to take the "economic career path" and does not care about fame and fortune.
[0056] The historical conversation information refers to the log records generated during the conversation between the user and the character, which is in text form. An example is as follows:
[0057] Lin Daiyu: (sighs lightly) There are so many people in this world!
[0058] User: Why do you say that?
[0059] Lin Daiyu: Look at the streets bustling with people, but how many of them can truly understand the meaning of life?
[0060] User: What do you think is the meaning of life?
[0061] Lin Daiyu: The meaning of life lies in the pursuit of true freedom and happiness.
[0062] User: How did you pursue it?
[0063] Lin Daiyu: I often go for walks in the mountains and forests alone, quietly thinking about the philosophy of life and looking for inner peace.
[0064] The dialogue content of "Lin Daiyu" is generated by the AI model of the character dialogue system, and the dialogue content of "user" is the dialogue content input by the user. The dialogue history information is obtained by sorting the multiple dialogue contents generated when the user and the character talk in chronological order.
[0065] S103: Determine the user's dialogue state information according to the role setting information, historical dialogue information and dialogue information.
[0066] Among them, the dialogue state information is used to represent the dialogue state between the user and the character. That is, based on the dialogue state information, the character preference setting, multiple dialogue history records generated in the multi-round dialogue, etc. can be determined, so as to ensure that the dialogue reply content generated in the next round of dialogue has better coherence.
[0067] S104. Input the dialogue state information into the pre-trained large language model to obtain the decision result output by the large language model.
[0068] Among them, the decision result is used to indicate the type of the reply information and the prompt word corresponding to the reply information.
[0069] In an implementable manner, for example, the obtained dialogue state information can be directly input into the large language model. The large language model performs intention recognition on the dialogue text information in the dialogue state information to obtain the user's dialogue intention, and based on the user's dialogue intention, obtains the type of the reply information; and, according to the user's dialogue intention, historical dialogue information and character setting information, obtains the prompt word corresponding to the reply information. For example, the converted dialogue text information is: Show me what you look like; after the large language model performs intention recognition on the converted dialogue text information, the type of the reply information obtained is the picture type; then, according to the user's dialogue intention, historical dialogue information and character setting information, the prompt word corresponding to the reply information is: A photo of Lin Daiyu. In this way, it can be ensured that the obtained decision result has a more anthropomorphic effect.
[0070] Specifically, the decision result output by the large language model can be expressed as {"decide": "image", "prompt": " <prompt>"}. Among them, "decide" indicates the type of the reply information, and the type of the reply information is "image", that is, the image generation type; "prompt" indicates the prompt word corresponding to the reply information, that is <prompt>It is the name of the photo to be generated. For example, the photo name is "Photo of Lin Daiyu".
[0071] For another example, the decision result output by the large language model can also be {"decide": "music", "song_name": <song_name>}. Among them, "decide" indicates the type of the reply information, and the type of the reply information is "music", that is, the music generation type; "song_name" indicates the prompt word corresponding to the reply information, that is, <song_name> is the name of the song to be played.
[0072] S105. Generate and output the target dialogue reply information according to the type of the reply information and the prompt word corresponding to the reply information.
[0073] Among them, the target dialogue reply information is information of multiple modalities. Exemplarily, the target dialogue reply information can be text reply information, picture reply information, voice reply information or emoji reply information, etc., that is, the types of the dialogue reply information are also in different forms.
[0074] In an implementable manner, for example, taking the decision result output by the large language model as {"decide": "image", "prompt": "Photo of Lin Daiyu"} as an example, the image generation type indicated in the decision result is called, and "Photo of Lin Daiyu" is used as the prompt word and input to the image generation type to generate the corresponding target picture, and the target picture is returned to the user as the target dialogue reply information, which improves the freedom and immersion when the user converses with the character role.
[0075] In summary, the embodiment of the present application provides a multi-modal dialogue processing method, which includes: obtaining the dialogue information input by the user; obtaining the role setting information and historical dialogue information of the character conversing with the user; determining the dialogue status information of the user according to the role setting information, historical dialogue information and dialogue information; inputting the dialogue status information into a pre-trained large language model to obtain the decision result output by the large language model, and the decision result is used to indicate the type of the reply information and the prompt word corresponding to the reply information; generating and outputting the target dialogue reply information according to the type of the reply information and the prompt word corresponding to the reply information. In this solution, in order to solve the problem that the related technology does not support the input / output of multi-modal information, it is proposed that after obtaining the multi-modal dialogue information input by the user, and based on the dialogue information, role setting information and historical dialogue information, the dialogue status information of the user is determined, and the large language model is used to understand the dialogue intention of the dialogue information, role setting information and historical dialogue information input by the user, so that the generated decision result has a more anthropomorphic effect, and the generated target dialogue reply information is output in a multi-modal form, which improves the freedom and immersion when the user converses with the character role.
[0076] Optionally, step S103 above includes:
[0077] Convert the dialogue information to obtain dialogue text information; determine the user's dialogue status information according to the role setting information, historical dialogue information, and dialogue text information.
[0078] In an implementable manner, in order to improve the processing efficiency of the multimodal dialogue information input by the user, it is proposed that the multimodal dialogue information can be first uniformly converted into dialogue text information, and then based on the role setting information, historical dialogue information, and dialogue text information, the user's dialogue status information can be further determined.
[0079] Optionally, referring to Figure 2 As shown, the above step of converting the dialogue information to obtain dialogue text information includes:
[0080] S201. If the type of the dialogue information is not the text type, determine the target conversion model according to the type of the dialogue information.
[0081] S202. Use the target conversion model to convert the dialogue information to obtain dialogue text information.
[0082] Optionally, in order to facilitate the processing of the multimodal dialogue information input by the user, the dialogue information is uniformly parsed into a description of the text content. Among them, if the type of the dialogue information is the text type, no conversion processing is required; if the type of the dialogue information is not the text type, it needs to be converted into text content.
[0083] Specifically, the target conversion model can be determined according to the type of the dialogue information. For example, if the dialogue information is a piece of voice data, that is, the type of the dialogue information is the voice type, the target conversion model can be determined as the speech recognition model, and the speech recognition model converts this piece of voice data into text, and the converted text is used as the dialogue text information. Thus, the parsing of the input multimodal information is realized.
[0084] Optionally, referring to Figure 3 As shown, the above step S201 includes:
[0085] S301. If the type of the dialogue information is the expression type or the picture type, determine that the target conversion model corresponding to the dialogue information is the picture conversion model.
[0086] The above step S202 includes:
[0087] S302. Input the dialogue information into the picture conversion model, and the picture conversion model describes the dialogue information in text form to obtain dialogue text information.
[0088] In an implementable manner, for example, if the dialogue information input by the user is a picture, that is, the type of the dialogue information is the picture type, the target conversion model can be determined to be the picture conversion type; input the picture into the picture conversion type, that is, the picture conversion type describes the picture content shown in the picture in text form to obtain the dialogue text information. The text description example is "This picture depicts a dog running cheerfully on the grass", which realizes the conversion processing of different-modal dialogue information input by the user and solves the problem that multi-modal information input is not supported in the related technology.
[0089] Optionally, referring to Figure 4 As shown, the above step S104 includes:
[0090] S401. Perform semantic analysis on the dialogue text information to obtain the semantic result corresponding to the dialogue text information.
[0091] S402. Determine the type of the reply information according to the semantic result.
[0092] S403. Determine the prompt word corresponding to the reply information according to the type of the reply information, the role setting information, the historical dialogue information, and the dialogue text information.
[0093] Optionally, continue with an example. For example, the dialogue text information input by the user is: Show me what you look like. After the large language model performs semantic analysis on the dialogue text information of "Show me what you look like", the obtained semantic result is: Give, me, show, you, of, look; and it is determined that the semantics of "look" is similar to words such as "photo" in the semantic library, then the type of the reply information can be further determined to be the picture generation type, that is, "decide": "image"; then, according to the role setting information, it is determined that "me" in the semantic result is the user and "you" is Lin Daiyu, then the prompt word corresponding to the reply information can be obtained as the photo of Lin Daiyu.
[0094] Another example, the dialogue text information input by the user is: What's your father's name. After the large language model performs semantic analysis on the dialogue text information of "What's your father's name", the obtained semantic result is: You, father, call, what, name; and it is determined that the semantics of "name" is similar to words such as "surname" in the semantic library, then the type of the reply information can be further determined to be the text generation type, that is, "decide": "text"; then, according to the role setting information and the historical dialogue information, it is determined that "me" in the semantic result is the user and "you" is Lin Daiyu, then the prompt word corresponding to the reply information can be obtained as the name of Lin Daiyu's father.
[0095] Optionally, referring to Figure 5 As shown, the above step S403 includes:
[0096] S501. If the type of the reply information is text type, determine the target query library according to the role setting information, historical conversation information, and conversation text information.
[0097] Among them, the target query library includes: knowledge base or memory bank.
[0098] Optionally, in this solution, in order to meet the requirement of having multi-round conversations with users that conform to the role settings, having long-term interaction memories with users, searching and replying based on external knowledge bases, and having rich modal output forms such as text, pictures, and expressive actions in the interaction with users and other conversation requirements, it is also possible to generate conversation reply information with the help of the target query library.
[0099] Among them, the knowledge base records more information about the character role settings. Exemplarily, the knowledge base of Lin Daiyu includes:
[0100] 1. Lin Daiyu's father is Lin Ruhai, a descendant of the Lin family in Gusu, surnamed Lin, named Ruhai, with the courtesy name Ruhai. His native place is Gusu.
[0101] 2. Lin Daiyu's mother is Jia Min, who is the daughter of Jia Daishan and Lady Shi.
[0102] 3. Jia Baoyu and Lin Daiyu are cousins. Jia Baoyu's father, Jia Zheng, and Lin Daiyu's mother, Jia Min, are siblings, so Jia Baoyu is Lin Daiyu's cousin.
[0103] The memory bank records the interaction information between the user and the character role. Exemplarily, the memory bank of Lin Daiyu includes:
[0104] 1. The user and Lin Daiyu made an appointment to go to the temple to pray for blessings.
[0105] 2. The user chatted with Lin Daiyu about poetry, songs, and odes.
[0106] 3. The user thought that Lin Daiyu was too pessimistic.
[0107] S502. Generate a prompt word corresponding to the reply information according to the target query library.
[0108] Among them, the prompt word is used to indicate the target query library.
[0109] Among them, if the type of the reply information is text type, the decision result output by the large language model is: {"decide":"text","knowledge_query":<knowledge_query>,"memory_query":<memory_query>}. Among them, <knowledge_query> and <memory_query> are prompt words.
[0110] 1. Based on the role setting information, historical dialogue information, and dialogue text information, determine whether it is necessary to retrieve information through the knowledge base. If so, <knowledge_query> is the retrieval input generated by the large language model, which is also called the prompt word corresponding to the reply information; if not, <knowledge_query> is an empty string, that is, the current dialogue state does not require knowledge retrieval.
[0111] Exemplarily, take the following historical dialogue as an example:
[0112] Lin Daiyu: (Sighing softly) There are so many people in this world!
[0113] User: Why do you say so?
[0114] Lin Daiyu: Look at the bustling people on the street. But how many of them can truly understand the meaning of life?
[0115] User: What do you think the meaning of life is?
[0116] Lin Daiyu: The meaning of life lies in the pursuit of true freedom and happiness, I suppose.
[0117] User: What's your father's name?
[0118] That is to say, the current dialogue text information input by the user is: What's your father's name. Then, based on the role setting information, historical dialogue information, and dialogue text information, it is determined that it is necessary to retrieve "Lin Daiyu's father's name" through the knowledge base. That is, it can be determined that the target query library is the knowledge base, and "knowledge_query": Lin Daiyu's father's name.
[0119] 2. Based on the role setting information, historical dialogue information, and dialogue text information, determine whether it is necessary to retrieve information through the memory bank. If so, <memory_query> is the retrieval input generated by the large language model, which is also called the prompt word corresponding to the reply information; if not, <memory_query> is an empty string, that is, the current dialogue state does not require memory retrieval.
[0120] Exemplarily, take the following historical dialogue as an example:
[0121] Lin Daiyu: (Sighing softly) There are so many people in this world!
[0122] User: Why do you say so?
[0123] Lin Daiyu: Look at the bustling people on the street. But how many of them can truly understand the meaning of life?
[0124] User: What do you think the meaning of life is?
[0125] Lin Daiyu: The meaning of life lies in the pursuit of true freedom and happiness, right?
[0126] User: Do you still remember our agreement?
[0127] That is, the current input dialogue text information of the user is: Do you still remember our agreement? Then, according to the role setting information, historical dialogue information, and dialogue text information, it is necessary to retrieve "the agreement between Lin Daiyu and the user" through the memory library. That is, it can be determined that the target query library is the memory library, and "knowledge_query": the agreement between Lin Daiyu and the user.
[0128] Optionally, referring to Figure 6 As shown, the above step S105 includes:
[0129] S601. If the type of the reply information is text type, and the prompt word indicates that the target query library is the knowledge library, then call the text generation model, use the prompt word as the retrieval condition, and screen out the target retrieval results that meet the preset conditions from the knowledge library.
[0130] Continuing with the example, based on the above embodiment, the decision result output by the large language model is: {"decide": "text", "knowledge_query": "the name of Lin Daiyu's father", "memory_query": NULL}. Then call the text generation model, and use "the name of Lin Daiyu's father" as the retrieval condition. The target retrieval result screened out from the knowledge library is: "Lin Daiyu's father is Lin Ruhai, a descendant of the Lin family in Gusu, surnamed Lin, named Hai, with the courtesy name Ruhai, originally from Gusu".
[0131] S602. If the type of the reply information is text type, and the prompt word indicates that the target query library is the memory library, then call the text generation model, use the prompt word as the retrieval condition, and screen out the target retrieval results that meet the preset conditions from the memory library.
[0132] Based on the above embodiment, the decision result output by the large language model is: {"decide": "text", "knowledge_query": NULL, "memory_query": "the agreement between Lin Daiyu and the user"}. Then call the text generation model, and use "the agreement between Lin Daiyu and the user" as the retrieval condition. The target retrieval result screened out from the memory library is: "The user and Lin Daiyu agreed to go to the temple to pray for blessings together".
[0133] It should be noted that for knowledge retrieval and memory retrieval, the specific implementation process of retrieving from the knowledge library (or memory library) is as follows:
[0134] 1. Convert each piece of knowledge in the knowledge base (or memory bank) into a vector representation through a vector encoding model, i.e., obtain a knowledge vector library (or knowledge or memory vector library).
[0135] 2. Use the same vector encoding model to convert the retrieval condition into a retrieval input vector.
[0136] 3. Calculate the similarity between the retrieval input vector and all vectors in the knowledge vector library, and filter out the target retrieval results that meet the similarity threshold.
[0137] S603. Generate and output target dialogue reply information according to the target retrieval results.
[0138] In an implementable manner, the target dialogue reply information can be generated and output according to the target retrieval results and the dialogue status information. For example, the target dialogue reply information is: Go to the temple to pray for blessings together, and the target dialogue reply information is output in text form.
[0139] Optionally, for the multi-modal dialogue processing method provided in this application, in order to ensure the coherence of the dialogue content generated during the multi-round dialogue between the user and the character role, it is proposed that the retrieval condition can be used to retrieve the target dialogue reply information from the knowledge base and the memory bank, which is more in line with the best reply that sometimes needs to be obtained with the help of memory or knowledge during the communication process between people, achieving a more anthropomorphic effect.
[0140] Optionally, as shown in Figure 7 Step S603 above includes:
[0141] S701. Generate target dialogue text information according to the target retrieval results and the dialogue status information.
[0142] Optionally, the role setting information and historical dialogue information in the dialogue status information can be used to eliminate the irrelevant content in the target retrieval results. For example, if the target retrieval result is "The user and Lin Daiyu agreed to go to the temple to pray for blessings together", then according to the dialogue status information, it can be determined that the objects participating in the dialogue include: the user and Lin Daiyu, and according to the historical dialogue information, it can be determined that the next dialogue status is to reply to the user, that is, it can be determined that the generated target dialogue text information is: Go to the temple to pray for blessings together.
[0143] S702. If the target dialogue text information includes an expression status description, select the target expression that matches the expression status description from the preset expression library.
[0144] S703. Generate target dialogue reply information according to the target expression and the target dialogue text information.
[0145] Optionally, if the generated target dialogue text information includes an expression status description, the text of the expression description is extracted (usually the content within parentheses in the reply text).
[0146] For example, the generated target dialogue text information is: "(Smile) I like playing games that require thinking and are full of challenges". Among them, "(Smile)" is the expression status description. According to the expression status description of "(Smile)", the target expression matching "(Smile)" is queried from the expression library: Smile. Among them, during the matching process, the similarity between the expression status description of the target expression and the expression status description of "(Smile)" reaches more than 0.7, and the closest expression is used as the target expression. Then, based on the target expression and the target dialogue text information, the target dialogue reply information is generated and output. That is, this solution supports the output of dialogue reply information in the form of expressions, achieving the effect of enhancing the interest and attractiveness of the reply.
[0147] In another implementable way, for example, the user can have a conversation with a non-player character or a robot character in the dialogue system. The non-player character or the robot character can be named "Xiaohua". During the conversation, specifically as follows:
[0148] The dialogue information input by the user is: Hello, Xiaohua, help me query what the surrounding tourist attractions are?
[0149] The dialogue reply information output by Xiaohua is: (Received), the tourist attractions include museums, shopping malls, etc.
[0150] Among them, "(Received)" is the received interactive expression status description.
[0151] Optionally, as shown in Figure 8 Step S603 includes:
[0152] S801. Generate target dialogue text information according to the target retrieval result and the dialogue status information.
[0153] S802. If the target dialogue text information includes an action status description, screen out the target action matching the action status description from the preset action library.
[0154] S803. Generate target dialogue reply information according to the target action and the target dialogue text information.
[0155] In another implementable way, if the generated target dialogue text information includes an action status description, the text of the action status description is extracted (usually the content within parentheses in the reply text). For example, the generated target dialogue text information is: "(Sighing softly) There are so many people in this world!". Among them, "(Sighing softly)" is the action status description. According to the action status description of "(Sighing softly)", the target action matching "(Sighing softly)" is queried from the action library as: smiling. Among them, during the matching process, the similarity between the action status description of the target action and that of "(Smiling)" reaches more than 0.7, and the closest action is used as the target action; then, based on the target action and the target dialogue text information, the target dialogue reply information is generated and output. That is to say, this solution supports the output of dialogue reply information in the form of actions.
[0156] In one implementable way, for example, a user can have a conversation with a non-player character or a robot character in a dialogue system. During the conversation, specifically as follows:
[0157] The dialogue information input by the user is: Hello, Xiaohua, sing a song?
[0158] The dialogue reply information output by Xiaohua is: (Dancing),,, etc.
[0159] Among them, "(Dancing)" is the action expression status description of dancing.
[0160] In another implementable way, after generating the target dialogue text information according to the target retrieval result and the dialogue status information, if the output format of the dialogue reply information input by the user is in audio format, the speech synthesis model can be called, and the target dialogue text information is input into the speech synthesis model. The speech synthesis model converts the target dialogue text information into a voice file, and the voice file is output as the target dialogue reply information.
[0161] Optionally, the above step S105 includes:
[0162] If the type of the reply information is a picture type, the picture generation model is called, the prompt is input into the picture generation model, and the target reply picture is obtained, and the target reply picture is used as the target dialogue reply information.
[0163] In one implementable way, if the decision result output by the large language model is: {"decide": "image", "prompt": "A photo of Lin Daiyu"}, the picture generation model is called, and "A photo of Lin Daiyu" is input into the picture generation model. The target reply picture is obtained, the target reply picture is used as the target dialogue reply information, and it is directly output.
[0164] Optionally, step S105 above includes:
[0165] If the type of the reply information is audio, call the audio generation model, use the prompt as the retrieval condition, screen out the target audio data from the preset music library, and use the target audio data as the target dialogue reply information.
[0166] In an implementable manner, if the decision result output by the large language model is: {"decide": "music", "song_name": "ABC"}, call the audio generation model, input the prompt "ABC" into the audio generation model. If the target audio data is screened out from the preset music library, use the target audio data as the target dialogue reply information; if the target audio data is not screened out, generate the target dialogue reply information as: I don't know how to sing this song yet.
[0167] In summary, the multi-modal dialogue processing method provided by this solution supports the output of multi-modal dialogue reply information, that is, the type of the output dialogue reply information can be text type, voice type, picture type, expression type, action type, etc., ensuring the richness of the dialogue content style when the user converses with the character role, achieving a more anthropomorphic effect, and meeting the needs of human-computer interaction in complex scenarios.
[0168] Optionally, as shown in Figure 9 the following is a schematic flowchart of the process for obtaining the decision result based on the large language model provided by the embodiment of the present application; this method includes:
[0169] The first step: Obtain the dialogue text information converted from the user input, the current historical dialogue between the user and the character, and the character setting information;
[0170] The second step: Input the dialogue text information converted from the user input, the current dialogue history between the user and the character, and the character setting into the function distribution engine trained based on the large language model to obtain the decision result output by the function distribution engine.
[0171] Among them, the decision result is used to indicate the type of the reply information and the prompt corresponding to the reply information.
[0172] The types of the reply information include: text generation model, picture generation model or music generation model.
[0173] Based on the same inventive concept, the embodiment of the present application also provides a multi-modal dialogue processing device corresponding to the multi-modal dialogue processing method. Since the principle of solving problems by the device in the embodiment of the present application is similar to that of the above multi-modal dialogue processing method in the embodiment of the present application, the implementation of the device can refer to the implementation of the method, and the repeated parts will not be described again.
[0174] Figure 10 Schematic diagram of the multimodal dialogue processing device provided by the embodiment of the present application, refer to Figure 10 As shown, the device includes:
[0175] An acquisition module 901, configured to acquire the dialogue information input by the user, where the dialogue information is information of multiple modalities; acquire the role setting information and historical dialogue information of the user's dialogue;
[0176] A determination module 902, configured to determine the dialogue status information of the user according to the role setting information, the historical dialogue information, and the dialogue information;
[0177] A decision-making module 903, configured to input the dialogue status information into a pre-trained large language model to obtain a decision result output by the large language model, where the decision result is used to indicate the type of the reply information and the prompt word corresponding to the reply information;
[0178] A generation module 904, configured to generate and output a target dialogue reply information according to the type of the reply information and the prompt word corresponding to the reply information, where the target dialogue reply information is information of multiple modalities.
[0179] Optionally, the determination module 902 is specifically configured to:
[0180] Convert the dialogue information to obtain dialogue text information;
[0181] Determine the dialogue status information of the user according to the role setting information, the historical dialogue information, and the dialogue text information.
[0182] Optionally, the determination module 902 is specifically configured to:
[0183] If the type of the dialogue information is not the text type, determine a target conversion model according to the type of the dialogue information;
[0184] Use the target conversion model to convert the dialogue information to obtain the dialogue text information.
[0185] Optionally, the determination module 902 is specifically configured to:
[0186] If the type of the dialogue information is the expression type or the picture type, determine that the target conversion model corresponding to the dialogue information is a picture conversion model;
[0187] The using the target conversion model to convert the dialogue information to obtain dialogue text information includes:
[0188] Input the conversation information into the image conversion model, and the image conversion model describes the conversation information in text form to obtain the conversation text information.
[0189] Optionally, the determining module 902 is specifically configured to:
[0190] Perform semantic analysis on the conversation text information to obtain a semantic result corresponding to the conversation text information;
[0191] Determine the type of the reply information according to the semantic result;
[0192] Determine a prompt word corresponding to the reply information according to the type of the reply information, the role setting information, the historical conversation information, and the conversation text information.
[0193] Optionally, the determining module 902 is specifically configured to:
[0194] If the type of the reply information is a text type, determine a target query library according to the role setting information, the historical conversation information, and the conversation text information, where the target query library includes: a knowledge base or a memory bank;
[0195] Generate a prompt word corresponding to the reply information according to the target query library, where the prompt word is used to indicate the target query library.
[0196] Optionally, the generating module 904 is specifically configured to:
[0197] If the type of the reply information is a text type, and the prompt word indicates that the target query library is a knowledge base, call a text generation model, use the prompt word as a retrieval condition, and screen out target retrieval results that meet preset conditions from the knowledge base;
[0198] If the type of the reply information is a text type, and the prompt word indicates that the target query library is a memory bank, call a text generation model, use the prompt word as a retrieval condition, and screen out target retrieval results that meet preset conditions from the memory bank;
[0199] Generate the target conversation reply information according to the target retrieval result.
[0200] Optionally, the generating module 904 is specifically configured to:
[0201] Generate target conversation text information according to the target retrieval result and the conversation status information;
[0202] If the target conversation text information includes an expression status description, screen out a target expression that matches the expression status description from a preset expression library;
[0203] Generate the target dialogue response information according to the target expression and the target dialogue text information.
[0204] Optionally, the generating module 904 is specifically configured to:
[0205] Generate target dialogue text information according to the target retrieval result and the dialogue status information;
[0206] If the target dialogue text information includes an action status description, filter out a target action that matches the action status description from a preset action library;
[0207] Generate and output the target dialogue response information according to the target action and the target dialogue text information.
[0208] Optionally, the generating module 904 is specifically configured to:
[0209] If the type of the response information is a picture type, call the picture generation model, input the prompt into the picture generation model to obtain a target response picture, and use the target response picture as the target dialogue response information.
[0210] Optionally, the generating module 904 is specifically configured to:
[0211] If the type of the response information is an audio type, call the audio generation model, use the prompt as a retrieval condition to filter out target audio data from a preset music library, and use the target audio data as the target dialogue response information.
[0212] The above device is used to execute the method provided in the foregoing embodiment, and its implementation principle and technical effects are similar, which will not be elaborated here.
[0213] The above modules may be one or more integrated circuits configured to implement the above methods, such as: one or more Application Specific Integrated Circuits (ASICs), or, one or more digital signal processors (DSPs), or, one or more Field Programmable Gate Arrays (FPGAs), etc. For another example, when a certain module above is implemented in the form of a processing element scheduler code, the processing element may be a general-purpose processor, such as a Central Processing Unit (CPU) or other processors that can call program code. For another example, these modules may be integrated together and implemented in the form of a system-on-a-chip (SOC).
[0214] Figure 11 FIG. is a schematic structural diagram of an electronic device provided by an embodiment of the present application. The electronic device includes: a processor 1001, a storage medium 1002, and a bus 1003. The storage medium 1002 stores machine-readable instructions executable by the processor 1001. When the electronic device runs, the processor 1001 communicates with the storage medium 1002 through the bus 1003. The processor 1001 executes the machine-readable instructions to perform the following steps:
[0215] Obtain the conversation information input by the user, where the conversation information is multi-modal information;
[0216] Obtain the role setting information and historical conversation information of the conversation with the user;
[0217] Determine the conversation status information of the user according to the role setting information, the historical conversation information, and the conversation information;
[0218] Input the conversation status information into a pre-trained large language model to obtain a decision result output by the large language model. The decision result is used to indicate the type of the reply information and the prompt word corresponding to the reply information;
[0219] Generate and output a target conversation reply information according to the type of the reply information and the prompt word corresponding to the reply information. The target conversation reply information is multi-modal information.
[0220] Optionally, when the processor 1001 executes to determine the conversation status information of the user according to the role setting information, the historical conversation information, and the conversation information, it is specifically used for:
[0221] Convert the conversation information to obtain conversation text information;
[0222] Determine the conversation status information of the user according to the role setting information, the historical conversation information, and the conversation text information.
[0223] Optionally, when the processor 1001 executes the conversion of the conversation information to obtain conversation text information, it is specifically used for:
[0224] If the type of the conversation information is not the text type, determine the target conversion model according to the type of the conversation information;
[0225] Use the target conversion model to convert the conversation information to obtain the conversation text information.
[0226] Optionally, when the processor 1001 executes the determination of the target conversion model according to the conversation type of the conversation information, it is specifically used for:
[0227] If the type of the conversation information is the expression type or the picture type, determine that the target conversion model corresponding to the conversation information is the picture conversion model.
[0228] The use of the target conversion model to convert the conversation information to obtain conversation text information is specifically used for:
[0229] Input the conversation information into the picture conversion model, and the picture conversion model describes the conversation information in text form to obtain conversation text information.
[0230] Optionally, when the processor 1001 executes the input of the conversation status information into the pre-trained large language model to obtain the decision result output by the large language model, it is specifically used for:
[0231] Perform semantic analysis on the conversation text information to obtain the semantic result corresponding to the conversation text information;
[0232] Determine the type of the reply information according to the semantic result;
[0233] Determine the prompt word corresponding to the reply information according to the type of the reply information, the role setting information, the historical conversation information, and the conversation text information.
[0234] Optionally, when the processor 1001 executes the determination of the prompt word corresponding to the reply information according to the type of the reply information, the character role setting information, the historical conversation information, and the conversation text information, it is specifically used for:
[0235] If the type of the reply information is text type, determine a target query library according to the role setting information, the historical conversation information, and the conversation text information, where the target query library includes: a knowledge base or a memory bank;
[0236] Generate a prompt word corresponding to the reply information according to the target query library, where the prompt word is used to indicate the target query library.
[0237] Optionally, when the processor 1001 executes generating and outputting a target conversation reply information according to the type of the reply information and the prompt word corresponding to the reply information, it is specifically used for:
[0238] If the type of the reply information is text type and the prompt word indicates that the target query library is a knowledge base, call a text generation model, use the prompt word as a retrieval condition, and screen out target retrieval results that meet preset conditions from the knowledge base;
[0239] If the type of the reply information is text type and the prompt word indicates that the target query library is a memory bank, call a text generation model, use the prompt word as a retrieval condition, and screen out target retrieval results that meet preset conditions from the memory bank;
[0240] Generate and output the target conversation reply information according to the target retrieval results.
[0241] Optionally, when the processor 1001 executes generating the target conversation reply information according to the target retrieval results, it is specifically used for:
[0242] Generate target conversation text information according to the target retrieval results and the conversation status information;
[0243] If the target conversation text information includes an expression status description, screen out a target expression that matches the expression status description from a preset expression library;
[0244] Generate the target conversation reply information according to the target expression and the target conversation text information.
[0245] Optionally, when the processor 1001 executes generating the target conversation reply information according to the target retrieval results, it is specifically used for:
[0246] Generate target conversation text information according to the target retrieval results and the conversation status information;
[0247] If the target conversation text information includes an action status description, screen out a target action that matches the action status description from a preset action library;
[0248] Generate the target dialogue reply information according to the target action and the target dialogue text information.
[0249] Optionally, when the processor 1001 executes generating the target dialogue reply information according to the type of the reply information and the prompt word corresponding to the reply information, it is specifically used for:
[0250] If the type of the reply information is a picture type, call the picture generation model, input the prompt word into the picture generation model, obtain the target reply picture, and use the target reply picture as the target dialogue reply information.
[0251] Optionally, when the processor 1001 executes generating the target dialogue reply information according to the type of the reply information and the prompt word corresponding to the reply information, it is specifically used for:
[0252] If the type of the reply information is an audio type, call the audio generation model, use the prompt word as the retrieval condition, screen out the target audio data from the preset music library, and use the target audio data as the target dialogue reply information.
[0253] Optionally, the present invention further provides a program product, such as a computer-readable storage medium, including a program, which is used for execution when being executed by a processor. The processor executes the following steps:
[0254] Obtain the dialogue information input by the user, where the dialogue information is information of multiple modalities;
[0255] Obtain the role setting information and the historical dialogue information of the user's dialogue;
[0256] Determine the dialogue status information of the user according to the role setting information, the historical dialogue information and the dialogue information;
[0257] Input the dialogue status information into a pre-trained large language model to obtain a decision result output by the large language model, where the decision result is used to indicate the type of the reply information and the prompt word corresponding to the reply information;
[0258] Generate and output the target dialogue reply information according to the type of the reply information and the prompt word corresponding to the reply information, where the target dialogue reply information is information of multiple modalities.
[0259] Optionally, when the processor executes determining the dialogue status information of the user according to the role setting information, the historical dialogue information and the dialogue information, it is specifically used for:
[0260] Convert the dialogue information to obtain dialogue text information;
[0261] Determine the conversation status information of the user according to the character setting information, the historical conversation information, and the conversation text information.
[0262] Optionally, when the processor executes the conversion of the conversation information to obtain conversation text information, it is specifically used for:
[0263] If the type of the conversation information is not the text type, determine the target conversion model according to the type of the conversation information;
[0264] Use the target conversion model to convert the conversation information to obtain the conversation text information.
[0265] Optionally, when the processor executes the determination of the target conversion model according to the conversation type of the conversation information, it is specifically used for:
[0266] If the type of the conversation information is the emoji type or the picture type, determine that the target conversion model corresponding to the conversation information is the picture conversion model.
[0267] The use of the target conversion model to convert the conversation information to obtain conversation text information is specifically used for:
[0268] Input the conversation information into the picture conversion model, and the picture conversion model describes the conversation information in text form to obtain the conversation text information.
[0269] Optionally, when the processor executes the input of the conversation status information into the pre-trained large language model to obtain the decision result output by the large language model, it is specifically used for:
[0270] Perform semantic analysis on the conversation text information to obtain the semantic result corresponding to the conversation text information;
[0271] Determine the type of the reply information according to the semantic result;
[0272] Determine the prompt word corresponding to the reply information according to the type of the reply information, the character setting information, the historical conversation information, and the conversation text information.
[0273] Optionally, when the processor executes the determination of the prompt word corresponding to the reply information according to the type of the reply information, the character role setting information, the historical conversation information, and the conversation text information, it is specifically used for:
[0274] If the type of the reply information is the text type, determine the target query library according to the character setting information, the historical conversation information, and the conversation text information. The target query library includes: the knowledge base or the memory library;
[0275] Generate a prompt word corresponding to the reply information according to the target query library, where the prompt word is used to indicate the target query library.
[0276] Optionally, when the processor executes generating and outputting a target dialogue reply information according to the type of the reply information and the prompt word corresponding to the reply information, it specifically is used for:
[0277] If the type of the reply information is text type and the prompt word indicates that the target query library is a knowledge base, call a text generation model, use the prompt word as a retrieval condition, and screen out a target retrieval result that meets a preset condition from the knowledge base;
[0278] If the type of the reply information is text type and the prompt word indicates that the target query library is a memory library, call a text generation model, use the prompt word as a retrieval condition, and screen out a target retrieval result that meets a preset condition from the memory library;
[0279] Generate and output the target dialogue reply information according to the target retrieval result.
[0280] Optionally, when the processor executes generating the target dialogue reply information according to the target retrieval result, it specifically is used for:
[0281] Generate target dialogue text information according to the target retrieval result and the dialogue state information;
[0282] If the target dialogue text information includes an expression state description, screen out a target expression that matches the expression state description from a preset expression library;
[0283] Generate the target dialogue reply information according to the target expression and the target dialogue text information.
[0284] Optionally, when the processor executes generating the target dialogue reply information according to the target retrieval result, it specifically is used for:
[0285] Generate target dialogue text information according to the target retrieval result and the dialogue state information;
[0286] If the target dialogue text information includes an action state description, screen out a target action that matches the action state description from a preset action library;
[0287] Generate the target dialogue reply information according to the target action and the target dialogue text information.
[0288] Optionally, when the processor executes generating the target dialogue reply information according to the type of the reply information and the prompt word corresponding to the reply information, it specifically is used for:
[0289] If the type of the reply information is a picture type, then call the picture generation model, input the prompt into the picture generation model to obtain a target reply picture, and use the target reply picture as the target dialogue reply information.
[0290] Optionally, when the processor executes generating the target dialogue reply information according to the type of the reply information and the prompt corresponding to the reply information, it is specifically configured to:
[0291] If the type of the reply information is an audio type, then call the audio generation model, use the prompt as a retrieval condition to screen out target audio data from a preset music library, and use the target audio data as the target dialogue reply information.
[0292] In several embodiments provided by the present invention, it should be understood that the disclosed device and method can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of the device or unit can be in an electrical, mechanical or other form.
[0293] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0294] In addition, each functional unit in various embodiments of the present invention can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated unit can be implemented in the form of hardware, or in the form of a hardware plus a software functional unit.
[0295] The integrated unit implemented in the form of software functional units can be stored in a computer-readable storage medium. The above-mentioned software functional units are stored in a storage medium, including several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) or a processor (English: processor) to execute some steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (English: Read-Only Memory, abbreviated as: ROM), random access memories (English: Random Access Memory, abbreviated as: RAM), magnetic disks, or optical discs.< / prompt> < / prompt>
Claims
1. A multimodal dialogue processing method, characterized in that The method includes: Obtaining the conversation information input by the user, where the conversation information is information of multiple modalities; Obtaining the role setting information and historical conversation information of the conversation with the user; Determining the conversation status information of the user according to the role setting information, the historical conversation information, and the conversation information; Inputting the conversation status information into a pre-trained large language model to obtain a decision result output by the large language model, where the decision result is used to indicate the type of the reply information and the prompt word corresponding to the reply information; Generating and outputting a target conversation reply information according to the type of the reply information and the prompt word corresponding to the reply information, where the target conversation reply information is information of multiple modalities.
2. The method according to claim 1, wherein The determining the conversation status information of the user according to the role setting information, the historical conversation information, and the conversation information includes: Converting the conversation information to obtain conversation text information; Determining the conversation status information of the user according to the role setting information, the historical conversation information, and the conversation text information.
3. The method according to claim 2, wherein The converting the conversation information to obtain conversation text information includes: If the type of the conversation information is not the text type, determining a target conversion model according to the type of the conversation information; Using the target conversion model to convert the conversation information to obtain the conversation text information.
4. The method according to claim 2, wherein The determining a target conversion model according to the conversation type of the conversation information includes: If the type of the conversation information is the expression type or the picture type, determining that the target conversion model corresponding to the conversation information is a picture conversion model; The using the target conversion model to convert the conversation information to obtain the conversation text information includes: Inputting the conversation information into the picture conversion model, and the picture conversion model describes the conversation information in text form to obtain the conversation text information.
5. The method according to claim 1, wherein The inputting the conversation status information into a pre-trained large language model to obtain a decision result output by the large language model includes: Performing semantic analysis on the conversation text information to obtain a semantic result corresponding to the conversation text information; Determining the type of the reply information according to the semantic result; Determining the prompt word corresponding to the reply information according to the type of the reply information, the role setting information, the historical conversation information, and the conversation text information.
6. The method according to claim 5, wherein The determining the prompt word corresponding to the reply information according to the type of the reply information, the character role setting information, the historical conversation information, and the conversation text information includes: If the type of the reply information is the text type, determining a target query library according to the role setting information, the historical conversation information, and the conversation text information, where the target query library includes: a knowledge base or a memory bank; Generating the prompt word corresponding to the reply information according to the target query library, where the prompt word is used to indicate the target query library.
7. The method according to claim 1, characterized in that, The generating a target conversation reply information according to the type of the reply information and the prompt word corresponding to the reply information includes: If the type of the reply information is text type and the prompt word indicates that the target query library is the knowledge base, call the text generation model, use the prompt word as the retrieval condition, and screen out the target retrieval results that meet the preset conditions from the knowledge base; If the type of the reply information is text type and the prompt word indicates that the target query library is the memory bank, call the text generation model, use the prompt word as the retrieval condition, and screen out the target retrieval results that meet the preset conditions from the memory bank; Generate the target dialogue reply information according to the target retrieval results.
8. The method according to claim 7, wherein The generating the target dialogue reply information according to the target retrieval results includes: Generate target dialogue text information according to the target retrieval results and the dialogue status information; If the target dialogue text information includes an expression status description, screen out the target expression that matches the expression status description from the preset expression library; Generate the target dialogue reply information according to the target expression and the target dialogue text information.
9. The method according to claim 7, wherein The generating the target dialogue reply information according to the target retrieval results includes: Generate target dialogue text information according to the target retrieval results and the dialogue status information; If the target dialogue text information includes an action status description, screen out the target action that matches the action status description from the preset action library; Generate the target dialogue reply information according to the target action and the target dialogue text information.
10. The method according to claim 1, wherein The generating the target dialogue reply information according to the type of the reply information and the prompt word corresponding to the reply information includes: If the type of the reply information is picture type, call the picture generation model, input the prompt word into the picture generation model, obtain the target reply picture, and use the target reply picture as the target dialogue reply information.
11. The method according to claim 1, wherein The generating the target dialogue reply information according to the type of the reply information and the prompt word corresponding to the reply information includes: If the type of the reply information is audio type, call the audio generation model, use the prompt word as the retrieval condition, screen out the target audio data from the preset music library, and use the target audio data as the target dialogue reply information.
12. A multimodal dialogue processing device, characterized in that, The device includes: An acquisition module, configured to acquire the dialogue information input by the user, where the dialogue information is information of multiple modalities; acquire the role setting information and historical dialogue information of the conversation with the user; A determination module, configured to determine the dialogue status information of the user according to the role setting information, the historical dialogue information, and the dialogue information; A decision module, configured to input the dialogue status information into a pre-trained large language model to obtain a decision result output by the large language model, where the decision result is used to indicate the type of the reply information and the prompt word corresponding to the reply information; A generation module, configured to generate and output target dialogue reply information according to the type of the reply information and the prompt word corresponding to the reply information, where the target dialogue reply information is information of multiple modalities.
13. An electronic device, characterized in that, including: A processor, a storage medium, and a bus, wherein the storage medium stores machine-readable instructions executable by the processor. When the electronic device is running, the processor communicates with the storage medium via the bus, and the processor executes the machine-readable instructions to perform the steps of the method according to any one of claims 1-11.
14. A computer-readable storage medium, characterized in that, A computer program is stored on the storage medium, and when the computer program is run by the processor, it executes the method according to any one of claims 1-11.
Citation Information
Cited By
Game question and answer method and device based on large language model
CN120632051A
Session processing method, system, device, product and medium
CN121603481A
Multi-modal digital human interaction method and system based on large language model
CN122334324A