Man-machine conversation processing method and device based on multi-modal message, equipment and medium

By introducing multimodal message processing methods into the human-computer dialogue system, identifying and processing non-text messages such as images and videos, the problem that existing systems cannot process multimodal messages is solved, and dialogue efficiency and accuracy are improved.

CN119938995APending Publication Date: 2025-05-06LINGXI TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510105888.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-23
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The existing human-computer dialogue system cannot recognize and process images, videos or other types of messages, resulting in users needing to summarize and edit text messages by themselves in these scenarios, and the system is difficult to ensure accurate reply, which reduces the efficiency of human-computer dialogue.

Method used

By introducing a multimodal message processing method in the human-computer dialogue system, the corresponding target mode message recognition model is determined based on the message modality input by the user, the message content is identified, and the accurate reply is generated in combination with historical message records.

Benefits of technology

The identification and processing of multimodal messages is realized, the efficiency and accuracy of human-computer dialogue are improved, and the input burden of users is reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119938995A_ABST
    Figure CN119938995A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a man-machine conversation processing method and device based on a multi-modal message, equipment and a medium, and relates to the technical field of artificial intelligence. The multi-modal message-based man-machine conversation processing method comprises the following steps: determining a target modal message identification model corresponding to a message modal according to the message modal of a first message input by a user; inputting the first message into the target modal message identification model to obtain an identification result; obtaining a target message record according to the first message, the identification result and a historical message record of the session; and inputting the target message record into a pre-constructed reply message generation model to obtain a second message, and returning the second message to the user. According to the embodiment of the invention, the technical effects of supporting recognition of multi-modal messages for man-machine conversation and improving the man-machine conversation efficiency can be achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence technology, and more specifically, to a method, device, equipment and medium for processing human-computer dialogue based on multimodal messages. Background Art

[0002] At present, human-computer dialogue systems are widely used in various fields to provide users with various services such as consultation and chat. Existing human-computer dialogue systems rely on technologies such as natural language understanding, and can usually only recognize text messages entered by users and determine reply messages based on text messages and historical messages.

[0003] Since the existing human-computer dialogue system cannot support the recognition of images, videos or other types of messages, in the scenario where the user initiates a dialogue regarding the content in an image, video or other type of message, not only is the user required to summarize the content in the image, video or other type of message and edit text-type messages and input them into the existing human-computer dialogue system, but the existing human-computer dialogue system is also difficult to guarantee accurate replies to messages summarized and edited by users themselves, greatly reducing the efficiency of human-computer dialogue. Summary of the invention

[0004] The purpose of the embodiments of the present application is to provide a method, device, equipment and medium for processing human-computer dialogue based on multimodal messages, so as to achieve the technical effect of supporting the recognition of multimodal messages for human-computer dialogue and improving the efficiency of human-computer dialogue.

[0005] In a first aspect, an embodiment of the present application provides a method for processing a human-computer dialogue based on a multimodal message, comprising:

[0006] Determining a target modality message recognition model corresponding to the message modality according to the message modality of the first message input by the user;

[0007] Inputting the first message into the target modal message recognition model to obtain a recognition result;

[0008] Obtaining a target message record according to the first message, the recognition result, and the historical message record of this conversation;

[0009] The target message record is input into a pre-built reply message generation model to obtain a second message, and the second message is returned to the user.

[0010] In the above implementation process, by determining the target modal message recognition model corresponding to the message modality according to the message modality of the first message input by the user, the first message is input into the target modal message recognition model to obtain a recognition result, and the target message record is obtained according to the first message, the recognition result and the historical message record of this conversation, and the target message record is input into a pre-built reply message generation model to obtain a second message, and the second message is returned to the user. The target modal message recognition model corresponding to the message modality of the first message can be used to specifically identify the content in the first message, and the reply message generation model can be used to comprehensively combine the first message, the recognition result and the historical message record of this conversation to accurately generate the second message, thereby supporting the recognition of multi-modal messages for human-computer dialogue and improving the efficiency of human-computer dialogue.

[0011] Further, the message modality is text, document, image, video or audio;

[0012] The determining, according to the message modality of the first message input by the user, a target modality message recognition model corresponding to the message modality includes:

[0013] In the case where the message modality is text, selecting a text recognition model from a plurality of pre-constructed modality message recognition models as the target modality message recognition model;

[0014] In the case where the message modality is a document, selecting a document recognition model from the multiple modality message recognition models as the target modality message recognition model;

[0015] When the message modality is an image, selecting an image recognition model from the multiple modality message recognition models as the target modality message recognition model;

[0016] When the message modality is a video, selecting a video recognition model from the multiple modality message recognition models as the target modality message recognition model;

[0017] When the message modality is audio, a speech recognition model is selected from the multiple modality message recognition models as the target modality message recognition model.

[0018] In the above implementation process, by pre-constructing multiple modal message recognition models, selecting the target modal message recognition model corresponding to the message modality of the first message from the multiple modal message recognition models, it can be ensured that the target modal message recognition model can accurately recognize the content in the first message, thereby supporting the recognition of multimodal messages for human-computer dialogue and improving the efficiency of human-computer dialogue.

[0019] Furthermore, the document recognition model includes a first language model GPT-4Turbo, and the image recognition model and the video recognition model include a second language model GPT-4o.

[0020] In the above implementation process, by selecting the first language model GPT-4Turbo to build a document recognition model, and selecting the second language model GPT-4o to build an image recognition model and a video recognition model, it is possible to ensure that when the message modality of the first message is a complex document, image or video, the content in the first message can be quickly and accurately identified, thereby supporting the recognition of multimodal messages for human-computer dialogue and improving the efficiency of human-computer dialogue.

[0021] Further, obtaining a target message record according to the first message, the recognition result and the historical message record of this conversation includes:

[0022] The first message, the recognition result and the historical message record are spliced ​​in a preset splicing order to obtain the target message record.

[0023] In the above implementation process, by splicing the first message, the recognition result and the historical message record of this conversation in a preset splicing order, the target message record is obtained, which can ensure that the reply message generation model accurately locates the first message, the recognition result and the historical message record of this conversation for processing, thereby supporting the recognition of multimodal messages for human-computer dialogue and improving the efficiency of human-computer dialogue.

[0024] Furthermore, before obtaining the target message record according to the first message, the recognition result and the historical message record of the current conversation, the method further includes:

[0025] The historical message record is obtained according to the current session identifier carried in the first message; wherein the current session identifier is generated when the current session is established and is stored in association with the historical message record.

[0026] In the above implementation process, by obtaining the historical message records of the current session stored in association with the current session identifier according to the current session identifier carried by the first message, it is possible to ensure that the historical message records of the current session are obtained quickly and accurately, thereby supporting the recognition of multimodal messages for human-computer dialogue and improving the efficiency of human-computer dialogue.

[0027] Furthermore, the reply message generation model includes a question extraction agent, an intention recognition agent and a message generation agent;

[0028] The step of inputting the target message record into a pre-built reply message generation model to obtain a second message includes:

[0029] Extracting at least one user question from the target message record by the question extraction agent;

[0030] identifying the user intent according to the at least one user question by the intent recognition agent;

[0031] An intelligent agent is generated through the message to determine an answer that meets the user's intention, and the second message is generated according to the answer.

[0032] In the above implementation process, by pre-building a reply message generation model including a question extraction agent, an intent recognition agent and a message generation agent, at least one user question is extracted from the target message record through the question extraction agent, the user intent is recognized based on at least one user question through the intent recognition agent, the answer that meets the user intent is determined through the message generation agent, and a second message is generated based on the answer. It is possible to coordinate multiple agents to finely identify the user intent in the target message record and determine the answer that meets the user intent to generate the second message, thereby supporting the recognition of multimodal messages for human-computer dialogue and improving the efficiency of human-computer dialogue.

[0033] Furthermore, the method further comprises:

[0034] Inserting the first message and the second message into a message queue;

[0035] The message queue is consumed, the historical message record is updated, and the updated historical message record is stored.

[0036] In the above implementation process, by inserting the first message and the second message into the message queue, consuming the message queue, updating the historical message record, and storing the updated historical message record, the historical message record can be updated and stored asynchronously, effectively reducing the pressure of human-computer dialogue processing, thereby supporting the recognition of multimodal messages for human-computer dialogue and improving the efficiency of human-computer dialogue.

[0037] In a second aspect, an embodiment of the present application provides a human-computer dialogue processing device based on multimodal messages, comprising:

[0038] A target model determination module, used to determine a target modality message recognition model corresponding to the message modality according to the message modality of the first message input by the user;

[0039] A first message recognition module, used for inputting the first message into the target modality message recognition model to obtain a recognition result;

[0040] A message record acquisition module, configured to obtain a target message record according to the first message, the recognition result and the historical message record of this conversation;

[0041] The second message returning module is used to input the target message record into a pre-built reply message generation model to obtain a second message, and return the second message to the user.

[0042] In a third aspect, an embodiment of the present application provides an electronic device, comprising a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor; when the processor executes the computer program, the method as described above is implemented.

[0043] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium includes a stored computer program; wherein, when the computer program is running, the device where the computer-readable storage medium is located is controlled to execute the method described above. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for use in the embodiments of the present application will be briefly introduced below. It should be understood that the following drawings only show certain embodiments of the present application and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without paying creative work.

[0045] Figure 1 A flowchart of a method for processing a human-computer dialogue based on multimodal messages provided in the first embodiment of the present application;

[0046] Figure 2 A schematic diagram of the structure of a human-computer dialogue processing device based on multimodal messages provided in the second embodiment of the present application;

[0047] Figure 3 A schematic diagram of the structure of an electronic device provided in the third embodiment of the present application. DETAILED DESCRIPTION

[0048] The technical solutions in the embodiments of the present application will be described below in conjunction with the drawings in the embodiments of the present application.

[0049] It should be noted that in the description of this application, the terms "first", "second", etc. are only used to distinguish descriptions and cannot be understood as indicating or implying relative importance. At the same time, the step numbers in the text are only for the convenience of explaining the embodiments of this application and do not serve to limit the order of execution of the steps. The method provided in the embodiments of this application can be executed by relevant terminal devices, and the following description will be given by taking a human-machine dialogue terminal equipped with a human-machine dialogue system as an example of the execution subject.

[0050] Please see Figure 1 , Figure 1The first embodiment of the present application provides a method for processing a human-computer dialogue based on multimodal messages, comprising steps S101 to S104:

[0051] S101, determining a target modality message recognition model corresponding to the message modality according to the message modality of a first message input by a user;

[0052] S102, inputting the first message into a target modal message recognition model to obtain a recognition result;

[0053] S103, obtaining a target message record according to the first message, the recognition result and the historical message record of this conversation;

[0054] S104: Input the target message record into a pre-built reply message generation model to obtain a second message, and return the second message to the user.

[0055] As an exemplary embodiment, the human-machine dialogue terminal can establish a session with any user according to actual business needs, and when the session is established, conduct multiple rounds of dialogue with the user to complete the session.

[0056] During each round of dialogue between the human-machine dialogue terminal and the user, the user is allowed to input multimodal messages to the human-machine dialogue terminal according to his or her own service needs.

[0057] It should be noted that multimodal messages refer to messages of various types, including text, documents, images, animated images, videos, and audio.

[0058] During each round of dialogue with the user, the human-computer dialogue terminal obtains the first message input by the user in this round, determines the message modality of the first message, and determines the target modality message recognition model corresponding to the message modality based on the message modality of the first message.

[0059] After determining the target modal message recognition model, the human-computer dialogue terminal inputs the first message into the target modal message recognition model, uses the target modal message recognition model to recognize the content in the first message, and obtains a recognition result.

[0060] After obtaining the recognition result and the historical message record of the current conversation, the human-machine dialogue terminal obtains the target message record according to the first message, the recognition result and the historical message record of the current conversation.

[0061] It should be noted that the historical message record of this conversation is used to store the message records of the historical rounds of conversation between the human-machine dialogue terminal and the user before the current round of conversation in this conversation.

[0062] After obtaining the target message record, the human-computer dialogue terminal inputs the target message record into a pre-built reply message generation model to obtain a second message, and returns the second message to the user to complete this round of dialogue.

[0063] In practical applications, a large model can be used as the reply message generation model. A large model refers to a machine learning model with large-scale parameters and complex computing structure, which can process multiple different types of data in parallel on a large scale.

[0064] The embodiment of the present application determines a target modal message recognition model corresponding to the message modality according to the message modality of the first message input by the user, inputs the first message into the target modal message recognition model to obtain a recognition result, obtains a target message record according to the first message, the recognition result and the historical message record of this conversation, inputs the target message record into a pre-built reply message generation model to obtain a second message, and returns the second message to the user. The target modal message recognition model corresponding to the message modality of the first message can be used to specifically identify the content in the first message, and the reply message generation model can be used to comprehensively combine the first message, the recognition result and the historical message record of this conversation to accurately generate the second message, thereby supporting the recognition of multi-modal messages for human-computer dialogue and improving the efficiency of human-computer dialogue.

[0065] In an optional embodiment, the message modality is text, document, image, video or audio; the target modality message recognition model corresponding to the message modality is determined according to the message modality of the first message input by the user, including: when the message modality is text, selecting a text recognition model from a plurality of pre-built modality message recognition models as the target modality message recognition model; when the message modality is a document, selecting a document recognition model from a plurality of modality message recognition models as the target modality message recognition model; when the message modality is an image, selecting an image recognition model from a plurality of modality message recognition models as the target modality message recognition model; when the message modality is a video, selecting a video recognition model from a plurality of modality message recognition models as the target modality message recognition model; when the message modality is an audio, selecting a speech recognition model from a plurality of modality message recognition models as the target modality message recognition model.

[0066] As an example, in order to ensure that the human-computer dialogue terminal accurately recognizes multimodal messages such as text, documents, images, videos and audio input by the user for human-computer dialogue, one or more modal message recognition models are pre-built for each type of message in multiple types. Specifically, for text type messages, one or more text recognition models are pre-built; for document type messages, one or more document recognition models are pre-built; for image type messages, one or more image recognition models are pre-built; for video type messages, one or more video recognition models are pre-built; and for audio type messages, one or more audio recognition models are pre-built, thereby pre-building multiple modal message recognition models.

[0067] In the process of each round of dialogue with the user, the human-computer dialogue terminal obtains the first message input by the user in this round and determines the message modality of the first message, wherein the message modality of the first message is text, document, image, video or audio. When the message modality of the first message is text, a text recognition model is selected from multiple modal message recognition models as the target modal message recognition model; when the message modality of the first message is a document, a document recognition model is selected from multiple modal message recognition models as the target modal message recognition model; when the message modality of the first message is an image, an image recognition model is selected from multiple modal message recognition models as the target modal message recognition model; when the message modality of the first message is a video, a video recognition model is selected from multiple modal message recognition models as the target modal message recognition model; when the message modality of the first message is an audio, a speech recognition model is selected from multiple modal message recognition models as the target modal message recognition model, so as to input the first message into the target modal message recognition model, and use the target modal message recognition model to recognize the content in the first message and obtain a recognition result.

[0068] The embodiment of the present application pre-constructs multiple modal message recognition models, and selects a target modal message recognition model corresponding to the message modality of the first message from the multiple modal message recognition models, thereby ensuring that the target modal message recognition model can accurately recognize the content in the first message, thereby supporting the recognition of multi-modal messages for human-computer dialogue and improving the efficiency of human-computer dialogue.

[0069] In an optional embodiment, the document recognition model includes a first language model GPT-4Turbo, and the image recognition model and the video recognition model include a second language model GPT-4o.

[0070] As an example, considering the different characteristics of document, image and video type messages, the first language model GPT-4Turbo is selected to build a document recognition model, and the second language model GPT-4o is selected to build an image recognition model and a video recognition model.

[0071] The first language model GPT-4Turbo and the second language model GPT-4o are language models released for the chatbot ChatGPT.

[0072] The first language model, GPT-4Turbo, not only has better text generation and comprehension capabilities, but can also process multiple documents in parallel on a large scale. It also optimizes the use of computing resources, significantly reduces computing resource consumption, and improves model operation efficiency. The second language model, GPT-4o, has multimodal processing capabilities, supports processing inputs such as text, images, and videos, can retain more features of multimodal inputs, accurately understand inputs, and quickly respond to outputs.

[0073] The embodiment of the present application selects the first language model GPT-4Turbo to build a document recognition model, and selects the second language model GPT-4o to build an image recognition model and a video recognition model. This can ensure that when the message modality of the first message is a complex document, image or video, the content of the first message can be quickly and accurately identified, thereby supporting the recognition of multimodal messages for human-computer dialogue and improving the efficiency of human-computer dialogue.

[0074] In an optional embodiment, obtaining the target message record based on the first message, the recognition result and the historical message record of this conversation includes: splicing the first message, the recognition result and the historical message record in a preset splicing order to obtain the target message record.

[0075] As an example, in order to ensure that the reply message generation model accurately locates the first message, the recognition result and the historical message record of this conversation for processing, a splicing order is pre-set for the first message, the recognition result and the historical message record of this conversation.

[0076] After obtaining the first message, the recognition result and the historical message record of this conversation, the human-computer dialogue terminal splices the first message, the recognition result and the historical message record of this conversation in a pre-set splicing order to obtain a target message record, and inputs the target message record into a reply message generation model to obtain a second message.

[0077] For example, assuming that based on the context order, the splicing order is pre-set to "historical message record of this conversation - first message - recognition result", after the human-computer dialogue terminal obtains the first message A, the recognition result a and the historical message record B of this conversation, it splices the first message A, the recognition result a and the historical message record B of this conversation according to the pre-set splicing order to obtain the target message record "B-A-a".

[0078] The embodiment of the present application obtains the target message record by splicing the first message, the recognition result and the historical message record of this conversation in a preset splicing order, so as to ensure that the reply message generation model accurately locates the first message, the recognition result and the historical message record of this conversation for processing, thereby supporting the recognition of multimodal messages for human-computer dialogue and improving the efficiency of human-computer dialogue.

[0079] In an optional embodiment, before obtaining the target message record based on the first message, the recognition result and the historical message record of this conversation, it also includes: obtaining the historical message record based on the current conversation identifier carried by the first message; wherein the current conversation identifier is generated when the current conversation is established and stored in association with the historical message record.

[0080] As an exemplary example, when establishing a current session with a user, the human-machine dialogue terminal generates a unique current session identifier, so that the first message input by the user in each round in the current session carries the current session identifier.

[0081] After obtaining the first message input by the user in this round, the human-computer dialogue terminal obtains the historical message record of this session stored in association with the current session identifier according to the current session identifier carried in the first message.

[0082] The embodiment of the present application obtains the historical message record of the current session stored in association with the current session identifier according to the current session identifier carried by the first message, thereby ensuring that the historical message record of the current session is obtained quickly and accurately, thereby supporting the recognition of multimodal messages for human-computer dialogue and improving the efficiency of human-computer dialogue.

[0083] In an optional embodiment, the reply message generation model includes a question extraction agent, an intent recognition agent and a message generation agent; the target message record is input into a pre-built reply message generation model to obtain a second message, including: extracting at least one user question from the target message record through a question extraction agent; identifying the user intent based on at least one user question through an intent recognition agent; determining an answer that meets the user intent through a message generation agent, and generating a second message based on the answer.

[0084] As an example, considering that the data composition of the target message record is relatively complex, in order to ensure accurate replies to users, a reply message generation model is pre-built that includes a question extraction agent, an intent recognition agent, and a message generation agent.

[0085] An agent is an agent that can perceive the environment and take actions to achieve specific goals. It can be software, hardware or a system with autonomy, adaptability and interaction. The agent perceives changes in the environment (such as through sensors or data input), makes judgments and decisions based on the knowledge and algorithms it has learned, and then performs actions to affect the environment or achieve predetermined goals.

[0086] The question extraction agent is an agent trained using a large amount of sample data, such as sample message records, and at least one user question corresponding to the sample message records, and is used to achieve the goal of extracting at least one user question from the target message record.

[0087] The intent recognition agent is an agent trained using a large amount of sample data, such as at least one user question corresponding to a sample message record, and the user intent corresponding to these user questions, and is used to achieve the goal of identifying the user intent based on at least one user question.

[0088] The message generation agent is an agent trained using a large amount of sample data, such as the user intent corresponding to these user questions, and the knowledge answers in a pre-built business knowledge base, to achieve the goal of generating reply messages based on answers that meet the user intent.

[0089] The training method of the above-mentioned intelligent agent can refer to the existing intelligent agent training method, which will not be elaborated here.

[0090] After obtaining the target message record, the human-computer dialogue terminal inputs the target message record into a pre-built reply message generation model, extracts at least one user question from the target message record through a question extraction agent, identifies the user intent based on at least one user question through an intent recognition agent, determines an answer that meets the user intent through a message generation agent, and generates a second message based on the answer.

[0091] In practical applications, the question extraction agent can locate and extract the first message, recognition results, and historical message records of this conversation from the target message records, extract features from the first message, recognition results, and historical message records of this conversation respectively, cluster the multiple features obtained, and understand the internal logic of each feature to sort out each user's questions.

[0092] The intent recognition agent can extract keywords from each user question, perform preprocessing such as deduplication on multiple keywords, match the preprocessed keywords with each candidate intent in the pre-built candidate intent library, and determine the candidate intent that matches the most keywords as the user intent.

[0093] The message generation agent can query the answer that meets the user's intention from the pre-built business knowledge base, and fill the answer into the message template corresponding to the user's intention to generate a second message.

[0094] The embodiment of the present application pre-constructs a reply message generation model including a question extraction agent, an intent recognition agent and a message generation agent. The question extraction agent extracts at least one user question from the target message record, the intent recognition agent recognizes the user intent based on at least one user question, the message generation agent determines an answer that meets the user intent, and generates a second message based on the answer. The embodiment of the present application can coordinate multiple agents to finely identify the user intent in the target message record and determine an answer that meets the user intent to generate a second message, thereby supporting the recognition of multimodal messages for human-computer dialogue and improving the efficiency of human-computer dialogue.

[0095] In an optional embodiment, the method further includes steps S105-S106:

[0096] S105, inserting the first message and the second message into the message queue;

[0097] S106: Consume the message queue, update the historical message record, and store the updated historical message record.

[0098] As an exemplary embodiment, after obtaining the second message, the human-machine dialogue terminal sequentially inserts the previously obtained first message and the second message into the message queue.

[0099] After inserting the first message and the second message into the message queue, the human-machine dialogue terminal consumes the message queue, updates the historical message record, and stores the updated historical message record, so as to continue the next round of dialogue and persist the updated historical message record.

[0100] Message Queue (MQ) refers to a container that stores messages during the transmission process. It is a message middleware used to implement asynchronous communication. When a thread of a human-computer dialogue terminal inserts the first message and the second message into the message queue, the thread will directly return to perform other operations. The message queue can be connected to another thread or other process of the human-computer dialogue terminal. Another thread or other process of the human-computer dialogue terminal can consume the message queue, obtain the first message and the second message to update the historical message record of this session, and store the updated historical message record, so as to achieve asynchronous update of the historical message record and storage of the updated historical message record.

[0101] The message queue follows the "first in, first out" principle. By obtaining the first message and the second message from the message queue to update the historical message record, the updated historical message record can be obtained, which can ensure that the first message and the second message are recorded in sequence and the historical message record is accurately updated.

[0102] In practical applications, the updated historical message records may be stored in the target cache.

[0103] Cache refers to a high-speed memory with a faster access speed than general random access memory (RAM). By storing the updated historical message records in the target cache, the historical message records of this conversation can be instantly obtained in each conversation after this round of conversation in this conversation.

[0104] In practical applications, Redis cache can be used as the target cache.

[0105] Redis (Remote Dictionary Server) cache is an open source log-type, key-value database written in ANSI C language, supporting the network, memory-based and persistent, and providing APIs (Application Programming Interface) in multiple languages.

[0106] By selecting Redis cache as the target cache and storing the updated historical message records in Redis cache, the historical message records of this session can be instantly obtained in each round of conversation after this round of conversation. The session identifier can also be used as the key and the historical message record as the corresponding value to associate the two for storage, which is conducive to improving the processing speed of each round of conversation.

[0107] The embodiment of the present application inserts the first message and the second message into the message queue, consumes the message queue, updates the historical message record, and stores the updated historical message record. It can asynchronously update the historical message record and store the updated historical message record, effectively reducing the processing pressure of human-computer dialogue, thereby supporting the recognition of multimodal messages for human-computer dialogue and improving the efficiency of human-computer dialogue.

[0108] Please see Figure 2 , Figure 2A schematic diagram of the structure of a human-computer dialogue processing device based on multimodal messages provided for the second embodiment of the present application. The second embodiment of the present application provides a human-computer dialogue processing device based on multimodal messages, including: a target model determination module 201, used to determine the target modal message recognition model corresponding to the message modality according to the message modality of the first message input by the user; a first message recognition module 202, used to input the first message into the target modal message recognition model to obtain a recognition result; a message record acquisition module 203, used to obtain a target message record according to the first message, the recognition result and the historical message record of this conversation; a second message return module 204, used to input the target message record into a pre-built reply message generation model to obtain a second message, and return the second message to the user.

[0109] In an optional embodiment, the message modality is text, document, image, video or audio; the target modality message recognition model corresponding to the message modality is determined according to the message modality of the first message input by the user, including: when the message modality is text, selecting a text recognition model from a plurality of pre-built modality message recognition models as the target modality message recognition model; when the message modality is a document, selecting a document recognition model from a plurality of modality message recognition models as the target modality message recognition model; when the message modality is an image, selecting an image recognition model from a plurality of modality message recognition models as the target modality message recognition model; when the message modality is a video, selecting a video recognition model from a plurality of modality message recognition models as the target modality message recognition model; when the message modality is an audio, selecting a speech recognition model from a plurality of modality message recognition models as the target modality message recognition model.

[0110] In an optional embodiment, the document recognition model includes a first language model GPT-4Turbo, and the image recognition model and the video recognition model include a second language model GPT-4o.

[0111] In an optional embodiment, obtaining the target message record based on the first message, the recognition result and the historical message record of this conversation includes: splicing the first message, the recognition result and the historical message record in a preset splicing order to obtain the target message record.

[0112] In an optional embodiment, before obtaining the target message record based on the first message, the recognition result and the historical message record of this conversation, it also includes: obtaining the historical message record based on the current conversation identifier carried by the first message; wherein the current conversation identifier is generated when the current conversation is established and stored in association with the historical message record.

[0113] In an optional embodiment, the reply message generation model includes a question extraction agent, an intent recognition agent and a message generation agent; the target message record is input into a pre-built reply message generation model to obtain a second message, including: extracting at least one user question from the target message record through a question extraction agent; identifying the user intent based on at least one user question through an intent recognition agent; determining an answer that meets the user intent through a message generation agent, and generating a second message based on the answer.

[0114] In an optional embodiment, the device further includes: a message record update module, which is used to: insert the first message and the second message into the message queue; consume the message queue, update the historical message record, and store the updated historical message record.

[0115] The implementation process of the functions and effects of each module in the above-mentioned device is specifically described in the implementation process of the corresponding steps in the above-mentioned method, which will not be repeated here.

[0116] Please see Figure 3 , Figure 3 The third embodiment of the present application provides an electronic device 30, comprising a processor 301, a memory 302, and a computer program stored in the memory 302 and configured to be executed by the processor 301; when the processor 301 executes the computer program, the method described in the first embodiment of the present application is implemented and the same beneficial effects can be achieved.

[0117] The processor 301 may read the computer program from the memory 302 through the bus 303 and execute the computer program to implement any of the embodiments of the method described in the first embodiment of the present application.

[0118] Processor 301 can process digital signals and can include various computing structures, such as complex instruction set computer structure, reduced instruction set computer structure, or a structure that implements a combination of multiple instruction sets. In some examples, processor 301 can be a microprocessor.

[0119] The memory 302 may be used to store instructions executed by the processor 301 or data related to the execution of instructions. These instructions and / or data may include codes for implementing some or all functions of one or more modules described in the embodiments of the present application. The processor 301 of the disclosed embodiment may be used to execute instructions in the memory 302 to implement the method described in the first embodiment of the present application. The memory 302 includes a dynamic random access memory, a static random access memory, a flash memory, an optical memory, or other memory known to those skilled in the art.

[0120] The fourth embodiment of the present application provides a computer-readable storage medium, which includes a stored computer program; wherein, when the computer program is running, the device where the computer-readable storage medium is located is controlled to execute the method described in the first embodiment of the present application, and can achieve the same beneficial effects as the method.

[0121] In summary, the embodiments of the present application provide a method, apparatus, device and medium for processing human-computer dialogue based on multimodal messages. The method for processing human-computer dialogue based on multimodal messages includes: determining a target modal message recognition model corresponding to the message modality according to the message modality of the first message input by the user; inputting the first message into the target modal message recognition model to obtain a recognition result; obtaining a target message record according to the first message, the recognition result and the historical message record of this conversation; inputting the target message record into a pre-built reply message generation model to obtain a second message, and returning the second message to the user. The embodiment of the present application determines a target modal message recognition model corresponding to the message modality according to the message modality of the first message input by the user, inputs the first message into the target modal message recognition model to obtain a recognition result, obtains a target message record according to the first message, the recognition result and the historical message record of this conversation, inputs the target message record into a pre-built reply message generation model to obtain a second message, and returns the second message to the user. The target modal message recognition model corresponding to the message modality of the first message can be used to specifically identify the content in the first message, and the reply message generation model can be used to comprehensively combine the first message, the recognition result and the historical message record of this conversation to accurately generate the second message, thereby supporting the recognition of multi-modal messages for human-computer dialogue and improving the efficiency of human-computer dialogue.

[0122] In several embodiments provided in the present application, it should be understood that the disclosed devices and methods can also be implemented in other ways. The device embodiments described above are merely schematic. For example, the flowcharts and block diagrams in the accompanying drawings show the possible architecture, functions and operations of the devices, methods and computer program products according to multiple embodiments of the present application. In this regard, each box in the flowchart or block diagram can represent a module, a program segment or a part of a code, and the module, a program segment or a part of a code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order from the order marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of boxes in the block diagram and / or flowchart can be implemented with a dedicated hardware-based system that performs a specified function or action, or can be implemented with a combination of dedicated hardware and computer instructions.

[0123] In addition, the functional modules in the various embodiments of the present application may be integrated together to form an independent part, or each module may exist separately, or two or more modules may be integrated to form an independent part.

[0124] If the functions are implemented in the form of software function modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application can be essentially or partly embodied in the form of a software product that contributes to the prior art. The computer software product is stored in a storage medium, including several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the methods described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.

[0125] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art who is familiar with the present technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.

Claims

1. A method for processing human-computer dialogue based on multimodal messages, characterized in that: include: Determining a target modality message recognition model corresponding to the message modality according to the message modality of the first message input by the user; Inputting the first message into the target modal message recognition model to obtain a recognition result; Obtaining a target message record according to the first message, the recognition result, and the historical message record of this conversation; The target message record is input into a pre-built reply message generation model to obtain a second message, and the second message is returned to the user.

2. The method according to claim 1, characterized in that The message modality is text, document, image, video or audio; The determining, according to the message modality of the first message input by the user, a target modality message recognition model corresponding to the message modality includes: In the case where the message modality is text, selecting a text recognition model from a plurality of pre-constructed modality message recognition models as the target modality message recognition model; In the case where the message modality is a document, selecting a document recognition model from the multiple modality message recognition models as the target modality message recognition model; When the message modality is an image, selecting an image recognition model from the multiple modality message recognition models as the target modality message recognition model; When the message modality is a video, selecting a video recognition model from the multiple modality message recognition models as the target modality message recognition model; When the message modality is audio, a speech recognition model is selected from the multiple modality message recognition models as the target modality message recognition model.

3. The method according to claim 2, characterized in that The document recognition model includes a first language model GPT-4Turbo, and the image recognition model and the video recognition model include a second language model GPT-4o.

4. The method according to claim 1, characterized in that: The obtaining of a target message record according to the first message, the recognition result and the historical message record of the current conversation includes: The first message, the recognition result and the historical message record are spliced ​​in a preset splicing order to obtain the target message record.

5. The method according to claim 1, characterized in that: Before obtaining the target message record according to the first message, the recognition result and the historical message record of the current conversation, the method further includes: The historical message record is obtained according to the current session identifier carried in the first message; wherein the current session identifier is generated when the current session is established and is stored in association with the historical message record.

6. The method according to claim 1, characterized in that The reply message generation model includes a question extraction agent, an intention recognition agent and a message generation agent; The step of inputting the target message record into a pre-built reply message generation model to obtain a second message includes: Extracting at least one user question from the target message record by the question extraction agent; identifying the user intent according to the at least one user question by the intent recognition agent; An intelligent agent is generated through the message to determine an answer that meets the user's intention, and the second message is generated according to the answer.

7. The method according to any one of claims 1 to 6, characterized in that: The method further comprises: Inserting the first message and the second message into a message queue; The message queue is consumed, the historical message record is updated, and the updated historical message record is stored.

8. A human-computer dialogue processing device based on multimodal messages, characterized in that: include: A target model determination module, used to determine a target modality message recognition model corresponding to the message modality according to the message modality of the first message input by the user; A first message recognition module, used for inputting the first message into the target modality message recognition model to obtain a recognition result; A message record acquisition module, configured to obtain a target message record according to the first message, the recognition result and the historical message record of this conversation; The second message returning module is used to input the target message record into a pre-built reply message generation model to obtain a second message, and return the second message to the user.

9. An electronic device, characterized in that: The method comprises a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor; when the processor executes the computer program, the human-computer dialogue processing method according to any one of claims 1 to 7 is implemented.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium includes a stored computer program; wherein, when the computer program is running, the device where the computer-readable storage medium is located is controlled to execute the human-computer dialogue processing method according to any one of claims 1 to 7.