Interaction method and device based on artificial intelligence, equipment and intelligent agent

By combining search enhancement technology with multimodal big models, using external knowledge bases to enhance the knowledge reserve of multimodal big models, the problem of insufficient reliability, timeliness and professionalism when generating content in multimodal big models is solved, and higher quality content generation is achieved.

CN120144837APending Publication Date: 2025-06-13BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510264901.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-06
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

Existing multimodal large models have problems of insufficient reliability, timeliness and professionalism when generating content, including opaque information sources, possible hallucinations, misleading information, poor knowledge timeliness and insufficient mastery of fine-grained or professional knowledge.

Method used

Combining search enhancement technology and multimodal big models, by retrieving external knowledge related to problems from an accurate, real-time updated and professional external knowledge base, we enhance and supplement the knowledge reserve of multimodal big models, thereby improving the reliability, timeliness and professionalism of generated content.

Benefits of technology

It effectively improves the reliability, timeliness and professionalism of multimodal large-scale model generation content, reduces the risk of information misleading, and improves the mastery of fine-grained and professional knowledge.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120144837A_ABST
    Figure CN120144837A_ABST
Patent Text Reader

Abstract

The invention provides an interaction method and device based on artificial intelligence, equipment, a medium, a program product and an intelligent agent, relates to the technical field of artificial intelligence, in particular to the technical fields of computer vision, deep learning, large models and the like, and can be applied to scenes such as AIGC content generation based on artificial intelligence. The method comprises the steps that a multi-modal problem is acquired, and the multi-modal problem comprises a text and an image; information matched with the text and the image is retrieved, and multi-source retrieval information is obtained; performing multi-task processing on the multi-source retrieval information based on a multi-modal problem by utilizing a multi-modal knowledge extraction large model to obtain retrieval enhancement information; and processing the retrieval enhancement information based on the multi-modal problem by using the multi-modal interaction large model to obtain reply content aiming at the multi-modal problem.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence technology, and particularly to technical fields such as computer vision, deep learning, large models, etc., and can be applied to scenarios such as AIGC (Artificial Intelligence Generated Content). More specifically, it relates to an interaction method, device, equipment, medium, program product and intelligent agent based on artificial intelligence. Background Art

[0002] With the rapid development of artificial intelligence technology, the processing ability of large models for multimodal information has gradually attracted attention, and multimodal interaction large models have become a new development direction. In this context, how to improve the generation quality of multimodal interaction large models is a direction worthy of exploration. Summary of the Invention

[0003] The present disclosure provides an interaction method, device, equipment, medium, program product and intelligent agent based on artificial intelligence.

[0004] According to one aspect of the present disclosure, there is provided an interaction method based on artificial intelligence, including obtaining a multimodal problem, where the multimodal problem includes text and image; respectively retrieving information matching the text and the image to obtain multi-source retrieval information; using a multimodal knowledge refinement large model to perform multitask processing on the multi-source retrieval information based on the multimodal problem to obtain retrieval enhanced information; and using a multimodal interaction large model to process the retrieval enhanced information based on the multimodal problem to obtain a reply content for the multimodal problem.

[0005] According to another aspect of the present disclosure, there is provided an interaction device based on artificial intelligence, including an obtaining module for obtaining a multimodal problem, where the multimodal problem includes text and image; a retrieval module for respectively retrieving information matching the text and the image to obtain multi-source retrieval information; a retrieval enhancement module for using a multimodal knowledge refinement large model to perform multitask processing on the multi-source retrieval information based on the multimodal problem to obtain retrieval enhanced information; and a reply module for using a multimodal interaction large model to process the retrieval enhanced information based on the multimodal problem to obtain a reply content for the multimodal problem.

[0006] According to another aspect of the present disclosure, there is provided an intelligent agent based on artificial intelligence, configured to execute the method as described in the present disclosure.

[0007] According to another aspect of the present disclosure, there is provided an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method as disclosed in the present disclosure.

[0008] According to another aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to execute the method as disclosed in the present disclosure.

[0009] According to another aspect of the present disclosure, there is provided a computer program product, including a computer program which, when executed by a processor, implements the method as disclosed in the present disclosure.

[0010] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:

[0012] Figure 1 Schematically shows an exemplary system architecture to which an artificial intelligence-based interaction method and apparatus can be applied according to an embodiment of the present disclosure;

[0013] Figure 2 Schematically shows a flowchart of an artificial intelligence-based interaction method according to an embodiment of the present disclosure;

[0014] Figure 3 Schematically shows a schematic diagram of an artificial intelligence-based interaction method according to an embodiment of the present disclosure;

[0015] Figure 4 Schematically shows a schematic diagram of an artificial intelligence-based interaction method according to an embodiment of the present disclosure;

[0016] Figure 5 Schematically shows a schematic diagram of obtaining multi-source retrieval information through multi-source multi-modal retrieval according to an embodiment of the present disclosure;

[0017] Figure 6 Schematically shows a block diagram of an artificial intelligence-based interaction apparatus according to an embodiment of the present disclosure;

[0018] Figure 7 Schematically shows a block diagram of an artificial intelligence-based agent according to an embodiment of the present disclosure; and

[0019] Figure 8 A block diagram of an electronic device suitable for implementing an artificial intelligence-based interaction method according to an embodiment of the present disclosure is schematically shown. Detailed implementation manners

[0020] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, descriptions of well-known functions and structures are omitted below for clarity and conciseness.

[0021] In the technical solution of the present disclosure, the processing of the collection, storage, use, processing, transmission, provision, disclosure, and application of the user's personal information involved all comply with the provisions of relevant laws and regulations, necessary confidentiality measures are taken, and public order and good customs are not violated.

[0022] In the technical solution of the present disclosure, the authorization or consent of the user is obtained before obtaining or collecting the user's personal information.

[0023] With the rapid development of artificial intelligence technology, the processing ability of large models for multimodal information has gradually attracted attention, and multimodal large models have become a new development direction. In this context, how to improve the generation quality of multimodal large models is a direction worthy of exploration.

[0024] According to an embodiment of the present disclosure, a multimodal large language model (MLLM) can be used to process various modalities of information (such as text, images, audio, video, etc.), and has powerful multimodal understanding and reasoning capabilities.

[0025] In implementing the inventive concept of the present disclosure, the inventors found that although multimodal large models can achieve complex multimodal information understanding, reasoning, and generation, the following problems still exist in the generated content of related multimodal large models: First, there is a lack of reliability. The information sources of the content generated by multimodal large models are not transparent enough, and at the same time, multimodal large models may have hallucinations, which may lead to the risk of information misleading in the generated content; Second, there is a lack of timeliness. Since the timeliness of knowledge in multimodal large models depends on the training data, and the knowledge in multimodal large models is solidified after training and cannot be updated in real time during reasoning, this may lead to poor timeliness of the generated content; Third, there is a lack of professionalism. Since multimodal large models are usually trained based on general training sets, the multimodal large models lack the mastery of fine-grained or professional knowledge.

[0026] According to an embodiment of the present disclosure, the retrieval enhancement technology can combine information retrieval technology with a large model. By retrieving relevant external knowledge from an external knowledge base, the large model can generate more accurate and richer content based on the retrieved external knowledge, effectively improving the generation quality of the large model.

[0027] Based on this, the inventors innovatively proposed to combine the retrieval enhancement technology with a multimodal large model to enhance and supplement the knowledge reserve of the multimodal large model by mining and retrieving external knowledge related to the problem from an accurate, real-time updated, and professional external knowledge base, thereby improving the reliability, timeliness, and professionalism of the content generated by the multimodal large model. However, further, the inventors also found that the relevant retrieval enhancement technology is mainly concentrated on single-modal models, and there is insufficient research on multimodal retrieval enhancement generation applicable to multimodal large models, and there are problems such as a single information source, a single retrieval method, poor knowledge flexibility, and noise in the retrieved external knowledge in the specific implementation.

[0028] In view of this, an embodiment of the present disclosure provides an artificial intelligence-based interaction method, device, device, medium, program product, and intelligent agent. The artificial intelligence-based interaction method includes: obtaining a multimodal problem, where the multimodal problem includes text and an image; respectively retrieving information matching the text and the image to obtain multi-source retrieval information; using a multimodal knowledge refinement large model to perform multitask processing on the multi-source retrieval information based on the multimodal problem to obtain retrieval enhancement information; and using a multimodal interaction large model to process the retrieval enhancement information based on the multimodal problem to obtain a reply content for the multimodal problem.

[0029] Figure 1 Schematically shows an exemplary system architecture to which the artificial intelligence-based interaction method and device according to an embodiment of the present disclosure can be applied.

[0030] It should be noted that Figure 1 The shown is only an example of the system architecture to which the embodiments of the present disclosure can be applied, to help those skilled in the art understand the technical content of the present disclosure, but it does not mean that the embodiments of the present disclosure cannot be used in other devices, systems, environments, or scenarios. For example, in another embodiment, the exemplary system architecture to which the artificial intelligence-based interaction method and device can be applied may include a terminal device, but the terminal device can implement the artificial intelligence-based interaction method and device provided by the embodiments of the present disclosure without interacting with the server.

[0031] Such as Figure 1As shown, the system architecture 100 according to this embodiment may include terminal devices 101, 102, 103, a network 104, and a server 105. The network 104 is used to provide a medium for communication links between the terminal devices 101, 102, 103 and the server 105. The network 104 may include various connection types, such as wired and / or wireless communication links, etc.

[0032] Users can use the terminal devices 101, 102, 103 to interact with the server 105 through the network 104 to receive or send messages, etc. Various communication client applications may be installed on the terminal devices 101, 102, 103, such as knowledge reading applications, web browser applications, search applications, instant messaging tools, email clients, and / or social platform software, etc. (for example only).

[0033] The terminal devices 101, 102, 103 may be various electronic devices with a display screen and supporting web browsing, including but not limited to smart phones, tablet computers, laptop portable computers, and desktop computers, etc.

[0034] The server 105 may be a server that provides various services, such as a background management server that provides support for the content browsed by users using the terminal devices 101, 102, 103 (for example only). The background management server may analyze and process data such as received user requests, and feedback the processing results (such as web pages, information, or data, etc. obtained or generated according to user requests) to the terminal devices.

[0035] It should be noted that the artificial intelligence-based interaction method provided by the embodiments of the present disclosure can generally be executed by the terminal devices 101, 102, or 103. Correspondingly, the artificial intelligence-based interaction device provided by the embodiments of the present disclosure can also be set in the terminal devices 101, 102, or 103.

[0036] Alternatively, the artificial intelligence-based interaction method provided by the embodiments of the present disclosure can generally also be executed by the server 105. Correspondingly, the artificial intelligence-based interaction device provided by the embodiments of the present disclosure can generally be set in the server 105. The artificial intelligence-based interaction method provided by the embodiments of the present disclosure can also be executed by a server or a server cluster different from the server 105 and capable of communicating with the terminal devices 101, 102, 103 and / or the server 105. Correspondingly, the artificial intelligence-based interaction device provided by the embodiments of the present disclosure can also be set in a server or a server cluster different from the server 105 and capable of communicating with the terminal devices 101, 102, 103 and / or the server 105.

[0037] It should be understood, Figure 1The numbers of the terminal devices, networks, and servers in [it] are merely illustrative. According to the implementation requirements, there can be any number of terminal devices, networks, and servers.

[0038] It should be noted that the serial numbers of the respective operations in the following methods are only used as representations of the operations for description purposes and should not be regarded as indicating the execution order of the respective operations. Unless explicitly stated, this method does not need to be executed exactly in the order shown.

[0039] Figure 2 A flowchart of an artificial intelligence-based interaction method according to an embodiment of the present disclosure is schematically shown.

[0040] As Figure 2 shown, the method 200 includes operations S210 to S240.

[0041] In operation S210, a multi-modal problem is obtained, where the multi-modal problem includes text and an image.

[0042] In operation S220, information respectively matching the text and the image is retrieved to obtain multi-source retrieval information.

[0043] In operation S230, the multi-modal knowledge refinement large model performs multi-task processing on the multi-source retrieval information based on the multi-modal problem to obtain retrieval-enhanced information.

[0044] In operation S240, the multi-modal interaction large model processes the retrieval-enhanced information based on the multi-modal problem to obtain a reply content for the multi-modal problem.

[0045] According to an embodiment of the present disclosure, the multi-modal problem may include information of multiple modalities. For example, the multi-modal problem may include text (e.g., denoted as Q) and an image (e.g., denoted as I). Exemplarily, the modality types may include text, image, audio, video, etc., which are not limited herein.

[0046] In one embodiment, the multi-modal problem may be information input by a user including an image and a natural language description. For example, the multi-modal problem may be "Please recommend several mobile phones with high cost performance + an image of a mobile phone". Another example, the multi-modal problem may be "Please analyze the force situation of a small ball + an image of a small ball".

[0047] According to embodiments of the present disclosure, information respectively matching text Q and image I can be retrieved to obtain multi-source retrieval information (which can be denoted as R for example). Exemplarily, based on at least one text-based retrieval tool, first retrieval information matching text Q can be retrieved. Based on at least one image-based retrieval tool, second retrieval information matching image I can be retrieved. The multi-source retrieval information R can include the first retrieval information and the second retrieval information. In one embodiment, the first retrieval information, the second retrieval information, and the multi-source retrieval information R can all be in text modality.

[0048] According to embodiments of the present disclosure, the text-based retrieval tool can perform text-based retrieval according to text Q in the multi-modal question to obtain the first retrieval information. The image-based retrieval tool can perform image-based retrieval according to image I in the multi-modal question to obtain the second retrieval information.

[0049] Merely as an example, the text-based retrieval tool can include, for example, but not limited to, search engines, preset knowledge bases, encyclopedia websites, text-to-image search tools, etc., and the image-based retrieval tool can include, for example, but not limited to, image recognition tools, preset knowledge bases, image-to-image search tools, image-to-text search tools, etc. Those skilled in the art can adopt at least one retrieval tool of any type, any structure, and any retrieval method according to actual needs or application scenarios, etc., to perform text-based retrieval and image-based retrieval respectively to obtain the multi-source retrieval information R, which is not limited herein.

[0050] It can be understood that the multi-source retrieval information R contains retrieval information from multiple information sources and matching the multi-modal information respectively. Compared with existing retrieval enhancement technologies, by performing multi-source multi-modal retrieval for text Q and image I respectively from multiple information sources, the multi-source retrieval information R can provide more comprehensive, richer, and more accurate external knowledge related to the multi-modal question.

[0051] According to embodiments of the present disclosure, a multi-modal knowledge refinement large model can be used to perform multi-task processing on the multi-source retrieval information based on the multi-modal question to obtain retrieval enhancement information (which can be denoted as K for example). Among them, the retrieval enhancement information K can be used to assist the multi-modal interaction large model in performing a generation task for the multi-modal question to improve the generation quality of the multi-modal interaction large model.

[0052] According to embodiments of the present disclosure, the multi-modal knowledge refinement large model can be used to refine the retrieved multi-source retrieval information R, discard the irrelevant and redundant knowledge therein, so as to enhance the authority, timeliness, and professionalism of the output result of the multi-modal interaction large model.

[0053] Merely as an example, the above multi-task processing can include the following aspects:

[0054] Filtering, for example, can remove the parts that are irrelevant and redundant to the multimodal problem from the multi-source retrieval information R, and only retain the more relevant and effective parts among them;

[0055] Verification, for example, can verify the multi-source retrieval information R based on the multimodal problem to ensure the accuracy and consistency of the multi-source retrieval information R and avoid introducing misleading information;

[0056] Refinement, for example, can simplify, fuse, reorganize, etc. the multi-source retrieval information R, and refine the key information therein as the retrieval enhancement information K to enhance the logic and coherence of the retrieval enhancement information K, so as to improve the generation efficiency of the subsequent multimodal interaction large model.

[0057] According to an embodiment of the present disclosure, the first prompt and the multi-source retrieval information R can be input into the multimodal knowledge refinement large model to obtain the retrieval enhancement information K. Among them, the first prompt can be used to guide the multimodal knowledge refinement large model to perform multi-task processing on the multi-source retrieval information R based on the text Q and the image I in the multimodal problem to obtain the retrieval enhancement information K. Among them, the multi-source retrieval information R can be in text modality, the text Q and the image I are in text modality and image modality respectively. Through the first prompt, the multimodal knowledge refinement large model can be guided to use its powerful multimodal processing ability to perform multi-task processing to obtain the retrieval enhancement information K with higher refinement, logic and relevance. Exemplarily, the retrieval enhancement information K can be in text modality.

[0058] Only as an example, the first prompt can be "Please summarize the information in {R} related to {Q} and {I}, and answer 'irrelevant' if neither is relevant".

[0059] In one embodiment, the retrieval enhancement information K can be in text modality. In the structural design of the retrieval enhancement information K, it can include but is not limited to cause + process + result, conclusion + reason, question + answer, phenomenon + analysis + suggestion, general + sub + general, key point listing, etc., which are not specifically limited herein.

[0060] According to the embodiment of the present disclosure, by using the multimodal knowledge refinement model to perform multi-task processing on the multi-source retrieval information R based on the multimodal problem, the refinement and logic of the retrieval enhancement information K, as well as the relevance between the retrieval enhancement information K and the multimodal problem, can be effectively improved, thereby improving the generation efficiency of the multimodal interaction large model and the accuracy of the generated content.

[0061] According to an embodiment of the present disclosure, a multi-modal interaction large model can be utilized to process retrieved enhanced information based on a multi-modal question to obtain a response content for the multi-modal question. Among them, the retrieved enhanced information K can provide the multi-modal interaction large model with external knowledge having high reliability, strong timeliness, and strong professionalism, so as to assist the multi-modal interaction large model in performing a generation task for the multi-modal question. The retrieved enhanced information K can effectively improve the reliability, timeliness, and professionalism of the response content, thereby improving the generation quality of the multi-modal interaction large model.

[0062] According to an embodiment of the present disclosure, a second prompt and the retrieved enhanced information K can be input into the multi-modal interaction large model to obtain a response content (which can be denoted as A, for example). Among them, the second prompt can be used to guide the multi-modal interaction large model to perform a generation task based on the text Q and the image I in the multi-modal question, as well as the external knowledge provided by the retrieved enhanced information K, so as to obtain a response content A for the multi-modal question. Among them, the retrieved enhanced information K can be in text modality, the text Q and the image I are in text modality and image modality respectively. Through the second prompt, the multi-modal interaction large model can be guided to utilize its powerful multi-modal processing ability to perform a generation task to obtain a response content A with higher reliability, timeliness, and professionalism.

[0063] According to an embodiment of the present disclosure, the second prompt can be determined based on the multi-modal question. As an example, by performing intent recognition on the multi-modal question, the second prompt can be determined according to the intent of the multi-modal question. For example, the second prompt can be "Please analyze the force on the small ball according to {External Knowledge 1} and {the image of the small ball}, and detailed analysis results are required to be output". For example, the second prompt can be "Please recommend several mobile phones with high cost performance according to {External Knowledge 2} and {the image of the mobile phone}, and specific recommendation results and corresponding links are required to be output".

[0064] According to an embodiment of the present disclosure, through multi-source multi-modal retrieval for the text Q and the image I respectively from multiple information sources, the multi-source retrieval information R can provide more comprehensive, richer, and more accurate external knowledge related to the multi-modal question. By utilizing a multi-modal knowledge refinement model to perform multi-task processing on the multi-source retrieval information R based on the multi-modal question, the refinement and logic of the retrieved enhanced information K, as well as the relevance between the retrieved enhanced information K and the multi-modal question, can be effectively improved, thereby improving the generation efficiency of the multi-modal interaction large model and the accuracy of the generated content. The retrieved enhanced information K can provide the multi-modal interaction large model with external knowledge having high reliability, strong timeliness, and strong professionalism. By utilizing the multi-modal interaction large model to process the retrieved enhanced information K based on the multi-modal question, the reliability, timeliness, and professionalism of the response content A can be effectively improved, thereby improving the generation quality of the multi-modal interaction large model.

[0065] It should be noted that the AI-based interaction method provided in the embodiments of the present disclosure can be applied to any generation scenario based on a multi-modal interaction large model to help improve the reliability, timeliness, and professionalism of the generated content of the multi-modal interaction large model, thereby improving the generation quality of the multi-modal interaction large model. Exemplarily, the AI-based method provided in the embodiments of the present disclosure can be applied to scenarios such as intelligent visual question answering, intelligent medical and health, cross-modal content search, and cross-modal recommendation systems based on the text-image multi-modal large model, which is not specifically limited herein.

[0066] Figure 3 A schematic diagram of an AI-based interaction method according to an embodiment of the present disclosure is schematically shown.

[0067] As Figure 3 shown, multi-source retrieval information 310, predetermined prompt information 320 (the first prompt word in the above text), and multi-modal question 330 can be input into the multi-modal knowledge refinement large model M310 to obtain retrieval enhanced information 340.

[0068] As Figure 3 shown, the retrieval enhanced information 340 and the multi-modal question 330 can be input into the multi-modal interaction large model M320 to obtain a response content 350.

[0069] According to an embodiment of the present disclosure, using the multi-modal knowledge refinement large model to perform multi-task processing on multi-source retrieval information based on a multi-modal question to obtain retrieval enhanced information includes: using the multi-modal knowledge refinement large model to denoise the multi-source retrieval information based on the multi-modal question to obtain denoised multi-source retrieval information; and using the multi-modal knowledge refinement large model to refine the denoised multi-source retrieval information based on the multi-modal question to obtain retrieval enhanced information.

[0070] According to an embodiment of the present disclosure, using the multi-modal knowledge refinement large model to perform multi-task processing on multi-source retrieval information R based on a multi-modal question may include a denoising sub-task and a refinement sub-task. For example, the denoising sub-task may include using the multi-modal knowledge refinement large model to denoise the multi-source retrieval information R based on the multi-modal question to obtain denoised multi-source retrieval information R'. For example, the refinement sub-task may include using the multi-modal knowledge refinement large model to refine the denoised multi-source retrieval information R' based on the multi-modal question to obtain retrieval enhanced information K.

[0071] In one embodiment, multi-source retrieval information R and a multimodal question can be input into a multimodal knowledge refinement large model based on a first prompt. The multimodal knowledge refinement large model can perform a noise reduction subtask to reduce the noise of the multi-source retrieval information R based on the multimodal question, and obtain the denoised multi-source retrieval information R'. The multimodal knowledge refinement large model can perform a refinement subtask to refine the denoised multi-source retrieval information R' based on the multimodal question, and obtain the retrieval enhanced information K.

[0072] According to an embodiment of the present disclosure, since the multi-source retrieval information R is obtained by performing multimodal retrieval based on multiple information sources, it is easy to cause the multi-source retrieval information R retrieved to possibly contain noise information, for example, it may contain information that is irrelevant or redundant to the multimodal question. To ensure the matching degree and relevance between the retrieval enhanced information K and the multimodal question, the multimodal knowledge refinement large model can be used to reduce the noise of the multi-source retrieval information R based on the multimodal question, so as to discard the irrelevant and redundant noisy multi-source retrieval information r and obtain the denoised multi-source retrieval information R'. Among them, the denoised multi-source retrieval information R' can be understood as the retrieval information in the multi-source retrieval information R with a relatively high matching degree to the multimodal question.

[0073] According to an embodiment of the present disclosure, there may be multiple denoised multi-source retrieval information R'. To ensure the refinement and logic of the retrieval enhanced information K, the multimodal knowledge refinement large model can be used to refine the denoised multi-source retrieval information R' based on the multimodal question, so as to further condense, streamline, and refine the information in R' into the retrieval enhanced information K, so that the multimodal interaction large model can efficiently obtain effective information based on the retrieval enhanced information K, thereby helping to improve the generation efficiency and generation quality of the multimodal interaction large model.

[0074] According to an embodiment of the present disclosure, using the multimodal knowledge refinement large model to reduce the noise of the multi-source retrieval information based on the multimodal question to obtain the denoised multi-source retrieval information includes: determining the relevance between the multimodal question and the multi-source retrieval information to obtain a relevance recognition result; and in the case where the relevance recognition result indicates that the multi-source retrieval information is relevant to the multimodal question, obtaining the denoised multi-source retrieval information based on the multi-source retrieval information.

[0075] According to an embodiment of the present disclosure, the relevance between the multimodal question and the multi-source retrieval information R can be determined to obtain a relevance recognition result. In one embodiment, the relevance between the question Q and the image I in the multimodal question and the multi-source retrieval information R can be determined respectively to obtain a relevance recognition result. Exemplarily, the text similarity between the question Q and the multi-source retrieval information R can be determined to obtain a first relevance result. The semantic similarity between the image I and the multi-source retrieval information R can be determined to obtain a second relevance result. The relevance recognition result can be determined according to the first relevance result and the second relevance result.

[0076] For example, the multi-source retrieval information R can be encoded to obtain a first vector. The question Q can be encoded to obtain a second vector. The image I can be semantically recognized to obtain a third vector. The similarity between the first vector and the second vector can be calculated to obtain a first correlation result. The similarity between the third vector and the second vector can be calculated to obtain a second correlation result. The weighted result of the first correlation result and the second correlation result can be determined as the correlation recognition result. It should be noted that the calculation methods of similarity can include but are not limited to cosine similarity, Euclidean distance, etc., and are not limited here.

[0077] According to an embodiment of the present disclosure, the correlation recognition result can characterize the correlation relationship between the multi-source retrieval information and the multi-modal question. Exemplarily, the correlation relationship between the multi-source retrieval information and the multi-modal question can be determined based on the similarity recognition result and a preset similarity threshold. For example, when the similarity recognition result is greater than or equal to the preset similarity threshold, the correlation recognition result indicates that the multi-source retrieval information is relevant to the multi-modal question. For example, when the similarity recognition result is less than the preset similarity threshold, the correlation recognition result indicates that the multi-source retrieval information is not relevant to the multi-modal question.

[0078] According to an embodiment of the present disclosure, when the correlation recognition result indicates that the multi-source retrieval information R is relevant to the multi-modal question, the denoised multi-source retrieval information can be obtained based on the multi-source retrieval information.

[0079] According to an embodiment of the present disclosure, by determining the correlation between the multi-modal question and the multi-source retrieval information to obtain the correlation recognition result, when the correlation recognition result indicates that the multi-source retrieval information is relevant to the multi-modal question, the denoised multi-source retrieval information can be obtained based on the multi-source retrieval information. Thus, it is possible to remove the parts that are not relevant and redundant to the multi-modal question from the multi-source retrieval information R, and only retain the effective parts with higher correlation, thereby ensuring the correlation between the retrieved enhanced information K and the multi-modal question.

[0080] According to an embodiment of the present disclosure, the multi-modal knowledge refinement large model refines the denoised multi-source retrieval information based on the multi-modal question to obtain the retrieved enhanced information, including: sequentially performing content refinement and logical arrangement on the denoised multi-source retrieval information to obtain the retrieved enhanced information.

[0081] According to an embodiment of the present disclosure, the multi-modal knowledge refinement large model can sequentially perform content refinement and logical arrangement on the denoised multi-source retrieval information R´ to obtain the retrieved enhanced information K.

[0082] Exemplarily, the multi-modal knowledge extraction large model can extract keywords from the denoised multi-source retrieval information R´, and generate a summary based on the keywords. The purpose of content extraction is to summarize relatively fragmented and complex information into a short and concise summary. The multi-modal knowledge extraction large model can organize the generated summary in a certain logical order and organize it into a retrieval-enhanced information K with higher organization and logic. The purpose of logical organization is to make the retrieval-enhanced information K clearer and easier to understand. Thus, the retrieval-enhanced information K obtained after content extraction and logical organization can have higher comprehensiveness, accuracy, logic, and organization.

[0083] According to an embodiment of the present disclosure, by using the powerful semantic understanding ability of the multi-modal knowledge extraction large model to perform content extraction and logical organization on the denoised multi-source retrieval information R´, the refinement and logic of the retrieval-enhanced information K can be effectively improved, so that the multi-modal interaction large model can efficiently obtain effective information based on the retrieval-enhanced information K, thereby helping to improve the generation efficiency and generation quality of the multi-modal interaction large model.

[0084] According to an embodiment of the present disclosure, the interaction method may further include: in the case where the relevance recognition result indicates that all multi-source retrieval information is not relevant to the multi-modal problem, determining the multi-source retrieval information as noise multi-source retrieval information. In the case where the multi-source retrieval information is determined to be noise multi-source retrieval information, using the multi-modal interaction large model to process the multi-modal problem to obtain a response content.

[0085] Specifically, the multi-modal knowledge extraction large model can be used to perform a relevance recognition on the multi-source retrieval information to determine a relevance recognition result indicating whether there is a relevance between the multi-modal problem and the multi-source retrieval information. Exemplarily, in the case where there are multiple pieces of multi-source retrieval information, the non-relevance between the multi-source retrieval information and the multi-modal problem includes: all multiple pieces of multi-source retrieval information are not relevant to the multi-modal problem.

[0086] According to an embodiment of the present disclosure, in the case where the relevance recognition result indicates that all multi-source retrieval information R is not relevant to the multi-modal problem, the multi-source retrieval information R can be determined as noise multi-source retrieval information r.

[0087] According to an embodiment of the present disclosure, in the case where the multi-source retrieval information R is determined to be noise multi-source detection information r, the multi-modal interaction large model can be used to process the multi-modal problem to obtain a response content A.

[0088] Exemplarily, in the case where the relevance recognition result indicates that all multi-source retrieval information R is not relevant to the multi-modal problem, the multi-modal problem can be directly input into the multi-modal interaction large model, and the multi-modal interaction large model can be used to process the multi-modal problem to obtain a response content A.

[0089] According to an embodiment of the present disclosure, when it is determined that the multi-source retrieval information R is noise multi-source detection information r, if the multi-modal interaction large model continues to rely on the noise multi-source detection information r for generation, it may lead to inaccurate or irrelevant generated content. In this regard, a non-retrieval-enhanced generation method can be adopted. By directly using the multi-modal interaction large model to process the multi-modal problem to obtain the reply content A, it is possible to avoid the misguidance or generation failure that may be caused by noise information, ensure that the multi-modal interaction large model can answer multi-modal problems, and thus facilitate ensuring the reliability and stability of the multi-modal interaction large model. In addition, the non-retrieval-enhanced generation method enables the multi-modal interaction large model to more flexibly handle various types of multi-modal problems, especially for problems lacking clear external knowledge support, thereby further expanding the application scope of the multi-modal interaction large model.

[0090] Figure 4 Schematically shows a schematic diagram of an artificial intelligence-based interaction method according to an embodiment of the present disclosure.

[0091] As Figure 4 shown, based on the correlation recognition result between the multi-source retrieval information and the multi-modal problem, the multi-modal interaction large model can be used to generate the reply content for the multi-modal problem. In operation S401, it can be determined whether the multi-source retrieval information is relevant to the multi-modal problem.

[0092] As Figure 4 shown, when the multi-source retrieval information is relevant to the multi-modal problem, operation S402 can be executed. In operation S402, based on the multi-source retrieval information, the denoised multi-source retrieval information can be obtained. In operation S403, the multi-modal knowledge refinement large model can be used to refine the denoised multi-source retrieval information based on the multi-modal problem to obtain the retrieval-enhanced information. In operation S404, the multi-modal interaction large model can be used to process the retrieval-enhanced information based on the multi-modal problem to obtain the reply content for the multi-modal problem.

[0093] As Figure 4 shown, when the multi-source retrieval information is not relevant to the multi-modal problem, operation S405 can be executed. In operation S405, the multi-modal interaction large model can be used to process the multi-modal problem to obtain the reply content.

[0094] According to an embodiment of the present disclosure, using the multi-modal interaction large model to process the retrieval-enhanced information based on the multi-modal problem to obtain the reply content for the multi-modal problem includes: using the multi-modal interaction large model to process the retrieval-enhanced information based on the multi-modal problem to obtain the initial reply content; and optimizing the initial reply content based on the intention of the multi-modal problem to obtain the reply content.

[0095] According to an embodiment of the present disclosure, the retrieval enhancement information K and the multimodal question can be input into the multimodal interaction large model. The multimodal interaction large model can process the retrieval enhancement information K based on the multimodal question to obtain the initial reply content A'. It can be understood that different multimodal questions may correspond to different intents, and the initial reply content can be optimized based on the intent of the multimodal question to obtain the reply content.

[0096] According to an embodiment of the present disclosure, the output form of the reply content can be determined based on the intent of the multimodal question. The initial reply content A' can be optimized according to the foregoing output form to obtain the reply content A. Exemplarily, the intent of the multimodal question can include a question-and-answer intent, a recommendation intent, an execution intent, etc. By identifying the intent of the multimodal question, the multimodal interaction large model can determine an answer output form that more conforms to the question intent, so that the initial reply content A' can be optimized into the reply content A according to the output form to better meet the user's needs and improve the user experience.

[0097] In an alternative embodiment, the retrieval enhancement information K and the multimodal question can be input into the multimodal interaction large model based on the second prompt. The multimodal interaction large model can process the retrieval enhancement information K based on the multimodal question to obtain the initial reply content A'. As an example, the second prompt can also be used to determine the output form of the reply content. For example, the intent of the multimodal question can be identified, and the second prompt can be determined according to the intent of the multimodal question, where the second prompt indicates the output form of the reply content. The initial reply content A' can be optimized based on the second prompt to obtain the reply content A.

[0098] For example, when the intent of the multimodal question is a question-and-answer intent, the output form of the reply content can include, for example, but not limited to, text, graphics and text, tables, etc. For example, when the intent of the multimodal question is a recommendation intent, the output form of the reply content can include, for example, but not limited to, recommendation result cards, recommendation result lists, etc. It should be noted that those skilled in the art can determine the output forms corresponding to different question intents according to actual needs or application scenarios, etc., and no specific limitation is made here.

[0099] According to the embodiment of the present disclosure, the output form of the reply content A can be determined based on the intent of the multimodal question, and the initial reply content A' can be optimized based on the output form to obtain the reply content A, thereby improving the diversity and richness of the reply content A, and effectively improving the matching degree between the reply content A and the multimodal question to better meet the user's needs and improve the user experience.

[0100] According to an embodiment of the present disclosure, based on the intention of the multimodal question, the initial response content is optimized to obtain the response content, including: when the intention includes a recommendation intention, determining target content that conforms to a predetermined format from the initial response content; and updating the target content based on multi-source retrieval information corresponding to the target content to obtain the response content.

[0101] According to an embodiment of the present disclosure, when the intention of the multimodal question is a recommendation intention, target content that conforms to a predetermined format can be determined from the initial response content A´. Among them, the target content that conforms to the predetermined format can, for example, include a recommendation result card. Exemplarily, the recommendation result card can include a theme name, an abstract, and a picture, and key information such as the theme name, the abstract, and the picture can be combined based on the predetermined format to obtain the recommendation result card. Among them, the predetermined format can, for example, include a layout method (such as a flow layout, a list layout, a card layout, etc.), a font selection, an icon style, a color scheme, an interaction method, etc.

[0102] According to an embodiment of the present disclosure, the target content can be updated based on the multi-source retrieval information R corresponding to the target content to obtain the response content A, thereby improving the accuracy, timeliness, and interactivity of the response content A.

[0103] Only as an example, the target content that conforms to the predetermined format can, for example, include a recommendation card for mobile phone A. The recommendation card for mobile phone A can be obtained by combining the name of mobile phone A, the summary information of mobile phone A, and the picture of mobile phone A based on the predetermined format. The multi-source retrieval information R corresponding to the recommendation card for mobile phone A can, for example, include the promotion page of mobile phone A, which can, for example, include a promotion page link, a purchase link for mobile phone A, etc. The recommendation card for mobile phone A can be updated based on the promotion page of mobile phone A. For example, an interactive promotion page link and / or a purchase link for mobile phone A can be added to the recommendation card for mobile phone A to obtain the response content A. Users can quickly view the promotion page of mobile phone A or purchase mobile phone A by interacting with the response content A, thereby providing users with a one-stop information acquisition and purchase experience, and further improving the user's interaction experience.

[0104] According to an embodiment of the present disclosure, by based on the multi-source retrieval information R corresponding to the target content, the target content can be updated to obtain the response content A, thereby improving the accuracy, timeliness, and interactivity of the response content A. The response content A can provide users with a one-stop information acquisition and purchase experience, and further improve the user's interaction experience.

[0105] According to an embodiment of the present disclosure, based on the intention of the multimodal question, the initial response content is optimized to obtain the response content, including: when the intention includes a question-and-answer intention, performing format conversion on the initial response content to obtain the response content.

[0106] According to an embodiment of the present disclosure, when the intention of the multimodal question is a question-and-answer intention, the initial reply content A´ can be format-converted to obtain the reply content A. Exemplarily, the format-conversion methods may include but are not limited to structuring (such as in the form of a list or a table), highlighting (such as optimizing the layout, increasing the font size, etc.), visualization (such as combining text and graphics), etc., which are not limited herein.

[0107] According to an embodiment of the present disclosure, by format-converting the initial reply content A´ to obtain the reply content A, the intuitiveness and readability of the reply content A can be effectively improved, so that the user can quickly browse and understand it.

[0108] According to an embodiment of the present disclosure, retrieving information matching the text includes: obtaining a candidate content set matching the text; and determining, based on the semantic similarity between the text and the candidate content in the candidate content set, the information matching the text from the candidate content set.

[0109] According to an embodiment of the present disclosure, based on at least one text-based retrieval tool, text-based retrieval can be performed on the text Q in the multimodal question to obtain the first retrieval information. Exemplarily, the text-based retrieval tool may include but is not limited to a search engine, a preset knowledge base, an encyclopedia website, a text-to-image search tool, etc.

[0110] According to an embodiment of the present disclosure, a knowledge set can be constructed based on the above at least one text-based retrieval tool, and a candidate content set matching the text Q can be obtained based on the knowledge set. The candidate content set may include multiple candidate contents.

[0111] According to an embodiment of the present disclosure, based on the semantic similarity between the text Q and the multiple candidate contents, the information matching the text Q can be determined from the multiple candidate contents as the first retrieval information. Exemplarily, the semantic similarity between the text Q and each of the multiple candidate contents can be calculated, the multiple candidate contents can be sorted in descending order of semantic similarity, and the top N candidate contents in the sorting can be determined as the first retrieval information.

[0112] For example, the text Q can be input into a search engine tool to retrieve several candidate contents, and the summary of the first three candidate contents can be taken as the first retrieval information obtained from the search engine tool, which can be denoted as R, for example. 1 .

[0113] According to an embodiment of the present disclosure, obtaining a candidate content set matching the text includes: extracting the main content in the text; and obtaining a candidate content set matching the text based on the main content and the intention of the multimodal question.

[0114] According to an embodiment of the present disclosure, the main content in text Q can be extracted based on natural language processing technology, and the main content can be used to characterize the core content such as semantic information and structural information of text Q. Exemplarily, the main content can include triples, namely, Subject, Predicate, and Object. The triple structure can be used to describe the relationship or attribute between entities, and is usually expressed as (Subject, Predicate, Object).

[0115] According to an embodiment of the present disclosure, based on the main content and the intention of the multimodal question, a candidate content set matching text Q can be screened out from the knowledge set. Exemplarily, the knowledge set can include candidate content sets in multiple different fields. For example, if the main content is (mobile phone, has, high cost performance) and the intention of the multimodal question is a recommendation intention, then the candidate content set corresponding to the mobile phone field can be screened out.

[0116] According to an embodiment of the present disclosure, by extracting the main content in the text, the core content such as semantic information and structural information of text Q can be effectively captured. By screening out the candidate content set based on the main content and the intention of the multimodal question, it can be ensured that the candidate content set has a high degree of matching and relevance with the multimodal question, so as to improve the accuracy and precision of text-based retrieval, thereby improving the accuracy of the first retrieval information and its relevance to the multimodal question. At the same time, it is possible to avoid retrieving candidate content sets with low relevance to the multimodal question, so as to reduce the consumption of computing resources and improve the efficiency of retrieval.

[0117] According to an embodiment of the present disclosure, information matching the image can be retrieved through at least one of the image features of the image, the image, and the objects in the image.

[0118] According to an embodiment of the present disclosure, based on at least one image-based retrieval tool, image-based retrieval can be performed on image I in the multimodal question to obtain second retrieval information. Exemplarily, the image-based retrieval tool can include, but is not limited to, image recognition tools, preset knowledge bases, image search by image tools, image search by text tools, etc.

[0119] As an example, information matching the image can be retrieved through the image features of the image to obtain second retrieval information. As another example, information matching the image can be retrieved through the image to obtain second retrieval information. As another example, information matching the image can be retrieved through the objects in the image to obtain second retrieval information. As another example, information matching the image can be retrieved respectively through the image features of the image, the image, and the objects in the image to obtain second retrieval information.

[0120] It can be understood that the image features, the image, and the objects in the image of an image involve different dimensions of the image, which can cover a wider range of visual and semantic information. In addition, various different image-based retrieval tools can be adopted for different dimensions of the image to retrieve information matching the image, thereby improving the diversity, comprehensiveness, and richness of the second retrieved information.

[0121] According to an embodiment of the present disclosure, image-based retrieval can be performed for at least one dimension of the image, covering a wider range of visual and semantic information, so as to improve the diversity, comprehensiveness, and richness of the second retrieved information.

[0122] According to an embodiment of the present disclosure, retrieving information matching the image may include: determining a similar retrieved image matching the image from multimodal information by using image similarity, where the multimodal information stored in the multimodal knowledge base includes retrieved images and retrieved texts; and determining a retrieved text for describing the similar retrieved image from the multimodal information as the information matching the image.

[0123] According to an embodiment of the present disclosure, information matching the image can be retrieved through the image to obtain second retrieved information. As an example, the multimodal knowledge base may store multiple pieces of multimodal information, and the multimodal information may include retrieved images and retrieved texts, where the retrieved texts can be used to describe the retrieved images. For example, the multimodal knowledge base can be the knowledge base of an image search by image tool, and the multimodal information can be a web page containing an image. The image in the web page can be used as the retrieved image, and the title of the web page and the context content corresponding to the image in the web page can be used as the retrieved text.

[0124] According to an embodiment of the present disclosure, based on the image similarity between image I and multiple pieces of multimodal information in the multimodal knowledge base, a similar retrieved image matching image I can be determined from the multiple pieces of multimodal information. Exemplarily, the image similarity between image I and multiple retrieved images can be calculated, the multiple retrieved images can be sorted according to the descending order of the image similarity, and the TopM retrieved images in the sorting can be determined as the similar retrieved images. Further, a retrieved text for describing the similar retrieved image can be determined from the multiple pieces of multimodal information, and the retrieved text can be used as the second retrieved information.

[0125] For example, image I can be input into an image search by image tool to retrieve the web pages where several similar images are located. For example, the first five web page titles and the context content corresponding to the images in the web pages can be taken as the second retrieved information obtained by using the image search by image tool, which can be denoted as R 2 。

[0126] According to an embodiment of the present disclosure, the interaction method may further include: performing subject recognition on an image to obtain a subject recognition result; performing object recognition on the image to obtain an object recognition result; and obtaining an object in the image based on the subject recognition result and the object recognition result.

[0127] According to an embodiment of the present disclosure, an image recognition operation may be performed on image I to obtain an object in the image. Exemplarily, the image recognition operation may include a subject recognition operation and an object recognition operation. For example, a subject recognition operation may be performed on image I to obtain a subject recognition result. An object recognition operation may be performed on image I to obtain an object recognition result. An object in image I may be obtained based on the subject recognition result and the object recognition result.

[0128] Among them, the subject recognition operation can be understood as a generalized coarse-grained recognition, and the object recognition operation can be understood as a single fine-grained recognition. For example, for image I, the subject recognition result obtained by performing a subject recognition operation on image I may be a dog, while the object recognition result obtained by performing an object recognition operation on image I may be a Pekingese. It should be noted that the subject recognition operation and the object recognition operation may be two operations processed in parallel, and there is no limitation here.

[0129] For example, image I may be input into a subject recognition tool for subject recognition to obtain a subject recognition result, which may be denoted as R for example. 3 For example, image I may be input into an object recognition tool for object recognition to obtain an object recognition result. Among them, the object recognition tool may include, for example, but not limited to, a text recognition operator, an animal and plant recognition operator, a vehicle recognition operator, a LOGO recognition operator, etc. For example, the text content output by the text recognition operator may be denoted as R 4 ; for example, the content output by the animal and plant recognition operator, the vehicle recognition operator, the LOGO recognition operator, etc. may be denoted as R 5 .

[0130] According to an embodiment of the present disclosure, by performing subject recognition operations and object recognition operations with different granularities on image I respectively, the coverage range and dimension of image class retrieval can be expanded, so that the recognized objects have rich granularity, thereby improving the diversity, comprehensiveness, and richness of the second retrieval information.

[0131] According to an embodiment of the present disclosure, retrieving information matching the image includes: obtaining text for describing the object as the information matching the image.

[0132] According to an embodiment of the present disclosure, information matching the image can be retrieved through an object in the image to obtain second retrieval information. Exemplarily, text for describing the object can be obtained and used as the second retrieval information. Merely as an example, the text for describing the object can be obtained from a website, for example.

[0133] For example, R in the above text can be 3 and R 5 input into the website, and the retrieved content of the entry for describing the object can be used as the second retrieval information, which can be denoted as R for example 6 .

[0134] According to an embodiment of the present disclosure, retrieving information matching the image includes: determining the feature similarity between the image features of the image and the reference features in the reference feature set to obtain a feature similarity set, where the reference feature set is obtained by performing feature extraction on the information in the knowledge base, and the image features are obtained by performing image feature extraction on the image; and determining the information matching the image from the knowledge base based on the feature similarity set.

[0135] According to an embodiment of the present disclosure, information matching the image can be retrieved through the image features of the image to obtain second retrieval information. As an example, the knowledge base can store multiple image description texts, which can include, for example, image description texts of multiple categories such as poster stills, anime characters, food, cars, etc. Merely as an example, the structure of the knowledge base can be as shown in Table 1 below.

[0136]

[0137] Table 1

[0138] According to an embodiment of the present disclosure, multiple image description texts in the knowledge base can be pre - feature - extracted according to image category labels respectively to obtain a reference feature set, which can be denoted as C for example t . Image feature extraction can be performed on image I to obtain the image features of the image, which can be denoted as C for example i . The feature similarity between the image features C i and the reference features in the reference feature set C t can be determined to obtain a feature similarity set. Exemplarily, the feature similarity between the image features C i and the reference features in the reference feature set C t can be calculated, the multiple image description texts in the knowledge base can be sorted according to the descending order of the feature similarity, and the TopS image description texts in the sorting can be determined as the ones matching image I as the second retrieval information, which can be denoted as R for example 7 .

[0139] According to an embodiment of the present disclosure, information related to the image I and the text Q can be aggregated to have [R 1 , R 2 , R 4 ,R 6 , R 7 , and the multi-source retrieval information R can be obtained. By performing multi-source multi-modal retrieval on the text Q and the image I respectively from multiple information sources, the multi-source retrieval information R can provide more comprehensive, richer, and more accurate external knowledge related to the multi-modal problem.

[0140] It should be noted that one or more of the preset knowledge base, multi-modal knowledge base, knowledge set, knowledge base, etc. mentioned above can be combined or replaced arbitrarily. The above is only a distinction of the knowledge base in terms of expression, rather than a limitation on the type, structure, etc. of the knowledge base.

[0141] Figure 5 Schematically shows a schematic diagram of obtaining multi-source retrieval information through multi-source multi-modal retrieval according to an embodiment of the present disclosure.

[0142] As Figure 5 shown, based on the image feature 511 of the image I510 and the object 512 in the image I510, as well as the main content 521 of the text Q420 and the text feature 522 of the text Q520, multi-source multi-modal retrieval can be performed respectively from the knowledge base 530 to obtain the multi-source retrieval information 540.

[0143] According to an embodiment of the present disclosure, the multi-modal knowledge refinement large model is trained through the following training samples: sample multi-modal problems and reference answers, where the sample multi-modal problems include sample texts and sample images, and the reference answers include sample multi-source retrieval information and labels of the sample multi-source retrieval information, and the labels are used to characterize whether the sample multi-source retrieval information is noise.

[0144] According to an embodiment of the present disclosure, the multi-modal knowledge refinement large model can be trained through sample multi-modal problems and reference answers. Among them, the sample multi-modal problems can include the sample text Q train and the sample image I train . The reference answers can include the sample multi-source retrieval information R train and the label T of the sample multi-source retrieval information Rtrain, and the label T can be used to characterize the sample multi-source retrieval information R train whether it is noise. Exemplarily, the sample multi-source retrieval information R train can include the sample denoised multi-source retrieval information R´ train and the sample noisy multi-source retrieval information r train , and the sample denoised multi-source retrieval information R´train With label T´, the sample noise multi-source retrieval information r train Has label t.

[0145] According to an embodiment of the present disclosure, manual annotation or machine learning can be used to label the sample multi-source retrieval information R train To obtain the sample denoised multi-source retrieval information R´ with label T´ train And the sample noise multi-source retrieval information r with label t train .

[0146] According to an embodiment of the present disclosure, a multi-modal knowledge extraction large model can be obtained by performing instruction fine-tuning on a multi-modal large language model. Exemplarily, the method of instruction fine-tuning can be as follows:

[0147] A training sample set can be constructed based on multiple sample multi-modal problems, and the sample multi-modal problems can be denoted as {Q train , I train}. By separately retrieving the information that matches the sample text Q train and the sample image I train respectively, the sample multi-source retrieval information R train can be obtained. {Q train , I train , R train , Prompt} can be input into the multi-modal large language model for extraction, and the sample retrieval enhanced information K train is output. Among them, Prompt can be used to guide the multi-modal large language model to summarize the relevant content related to the sample problem Q train and the sample image I train by understanding the sample multi-source retrieval information R train , and summarize and extract the sample retrieval enhanced information K train . {Q train , I train , R train , K train , Prompt} can be used to perform instruction fine-tuning training on the multi-modal large language model to obtain the multi-modal knowledge extraction large model. Specifically, the sample retrieval enhanced information K train can be used as the predicted value, and the sample denoised multi-source retrieval information R´ train can be used as the reference answer, and the loss function is used to process the predicted value and the reference answer, and thus training is performed. Among them, those skilled in the art can select a suitable loss function according to actual needs or application scenarios, etc., and no limitation is made here.

[0148] According to an embodiment of the present disclosure, the sample multi-source retrieval information R trainIt can be obtained through the following method: by separately retrieving the information that matches the sample text Q train and the sample image I train respectively, to obtain the sample multi-source retrieval information R train .

[0149] Among them, retrieving the information that matches the sample text Q train may include: obtaining a set of sample candidate contents that match the sample text Q train ; and determining, based on the semantic similarity between the sample text Q train and the sample candidate contents in the set of sample candidate contents, the information that matches the sample text Q train .

[0150] Among them, obtaining the set of candidate contents that match the sample text Q train may include: extracting the sample main content in the sample text Q train ; and obtaining a set of sample candidate contents that match the sample text Q train based on the sample main content and the intention of the sample multi-modal question.

[0151] Among them, the information that matches the sample image I train can be retrieved through at least one of the following: the sample image features of the sample image I train , the sample image I train , the sample object in the sample image I train .

[0152] Among them, retrieving the information that matches the sample image I train may include: using image similarity to determine, from the sample multi-modal information, a sample similar retrieval image that matches the sample image I train , where the sample multi-modal information stored in the sample multi-modal knowledge base includes sample retrieval images and sample retrieval texts; and determining, from the sample multi-modal information, the sample retrieval text used to describe the sample similar retrieval image as the information that matches the sample image I train .

[0153] Among them, retrieving the information that matches the sample image I train through at least one of the sample image features of the sample image I train , the sample image I train , and the sample object in the sample image I train also includes:

[0154] performing main body recognition on the sample image I train to obtain a sample main body recognition result; performing main body recognition on the sample image I trainPerform object recognition to obtain the sample object recognition result; and based on the sample subject recognition result and the sample object recognition result, obtain the sample object in the sample image I train in the sample image I

[0155] Among them, retrieving information matching the sample image I train may include: obtaining the text for describing the sample object as the information matching the sample image I train in the sample image I

[0156] Among them, retrieving information matching the sample image I train may include: determining the feature similarity between the sample image features of the sample image I train and the sample reference features in the sample reference feature set, obtaining the sample feature similarity set, where the sample reference feature set is obtained by extracting features from the sample information in the sample knowledge base, and the sample image features are obtained by performing image feature extraction on the sample image I train ; and based on the sample feature similarity set, determining the information matching the sample image I train from the sample knowledge base

[0157] Figure 6 Schematically shows a block diagram of an artificial intelligence-based interaction device according to an embodiment of the present disclosure

[0158] As Figure 6 shown, the artificial intelligence-based interaction device 600 may include an acquisition module 610, a retrieval module 620, a retrieval enhancement module 630, and a first reply module 640

[0159] The acquisition module 610 is configured to acquire a multimodal question, where the multimodal question includes text and an image

[0160] The retrieval module 620 is configured to respectively retrieve the information matching the text and the image, and obtain multi-source retrieval information

[0161] The retrieval enhancement module 630 is configured to use the multimodal knowledge refinement large model to perform multitask processing on the multi-source retrieval information based on the multimodal question, and obtain retrieval enhancement information

[0162] The first reply module 640 is configured to use the multimodal interaction large model to process the retrieval enhancement information based on the multimodal question, and obtain the reply content for the multimodal question

[0163] According to an embodiment of the present disclosure, the retrieval enhancement module 630 may include a noise reduction sub-module and a refinement sub-module

[0164] A noise reduction sub-module for using multimodal knowledge to refine the large model to perform noise reduction on multi-source retrieval information based on the multimodal problem, and obtaining the noise-reduced multi-source retrieval information; and

[0165] A refinement sub-module for using multimodal knowledge to refine the large model to perform refinement on the noise-reduced multi-source retrieval information based on the multimodal problem, and obtaining the retrieval-enhanced information.

[0166] According to an embodiment of the present disclosure, the noise reduction sub-module may include a first noise reduction unit and a second noise reduction unit.

[0167] The first noise reduction unit is configured to determine the correlation between the multimodal problem and the multi-source retrieval information, and obtain a correlation recognition result.

[0168] The second noise reduction unit is configured to, when the correlation recognition result indicates that the multi-source retrieval information is relevant to the multimodal problem, obtain the noise-reduced multi-source retrieval information based on the multi-source retrieval information.

[0169] According to an embodiment of the present disclosure, the refinement sub-module may include a refinement unit.

[0170] The refinement unit is configured to sequentially perform content refinement and logical arrangement on the noise-reduced multi-source retrieval information, and obtain the retrieval-enhanced information.

[0171] According to an embodiment of the present disclosure, the artificial intelligence-based interaction device 600 may further include a noise information determination module and a second reply module.

[0172] The noise information determination module is configured to, when the correlation recognition result indicates that the multi-source retrieval information is not relevant to the multimodal problem, determine that the multi-source retrieval information is noise multi-source retrieval information.

[0173] The second reply module is configured to, when it is determined that the multi-source retrieval information is noise multi-source retrieval information, use the multimodal interaction large model to process the multimodal problem, and obtain a reply content.

[0174] According to an embodiment of the present disclosure, the first reply module 640 may include a reply sub-module and an optimization sub-module.

[0175] The reply sub-module is configured to use the multimodal interaction large model to process the retrieval-enhanced information based on the multimodal problem, and obtain an initial reply content.

[0176] The optimization sub-module is configured to optimize the initial reply content based on the intention of the multimodal problem, and obtain a reply content.

[0177] According to an embodiment of the present disclosure, the optimization sub-module may include a target content determination unit and an update unit.

[0178] A target content determination unit, configured to determine target content meeting a predetermined format from initial response content when the intent includes a recommendation intent.

[0179] An update unit, configured to update the target content based on multi-source retrieval information corresponding to the target content to obtain response content.

[0180] According to an embodiment of the present disclosure, the optimization sub-module may include a format conversion unit.

[0181] The format conversion unit is configured to perform format conversion on the initial response content to obtain response content when the intent includes a question-and-answer intent.

[0182] According to an embodiment of the present disclosure, the retrieval module 620 may include an acquisition sub-module and a matching sub-module.

[0183] The acquisition sub-module is configured to acquire a candidate content set matching the text.

[0184] The matching sub-module is configured to determine information matching the text from the candidate content set based on the semantic similarity between the text and the candidate content in the candidate content set.

[0185] According to an embodiment of the present disclosure, the acquisition sub-module may include an extraction unit and an acquisition unit.

[0186] The extraction unit is configured to extract the main content in the text.

[0187] The acquisition unit is configured to acquire a candidate content set matching the text based on the main content and the intent of the multi-modal question.

[0188] According to an embodiment of the present disclosure, the retrieval module 620 may include at least one of a first image retrieval sub-module, a second image retrieval sub-module, and a third image retrieval sub-module.

[0189] The first image retrieval sub-module is configured to retrieve information matching the image by the image.

[0190] The second image retrieval sub-module is configured to retrieve information matching the image by an object in the image.

[0191] The third image retrieval sub-module is configured to retrieve information matching the image by image features of the image.

[0192] According to an embodiment of the present disclosure, the first image retrieval sub-module may include a first retrieval unit and a first determination unit.

[0193] The first retrieval unit is configured to use image similarity to determine a similar retrieval image matching the image from multi-modal information, where the multi-modal information stored in the multi-modal knowledge base includes retrieval images and retrieval texts.

[0194] A first determination unit, configured to determine, from the multimodal information, a retrieval text for describing a similar retrieved image as information matching the image.

[0195] According to an embodiment of the present disclosure, the retrieval module 620 may further include a subject recognition module, an object recognition module, and an object determination module.

[0196] The subject recognition module is configured to perform subject recognition on the image to obtain a subject recognition result.

[0197] The object recognition module is configured to perform object recognition on the image to obtain an object recognition result.

[0198] The object determination module is configured to obtain an object in the image based on the subject recognition result and the object recognition result.

[0199] According to an embodiment of the present disclosure, the second image sub-retrieval module may include a second retrieval unit.

[0200] The second retrieval unit is configured to obtain a text for describing the object as information matching the image.

[0201] According to an embodiment of the present disclosure, the third image sub-retrieval module may include a third determination unit and a third retrieval unit.

[0202] The third determination unit is configured to determine a feature similarity between the image feature of the image and a reference feature in a reference feature set, to obtain a feature similarity set, where the reference feature set is obtained by performing feature extraction on the information in the knowledge base, and the image feature is obtained by performing image feature extraction on the image.

[0203] The third retrieval unit is configured to determine, based on the feature similarity set, information matching the image from the knowledge base.

[0204] According to an embodiment of the present disclosure, the multimodal knowledge refinement large model is trained using the following training samples:

[0205] Sample multimodal questions and reference answers, where the sample multimodal questions include sample texts and sample images, and the reference answers include sample multi-source retrieval information and labels of the sample multi-source retrieval information, and the labels are used to characterize whether the sample multi-source retrieval information is noise.

[0206] It should be noted that in the embodiments of the present disclosure, the part of the artificial intelligence-based interaction device corresponds to the part of the artificial intelligence-based interaction method in the embodiments of the present disclosure. For the description of the part of the artificial intelligence-based interaction device, refer specifically to the part of the artificial intelligence-based interaction method, and details are not described herein again.

[0207] Figure 7A block diagram of an AI-based agent according to an embodiment of the present disclosure is schematically shown.

[0208] In an embodiment of the present disclosure, as Figure 7 shown, an AI-based agent such as AI agent 700 may include an input module 710, a processing module 720, and an output module 730.

[0209] The input module 710 is configured to receive multimodal questions.

[0210] The processing module 720 is configured to, based on the multimodal questions received by the input module, execute the above-described AI-based interaction method by invoking a large model to obtain a reply content.

[0211] The output module 730 is configured to output the reply content obtained by the processing module.

[0212] According to an embodiment of the present disclosure, the input module 710 is responsible for receiving or perceiving information such as queries, requests, instructions, signals, or data from the outside world (e.g., users or the external environment), and converting it into a format that the AI agent 700 can understand and process. The input module 710 is the primary link for the AI agent 700 to interact with the outside world, enabling the AI agent 700 to efficiently and accurately obtain necessary "sensory" information from the outside world and respond to this information.

[0213] In an example, the input module 710 may input the multimodal questions described above.

[0214] In an example, the processing module 720 is the core support for the AI agent 700's ability to handle complex tasks. The processing module 720 may execute the AI-based interaction method described above.

[0215] In an example, the performance of the processing module 720 may be closely related to the large model on which the AI agent 700 is based. To fully utilize the capabilities of the large model, the internal structure of the processing module 720 can be designed to be highly configurable and extensible to handle various different types of tasks and requirements in real-world scenarios.

[0216] In an example, after the AI agent 700 obtains a multimodal question, the processing module 720 may use the large model to separately retrieve information that matches the text and images in the multimodal question, respectively, to obtain multi-source retrieval information. The processing module 720 may invoke a multimodal knowledge refinement large model to perform multitask processing on the multi-source retrieval information based on the multimodal question to obtain retrieval-enhanced information.

[0217] The processing module 720 may call the multi-modal interaction large model to process the retrieved enhanced information based on the multi-modal question, obtain the reply content for the multi-modal question, and transmit the reply content to the output module 730.

[0218] It can be understood that although the large model has excellent language understanding and generation capabilities, like humans, without any tools, the tasks it can solve are very limited. When the AI agent 1000 is given the ability to call tools, it can perform tasks such as completing mathematical operations with the help of a calculator, completing data analysis with the help of the Python language, and completing weather forecasts with the help of a search engine.

[0219] In the example, the output module 730 may output the reply content described above.

[0220] The AI agent 700 according to the embodiments of the present disclosure can simply and effectively improve the degree of intelligence, and improve flexibility and versatility.

[0221] According to the embodiments of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0222] According to an embodiment of the present disclosure, an electronic device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method as in the embodiments of the present disclosure.

[0223] According to an embodiment of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to execute the method as in the embodiments of the present disclosure.

[0224] According to an embodiment of the present disclosure, a computer program product includes a computer program, and the computer program implements the method as in the embodiments of the present disclosure when executed by a processor.

[0225] Figure 8 FIG. shows a schematic block diagram of an exemplary electronic device 800 that can be used to implement the embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely exemplary and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0226] AsFigure 8 As shown in Figure 8 , device 800 includes a computing unit 801 that can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 802 or a computer program loaded from a storage unit 808 into a random access memory (RAM) 803. In the RAM 803, various programs and data required for the operation of device 800 can also be stored. The computing unit 801, ROM 802, and RAM 803 are connected to each other via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.

[0227] Multiple components in device 800 are connected to the input / output (I / O) interface 805, including: an input unit 806, such as a keyboard, mouse, etc.; an output unit 807, such as various types of displays, speakers, etc.; a storage unit 808, such as a magnetic disk, optical disc, etc.; and a communication unit 809, such as a network card, modem, wireless communication transceiver, etc. The communication unit 809 allows device 800 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0228] The computing unit 801 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 801 include but are not limited to a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The computing unit 801 executes the various methods and processes described above, such as the artificial intelligence-based interaction method. For example, in some embodiments, the artificial intelligence-based interaction method can be implemented as a computer software program that is tangibly contained in a machine-readable medium, such as the storage unit 808. In some embodiments, part or all of the computer program can be loaded and / or installed onto device 800 via the ROM 802 and / or the communication unit 809. When the computer program is loaded into the RAM 803 and executed by the computing unit 801, one or more steps of the artificial intelligence-based interaction method described above can be executed. Alternatively, in other embodiments, the computing unit 801 can be configured to execute the artificial intelligence-based interaction method in any other appropriate manner (e.g., by means of firmware).

[0229] The various embodiments of the systems and techniques described above in this specification can be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGA), application specific integrated circuits (ASIC), application specific standard products (ASSP), systems on a chip (SOC), complex programmable logic devices (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which may be a special-purpose or general-purpose programmable processor that receives data and instructions from, and transmits data and instructions to, a storage system, at least one input device, and at least one output device.

[0230] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general purpose computer, special purpose computer, or other programmable data processing device, such that the program codes, when executed by the processor or controller, cause the functions / operations specified in the flowchart and / or block diagram to be implemented. The program code may execute entirely on the machine, partly on the machine, as a stand-alone software package partly on the machine and partly on a remote machine or entirely on the remote machine or server.

[0231] In the context of the present disclosure, a machine-readable medium may be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0232] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and a pointing device (e.g., a mouse or a trackball) through which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, speech input, or tactile input).

[0233] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), and the Internet.

[0234] A computer system can include a client and a server. The client and the server are generally remote from each other and typically interact through a communication network. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, a server of a distributed system, or a server incorporating a blockchain.

[0235] It should be understood that various forms of the processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and no limitation is imposed herein.

[0236] The above specific embodiments do not constitute a limitation on the protection scope of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the protection scope of this disclosure.

Claims

1. An interactive method based on artificial intelligence, comprising: Acquire a multimodal question, wherein the multimodal question includes text and an image; Retrieve information matching the text and the image respectively to obtain multi-source retrieval information; Using a multimodal knowledge extraction large model to perform multi-task processing on the multi-source retrieval information based on the multimodal problem to obtain retrieval enhancement information; and The retrieval enhancement information is processed based on the multimodal question using a multimodal interaction macro model to obtain response content for the multimodal question.

2. The method according to claim 1, wherein: The method of using the multimodal knowledge to extract a large model to perform multi-task processing on the multi-source retrieval information based on the multimodal problem to obtain retrieval enhancement information includes: Utilizing the multimodal knowledge extraction large model to reduce noise on the multi-source search information based on the multimodal problem, to obtain the reduced noise multi-source search information; and The multimodal knowledge extraction large model is used to extract the denoised multi-source retrieval information based on the multimodal problem to obtain the retrieval enhancement information.

3. The method according to claim 2, wherein: The method of using the multimodal knowledge extraction large model to reduce noise on the multi-source search information based on the multimodal problem to obtain the reduced noise multi-source search information includes: Determining the correlation between the multimodal question and the multi-source retrieval information to obtain a correlation identification result; and In a case where the correlation identification result indicates that the multi-source retrieval information is correlated with the multimodal problem, the denoised multi-source retrieval information is obtained based on the multi-source retrieval information.

4. The method according to claim 2 or 3, wherein: The method of refining the denoised multi-source retrieval information based on the multi-modal problem using the multi-modal knowledge refining large model to obtain the retrieval enhancement information includes: The denoised multi-source retrieval information is sequentially subjected to content extraction and logic arrangement to obtain the retrieval enhancement information.

5. The method according to claim 3, further comprising: In a case where the correlation identification result indicates that there is no correlation between the multi-source retrieval information and the multimodal problem, determining that the multi-source retrieval information is noisy multi-source retrieval information; When it is determined that the multi-source retrieval information is noisy multi-source retrieval information, the multi-modal question is processed using the multi-modal interaction large model to obtain the reply content.

6. The method according to any one of claims 1 to 5, wherein: The method of processing the search enhancement information based on the multimodal question using the multimodal interaction macro model to obtain a response content for the multimodal question includes: Processing the search enhancement information based on the multimodal question using the multimodal interaction macro model to obtain initial reply content; and Based on the intent of the multimodal question, the initial response content is optimized to obtain the response content.

7. The method according to claim 6, wherein: The step of optimizing the initial response content based on the intention of the multimodal question to obtain the response content includes: In a case where the intent includes a recommendation intent, determining target content that complies with a predetermined format from the initial reply content; and Based on the multi-source search information corresponding to the target content, the target content is updated to obtain the reply content.

8. The method according to claim 6 or 7, wherein: The step of optimizing the initial response content based on the intention of the multimodal question to obtain the response content includes: In the case where the intent includes a question-and-answer intent, the initial reply content is format-converted to obtain the reply content.

9. The method according to any one of claims 1 to 8, wherein: Retrieve information that matches the text, including: Obtaining a set of candidate content matching the text; and Based on the semantic similarity between the text and candidate content in the candidate content set, information matching the text is determined from the candidate content set.

10. The method according to claim 9, wherein: The step of obtaining a candidate content set matching the text includes: extracting the main content of the text; and Based on the main content and the intention of the multimodal question, the candidate content set matching the text is obtained.

11. The method according to any one of claims 1 to 10, wherein: Retrieve information matching the image by at least one of the following: image features of the image, the image, and an object in the image.

12. The method according to claim 11, wherein: Retrieve information matching the image, including: Determining a similar retrieval image matching the image from multimodal information using image similarity, wherein the multimodal information stored in a multimodal knowledge base includes a retrieval image and a retrieval text; and A search text for describing the similar search image is determined from the multimodal information as information matching the image.

13. The method according to claim 11 or 12, further comprising: Performing subject recognition on the image to obtain a subject recognition result; Performing object recognition on the image to obtain an object recognition result; as well as Based on the subject recognition result and the object recognition result, the object in the image is obtained.

14. The method according to claim 11 or 13, wherein: Retrieve information matching the image, including: A text describing the object is obtained as information matched with the image.

15. The method according to any one of claims 11 to 14, wherein: Retrieve information matching the image, including: Determine feature similarity between image features of the image and reference features in a reference feature set to obtain a feature similarity set, wherein the reference feature set is obtained by extracting features from information in a knowledge base, and the image features are obtained by extracting image features from the image; and Based on the feature similarity set, information matching the image is determined from the knowledge base.

16. The method according to any one of claims 1 to 15, wherein: The multimodal knowledge extraction model is obtained by training with the following training samples: Sample multimodal questions and reference answers, wherein the sample multimodal questions include sample text and sample images, and the reference answers include sample multi-source retrieval information and a label of the sample multi-source retrieval information, wherein the label is used to characterize whether the sample multi-source retrieval information is noise.

17. An interactive device based on artificial intelligence, comprising: An acquisition module, used for acquiring a multimodal question, wherein the multimodal question includes text and an image; A retrieval module, used to retrieve information matching the text and the image respectively, to obtain multi-source retrieval information; A retrieval enhancement module, configured to perform multi-task processing on the multi-source retrieval information based on the multi-modal problem using a multi-modal knowledge extraction macro model to obtain retrieval enhancement information; and The reply module is used to process the retrieval enhancement information based on the multimodal question using a multimodal interaction macro model to obtain reply content for the multimodal question.

18. An artificial intelligence based agent configured to perform the method according to any one of claims 1 to 16.

19. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 16.

20. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 16.

21. A computer program product comprising a computer program which, when executed by a processor, implements the method according to any one of claims 1 to 16.

Citation Information

Cited By

  • Multi-source information retrieval fusion method, device and equipment and readable storage medium

    CN120541310A

  • Interaction method and device based on large model, intelligent agent and storage medium

    CN120706575A

  • Voice interaction method based on retrieval enhancement generation and multi-model collaboration and application thereof

    CN121545518A

  • An ai union wisdom hub system and method based on multi-modal interaction

    CN122736565A