Question answering method and device based on artificial intelligence, intelligent agent, equipment and medium

By introducing reflection-based question-and-answer methods into the big model, the problem of insufficient reliability, timeliness and professionalism in generating answers is solved, and more accurate and comprehensive knowledge processing is achieved, and more valuable and credible information is provided.

CN120197706APending Publication Date: 2025-06-24BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510337954.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-20
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

Large models have problems with insufficient reliability, timeliness and professionalism when generating answers, especially when there is little knowledge and poor timeliness.

Method used

By introducing a reflection-based question-and-answer method in the big model, the questions are obtained and the relevant knowledge set is retrieved, and the reflection results are used to determine whether the current amount of knowledge is sufficient to answer the questions, and the subset of knowledge is updated when there is insufficient.

Benefits of technology

It improves the authority and professionalism of the answers generated by the big model, so that it can grasp the details of knowledge more accurately and cover the knowledge field more comprehensively, thereby providing more valuable and credible information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120197706A_ABST
    Figure CN120197706A_ABST
Patent Text Reader

Abstract

The invention provides a question answering method and device based on artificial intelligence, an intelligent agent, electronic equipment, a storage medium and a program product, and relates to the technical field of artificial intelligence, in particular to the technical fields of computer vision, deep learning, large models and the like. The method can be applied to scenes such as AIGC content generation based on artificial intelligence. According to the specific implementation scheme, the method comprises the steps of obtaining a question and a knowledge set obtained based on question retrieval; calling the large model to execute the following operations based on the problem and the knowledge set: reflecting the problem and the current knowledge subset to obtain a reflecting result, the knowledge set comprising knowledge in the current knowledge subset; under the condition that the reflection result represents that the current knowledge subset reaches the knowledge quantity for answering the question, obtaining an answer based on the question and the current knowledge subset; under the condition that the reflection result represents that the current knowledge subset does not reach the knowledge quantity used for answering the question, the current knowledge subset is updated based on the remaining knowledge subset, and the remaining knowledge subset comprises knowledge in the knowledge set except the current knowledge subset.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence technology, especially to computer vision, deep learning, large models and other technical fields; it can be applied to scenarios such as AIGC (Artificial Intelligence Generated Content, content generation based on artificial intelligence). Specifically, it relates to a question-answering method, device, intelligent agent, electronic device, storage medium and program product based on artificial intelligence. Background Art

[0002] Large models can understand and generate natural language and can handle a variety of natural language tasks, such as text generation, knowledge question answering, reasoning calculation, reading comprehension, etc. However, the model knowledge learned by large models themselves is often small in amount and poor in timeliness, so they need to use retrieval technology to obtain external resources to assist in task execution. How to effectively use external resources has become a research focus. Summary of the invention

[0003] The present disclosure provides a question-answering method, device, intelligent agent, electronic device, storage medium and program product based on artificial intelligence.

[0004] According to one aspect of the present disclosure, there is provided an artificial intelligence-based question-answering method, comprising: obtaining a question and a knowledge set retrieved based on the question; calling a large model to perform the following operations based on the question and the knowledge set: reflecting on the question and the current knowledge subset to obtain a reflection result, the knowledge set including the knowledge in the current knowledge subset; obtaining an answer based on the question and the current knowledge subset when the reflection result indicates that the current knowledge subset has reached the amount of knowledge for answering the question; and updating the current knowledge subset based on the remaining knowledge subset when the reflection result indicates that the current knowledge subset has not reached the amount of knowledge for answering the question. The remaining knowledge subset includes the knowledge in the knowledge set other than the current knowledge subset.

[0005] According to another aspect of the present disclosure, there is provided an artificial intelligence-based question-answering device, including: an acquisition module configured to acquire a question and a knowledge set retrieved based on the above question; a calling module configured to call a large model to perform the following operations based on the above question and the above knowledge set: reflect on the above question and the current knowledge subset to obtain a reflection result, where the knowledge set includes the knowledge in the above current knowledge subset; when the reflection result indicates that the current knowledge subset has reached the amount of knowledge for answering the above question, obtain an answer based on the above question and the above current knowledge subset; and when the reflection result indicates that the current knowledge subset has not reached the amount of knowledge for answering the above question, update the current knowledge subset based on the remaining knowledge subset, where the remaining knowledge subset includes the knowledge in the knowledge set except the above current knowledge subset.

[0006] According to another aspect of the present disclosure, there is provided an artificial intelligence-based agent, where the above agent is configured to execute the method as described above.

[0007] According to another aspect of the present disclosure, there is provided an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; where the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method as described above.

[0008] According to another aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, where the above computer instructions are used to cause the computer to execute the method as described above.

[0009] According to another aspect of the present disclosure, there is provided a computer program product including a computer program, where the computer program implements the method as described above when executed by a processor.

[0010] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understandable through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:

[0012] Figure 1 Schematically shows an exemplary system architecture to which the artificial intelligence-based question-answering method and device according to the embodiments of the present disclosure can be applied;

[0013] Figure 2 Schematically shows a flowchart of the artificial intelligence-based question-answering method according to the embodiments of the present disclosure;

[0014] Figure 3 Schematically shows a flowchart of an artificial intelligence-based question-answering method according to another embodiment of the present disclosure;

[0015] Figure 4 Schematically shows a flowchart of multimodal knowledge retrieval according to an embodiment of the present disclosure;

[0016] Figure 5 Schematically shows a schematic diagram of a multimodal knowledge unit according to an embodiment of the present disclosure;

[0017] Figure 6 Schematically shows a flowchart of an artificial intelligence-based question-answering method according to an embodiment of the present disclosure;

[0018] Figure 7 Schematically shows a block diagram of an artificial intelligence-based question-answering device according to an embodiment of the present disclosure;

[0019] Figure 8 Schematically shows a block diagram of an intelligent agent of artificial intelligence according to an embodiment of the present disclosure; and

[0020] Figure 9 Schematically shows a block diagram of an electronic device suitable for implementing an artificial intelligence-based question-answering method according to an embodiment of the present disclosure. Detailed Embodiments

[0021] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, descriptions of well-known functions and structures are omitted below for clarity and conciseness.

[0022] In recent years, with the rapid development of artificial intelligence, large models such as vision-language large models (LVLMs) have shown great potential in multimodal aspects such as understanding of text and image content and human conversations. Nevertheless, there are still obvious drawbacks in generating answers solely relying on the knowledge of large models. One is the lack of reliability. The content information sources generated by the model relying on its own knowledge are uncertain, and there may be hallucination problems, thus posing a risk of information misleading. The second is the lack of timeliness. The timeliness of knowledge in large models depends on the training data. After the large model is trained, the knowledge is solidified and cannot be updated in real time during inference. The third is the lack of professionalism. Large models are generally trained on general datasets, resulting in the knowledge they master being broad but not refined, and lacking in-depth understanding of fine-grained or specialized knowledge.

[0023] In view of this, the present disclosure provides a question-answering method, apparatus, intelligent agent, electronic device, storage medium, and program product based on artificial intelligence. It is expected to use a large model to perform progressive reflective thinking chain operations on the retrieved knowledge set, accurately deliver fine-grained, comprehensive, and in-depth long-tail knowledge to the large model, enhance the authority and professionalism of the answers generated by the large model, enable it to more accurately grasp knowledge details, and more comprehensively cover knowledge fields, so as to provide users with more valuable and credible information.

[0024] Figure 1 Schematically shows an exemplary system architecture to which the question-answering method and apparatus based on artificial intelligence according to embodiments of the present disclosure can be applied.

[0025] It should be noted that Figure 1 The shown is only an example of the system architecture to which embodiments of the present disclosure can be applied, to help those skilled in the art understand the technical content of the present disclosure, but it does not mean that embodiments of the present disclosure cannot be used in other devices, systems, environments, or scenarios. For example, in another embodiment, the exemplary system architecture to which the question-answering method and apparatus based on artificial intelligence can be applied may include terminal devices, but the terminal devices can implement the question-answering method and apparatus provided by embodiments of the present disclosure without interacting with the server.

[0026] As Figure 1 shown, the system architecture 100 according to this embodiment may include terminal devices 101, 102, 103, a network 104, and a server 105. The network 104 is used to provide a medium for communication links between the terminal devices 101, 102, 103 and the server 105. The network 104 may include various connection types, such as wired and / or wireless communication links, etc.

[0027] Users can use the terminal devices 101, 102, 103 to interact with the server 105 through the network 104 to receive or send messages, etc. Various communication client applications may be installed on the terminal devices 101, 102, 103, such as knowledge reading applications, web browser applications, search applications, instant messaging tools, email clients, and / or social platform software, etc. (only as examples).

[0028] The terminal devices 101, 102, 103 may be various electronic devices with a display screen and supporting web browsing, including but not limited to smartphones, tablets, laptop computers, and desktop computers, etc.

[0029] Server 105 may be a server that provides various services, such as a background management server (for example only) that supports the content browsed by users using terminal devices 101, 102, and 103. The background management server may analyze and process data such as user requests received, and feedback the processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal devices.

[0030] It should be noted that the artificial intelligence-based question-and-answer method provided by the embodiments of the present disclosure can generally be executed by terminal devices 101, 102, or 103. Correspondingly, the artificial intelligence-based question-and-answer device provided by the embodiments of the present disclosure can also be set in terminal devices 101, 102, or 103.

[0031] Alternatively, the artificial intelligence-based question-and-answer method provided by the embodiments of the present disclosure can generally also be executed by server 105. Correspondingly, the artificial intelligence-based question-and-answer device provided by the embodiments of the present disclosure can generally be set in server 105. The artificial intelligence-based question-and-answer method provided by the embodiments of the present disclosure can also be executed by a server or a server cluster different from server 105 and capable of communicating with terminal devices 101, 102, 103 and / or server 105. Correspondingly, the artificial intelligence-based question-and-answer device provided by the embodiments of the present disclosure can also be set in a server or a server cluster different from server 105 and capable of communicating with terminal devices 101, 102, 103 and / or server 105.

[0032] It should be understood that Figure 1 the numbers of the terminal devices, networks, and servers in

[0033] In the technical solution of the present disclosure, the processing of collection, storage, use, processing, transmission, provision, disclosure, and application of the user's personal information involved all comply with the provisions of relevant laws and regulations, take necessary confidentiality measures, and do not violate public order and good customs.

[0034] In the technical solution of the present disclosure, before obtaining or collecting the user's personal information, the authorization or consent of the user has been obtained.

[0035] It should be noted that the sequence numbers of the respective operations in the following methods are only used as representations of the operations for description, and should not be regarded as indicating the execution sequence of the respective operations. Unless explicitly stated, the method does not need to be executed exactly in the order shown.

[0036] Figure 2 Schematically shows a flowchart of an artificial intelligence-based question-and-answer method according to an embodiment of the present disclosure.

[0037] AsFigure 2 As shown, the method includes operations S210 - S220, S221 - S223.

[0038] In operation S210, a question and a knowledge set retrieved based on the question are obtained.

[0039] In operation S220, a large model is called to perform the following operations based on the question and the knowledge set:

[0040] In operation S221, the question and the current knowledge subset are reflected on to obtain a reflection result, where the knowledge set includes the knowledge in the current knowledge subset.

[0041] In operation S222, when the reflection result indicates that the current knowledge subset has reached the amount of knowledge for answering the question, an answer is obtained based on the question and the current knowledge subset.

[0042] In operation S223, when the reflection result indicates that the current knowledge subset has not reached the amount of knowledge for answering the question, the current knowledge subset is updated based on the remaining knowledge subset, where the remaining knowledge subset includes the knowledge in the knowledge set other than the current knowledge subset.

[0043] A question, which can also be called a query, can be input through the input box of the human - machine interaction interface. The type of the question is not limited, for example, it includes one or more of text, image, video, and audio.

[0044] The knowledge set can include relevant knowledge for answering the question. The type of the knowledge in the knowledge set is not limited, for example, it includes one or more of text, image, video, and audio, as long as it is knowledge related to the question.

[0045] For example, if the question involves the target object A, the knowledge in the knowledge set is also related to the target object A.

[0046] For the large model, the network structure of the large model is not limited. For example, it can include a visual - language large model (LVLM, Large Visual - Language Models) or a large - language model (LLM, Large Language Models), as long as it can perform the dual functions of reflection and answering questions.

[0047] Reflection can refer to thinking or judgment, which is an operation of judging or thinking about whether the amount of knowledge in the current knowledge subset meets the requirement for answering the question and obtaining a result.

[0048] Knowledge quantity, an indicator for measuring the amount of knowledge. The knowledge quantity can not only represent the content quantity of knowledge, but also represent the richness of knowledge. The knowledge quantity required to answer a question can be used as the knowledge quantity threshold, and the knowledge quantity threshold can be used as a measurement standard or a trigger condition. When the knowledge quantity of the current knowledge subset reaches the knowledge quantity threshold, it is determined that the operation of generating an answer can be performed. When the knowledge quantity of the current knowledge subset does not reach the knowledge quantity threshold, it is determined that the current knowledge subset does not meet the trigger condition for answering the question, and the current knowledge subset can be updated based on the knowledge in the remaining knowledge subsets of the knowledge set.

[0049] For example, based on a question about the target object A, a knowledge set is retrieved. The knowledge set includes knowledge A, knowledge B, knowledge C, and knowledge D. Knowledge A is used as the current knowledge subset, and knowledge B, knowledge C, and knowledge D are used as the remaining knowledge subsets. The large model is used to reflect on knowledge A and the question to obtain the first-round reflection result. When the first-round reflection result indicates that knowledge A reaches the knowledge quantity required to answer the question, an answer generation operation is performed based on knowledge A and the question to obtain an answer. Otherwise, at least one piece of knowledge from the remaining knowledge subsets is added to the current knowledge subset as the new current knowledge subset. The reflection process is repeated until the knowledge quantity of the current knowledge subset can meet the trigger condition for answering the question, and then an answer is generated based on the current knowledge subset and the question.

[0050] The answer can be displayed on the human-computer interaction interface to complete the current human-computer question-and-answer interaction.

[0051] Optionally, the above question-and-answer method can also be applied to many application scenarios such as intelligent medical health, cross-modal information search, and cross-modal recommendation systems.

[0052] Using the question-and-answer method provided by the embodiments of the present disclosure, it is possible to gradually determine whether the knowledge quantity in the current knowledge subset can answer the question through the reflection operation, thereby accurately grasping the knowledge quantity of the reference knowledge for answering the question. Furthermore, while ensuring the reliability, timeliness, and professionalism of the answer generated by the large model, it avoids the redundancy and waste of knowledge, thereby improving the execution efficiency of the large model and reducing the response time of human-computer interaction.

[0053] The above provides a general description of the question-and-answer method, and the following will introduce each operation in detail.

[0054] The following will detail how to update the current knowledge subset.

[0055] According to the embodiments of the present disclosure, for Figure 2The operation S223 shown, which updates the current knowledge subset based on the remaining knowledge subset, may include: determining a target remaining knowledge with the least amount of knowledge from the remaining knowledge subset. Adding the target remaining knowledge to the current knowledge subset.

[0056] The knowledge in the remaining knowledge subset can be sorted in ascending order of the amount of knowledge to obtain a sorting result. The target remaining knowledge is determined from the remaining knowledge subset according to the sorting result. The current knowledge subset is updated with the target remaining knowledge.

[0057] For example, based on a question about the target object A, a knowledge set is retrieved. The knowledge set includes knowledge A, knowledge B, knowledge C, and knowledge D. Knowledge A with the least amount of knowledge can be used as the current knowledge subset, and the knowledge in the remaining knowledge subset is arranged in ascending order of the amount of knowledge, such as knowledge B, knowledge C, and knowledge D. The large model is used to reflect based on knowledge A and the question to obtain a reflection result. In the case where the reflection result indicates that knowledge A does not reach the amount of knowledge required to answer the question, knowledge B in the remaining knowledge subset is added to the current knowledge subset as a new round of the current knowledge subset.

[0058] Optionally, any knowledge in the remaining knowledge subset can also be randomly added to the current knowledge subset. For example, knowledge D with the most amount of knowledge is added to the current knowledge subset.

[0059] Compared with the random update method, the method of updating the knowledge with the least amount of knowledge to the current knowledge subset can more finely control the amount of knowledge of the knowledge used to answer the question, improve the effect of knowledge screening, integration, and application, and further avoid knowledge redundancy and waste from the perspective of knowledge amount screening.

[0060] According to another embodiment of the present disclosure, updating the current knowledge subset based on the remaining knowledge subset may further include: determining a target remaining knowledge from the remaining knowledge subset based on the knowledge type of the knowledge used to answer the question. Adding the target remaining knowledge to the current knowledge subset.

[0061] The knowledge type can be understood as the form of knowledge. For example, it may include image type, text type, video type, or audio type, etc.

[0062] For example, if the knowledge type of the knowledge used to answer the question includes the image type, the knowledge of the image type can be determined from the remaining knowledge subset as the target remaining knowledge. The target remaining knowledge is added to the current knowledge subset.

[0063] Filtering out the knowledge of the knowledge type used to answer questions from the remaining knowledge subset can improve the current knowledge subset as reference knowledge for answering questions, improve the accuracy of the answer, improve the effectiveness of knowledge filtering, and further avoid knowledge redundancy and waste from the perspective of knowledge type filtering.

[0064] According to an embodiment of the present disclosure, before determining the target remaining knowledge from the remaining knowledge subset based on the knowledge type, the question-and-answer method may further include an operation: identifying the question and determining the knowledge type of the knowledge used to answer the question.

[0065] There is no limitation on the way of identifying the question. For example, it can be semantic recognition, but it is not limited to this. It can also be image recognition. As long as the question is analyzed to determine the knowledge type of the knowledge used to answer the question.

[0066] For example, in the case where the question includes a question image, the question image can be image-recognized to determine the clarity or object integrity of the question image. In the case where it is determined that the question image is blurred or partial information of the target object is incomplete, it can be determined that the knowledge type of the knowledge used to answer the question includes the image type.

[0067] Also, for example, in the case where the question includes a question text, the question text can be semantically recognized or intention-recognized. In the case where it is determined that the question text represents the attribute information used to determine the target object, it can be determined that the knowledge type of the knowledge used to answer the question includes the text type, and it is the knowledge of the text type regarding the attribute information of the target object.

[0068] Specifically, the question text may include "What is the price of Model B of Brand A cars?", then it is determined that the knowledge type of the knowledge used to answer the question includes the text type, specifically the knowledge of the text type regarding the price of Model B of Brand A cars.

[0069] By invoking the large model to identify the question, determining the knowledge type of the knowledge used to answer the question, and then filtering the target remaining knowledge, the accuracy and effectiveness of knowledge filtering are improved, and thus the execution efficiency of the large model is improved.

[0070] The above has described in detail how to update the current knowledge subset. The following will explain whether to obtain the knowledge set.

[0071] According to an embodiment of the present disclosure, before performing the operation S210 as Figure 2 shown, the question-and-answer method may further include the following operations.

[0072] For example, call a large model to reflect on the question and obtain a primary reflection result. In the case where the model knowledge represented by the primary reflection result does not reach the amount of knowledge required to answer the question, perform a retrieval based on the question to obtain a knowledge set.

[0073] Before calling the large model for a retrieval operation, a reflection operation can be used to reflect on the model knowledge of the large model itself to determine the primary reflection result of whether the model knowledge represented by the large model reaches the amount of knowledge required to answer the question. In the case where it reaches, directly generate an answer based on the question. In the case where it does not reach, perform a retrieval based on the question to obtain a knowledge set.

[0074] By reflecting on the model knowledge of the large model, it is possible to pre-judge whether to perform a retrieval, and different primary reflection results perform different operations, thereby improving the flexibility of answer generation in the question-and-answer interaction scenario while ensuring the effectiveness of the answer.

[0075] The following will Figure 3 describe the multi-level reflection operation of the large model.

[0076] Figure 3 Schematically shows a flowchart of an artificial intelligence-based question-and-answer method according to another embodiment of the present disclosure.

[0077] As Figure 3 shown, the method includes operations S310 to S380.

[0078] In operation S310, call a large model to reflect on the question and obtain a primary reflection result.

[0079] In operation S320, determine whether the primary reflection result represents that the model knowledge reaches the amount of knowledge required to answer the question, that is, the knowledge amount threshold. In the case where the primary reflection result represents that the model knowledge of the large model does not reach the knowledge amount threshold, perform operation S330, otherwise, perform operation S380.

[0080] In operation S330, perform a retrieval based on the question to obtain a knowledge set.

[0081] In operation S340, call a large model to reflect on the question and the current knowledge subset to obtain a reflection result.

[0082] In operation S350, determine whether the reflection result represents that the current knowledge subset reaches the knowledge amount threshold. In the case where the reflection result represents that the current knowledge subset does not reach the knowledge amount threshold, perform operation S360, otherwise, perform operation S370.

[0083] In operation S360, update the current knowledge subset based on the remaining knowledge subset.

[0084] In operation S370, an answer is obtained based on the question and the current subset of knowledge.

[0085] In operation S380, an answer is obtained based on the question.

[0086] The above text has described multi-level reflection and whether to obtain a knowledge set. The following will explain how to obtain a knowledge set.

[0087] Optionally, the question may include a multi-modal question, for example, it may include question text and a question image.

[0088] In the case where the question includes a question image and question text, at least one of the question image and the question text can be used for retrieval, such as multi-modal knowledge retrieval, to obtain a knowledge set.

[0089] Multi-modal knowledge retrieval may include retrieving multiple knowledges of different knowledge types.

[0090] Exemplarily, before performing operation S210 as shown in Figure 2 , the question-answering method may include an operation: performing multi-modal knowledge retrieval based on the question image in the question to obtain a knowledge set.

[0091] The retrieval can be performed in a combination of image search by image and image search by text to achieve multi-modal knowledge retrieval and obtain a multi-modal knowledge set.

[0092] For example, the image features of the question image can be extracted and respectively subjected to similarity matching with the image features of the knowledge images and the text features of the knowledge texts in the multi-modal knowledge base to obtain a set of matching results. Based on the set of matching results, the knowledge images and the knowledge texts that match the question image are determined from the multi-modal knowledge base to obtain a knowledge set.

[0093] Optionally, multi-modal knowledge retrieval can also be performed based on the question text to obtain a knowledge set.

[0094] Specifically, the text features of the question text can be extracted. The method of text search by image and text search by text using the text features from the multi-modal knowledge base is similar to the method of image search by image and image search by text using the image features, and will not be elaborated here.

[0095] Compared with the method of multi-modal knowledge retrieval based on the question text, the method of multi-modal knowledge retrieval based on the question image can, in the case where the effective information contained in the question text for answering the question is very little, such as when the question text includes "Please introduce the items in the image", retrieve relevant information using the question image, improve the reliability of the knowledge in the knowledge set, and then apply it to problems such as fine-grained recognition to improve the effect of answer generation.

[0096] Exemplarily, in the case where the question includes a question image and question text, obtaining the knowledge set may further adopt the following operations: Based on the image features of the question image and the text features of the question text, a fused feature is obtained. Based on the fused feature, multi-modal knowledge retrieval is performed to obtain the knowledge set.

[0097] The image features of the question image and the text features of the question text may be concatenated in the channel dimension to obtain the fused feature. However, it is not limited thereto. The attention mechanism or the gated unit may also be used to fuse the image features of the question image and the text features of the question text to obtain the fused feature. As long as the fused feature combines the image features of the question image and the text features of the question text.

[0098] The manner of using the fused feature to perform knowledge search from the multi-modal knowledge base is similar to the manner of using the image feature to perform knowledge search from the multi-modal knowledge base, and will not be elaborated herein.

[0099] Using the fused feature provided by the embodiments of the present disclosure for multi-modal knowledge retrieval can utilize the characteristics that the fused feature combines the image features of the question image and the text features of the question text, simplify the retrieval process, and improve the effectiveness and reliability of the retrieved knowledge.

[0100] The above introduces various implementation manners of how to retrieve the knowledge set. The following will illustrate retrieving the knowledge set from the multi-modal knowledge base using the question image.

[0101] According to the embodiments of the present disclosure, performing multi-modal knowledge retrieval based on the question image to obtain the knowledge set may include: Retrieving knowledge text and knowledge images respectively matching the question image from the multi-modal knowledge base. The multi-modal knowledge base includes at least one multi-modal knowledge unit, and the multi-modal knowledge unit includes knowledge text and knowledge images. In the case where it is determined that the knowledge text and knowledge images matching the question image belong to the same target multi-modal knowledge unit, based on the target multi-modal knowledge unit, the knowledge set is obtained.

[0102] It should be noted that the above only illustrates, by way of example, performing multi-modal retrieval from the multi-modal knowledge base using the question image to obtain the knowledge set. However, it is not limited thereto. The retrieval manner similar to the question image may also be adopted to perform multi-modal retrieval from the multi-modal knowledge base using the fused feature to obtain the knowledge set. This will not be elaborated herein.

[0103] The following will Figure 4 explain how to screen the retrieved knowledge to improve the reliability and effectiveness of the knowledge in the knowledge set.

[0104] Figure 4A flowchart of multimodal knowledge retrieval according to an embodiment of the present disclosure is schematically shown.

[0105] As Figure 4 shown, multimodal knowledge retrieval is performed in a multimodal knowledge base using a problem image such as Picture I. The multimodal knowledge base stores multiple multimodal knowledge units, and each multimodal knowledge unit includes a knowledge image such as an example diagram, knowledge text such as a category name and knowledge content. It is possible to use image search to determine the top 3 example diagrams that match Picture I, for example, with the highest similarity, including Example Diagram d, Example Diagram a, and Example Diagram e. Using image-to-text search, determine the top 3 category names that match Picture I, for example, with the highest similarity, including Category Name a, Category Name b, and Category Name c. Based on the dual image-text verification strategy, find the intersection of the multimodal knowledge units determined based on the category names and the multimodal knowledge units determined based on the example diagrams, and only retain the intersection to determine the target multimodal knowledge unit, which includes the multimodal knowledge unit with Category Name a, Example Diagram a, and Knowledge Content a.

[0106] If there is no intersection among the top 3 multimodal knowledge units with the highest similarity, the retrieved content is empty.

[0107] The multimodal knowledge base includes at least one multimodal knowledge unit, and each multimodal knowledge unit includes combined knowledge of knowledge text and knowledge image. It is possible to combine the feature that the knowledge in the multimodal knowledge base is stored in the form of unit combinations, and perform intersection screening on the retrieved knowledge, thereby improving the effectiveness and reliability of the knowledge in the knowledge set.

[0108] The above explains how to retrieve target multimodal knowledge units from the multimodal knowledge base. The following will explain how to construct a multimodal knowledge base.

[0109] According to an embodiment of the present disclosure, the multimodal knowledge units in the multimodal knowledge base include knowledge images.

[0110] Retrieval enhancement technology can retrieve the multimodal knowledge base through a large model, mine information related to the problem from a knowledge base constructed by humans, accurate, proprietary, and real-time updated, to enhance and supplement the knowledge reserve of the large model, and improve the reliability, timeliness, and professionalism of the answers generated by the large model.

[0111] The acquisition method of the knowledge image can include the following operations: Obtain multiple different images of the same target object. Cluster multiple image features of the multiple images to obtain a cluster center. Based on the image with the closest image feature distance to the cluster center, determine the knowledge image.

[0112] The category of the target object is not limited. For example, it can include animals, plants, vehicles, commodities, anime characters, etc. Multiple images of the same target object with different angles, different sizes, different colors, different backgrounds, or different actions can be obtained. The image features of the multiple images are respectively extracted to obtain multiple image features. The multiple image features are clustered to obtain a cluster center. The cluster center is used to represent the most representative central image feature. The image among the multiple image features that is closest to the cluster center, for example, has the highest feature matching degree, is used as the knowledge image, such as the example image.

[0113] The clustering method is not limited. For example, the K-means clustering algorithm. As long as it can cluster the image features.

[0114] The knowledge image obtained through the above method can include the representative features of the target object. Using the knowledge image for retrieval can improve the retrieval accuracy and at the same time improve the generation efficiency and generation effect of the subsequent large model using the knowledge image to obtain answers.

[0115] According to an embodiment of the present disclosure, the multimodal knowledge unit may further include knowledge text.

[0116] The knowledge text can be obtained through the following method.

[0117] For example, based on the target object in the knowledge image, content of a predetermined type is obtained to obtain the knowledge text.

[0118] The content of the predetermined type can be content divided according to the content form. For example, it can include content types such as text, images, tables, etc. However, it is not limited thereto. It can also be a type divided according to the amount of knowledge of the content. It can also be divided according to different categories of the content. For example, object category knowledge text, object attribute knowledge text, etc.

[0119] The content of the predetermined type is pre-divided for the knowledge text. By setting the predetermined type, a judgment criterion can be set for different knowledge texts, and thus the effectiveness and richness of the collected knowledge texts can be ensured.

[0120] According to an embodiment of the present disclosure, in the case where the knowledge text includes object category knowledge text, the predetermined type can be determined as the content of multiple object fields and object names at the same level. Multiple object fields and object names at different levels are obtained to obtain the knowledge text.

[0121] For example, the knowledge text includes object category knowledge text. Specifically, it includes the first object field, the second object field, and the object name. The levels between the first object field and the second object field are different.

[0122] Taking the tricolored heron as an example, the tricolored heron is the object name. The first object area can include animals. The second object area can include birds. Multiple different text contents can be distinguished by dashes. The object name can also include aliases. Multiple names are separated by slashes.

[0123] Forming a knowledge text through multiple object areas and object names at different levels can make the category knowledge of each target object be reflected completely and effectively, avoiding the problem of inaccurate retrieval caused by incomplete reflection of category knowledge.

[0124] According to an embodiment of the present disclosure, when the knowledge text includes an object attribute knowledge text, text content of a predetermined type can be obtained based on the target object. The text content is subjected to generative processing to obtain the knowledge text.

[0125] The object attribute knowledge text can include knowledge related to object categories, but the text structures are different. In addition, the object attribute knowledge text can also include other knowledge related to the target object. For example, a detailed introduction knowledge text. As long as it is knowledge related to the target object.

[0126] There may be problems such as redundancy and repetition among multiple text contents collected from different channels. The text content can be subjected to generative processing, for example, input into a large model, so as to use the large model to perform operations such as deduplication, logical sorting, and structural information reconstruction on the multiple text contents to obtain a knowledge text that conforms to predetermined rules.

[0127] Using the generative processing method to process the knowledge content can improve the processing efficiency, enhance the recombination effect of the knowledge text, and improve the effectiveness of the knowledge text as reference knowledge for answering questions.

[0128] According to another embodiment of the present disclosure, when the knowledge text includes an object attribute knowledge text, initial content of a predetermined type can also be determined based on the target object. When it is determined that the initial content includes non-text content, the initial content is converted into text content to obtain the knowledge text.

[0129] The type of the knowledge text is text. When the obtained initial content of the predetermined type does not belong to text, for example, in the case of an image or a table, the initial content is converted to obtain text content. Thereby improving the normativity of the knowledge text.

[0130] According to an embodiment of the present disclosure, the predetermined type can be randomly specified. However, it is not limited to this. It can also be determined based on a historical question set.

[0131] For example, taking vehicle A as the target object. The historical question set can include the price of vehicle A, the vehicle performance of vehicle A, the fuel consumption per kilometer of vehicle A, the depreciation of vehicle A, and the insurance of vehicle A.

[0132] Based on the above historical problem set, it can be determined that the predetermined types include: price, performance, fuel consumption, depreciation, insurance, etc.

[0133] By summarizing the problems in the historical problem set, determine the relevant knowledge types that are popular or highly concerned, and determine this as the predetermined type. Thereby improving the effectiveness and rationality of using the text knowledge of this predetermined type as reference knowledge for answering questions, and further improving the intelligence of the question-and-answer interaction.

[0134] According to an embodiment of the present disclosure, before performing the operation S210 as Figure 2 shown, the question-and-answer method based on artificial intelligence can also operate: generating a multi-modal knowledge unit based on a knowledge image and multiple knowledge texts with different amounts of knowledge.

[0135] Figure 5 Schematically shows a schematic diagram of a multi-modal knowledge unit according to an embodiment of the present disclosure.

[0136] As Figure 5 shown, the multi-modal knowledge unit may include a knowledge image 520, such as a flower as the target object. The multiple knowledge texts may include the object category knowledge text of the flower and the object attribute knowledge text of the flower.

[0137] As Figure 5 shown, the object category knowledge text 510 includes relevant information such as plants, flowers, and flower names such as sunflowers.

[0138] As Figure 5 shown, the object attribute knowledge text 530 includes relevant knowledge such as the growth environment, common types, and fruits of this flower.

[0139] It can be known by comparing the number of words, types of content, etc. that the amount of knowledge in the object category knowledge text is less than the amount of knowledge in the object attribute knowledge text.

[0140] The above gives an example of the multi-modal knowledge unit. The following will give an example of the question-and-answer method in combination with the multi-modal knowledge unit.

[0141] Figure 6 Schematically shows a flowchart of the question-and-answer method based on artificial intelligence according to an embodiment of the present disclosure.

[0142] As Figure 6As shown in the figure, the large model is called to perform a progressive reflective thinking chain operation on the problem image I and the problem text Q to obtain a first-level reflection result 610. It is determined whether the first-level reflection result 610 represents that the amount of knowledge of the model's own model knowledge reaches the knowledge amount threshold. That is, it is determined whether external knowledge is needed. In the case where the first-level reflection result 610 represents that the amount of knowledge of the model knowledge does not reach the knowledge amount threshold, indicating that external knowledge is needed, the large model is called to perform a progressive reflective thinking chain operation on the problem image I, the problem text Q, the object category knowledge text such as category name a, and the knowledge image such as example diagram a to obtain a second-level reflection result 620. In the case where external knowledge is not needed, based on the problem and the model knowledge, the answer 630 is obtained. It is determined whether the second-level reflection result 620 represents that the amount of knowledge of category name a and example diagram a reaches the knowledge amount threshold. That is, it is determined whether external knowledge is needed. In the case where the second-level reflection result 620 represents that the amount of knowledge of category name a and example diagram a does not reach the knowledge amount threshold, indicating that external knowledge is needed, based on the problem image I, the problem text Q, category name a, example diagram a, and the object attribute knowledge text such as knowledge content a, the answer 630 is generated. In the case where external knowledge is not needed, based on the problem image I, the problem text Q, category name a, and example diagram a, the answer 630 is obtained.

[0143] In summary, the question-answering method provided by the embodiments of the present disclosure establishes a multi-modal knowledge base, in which the object category knowledge text, the knowledge image, and the object attribute knowledge text are matched one by one to form a number of multi-modal knowledge units. In addition, a retrieval scheme for image search in image and image search in text is designed, and the screening method of taking the intersection of the image search in image result and the image search in text result is used to further improve the effectiveness of the knowledge set and the strong relevance to the question. In the answer generation stage, a progressive reflective thinking chain is designed, enabling the large model to more efficiently screen, integrate, and apply knowledge, avoiding redundancy and waste of knowledge, significantly improving the overall performance, and finally the large model combines the knowledge set to generate the final answer. Provide fine-grained and comprehensive long-tail knowledge for the large model, effectively improving the authority, reliability, and professionalism of the content generated by the large model.

[0144] Figure 7 A block diagram of an artificial intelligence-based question-answering device according to an embodiment of the present disclosure is schematically shown.

[0145] As Figure 7 shown, the artificial intelligence-based question-answering device may include an acquisition module 710 and a call module 720.

[0146] The acquisition module 710 is used to acquire the question and the knowledge set retrieved based on the question.

[0147] The calling module 720 is used to call the large model to perform the following operations based on the question and the knowledge set. Among them, the calling module 720 includes a reflection sub-module 721, an answer sub-module 722, and an update sub-module 723.

[0148] The reflection sub-module 721 is used to reflect on the question and the current knowledge subset to obtain a reflection result, and the knowledge set includes the knowledge in the current knowledge subset.

[0149] The answer sub-module 722 is used to obtain an answer based on the question and the current knowledge subset when the reflection result represents that the current knowledge subset reaches the knowledge amount for answering the question.

[0150] The update sub-module 723 is used to update the current knowledge subset based on the remaining knowledge subset when the reflection result represents that the current knowledge subset does not reach the knowledge amount for answering the question. The remaining knowledge subset includes the knowledge in the knowledge set except the current knowledge subset.

[0151] According to an embodiment of the present disclosure, the update sub-module includes: a knowledge amount determination unit and an addition unit.

[0152] The knowledge amount determination unit is used to determine the target remaining knowledge with the least knowledge amount from the remaining knowledge subset.

[0153] The first addition unit is used to add the target remaining knowledge to the current knowledge subset.

[0154] According to an embodiment of the present disclosure, the update sub-module includes: a knowledge type determination unit and a second addition unit.

[0155] The knowledge type determination unit is used to determine the target remaining knowledge from the remaining knowledge subset based on the knowledge type of the knowledge for answering the question.

[0156] The second addition unit is used to add the target remaining knowledge to the current knowledge subset.

[0157] According to an embodiment of the present disclosure, the question-answering device further includes: an identification module.

[0158] The identification module is used to identify the question and determine the knowledge type of the knowledge for answering the question.

[0159] According to an embodiment of the present disclosure, the question includes a question image.

[0160] The question-answering device further includes: a first retrieval module.

[0161] The first retrieval module is used to perform multi-modal knowledge retrieval based on the question image to obtain a knowledge set.

[0162] According to an embodiment of the present disclosure, the first retrieval module includes: a first retrieval sub-module and a second retrieval sub-module.

[0163] The first retrieval sub-module is configured to retrieve, from a multi-modal knowledge base, a knowledge text and a knowledge image that respectively match the problem image. The multi-modal knowledge base includes at least one multi-modal knowledge unit, and the multi-modal knowledge unit includes a knowledge text and a knowledge image.

[0164] The second retrieval sub-module is configured to obtain a knowledge set based on the target multi-modal knowledge unit when it is determined that the knowledge text and the knowledge image that match the problem image belong to the same target multi-modal knowledge unit.

[0165] According to an embodiment of the present disclosure, the question-and-answer device further includes: an image acquisition module, a clustering module, and an image determination module.

[0166] The image acquisition module is configured to acquire multiple different images of the same target object.

[0167] The clustering module is configured to cluster multiple image features of the multiple images to obtain a clustering center.

[0168] The image determination module is configured to determine the knowledge image based on the image whose image features are closest to the clustering center.

[0169] According to an embodiment of the present disclosure, the question-and-answer device further includes: a modal unit generation module.

[0170] The modal unit generation module is configured to generate a multi-modal knowledge unit based on the knowledge image and multiple knowledge texts with different knowledge amounts.

[0171] According to an embodiment of the present disclosure, the question-and-answer device further includes: a content acquisition module.

[0172] The content acquisition module is configured to acquire content of a predetermined type based on the target object in the knowledge image to obtain a knowledge text.

[0173] According to an embodiment of the present disclosure, the content acquisition module includes: a first acquisition sub-module.

[0174] The first acquisition sub-module is configured to, when the knowledge text includes an object category knowledge text, acquire multiple object domains and object names at different levels of the target object to obtain a knowledge text.

[0175] According to an embodiment of the present disclosure, the question-and-answer device further includes: a history acquisition module and a type determination module.

[0176] The history acquisition module is configured to acquire a set of historical questions.

[0177] The type determination module is configured to determine a predetermined type based on the set of historical questions.

[0178] According to an embodiment of the present disclosure, the content acquisition module includes: a second acquisition sub-module and a generation processing sub-module.

[0179] The second acquisition sub-module is configured to acquire text content of a predetermined type based on a target object.

[0180] The generation processing sub-module is configured to perform generative processing on the text content to obtain knowledge text.

[0181] According to an embodiment of the present disclosure, the content acquisition module includes: a third acquisition sub-module and a conversion sub-module.

[0182] The third acquisition sub-module is configured to determine initial content of a predetermined type based on a target object.

[0183] The conversion sub-module is configured to convert the initial content into text content to obtain knowledge text when it is determined that the initial content includes non-text content.

[0184] According to an embodiment of the present disclosure, the question includes a question image and question text.

[0185] The question-answering device further includes: a fusion module and a second retrieval module.

[0186] The fusion module is configured to obtain a fusion feature based on the image feature of the question image and the text feature of the question text.

[0187] The second retrieval module is configured to perform multimodal knowledge retrieval based on the fusion feature to obtain a knowledge set.

[0188] According to an embodiment of the present disclosure, the question-answering device further includes: a primary reflection module and a third retrieval module.

[0189] The primary reflection module is configured to call a large model to reflect on the question to obtain a primary reflection result.

[0190] The third retrieval module is configured to perform retrieval based on the question to obtain a knowledge set when the primary reflection result indicates that the model knowledge of the large model does not reach the amount of knowledge for answering the question.

[0191] Figure 8 A structural block diagram of an agent of artificial intelligence according to an embodiment of the present disclosure is schematically shown.

[0192] In an embodiment of the present disclosure, as Figure 8 shown, the agent 800 may include an input module 810, a processing module 820, and an output module 830.

[0193] The input module 810 is configured to receive input information.

[0194] A processing module 820 is configured to determine a target task based on the input information received by the input module, determine a large model based on the target task, and execute the artificial intelligence-based question-and-answer method provided in the embodiments of the present disclosure by invoking the large model.

[0195] An output module 830 is configured to output the output information obtained by the processing module.

[0196] According to an embodiment of the present disclosure, the input module 810 is responsible for receiving or sensing information such as queries, requests, instructions, signals, or data from the outside world (such as a user or an external environment), and converting it into a format that the agent 800 can understand and process. The input module 810 is the primary link for the agent 800 to interact with the outside world, enabling the agent 800 to efficiently and accurately obtain necessary "sensory" information from the outside world and respond to this information.

[0197] In an example, the input module 810 can input the questions described above.

[0198] In an example, the processing module 820 is the core support for the agent 800's ability to handle complex tasks. The processing module 820 can execute the artificial intelligence-based question-and-answer method described above.

[0199] In an example, the performance of the processing module 820 can be closely related to the large model on which the agent 800 is based. To fully utilize the capabilities of the large model, the internal structure of the processing module 820 can be designed to be highly configurable and extensible to handle various different types of tasks and requirements in real-world scenarios.

[0200] In an example, after the agent 800 obtains a question, the processing module 820 can use the large model to process the question and the knowledge set, obtain an answer, and pass it to the output module 830.

[0201] In an example, the output module 830 can output the answer described above.

[0202] The agent 800 according to the embodiments of the present disclosure can simply and effectively improve the degree of intelligence, and improve flexibility and versatility.

[0203] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0204] According to an embodiment of the present disclosure, an electronic device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method as described above.

[0205] According to an embodiment of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to execute the method as described above.

[0206] According to an embodiment of the present disclosure, a computer program product includes a computer program that, when executed by a processor, implements the method as described above.

[0207] Figure 9 A schematic block diagram of an example electronic device 900 that can be used to implement the embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as, for example, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, for example, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0208] As Figure 9 shown, the device 900 includes a computing unit 901 that can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 902 or a computer program loaded from a storage unit 908 into a random access memory (RAM) 903. In the RAM 903, various programs and data required for the operation of the device 900 can also be stored. The computing unit 901, the ROM 902, and the RAM 903 are connected to each other via a bus 904. An input / output (I / O) interface 905 is also connected to the bus 904.

[0209] A plurality of components in the device 900 are connected to the input / output (I / O) interface 905, including: an input unit 906, such as a keyboard, a mouse, etc.; an output unit 907, such as various types of displays, speakers, etc.; a storage unit 908, such as a magnetic disk, an optical disk, etc.; and a communication unit 909, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 909 allows the device 900 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0210] The computing unit 901 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 901 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 901 executes the various methods and processes described above, such as the artificial intelligence-based question-answering method. For example, in some embodiments, the artificial intelligence-based question-answering method can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 908. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 900 via the ROM 902 and / or the communication unit 909. When the computer program is loaded into the RAM 903 and executed by the computing unit 901, one or more steps of the artificial intelligence-based question-answering method described above can be executed. Alternatively, in other embodiments, the computing unit 901 can be configured to execute the artificial intelligence-based question-answering method in any other suitable manner (e.g., by means of firmware).

[0211] Various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGA), application-specific integrated circuits (ASIC), application-specific standard products (ASSP), systems-on-a-chip (SOC), complex programmable logic devices (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a dedicated or general-purpose programmable processor, receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting the data and instructions to the storage system, the at least one input device, and the at least one output device.

[0212] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to the processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The program code can be executed entirely on the machine, partially on the machine, as an independent software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0213] In the context of this disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0214] In order to provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0215] The systems and techniques described herein can be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.

[0216] A computer system can include a client and a server. The client and the server are generally far apart from each other and typically interact through a communication network. The client-server relationship is generated by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, or a server of a distributed system, or a server incorporating a blockchain.

[0217] It should be understood that the various forms of processes shown above can be used, with steps reordered, added or deleted. For example, the steps described in the present disclosure can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solutions disclosed in the present disclosure can be achieved, and no limitations are imposed herein.

[0218] The above specific embodiments do not constitute a limitation on the protection scope of the present disclosure. Those skilled in the art should understand that various modifications, combinations, sub - combinations and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions and improvements made within the spirit and principle of the present disclosure shall be included within the protection scope of the present disclosure.

Claims

1. A question-answering method based on artificial intelligence, comprising: Obtaining a question and a knowledge set retrieved based on the question; The large model is called to perform the following operations based on the question and the knowledge set: Reflecting on the problem and the current knowledge subset to obtain a reflection result, wherein the knowledge set includes the knowledge in the current knowledge subset; When the reflection result indicates that the current knowledge subset reaches the knowledge amount for answering the question, obtaining an answer based on the question and the current knowledge subset; as well as In the case where the reflection result indicates that the current knowledge subset does not reach the knowledge amount for answering the question, the current knowledge subset is updated based on the remaining knowledge subset, where the remaining knowledge subset includes the knowledge in the knowledge set except the current knowledge subset.

2. The method according to claim 1, wherein: The updating of the current knowledge subset based on the remaining knowledge subset comprises: Determining target remaining knowledge with the least amount of knowledge from the remaining knowledge subset; and The target remaining knowledge is added to the current knowledge subset.

3. The method according to claim 1, wherein: The updating of the current knowledge subset based on the remaining knowledge subset comprises: determining target remaining knowledge from the remaining knowledge subset based on the knowledge type of the knowledge used to answer the question; and The target remaining knowledge is added to the current knowledge subset.

4. The method according to claim 3, further comprising: The question is identified, and the type of knowledge used to answer the question is determined.

5. The method according to any one of claims 1 to 4, wherein: The question includes a question image; The method further comprises: Multimodal knowledge retrieval is performed based on the problem image to obtain the knowledge set.

6. The method according to claim 5, wherein: The multimodal knowledge retrieval based on the problem image to obtain the knowledge set includes: Retrieving knowledge text and knowledge image respectively matching the question image from a multimodal knowledge base, the multimodal knowledge base comprising at least one multimodal knowledge unit, the multimodal knowledge unit comprising knowledge text and knowledge image; and When it is determined that the knowledge text and the knowledge image matching the question image belong to the same target multimodal knowledge unit, the knowledge set is obtained based on the target multimodal knowledge unit.

7. The method according to claim 6, further comprising: Acquire multiple different images of the same target object; Clustering multiple image features of multiple images to obtain cluster centers; as well as The knowledge image is determined based on the image whose image feature is closest to the cluster center.

8. The method according to claim 6 or 7, further comprising: The multimodal knowledge unit is generated based on the knowledge image and a plurality of the knowledge texts with different amounts of knowledge.

9. The method according to any one of claims 6 to 8, further comprising: Based on the target object in the knowledge image, a predetermined type of content is acquired to obtain the knowledge text.

10. The method according to claim 9, wherein: The step of acquiring a predetermined type of content based on the target object in the knowledge image to obtain the knowledge text includes: In the case where the knowledge text includes an object category knowledge text, a plurality of object fields and object names at different levels of the target object are acquired to obtain the knowledge text.

11. The method according to claim 9 or 10, further comprising: Get a collection of historical questions; as well as Based on a set of historical questions, the predetermined type is determined.

12. The method according to any one of claims 9 to 11, wherein: The step of acquiring a predetermined type of content based on the target object in the knowledge image to obtain the knowledge text includes: Based on the target object, acquiring the text content of the predetermined type; and Generative processing is performed on the text content to obtain the knowledge text.

13. The method according to any one of claims 9 to 11, wherein: The step of acquiring a predetermined type of content based on the target object in the knowledge image to obtain the knowledge text includes: Based on the target object, determining the initial content of the predetermined type; and In the case where it is determined that the initial content includes non-text content, the initial content is converted into text content to obtain the knowledge text.

14. The method according to any one of claims 1 to 4, wherein: The question includes a question image and a question text; The method further comprises: Obtaining fusion features based on the image features of the question image and the text features of the question text; and Multimodal knowledge retrieval is performed based on the fusion features to obtain the knowledge set.

15. The method according to any one of claims 1 to 14, further comprising: Calling the large model to reflect on the problem and obtain a primary reflection result; as well as In the case where the primary reflection result indicates that the model knowledge of the large model does not reach the amount of knowledge required to answer the question, a search is performed based on the question to obtain the knowledge set.

16. A question-answering device based on artificial intelligence, comprising: An acquisition module, used for acquiring a question and a knowledge set retrieved based on the question; The calling module is used to call the large model to perform the following operations based on the problem and the knowledge set: Reflecting on the problem and the current knowledge subset to obtain a reflection result, wherein the knowledge set includes the knowledge in the current knowledge subset; When the reflection result indicates that the current knowledge subset reaches the knowledge amount for answering the question, obtaining an answer based on the question and the current knowledge subset; as well as In the case where the reflection result indicates that the current knowledge subset does not reach the knowledge amount for answering the question, the current knowledge subset is updated based on the remaining knowledge subset, where the remaining knowledge subset includes the knowledge in the knowledge set except the current knowledge subset.

17. An artificial intelligence-based intelligent agent, wherein: The agent is configured to perform the method according to any one of claims 1 to 15.

18. An electronic device, comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 15.

19. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1-15.

20. A computer program product comprising a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 15.