Visual question and answer method and device, electronic equipment, storage medium and computer program product

By generating and extending programs in the visual question-and-answer system to record and interpret the reasoning process, the problem of lack of transparency in the existing system is solved, and higher user trust and the convenience of technology promotion and application are achieved.

CN120146205AActive Publication Date: 2025-06-13INST OF AUTOMATION CHINESE ACAD OF SCI

Patent Information

Application Number
CN202510621772.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-14
Publication Date
2025-06-13
Estimated Expiration
2045-05-14

AI Technical Summary

Technical Problem

The existing visual question and answer system can only provide the final answer, and the lack of explanation of the internal reasoning process leads to lower transparency of the model and reduced user trust, which is not conducive to the promotion and application of technology.

Method used

Generate an extension by obtaining the target image and problem, generating the initial program, and adding object code to it that records the program execution process. Enter the target image into the extension program to obtain predicted answers, execution process information and screenshot images, and generate a multimodal form of explanation for the predicted answers based on this information.

Benefits of technology

The explanation of the reasoning process of predictive answers is added, allowing users to intuitively understand the correspondence between image features and semantic reasoning, improve the transparency of reasoning and decision-making credibility, and overcome the black box limitations of traditional visual question-and-answer models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120146205A_ABST
    Figure CN120146205A_ABST
Patent Text Reader

Abstract

The invention relates to a visual question and answer method and device, electronic equipment, a storage medium and a computer program product. The method comprises the steps of obtaining a target image and a target question for the target image; generating an initial program based on the target problem; adding a target code for recording a program execution process to the initial program; inputting the target image into an extension program, and obtaining a prediction answer for the target question, execution process information of the extension program and a screenshot image; based on the execution process information and the screenshot image, an interpretation in a multi-modal form for the predicted answer is generated. Therefore, the decision basis picture and the semantic association analysis can be synchronously generated while the prediction answer is output, that is, the explanation of the reasoning process of the prediction answer can be increased and output, so that a user can intuitively understand the corresponding relationship between the image features and semantic reasoning, the reasoning transparency and the decision credibility can be improved, and the user experience can be improved. Therefore, popularization and application of the visual question-answering technology are facilitated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technologies, and more particularly, to a visual question answering method, apparatus, electronic device, storage medium, and computer program product. Background Art

[0002] Visual Question Answering (VQA) combines computer vision technology and natural language processing technology, and can provide users with a more intelligent image understanding and interaction experience. Specifically, it enables a computer to "see" an image and answer questions about the image, that is, given an image and a question about the image, a visual question answering software can output an answer to the given question based on the input image. This requires the visual question answering software to have the ability of multimodal understanding, and currently most mainstream methods are implemented based on multimodal pre-training. The VQA technology can be applied to multiple fields, including but not limited to: medical, education, monitoring, entertainment and other fields.

[0003] However, the visual question answering systems in related technologies usually can only provide the final answer, and lack an explanation of the internal reasoning process, that is, the model transparency is relatively low, which will lead to a decrease in user trust and is not conducive to the popularization and application of the visual question answering technology. Summary of the Invention

[0004] The present disclosure provides a visual question answering method, apparatus, electronic device, storage medium, and computer program product to at least solve the problem in the above related technologies that the visual question answering system usually can only provide the final answer and lacks an explanation of the internal reasoning process, resulting in relatively low model transparency.

[0005] According to a first aspect of an embodiment of the present disclosure, a visual question answering method is provided, including: obtaining a target image and a target question for the target image; generating an initial program based on the target question, where the initial program is used to perform reasoning to obtain an answer to the target question; adding a target code for recording the program execution process to the initial program to obtain an extended program; inputting the target image into the extended program to obtain a predicted answer to the target question, execution process information of the extended program, and a screenshot image, where the screenshot image is a screenshot image associated with a target object pointed to by the target question and intercepted from the target image; generating a multimodal form of explanation for the predicted answer based on the execution process information and the screenshot image, where the multimodal form at least fuses a text form and an image form.

[0006] Optionally, generating an initial program based on the target problem includes: obtaining a preset program prompt, where the preset program prompt includes a structural thinking chain construction method and program prompt examples. The structural thinking chain construction method is used to indicate how to construct a corresponding structural thinking chain based on the problem, and the program prompt examples are used to indicate how to generate a corresponding program based on the structural thinking chain; constructing a target structural thinking chain corresponding to the target problem based on the structural thinking chain construction method; and generating the initial program corresponding to the target structural thinking chain based on the program prompt examples.

[0007] Optionally, the preset program prompt further includes an application programming interface description, where the application programming interface description is used to indicate how to use the program to call open-world tools; and obtaining the predicted answer for the target problem, the execution process information of the extended program, and the screenshot image by inputting the target image into the extended program includes: by executing the extended program, calling the open-world tools based on the application programming interface description to process the target image, and obtaining the predicted answer, the execution process information, and the screenshot image.

[0008] Optionally, the open-world tools include the following items: an image object detector, where the image object detector is used to detect objects of a specified type in the image; an image cropping tool, where the image cropping tool is used to crop a part or all of the regions included in the image; and a customized Python function, where the customized Python function is used to perform at least one of specific types of calculation tasks, format conversion tasks, and logical judgment tasks to obtain the predicted answer and the execution process information.

[0009] Optionally, generating a multimodal form of explanation for the predicted answer based on the execution process information and the screenshot image includes: obtaining a preset explanation prompt example, where the preset explanation prompt example is used to indicate how to generate a corresponding natural language explanation based on the program execution process; generating a natural language target text explanation for the predicted answer based on the preset explanation prompt example and the execution process information; and generating a multimodal form of explanation for the predicted answer based on the target text explanation and the screenshot image.

[0010] Optionally, generating a natural language target text explanation for the predicted answer based on the preset explanation prompt example and the execution process information includes: inputting the preset explanation prompt example and the execution process information into a large language model to obtain a natural language target text explanation for the predicted answer.

[0011] According to a second aspect of the embodiments of the present disclosure, a visual question answering device is provided, including: an image and question acquisition module configured to acquire a target image and a target question for the target image; an initial program generation module configured to generate an initial program based on the target question, where the initial program is used to perform reasoning to obtain an answer to the target question; a program extension module configured to add target code for recording the program execution process to the initial program to obtain an extended program; an answer prediction module configured to input the target image into the extended program to obtain a predicted answer for the target question, execution process information of the extended program, and a screenshot image, where the screenshot image is a screenshot image associated with the target object pointed to by the target question and intercepted from the target image; a multimodal explanation generation module configured to generate a multimodal form of explanation for the predicted answer based on the execution process information and the screenshot image, where the multimodal form at least fuses a text form and an image form.

[0012] Optionally, the initial program generation module is configured to: acquire a preset program prompt, where the preset program prompt includes a structural thinking chain construction method and a program prompt example, the structural thinking chain construction method is used to indicate how to construct a corresponding structural thinking chain based on a question, and the program prompt example is used to indicate how to generate a corresponding program based on the structural thinking chain; construct a target structural thinking chain corresponding to the target question based on the structural thinking chain construction method; generate the initial program corresponding to the target structural thinking chain based on the program prompt example.

[0013] Optionally, the preset program prompt further includes an application programming interface description, where the application programming interface description is used to indicate how to use a program to call open-world tools; the answer prediction module is configured to: by executing the extended program, call the open-world tools to process the target image based on the application programming interface description to obtain the predicted answer, the execution process information, and the screenshot image.

[0014] Optionally, the open-world tools include the following items: an image object detector, where the image object detector is used to detect objects of a specified type in an image; an image cropping tool, where the image cropping tool is used to crop a part or all of the regions included in the image; a customized Python function, where the customized Python function is used to perform at least one of specific types of calculation tasks, format conversion tasks, and logical judgment tasks to obtain the predicted answer and the execution process information.

[0015] Optionally, the multimodal explanation generation module is configured to: obtain a preset explanation prompt example, where the preset explanation prompt example is used to indicate how to generate a corresponding natural language explanation based on the program execution process; generate a target text explanation in natural language for the predicted answer based on the preset explanation prompt example and the execution process information; and generate an explanation in multimodal form for the predicted answer based on the target text explanation and the screenshot image.

[0016] Optionally, the multimodal explanation generation module is configured to: input the preset explanation prompt example and the execution process information into a large language model to obtain a target text explanation in natural language for the predicted answer.

[0017] According to a third aspect of the embodiments of the present disclosure, there is provided an electronic device, including: a processor; a memory for storing executable instructions of the processor; wherein, the processor is configured to execute the instructions to implement the visual question answering method according to the present disclosure.

[0018] According to a fourth aspect of the embodiments of the present disclosure, there is provided a computer-readable storage medium, when the instructions in the computer-readable storage medium are executed by a processor of an electronic device, enabling the electronic device to execute the visual question answering method according to the present disclosure.

[0019] According to a fifth aspect of the embodiments of the present disclosure, there is provided a computer program product, including a computer program, where the computer program implements the visual question answering method according to the present disclosure when executed by a processor.

[0020] The technical solutions provided by the embodiments of the present disclosure at least bring the following beneficial effects: In the present disclosure, a decision basis picture and semantic association analysis can be synchronously generated while outputting the predicted answer, that is, an explanation of the reasoning process for the predicted answer can be added, enabling the user to intuitively understand the correspondence between image features and semantic reasoning, thereby improving the reasoning transparency and decision credibility, that is, overcoming the defects of low model transparency and low user trust caused by the black box limitation of traditional visual question answering models, which is conducive to the popularization and application of visual question answering technology. And, by outputting an explanation in multimodal form for the predicted answer, the display of the explanation is clearer and more intuitive, reducing the understanding difficulty and making it easy for the user to quickly and accurately understand the entire reasoning process.

[0021] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] The accompanying drawings here are incorporated into the specification and form a part of this specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure, and do not constitute an improper limitation of the present disclosure.

[0023] Figure 1 is a flowchart showing a visual question - answering method according to an exemplary embodiment of the present disclosure; Figure 2 is an example diagram showing a target image and a question asked about the target image according to an exemplary embodiment of the present disclosure; Figure 3 is an example diagram showing a preset program prompt according to an exemplary embodiment of the present disclosure; Figure 4 is an example diagram showing an example of a preset explanation prompt according to an exemplary embodiment of the present disclosure; Figure 5 is an example diagram showing a specific implementation process of a visual question - answering method according to an exemplary embodiment of the present disclosure; Figure 6 is a block diagram showing a visual question - answering device according to an exemplary embodiment of the present disclosure; Figure 7 is a block diagram showing an electronic device according to an exemplary embodiment of the present disclosure. Detailed implementation manners

[0024] In order to enable those of ordinary skill in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings.

[0025] It should be noted that the terms "first", "second", etc. in the specification and claims of the present disclosure and the above - mentioned accompanying drawings are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such used data can be interchanged under appropriate circumstances so that the embodiments of the present disclosure described here can be implemented in an order different from those illustrated or described here. The implementation manners described in the following embodiments do not represent all implementation manners consistent with the present disclosure. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.

[0026] It should be noted here that "at least one of several items" as used in this disclosure means that it includes three parallel cases: "any one of the several items", "any combination of multiple items among the several items", and "all of the several items". For example, "including at least one of A and B" includes the following three parallel cases: (1) including A; (2) including B; (3) including both A and B. Another example, "performing at least one of Step 1 and Step 2" means the following three parallel cases: (1) performing Step 1; (2) performing Step 2; (3) performing both Step 1 and Step 2.

[0027] As mentioned above, the visual question answering systems in the related art usually can only provide the final answer, lacking an explanation of the internal reasoning process, that is, the model transparency is relatively low, which will lead to a decrease in user trust. Moreover, the current visual question answering systems also have deficiencies in example dependence, the flexibility of code generation, and the integration of external tools. Especially in the scenario of relying on only a single example (one-shot), logical redundancy and unstable execution are likely to occur in code generation.

[0028] To solve the above problems existing in the related art, the visual question answering method, device, electronic device, storage medium, and computer program product provided by this disclosure can synchronously generate a decision basis picture and semantic association analysis while outputting a predicted answer, that is, it can add an explanation of the reasoning process for the predicted answer, enabling users to intuitively understand the correspondence between image features and semantic reasoning, and further improving the reasoning transparency and decision credibility. That is, it can overcome the defects of low model transparency and low user trust caused by the black box limitation of traditional visual question answering models, thus facilitating the popularization and application of visual question answering technology. Moreover, by outputting an explanation in a multimodal form for the predicted answer, the display of the explanation is clearer and more intuitive, reducing the understanding difficulty and making it easy for users to quickly and accurately understand the entire reasoning process.

[0029] Figure 1 It is a flowchart showing a visual question answering method according to an exemplary embodiment of the present disclosure.

[0030] Referring to Figure 1 , in step 101, a target image and a target question for the target image can be obtained. The target image and the target question can form multimodal question data, and its content and format can vary flexibly according to the actual application scenario.

[0031] Figure 2 It is an example diagram showing a target image and a question for the target image according to an exemplary embodiment of the present disclosure. Referring to Figure 2, the target image can be a field picture, which contains a total of 4 puppies. And the question data related to the target image can be any question about the image content. Exemplarily, in Figure 2 , the question data can be "What color is the dog on the far right?". And this question data can be either automatically generated by the system or manually input. The present disclosure does not make specific limitations on the content, format, and generation method of the question data.

[0032] It should be noted that in the present disclosure, simply installing visual question answering software on the terminal can execute the visual question answering method according to the present disclosure. Exemplarily, the terminal in the present disclosure can be, but is not limited to: mobile phones, tablets, laptops, wearable devices, etc.

[0033] In step 102, an initial program can be generated based on the target question, where the initial program is used to perform reasoning to obtain the answer to the target question.

[0034] According to an exemplary embodiment of the present disclosure, a preset program prompt can be obtained, where the preset program prompt can include a structural thinking chain construction method and program prompt examples. The "structural thinking chain construction method" can be used to indicate how to construct a corresponding structural thinking chain based on the question; the "program prompt examples" can be used to indicate how to generate a corresponding program based on the structural thinking chain, that is, the program prompt examples can show the correspondence between the question and the program. In addition, the preset program prompt in the present disclosure can be a single-sample program prompt.

[0035] Then, based on the structural thinking chain construction method, the target structural thinking chain corresponding to the target question can be constructed. Next, based on the program prompt examples, the initial program corresponding to the target structural thinking chain can be generated, that is, a structured initial logical program can be generated.

[0036] In this way, by setting the preset program prompt, the model used to generate the initial program can quickly learn the format of the program to be generated, and thus can ensure that the generated initial program meets the format and logical requirements. At the same time, it can also improve the program generation efficiency. Further, by adopting the learning method of single-sample program prompts, it is also possible to avoid over-reliance on a large amount of training data, thereby improving the generalization ability of the model used to generate the initial program.

[0037] Further, in the present disclosure, a preset program prompt and the target question can also be input into a large language model, and then the large language model generates and outputs an initial program. Exemplarily, the large language model can be, but is not limited to: GPT-4o. The initial program may consist of one or more program modules, and each module can correspond to processing a part of the logic in the target question. Specifically, the large language model can decompose the target question into several program steps, and these program steps can then be executed together with the target image, and then an answer to the question can be generated.

[0038] Figure 3 FIG. is an exemplary diagram showing a preset program prompt according to an exemplary embodiment of the present disclosure. Refer to Figure 3 , the preset program prompt gives the writing formats of the question "Question", the structured chain of thought "SCoT", and the program "Program", and the large language model can generate an initial program accordingly. Exemplarily, still taking the target question in Figure 2 as an example, after inputting the target question: "What color is the dog on the far right?" in Figure 2 and the preset program prompt into the large language model, the large language model can first generate the following structured chain of thought: "Structured Chain-of-Thought: Input: image Output: result: str 1: Initialize ImageBox from the image: image_box 2: Locate dog that exists definitely in image_box: dog_boxes 3: If no dog is found in image_box: 4:return "No dog found" 5: Select the right dog found in image_box: right_dog 6: Ask a VQA model what the color of the dog in right_dog: result 7: Return result”.

[0039] Next, the large language model can also generate the following initial program based on the above chain of structured thinking: “def inference(image): image_crop = GetCrop(image) dog_crops = image_crop.loc("dog") if not dog_crops: return"No dog found" dog_crops.sort(key=lambda x: x.horizontal_center, reverse=True) right_dog = dog_crops[0] color = right_dog.simple_vqa(“What color?") return color”.

[0040] In step 103, target code for recording the program execution process can be added to the initial program, that is, the initial program can be extended for interpretability to obtain an extended program. Specifically, the content of the initial program can be analyzed to understand the functions of program statements. Then, the code functions of the initial program can be extended to generate data structures during the program execution stage to save basic execution information.

[0041] Exemplarily, still taking Figure 2 as an example, the target code added to the initial program for recording the program execution process can be: “image-crop.save()”, “dog-crops.save()” and “right-dog.save()”. Specifically, “image-crop.save()” is used to indicate reading and saving the entire target image, that is, it is used to indicate reading and saving the above target image containing 4 dogs; “dog-crops.save()” is used to indicate saving the local image of each of the 4 dogs detected from the target image. Referring to Figure 2 , Figure 2 the right side of which shows the local images of each of the 4 dogs; “right-dog.save()” is used to indicate saving the local image of the dog on the far right.

[0042] In step 104, the target image can be input into the extension program to obtain a predicted answer to the target question, the execution process information of the extension program, and a screenshot image. The screenshot image can be a screenshot image associated with the target object pointed to by the target question and intercepted from the target image.

[0043] Exemplarily, still taking Figure 2 as an example, Figure 2 the target question shown in is: "What color is the dog on the far right?", at this time, the screenshot image can be a screenshot image associated with the target object: "the dog on the far right" pointed to by the target question "What color is the dog on the far right?" and intercepted from the target image. The execution process information can be expressed in the form of a dictionary, which can include the intermediate variable values and visual box variable values during the execution of the extension program. The "intermediate variable values" can include but are not limited to: how many times the extension program has executed a loop, where the return is made from the extension program, etc.; the "visual box variable values" can be the coordinate values of the area where the object detected from the target image is located. Exemplarily, still taking Figure 2 as an example, the "visual box variable values" can be the coordinate values of the area where the dog detected from the target image is located.

[0044] According to an exemplary embodiment of the present disclosure, the preset program prompt may further include an application programming interface description, where the application programming interface description can be used to indicate how to use the program to call open-world tools, that is, the application programming interface description can show how to call various types of specific functions.

[0045] By executing the extension program, the open-world tools can be called based on the application programming interface description to process the target image, and then a predicted answer, execution process information, and a screenshot image can be obtained. Specifically, during the execution of the extension program, the open-world multimodal tool can be used to convert the target question into an execution process composed of predefined program modules. For example, the key visual targets in the target image can be detected and marked to ensure that the captured execution process data can comprehensively reflect the reasoning process, and the intermediate variable values and the values of the visual box can be recorded synchronously. Exemplarily, still taking Figure 2 as an example, the intermediate process data including all dog images and the separately extracted image of the dog on the far right can be obtained after executing the extension program.

[0046] It should be noted that the open-world multimodal tool can usually be invoked in the form of program modules. For example, the open-world object detection model GroundingDINO can be used for object detection. GroundingDINO is an Open-Vocabulary Object Detection (OVOD) model that combines DINO (a DETR-based detection architecture) with a Grounding mechanism to achieve object detection based on text descriptions.

[0047] According to an exemplary embodiment of the present disclosure, the above open-world tool may include the following items: an image object detector, wherein the image object detector can be used to detect objects of a specified type in an image, that is, the image object detector can be used to identify and locate specific types of targets in the image; an image cropping tool, wherein the image cropping tool can be used to crop a specific area in the image, that is, the image cropping tool can be used to crop part or all of the area included in the image; a customized Python function, wherein the customized Python function can be used to perform at least one of specific types of calculation tasks, format conversion tasks, and logical judgment tasks to obtain a prediction answer and execution process information, and can automatically adjust the parameters of the open-world tool according to the detection result.

[0048] It should be noted that in addition to the above tools, the open-world tool may also include other types of tools, such as a vision language model, etc. The foregoing embodiments are merely an exemplary illustration.

[0049] In step 105, a multimodal form of explanation for the prediction answer can be generated based on the execution process information and the screenshot image, wherein the multimodal form can at least fuse text form and image form.

[0050] According to an exemplary embodiment of the present disclosure, a preset explanation prompt example can be obtained, wherein the preset explanation prompt example can be used to indicate how to generate a corresponding natural language explanation based on the program execution process. Then, based on the preset explanation prompt example and the execution process information, a natural language target text explanation for the prediction answer can be generated. Next, based on the target text explanation and the screenshot image, a multimodal form of explanation for the prediction answer can be generated.

[0051] Figure 4 is an example diagram showing a preset explanation prompt example according to an exemplary embodiment of the present disclosure. Refer to Figure 4 , the preset explanation prompt example gives the writing format of the program execution process, and text explanation data can be generated accordingly.

[0052] According to an exemplary embodiment of the present disclosure, preset explanation prompt examples and execution process information can also be input into a large language model to obtain a target text explanation in natural language for the predicted answer. Exemplarily, the large language model can be, but is not limited to, the GPT-4o model. Specifically: The execution process information of the extended program and a single-sample explanation prompt example can also be input into the large language model. Then, the large language model converts the execution process information into a coherent target text explanation according to the guidance of the single-sample explanation prompt example. The format of the target text explanation can be the same as that of the single-sample explanation prompt example. Moreover, the target text explanation can be composed of a specific data structure. For example, it can include, but is not limited to, execution process parameter values, multi-modal data names, and numerical values of positioning boxes. Additionally, the original numerical values of the visual box variables may not be directly presented in the target text explanation, but their positioning information can be reflected in a marked form. Specifically, the target text explanation can be composed of natural language, and the positioning information can be embedded in the corresponding visual target name in the target text explanation in a marked form.

[0053] For example, still taking Figure 2 as an example, for the two variables of the visual targets "the rightmost dog (right_dog)" and "image crop area (image_crop)", the marks "[crop]" or "[box]" can be embedded in the target text explanation, and then the following example can be formed: "Black. Because the dog [box] closest to the right in the image [box] is black. Boxes: BOX[0]=[10, 258, 206, 380] BOX[1]=[0, 0, 250, 400]".

[0054] Among them, [10, 258, 206, 380] and [0, 0, 250, 400] can represent the coordinate positions of the corresponding image crop areas. For example, a coordinate system can be established with the lower left corner of the entire target image as the origin. At this time, for [10, 258, 206, 380], the "10" and "206" inside can respectively represent the abscissa value of the lower left vertex and the abscissa value of the lower right vertex of the image crop area; the "258" and "380" inside can respectively represent the ordinate value of the lower left vertex and the ordinate value of the upper left vertex of the image crop area.

[0055] In this way, by adopting single-sample explanation prompt examples, not only can the format of the generated target text explanations be standardized, but also the over-reliance on a large amount of training data can be avoided, thereby improving the generalization ability of the model for generating natural language text explanations for predicted answers.

[0056] Figure 5 FIG. is an example diagram showing the specific implementation process of a visual question answering method according to an exemplary embodiment of the present disclosure. Referring to Figure 5 , the target question for the target image and the single-sample program prompt can be input into the large language model, and then the large language model can generate an initial program corresponding to the target question according to the guidance of the single-sample program prompt. Then, several lines of code for recording the program execution process can be added to the initial program to obtain an extended program.

[0057] Next, the target image can be input into the extended program, and then the extended program can process the target image by calling open-world tools based on the application programming interface description included in the above single-sample program prompt to obtain a predicted answer, execution process information, and a screenshot image.

[0058] Then, the predicted answer, execution process information, and a preset single-sample explanation prompt can also be input into the large language model to obtain a target text explanation in natural language form for the predicted answer. Next, the visual target can be embedded at the corresponding position of the target text explanation to obtain an intuitive and easy-to-understand multi-modal form of explanation.

[0059] The present disclosure provides a visual question answering method based on a single-example code-driven large model. By introducing preset program prompts and explanation prompts, the model can provide both predicted answers and output multi-modal explanations of the reasoning process without retraining the model. In this way, the interpretability and credibility of the predicted answers can be improved, thereby improving the model transparency and user trust.

[0060] Figure 6 FIG. is a block diagram showing a visual question answering device 600 according to an exemplary embodiment of the present disclosure.

[0061] Referring to Figure 6 , the visual question answering device 600 may include an image and question acquisition module 601, an initial program generation module 602, a program extension module 603, an answer prediction module 604, and a multi-modal explanation generation module 605.

[0062] The image and question acquisition module 601 can acquire a target image and a target question for the target image. The target image and the target question can form multi-modal question data, and its content and format can be flexibly changed according to the actual application scenario.

[0063] It should be noted that in the present disclosure, the visual question answering method according to the present disclosure can be executed by simply installing visual question answering software on the terminal. Exemplarily, the terminal in the present disclosure can be, but is not limited to: mobile phones, tablet computers, laptop computers, wearable devices, etc.

[0064] The initial program generation module 602 can generate an initial program based on the target question, where the initial program is used to perform inference to obtain the answer to the target question.

[0065] According to an exemplary embodiment of the present disclosure, the initial program generation module 602 can obtain a preset program hint, where the preset program hint can include a structural thinking chain construction method and a program hint example. The "structural thinking chain construction method" can be used to indicate how to construct a corresponding structural thinking chain based on the question; the "program hint example" can be used to indicate how to generate a corresponding program based on the structural thinking chain, that is, the program hint example can show the correspondence between the question and the program. Additionally, the preset program hint in the present disclosure can be a single-sample program hint.

[0066] Then, the initial program generation module 602 can construct a target structural thinking chain corresponding to the target question based on the structural thinking chain construction method. Next, the initial program generation module 602 can generate an initial program corresponding to the target structural thinking chain based on the program hint example, that is, a structured initial logical program can be generated.

[0067] In this way, by setting the preset program hint, the model used to generate the initial program can quickly learn the format of the program to be generated, and thus it can be ensured that the generated initial program meets the format and logical requirements. At the same time, the program generation efficiency can also be improved. Further, by adopting the learning method of single-sample program hint, the over-reliance on a large amount of training data can be avoided, thereby improving the generalization ability of the model used to generate the initial program.

[0068] The program extension module 603 can add target code for recording the program execution process to the initial program, that is, the initial program can be interpretably extended to obtain an extended program. Specifically, the content of the initial program can be analyzed to understand the function of the program statements. Then, the code function of the initial program can be extended to generate a data structure during the program execution stage to save basic execution information.

[0069] The answer prediction module 604 can input the target image into the extended program to obtain a predicted answer to the target question, the execution process information of the extended program, and a screenshot image. The screenshot image can be a screenshot image associated with the target object pointed to by the target question and intercepted from the target image.

[0070] According to an exemplary embodiment of the present disclosure, the preset program prompt may further include an application programming interface description, where the application programming interface description can be used to indicate how to use a program to call an open-world tool, that is, the application programming interface description can show how to call various types of specific functions.

[0071] By executing an extension program, the answer prediction module 604 can call an open-world tool based on the application programming interface description to process the target image, and then can obtain a predicted answer, execution process information, and a screenshot image. Specifically, during the execution of the extension program, an open-world multimodal tool can be used to convert the target question into an execution process composed of predefined program modules. For example, key visual targets in the target image can be detected and marked to ensure that the captured execution process data can comprehensively reflect the reasoning process, and the values of intermediate variables and visual box variables can be synchronously recorded.

[0072] It should be noted that the open-world multimodal tool can usually be called in the form of a program module. For example, an open-world object detection model GroundingDINO can be used for object detection. GroundingDINO is an open-vocabulary object detection model that combines the DINO and Grounding mechanisms to achieve object detection based on text descriptions.

[0073] According to an exemplary embodiment of the present disclosure, the above open-world tool may include the following items: an image object detector, where the image object detector can be used to detect objects of a specified type in an image, that is, the image object detector can be used to identify and locate specific types of targets in the image; an image cropping tool, where the image cropping tool can be used to crop a specific area in the image, that is, the image cropping tool can be used to crop part or all of the area included in the image; a customized Python function, where the customized Python function can be used to perform at least one of specific types of calculation tasks, format conversion tasks, and logical judgment tasks to obtain a predicted answer and execution process information, and can automatically adjust the parameters of the open-world tool according to the detection results.

[0074] It should be noted that in addition to the above tools, the open-world tool may further include other types of tools, such as a vision-language model, etc. The foregoing embodiments are merely an exemplary illustration.

[0075] The multimodal explanation generation module 605 can generate a multimodal form of explanation for the predicted answer based on the execution process information and the screenshot image, where the multimodal form can at least fuse text form and image form.

[0076] According to an exemplary embodiment of the present disclosure, the multimodal explanation generation module 605 may obtain a preset explanation hint example, where the preset explanation hint example may be used to indicate how to generate a corresponding natural language explanation based on the program execution process. Then, the multimodal explanation generation module 605 may generate a target text explanation in natural language for the predicted answer based on the preset explanation hint example and the execution process information. Next, the multimodal explanation generation module 605 may generate an explanation in multimodal form for the predicted answer based on the target text explanation and the screenshot image.

[0077] According to an exemplary embodiment of the present disclosure, the multimodal explanation generation module 605 may also input the preset explanation hint example and the execution process information into a large language model to obtain a target text explanation in natural language for the predicted answer. Exemplarily, the large language model may be, but is not limited to, the GPT-4o model. Specifically: The multimodal explanation generation module 605 may also input the execution process information of the extended program and a single-sample explanation hint example into the large language model. Then, the large language model may convert the execution process information into a coherent target text explanation according to the guidance of the single-sample explanation hint example. The format of the target text explanation may be the same as that of the single-sample explanation hint example. Moreover, the target text explanation may be composed of a specific data structure. For example, it may include, but is not limited to, execution process parameter values, multimodal data names, and numerical values of positioning boxes. In addition, the original numerical values of the visual box variables may not be directly presented in the target text explanation, but their positioning information may be reflected in the form of tags. Specifically, the target text explanation may be composed of natural language, and the positioning information may be embedded in the corresponding visual target names in the target text explanation in the form of tags.

[0078] In this way, by adopting the single-sample explanation hint example, not only can the format of the generated target text explanation be standardized, but also the over-reliance on a large amount of training data can be avoided, thereby improving the generalization ability of the model for generating text explanations in natural language for the predicted answer.

[0079] Figure 7 FIG. is a block diagram showing an electronic device 700 according to an exemplary embodiment of the present disclosure.

[0080] Referring to Figure 7 , the electronic device 700 includes at least one memory 701 and at least one processor 702. Instructions are stored in the at least one memory 701. When the instructions are executed by the at least one processor 702, the visual question answering method according to the exemplary embodiment of the present disclosure is executed.

[0081] As an example, the electronic device 700 can be a PC computer, a tablet device, a personal digital assistant, a smart phone, or other devices capable of executing the above instructions. Here, the electronic device 700 does not have to be a single electronic device, but can also be a collection of devices or circuits that can execute the above instructions (or instruction sets) individually or jointly. The electronic device 700 can also be a part of an integrated control system or system manager, or can be configured as a portable electronic device that interfaces with a local or remote (e.g., via wireless transmission).

[0082] In the electronic device 700, the processor 702 can include a central processing unit (CPU), a graphics processing unit (GPU), a programmable logic device, a dedicated processor system, a microcontroller, or a microprocessor. By way of example and not limitation, the processor can also include an analog processor, a digital processor, a microprocessor, a multi-core processor, a processor array, a network processor, and so on.

[0083] The processor 702 can run instructions or code stored in the memory 701, where the memory 701 can also store data. The instructions and data can also be sent and received over a network via a network interface device, where the network interface device can use any known transmission protocol.

[0084] The memory 701 can be integrated with the processor 702, for example, by arranging RAM or flash memory within an integrated circuit microprocessor, etc. In addition, the memory 701 can include a separate device, such as an external disk drive, a storage array, or other storage devices that can be used by any database system. The memory 701 and the processor 702 can be operatively coupled, or can communicate with each other, for example, through an I / O port, a network connection, etc., so that the processor 702 can read files stored in the memory.

[0085] In addition, the electronic device 700 can also include a video display (such as a liquid crystal display) and a user interaction interface (such as a keyboard, a mouse, a touch input device, etc.). All components of the electronic device 700 can be connected to each other via a bus and / or a network.

[0086] According to an exemplary embodiment of the present disclosure, a computer-readable storage medium may also be provided. When instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device can execute the above-described visual question answering method. Examples of such computer-readable storage media include: read-only memory (ROM), programmable read-only memory (PROM), electrically erasable programmable read-only memory (EEPROM), random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), flash memory, non-volatile memory, CD-ROM, CD-R, CD+R, CD-RW, CD+RW, DVD-ROM, DVD-R, DVD+R, DVD-RW, DVD+RW, DVD-RAM, BD-ROM, BD-R, BD-R LTH, BD-RE, Blu-ray or optical disc memory, hard disk drive (HDD), solid state drive (SSD), cartridge memory (such as, multimedia card, secure digital (SD) card or extreme digital (XD) card), magnetic tape, floppy disk, magneto-optical data storage device, optical data storage device, hard disk, solid state disk, and any other device configured to store a computer program and any associated data, data files, and data structures in a non-transitory manner and provide the computer program and any associated data, data files, and data structures to a processor or computer such that the processor or computer can execute the computer program. The computer program in the above computer-readable storage medium may run in an environment deployed in computer devices such as clients, hosts, proxy devices, servers, etc. In addition, in one example, the computer program and any associated data, data files, and data structures are distributed on a networked computer system such that the computer program and any associated data, data files, and data structures are stored, accessed, and executed in a distributed manner by one or more processors or computers.

[0087] According to an exemplary embodiment of the present disclosure, a computer program product may also be provided, including a computer program that, when executed by a processor, implements the visual question answering method according to the present disclosure.

[0088] According to the visual question answering method, device, electronic device, storage medium, and computer program product of the present disclosure, a decision-making basis picture and semantic association analysis can be generated synchronously while outputting the predicted answer, that is, an explanation of the reasoning process of the predicted answer can be added, enabling users to intuitively understand the correspondence between image features and semantic reasoning. Furthermore, the reasoning transparency and decision-making credibility can be improved, that is, the defects of low model transparency and low user trust caused by the black box limitation of traditional visual question answering models can be overcome, thus facilitating the popularization and application of visual question answering technology. Moreover, by outputting multi-modal explanations for the predicted answer, the display of the explanations is clearer and more intuitive, reducing the understanding difficulty and making it easy for users to quickly and accurately understand the entire reasoning process.

[0089] According to an exemplary embodiment of the present disclosure, by setting a preset program prompt, the model used to generate the initial program can quickly learn the format of the program to be generated, and thus the generated initial program can meet the format and logical requirements. At the same time, the program generation efficiency can also be improved. Further, by adopting the learning method of single-sample program prompts, the over-reliance on a large amount of training data can be avoided, thereby improving the generalization ability of the model used to generate the initial program.

[0090] According to an exemplary embodiment of the present disclosure, by adopting a single-sample explanation prompt example, not only can the format of the generated target text explanation be standardized, but also the over-reliance on a large amount of training data can be avoided, thereby improving the generalization ability of the model used to generate the natural language text explanation for the predicted answer.

[0091] Those skilled in the art will readily conceive of other embodiments of the present disclosure after considering the specification and practicing the invention disclosed herein. The present disclosure is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include well-known knowledge or conventional technical means in the technical field not disclosed herein. The specification and examples are only regarded as exemplary, and the true scope and spirit of the present disclosure are pointed out by the following claims.

[0092] It should be understood that the present disclosure is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is only limited by the appended claims.

Claims

1. A visual question answering method, characterized in that: include: Acquire a target image and a target question for the target image; Based on the target question, generating an initial program, wherein the initial program is used to perform reasoning to obtain an answer to the target question; Adding object code for recording the program execution process to the initial program to obtain an extended program; Inputting the target image into the extension program to obtain a predicted answer to the target question, execution process information of the extension program, and a screenshot image, wherein the screenshot image is a screenshot image associated with a target object pointed to by the target question and obtained by intercepting the target image; Based on the execution process information and the screenshot image, a multimodal explanation for the predicted answer is generated, wherein the multimodal form at least integrates a text form and an image form.

2. The visual question answering method according to claim 1, wherein: The generating of an initial program based on the target problem comprises: Obtaining a preset program prompt, wherein the preset program prompt includes a structured thinking chain construction method and a program prompt example, wherein the structured thinking chain construction method is used to indicate how to construct a corresponding structured thinking chain based on a problem, and the program prompt example is used to indicate how to generate a corresponding program based on the structured thinking chain; Based on the structural thinking chain construction method, construct a target structural thinking chain corresponding to the target problem; Based on the program prompt example, the initial program corresponding to the target structure thinking chain is generated.

3. The visual question answering method according to claim 2, wherein: The preset program prompt further includes an application programming interface description, wherein the application programming interface description is used to indicate how to use the program to call the open world tool; The step of inputting the target image into the extension program to obtain a predicted answer to the target question, execution process information of the extension program, and a screenshot image includes: By executing the extension program, the open world tool is called based on the application programming interface description to process the target image, thereby obtaining the predicted answer, the execution process information and the screenshot image.

4. The visual question answering method according to claim 3, wherein: The open world tools include the following: An image object detector, wherein the image object detector is used to detect objects of a specified type in an image; An image cropping tool, wherein the image cropping tool is used to crop a part or all of an area contained in an image; A customized Python function, wherein the customized Python function is used to perform at least one of a specific type of computing task, format conversion task, and logic judgment task to obtain the predicted answer and the execution process information.

5. The visual question answering method according to claim 1, wherein: The generating a multimodal explanation for the predicted answer based on the execution process information and the screenshot image comprises: Obtaining a preset explanation prompt example, wherein the preset explanation prompt example is used to indicate how to generate a corresponding natural language explanation based on a program execution process; Based on the preset explanation prompt example and the execution process information, generating a target text explanation in natural language for the predicted answer; Based on the target text explanation and the screenshot image, a multimodal explanation for the predicted answer is generated.

6. The visual question answering method according to claim 5, characterized in that: The step of generating a target text explanation in a natural language for the predicted answer based on the preset explanation prompt example and the execution process information includes: The preset explanation prompt example and the execution process information are input into a large language model to obtain a target text explanation in natural language for the predicted answer.

7. A visual question-answering device, characterized in that: include: An image and question acquisition module, configured to acquire a target image and a target question for the target image; An initial program generation module is configured to generate an initial program based on the target problem, wherein the initial program is used to perform reasoning to obtain an answer to the target problem; A program extension module is configured to add a target code for recording the program execution process to the initial program to obtain an extended program; an answer prediction module, configured to input the target image into the extension program, obtain a predicted answer to the target question, execution process information of the extension program, and a screenshot image, wherein the screenshot image is a screenshot image associated with a target object pointed to by the target question and obtained by intercepting the target image; The multimodal explanation generation module is configured to generate a multimodal explanation for the predicted answer based on the execution process information and the screenshot image, wherein the multimodal form at least integrates a text form and an image form.

8. An electronic device, characterized in that: include: processor; a memory for storing instructions executable by the processor; The processor is configured to execute the instructions to implement the visual question answering method as described in any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that: When the instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is enabled to execute the visual question answering method as claimed in any one of claims 1 to 6.

10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the visual question answering method according to any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Program chain-based visual question and answer method, device and equipment, medium and product

    CN118365919A

  • Visual question and answer method, device, equipment, medium and product

    CN118798372A

  • Combined visual question and answer method based on core-to-global semantic fusion reasoning

    CN119397384A

  • Method and apparatus for visual question answering, computer device and medium

    US20210406619A1

Cited By

  • Data set creation method, computer equipment and storage medium

    CN120806169A