Confidence-based interactive neurosymbolic visual questions and answers

By adopting a confidence-based neural symbolic method in visual question-and-answer tasks, combining AI scene perception and problem analysis models, the problems of high interpretability and computational cost in the prior art are solved, and more accurate and reliable answer generation is achieved.

CN119998799APending Publication Date: 2025-05-13SAMSUNG ELECTRONICS CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202380070720.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-11-07
Filing Date
2023-11-09
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The prior art lacks interpretability, high computational cost, and demands on large amounts of data in visual question-and-answer (VQA) tasks, and data-driven models may take advantage of deviations in the dataset rather than performing advanced reasoning, and cannot maintain inference consistency when answering combination questions and their sub-questions.

Method used

A confidence-based neural symbolic method is adopted to generate feature predictions through AI scene-aware model, and symbolic programs and confidence scores are generated through AI problem analysis model. The symbolic program with the highest confidence is selected for logical operations to determine natural language answers, and users are allowed to interact to adjust confidence scores.

Benefits of technology

Improve the interpretability and accuracy of the visual Q&A task, reduce calculation costs, and improve the reliability and confidence of the answers through user interaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119998799A_ABST
    Figure CN119998799A_ABST
Patent Text Reader

Abstract

A method of performing a visual question and answer (VQA) includes: obtaining an image and a question corresponding to the image; generating a plurality of feature predictions with respect to at least one object included in the image by providing the image to an artificial intelligence (AI) scene awareness model; generating a plurality of symbolic programs and a plurality of program confidence scores by providing the problem to the AI problem analytic model; selecting a symbolic program associated with the highest one of the plurality of program confidence scores; executing the selected symbolic program by performing a logical operation set included in the selected symbolic program on the plurality of feature predictions; and determining a natural language answer to the question based on a result of the logical operation set.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to a method and a device for performing a visual question answering (VQA) task, and more particularly, to an interactive confidence-based neural symbolic method and a device for performing a VQA task. Background Art

[0002] Visual Question Answering (VQA) refers to the task of providing accurate natural language answers given an image context and a natural language question about the image. VQA is becoming increasingly important in a wide range of applications, including intelligent assistants, information retrieval, and assistance for users with visual impairments. Visual questions can selectively target different regions and aspects of an image and may require detailed understanding and complex reasoning about the image.

[0003] While data-driven methods (e.g., deep learning) can work in an end-to-end manner, their lack of interpretability, high computational cost, and requirement for large amounts of data may hinder their application in the real world. Furthermore, data-driven visual question answering (VQA) models may tend to exploit biases in the dataset to find shortcuts instead of performing high-level reasoning, and may fail to maintain reasoning consistency when answering a combined question and its sub-questions.

[0004] Neural-symbolic (NS) learning provides an effective approach for VQA by combining the advantages of neural network learning and symbolic reasoning. Performing VQA with NS introduces transparency into the reasoning process and allows diagnosis of each execution step. However, the neural network (NN) in some NS methods may be regarded as a black box model and cannot provide information for user interaction. Summary of the invention

[0005] Example embodiments address at least the above problems and / or disadvantages and other disadvantages not described above. Also, example embodiments are not required to overcome the disadvantages described above, and may not overcome any of the problems described above.

[0006] According to one aspect of the present disclosure, a method for performing visual question answering (VQA) includes: obtaining an image and a question corresponding to the image; generating multiple feature predictions about at least one object included in the image by providing the image to an artificial intelligence (AI) scene perception model; generating multiple symbolic programs and multiple program confidence scores by providing the question to an AI question parsing model, wherein each symbolic program in the multiple symbolic programs includes a set of logical operations and is associated with a corresponding program confidence score in a plurality of program confidence scores; selecting a symbolic program associated with a highest program confidence score in the multiple program confidence scores; running the selected symbolic program by executing the set of logical operations included in the selected symbolic program on the multiple feature predictions; and determining a natural language answer to the question based on a result of the set of logical operations.

[0007] The image and the question may be received from a user, and the method may further include providing a natural language answer to the user as a response to the question.

[0008] Multiple feature predictions may be associated with multiple feature confidence scores generated by the AI ​​scene perception model.

[0009] A set of logical operations included in the selected symbolic program may be associated with a plurality of operation confidence scores, and each logical operation in the set of logical operations may be associated with an operation confidence score from the plurality of operation confidence scores, the operation confidence score being determined based on the program confidence score and at least one of the plurality of feature confidence scores.

[0010] Based on at least one confidence score among the multiple program confidence scores, the multiple feature confidence scores, and the multiple operation confidence scores being below a threshold: the method may also include obtaining user input corresponding to the at least one confidence score; and adjusting the at least one confidence score based on the user input.

[0011] The method may also include determining an answer confidence score corresponding to the natural language answer based on the plurality of operational confidence scores.

[0012] The method may also include generating enhanced training data based on multiple symbolic programs; and training the AI ​​scene perception model based on the enhanced training data.

[0013] The step of generating enhanced training data may include: generating a first ranking of multiple symbolic programs based on multiple program confidence scores; generating a second ranking of the multiple symbolic programs based on multiple consistency losses between the multiple program confidence scores and the problem; selecting a subset of the multiple symbolic programs based on the first ranking and the second ranking; and generating enhanced training data based on outputs of the subset of the multiple symbolic programs.

[0014] According to aspects of the present disclosure, a device for performing VQA includes: a memory configured to store instructions; and at least one processor configured to execute the instructions to perform the following operations: obtain an image and a question corresponding to the image; generate multiple feature predictions about at least one object included in the image by providing the image to an AI scene perception model; generate multiple symbolic programs and multiple program confidence scores by providing the question to an AI question parsing model, wherein each symbolic program in the multiple symbolic programs includes a set of logical operations and is associated with a corresponding program confidence score in a plurality of program confidence scores; select a symbolic program associated with a highest program confidence score in the multiple program confidence scores; run the selected symbolic program by executing the set of logical operations included in the selected symbolic program on the multiple feature predictions; and determine a natural language answer to the question based on the result of the set of logical operations.

[0015] An image and a question are received from a user, and the at least one processor may be further configured to execute instructions to provide a natural language answer to the user as a response to the question.

[0016] Multiple feature predictions may be associated with multiple feature confidence scores generated by the AI ​​scene perception model.

[0017] A set of logical operations included in the selected symbolic program may be associated with a plurality of operation confidence scores, and each logical operation in the set of logical operations may be associated with an operation confidence score from the plurality of operation confidence scores, the operation confidence score being determined based on the program confidence score and at least one of the plurality of feature confidence scores.

[0018] At least one processor may also be configured to run instructions to perform the following operations: based on at least one confidence score among multiple program confidence scores, multiple feature confidence scores, and multiple operation confidence scores being below a threshold: obtaining user input corresponding to at least one confidence score; and adjusting at least one confidence score based on the user input.

[0019] The at least one processor may be further configured to execute instructions to determine an answer confidence score corresponding to the natural language answer based on the plurality of operational confidence scores.

[0020] At least one processor may also be configured to execute instructions to perform the following operations: generate enhanced training data based on multiple symbolic programs; and train the AI ​​scene perception model based on the enhanced training data.

[0021] To generate enhanced training data, at least one processor may also be configured to: generate a first ranking of multiple symbolic programs based on multiple program confidence scores; generate a second ranking of multiple symbolic programs based on multiple consistency losses between the multiple program confidence scores and the problem; select a subset of the multiple symbolic programs based on the first ranking and the second ranking; and generate enhanced training data based on the output of the subset of the multiple symbolic programs.

[0022] According to aspects of the present disclosure, a non-transitory computer-readable medium storing instructions, wherein the instructions, when executed by at least one processor of an apparatus for performing VQA, cause the at least one processor to perform the following operations: obtain an image and a question corresponding to the image; generate multiple feature predictions about at least one object included in the image by providing the image to an AI scene perception model; generate multiple symbolic programs and multiple program confidence scores by providing the question to an AI question parsing model, wherein each of the multiple symbolic programs includes a set of logical operations and is associated with a corresponding program confidence score in a plurality of program confidence scores; select a symbolic program associated with a highest program confidence score in the plurality of program confidence scores; run the selected symbolic program by executing the set of logical operations included in the selected symbolic program on the multiple feature predictions; and determine a natural language answer to the question based on a result of the set of logical operations.

[0023] Multiple feature predictions may be associated with multiple feature confidence scores generated by the AI ​​scene perception model.

[0024] A set of logical operations included in the selected symbolic program may be associated with a plurality of operation confidence scores, and each logical operation in the set of logical operations may be associated with an operation confidence score from the plurality of operation confidence scores, the operation confidence score being determined based on the program confidence score and at least one of the plurality of feature confidence scores.

[0025] The instructions may also be configured to cause at least one processor to perform the following operations: based on at least one confidence score among a plurality of program confidence scores, a plurality of feature confidence scores, and a plurality of operation confidence scores being below a threshold: obtaining user input corresponding to at least one confidence score; and adjusting at least one confidence score based on the user input. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] The above and other aspects, features and aspects of the embodiments of the present disclosure will become more apparent through the following description in conjunction with the accompanying drawings, in which: Figure 1 is a diagram illustrating a system for performing visual question answering according to an embodiment of the present disclosure; Figure 2is a flowchart illustrating a method of performing visual question answering according to an embodiment of the present disclosure; FIG. 3A to FIG. 3D An example of interactive visual question answering according to an embodiment of the present disclosure is shown; Figure 4A is a diagram of a training framework for training a system for performing visual question answering according to an embodiment of the present disclosure; Figure 4B is a flowchart illustrating a method of training a system for performing visual question answering according to an embodiment of the present disclosure; Figure 5A is a diagram of a training framework for training a question parsing module according to an embodiment of the present disclosure; Figure 5B is a flow chart illustrating a method of training a question parsing module according to an embodiment of the present disclosure; FIG. 6A to FIG. 6C An example of interactive visual question answering according to an embodiment of the present disclosure is shown; 7A to 7C An example of interactive visual question answering according to an embodiment of the present disclosure is shown; FIG. 8A to FIG. 8B An example of interactive visual question answering according to an embodiment of the present disclosure is shown; 9A to 9C An example of interactive visual question answering according to an embodiment of the present disclosure is shown; Fig.10 is a flowchart illustrating a method of performing visual question answering according to an embodiment of the present disclosure; Fig.11 is a diagram of an electronic device for performing a multimodal retrieval task according to an embodiment of the present disclosure; and Fig.12 According to the embodiment of the present disclosure Fig.11 An illustration of components of one or more electronic devices. DETAILED DESCRIPTION

[0027] Example embodiments are described in more detail below with reference to the accompanying drawings.

[0028] In the following description, the same drawing reference numerals are used for the same elements even in different drawings. Matters defined in the specification, such as detailed structures and elements, are provided to assist in a comprehensive understanding of the example embodiments. However, it is apparent that the example embodiments can be practiced without those specifically defined matters. In addition, well-known functions or structures are not described in detail because they would obscure the description with unnecessary detail.

[0029] Expressions such as “at least one of…” when following a list of elements modify the entire list of elements and do not modify the individual elements of the list. For example, the expression “at least one of a, b, and c” should be understood to include only a, only b, only c, both a and b, both a and c, both b and c, and all or any variation of the foregoing examples of a, b, and c.

[0030] Although terms such as "first", "second", etc. may be used to describe various elements, these elements must not be limited to the above terms. The above terms may be used only to distinguish one element from another.

[0031] The term "module" is intended to be broadly interpreted as hardware, software, firmware, or any combination thereof.

[0032] It is apparent that the systems and / or methods described herein may be implemented in various forms of hardware, firmware, or a combination of hardware and software. The actual dedicated control hardware or software code used to implement these systems and / or methods does not limit the implementation. Therefore, the operation and behavior of the systems and / or methods are described herein without reference to specific software code - it should be understood that software and hardware can be designed to implement the systems and / or methods based on the description herein.

[0033] Even though particular combinations of features are recited in the claims and / or disclosed in the specification, these combinations are not intended to limit the possible enabling disclosure. In fact, many of these features may be combined in ways not specifically recited in the claims and / or disclosed in the specification. Although each dependent claim listed below may only be directly dependent on one claim, the possible enabling disclosure includes each dependent claim in combination with every other claim in the claim set.

[0034] Unless explicitly described, none of the elements, actions, or instructions used herein should be interpreted as critical or essential. In addition, as used herein, the singular form is intended to include one or more items and can be used interchangeably with "one or more". In addition, as used herein, the term "set" is intended to include one or more items (e.g., related items, unrelated items, combinations of related and unrelated items, etc.), and can be used interchangeably with "one or more". In cases where only one item is involved, the term "one" or similar language is used. In addition, as used herein, the terms "having" and the like are intended to be open terms. In addition, unless otherwise expressly stated, the phrase "based on" is intended to mean "based at least in part on".

[0035] In order to improve learning efficiency and interpretability, neural symbolic (NS) learning has been studied to combine the high interpretability, provable correctness, and convenience of using human expert knowledge of symbolic manipulation with the advantages of neural networks (NNs). For visual question answering (VQA), NS methods can use NNs to extract object concepts (which can be called scene perception) and convert questions into symbolic programs (which can be called question parsing), and perform symbolic manipulation by executing the programs on the concepts.

[0036] However, NS methods cannot eliminate the shortcomings of NNs. For example, there are inevitable uncertainties in NNs due to the probability variations in random events or the lack of understanding of the process. Some methods for NS VQA may focus on reducing the requirements for symbolic labels (e.g., neural symbolic concept learner (NS-CL)), learning new symbols (e.g., meta-concept learner), or increasing the complexity of the task (e.g., video question answering, requiring machines to understand the laws of physics) without considering the uncertainty propagation on the reasoning path. The lack of uncertainty awareness in reasoning fails to consider the long-tail distribution of visual concepts and the unequal importance of reasoning steps in real data, and cannot provide information for user interaction, which can lead to intolerable errors for safety-critical applications.

[0037] Thus, embodiments may provide uncertainty awareness, which may account for large variance in predictions for concepts at the end of heavy-tailed distributions, set the importance of reasoning steps based on uncertainty quantification of concepts and programs, and alert humans to potentially incorrect reasoning for interaction. For example, one or more example embodiments may provide an interactive confidence-based NS (CBNS) framework to evaluate the confidence of NN modules and perform reasoning based on confidence assessments to perform tasks such as VQA. In an embodiment, the confidence assessment may also be used as a signal for user interaction. One or more example embodiments may provide a probabilistic question parser that may not use a resource-intensive reinforcement (REINFORCE) learning process and may generate multiple program candidates with confidence assessments. One or more example embodiments may also include a probabilistic scene perception module that may provide an object-based scene representation and a confidence assessment for each attribute of one or more objects in an image. According to one or more example embodiments, the object-based scene representation and the program with confidence assessment may be used to evaluate the confidence of the answer during the reasoning process, which may allow users to interact and provide feedback on weak links based on confidence levels to improve the reliability of the answer. Embodiments may be model agnostic and compatible with other NSVQA frameworks.

[0038] One or more example embodiments may consider uncertainty in both scene perception and question parsing in the context of VQA. Some NS-CL methods may use differentiable program runs to train visual representations and use REINFORCE to train question parsers to avoid the requirement for real concepts and programs. While it is feasible to quantify uncertainty in NS-CL and inherit the advantages of NS-CL, the joint uncertainty in both scene perception and question parsing may exacerbate the training overhead of REINFORCE because gradients from the probability module may be passed into the program runner, adding noise to parameter updates.

[0039] On the contrary, according to one or more example embodiments, the learning efficiency of the question parser can be improved by adding reconstruction loss, consistency loss and variational discarding, so that the question parser can achieve high accuracy using a limited amount of data. In addition, data augmentation rules can be used to select predicted programs through confidence assessment, which can be based on uncertainty quantification through variational discarding, so that the selected program is likely to be accurate. Then, the selected program can be used to train the scene perception module without the need for a real program. In addition, the uncertainty of the scene perception module can be quantified to evaluate the confidence of the object concept prediction of the scene perception module (which can be referred to as concept quantization in NS-CL). The concept quantization with confidence assessment can be input into the predicted program candidate for confidence-based reasoning.

[0040] According to one or more example embodiments, CBNS VQA can provide benefits by allowing efficient user interaction. For example, in some NS VQA methods, only one representation for an image and one deterministic procedure for an associated question may be predicted, and a single answer may be provided at the end. In contrast, one or more example embodiments may provide a confidence assessment for each step of reasoning starting from scene perception, which may be used to trigger further investigation or even user interaction. For example, whenever the confidence level of a particular step is too low (e.g., below a threshold), the CBNS VQA system may request the user to review and correct (if necessary) the reasoning of the particular step, thereby improving the accuracy and confidence of the final answer.

[0041] Figure 1 is a diagram illustrating a system for performing visual question answering according to an embodiment of the present disclosure.

[0042] For example, Figure 1 The CBNS system 100 for performing VQA shown may include a scene perception module 102 , a question parsing module 104 , a user interaction module 106 , and a program execution module 108 .

[0043] The scene perception module 102 can extract object-based image representations and provide confidence assessments based on uncertainty quantification. For example, a pre-trained mask region convolutional neural network (R-CNN) model can be used to generate object proposals. The bounding box for each proposal can then be paired with the original input image and sent to a ResNet-34 model to extract region-based features and image-based features. The concatenation of the region-based features and the image-based features is then used for concept quantification and uncertainty quantification.

[0044] The question parsing module 104 may translate a natural language question into a plurality of programs associated with confidence scores. The plurality of programs may facilitate the search for an accurate program. In an embodiment, the question parsing module 104 may be referred to as a question parser.

[0045] The program execution module 108 can execute the program from the question parsing module 104 based on the concept quantification from the scene perception module 102 and the confidence assessment for answer prediction. In some embodiments, when a type mismatch occurs between the input and output of adjacent operations in execution, an error flag can be raised for the program. In an embodiment, the program execution module 108 can output the answer of the executable program with the highest confidence score. In addition, when there is an error in all program candidates for the question, the program execution module 108 can randomly sample answers from all possible outputs of the final operation.

[0046] The user interaction module 106 may be used to verify and, if necessary, correct the reasoning at each step of the reasoning. In an embodiment, the program execution module 108 may track the confidence of the intermediate answers and may trigger the user interaction module 106 to check the reasoning when the confidence score is below a threshold. In an embodiment, the threshold may be determined based on a trade-off between answer accuracy and interaction requirements.

[0047] Therefore, in one or more example embodiments, the question parsing module 104 may convert the question into a program and provide a confidence score for each candidate for a plurality of program candidates. The scene perception module 102 may provide an object-based image representation and a confidence evaluation for each concept prediction. The program execution module 108 may execute the program on the object-based representation and the confidence evaluation for each logical operation.

[0048] In order to make the execution output fully differentiable with respect to the parameters in the scene perception module 102 for concept learning, the program execution module 108 may be a quasi-symbolic execution module, which may mean that the intermediate results of the program may be represented as an attention mask for all objects in the scene. For example, each element of the mask The scene can be represented The probability that an object belongs to an intermediate result. However, instead of using REINFORCE to train the question parsing module 104, semi-supervised learning can be used to improve learning efficiency. Therefore, in an embodiment, a sufficiently accurate question parsing module 104 with uncertainty quantification can be obtained from limited fully annotated data.

[0049] Furthermore, in order to avoid the requirement for real programs or programs from REINFORCE's exploration for training concept learners, a program with a high probability of being correct may be selected from the predicted programs by the question parsing module 104 for concept learning. For example, a program may be selected based on uncertainty quantification by predicting the confidence evaluation of the program, and the scene perception module 102 with uncertainty quantification may be trained using the program. In this way, in an end-to-end training approach, the uncertainty of training the question parsing module 104 may not interfere with the training of the scene perception module 102. For reasoning, the confidence evaluation along the reasoning path of the scene perception module 102 and the question parsing module 104 may be used to determine the final answer to the question by the program execution module 108 and the user interaction module 106 may request user interaction.

[0050] Figure 2 is a flowchart illustrating a method for performing visual question answering according to an embodiment of the present disclosure, FIG. 3A to FIG. 3D An example of an interactive visual question-answering system according to an embodiment of the present disclosure is shown. In particular, according to an embodiment of the present disclosure, Figure 3A An example of an input image is shown, Figure 3B shows an example of the output of the scene perception module, Figure 3C shows an example of the output of the question parsing module, Figure 3D An example of the output of running the module is shown below. FIG. 3A to FIG. 3D To describe Figure 2 However, the embodiment is not limited thereto. For example, according to the embodiment, different inputs may be provided, and the intermediate output and the final output may be in different forms or include different information.

[0051] In an embodiment, the Figure 1 CBNS system implementation Figure 2operations. For example, as discussed in more detail below, operations S201 to S204 may be performed by the scene perception module 102 or using the scene perception module 102, operations S207 to S209 may be performed by the question parsing module 104 or using the question parsing module 104, operations S205 to S206 and S210-S211 may be performed by the user interaction module 106 or using the user interaction module 106, and operations S212 to S214 may be performed by the program execution module 108 or using the program execution module 108, however, the embodiments are not limited thereto. For example, in an embodiment, operations S205 to S206 and S210-S211 may be performed by one or more of the scene perception module 102, the question parsing module 104, and the program execution module 108 or using one or more of the scene perception module 102, the question parsing module 104, and the program execution module 108.

[0052] like Figure 2 As shown, the method 200 may include an operation S201 of receiving image data, and an operation S202 of generating an object suggestion corresponding to the image data. Figure 3A An example of an input image 300 labeled based on object proposals is shown. Figure 3A As shown, objects 654-0 through 654-6 may be detected in the input image 300 and associated with bounding boxes and labels.

[0053] Method 200 may further include an operation S203 of extracting image features and quantifying concepts. In an embodiment, the features and concepts may correspond to attributes of an object. Figure 3B As shown, attributes (such as color, material, shape, and size) may be determined for each of objects 654-0 to 654-6. As another example, the size attribute may include two concepts, such as "large" and "small". Figure 3B As shown, for object 654-0, the concept of "blue" may be determined for the color attribute, the concept of "metal" may be determined for the material attribute, the concept of "cylinder" may be determined for the shape attribute, and the concept of "large" may be determined for the size attribute.

[0054] Method 200 may also include an operation S204 of measuring the confidence corresponding to the object. In an embodiment, the confidence may include a confidence score for each concept corresponding to each object. Figure 3BAs shown, for object 654-0, a confidence score of 0.9644 may be calculated for the concept "blue", a confidence score of 0.9995 may be calculated for the concept "metal", a confidence score of 0.9981 may be calculated for the concept "cylinder", and a confidence score of 0.9995 may be calculated for the concept "large". In addition, the calculated confidence may include the total confidence score of each object on all its corresponding concepts, and the total confidence score of each attribute on all its corresponding concepts. For example, Figure 3B As shown, based on the confidence scores of the concepts discussed above, the total confidence score of object 654-0 can be calculated as 0.9617. In addition, in the case where a color attribute corresponds to each of objects 654-0 to 654-6, the total confidence score for the color attribute can be calculated as 0.9507 based on the confidence scores for the concepts. In addition, the total confidence for the input image can be determined based on the total confidence score for the attribute and the total confidence score for the object. For example, Figure 3B As shown, a total confidence score of 0.9473 may be calculated for the input image 300 .

[0055] The method 200 may also include an operation S205 of determining whether the confidence score associated with the concept is too low (e.g., below a confidence threshold). Based on the confidence score being too low (yes at operation S205), the method 200 proceeds to operation S206, which may include at least one of requesting user interaction and triggering further analysis. Based on the confidence score being satisfactory (no at operation S205), e.g., greater than or equal to the confidence threshold, the method 200 may proceed to operation S212.

[0056] The method 200 may further include an operation S207 of receiving one or more natural language questions, and an operation S208 of mapping the natural language questions into a program. Figure 3C As shown, program candidates (e.g., candidate 1, candidate 2, and candidate 3) may be determined based on the input question "Is there a green rubber cylinder and a blue shiny cylinder behind it?" In an embodiment, each program candidate may include one or more logical operations that may be performed on the attributes discussed above.

[0057] Method 200 may also include an operation S209 of measuring the confidence level corresponding to the program. Figure 3C As shown, a confidence score of 0.9424 may be calculated for candidate 1, a confidence score of 0.9610 may be calculated for candidate 2, and a confidence score of 0.9558 may be calculated for candidate 3.

[0058] Method 200 may also include an operation S210 of determining whether a confidence score associated with the program candidate is too low (e.g., below a confidence threshold). Based on the confidence score being too low (yes at operation S210), method 200 proceeds to operation S211, which may include at least one of requesting user interaction and triggering further analysis. Based on the confidence score being satisfactory (no at operation S210), e.g., greater than or equal to the confidence threshold, method 200 may proceed to operation S212. In an embodiment, the confidence threshold used in operation S210 may be the same as or different from the confidence threshold used in operation S205.

[0059] Method 200 may also include an operation S212 of running one or more programs on the object-based representations (e.g., the concepts and attributes discussed above). Method 200 may also include an operation S213 of evaluating the confidence of the logical operations corresponding to the programs, and an operation S214 of mapping the results of the logical operations to natural language answers. For example, Figure 3D As shown, based on the confidence score of candidate 2 being higher than the confidence scores of candidate 1 and candidate 3, candidate 2 may be selected. Then, the logical operation of candidate 2 may be performed on the concepts discussed above to obtain a natural language answer of "no". When the logical operation of candidate 2 is performed on the attributes discussed above, a confidence score may be calculated for each logical operation included in candidate 2, and then a total confidence score of 0.9483 may be calculated for the natural language answer "no".

[0060] Figure 4A is a diagram of a training framework for training a system for performing visual question answering according to an embodiment of the present disclosure. Figure 4B is a flowchart illustrating a method of training a system for performing visual question answering according to an embodiment of the present disclosure.

[0061] like FIG. 4A to FIG. 4B As shown, the method 400B corresponding to the training framework 400A may include operation S401 of using limited fully annotated data to train the question parsing module 104. For example, the question parsing module 104 may be trained based on a relatively small fully annotated training data set. For example, the fully annotated training data may be fully labeled with a true label for at least one of an object, a concept, a question, a procedure, a logical operation, and an answer.

[0062] Method 400B may also include operation S402 of predicting and evaluating a program generated based on a question sampled from a training set without using a real program. For example, question parsing module 104 may receive a relatively large amount of partially annotated data. The partially annotated data may include only high-level real labels (such as labels for questions and answers).

[0063] Method 400B may also include, for example, an operation S403 of selecting data based on a data enhancement rule. For example, based on the partially annotated data, the question parsing module 104 may generate a pseudo-label corresponding to the partially annotated data, and a confidence score corresponding to the pseudo-label. In an embodiment, the partially annotated data and the pseudo-label may be referred to as enhanced data.

[0064] The method 400B may further include an operation S404 of training the scene perception module 102 based on the enhanced data. For example, based on the confidence score, some of the enhanced data may be selected, and the scene perception module 102 may be trained based on the selected enhanced data.

[0065] FIG. 5A to FIG. 5B is a diagram of a training framework for training a question parsing module according to an embodiment of the present disclosure.

[0066] In the discussion below with reference to training frameworks 500A and 500B, It can represent the input image, Representable objects, can represent properties of an object, and Can represent concepts. In addition, The set of all concepts that can be represented, and The number of concepts that can be represented. properties describe each object, and the properties Can include For example, in the CLEVR dataset, each object has five attributes (e.g., color, material, shape, size, and position), and the size attribute may include two concepts (i.e., large and small). The scene annotation of the represented image may include a description of the object concept and location. Can represent the a problem, and It can represent the first In addition, It can be expressed as words, and It can represent the first operations. Additionally, can be expressed as an estimate of the variable. The answer can be expressed as , and the confidence score can be expressed as .

[0067] like FIG. 5A to FIG. 5BAs shown, the question parsing module 104 may include a machine learning (ML) model (such as an attention-based sequence-to-sequence (seq2seq) model) including an encoder 104-1 and a decoder 104-2 that may be used to convert a question into a symbolic program.

[0068] For example, encoder 104-1 may be represented by a bidirectional long short-term memory (LSTM) network that takes a variable-length question as input and trains the encoder 104-1 at time steps according to Equations 1 and 2 below. Output encoding vector : Equation (1) Equation (2) In Equation 1 and Equation 2, may represent the jointly trained word embedding for encoder 104-1, and Denote the output state and hidden state of the forward network and the backward network, respectively. Decoder 104-2 may be a similar LSTM network having an output according to Equation 3 below: Equation (3) In equation 3, can represent the previous token of the output sequence, and can represent the decoder word embedding, which is then fed to an attention layer with a constant attention matrix to obtain the context vector via the following equation 4 As the encoding state The weighted sum of: Equation (4) Then, you can Passed to a fully connected layer with softmax activation to get the predicted label The conditional distribution of .

[0069] To train the question parsing module 104, a reconstructor 501 may be used to reconstruct the question from the hidden layer of the decoder 104-2 to ensure that the program preserves the information in the question. The reconstructor 501 may be a similar decoder having an output according to the following equation 5 : Equation (5) The output can then be fed to the attention layer according to Equation 6 below: Equation (6) In equation 6, represents the attention weight matrix of the reconstructor 501. To obtain the distribution of predicted labels. Then, the reconstruction loss can be determined according to the following equation 7: Equation (7) In equation 7, Can be expressed for The sequence of hidden states of the decoder of the problem. In addition, the prediction of both the seq2seq model of encoder 104-1 and the reconstructor 501 at the current time step can be based on the prediction at the previous time step. In order to perform sequence-level consistency, it can be determined according to the following equation 8 FIG. 5A to FIG. 5B The sequence consistency loss shown in: Equation (8) In equation 8, and may represent the hidden states of encoder 104-1 and decoder 104-2 at the last time step, respectively, and Represents the dimension of the hidden state.

[0070] To quantify the model uncertainty of the question parsing module 104, variational discarding (VD) and local reparameterization can be used. Scale-invariant log-uniform prior and the factorized Gaussian approximate posterior Can be used with parameters In an embodiment, it can be learned by maximizing according to the following equations 9 and 10: : Equation (9) Equation (10) In Equation 9 and Equation 10, it can be approximated by an unbiased differentiable mini-batch-based Monte Carlo estimator according to the following Equation 11 : Equation (11) In equation 11, we can get sampling .For example, ,and ,in, Can represent the weights, ,and It can represent the discard rate. FIG. 5A to FIG. 5B As shown, variational dropout in the decoder 104-2 of the question parsing module 104 can be used for uncertainty quantification. For example, using the decoder output and context vector , the distribution of the predicted markers can be obtained according to the following equation 12: Equation (12) In an embodiment, the deterministic weight matrix Can be used to label predictions to reduce model complexity.

[0071] In an embodiment, the question parsing module 104 may be used to perform uncertainty-aware reasoning. Using Monte Carlo (MC) sampling, multiple models can be obtained and model averaging can be performed for procedural generation. For example, a given Multiple outputs , and the output can be used The markers are estimated according to the following equation 13 The conditional distribution of : Equation (13) In an embodiment, multiple programs may be generated to exploit and handle the uncertainty learned from data in the question parsing module 104. Beam search (BS) may refer to a test-time decoding algorithm in neural machine translation, which may lack diversity. Methods to enhance diversity have been proposed. However, it can be Determine the diversity of BS. When the distribution is close to uniform, BS can generate different sequences; if Close to one-hot encoding, naively performing diversity may increase the difference between the training and testing process of BS and thus reduce the decoding performance. In addition, minimizing the negative log-likelihood during training may lead to overconfident predictions and uncalibrated uncertainties, which may not match the model error. Using variational dropout to consider model uncertainty can alleviate this problem.

[0072] To generate A program, at each time step of decoding The symbol module can be stored in bundle candidates, among which, represents the beam width and can be expressed by Sort the candidates, where At the next time step, all possible single-label extensions of these bundles can be considered, and one can choose The process can be repeated until the maximum time Then, the most likely A sequence.

[0073] However, due to the errors / uncertainties in the model, the log probability of the bundle may not correspond well to the probability that the program candidate is correct. Therefore, the uncertainty of the model can be taken into account to determine the most promising program. The probability of the bundle can be calibrated by averaging the estimates with a variance penalty according to the following equation 14: Equation (14) In an embodiment, a confidence score of a program candidate may be calculated. For example, in order to achieve aggregation of uncertainty quantification of various modules for confidence-based interactive reasoning, a confidence score based on predicted uncertainty may be used. The confidence score of the model may be evaluated according to the following equation 15: : Equation (15) In Equation 15, the MC method can be used to obtain Extract weights to estimate The value of , may represent a tuning hyperparameter for controlling the difference between confidence values ​​of programs with large differences. In an embodiment, the confidence score can be between 0 and 1, since the variance of the probability may be no greater than the corresponding expectation. In addition, the confidence score can increase as uncertainty (described by the variance) decreases, and the confidence score can be compared with the calibrated probability discussed above Positive correlation.

[0074] In order to avoid the requirement of real programs or programs from REINFORCE exploration for training concept learners, a data augmentation rule is used to select programs with high correct probability from the predicted programs of the problems in the training set for concept learning. programs are sorted. In an embodiment, the reconstruction loss and consistency loss of each program candidate can be used to evaluate the candidates. However, the reconstruction loss may be affected by error propagation because the prediction of the current time step may be based on the prediction of the previous time step. In contrast, the consistency loss uses the hidden states of the encoder and decoder to measure the coverage of the program to the problem, wherein the hidden state summarizes the information of the problem and the program. Therefore, another ranking of the program candidates can be obtained based on the consistency loss between the candidates and the problem. Then, when the two rankings reach a consensus on the top 1 programs, the problem can be selected. Then, the selected questions can be sorted again by the calibrated probability of the top 1 programs, and the data set can be enhanced with the top-ranked questions associated with the top 1 predicted programs.

[0075] Since the accuracy of the procedures that pass the data augmentation rules can be high (e.g., 99.88%+), the selected procedures can be used as pseudo-real procedures to learn the parameters of the scene perception module 102. In order to quantify the uncertainty in the scene perception module, variational dropout can be applied to the object features.

[0076] To determine the object concept (which can be called concept quantization), a neural operator that maps the object representation to an embedding can be used. Then, the learned concept vector The attribute is determined by the cosine distance between the embedding of the object and the embedding of the object. For example, the attribute belonging to the object can be estimated according to the following equation 16 Properties The concept of probability: Equation (16) In Equation 16, Representable length The L1 normalized vector of can represent neural operators, and Representable Concepts In addition, can represent the softmax function, and can represent the cosine distance. In addition, and is a scalar constant used to scale and shift the similarity values.

[0077] When training the scene perception module 102, a model can be sampled from the posterior , and the calculation is used for In addition, the optimization goal of the scene perception module 102 can be to maximize the final answer The correct probability is shown in Equation 17 below: Equation (17) In Equation 17, may represent a program execution module 108, and Can represent parameters The scene perception module 102 (for example, including ResNet-34 for extracting object features, neural operators for attributes, and concept vectors). In addition, Can express the answer, can represent images, and Can be expressed from the A candidate pseudo-real program for this problem.

[0078] For evaluation, sampling Models that can be calculated Embed , and the probability can be calculated for all concepts of each embedding . In addition, softmax can be used to normalize the probabilities for all concepts. Then, confidence scores can be calculated and used to average the embeddings. The calculated probabilities of the concepts are weighted. The weighted probabilities can be used for concept quantification, and the prediction of the attribute value can be the concept with the highest weighted probability.

[0079] The following equation 18 can be used to calculate the Properties The average confidence score : Equation (18) For example, according to the following Equation 19, the minimum value of the confidence scores for the attributes of all objects in the image may be used as the confidence score for the attribute of the image: Equation (19) Then, the product of the confidence scores of all attributes is used as the confidence score of the image according to Equation 20 below: Equation (20) Using the confidence scores of the images, the most predictions with the highest confidence scores may be selected as scene annotations for data augmentation and fine-tuning the question parsing module 104. In an embodiment, the dataset used to train the question parsing module 104 and the scene perception module 102 may be augmented with the programs with the highest confidence scores.

[0080] In an embodiment, the confidence score may be used during the execution of the program, or, for example, to trigger user interaction. After the concept quantification of the image and the program generation of questions about the image with confidence evaluation, the confidence may be evaluated for each step of the program execution. For example, the confidence score for the attribute involved in the program may be calculated according to the following equation 21 No. Confidence score of the function operation : Equation (21) In Equation 21, The first The set of objects involved in the operation is represented by The confidence score of the answer obtained by running the program can then be determined using Equation 22 below: Equation (22) In Formula 22, It can represent the number of operations involving attributes in the program. can be used to normalize the scores, and may represent a tuning parameter for controlling the relative importance of the perception and program's final confidence scores. In an embodiment, one may select To achieve the maximum area under the curve (AUC) score to predict the correctness of the answer on the training set using the answer confidence score. Using confidence assessment, the CBNS system 100 can request user interaction at any stage of the reasoning process. In addition, interactions can be assigned to the weakest links based on the demand and supply of available resources. Because scene perception can provide information for answering questions and concepts can be synthetic, an incorrect prediction of a concept for an object in an image may lead to a wrong answer. By correcting possible incorrect predictions of concepts with confidence scores below a threshold, the interaction can correct errors and avoid entanglement of errors from various modules. After the interaction, the current confidence score can be set to 1 to continue calculating the confidence score.

[0081] Conceptual accuracy can be important for purely symbolic reasoning, and even one incorrect prediction for an attribute of an object in an image can lead to an incorrect answer to a question involving that object. End-to-end and quasi-symbolic methods can arrive at the correct answer even when intermediate answers are incorrect (which can be referred to as consistency of reasoning), however this can introduce more confusion into the reasoning process and make the model uninterpretable. Therefore, embodiments can introduce user interaction based on confidence estimates to correct inaccurate predictions, which can be effective and efficient for consistent, transparent, and correct reasoning. See below for more information. FIG. 6A to FIG. 6C and 7A to 7C Examples of confidence assessments for allowing users to assist in inference and reasoning are shown.

[0082] FIG. 6A to FIG. 6C An example of an interactive visual question-answering system according to an embodiment of the present disclosure is shown. In particular, according to an embodiment of the present disclosure, Fig. 6A An example of input is shown, Figure 6B shows an example of the output of the scene perception module, and Figure 6C An example of selected programs and the output of running the modules is shown.

[0083] FIG. 6A to FIG. 6C The example involves the scene perception module 102 calculating a low confidence score for a concept. Fig. 6A As shown, the input image 600 may include bounding boxes and labels for objects 2-0 through 2-7.

[0084] like Figure 6BAs shown, scene perception module 102 may generate a prediction for object 2-3 that results in a low confidence score of 0.4960 for the concept “gray” corresponding to the color attribute of object 2-3, which may result in a low overall confidence score of 0.4943 for object 2-3, and a low overall confidence score of 0.4941 for scene perception of image 600. For example, the confidence scores for the concept “gray,” object 2-3, and image 600 may be below one or more confidence thresholds.

[0085] Therefore, the CBNS system 100 may request user interaction after the scene perception module 102 performs scene perception in order to correct the prediction and improve the confidence. The user interaction module 106 may obtain input from the user indicating the concept "green" for the color attribute of the object 2-3. Therefore, the CBNS system 100 may update the concept of the color attribute of the object 2-3 to "green" and may set the corresponding confidence score to 1 or close to 1.

[0086] like Figure 6C As shown, the selected program can be determined based on the input question "There is a large metal object to the left of the tiny green object, what is its shape?" When the program is run without user interaction, the result of the run may be wrong, which may be caused by the incorrect perception of object 2-3, which produces a low confidence score. However, the user interaction discussed above can allow the program to run successfully by providing the answer "sphere".

[0087] 7A to 7C An example of an interactive visual question-answering system according to an embodiment of the present disclosure is shown. In particular, according to an embodiment of the present disclosure, Fig. 7A An example of an input image is shown, Figure 7B shows an example of the output of the scene perception module, and Figure 7C An example of selected programs and the output of running the modules is shown.

[0088] 7A to 7C The example involves the scene perception module 102 calculating a low confidence score for a concept. Fig. 7A As shown, input image 700 may include bounding boxes and labels for objects 67 - 1 and 67 - 6 .

[0089] like Figure 7BAs shown, scene perception module 102 may generate a prediction for object 67-1 and generate a prediction for object 67-6, wherein the prediction for object 67-1 results in a low confidence score of 0.4580 for the concept "yellow" corresponding to the color attribute of object 67-1, and the prediction for object 67-6 results in a low confidence score of 0.9301 for the concept "cube" corresponding to the shape attribute of object 67-6. For example, the confidence scores for the concept "yellow" for object 67-1 and the concept "cube" for object 67-6 may be lower than one or more confidence thresholds, such as a threshold of 0.94.

[0090] Therefore, the CBNS system 100 may request user interaction after the scene perception module 102 performs scene perception in order to correct the prediction and improve the confidence. The user interaction module 106 may obtain input from the user indicating the concept "brown" for the color attribute of the object 67-1. In addition, the user interaction module 106 may obtain input from the user indicating the concept "cylinder" for the shape attribute of the object 67-6 and confirming the concept "red" for the color attribute of the object 67-6. Therefore, the CBNS system 100 may update the concept of the color attribute of the object 67-1 to "brown", update the concept of the shape attribute of the object 67-6 to "cylinder", maintain the concept of the color attribute of the object 67-6 to "red", and set the corresponding confidence scores of all these concepts to 1.

[0091] like Figure 7C As shown, the selected program may be determined based on the input question "Is there a blue metal cube that is the same size as the red metal cylinder?" When the program is run with user interaction triggered by the threshold of 0.94 discussed above, the result of the run may be successful by providing the answer "no". However, a confidence threshold below 0.94 may result in uncorrected errors.

[0092] In an embodiment, the confidence score estimated by the CBNS system 100 can be used to trigger actions other than requesting user interaction, such as triggering further analysis using a more powerful ML model. For example, the CBNS system 100 can use a less complex CBNS-VQA model for general queries, and then when the confidence of the predicted attribute is low, the CBNS system 100 can adapt a more powerful ML model to correct the predicted attribute in order to obtain robust performance and improved efficiency.

[0093] Fig. 8A and Figure 8B An example of interactive visual question answering according to an embodiment of the present disclosure is shown. In particular, Fig. 8A shows an example of an input image according to an embodiment of the present disclosure, Figure 8BAn example of program candidates according to an embodiment of the present disclosure is shown.

[0094] Fig. 8A and Figure 8B Examples involving scene perception with high confidence and question parsing with low confidence. Figure 8B As shown, program candidates (e.g., Candidate 1, Candidate 2, and Candidate 3) may be generated based on the input question "How many red balls are to the left of the large shiny block to the left of the small brown object?" Candidate 1 may be missing information at operations 10-14, and Candidate 3 may be missing information at operation 6. In addition, the confidence scores of all three program candidates may be below a threshold of 0.88. Therefore, the user interaction module 106 may be triggered to obtain input from the user. For example, the user may provide missing information, which may increase the confidence score of one or more program candidates to above a threshold. As another example, the user may determine the correct program from the program candidates instead of providing the real program, which may demonstrate the advantage of considering multiple programs. In some embodiments, the confidence assessment may provide a confidence score for each program candidate, and the confidence score for each operation in the program may provide information for debugging.

[0095] 9A to 9C An example of interactive visual question answering according to an embodiment of the present disclosure is shown. In particular, Fig. 9A An example of an input image is shown, Fig. 9B shows an example of the output of the scene perception module, Fig. 9C An example of program candidates including the selected program is shown.

[0096] 9A to 9C This example involves a low confidence score calculated by the scene perception module 102 for a concept and a low confidence score calculated by the question parsing module 104 for a program candidate. Fig. 9A As shown, input image 900 may include bounding boxes and labels for objects 8727-1 and 8727-0.

[0097] like Fig. 9B As shown, the scene perception module 102 may generate a prediction for object 8727-1 and a prediction for object 8727-0, wherein the prediction for object 8727-1 results in a low confidence score of 0.5815 for the concept "cylinder" corresponding to the shape attribute of object 8727-1, and the prediction for object 8727-0 results in a confidence score of 0.9543 for the concept "cube" corresponding to the shape attribute of object 8727-0. For example, the confidence score for the concept "cylinder" corresponding to the shape attribute of object 8727-1 may be lower than one or more confidence thresholds, such as a threshold of 0.94.

[0098] Therefore, the CBNS system 100 may request user interaction after the scene perception module 102 performs scene perception in order to correct the prediction and improve the confidence. The user interaction module 106 may obtain input from the user indicating the concept "cube" for the shape attribute of the object 8727-1. In addition, the user interaction module 106 may obtain input from the user confirming the concept "cube" for the shape attribute of the object 8727-0. Therefore, the CBNS system 100 may update the concept of the shape attribute of the object 8727-1 to "cube", and may confirm the concept of the shape attribute of the object 8727-0 as "cube", and may set the corresponding confidence scores of the two concepts to 1.

[0099] like Fig. 9C As shown, program candidates (e.g., candidate 1, candidate 2, and candidate 3) may be determined based on the input question "There is a red cylinder in front of the cylinder to the left of the cylinder to the right of the tiny red matte cylinder. How big is it?". The confidence score of one or more program candidates may be below one or more confidence thresholds. Therefore, the CBNS system 100 may request user interaction after the question parsing module 102 performs question parsing in order to correct the program candidates and improve the confidence. The user interaction module 106 may obtain input from the user, which indicates that candidate 2 lacks key information in the question and indicates that candidate 3 includes meaningless operation 11. Therefore, candidate 1 may be selected as the selected program.

[0100] Fig.10 is a flowchart showing a method for performing visual question answering according to an embodiment of the present disclosure. In an embodiment, the CBNS system 100 and any element included therein and at least one of any other elements described above may be used to perform the visual question answering. Fig.10 Method 1000.

[0101] like Fig.10 As shown, in operation S1001, method 1000 may include obtaining an image and a question corresponding to the image.

[0102] In operation S1002, the method 1000 may further include generating a plurality of feature predictions about at least one object included in the image by providing the image to an artificial intelligence (AI) scene perception model. In an embodiment, the AI ​​scene perception model may correspond to the scene perception module 102 discussed above, and the plurality of feature predictions may correspond to at least one of the attributes and concepts discussed above.

[0103] In operation S1003, method 1000 may further include generating a plurality of symbolic programs and a plurality of program confidence scores by providing the question to an AI question parsing model, wherein each symbolic program in the plurality of symbolic programs includes a set of logical operations and is associated with a corresponding program confidence score in the plurality of program confidence scores. In an embodiment, the AI ​​question parsing model may correspond to the question parsing module 104 discussed above.

[0104] In operation S1004 , method 1000 may further include selecting a symbolic program associated with a highest program confidence score among the plurality of program confidence scores.

[0105] In operation S1005 , the method 1000 may further include running the selected symbolic program by performing a set of logical operations included in the selected symbolic program on the plurality of feature predictions.

[0106] In operation S1006 , the method 1000 may further include determining a natural language answer to the question based on a result of the set of logical operations.

[0107] In an embodiment, an image and a question may be received from a user, and the method may further include providing a natural language answer to the user as a response to the question.

[0108] In an embodiment, multiple feature predictions may be associated with multiple feature confidence scores generated by the AI ​​scene perception model.

[0109] In an embodiment, a set of logical operations included in a selected symbolic program may be associated with a plurality of operation confidence scores, and each logical operation in the set of logical operations may be associated with an operation confidence score from the plurality of operation confidence scores, the operation confidence score being determined based on a program confidence score and at least one of a plurality of feature confidence scores.

[0110] In an embodiment, based on at least one confidence score among multiple program confidence scores, multiple feature confidence scores, and multiple operation confidence scores being lower than a threshold: method 1000 may also include obtaining user input corresponding to at least one confidence score; and adjusting at least one confidence score based on the user input.

[0111] In an embodiment, method 1000 may further include determining an answer confidence score corresponding to the natural language answer based on the plurality of operational confidence scores.

[0112] In an embodiment, method 1000 may further include generating enhanced training data based on multiple symbolic programs; and training the AI ​​scene perception model based on the enhanced training data.

[0113] In an embodiment, the step of generating enhanced training data may include: generating a first ranking of multiple symbolic programs based on multiple program confidence scores; generating a second ranking of multiple symbolic programs based on multiple consistency losses between the multiple program confidence scores and the problem; selecting a subset of the multiple symbolic programs based on the first ranking and the second ranking; and generating enhanced training data based on the output of the subset of the multiple symbolic programs.

[0114] Thus, embodiments may provide improved answer accuracy. For example, the question parser architecture, confidence assessment, and training methods provided by embodiments may significantly improve question parser and overall system performance. Embodiments may also provide uncertainty quantification and confidence assessment: for example, confidence scores according to embodiments may effectively predict the correctness of inference results. Embodiments may also provide reduced computational costs: for example, by applying data enhancement methods according to embodiments, data- and computationally intensive REINFORCE methods may be avoided while achieving similar or even increased performance based on limited training data.

[0115] Thus, embodiments may provide improved systems for performing tasks such as VQA or other information retrieval tasks. For example, for safety-critical applications, confidence assessments provided by embodiments may be used to determine whether to take a machine-provided action. For error analysis of complex processes, confidence assessments for each step provided by embodiments may be used to track errors. For new data acquisition, uncertainty quantification provided by embodiments may be used to determine areas that are not well represented by the current data set. For decision making, multiple reasoning paths provided by embodiments may be used to select the most credible solution. For user interaction, where confidence assessment is enabled according to embodiments, the user may efficiently provide limited corrections to the system based on the estimated confidence. In addition, embodiments may be applied to devices such as augmented reality or smart glasses to help visually impaired patients better "visualize" the environment through question-and-answer methods.

[0116] Fig.11 is a diagram of an apparatus for performing an interactive CBNS VQA task according to an embodiment. Fig.11 Included are a user device 1110, a server 1120, and a communication network 1130. The user device 1110 and the server 1120 may be interconnected via a wired connection, a wireless connection, or a combination of a wired connection and a wireless connection.

[0117] The user device 1110 includes one or more devices (e.g., a processor 1111 and a data storage 1112) configured to acquire images corresponding to a search query. For example, the user device 1110 may include a computing device (e.g., a desktop computer, a laptop computer, a tablet computer, a handheld computer, a smart speaker, a server, etc.), a mobile phone (e.g., a smart phone, a wireless phone, etc.), a camera device, a wearable device (e.g., a pair of smart glasses, a smart watch, etc.), a home appliance (e.g., a robot vacuum cleaner, a smart refrigerator, etc.), or a similar device. The data storage 1112 of the user device 1110 may include one or more of the scene perception module 102, the question parsing module 104, the user interaction module 106, and the program execution module 108. Optionally, the user device 1110 stores one or more of the scene perception module 102, the question parsing module 104, the user interaction module 106, and the program execution module 108, and vice versa.

[0118] The server 1120 includes one or more devices (e.g., a processor 1121 and a data storage 1122) configured as one or more of the scene perception module 102, the question parsing module 104, the user interaction module 106, and the program execution module 108. The data storage 1122 of the server 1120 may include one or more of the scene perception module 102, the question parsing module 104, the user interaction module 106, and the program execution module 108. Optionally, the user device 1110 stores one or more of the scene perception module 102, the question parsing module 104, the user interaction module 106, and the program execution module 108, and vice versa.

[0119] The communication network 1130 includes one or more wired networks and / or wireless networks. For example, the network 1130 may include a cellular network, a public land mobile network (PLMN), a local area network (LAN), a wide area network (WAN), a metropolitan area network (MAN), a telephone network (e.g., a public switched telephone network (PSTN)), a private network, an ad hoc network, an intranet, the Internet, a fiber-based network, etc., and / or a combination of these or other types of networks.

[0120] Will Fig.11 The number and arrangement of devices and networks shown in the FIG. are provided as examples. In practice, there may be additional devices and / or networks, fewer devices and / or networks, different devices and / or networks, or different Fig.11 The devices and / or networks shown in FIG. 1 are arranged differently from the devices and / or networks shown in FIG. 1 . Fig.11 Two or more of the devices shown in the figure may be implemented in a single device, or Fig.11A single device shown in the may be implemented as multiple distributed devices. Additionally or alternatively, a group of devices (eg, one or more devices) may perform one or more functions described as being performed by another group of devices.

[0121] Fig.12 According to the embodiment Fig.11 An illustration of components of one or more electronic devices. Fig.12 The electronic device 1200 in the figure may correspond to the user device 1110 and / or the server 1120.

[0122] Fig.12 This is for illustration only, and other embodiments of the electronic device 1200 may be used without departing from the scope of the present disclosure. For example, the electronic device 1200 may correspond to a client device or a server.

[0123] The electronic device 1200 includes a bus 1210 , a processor 1220 , a memory 1230 , an interface 1240 , and a display 1250 .

[0124] The bus 1210 includes a circuit for mutually connecting the components 1220 to 1250. The bus 1210 serves as a communication system for transmitting data between the components 1220 to 1250 or between electronic devices.

[0125] The processor 1220 includes one or more of a central processing unit (CPU), a graphics processor unit (GPU), an accelerated processing unit (APU), an integrated many-core (MIC), a field programmable gate array (FPGA), or a digital signal processor (DSP). The processor 1220 is capable of performing control of any one or any combination of other components of the electronic device 1200, and / or performing operations related to communication or data processing. For example, the processor 1220 may execute methods 200, 400B, and 1000, as well as methods corresponding to frameworks 400A, 500A, and 500B, such as Figure 2 , FIG. 4A to FIG. 4B , FIG. 5A to FIG. 5B and Fig.10 The processor 1220 runs one or more programs stored in the memory 1230 .

[0126] The memory 1230 may include a volatile memory and / or a non-volatile memory. The memory 1230 stores information (such as one or more of commands, data, programs (one or more instructions), applications 1234, etc.) related to at least one other component of the electronic device 1200 and used to drive and control the electronic device 1200. For example, the command and / or data may formulate an operating system (OS) 1232. The information stored in the memory 1230 may be executed by the processor 1220.

[0127] Application 1234 includes the embodiments discussed above. These functions may be performed by a single application or multiple applications, each of which performs one or more of these functions. For example, application 1234 may include an artificial intelligence (AI) model for performing methods 200, 400B, and 1000 and methods corresponding to frameworks 400A, 500A, and 500B, such as Figure 2 , FIG. 4A to FIG. 4B , FIG. 5A to FIG. 5B and Fig.10 shown.

[0128] The display 1250 includes, for example, a liquid crystal display (LCD), a light emitting diode (LED) display, an organic light emitting diode (OLED) display, a quantum dot light emitting diode (QLED) display, a micro-electromechanical system (MEMS) display, or an electronic paper display. The display 1250 may also be a depth perception display (such as a multi-focal display). The display 1250 is capable of presenting, for example, various contents (such as text, images, videos, icons, and symbols).

[0129] Interface 1240 includes an input / output (I / O) interface 1242, a communication interface 1244, and / or one or more sensors 1246. I / O interface 1242 serves as an interface that can transmit commands and / or data between a user and / or other external devices and other components of electronic device 1200, for example.

[0130] The communication interface 1244 may enable communication between the electronic device 1200 and other external devices via a wired connection, a wireless connection, or a combination of a wired connection and a wireless connection. The communication interface 1244 may allow the electronic device 1200 to receive information from another device and / or provide information to another device. For example, the communication interface 1244 may include an Ethernet interface, an optical interface, a coaxial interface, an infrared interface, a radio frequency (RF) interface, a universal serial bus (USB) interface, a Wi-Fi interface, a cellular network interface, etc. The communication interface 1244 may receive video and / or video frames from an external device (such as a server).

[0131] The sensor 1246 of the interface 1240 can measure a physical quantity or detect the activation state of the electronic device 1200, and convert the measured or detected information into an electrical signal. For example, the sensor 1246 may include one or more cameras or other image sensors for capturing an image of a scene. The sensor 1246 may also include any one or any combination of a microphone, a keyboard, a mouse, and one or more buttons for touch input. The sensor 1246 may also include an inertial measurement unit. In addition, the sensor 1246 may include a control circuit for controlling at least one of the sensors included herein. Any of these sensors 1246 may be located within the electronic device 1200 or incorporated into the electronic device 1200. The sensor 1246 may receive a text and / or voice signal including one or more queries.

[0132] The interactive CBNS VQA process may be written as a computer executable program or instructions that may be stored in a medium.

[0133] The medium may store computer executable programs or instructions continuously, or temporarily store computer executable programs or instructions to run or download. In addition, the medium may be any of a variety of recording media or storage media in which a single or multiple hardware is combined, and the medium is not limited to a medium directly connected to the electronic device 1200, but may be distributed on a network. Examples of media include magnetic media (such as hard disks, floppy disks, and tapes) configured to store program instructions, optical recording media (such as CD-ROMs and DVDs), magneto-optical media (such as optical magnetic floppy disks), and ROM, RAM, and flash memory. Other examples of media include recording media and storage media managed by application stores that distribute applications or by websites, servers, etc. that supply or distribute various other types of software.

[0134] The interactive CBNS VQA process may be provided in the form of downloadable software. The computer program product may include a product in the form of a software program (e.g., a downloadable application) that is electronically distributed through a manufacturer or an electronic marketplace. For electronic distribution, at least a portion of the software program may be stored in a storage medium or may be temporarily generated. In this case, the storage medium may be a storage medium of a server or server 106.

[0135] The foregoing disclosure provides illustration and description, but is not intended to be exhaustive or to limit the implementations to the precise forms disclosed.Modifications and variations are possible in light of the above disclosure or may be acquired from practice of the implementations.

[0136] It is apparent that the systems and / or methods described herein may be implemented in various forms of hardware, firmware, or a combination of hardware and software. The actual dedicated control hardware or software code used to implement these systems and / or methods does not limit the implementation. Therefore, the operation and behavior of the systems and / or methods are described herein without reference to specific software code - it should be understood that software and hardware can be designed to implement the systems and / or methods based on the description herein.

[0137] Even though particular combinations of features are recited in the claims and / or disclosed in the specification, these combinations are not intended to limit the disclosure that may be implemented. In fact, many of these features may be combined in ways not specifically recited in the claims and / or disclosed in the specification. Although each dependent claim listed below may only be directly dependent on one claim, the disclosure that may be implemented includes each dependent claim in combination with every other claim in the claim set.

[0138] The model related to the above-mentioned neural network can be implemented via a software module. When the model is implemented via a software module (eg, a program module including instructions), the model can be stored in a computer-readable recording medium.

[0139] In addition, the model can be integrated in the form of a hardware chip to become a part of the above-mentioned electronic device 1200. For example, the model can be manufactured in the form of a dedicated hardware chip for artificial intelligence, or can be manufactured as a part of an existing general-purpose processor (e.g., a CPU or an application processor) or a graphics-specific processor (e.g., a GPU).

[0140] In addition, the model may be provided in the form of downloadable software. The computer program product may include a product in the form of a software program (e.g., a downloadable application) that is electronically distributed through a manufacturer or an electronic market. For electronic distribution, at least a portion of the software program may be stored in a storage medium or may be temporarily generated. In this case, the storage medium may be a server of the manufacturer or the electronic market, or a storage medium of a relay server.

[0141] Although embodiments of the present disclosure have been described with reference to the drawings, it will be understood by those skilled in the art that various changes in form and details may be made therein without departing from the spirit and scope defined by the following claims.

Claims

1. A method for performing visual question answering (VQA), the method comprising: obtaining an image and a question corresponding to the image; generating a plurality of feature predictions about at least one object included in the image by providing the image to an artificial intelligence (AI) scene perception model; generating a plurality of symbolic programs and a plurality of program confidence scores by providing the problem to an AI problem parsing model, wherein each symbolic program in the plurality of symbolic programs includes a set of logical operations and is associated with a corresponding program confidence score in the plurality of program confidence scores; selecting a symbolic program associated with a highest program confidence score among the plurality of program confidence scores; executing the selected symbolic program by performing a set of logical operations included in the selected symbolic program on the plurality of feature predictions; and A natural language answer to the question is determined based on the results of the set of logical operations.

2. The method of claim 1, wherein: receiving the image and the question from a user, and The method further includes providing the natural language answer to the user as a response to the question.

3. The method of claim 1, wherein: The multiple feature predictions are associated with multiple feature confidence scores generated by the AI ​​scene perception model.

4. The method of claim 3, wherein: The set of logical operations included in the selected symbolic procedure is associated with a plurality of operation confidence scores, and Each logic operation in the logic operation set is associated with an operation confidence score from the plurality of operation confidence scores, and the operation confidence score is determined based on the program confidence score and at least one of the plurality of feature confidence scores.

5. The method of claim 4, further comprising: Based on at least one confidence score of the plurality of program confidence scores, the plurality of feature confidence scores, and the plurality of operation confidence scores being below a threshold: obtaining user input corresponding to the at least one confidence score; and The at least one confidence score is adjusted based on the user input.

6. The method of claim 4, further comprising: An answer confidence score corresponding to the natural language answer is determined based on the plurality of operation confidence scores.

7. The method of claim 1, further comprising: generating enhanced training data based on the plurality of symbolic procedures; as well as The AI ​​scene perception model is trained based on the enhanced training data.

8. The method of claim 7, wherein: The step of generating the enhanced training data comprises: generating a first ranking of the plurality of symbolic programs based on the plurality of program confidence scores; generating a second ranking of the plurality of symbolic programs based on a plurality of consistency losses between the plurality of program confidence scores and the questions; selecting a subset of the plurality of symbolic programs based on the first ranking and the second ranking; and The augmented training data is generated based on outputs of a subset of the plurality of symbolic procedures.

9. A device for performing visual question answering (VQA), the device comprising: a memory configured to store instructions; as well as At least one processor is configured to execute the instructions to perform the following operations: obtaining an image and a question corresponding to the image; generating a plurality of feature predictions about at least one object included in the image by providing the image to an artificial intelligence (AI) scene perception model; generating a plurality of symbolic programs and a plurality of program confidence scores by providing the problem to an AI problem parsing model, wherein each symbolic program in the plurality of symbolic programs includes a set of logical operations and is associated with a corresponding program confidence score in the plurality of program confidence scores; selecting a symbolic program associated with a highest program confidence score among the plurality of program confidence scores; executing the selected symbolic program by performing a set of logical operations included in the selected symbolic program on the plurality of feature predictions; and A natural language answer to the question is determined based on the results of the set of logical operations.

10. The device of claim 9, wherein: The multiple feature predictions are associated with multiple feature confidence scores generated by the AI ​​scene perception model.

11. The device of claim 10, wherein: The set of logical operations included in the selected symbolic procedure is associated with a plurality of operation confidence scores, and Each logic operation in the logic operation set is associated with an operation confidence score from the plurality of operation confidence scores, and the operation confidence score is determined based on the program confidence score and at least one of the plurality of feature confidence scores.

12. The device of claim 11, wherein: The at least one processor is further configured to execute the instructions to perform the following operations: Based on at least one confidence score of the plurality of program confidence scores, the plurality of feature confidence scores, and the plurality of operation confidence scores being below a threshold: obtaining user input corresponding to the at least one confidence score; and The at least one confidence score is adjusted based on the user input.

13. The apparatus of claim 9, wherein: The at least one processor is further configured to execute the instructions to perform the following operations: generating enhanced training data based on the plurality of symbolic procedures; and The AI ​​scene perception model is trained based on the enhanced training data.

14. The device of claim 13, wherein: To generate the enhanced training data, the at least one processor is further configured to: generating a first ranking of the plurality of symbolic programs based on the plurality of program confidence scores; generating a second ranking of the plurality of symbolic programs based on a plurality of consistency losses between the plurality of program confidence scores and the questions; selecting a subset of the plurality of symbolic programs based on the first ranking and the second ranking; as well as The augmented training data is generated based on outputs of a subset of the plurality of symbolic procedures.

15. A computer-readable medium storing instructions, wherein: When the instructions are executed by at least one processor of an apparatus for performing visual question answering (VQA), the at least one processor is caused to perform the method according to any one of claims 1 to 8.

Citation Information

Cited By

  • Nerve symbol learning-based natural language problem programmed analysis method

    CN121859889A