Device, data structure and computer-implemented method for determining a quantity for evaluating the quality of a generative, in particular multi-modal, basic model, for example for optical inspection, for identifying sounds or technical parts, for natural language processing, for program generation, or for marking data
A method to evaluate generative models by defining relationships between answers enhances their quality assessment, ensuring reliable performance across domains by using confidence levels and reward systems.
Patent Information
- Authority / Receiving Office
- DE · DE
- Patent Type
- Applications
- Current Assignee / Owner
- ROBERT BOSCH GMBH
- Filing Date
- 2024-11-04
- Publication Date
- 2026-05-07
AI Technical Summary
Existing generative models, particularly large language models (LLMs), lack a reliable method to assess their quality across different domains, leading to potential unreliable performance in applications.
A computer-implemented method determines a metric for evaluating the quality of generative models by defining relationships between answers to multiple questions, using a confidence level or reward system based on these relationships, and blocking or training the model based on these metrics.
This approach provides a reliable assessment of model quality, ensuring its effective use only in domains where it is reliable, thereby improving performance and preventing unreliable applications.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
State of the art
[0001] The invention relates to a device, a data structure and a computer-implemented method for determining a quantity for evaluating the quality of a generative, in particular multi-modal, basic model, for example for optical inspection, for identifying sounds or technical parts, for natural language processing, for program generation, or for marking data.
[0002] Generative basic models, such as large language models (LLMs), particularly multimodal generative language models, are capable of answering questions on diverse topics, meaning across different domains. Examples of these domains include optical inspection, noise or component identification, natural language processing, program generation, and data labeling. Disclosure of the invention
[0003] A computer-implemented method for determining a metric to assess the quality of a generative, particularly multi-modal, basic model—for example, for optical inspection, noise or technical part identification, natural language processing, program generation, or data labeling—provides that the generative basic model is used to determine a first answer to a first question, and then a second answer to a second question. A relationship between the first and second answers is defined, and the metric is determined based on these answers and the relationship. The method provides a metric that allows the basic model to be automatically trained or tested.
[0004] For example, a confidence level is determined that the first and second answers satisfy the relation. This confidence level is determined, for example, by the fulfillment of various individual relations for each relation.
[0005] The confidence level is, for example, output to a user of the basic model. This informs the user about the goodness of fit of the basic model with respect to the relation.
[0006] The use of the basic model is blocked, for example, for an application area that includes the first question and / or the second question, if it is determined that the confidence falls below a threshold. This prevents the use of the basic model in areas where it is unreliable.
[0007] For example, a reward is determined if the first and second answers satisfy the relation, with the basic model being trained depending on the reward. The reward enables training that is dependent on the reward.
[0008] The basic model includes, for example, an artificial neural network with weights, where the weights are determined depending on the reward.
[0009] The relation can be defined differently, preferably as a metamorphic relation.
[0010] The first answer, for example, comprises a set of individual answers to the first question, where the second answer comprises a set of individual answers to the second question, where the relation is given that the first answer and the second answer are disjoint sets of individual answers, or that the first answer is a subset of the second answer.
[0011] The first question, for example, asks for a first program to complete a task on a first domain, the second question asks for a second program to complete the task on a second domain, where the first domain includes the second domain, and the relation is given that the first program and the second program solve the task on the second domain with the same result.
[0012] For optical inspection in particular, a first digital image representing an object is provided, the first digital image is transformed into a second digital image representing the object, the first question is about the position of the object in the first digital image, the second question is about the position of the object in the second digital image, specifying the relation how the first digital image is transformed to the second digital image.
[0013] In particular, when dealing with a noise, a first audio sequence containing the noise is provided, a second audio sequence containing the noise is provided, the first question asks for a time at which the noise occurs in the first audio sequence, the second question asks for a time at which the noise occurs in the second audio sequence, whereby the relation of the times is specified.
[0014] In particular, for testing programs automatically generated with the basic model for a computer, the first question is, for example, about the result of a concatenation of a first program with a second program to complete a task on an input, whereby the second question is about the result of a third program to complete the task on the input, whereby the relation is given that the third program solves the task with the same result as the concatenation of the first program with the second program.
[0015] It may be provided that a third answer to a third question is determined, for which the relation is expected to be that the third answer includes the first answer and the second answer, or that the third answer is identical to the union of the first answer and the second answer.
[0016] A device for determining a quantity for evaluating the quality of a generative, in particular multi-modal, basic model, for example for optical inspection, for identifying sounds or technical parts, for natural language processing, for program generation, or for marking data, provides that the device comprises at least one processor and at least one memory, wherein the at least one memory comprises instructions, the execution of which by the at least one processor on the device carries out the method.
[0017] A computer program may be provided, wherein the computer program comprises instructions executable by a computer, the execution of which by the computer on the computer causes the procedure to take place.
[0018] A data structure for determining a metric for evaluating the quality of a generative, in particular multi-modal, basic model, for example for optical inspection, for identifying sounds or technical parts, for natural language processing, for program generation, or for labeling data, provides that the data structure has at least one data field for a generative basic model, at least one data field for a first answer to a first question determined with the generative basic model, at least one data field for a second answer to a second question determined with the generative basic model, at least one data field for a relation between the first answer and the second answer, and at least one data field for the metric determined depending on the answers and the relation.
[0019] Further advantageous embodiments can be found in the following description and the drawing. The drawing shows: Fig. 1 a schematic representation of a device 100 for determining a quantity for evaluating the quality of a generative, in particular multi-modal, basic model, Fig. 2. A flowchart showing the steps of a procedure for determining the size for evaluating the quality of the generative, in particular multi-modal, basic model. Fig. 3 a schematic representation of a data structure for determining the size for evaluating the quality of the generative, in particular multi-modal, basic model.
[0020] The device 100 comprises at least one processor 102 and at least one memory 104. The at least one memory 104 stores instructions executable by the at least one processor 102, the execution of which by the at least one processor 102 on the device 100 carries out a method for determining the size for evaluating the quality of a generative, in particular multi-modal, basic model.
[0021] The device 100 may include an interface 106. The interface 106 includes, for example, a human-machine interface or a machine-machine interface.
[0022] Interface 106 is used, for example, for inputting questions and outputting answers. Interface 106 can also be configured to output hints related to the answers.
[0023] An input can be, for example, text or a file. The file can be, for example, text, a digital image, an audio signal, and / or a program, i.e., computer program code. The multimodal basic model is, for example, designed to process the input. An output can be, for example, text or a file. The file can be, for example, text, a digital image, an audio signal, and / or a program, i.e., computer program code. The multimodal basic model is, for example, designed to output the output.
[0024] The input and output are data, for example a natural language description and the program.
[0025] It may be provided that at least one memory module 104 comprises the basic model.
[0026] The basic model includes, for example, an artificial neural network with weights.
[0027] The generative basic model is used, for example, for optical inspection, for the identification of sounds or technical parts, for natural language processing, for program generation and / or for data labeling.
[0028] The generative basic model is, for example, trained to perform tasks in the domain of optical inspection, the identification of sounds or technical parts, natural language processing, program generation and / or data labeling.
[0029] The generative multi-modal basic model is designed to answer questions on at least two of the domains.
[0030] The basic model is, for example, a large language model (LLM).
[0031] In this example, the LLM is already trained to answer questions. This means the LLM is trained to output an answer to a question posed to it. The answer to the question is determined using the generative basic model.
[0032] The basic model is trained, for example, to take predefined quality characteristics into account when generating its output. The training is done, for example, unsupervised.
[0033] An example of training an LLM on the program generation domain is for a programming task in RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning (arXiv:2410.02089v1).
[0034] The procedure for determining the size used to evaluate the quality of the generative, and especially multimodal, basic model is based on questions for which an expected relation between the answers is specified. This relation is, for example, a metamorphic relation. The relation is described, for example, in a formal language.
[0035] The procedure is described using the example of a first question and a second question, and a relation between the first answer of the generative base model to the first question and the second answer of the generative base model to the second question. The example optionally describes the possibility of providing a third question instead of the relation between the first and second answers, and the relation between the first and second answers and a third answer of the base model to the third question.
[0036] The procedure is not limited to three questions and three answers, or to the relationship between two or three answers. More than three questions, more than three answers, and more than one relationship between these more than three answers are possible. Different relationships between the answers are possible. Different relationships between subsets of the answers are possible. At least one relationship between the answers and at least one relationship between subsets of the answers are possible.
[0037] According to an example where it is known that the correct answer to the first question comprises a first set of individual answers, and the correct answer to the second question comprises a set of individual answers disjoint from the first set, the relation is given, for example, that the first answer of the basic model and the second answer of the basic model are disjoint sets of individual answers.
[0038] According to an example where it is known that the correct answer to the first question comprises a first set of individual answers, and the correct answer to the second question comprises a set of individual answers encompassing the first set, the relation is given, for example, that the first answer of the basic model is a subset of individual answers from the set of individual answers in the second answer of the basic model.
[0039] To determine a program for completing a task on a first domain, the first question may ask for a first program capable of completing the task on that first domain, and the second question asks for a second program capable of completing the task on a second domain encompassing the first domain. To verify this, the requirement may be specified that the first program and the second program must solve the task on the second domain with the same result.
[0040] It may be stipulated that the third question is formulated in such a way that a correct answer to the third question encompasses both the first and second answers. This means that the third answer is defined as encompassing both the first and second answers.
[0041] It may be stipulated that the third question is formulated such that a correct answer to the third question is identical to the union of the first and second answers. This means that the third answer is defined as being identical to the union of the first and second answers.
[0042] To determine the first program, the initial question might ask for the result of concatenating the first program with a second program to complete a task based on a given input. The second question, for example, might ask for the result of a third program to complete the same task based on the input. For instance, the requirement might be that the third program solves the task with the same result as concatenating the first program with the second.
[0043] To determine the position of an object in a first digital image, the first question might ask for the object's position in that image, and the second question asks for its position in a second digital image. For example, the object is located in a first position in the first image and in a second position in the second image. The second position is located a certain distance away from the first position in a specific direction. The first answer represents a predicted first position of the object. The second answer represents a predicted second position of the object. For verification, the relationship is defined, for example, that the predicted second position in the second answer to the second question differs from the predicted first position in the first answer to the first question by a certain distance in that direction.
[0044] For example, in audio signals, a first audio sequence containing the noise and a second audio sequence containing the noise might be provided. The first question might ask, for example, at what time the noise occurs in the first audio sequence, and the second question at what time the noise occurs in the second audio sequence. For verification purposes, the relationship between these times might be specified.
[0045] Determining the relationship depends, for example, on the initial question. This determination can occur through an interaction between a specially trained computer program, such as a chatbot, and a user. The computer program includes, for example, the generative basic model.
[0046] The computer program is, for example, trained to capture the first question from a user.
[0047] The computer program is, for example, trained to ask the user about the relation in a follow-up question. The relation is, for example, a metamorphic relation.
[0048] The computer program is, for example, trained to determine the second question based on the relation and the first question. Optionally, the computer program is trained to determine the third question based on the relation and the first question.
[0049] The computer program is, for example, trained to ask the first and second questions of the generative basic model. Optionally, the computer program is trained to ask the generative basic model a third question.
[0050] The computer program is trained, for example, to determine in a check whether the answers to the questions asked satisfy the relation or not.
[0051] The computer program is, for example, trained to output the result of the check to the user. The computer program is, for example, trained to output the result of the check and the answers to the user. The computer program is, for example, trained to output the answers to the user if the check is successful, and not to output the answers to the user otherwise. The check is successful, for example, if the answers satisfy the relation.
[0052] For multiple relations, the check is successful, for example, if all relations are satisfied.
[0053] A first example output of the computer program is “My answer is [...], but I am not sure in this case because my internal check failed”, where [...] represents at least one of the answers.
[0054] A second exemplary output of the computer program is “My answer is [...], and I am quite sure because my internal check works,” where [...] represents at least one of the answers.
[0055] A first exemplary workflow of the computer program is described using the example of a query requesting a list of presidents of country X. In this example, four presidents of country X are known: President A, President B, President C, and President D. The last two presidents are President A and President D. The first and second user questions are entered into the computer program by a user. The response query is generated by the computer program, for example, from a template for a response query to a query requesting a list. The first, second, and third questions to the base model are generated by the computer program based on the first and second questions to the base model, for example, from a template for queries requesting a list. The query requesting the list is identified, for example, by a classifier that assigns the question to the templates.In this example, the template for the question is associated with the relation expected from the answers to the questions posed to the base model. The questions and answers are output to the user by the computer program. The computer program queries the base model with the respective question and outputs the corresponding answer from the base model. First user question: Name all the presidents of country X. First question to the basic model: Name all the presidents of country X. First answer: President A, President B, President C, President D. Follow-up question: I see you're looking for a list. Give me an example of a question that would help me find a shorter list. Second user question: Name the last two presidents of country X. Second question to the base model: Name the last two presidents of country X. Second answer: President A, President D. Third question regarding the basic model: Name all presidents of country X, excluding the last two presidents of country X. Third answer: President B, President C.
[0056] Relation: {second answer, third answer} == first answer. This means that the expected relation is that the union of the list of presidents from the second answer and the list of presidents from the third answer is equal to the list of presidents from the first answer.
[0057] A second exemplary execution of the computer program is described using the example of a question about a program to check whether a graph is a tree. In this example, a graph is a tree if it is connected and has no cycles.
[0058] In this example, a user enters questions into a computer program. The computer program generates the corresponding queries, for example, from a template for queries in response to a question about a program. The computer program then outputs the queries and answers to the user. It queries the base model with the respective question and outputs the corresponding answer from the base model. The question about the program is recognized, for example, by a classifier that assigns the question to the template. In this example, the template is associated with the relation expected from the answers to the questions posed to the base model. First user question: Write me a checker to determine if a graph is a tree. First question for the basic model: Write me a checker to determine if a graph is a tree.
[0059] Follow-up question: Give me a list of subtasks into which I can divide this task. First answer: Program 1 - Checker to see if a graph is a tree Second user question: 1. Write a checker to see if a graph is connected. 2. Write a checker to see if a graph has a cycle. Second question regarding the basic model: Write a checker to determine if a graph is connected. Second answer: Program 2 - Checker to see if a graph is connected Third question regarding the basic model: Write a checker to determine if a graph has a cycle. Third answer: Program 3 - Checker to see if a graph has a cycle
[0060] In this example, the relation refers to the result that the respective programs deliver for a test graph: Result 1: Application of program 1 to the test graph Result 2: Application of program 2 to the test graph Result 3: Application of program 3 to the test graph
[0061] In the example, the result 1 is true if the graph is a tree, and false otherwise.
[0062] In the example, result 2 is true if the graph is connected, and false otherwise.
[0063] In the example, the result 3 is true if the graph has no cycle, and false otherwise.
[0064] Relation: Result 2 & Result 3 == Result 1. This means the expected relation is that result 1 is true if result 2 and result 3 are both true, and result 1 is false otherwise.
[0065] In one example, a message is displayed to the user along with the answers. The message is determined differently, for example, depending on whether the answers fulfill the expected relationship or not.
[0066] Example of answers that do not meet the relation: "My result is: "Program 1". Here is also the code for the sub-steps: "Program 2, Program 3".
[0067] Unfortunately, I'm not sure, because I tested Program 1 for you using a test graph. The test revealed a difference between the result from Program 1 and the results from the sub-steps. Therefore, be careful if you continue using the code.
[0068] The verification of whether the questions fulfill the expected relation is used to determine a metric for evaluating the quality of the generative, especially multi-modal, basic model.
[0069] The procedure determines the metric used to assess the quality.
[0070] The procedure includes step 202.
[0071] In step 202, the first answer to the first question is determined using the generative basic model.
[0072] The procedure includes step 204.
[0073] In step 204, the second answer to the second question is determined using the generative basic model.
[0074] The procedure optionally includes step 206. Step 206 is executed, for example, when a relationship between three answers is to be checked.
[0075] In step 206, the third answer to the third question is determined using the generative basic model.
[0076] The procedure includes step 208.
[0077] In step 208, the relation is specified.
[0078] For example, the relationship between the first answer and the second answer is predefined. Optionally, if the third question is used, the relationship between the first answer, the second answer, and the third answer is also predefined.
[0079] The procedure includes step 210.
[0080] In step 210, the size is determined depending on the answers and the relationship.
[0081] It may be possible to pose several alternative formulations of the questions to the basic model. The relationship is then individually checked for each of the answers that the basic model provides for each alternative formulation of the questions.
[0082] The size is determined, for example, depending on the result of the respective inspection.
[0083] For example, a confidence level is determined that the first answer and the second answer satisfy the relation.
[0084] Optionally, it may be provided that, if the third question is used, a confidence level is determined as a measure that the first answer, the second answer, and the third answer satisfy the relation.
[0085] For example, a value greater than zero represents confidence and a value of zero represents no confidence.
[0086] For example, if multiple checks are performed, the confidence values from the individual checks are added or multiplied to form a total value, with an increasing total value representing an increasingly greater confidence.
[0087] For example, a reward is determined as a measure if the first answer and the second answer satisfy the relation.
[0088] Optionally, it can be provided that, if the third question is used, a reward is determined as a measure if the first answer, the second answer, and the third answer satisfy the relation.
[0089] For example, otherwise no reward is determined. For instance, a value greater than zero represents the reward, and a value of zero does not.
[0090] For example, if multiple checks are performed, the reward values from each check are added or multiplied to form a total value, with an increasing total value representing an increasingly larger reward.
[0091] The procedure includes step 212.
[0092] Step 212 may provide for the confidence to be output specifically to the user of the basic model.
[0093] Step 212 may provide for the use of the basic model to be blocked for an application area that includes the first question and / or the second question if it is determined that the confidence falls below a threshold.
[0094] Step 212 may include the provision that the basic model is trained depending on the reward.
[0095] For example, the weights are determined depending on the reward.
[0096] Depending on the outcome of the verification, it may be possible to activate a computer-controlled or regulated machine, such as a robotic system, vehicle, household appliance, machine tool, manufacturing machine, personal assistance system, or access control system, based on the first, second, and / or third response. For example, activation will occur based on the first, second, and / or third response if the verification is determined to be successful. Conversely, no activation will occur based on the first, second, and / or third response, particularly if the verification is determined to be unsuccessful.
[0097] For example, in a robotic system for grasping objects, the first digital image captures the object in its first position, while the second digital image captures the object in its second position. Depending on whether the condition is met that the object's position in the response to the second question differs from its position in the response to the first question by a certain distance in the direction, the object is grasped in the second position or not.
[0098] For example, in an autonomous vehicle, the first digital image captures the object in its first position, while the second digital image captures the object in its second position. For instance, the vehicle will avoid running over the object in the second position depending on whether the condition is met that the position in the response to the second question differs in direction from the position in the response to the first question by a certain distance.
[0099] The distance and direction are determined, for example, based on a predefined expected trajectory of the object. The trajectory is estimated, for example, based on several captured images.
[0100] For optical inspection, for example, a first digital image representing an object is provided. This first digital image is transformed into a second digital image. The second digital image represents the object. For example, the relationship between how the first digital image is transformed into the second digital image is specified. This means it is checked whether the transformation of the first image results in the expected representation of the object in the second image, or not.
[0101] The images are captured using, for example, a camera, a radar sensor, a LiDAR sensor, an infrared camera, a motion sensor or an ultrasonic sensor.
[0102] In Fig. Figure 3 is an example of a data structure 300 for determining the size, shown schematically.
[0103] The data structure 300 comprises at least one data field 302 for a generative basic model, at least one data field 302 for a first answer to a first question, which is determined with the generative basic model, at least one data field 302 for a second answer to a second question, which is determined with the generative basic model, at least one data field 302 for a relation between the first answer and the second answer, and at least one data field 302 for the size, which is determined depending on the answers and the relation.
Claims
[1] Computer-implemented method for determining a quantity for evaluating the quality of a generative, in particular multi-modal, basic model, for example for optical inspection, for identifying sounds or technical parts, for natural language processing, for program generation, or for labeling data, characterized by , that a first answer to a first question is determined using the generative basic model (202), wherein a second answer to a second question is determined using the generative basic model (204), wherein a relation between the first answer and the second answer is specified (208), wherein the size is determined depending on the answers and the relation (210). [2] Method according to claim 1, characterized by , that a confidence is determined as a quantity (210) that the first answer and the second answer satisfy the relation. [3] Method according to claim 2, characterized by, that the confidence is output to a user of the basic model (212). [4] Method according to claim 2 or 3, characterized by , that the use of the basic model for an application area that includes the first question and / or the second question will be blocked (212) if it is found that the confidence falls below a threshold. [5] Method according to any one of the preceding claims, characterized by , that a reward is determined as a quantity (210) if the first response and the second response satisfy the relation, and the basic model is trained depending on the reward (212). [6] Method according to claim 5, characterized by , that the basic model includes an artificial neural network with weights, the weights being determined depending on the reward (212). [7] Method according to any of the preceding claims, characterized by, that the first answer comprises a set of individual answers to the first question, wherein the second answer comprises a set of individual answers to the second question, wherein the relation is given (208) that the first answer and the second answer are disjoint sets of individual answers, or that the first answer is a subset of the second answer. [8] Method according to any one of the preceding claims, characterized by , that the first question asks for a first program to complete a task on a first domain (202), wherein the second question asks for a second program to complete the task on a second domain (204), wherein the first domain includes the second domain, wherein the relation is specified (208) that the first program and the second program solve the task on the second domain with the same result. [9] Method according to any one of claims 1 to 7, characterized by, that a first digital image representing an object is provided, the first digital image is transformed into a second digital image representing the object, the first question is asked about the position of the object in the first digital image (202), the second question is asked about the position of the object in the second digital image (204), specifying the relation (208) of how the first digital image is transformed into the second digital image. [10] Method according to any one of claims 1 to 7, characterized by, that a first audio sequence with a noise is provided, a second audio sequence with the noise is provided, the first question is about a time at which the noise occurs in the first audio sequence (202), the second question is about a time at which the noise occurs in the second audio sequence (204), the relation of the times is specified (208). [11] Method according to any one of claims 1 to 7, characterized by , that the first question asks for the result of a concatenation of a first program with a second program to complete a task on an input (202), wherein the second question asks for the result of a third program to complete the task on the input (204), wherein the relation is given (208) that the third program solves the task with the same result as the concatenation of the first program with the second program. [12] Method according to any one of claims 1 to 7, characterized by , that a third answer to a third question is determined (206), for which the relation is expected (208) that the third answer includes the first answer and the second answer, or that the third answer is identical to the union of the first answer and the second answer. [13] Device (100) for determining a quantity for evaluating the quality of a generative, in particular multi-modal, basic model, for example for optical inspection, for identifying sounds or technical parts, for natural language processing, for program generation, or for marking data, characterized by, that the device comprises at least one processor (102) and at least one memory (104), wherein the at least one memory (104) comprises instructions, the execution of which by the at least one processor (102) on the device (100) carries out the method according to one of claims 1 to 12. [14] Computer program, characterized by , that the computer program comprises instructions executable by a computer, the execution of which by the computer on the computer results in the method according to one of claims 1 to 12. [15] Data structure (300) for determining a quantity for evaluating the quality of a generative, in particular multi-modal, basic model, for example for optical inspection, for identifying sounds or technical parts, for natural language processing, for program generation, or for labeling data, characterized by, that the data structure (300) has at least one data field (302) for a generative basic model, at least one data field (302) for a first answer to a first question determined by the generative basic model, at least one data field (302) for a second answer to a second question determined by the generative basic model, at least one data field (302) for a relation between the first answer and the second answer, and at least one data field (302) for the size determined depending on the answers and the relation.