Evaluation data generation method, evaluation method and related product

By obtaining non-text objects to generate description text and using models to generate evaluation questions and answers, the problems of inefficient and high labor costs in the prior art are solved, and a variety of evaluation questions and answers are efficiently generated.

CN120430285APending Publication Date: 2025-08-05SHUXING TECH (BEIJING) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510500166.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-21
Publication Date
2025-08-05

AI Technical Summary

Technical Problem

The prior art generates evaluation questions and target answers for evaluating language models' ability to understand information on non-text object, inefficient and high labor costs.

Method used

By obtaining non-text objects, generating description text using the first model, and inputting text and target dimensions into the second model, generating evaluation questions and target answers, and using prompt word guidance model for evaluation.

Benefits of technology

It improves the efficiency of generating evaluation questions and target answers, reduces labor costs, and can generate multiple evaluation questions and answers based on different non-text objects and target dimensions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120430285A_ABST
    Figure CN120430285A_ABST
Patent Text Reader

Abstract

The invention discloses an evaluation data generation method, an evaluation method and a related product. The evaluation data generation method comprises the following steps: acquiring a first non-text object, wherein the modal of the first non-text object is a non-text; the first non-text object is input into a first model, a first text describing the content of the first non-text object is generated, the first model has the capability of generating a description text, the description text comprises the content of the non-text object, and the modal of the non-text object is the non-text; the first text and the target dimension are input into a second model, an evaluation question and a target answer to the evaluation question are generated, the second model has the ability to generate a question for evaluation and an answer to the question for evaluation based on the description text and the evaluation dimension, the evaluation question is related to the first non-text object, and the evaluation question is related to the second non-text object. The evaluation problem comprises a problem of performing evaluation based on the target dimension. Through the evaluation data generation method, the labor cost for generating the evaluation question and the target answer of the evaluation question can be reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence technology, and in particular to an evaluation data generation method, an evaluation method, and related products. Background Art

[0002] Thanks to their powerful performance, language models are increasingly being used in a wide range of applications, including using them to process non-text objects. Because language models are designed to process text, they have a strong understanding of textual information, but a weaker understanding of non-textual information. Therefore, it is necessary to evaluate the language model's ability to understand non-textual information. To evaluate the language model's understanding of non-textual information, it is necessary to generate evaluation questions and target answers based on non-textual information. This allows the language model's understanding of non-textual information to be evaluated based on the evaluation questions and target answers.

[0003] The current approach to generating test questions and target answers for evaluation is manual work. Specifically, humans generate test questions and target answers based on non-text objects. However, this approach is inefficient and labor-intensive. Summary of the Invention

[0004] The present application provides an evaluation data generation method, an evaluation method, and related products to reduce the human cost of generating evaluation questions and target answers to evaluation questions, and improve the efficiency of generating evaluation questions and target answers. The related products include an evaluation device, an electronic device, a computer-readable storage medium, and a computer program product.

[0005] In a first aspect, a method for generating evaluation data is provided, the method comprising:

[0006] Acquire a first non-text object, where the modality of the first non-text object is non-text;

[0007] Inputting the first non-text object into a first model to generate a first text describing the content of the first non-text object, wherein the first model is capable of generating a description text, the description text including the content of the non-text object, and the modality of the non-text object is the non-text;

[0008] The first text and target dimension are input into the second model to generate evaluation questions and target answers to the evaluation questions. The second model has the ability to generate questions for evaluation and answers to the questions for evaluation based on the text of the data describing the text and the evaluation dimension. The evaluation questions are related to the first non-text object and include questions for evaluation based on the target dimension.

[0009] In conjunction with any embodiment of the present application, inputting the first non-text object into the first model to generate a first text describing the content of the first non-text object includes:

[0010] The first prompt word and the first non-text object are input into the first model to generate the first text describing the content of the first non-text object from the preset perspective, wherein the first prompt word is used to guide the model to describe the modality as non-text data from the preset perspective.

[0011] In conjunction with any embodiment of the present application, inputting the first text and target dimension into the second model to generate an evaluation question and a target answer to the evaluation question includes:

[0012] The second prompt word, the first text and the target dimension are input into the second model to generate the evaluation question and the target answer. The second prompt word is used to guide the model to generate the question for evaluation and the answer to the question for evaluation based on the text and evaluation dimension input into the model.

[0013] In combination with any embodiment of the present application, after generating the evaluation question, the method further includes:

[0014] When the accuracy rate of the evaluation question is less than or equal to the first threshold, based on the incorrect questions in the evaluation question, modify the content of the second prompt word related to the question generated for evaluation to obtain a third prompt word;

[0015] The third prompt word, the first text and the target dimension are input into the second model to generate a modified evaluation question and a modified answer to the modified evaluation question, wherein the modified evaluation question is related to the first non-text object and the modified evaluation question includes a question evaluated based on the target dimension.

[0016] In a second aspect, an evaluation method is provided, the method comprising:

[0017] Obtaining an evaluation question, a target answer to the evaluation question, and a result to be evaluated, wherein the evaluation question and the target answer are generated by a second model based on a first text describing the content of the first non-text object and a target dimension, the evaluation question includes a question evaluated based on the target dimension, the evaluation question is related to the first non-text object, the modality of the first non-text object is non-text, and the result to be evaluated includes a result output by the model to be evaluated based on the first non-text object;

[0018] Based on the evaluation question and the target answer, the result to be evaluated is evaluated to obtain the evaluation result of the model to be evaluated.

[0019] In conjunction with any embodiment of the present application, before evaluating the result to be evaluated based on the evaluation question and the target answer to obtain the evaluation result of the model to be evaluated, the method further includes:

[0020] Obtaining an evaluation criterion for the target answer;

[0021] The step of evaluating the result to be evaluated based on the evaluation question and the target answer to obtain the evaluation result includes:

[0022] Based on the evaluation question, the target answer and the evaluation standard, the result to be evaluated is evaluated to obtain the evaluation result.

[0023] In combination with any embodiment of the present application, the result to be evaluated includes the result generated by the model to be evaluated based on a third prompt word, the third prompt word includes the first non-text object and instruction text, and the instruction text is used to instruct the model to be evaluated to perform the task in the instruction text based on the first non-text object.

[0024] In a third aspect, a device for generating evaluation data is provided, the device comprising:

[0025] An acquiring unit, configured to acquire a first non-text object, wherein the modality of the first non-text object is non-text;

[0026] a generating unit, configured to input the first non-text object into a first model and generate a first text describing the content of the first non-text object, wherein the first model has the ability to generate a description text, the description text including the content of the non-text object, and the modality of the non-text object is the non-text;

[0027] The generation unit is also used to input the first text and target dimension into the second model to generate evaluation questions and target answers to the evaluation questions. The second model has the ability to generate questions for evaluation and answers to the questions for evaluation based on the description text and the evaluation dimensions. The evaluation questions are related to the first non-text object, and the evaluation questions include questions evaluated based on the target dimension.

[0028] In combination with any embodiment of the present application, the generating unit is further configured to:

[0029] The first prompt word and the first non-text object are input into the first model to generate the first text describing the content of the first non-text object from the preset perspective, wherein the first prompt word is used to guide the model to describe the modality as non-text data from the preset perspective.

[0030] In combination with any embodiment of the present application, the generating unit is further configured to:

[0031] The second prompt word, the first text and the target dimension are input into the second model to generate the evaluation question and the target answer. The second prompt word is used to guide the model to generate the question for evaluation and the answer to the question for evaluation based on the text and evaluation dimension input into the model.

[0032] In combination with any embodiment of the present application, the generating unit is further configured to: when the accuracy rate of the evaluation question is less than or equal to a first threshold, modify the content of the second prompt word related to the question generated for evaluation based on incorrect questions in the evaluation question to obtain a third prompt word;

[0033] The third prompt word, the first text and the target dimension are input into the second model to generate a modified evaluation question and a modified answer to the modified evaluation question, wherein the modified evaluation question is related to the first non-text object and the modified evaluation question includes a question evaluated based on the target dimension.

[0034] In a fourth aspect, an evaluation device is provided, the evaluation device comprising:

[0035] an acquisition unit, configured to acquire an evaluation question, a target answer to the evaluation question, and a result to be evaluated, wherein the evaluation question and the target answer are generated by a second model based on a first text describing the content of the first non-text object and a target dimension, the evaluation question includes a question evaluated based on the target dimension, the evaluation question is related to the first non-text object, the modality of the first non-text object is non-text, and the result to be evaluated includes a result output by the model to be evaluated based on the first non-text object;

[0036] An evaluation unit is used to evaluate the result to be evaluated based on the evaluation question and the target answer to obtain an evaluation result of the model to be evaluated.

[0037] In combination with any embodiment of the present application, the obtaining unit is further configured to obtain an evaluation standard for the target answer;

[0038] The evaluation unit is further configured to evaluate the result to be evaluated based on the evaluation question, the target answer, and the evaluation standard to obtain the evaluation result.

[0039] In combination with any embodiment of the present application, the result to be evaluated includes the result generated by the model to be evaluated based on a third prompt word, the third prompt word includes the first non-text object and instruction text, and the instruction text is used to instruct the model to be evaluated to perform the task in the instruction text based on the first non-text object.

[0040] In a fifth aspect, an electronic device is provided, comprising: a processor and a memory, the memory being used to store computer program code, the computer program code comprising computer instructions, and when the processor executes the computer instructions, the electronic device executes the first aspect and any embodiment thereof, or the electronic device executes the second aspect and any embodiment thereof.

[0041] In the sixth aspect, another electronic device is provided, comprising: a processor, a sending device, an input device, an output device and a memory, wherein the memory is used to store computer program code, and the computer program code includes computer instructions. When the processor executes the computer instructions, the electronic device executes the first aspect and any embodiment thereof, or the electronic device executes the second aspect and any embodiment thereof.

[0042] In the seventh aspect, a computer-readable storage medium is provided, in which a computer program is stored. The computer program includes program instructions, and when the program instructions are executed by a processor, the processor is caused to execute the first aspect and any embodiment thereof, or the processor is caused to execute the second aspect and any embodiment thereof.

[0043] In an eighth aspect, a computer program product is provided, which includes a computer program or instructions, and when the computer program or instructions are run on a computer, the computer is caused to execute the above-mentioned first aspect and any embodiment thereof, or the computer is caused to execute the above-mentioned second aspect and any embodiment thereof.

[0044] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present application.

[0045] In an embodiment of the present application, the modality of the first non-text object is non-text. After acquiring the first non-text object, the evaluation device inputs the first non-text object into a first model to generate a first text describing the content of the first non-text object. The first text and the target dimension are then input into a second model to generate an evaluation question and a target answer to the evaluation question, wherein the evaluation question is related to the first non-text object and includes a question evaluated based on the target dimension. This reduces the human cost of generating the evaluation question and the target answer to the evaluation question.

[0046] In an embodiment of the present application, on the one hand, the evaluation questions generated by the second model include questions that are evaluated based on the target dimension, that is, there is a corresponding relationship between the evaluation questions and the target dimension. Therefore, for the same non-text object, by changing the target dimension, different evaluation questions related to the non-text object and the target answers to each evaluation question can be generated. For example, after generating evaluation questions based on the two evaluation questions of position and size, if you want to continue to evaluate the dimension of color, by setting the target dimension to color, you can generate evaluation questions for evaluating color. On the other hand, since the evaluation questions are related to non-text objects, different evaluation questions can be generated for different non-text objects. Therefore, by changing the non-text object, different evaluation questions related to the non-text object and the target answers to each evaluation question can be generated. Based on the above two aspects, the present application can improve the efficiency of generating evaluation questions and target answers to evaluation questions. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the background technology, the drawings required for use in the embodiments of the present application or the background technology will be described below.

[0048] The drawings herein are incorporated into and constitute a part of the specification. These drawings illustrate embodiments consistent with the present application and, together with the specification, are used to illustrate the technical solutions of the present application.

[0049] Figure 1 A flowchart of a method for generating evaluation data provided in an embodiment of the present application;

[0050] Figure 2 A flowchart of another method for generating evaluation data provided in an embodiment of the present application;

[0051] Figure 3 A flow chart of an evaluation method provided in an embodiment of the present application;

[0052] Figure 4 A schematic diagram of the structure of an evaluation data generating device provided in an embodiment of the present application;

[0053] Figure 5 A schematic diagram of the structure of an evaluation device provided in an embodiment of the present application;

[0054] Figure 6 A schematic diagram of the hardware structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0055] In order to enable those skilled in the art to better understand the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0056] The terms "first," "second," and the like in the specification and claims of this application and the accompanying drawings are used to distinguish between different objects, not to describe a particular order. Furthermore, the terms "including," "having," and any variations thereof, are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or apparatus comprising a series of steps or elements is not limited to the listed steps or elements but may optionally include steps or elements not listed, or may optionally include other steps or elements inherent to the process, method, product, or apparatus.

[0057] References herein to "embodiments" mean that a particular feature, structure, or characteristic described in connection with the embodiments may be included in at least one embodiment of the present application. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor does it constitute an independent or alternative embodiment that is mutually exclusive of other embodiments. It is understood, both explicitly and implicitly, by those skilled in the art that the embodiments described herein may be combined with other embodiments.

[0058] Recently, language models have made significant progress, demonstrating exceptional capabilities in areas such as instruction-following, problem-solving, and open-ended chatting. The target language model is a model with natural language processing capabilities. Optionally, the target language model includes a large language model (LLM). For example, the target LLM is one of the following: ChatGPT or Llama. Thanks to the advancement of language models, their processing of multimodal objects has also been significantly improved. Multimodal objects include text and non-text objects. For convenience, non-text objects will be referred to as non-text objects below. By processing multimodal objects, language models can utilize text and non-text objects to generate corresponding results. For example, a language model can execute a task in response to a prompt and obtain the corresponding result. The prompt may include text and non-text objects. The text can instruct the language model to execute a task (such as a question-and-answer task) based on non-text objects and generate the result. In the embodiment of the present application, the prompt word is used to provide the language model with input data of the language model and the context of the input data. The prompt word can be used to guide the language model to perform the task in the prompt word based on the input data in the prompt word and the context of the input data.

[0059] Since language models need to utilize both textual and non-textual information when processing multimodal objects, their ability to correctly understand this information will affect their effectiveness in processing multimodal objects. Because language models are designed to process text, they have a strong understanding of textual information but a weaker understanding of non-textual information. Therefore, it is necessary to evaluate the language model's understanding of non-textual information. To evaluate the language model's understanding of non-textual information, it is necessary to generate evaluation questions and target answers based on non-textual information. This evaluation question and target answer can then be used to evaluate the language model's understanding of non-textual information.

[0060] The current approach to generating assessment questions and target answers for evaluation is manual. Specifically, humans generate assessment questions and target answers based on non-text objects. However, this approach is inefficient and labor-intensive. Therefore, embodiments of the present application provide an assessment data generation method to improve the efficiency of generating assessment questions and target answers, while reducing labor costs.

[0061] The evaluation data generation method of the present embodiment is performed by an evaluation data generation device (hereinafter referred to as the generation device). The generation device can be any electronic device capable of executing the technical solution disclosed in the method embodiment of the present application. Optionally, the generation device can be one of the following: a computer or a server.

[0062] It should be understood that the method embodiment of the present application can also be implemented by a processor executing computer program code. The following describes the embodiment of the present application in conjunction with the drawings in the embodiment of the present application. Figure 1 , Figure 1 A flowchart of a method for generating evaluation data provided in an embodiment of the present application.

[0063] 101. Obtain a first non-text object, where a modality of the first non-text object is non-text.

[0064] In the embodiments of the present application, object modalities include text, image, audio, video, and the like. Non-text objects are modalities other than text. Objects whose modality is non-text are referred to as non-text objects. For example, the first non-text object may be an image, the first non-text object may be audio, or the first non-text object may be video.

[0065] In an implementation of obtaining the first non-text object, the generating device receives the first non-text object input by the user through an input component, wherein the input component includes: a mouse, a keyboard, a touch screen, a touchpad, and an audio input device.

[0066] In another implementation of obtaining the first non-text object, the generating device receives the first non-text object sent by the user through a terminal, wherein the terminal includes: a mobile phone, a computer, a tablet computer, and a smart wearable device.

[0067] In another implementation of obtaining the first non-text object, the generating device extracts the non-text object from the multimodal object to obtain the first non-text object. For example, the multimodal object is a document including text and an image, and the first non-text object is obtained by extracting the image from the multimodal object.

[0068] In another implementation of obtaining the first non-text object, the first non-text object is a non-text object in a third prompt, and the generation device obtains the first non-text object using the third prompt. Specifically, the third prompt includes instruction text and the first non-text object, and the instruction text is used to instruct the language model to be evaluated to perform the task specified in the instruction text based on the first non-text object. After the third prompt is input into the language model to be evaluated, the language model to be evaluated performs the task specified in the instruction text based on the first non-text object and obtains a result to be evaluated. After the generation device generates an evaluation question and a target answer for evaluating the language model to be evaluated based on the first non-text object, the language model's ability to understand the non-text object can be determined based on the evaluation question, the target answer, and the result to be evaluated. For example, if the first non-text object is an image and the instruction text is to answer question A based on the content of the image, the answer to question A generated by the language model to be evaluated based on the image is the result to be evaluated. After the generation device generates the evaluation question and target answer based on the image, the language model's ability to understand the image can be determined based on the evaluation question, the target answer, and the result to be evaluated.

[0069] 102. Input a first non-text object into a first model to generate a first text describing the content of the first non-text object.

[0070] In an embodiment of the present application, a first language model is capable of generating descriptive text, wherein the descriptive text includes the content of a non-text object, and the modality of the non-text object is non-text, i.e., the descriptive text is used to describe the content of the non-text object. In one possible implementation, the first model is an agent constructed based on the LLM. In another possible implementation, the first model is a language model. The first model can generate first text describing the content of the non-text object based on its understanding of the non-text object. Optionally, the first model's understanding of the non-text object meets preset requirements. For example, n non-text objects are input into the first model respectively, and n processing results are generated, wherein the non-text objects correspond one-to-one to the processing results. If the accuracy of the n processing results is greater than or equal to a third threshold, it is determined that the first model's understanding of the non-text object meets the preset requirements. If the accuracy of the n processing results is less than the third threshold, it is determined that the first model's understanding of the non-text object does not meet the preset requirements. The first model's understanding of the non-text object meets the preset requirements, indicating that the text describing the content of the non-text object generated by the first model is highly accurate. Therefore, the generating device inputs the first non-text object into the first model, and generates a first text describing the content of the first non-text object, so as to subsequently generate an evaluation question and a target answer based on the first text.

[0071] In one possible implementation, a generation device inputs a first prompt word and a first non-text object into a first model to generate a first text describing the content of the first non-text object, wherein the first prompt word guides the model in describing the non-text object. The first model generates the first text under the guidance of the first prompt word, thereby improving the accuracy of the first text.

[0072] Optionally, the first prompt word is used to guide the model to describe the non-text object from a preset perspective. Accordingly, the generating device inputs the first prompt word and the first non-text object into the first model to generate a first text describing the content of the first non-text object from the preset perspective. For example, the preset perspective includes an overall description, a background description, a character description, and a color description. By instructing the first model to generate the first text describing the first non-text object from the preset perspective, the first text can provide a more comprehensive description of the first non-text object.

[0073] Optionally, the first prompt word includes an example of a non-text object that describes the modality from a prediction perspective, so that the first model can generate the first text with reference to the example, thereby improving the accuracy of the first text.

[0074] Optionally, the first model is an intelligent agent constructed based on LLM, and the first prompt word is a system prompt word (system prompt) of the intelligent agent.

[0075] For example, the first prompt includes the following content: You are an expert in image content description. Your task is to describe the image given to you in detail and completely. Please follow the steps below to describe it:

[0076] 1. Overall description: First, briefly summarize the subject and overall atmosphere of the image (e.g., this is an image of a happy family gathering, this is a picture of a quiet rural landscape).

[0077] 2. Background description: Describe the background of the image in detail, from far to near, observe the elements in the background (such as buildings, mountains, rivers, sky, etc.), and describe in detail their position, color, shape and any obvious features.

[0078] 3. Subject Description: If there are people in the image, please describe their appearance, location, facial expressions, movements, etc. If they are interacting, please describe their actions and relationships (for example, shaking hands, smiling, chatting, playing football, etc.);

[0079] If the subject in the image is an object or animal, please describe its type (e.g., cat breed, electronic device brand, etc.), shape, size, color, movement, material, etc.;

[0080] If the subject in the image is an object or animal, please describe its type (e.g., cat breed, electronic device brand, etc.), shape, size, color, movement, material, etc.;

[0081] 4. Environmental interaction description: Describe the interactive relationship between the subject and the background, as well as the dynamic events that may exist in the image (such as birds flying, cars driving, etc.).

[0082] 5. Detail description: Pay attention to any additional details (such as object texture, light intensity, etc.).

[0083] 6. Atmosphere and Emotion: Describe the overall mood of the image and explain how you came to this conclusion (e.g., through the character's expression, use of color, lighting effects, etc.).

[0084] Please use clear and understandable language to describe the image. Make sure to fully cover all the above aspects, but do not add any content that is not in the image. If the image is complex, you can describe different areas separately to ensure that every detail is clearly expressed.

[0085] [Format example]:

[0086] Overall: This is an image of multiple chairs placed randomly.

[0087] Background: This image shows an interior environment, possibly a living room or studio. The overall atmosphere is bright and modern, with light filtering in through the white floor-to-ceiling windows.

[0088] Main body: There are a total of eight chairs and a ladder arranged in four rows. The first row, the farthest away, has a brown fabric camping chair on the left, a computer chair with a brown backrest and a black seat in the middle, and a cyan sofa chair on the right. The second row has a brown chair on the left, a transparent chair in the middle, and a black chair on the right. The third row has an orange transparent ladder on the left and a black S-shaped chair on the right. The closest row has cyan backless camping chairs. The brand of the chairs is not given in the image, and it is impossible to identify the brand through the image.

[0089] Environmental Interactions: There are no hidden environmental interactions on this map.

[0090] Details: A bit of the table's corner is visible on the far left of the image, and above the image is written in orange Chinese font, which reads "Which 9 chairs did 2 people buy?"

[0091] Atmosphere and emotion: The overall image mood is dark.

[0092] Below is the image you need to describe.

[0093] After the first prompt word and the first non-text object are input into the first model, the first model can generate a first text describing the content of the first non-text object under the instruction of the first prompt word, thereby improving the accuracy of the first text.

[0094] In the first prompt, the preset angles include overall description, background description, theme description, environmental interaction description, detail description, atmosphere, and emotional description. The "Format Example" in the first prompt includes examples of describing non-text objects from the predicted angles and the format of the first text.

[0095] Optionally, the preset angle matches the task performed by the model to be evaluated based on the first non-text object. For example, the first non-text object includes an image, and the task performed by the model to be evaluated based on the first non-text object includes describing the characters in the image, then the preset angle matches describing the characters in the image. For example, the preset angle may be describing the appearance of the characters in the image, or describing the clothes of the characters in the image. For another example, the first non-text object includes a video, and the task performed by the model to be evaluated based on the first non-text object includes identifying a menu in the video, then the preset angle matches identifying a menu in the video. For example, the preset angle may be describing the menu in the video, describing the names of the dishes appearing in the video, or describing the prices of the dishes in the video.

[0096] 103. Input the first text and the target dimension into the second model to generate evaluation questions and target answers to the evaluation questions, wherein the evaluation questions are related to the first non-text object and the evaluation questions include questions evaluated based on the target dimension.

[0097] In an embodiment of the present application, the second model has the ability to generate questions for evaluation and answers to the questions for evaluation based on the descriptive text and the dimensions of the evaluation. In one possible implementation, the second model is an intelligent agent built based on the LLM. In another possible implementation, the second model is a language model. The target dimension is the dimension of the evaluation. For example, if the target dimension includes color, then the dimension of the evaluation includes color. For another example, if the target dimension includes position, then the dimension of the evaluation includes position. For another example, if the target dimension includes details, then the dimension of the evaluation includes details. The evaluation question is related to the first non-text object, that is, the evaluation question can be used to evaluate whether the understanding of the first non-text object is correct. For example, if the first non-text object is an image including a red vehicle, then the evaluation question can be whether there is a vehicle in the image, and the evaluation question can also be what the color of the vehicle in the image is. The target answer is the answer to the evaluation question. For example, the evaluation question is what the color of the vehicle in the image is, and the target answer is red.

[0098] In step 103, the generation device inputs the first text and the target dimension into the second model, which then generates an evaluation question and a target answer based on the target dimension. The target answer can be used as a basis for determining whether the evaluation result is correct. In this way, the evaluation question can be evaluated based on the target dimension.

[0099] In one possible implementation, the generation device inputs the first text and the target dimension into the second model to generate an evaluation question, a target answer to the evaluation question, and evaluation criteria for the target answer. In this implementation, the second model not only generates the evaluation question and the target answer, but also generates the evaluation criteria for the target answer. The evaluation criteria can be used as a basis for judging the accuracy of the evaluation result. For example, if the target dimension is color, the evaluation question is what is the color of the vehicle in the image, and the target answer is red, the evaluation criteria for the target answer include 1 point for a correct description of the color, 0.5 points for a partially correct description, and 0 points for an incorrect description.

[0100] In one possible implementation, the generating device inputs the second prompt word and the first text into the second model to generate target answers to the test questions and the evaluation questions, wherein the second prompt word is used to guide the model to generate questions for evaluation and answers to the evaluation questions based on the text and evaluation dimensions input into the model.

[0101] Optionally, the second model is an intelligent agent constructed based on LLM, and the second prompt word is a system prompt word of the intelligent agent.

[0102] For example, the second prompt might include the following: "You are an AI assistant with visual understanding capabilities. Based on the input image description and specified evaluation dimensions, you are responsible for generating image-related evaluation questions, answers to these evaluation questions, and evaluation criteria for the user." Your task is to ensure that the generated evaluation questions effectively assess the key elements of the image description and provide clear and accurate descriptions of the answers to the evaluation questions for use in the evaluation.

[0103] Please understand and generate the corresponding output based on the following few-shot examples:

[0104] Input Example 1:

[0105] Image description: There is a brown puppy standing on the green grass with a blue lake behind it.

[0106] Evaluation dimension: location.

[0107] Output example 1:

[0108] Evaluation question: What separates the puppy from the lake?

[0109] Answer to the assessment question: Green grass.

[0110] Evaluation criteria: 1 point for correctly identifying the separated elements, 0.5 points for correctly mentioning the distance or direction of the puppy and the lake, and 0 points for errors.

[0111] Input Example 2:

[0112] Image description: A red sports car is parked on a city street with a black bicycle parked next to it.

[0113] Evaluation dimension: color.

[0114] Output example 2:

[0115] Evaluation question: What color is the car in the image?

[0116] Answer to the assessment question: There is a red car in the image.

[0117] Evaluation criteria: 1 point for correct color description, 0.5 point for partially correct description, and 0 point for incorrect description.

[0118] Input Example 3:

[0119] Image description: Against a clear blue sky, five colorful balloons fly into the sky. The balloons are, in order: a large red balloon, a slightly smaller yellow balloon, a green heart-shaped balloon, a blue balloon with a star pattern, and a purple long balloon. The green heart-shaped balloon and the blue balloon are entwined, with the large red balloon in the foreground, the long purple balloon in the distance, and the yellow balloon to the lower right of the red balloon.

[0120] Evaluation dimension: details.

[0121] Output example 3:

[0122] Assessment Question: In this image, which balloon is entangled with the blue star-patterned balloon, and how do the positions of these two balloons affect each other?

[0123] Answer to the assessment question: The green heart-shaped balloon is entwined with the blue star-patterned balloon. Because they are entwined, the two balloons are closely connected and move synchronously during flight, making them more conspicuous relative to the blue balloon.

[0124] Evaluation criteria: 1 point for correctly identifying the green balloon, 1 point for reasonably describing the interactive relationship, 0.5 points for partially correct answers, and 0 points for errors.

[0125] Optionally, the evaluation criteria match the task performed by the model to be evaluated based on the first non-text object. For example, if the first non-text object includes an image, and the task performed by the model to be evaluated based on the first non-text object includes describing a person in the image, then the generated evaluation criteria include evaluation criteria related to describing the person in the image, and the evaluation criteria related to describing the person in the image are more important than evaluation criteria not related to describing the person in the image. For example, the score of the evaluation criteria related to describing the person in the image is higher than the score of the evaluation criteria not related to describing the person in the image.

[0126] In an embodiment of the present application, the modality of the first non-text object is non-text. After acquiring the first non-text object, the evaluation device inputs the first non-text object into a first model to generate a first text describing the content of the first non-text object. The first text and the target dimension are then input into a second model to generate an evaluation question and a target answer to the evaluation question, wherein the evaluation question is related to the first non-text object and includes a question evaluated based on the target dimension. This reduces the human cost of generating the evaluation question and the target answer to the evaluation question.

[0127] Based on steps 101 to 103, it can be seen that, on the one hand, the evaluation questions generated by the second model include questions evaluated based on the target dimension. In other words, there is a corresponding relationship between the evaluation questions and the target dimension. Therefore, for the same non-text object, by changing the target dimension, different evaluation questions related to the non-text object can be generated, as well as target answers for each evaluation question. For example, after generating evaluation questions based on the two evaluation questions of position and size, if you want to continue evaluating the dimension of color, by setting the target dimension to color, you can generate evaluation questions for color. On the other hand, because the evaluation questions are related to non-text objects, different evaluation questions can be generated for different non-text objects. Therefore, by changing the non-text object, different evaluation questions related to non-text objects can be generated, as well as target answers for each evaluation question. Based on these two aspects, the evaluation data generation method of steps 101 to 103 can improve the efficiency of generating evaluation questions and target answers for evaluation questions, and increase the number of evaluation questions.

[0128] As an optional implementation manner, the generating device performs the following steps during the process of executing step 102:

[0129] 2001. Input the second prompt word, the first text and the target dimension into the second model to generate evaluation questions and target answers, wherein the second prompt word is used to guide the model to generate evaluation questions and answers to the evaluation questions based on the text and evaluation dimensions input into the model.

[0130] As an optional implementation, after the generation device generates the evaluation questions in step 2001, the generation device further performs the following steps:

[0131] 2002. When the accuracy rate of the evaluation question is less than or equal to the first threshold, based on the incorrect question in the evaluation question, modify the content related to the question generated for evaluation in the second prompt word to obtain a third prompt word.

[0132] If the evaluation question is related to the first non-text object and is a question evaluated based on the target dimension, it means that the evaluation question is correct. Otherwise, if the evaluation question is not related to the first non-text object, or if the evaluation question is not a question evaluated based on the target dimension, it means that the evaluation question is incorrect. Since the evaluation question is generated based on the second prompt word, the correctness of the evaluation question is affected by the second prompt word. If the accuracy of the evaluation question is low, it means that there is a high probability that there is an error in the content of the second prompt word. Accordingly, the second prompt word needs to be modified. Otherwise, if the accuracy of the evaluation question is high, it means that there is a low probability that there is an error in the content of the second prompt word. Accordingly, the second prompt word does not need to be modified. In the embodiment of the present application, the generation device determines whether the accuracy of the evaluation question is high or low based on the first threshold. Specifically, if the accuracy of the evaluation question is greater than the first threshold, it means that the accuracy of the evaluation question is high. Otherwise, if the accuracy of the evaluation question is less than or equal to the first threshold, it means that the accuracy of the evaluation question is low. Therefore, when the accuracy of the evaluation question is less than or equal to the first threshold, the generation device determines that the content in the second prompt word needs to be modified. Specifically, based on the incorrect question in the evaluation question, the content related to the generated evaluation question in the second prompt word is modified to obtain a third prompt word. In this way, the evaluation question is generated based on the third prompt word, and the incorrect question in the evaluation question can be corrected.

[0133] Optionally, the content related to generating questions for evaluation in the second prompt word includes an example of generating evaluation questions, and the example is incorrect. For example, the example of generating evaluation questions in the second prompt word includes: Input example: Image description information: There is a brown puppy standing on the green grass, with a blue lake behind it. Evaluation dimension: position. Output example: Evaluation question: What color elements are there between the puppy and the lake? The answer to the evaluation question: green. In this example, the evaluation dimension is position, but the evaluation dimension of the evaluation question is color. Therefore, this example gives the second model incorrect guidance, which causes the second model to output an incorrect evaluation question. By modifying the content related to generating questions for evaluation in the second prompt word, the example is modified correctly to obtain the third prompt word.

[0134] Optionally, the generating device inputs the first non-text object, the first text, and the incorrect question into the model, so that the model generates a reason why the incorrect question is incorrect. The generating device then modifies the content of the second prompt word related to the question generated for evaluation based on the reason why the incorrect question is incorrect, to obtain a third prompt word.

[0135] In one implementation of determining the accuracy of an evaluation question, after generating the evaluation question, the generating device displays the evaluation question and the first non-text object, allowing a person to determine whether the evaluation question is correct. After receiving information indicating whether the evaluation question is correct, the generating device determines the accuracy of the evaluation question based on the information.

[0136] In another implementation method for determining the accuracy of an evaluation question, after generating the evaluation question, the generation device generates a first feature vector of the evaluation question, a second feature vector of the first non-text object, and a third feature vector of the target dimension. A first similarity between the first feature vector and the second feature vector, and a second similarity between the first feature vector and the third feature vector are determined. If the first similarity is greater than or equal to a first similarity threshold, and the second similarity is greater than or equal to a second similarity threshold, the evaluation question is determined to be correct. If the first similarity is less than the first similarity threshold, or the second similarity is less than the second similarity threshold, the evaluation question is determined to be incorrect.

[0137] 2003. Input the third prompt word, the first text and the target dimension into the second model to generate a modified evaluation question and a modified answer to the modified evaluation question, wherein the modified evaluation question is related to the first non-text object and the modified evaluation question includes a question evaluated based on the target dimension.

[0138] In an embodiment of the present application, the third prompt word is used to guide the model to generate evaluation questions and answers to the evaluation questions based on the text and evaluation dimensions input to the model. After the third prompt word, the first text, and the target dimension are input into the second model, the second model can, under the guidance of the third prompt word, generate a modified evaluation question and a modified answer to the modified evaluation question based on the first text and the target dimension. Because the third prompt word is obtained by modifying the second prompt word based on an incorrect question, the second model can correct the incorrect question by generating a modified evaluation question based on the third prompt word, thereby improving the accuracy of the modified evaluation question.

[0139] See also Figure 2 , Figure 2 This is a flow chart of another method for generating evaluation data provided in an embodiment of the present application. Figure 2As shown, first, an evaluation question and a target answer are generated based on the prompt word and the first non-text object. The implementation process of this step can be seen in steps 101 to 103. Then, a determination is made as to whether the accuracy of the evaluation question is less than or equal to a first threshold. If the determination result is negative, the evaluation question and the target answer are added to the evaluation dataset, where the evaluation dataset is a collection of evaluation data, including the evaluation question and the target answer. If the determination result is positive, the prompt word is modified based on the incorrect question in the evaluation question to obtain a modified prompt word. The implementation process of this step can be seen in step 2002. Then, a modified evaluation question and a modified answer are generated based on the modified prompt word and the first non-text object. The implementation process of this step can be seen in step 2003. The following steps are then iteratively performed: a determination is made as to whether the accuracy of the modified evaluation question is less than or equal to the first threshold. If the determination result is negative, the modified evaluation question and the modified answer are added to the evaluation dataset. If the determination result is negative, the prompt word is modified based on the incorrect question and a modified evaluation question and a modified answer are generated based on the modified prompt word. Obtaining the evaluation data set in this way can improve the accuracy of the evaluation data in the evaluation data set.

[0140] As an optional embodiment, after the generation device generates the target answer by executing step 2001, the generation device further performs the following steps: if the accuracy of the target answer is less than or equal to the second threshold, based on the incorrect answer in the target answer, modify the content of the second prompt word related to the answer to the generated evaluation question to obtain a fourth prompt word. The fourth prompt word, the first text, and the target dimension are input into the second model to generate a new evaluation question and a new answer to the new evaluation question. The new evaluation question is related to the first non-text object and includes a question evaluated based on the target dimension. Because the fourth prompt word is obtained by modifying the second prompt word based on the incorrect answer, the second model can correct the incorrect answer by generating a new answer based on the fourth prompt word, thereby improving the accuracy of the new answer.

[0141] As an optional embodiment, the generation device performs the following steps during step 2001: inputting the second prompt word, the first text, and the target dimension into the second model to generate an evaluation question, a target answer, and a basis for generating the evaluation question, wherein the second prompt word is used to guide the model to generate the evaluation question, the answer to the evaluation question, and the basis for generating the evaluation question based on the text and evaluation dimension input into the model. After obtaining the basis for generating the evaluation question, the generation device determines the accuracy of the evaluation question by performing the following steps: determining the accuracy of the evaluation question based on the basis for generating the evaluation question, the first text, and the evaluation question.

[0142] Considering that the model is prone to generating information that is inconsistent with the facts or is erroneous, in this embodiment, the second prompt word can be used to guide the model to generate a question for evaluation, an answer to the question for evaluation, and a basis for generating the question for evaluation based on the text and evaluation dimension input into the model. Therefore, after the generation device inputs the second prompt word, the first text, and the target dimension into the second model, it can generate an evaluation question, a target answer, and a basis for generating the evaluation question. Based on the basis for generating the evaluation question, it can then be determined whether the evaluation question generated by the second model is correct, thereby determining the accuracy of the evaluation question. For example, the evaluation question is: What elements separate the puppy and the lake? The basis for generating the evaluation question is: The image includes a puppy and a lake, and the puppy and the lake are separated by green grass. Based on the basis for generating the evaluation question, the generation device can determine whether the evaluation question is correct by determining whether the image includes a puppy and a lake, and whether the puppy and the lake are separated by green grass.

[0143] As an optional embodiment, the first model is an agent constructed based on the first LLM, and the second model is an agent constructed based on the second LLM, wherein the first model is represented as π_{s}, the second model is represented as μ_{s'}, π represents the first LLM, s represents the first prompt word, μ represents the second LLM, s' represents the second prompt word, wherein the second prompt word includes a target dimension, the target dimension is represented as {D_{j}, j = 1, 2, ..., N}, D_{j} represents the jth target dimension, and N represents the number of target dimensions. Optionally, μ represents a generative pre-trained transformer (GPT) 4. The modality of the first non-text object is an image, and the set of the first non-text object is represented as {IMG_{i}, i = 1, 2, ..., M}, wherein IMG_{i} represents the i-th image in the set and M is the number of images.

[0144] First, the first prompt word and the first non-text object are input into the first model to generate a first text describing the content of the first non-text object from a preset perspective. This step can be expressed as: caption = π_{s}(IMG), where caption represents the first text. Then, the second prompt word, the first text, and the target dimension are input into the second model to generate an evaluation question, a target answer, and an evaluation criterion for the target answer. This step can be expressed as <evaluation question q, target answer ref, evaluation criterion rule> = μ_{s'}(caption, D).

[0145] The present application also provides an evaluation method that utilizes the evaluation questions and target answers generated by the evaluation data generation method described above to evaluate the evaluation results output by the language model to be evaluated, thereby obtaining the evaluation results of the language model to be evaluated. The evaluation method is performed by an evaluation device, which can be any electronic device capable of executing the technical solutions disclosed in the present application. Optionally, the evaluation device can be any of the following: a computer or a server.

[0146] See also Figure 3 , Figure 3 A flowchart of an evaluation method provided in an embodiment of the present application.

[0147] 301. Obtain the evaluation question, the target answer to the evaluation question, and the result to be evaluated.

[0148] In an embodiment of the present application, the evaluation question and the target answer are generated by the second model based on the first text and the target dimension describing the content of the first non-text object, the evaluation question includes a question evaluated based on the target dimension, the evaluation question is related to the first non-text object, and the modality of the first non-text object is non-text. Optionally, the evaluation question and the target answer to the evaluation question can be generated by the following steps: Obtain the first non-text object. Input the first non-text object into the first model to generate a first text describing the content of the first non-text object. Input the first text and the target dimension into the second model to generate the evaluation question and the target answer to the evaluation question, wherein the evaluation question is related to the first non-text object, and the evaluation question includes a question evaluated based on the target dimension.

[0149] The result to be evaluated includes the result output by the language model to be evaluated based on the first non-text object. In one possible implementation, the result to be evaluated includes the result generated by the language model to be evaluated based on the third prompt word. The third prompt word includes the first non-text object and the instruction text, wherein the instruction text is used to instruct the language model to be evaluated to perform the task in the instruction text based on the first non-text object. After the third prompt word is input into the language model to be evaluated, the language model to be evaluated, under the guidance of the third prompt word, performs the task in the instruction text based on the first non-text object to obtain the result to be evaluated. For example, the first non-text object is an image, and the instruction text is used to instruct the language model to be evaluated to recognize the menu in the image. Then the result to be evaluated is the result obtained by the language model to be evaluated recognizing the menu in the image.

[0150] 302. Based on the evaluation question and the target answer, the evaluation result to be evaluated is evaluated to obtain the evaluation result of the model to be evaluated.

[0151] In the embodiment of the present application, the evaluation result is an evaluation result of the ability of the model to be evaluated to understand non-text objects. The evaluation device can determine whether the evaluation result is correct based on the evaluation question and the target answer, and then obtain the evaluation result of the model to be evaluated.

[0152] In one possible implementation, if the evaluation device determines that the result to be evaluated is correct based on the evaluation question and the target answer, the evaluation device determines that the evaluation result includes that the ability of the model to be evaluated to understand non-text objects meets the requirements. If the evaluation device determines that the result to be evaluated is incorrect based on the evaluation question and the target answer, the evaluation device determines that the evaluation result includes that the ability of the model to be evaluated to understand non-text objects does not meet the requirements.

[0153] In the embodiment of the present application, the evaluation questions and target answers are generated based on the evaluation data generation method described above. After obtaining the evaluation questions, the target answers to the evaluation questions, and the results to be evaluated, the evaluation device evaluates the results to be evaluated based on the evaluation questions and the target answers, and obtains the evaluation results of the model to be evaluated, thereby evaluating the ability of the model to be evaluated to understand non-text objects. Since the evaluation questions and target answers are generated based on the evaluation data generation method, the efficiency of generating the evaluation questions and the target answers to the evaluation questions can be improved, the number of evaluation questions can be increased, and the evaluation results to be evaluated based on the evaluation questions and the target answers can be evaluated to obtain the evaluation results of the model to be evaluated, which can improve the efficiency of obtaining the evaluation results and the accuracy of the evaluation results.

[0154] As an optional embodiment, before executing step 302, the evaluation device further performs the following steps: obtaining the evaluation criteria for the target answer. After obtaining the evaluation criteria for the target answer, during the execution of step 302, the device further performs the following steps: evaluating the target evaluation result based on the evaluation question, the target answer, and the evaluation criteria to obtain the evaluation result.

[0155] In this embodiment, the evaluation results are evaluated based on the evaluation questions, target answers and evaluation criteria, and the scores of the evaluation results can be determined, and then the evaluation results can be determined based on the scores, wherein the scores represent the ability of the model to be evaluated to understand non-text objects.

[0156] Optionally, the evaluation device inputs the fifth prompt word, evaluation question, target answer, evaluation standard and evaluation result to be evaluated into the third model to generate the evaluation result of the model to be evaluated, wherein the fifth prompt word is used to guide the third model to evaluate the evaluation result based on the evaluation question, target answer and evaluation standard.

[0157] Optionally, the third model is an agent constructed based on the third LLM, wherein the third model is represented by β_{s''}, β represents the third LLM, and s'' represents the fifth prompt word. The fifth prompt word, evaluation question, target answer, evaluation standard and evaluation result to be evaluated are input into the third model to generate the evaluation result of the model to be evaluated. The implementation of the step can be expressed as: <Evaluation result> = β_{s''}(<q,ref,rule,ans> ), where q represents the evaluation question, ref represents the target answer, rule represents the evaluation standard, and ans represents the result to be evaluated.

[0158] The present application embodiment also provides another evaluation method, which is based on Figure 1 The evaluation data generation method generates evaluation questions and target answers, and then evaluates the evaluation results output by the language model to be evaluated based on the evaluation questions and target answers to obtain the evaluation results of the language model to be evaluated. The evaluation method includes the following steps:

[0159] 401. Obtain a first non-text object and a result to be evaluated, wherein the modality of the first non-text object is non-text, and the result to be evaluated includes a result output by the model to be evaluated based on the first non-text object.

[0160] The implementation of this step can be referred to step 101 and will not be described in detail here.

[0161] 402. Input the first non-text object into the first model to generate a first text describing the content of the first non-text object. The first model has the ability to generate a description text.

[0162] The implementation of this step can be referred to step 102 and will not be described in detail here.

[0163] 403. Input the first text and the target dimension into the second model to generate evaluation questions and target answers to the evaluation questions, wherein the second model has the ability to generate questions for evaluation and answers to the evaluation questions based on the descriptive text and the evaluation dimension, the evaluation questions are related to the first non-text object, and the evaluation questions include questions evaluated based on the target dimension.

[0164] The implementation of this step can be referred to step 103 and will not be described in detail here.

[0165] 404. Based on the evaluation question and the target answer, the evaluation result to be evaluated is evaluated to obtain the evaluation result of the model to be evaluated.

[0166] The implementation of this step can be referred to step 302 and will not be described in detail here.

[0167] In the implementation of this application, after generating the evaluation questions and target answers, the evaluation device can further evaluate the evaluation results based on the evaluation questions and target answers based on steps 401 to 404 to obtain the evaluation results of the model to be evaluated. This can achieve the effect of obtaining the evaluation results of the model to be evaluated when the first non-text object to be processed by the model to be evaluated and the result to be evaluated output by the model to be evaluated are input into the evaluation device.

[0168] Those skilled in the art will understand that in the above-mentioned method of the specific implementation method, the writing order of each step does not mean a strict execution order and does not constitute any limitation on the implementation process. The specific execution order of each step should be determined by its function and possible internal logic.

[0169] The above describes in detail the method of the embodiment of the present application, and the following provides an apparatus of the embodiment of the present application.

[0170] See also Figure 4 , Figure 4 This is a schematic diagram of the structure of an evaluation data generation device provided in an embodiment of the present application. The evaluation data generation device 1 includes: an acquisition unit 11 and a generation unit 12, wherein:

[0171] An acquiring unit 11 is configured to acquire a first non-text object, wherein the modality of the first non-text object is non-text;

[0172] a generating unit 12, configured to input the first non-text object into a first model and generate a first text describing the content of the first non-text object, wherein the first model is capable of generating a description text, the description text including the content of the non-text object, and the modality of the non-text object is the non-text;

[0173] The generation unit 12 is also used to input the first text and target dimension into the second model to generate evaluation questions and target answers to the evaluation questions. The second model has the ability to generate questions for evaluation and answers to the questions for evaluation based on the descriptive text and the evaluation dimension. The evaluation questions are related to the first non-text object, and the evaluation questions include questions evaluated based on the target dimension.

[0174] In combination with any embodiment of the present application, the generating unit 12 is further configured to:

[0175] The first prompt word and the first non-text object are input into the first model to generate the first text describing the content of the first non-text object from the preset perspective, wherein the first prompt word is used to guide the model to describe the modality as non-text data from the preset perspective.

[0176] In combination with any embodiment of the present application, the generating unit 12 is further configured to:

[0177] The second prompt word, the first text and the target dimension are input into the second model to generate the evaluation question and the target answer. The second prompt word is used to guide the model to generate the question for evaluation and the answer to the question for evaluation based on the text and evaluation dimension input into the model.

[0178] In combination with any embodiment of the present application, the generating unit 12 is further configured to: when the accuracy rate of the evaluation question is less than or equal to a first threshold, modify the content of the second prompt word related to the question generated for evaluation based on the incorrect question in the evaluation question to obtain a third prompt word;

[0179] The third prompt word, the first text and the target dimension are input into the second model to generate a modified evaluation question and a modified answer to the modified evaluation question, wherein the modified evaluation question is related to the first non-text object and the modified evaluation question includes a question evaluated based on the target dimension.

[0180] In an embodiment of the present application, the modality of the first non-text object is non-text. After acquiring the first non-text object, the evaluation device inputs the first non-text object into a first model to generate a first text describing the content of the first non-text object. The first text and the target dimension are then input into a second model to generate an evaluation question and a target answer to the evaluation question, wherein the evaluation question is related to the first non-text object and includes a question evaluated based on the target dimension. This reduces the human cost of generating the evaluation question and the target answer to the evaluation question.

[0181] See also Figure 5 , Figure 5 This is a schematic diagram of the structure of an evaluation device provided in an embodiment of the present application. The evaluation device 2 includes: an acquisition unit 21 and an evaluation unit 22, wherein:

[0182] an acquisition unit 21, configured to acquire an evaluation question, a target answer to the evaluation question, and a result to be evaluated, wherein the evaluation question and the target answer are generated by a second model based on a first text describing the content of the first non-text object and a target dimension, the evaluation question includes a question evaluated based on the target dimension, the evaluation question is related to the first non-text object, the modality of the first non-text object is non-text, and the result to be evaluated includes a result output by the model to be evaluated based on the first non-text object;

[0183] The evaluation unit 22 is configured to evaluate the result to be evaluated based on the evaluation question and the target answer to obtain an evaluation result of the model to be evaluated.

[0184] In combination with any embodiment of the present application, the acquisition unit 21 is further configured to acquire an evaluation standard for the target answer;

[0185] The evaluation unit 22 is further configured to evaluate the result to be evaluated based on the evaluation question, the target answer, and the evaluation standard to obtain the evaluation result.

[0186] In combination with any embodiment of the present application, the result to be evaluated includes the result generated by the model to be evaluated based on a third prompt word, the third prompt word includes the first non-text object and instruction text, and the instruction text is used to instruct the model to be evaluated to perform the task in the instruction text based on the first non-text object.

[0187] In the embodiment of the present application, the evaluation questions and target answers are generated based on the evaluation data generation method described above. After obtaining the evaluation questions, the target answers to the evaluation questions, and the results to be evaluated, the evaluation device evaluates the results to be evaluated based on the evaluation questions and the target answers, and obtains the evaluation results of the model to be evaluated, thereby evaluating the ability of the model to be evaluated to understand non-text objects. Since the evaluation questions and target answers are generated based on the evaluation data generation method, the efficiency of generating the evaluation questions and the target answers to the evaluation questions can be improved, the number of evaluation questions can be increased, and the evaluation results to be evaluated based on the evaluation questions and the target answers can be evaluated to obtain the evaluation results of the model to be evaluated, which can improve the efficiency of obtaining the evaluation results and the accuracy of the evaluation results.

[0188] In some embodiments, the functions or modules included in the device provided in the embodiments of the present application can be used to execute the method described in the above method embodiments. The specific implementation can refer to the description of the above method embodiments. For the sake of brevity, it will not be repeated here.

[0189] Figure 6 A schematic diagram of the hardware structure of an electronic device provided in an embodiment of the present application. The electronic device 3 includes a processor 31 and a memory 32. Optionally, the electronic device 3 also includes an input device 33 and an output device 34. The processor 31, the memory 32, the input device 33 and the output device 34 are coupled via a connector, and the connector includes various interfaces, transmission lines or buses, etc., which are not limited in the embodiments of the present application. It should be understood that in each embodiment of the present application, coupling refers to mutual connection in a specific manner, including direct connection or indirect connection through other devices, for example, connection through various interfaces, transmission lines, buses, etc.

[0190] The processor 31 may include one or more processors, for example, one or more central processing units (CPUs). In the case where the processor is a CPU, the CPU may be a single-core CPU or a multi-core CPU. Alternatively, the processor 31 may be a processor group consisting of multiple CPUs, wherein the multiple processors are coupled to each other via one or more buses. Alternatively, the processor may also be other types of processors, etc., which are not limited in the embodiments of the present application.

[0191] The memory 32 can be used to store computer program instructions and various computer program codes, including the program code for executing the solution of the present application. Optionally, the memory includes, but is not limited to, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), or portable compact disc read-only memory (CD-ROM), which is used for related instructions and data.

[0192] The input device 33 is used to input data and / or signals, and the output device 34 is used to output data and / or signals. The input device 33 and the output device 34 can be independent devices or an integrated device.

[0193] It can be understood that in the embodiment of the present application, the memory 32 can be used not only to store relevant instructions, but also to store relevant data. The embodiment of the present application does not limit the specific data stored in the memory.

[0194] It is understandable that Figure 6 Only a simplified design of an electronic device is shown. In actual applications, the electronic device may further include other necessary components, including but not limited to any number of input / output devices, processors, memories, etc., and all electronic devices that can implement the embodiments of the present application are within the scope of protection of the present application.

[0195] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0196] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here. Those skilled in the art will also clearly understand that the descriptions of the various embodiments of this application have different focuses. For the convenience and brevity of description, the same or similar parts may not be repeated in different embodiments. Therefore, for parts not described or not described in detail in a certain embodiment, reference can be made to the descriptions of other embodiments.

[0197] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0198] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0199] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.

[0200] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted via the computer-readable storage medium. The computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center via wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that includes one or more available media integrated therein. The available medium may be a magnetic medium (eg, a floppy disk, a hard disk, a magnetic tape), an optical medium (eg, a digital versatile disc (DVD)), or a semiconductor medium (eg, a solid state disk (SSD)).

[0201] Those skilled in the art will appreciate that all or part of the processes in the above-described method embodiments can be implemented by a computer program instructing related hardware to perform the processes. The program can be stored in a computer-readable storage medium, and when executed, the program can include the processes in the above-described method embodiments. The aforementioned storage medium includes various media capable of storing program code, such as read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

Claims

1. A method for generating evaluation data, characterized in that: The method comprises: Acquire a first non-text object, where the modality of the first non-text object is non-text; Inputting the first non-text object into a first model to generate a first text describing the content of the first non-text object, wherein the first model is capable of generating a description text, the description text including the content of the non-text object, and the modality of the non-text object is the non-text; The first text and target dimension are input into the second model to generate evaluation questions and target answers to the evaluation questions. The second model has the ability to generate questions for evaluation and answers to the questions for evaluation based on the descriptive text and the evaluation dimension. The evaluation questions are related to the first non-text object and include questions for evaluation based on the target dimension.

2. The method according to claim 1, characterized in that Inputting the first non-text object into a first model to generate a first text describing the content of the first non-text object includes: The first prompt word and the first non-text object are input into the first model to generate the first text describing the content of the first non-text object from the preset perspective, wherein the first prompt word is used to guide the model to describe the modality as non-text data from the preset perspective.

3. The method according to claim 1 or 2, characterized in that Inputting the first text and target dimension into the second model to generate an evaluation question and a target answer to the evaluation question includes: The second prompt word, the first text and the target dimension are input into the second model to generate the evaluation question and the target answer. The second prompt word is used to guide the model to generate the question for evaluation and the answer to the question for evaluation based on the text and evaluation dimension input into the model.

4. The method according to claim 1 or 2, characterized in that After generating the assessment question, the method further includes: When the accuracy rate of the evaluation question is less than or equal to the first threshold, based on the incorrect questions in the evaluation question, modify the content of the second prompt word related to the question generated for evaluation to obtain a third prompt word; The third prompt word, the first text and the target dimension are input into the second model to generate a modified evaluation question and a modified answer to the modified evaluation question, wherein the modified evaluation question is related to the first non-text object and the modified evaluation question includes a question evaluated based on the target dimension.

5. An evaluation method, characterized in that: The method comprises: Obtaining an evaluation question, a target answer to the evaluation question, and a result to be evaluated, wherein the evaluation question and the target answer are generated by a second model based on a first text describing the content of the first non-text object and a target dimension, the evaluation question includes a question evaluated based on the target dimension, the evaluation question is related to the first non-text object, the modality of the first non-text object is non-text, and the result to be evaluated includes a result output by the model to be evaluated based on the first non-text object; Based on the evaluation question and the target answer, the result to be evaluated is evaluated to obtain the evaluation result of the model to be evaluated.

6. The method according to claim 5, characterized in that Before evaluating the result to be evaluated based on the evaluation question and the target answer to obtain the evaluation result of the model to be evaluated, the method further includes: Obtaining an evaluation criterion for the target answer; The step of evaluating the result to be evaluated based on the evaluation question and the target answer to obtain the evaluation result includes: Based on the evaluation question, the target answer and the evaluation standard, the result to be evaluated is evaluated to obtain the evaluation result.

7. The method according to claim 5 or 6, characterized in that The result to be evaluated includes the result generated by the model to be evaluated based on the third prompt word, the third prompt word includes the first non-text object and instruction text, and the instruction text is used to instruct the model to be evaluated to perform the task in the instruction text based on the first non-text object.

8. A device for generating evaluation data, characterized in that: The evaluation data generating device includes: An acquiring unit, configured to acquire a first non-text object, wherein the modality of the first non-text object is non-text; a generating unit, configured to input the first non-text object into a first model and generate a first text describing the content of the first non-text object, wherein the first model has the ability to generate a description text, the description text including the content of the non-text object, and the modality of the non-text object is the non-text; The generation unit is also used to input the first text and target dimension into the second model to generate evaluation questions and target answers to the evaluation questions. The second model has the ability to generate questions for evaluation and answers to the questions for evaluation based on the descriptive text and the evaluation dimension. The evaluation questions are related to the first non-text object, and the evaluation questions include questions evaluated based on the target dimension.

9. An evaluation device, characterized in that: The evaluation device comprises: an acquisition unit, configured to acquire an evaluation question, a target answer to the evaluation question, and a result to be evaluated, wherein the evaluation question and the target answer are generated by a second model based on a first text describing the content of the first non-text object and a target dimension, the evaluation question includes a question evaluated based on the target dimension, the evaluation question is related to the first non-text object, the modality of the first non-text object is non-text, and the result to be evaluated includes a result output by the model to be evaluated based on the first non-text object; An evaluation unit is used to evaluate the result to be evaluated based on the evaluation question and the target answer to obtain an evaluation result of the model to be evaluated.

10. An electronic device, characterized in that: include: A processor and a memory, the memory being used to store computer program code, the computer program code including computer instructions, and when the processor executes the computer instructions, the electronic device executes the method as claimed in any one of claims 1 to 4, or the electronic device executes the method as claimed in any one of claims 5 to 7.

11. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, which includes program instructions. When the program instructions are executed by a processor, the processor is caused to execute the method described in any one of claims 1 to 4, or the processor is caused to execute the method described in any one of claims 5 to 7.

12. A computer program product, characterized in that The computer program product includes a computer program or instructions; when the computer program or instructions are run on a computer, the computer is enabled to execute the method described in any one of claims 1 to 4, or the computer is enabled to execute the method described in any one of claims 5 to 7.