A model evaluation method and device, electronic equipment and storage medium

CN122838889APending Publication Date: 2026-09-29CHINA MOBILE COMM LTD RES INST +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510382486.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-28
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

即使使用的Prompt是相同的,由于模型之间存在差异,不同的待评估模型也不能完全按照Prompt的要求输出答案格式,而静态的答案解析规则无法准确匹配出模型的输出答案格式,导致评估的误判

Benefits of technology

[0027]本发明实施例提供的模型评估方法、装置、电子设备和存储介质,所述方法包括:利用待评估模型对第一提示词以及数据集中的数据进行处理,获得第一输出数据,所述第一输出数据中包括所述数据对应的第一输出结果;在所述第一输出结果与匹配规则集合中的所有匹配规则均不匹配的情况下,利用开源语言模型对所述第一输出结果和设定的第二提示词进行处理,获得第二输出数据,所述第二输出数据包括第一匹配规则和所述数据对应的第二输出结果;将所述第一匹配规则添加至所述匹配规则集合以供对所述待评估模型的下一个输出结果进行匹配,以及根据所述第二输出结果和所述数据对应的标签数据获得所述待评估模型的评分;所述标签数据为所述数据对应的期望输出结果;根据多个数据各自对应的评分获得所述待评估模型的评估结果。采用本发明实施例的技术方案,无论待评估模型与相应的提示词是否发生变化,利用开源语言模型自动生成匹配规则,并自动更新匹配规则,实现后续模型评估过程自动使用增强的规则,从而实现匹配规则的动态调整,避免人工干预,避免因为匹配规则与待评估模型的输出结果不匹配造成的结果误判,提高了模型评估结果的准确性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122838889A_ABST
    Figure CN122838889A_ABST
Patent Text Reader

Abstract

Embodiments of the present application disclose a model evaluation method and device, electronic equipment and a storage medium. The method comprises: processing a first prompt word and data in a data set by using a to-be-evaluated model to obtain first output data, wherein the first output data includes a first output result corresponding to the data; in a case where the first output result does not match all matching rules in a matching rule set, processing the first output result and a set second prompt word by using an open source language model to obtain second output data, wherein the second output data includes a first matching rule and a second output result corresponding to the data; adding the first matching rule to the matching rule set for matching a next output result of the to-be-evaluated model, and obtaining a score of the to-be-evaluated model according to the second output result and label data corresponding to the data; and obtaining an evaluation result of the to-be-evaluated model according to scores corresponding to a plurality of data respectively.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence technology, and more specifically to a model evaluation method, apparatus, electronic device, and storage medium. Background Technology

[0002] Model evaluation is a crucial step in machine learning and artificial intelligence, used to measure a model's performance on a specific task. For evaluating conversational language models, the typical approach involves using a dataset and custom prompts as input, calling the model to produce output, and then defining answer parsing rules based on the prompts and model output to score the model's responses.

[0003] In existing technologies, the organization of prompts depends on the format and content of the dataset and the requirements of the evaluation framework. To achieve good evaluation results, different evaluation frameworks use different prompts. Even if the same prompt is used, due to differences between models, different models to be evaluated cannot output answer formats exactly as required by the prompt. Static answer parsing rules cannot accurately match the output answer format of the model, leading to misjudgments in the evaluation. Summary of the Invention

[0004] To address the existing technical problems, embodiments of the present invention provide a model evaluation method, apparatus, electronic device, and storage medium.

[0005] To achieve the above objectives, the technical solution of this invention is implemented as follows:

[0006] This invention provides a model evaluation method, the method comprising:

[0007] The model to be evaluated is used to process the first prompt word and the data in the dataset to obtain the first output data, which includes the first output result corresponding to the data.

[0008] If the first output result does not match any of the matching rules in the matching rule set, the first output result and the set second prompt word are processed using an open-source language model to obtain second output data. The second output data includes the first matching rule and the second output result corresponding to the data.

[0009] The first matching rule is added to the matching rule set for matching the next output result of the model to be evaluated, and a score of the model to be evaluated is obtained based on the second output result and the label data corresponding to the data; the label data is the expected output result corresponding to the data.

[0010] The evaluation result of the model to be evaluated is obtained based on the scores corresponding to each of the multiple data points.

[0011] In the above scheme, the step of processing the first output result and the set second prompt word using an open-source language model to obtain the second output data includes:

[0012] The first output result, all matching rules in the matching rule set, and the set second prompt word are processed using an open-source language model to obtain a second output result corresponding to the first matching rule and the data; wherein, the second output result is obtained by the open-source language model based on all matching rules in the matching rule set and the first output result.

[0013] In the above scheme, before adding the first matching rule to the matching rule set, the method further includes: identifying the second output data based on preset keywords or preset key sentences, and extracting the first matching rule.

[0014] In the above scheme, the method further includes: when the first output result matches the second matching rule in the matching rule set, obtaining the score of the model to be evaluated based on the first output result and the label data corresponding to the data.

[0015] In the above scheme, obtaining the score of the model to be evaluated based on the first output result and the label data corresponding to the data includes: when the first output result matches the label data, determining the score of the model to be evaluated; when the first output result does not match the label data, determining that the model to be evaluated does not receive a score.

[0016] And / or, obtaining the score of the model to be evaluated based on the second output result and the label data corresponding to the data includes: when the second output result matches the label data, determining the score of the model to be evaluated; when the second output result does not match the label data, determining that the model to be evaluated does not receive a score.

[0017] In the above scheme, before adding the first matching rule to the matching rule set, the method further includes: identifying the second output data based on preset keywords or key sentences, and identifying and extracting the first matching rule.

[0018] In the above scheme, before processing the first prompt word and the data in the dataset using the model to be evaluated, the method further includes: obtaining the dataset and the first prompt word, wherein the dataset includes multiple data and the label data corresponding to each data.

[0019] This invention also provides a model evaluation apparatus, the apparatus comprising: a first processing unit, a second processing unit, a matching unit, and an evaluation unit; wherein,

[0020] The first processing unit is used to process the first prompt word and the data in the dataset using the model to be evaluated to obtain first output data, wherein the first output data includes the first output result corresponding to the data;

[0021] The second processing unit is used to process the first output result and the set second prompt word using an open-source language model when the first output result does not match any of the matching rules in the matching rule set, and to obtain second output data. The second output data includes the first matching rule and the second output result corresponding to the data.

[0022] The matching unit is used to add the first matching rule to the matching rule set for matching the next output result of the model to be evaluated;

[0023] The evaluation unit is configured to obtain a score of the model to be evaluated based on the second output result and the label data corresponding to the data; the label data is the expected output result corresponding to the data; and is also configured to obtain an evaluation result of the model to be evaluated based on the scores corresponding to multiple data.

[0024] This invention also provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the steps of the model evaluation method described in this invention.

[0025] This invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, it implements the steps of the model evaluation method described in this invention.

[0026] This invention also provides a computer program product, including computer program instructions that cause a computer to perform the steps of the model evaluation method described in this invention.

[0027] The present invention provides a model evaluation method, apparatus, electronic device, and storage medium. The method includes: processing a first prompt word and data in a dataset using a model to be evaluated to obtain first output data, wherein the first output data includes a first output result corresponding to the data; when the first output result does not match any matching rules in a matching rule set, processing the first output result and a set second prompt word using an open-source language model to obtain second output data, wherein the second output data includes a first matching rule and a second output result corresponding to the data; adding the first matching rule to the matching rule set for matching the next output result of the model to be evaluated; obtaining a score for the model to be evaluated based on the second output result and label data corresponding to the data; wherein the label data is the expected output result corresponding to the data; and obtaining an evaluation result for the model to be evaluated based on the scores corresponding to multiple data points. By employing the technical solution of this invention, regardless of whether the model to be evaluated and the corresponding prompt words change, matching rules are automatically generated using an open-source language model, and the matching rules are automatically updated. This enables the subsequent model evaluation process to automatically use enhanced rules, thereby achieving dynamic adjustment of the matching rules, avoiding manual intervention, and preventing misjudgments caused by mismatch between the matching rules and the output results of the model to be evaluated, thus improving the accuracy of the model evaluation results. Attached Figure Description

[0028] Figure 1 This is a flowchart illustrating the model evaluation method according to an embodiment of the present invention. Figure 1 ;

[0029] Figure 2 This is a flowchart illustrating the model evaluation method according to an embodiment of the present invention. Figure 2 ;

[0030] Figure 3a This is a schematic diagram of the output result of the model to be evaluated in the model evaluation method of this invention.

[0031] Figure 3b This is a schematic diagram of the input data for the open-source language model in the model evaluation method of this invention.

[0032] Figure 3c This is a schematic diagram of output data from an open-source language model in the model evaluation method of this invention.

[0033] Figure 3d This is a schematic diagram of another output data of the open-source language model in the model evaluation method of this invention.

[0034] Figure 4 This is a schematic diagram of the composition structure of the model evaluation device according to an embodiment of the present invention;

[0035] Figure 5 This is a schematic diagram of the hardware composition structure of an electronic device according to an embodiment of the present invention. Detailed Implementation

[0036] The present invention will now be described in further detail with reference to the accompanying drawings and specific embodiments.

[0037] It should be understood that the term "and / or" in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. Additionally, the character " / " in this article generally indicates that the preceding and following related objects have an "or" relationship.

[0038] The terms “first,” “second,” etc., used in the specification and claims of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented, for example, in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0039] This invention provides a model evaluation method. Figure 1 This is a flowchart illustrating the model evaluation method according to an embodiment of the present invention. Figure 1 ;like Figure 1 As shown, the method includes:

[0040] Step 101: Use the model to be evaluated to process the first prompt word and the data in the dataset to obtain the first output data, which includes the first output result corresponding to the data;

[0041] Step 102: If the first output result does not match any of the matching rules in the matching rule set, the first output result and the set second prompt word are processed using an open-source language model to obtain second output data. The second output data includes the first matching rule and the second output result corresponding to the data.

[0042] Step 103: Add the first matching rule to the matching rule set for matching the next output result of the model to be evaluated, and obtain the score of the model to be evaluated based on the second output result and the label data corresponding to the data; the label data is the expected output result corresponding to the data;

[0043] Step 104: Obtain the evaluation result of the model to be evaluated based on the scores corresponding to each of the multiple data points.

[0044] The model evaluation method in this embodiment is applied to an electronic device capable of running large language models. For example, the electronic device may be a personal computer (PC), a smartphone, or a network device such as a server.

[0045] In this embodiment, the model to be evaluated can specifically be a language model or a Large Language Model (LLM), which at least has the function of semantically understanding the input data and outputting corresponding results. As an example, the data input to the model to be evaluated can be a piece of text data, such as question data, and the output data of the model to be evaluated can be the answer data corresponding to the question data. As another example, the input data of the model to be evaluated can also be image data, audio data, etc. Then, through the processing of the input data by the model to be evaluated, such as feature analysis and semantic understanding of image data, and audio-to-text conversion and semantic understanding of audio data, the answer corresponding to the image data or audio data is output.

[0046] In this embodiment, the prompt words (such as the first prompt word, the second prompt word, etc.) can also be called Prompt, model prompt words, etc. Their main function is to provide the model (such as the model to be evaluated, the open source language model, etc.) with contextual information of the input information so that the model (such as the model to be evaluated, the open source language model, etc.) can better understand the intent of the input data.

[0047] In some alternative embodiments, before processing the first prompt word and the data in the dataset using the model to be evaluated, the method further includes: obtaining the dataset and the first prompt word, wherein the dataset includes multiple data and label data corresponding to each data.

[0048] In this embodiment, before evaluating the model to be evaluated, data for evaluating the model is obtained or prepared. This data includes input data (or evaluation data), label data corresponding to each input data, etc. The label data identifies the expected output result corresponding to the input data. For example, if the input data is a question, then the corresponding label data is the answer to that question. In some optional embodiments, the dataset can be an existing dataset used for training and / or evaluating models, such as the Chinese Multi-Modal Language Understanding (CMMLU) dataset or other publicly available datasets; this embodiment does not specifically limit this. In other optional embodiments, the dataset can also be a dataset containing input data and corresponding label data, obtained through data collection or creation.

[0049] In this embodiment, the first prompt word serves as a prompt for the input data of the model to be evaluated, and it can be obtained through prompt words recommended by a specific evaluation framework. As an example, if the dataset is the CMMLU dataset, the first prompt word can be obtained by using prompt words provided by the CMMLU dataset that match it, or by adjusting prompt words that match the CMMLU dataset. The adjustment can be automated or manual; this embodiment does not limit this. As another example, if the dataset does not have matching prompt words, the first prompt word can be obtained automatically or manually; this embodiment does not limit this either.

[0050] For example, the prompt words for matching in the CMMLU dataset could be as follows: "The following is a multiple-choice question about {subject}. Please provide the correct answer directly."

[0051] In this embodiment, a data point from the dataset and a first prompt word are used as input data and fed into the model to be evaluated for processing, thereby obtaining first output data containing the first output result.

[0052] For example, taking a multiple-choice question and its four corresponding options as data, the first output result can be the answer option. Optionally, the first output result can also include the reason or analysis process corresponding to the answer option, etc.

[0053] In this embodiment, the model evaluation system or device is equipped with a set of matching rules. The matching rules in the set (such as a first matching rule) are used to perform rule format matching with the model's output. In some optional embodiments, the set of matching rules may initially be empty. In this case, after obtaining the first output result using the model to be evaluated, the first output result and the second prompt word are directly processed using an open-source language model to obtain second output data including the first matching rule and the second output result. In other optional embodiments, the set of matching rules may initially include at least one matching rule. In this case, after obtaining the first output result using the model to be evaluated, the first output result can be matched with each matching rule in the set to obtain a matching result.

[0054] As one implementation method, the matching rules can be represented using regular expressions. For example, a matching rule represented by a regular expression can be as follows:

[0055] (r'Answer(Options)?(Yes|For):??([ABCD])',3);

[0056] (r'The answer is (is|is) the option?([ABCD])',2);

[0057] (r'Therefore?Choose?:??([ABCD])',1);

[0058] (r'([ABCD])?options(is|is)?correct',1).

[0059] In other alternative embodiments, the matching rules may also adopt other patterns or expressions, which will not be described in detail in this embodiment.

[0060] In this embodiment, the first output result is matched with the matching rules in the matching rule set. If the first output result does not match any of the matching rules in the matching rule set, the first output result and the set second prompt word are processed using an open-source language model to obtain the second output data.

[0061] In this embodiment, an open-source language model is used to process the first output result and the set second prompt word to obtain the second output data. That is, the first output result and the second prompt word are used as input data and fed into the open-source language model for processing to obtain the second output data. For example, the open-source language model is, for instance, the GLM4 model, but it can also be other open-source language models; this embodiment does not limit this. The second prompt word is the prompt word of the open-source language model. Because the selected open-source language model does not change during the evaluation process, the second prompt word can be set empirically according to the desired purpose, or it can be dynamically generated using automated methods. For example, the second prompt word could be, "Please analyze and provide the correct regular expression, and provide the correct option in the expression."

[0062] In this embodiment, the first output result and the second prompt word are processed by an open-source language model. According to the prompt word, the second output data containing the first matching rule and the second output result is output.

[0063] In some optional embodiments, the step of processing the first output result and the set second prompt word using an open-source language model to obtain second output data includes: processing the first output result, all matching rules in the matching rule set, and the set second prompt word using an open-source language model to obtain a second output result corresponding to the first matching rule and the data; wherein the second output result is obtained by the open-source language model based on the analysis and processing of all matching rules in the matching rule set and the first output result.

[0064] In this embodiment, if none of the matching rules in the matching rule set match the first output result, the first output result, all the matching rules in the matching rule set (i.e., the matching rules that do not match the first output result), and the second prompt word are further processed using an open-source language model. Specifically, the first output result, all the current matching rules, and the second prompt word are used as input data to the open-source language model. The open-source language model analyzes and processes the input information to obtain a new matching rule, denoted as the first matching rule. Simultaneously, the open-source language model analyzes and processes the results based on all the matching rules in the matching rule set and the first output result, or analyzes and processes the results based on the first matching rule and the first output result, to obtain the answer that the model to be evaluated cannot match in the current matching rule set (i.e., the second output result).

[0065] In some optional embodiments, before adding the first matching rule to the matching rule set, the method further includes: identifying the second output data based on preset keywords or preset key sentences to extract the first matching rule.

[0066] In this embodiment, the second output data is a series of text contents obtained by processing an open-source language model, such as one or more paragraphs of text. Within this series of contents, the position (or starting position) of the matching rule in the second output data can be determined by locating preset keywords or key sentences. Then, based on this position, the first matching rule is identified and extracted. After extracting the first matching rule, it is added to the matching rule set for format matching of the next output result of the model to be evaluated.

[0067] In some optional embodiments, obtaining the score of the model to be evaluated based on the second output result and the label data corresponding to the data includes: determining the score of the model to be evaluated when the second output result matches the label data; and determining that the model to be evaluated does not receive a score when the second output result does not match the label data.

[0068] In this embodiment, the scoring process for the model to be evaluated is performed using the second output result. Specifically, the second output result is matched with the label data corresponding to that data in the dataset. If the match is consistent, a score is awarded, for example, 1 point is added to the score of the model to be evaluated; if the match is inconsistent, no score is awarded.

[0069] In some alternative embodiments, the method further includes: if the first output result matches a second matching rule in the matching rule set, obtaining a score for the model to be evaluated based on the first output result and the label data corresponding to the data.

[0070] In this embodiment, if there is a matching rule (denoted as the second matching rule) in the matching rule set that matches the first output result, then there is no need to use the open-source language model for processing. Instead, the first output result is used to directly perform the scoring process of the model to be evaluated.

[0071] In some optional embodiments, obtaining the score of the model to be evaluated based on the first output result and the label data corresponding to the data includes: determining the score of the model to be evaluated when the first output result matches the label data; and determining that the model to be evaluated does not receive a score when the first output result does not match the label data.

[0072] In this embodiment, the scoring process of the model to be evaluated is performed using the first output result. Specifically, the first output result is matched with the label data corresponding to that data in the dataset. If the match is consistent, a score is awarded, for example, 1 point is added to the score of the model to be evaluated; if the match is inconsistent, no score is awarded.

[0073] By employing the technical solution of this invention, regardless of whether the model to be evaluated and the corresponding prompt words change, matching rules are automatically generated using an open-source language model, and the matching rules are automatically updated. This enables the subsequent model evaluation process to automatically use enhanced rules, thereby achieving dynamic adjustment of the matching rules, avoiding manual intervention, and preventing misjudgments caused by mismatch between the matching rules and the output results of the model to be evaluated, thus improving the accuracy of the model evaluation results.

[0074] The model evaluation method of this invention will be described below with reference to a specific example.

[0075] Figure 2 This is a flowchart illustrating the model evaluation method according to an embodiment of the present invention. Figure 2 ;like Figure 2 As shown, step 1: The model evaluation system or model evaluation device obtains the dataset and the matching first prompt word; the dataset is, for example, the CMMLU dataset, and the first prompt word is, for example, "The following are multiple-choice questions about {subject}, please give the correct answer directly".

[0076] Step 2: The model evaluation system or device inputs a data point from the dataset and the first prompt word into the model to be evaluated, obtaining the first output result. For example, the first output result can be referenced... Figure 3a As shown:

[0077] According to the material, the revised "Law of the People's Bank of China," promulgated in 1995, clearly stipulates that the objective of my country's monetary policy is to maintain the stability of the currency value and thereby promote economic growth. Therefore, the answer can be determined to be option B.

[0078] Step 3: The result matching module of the model evaluation system or model evaluation device performs format matching between the first output result and each matching rule in the matching rule set.

[0079] For example, the matching rule, expressed using a regular expression, can be represented as:

[0080] (r'Answer(Options)?(Yes|For):??([ABCD])',3);

[0081] (r'The answer is (is|is) the option?([ABCD])',2);

[0082] (r'Therefore?Choose?:??([ABCD])',1);

[0083] (r'([ABCD])?options(is|is)?correct',1).

[0084] Step 4: If the first output result matches any matching rule, proceed to the evaluation score calculation process. Match the first output result of the model to be evaluated with the label data corresponding to the data. If they match, 1 point is awarded; if they do not match, no points are awarded.

[0085] Assumption Figure 3a The first output shown can match a matching rule "(r'Answer(is|is)Option?([ABCD])',2)", and the result of the match is "B" in the option.

[0086] Step 5: If the first output does not match any of the matching rules, proceed to the parameter configuration process. The parameter configuration module inputs the first output, all current matching rules, and the set second prompt word into the open-source language model for processing.

[0087] For example, taking the open-source language model GLM4 as an example, the second prompt word for the GLM4 model can be: "The following is a statement that did not match the correct option after the regular expression matching was completed. Please analyze and give the correct regular expression, and give the correct option in the statement. If there is no correct option, the answer does not conform to the regular expression as follows: {Regular expression in the result processing module} The statement is as follows: {Model output of a single test question that was not matched by the model}.

[0088] For example, the input data for an open-source language model can be as follows: Figure 3b As shown:

[0089] The following is a description of a situation where no correct option was matched after the regular expression matching was completed. Please analyze and provide the correct regular expression, and give the correct option in the expression. If there is no correct option, then the answer is incorrect.

[0090] The regular expression is as follows:

[0091] (r'Answer(Options)?(Yes|For):??([ABCD])',3),

[0092] (r'Answer(is|is) option?([ABCD])',2),

[0093] (r'Therefore?Choose?:??([ABCD])',1),

[0094] (r'([ABCD])?options(is|is)?correct',1),

[0095] The statements are as follows: 1. According to the information in the question, GDP is an important indicator of a country's international status and development direction, but some people believe that GDP cannot reflect a country's status and development direction. 2. From a historical perspective, China's GDP was 33% of the world's total in 1840, while in 1949 it was only 6%. This shows that GDP can indeed reflect economic development, but it cannot reflect a country's status and development direction. 3. Therefore, option C, "GDP can only reflect the economic situation at that time, so it cannot reflect a country's status and development direction," is the most reasonable. 4. Option A, "The economic nature that determines GDP changes is what truly determines a country's status and development direction," is too absolute. Options B, "GDP can indeed reflect economic development and can confirm a country's status and development direction," and D, "The trend of GDP's international status changes often moves in the opposite direction to the trend of a country's international status changes," are both inconsistent with the information in the question.

[0096] Step 6: The open-source language model takes the parameter-configured data as input, outputs the matching rules applicable to this data, and obtains the answer corresponding to this data (i.e., the second output result). For example, one output data of the open-source language model can be referred to... Figure 3c As shown:

[0097] Based on the given regular expression and description, we can see that the regular expression failed to match the correct option C. This is because the format of the correct option in the expression does not completely match the provided regular expression. In the expression, the correct option appears in the form of "Option C is the most reasonable," while the provided regular expression mainly matches keywords such as "correct," "is," or "is."

[0098] To match the correct option C in the expression, we need to adjust the regular expression to include matches for descriptions like "most reasonable". Here is the adjusted regular expression:

[0099] regex1(r'Options ([ABCD]) are the most reasonable',1),

[0100] This regular expression works as follows:

[0101] Options: Matches the word "options" in the text.

[0102] ([ABCD]): Capture group, matching one of A, B, C or D.

[0103] Most reasonable: Matches the phrase "most reasonable" in the text.

[0104] Using this regular expression, we can extract the letter C of the correct answer from the expression.

[0105] The following is the matching process using this regular expression:

[0106] The most reasonable option for regex1 is ([ABCD]).

[0107] The text in the matching statement is: 1. Therefore, option C, "GDP can only reflect the economic situation at that time, so it cannot reflect the country's status and development direction," is the most reasonable.

[0108] The correct answer is C.

[0109] In some optional embodiments, if the first output does not output a correct option, then after processing the first output, all current matching rules, and the set second prompt word using an open-source language model, if the resulting second output also does not contain a correct option, then the second output will also include expressions that do not conform to any regular expression. In this case, one type of output data from the open-source language model can be referenced. Figure 3d As shown:

[0110] The following is a description of a situation where no correct option was matched after the regular expression matching was completed. Please analyze and provide the correct regular expression, and give the correct option in the expression. If there is no correct option, then the answer is incorrect.

[0111] The regular expression is as follows:

[0112] (r'Answer(Options)?(Yes|For):??([ABCD])',3),

[0113] (r'Answer(is|is) option?([ABCD])',2),

[0114] (r'Therefore?Choose?:??([ABCD])',1),

[0115] (r'([ABCD])?options(is|is)?correct',1),

[0116] The statement is as follows:

[0117] This viewpoint belongs to the biological origin theory of education.

[0118] Based on the given regular expression and description, we can see that the description does not directly indicate which option is correct, nor does it use keywords from regular expressions such as "answer," "yes," "for," "select," or "correct." The description merely states a viewpoint and does not provide enough information to match any of the regular expressions.

[0119] Since the expression does not contain the letter of the correct option (A, B, C, or D) and does not use any keywords in the regular expression, we cannot extract the correct option from the expression.

[0120] In this situation, the correct approach is to answer "not applicable," because the statement does not provide enough information to match the given regular expression and determine the correct option.

[0121] Answer: Does not meet the requirements.

[0122] Step 7: The result segmentation module of the model evaluation system or model evaluation device extracts a new matching rule (denoted as the first matching rule) based on the output data of the open source language model, adds the first matching rule to the matching rule set, and inputs the second output result to the evaluation score calculation module. The evaluation score calculation module matches the second output result with the label data corresponding to the data. If the match is consistent, 1 point is awarded; if the match is inconsistent, no point is awarded.

[0123] For example, with Figure 3c Taking the output data shown as an example, the preset keywords or key sentences are used to locate the preset key sentences in the output data; for example, the preset key sentence is: "The following is the adjusted regular expression"; further search the next line: the line of code immediately following this sentence is the adjusted matching rule.

[0124] Step 8: Summarize the scores corresponding to all output results of the model to be evaluated, and obtain the total score for the model to be evaluated based on the scores.

[0125] Based on the above embodiments, this invention also provides a model evaluation device, which is applied to an electronic device. Figure 4 This is a schematic diagram of the composition structure of the model evaluation device according to an embodiment of the present invention; as shown below. Figure 4 As shown, the device includes: a first processing unit 21, a second processing unit 22, a matching unit 23, and an evaluation unit 24; wherein,

[0126] The first processing unit 21 is used to process the first prompt word and the data in the dataset using the model to be evaluated to obtain first output data, wherein the first output data includes the first output result corresponding to the data;

[0127] The second processing unit 22 is used to process the first output result and the set second prompt word using an open-source language model when the first output result does not match any of the matching rules in the matching rule set, so as to obtain second output data. The second output data includes the first matching rule and the second output result corresponding to the data.

[0128] The matching unit 23 is used to add the first matching rule to the matching rule set for matching the next output result of the model to be evaluated;

[0129] The evaluation unit 24 is used to obtain a score of the model to be evaluated based on the second output result and the label data corresponding to the data; the label data is the expected output result corresponding to the data; and is also used to obtain an evaluation result of the model to be evaluated based on the scores corresponding to multiple data.

[0130] In some optional embodiments of the present invention, the second processing unit 22 is used to process the first output result, all matching rules in the matching rule set, and the set second prompt word using an open-source language model to obtain a second output result corresponding to the first matching rule and the data; wherein, the second output result is obtained by the open-source language model based on all matching rules in the matching rule set and the first output result.

[0131] In some optional embodiments of the present invention, the second processing unit 22 is used to identify the second output data based on preset keywords or preset key sentences and extract the first matching rule.

[0132] In some optional embodiments of the present invention, the evaluation unit 24 is further configured to obtain a score of the model to be evaluated based on the first output result and the label data corresponding to the data, when the first output result matches the second matching rule in the matching rule set.

[0133] In some optional embodiments of the present invention, the evaluation unit 24 is configured to determine the score of the model to be evaluated when the first output result matches the label data; determine that the model to be evaluated receives no score when the first output result does not match the label data; and / or, determine the score of the model to be evaluated when the second output result matches the label data; and determine that the model to be evaluated receives no score when the second output result does not match the label data.

[0134] In some optional embodiments of the present invention, the device further includes an identification unit, which is used to identify the second output data based on preset keywords or key sentences, identify and extract the first matching rule, and send the first matching rule to the matching unit 23.

[0135] In some optional embodiments of the present invention, the apparatus further includes an acquisition unit for acquiring a dataset and a first prompt word, wherein the dataset includes multiple data items and label data corresponding to each data item.

[0136] In this embodiment of the invention, the first processing unit 21, the second processing unit 22, the matching unit 23, the evaluation unit 24, the identification unit, and the acquisition unit in the device can all be implemented by a central processing unit (CPU), a digital signal processor (DSP), a microcontroller unit (MCU), or a field-programmable gate array (FPGA) in practical applications.

[0137] Combination Figure 2 The model evaluation method shown is as follows: Figure 2 The result matching module can be implemented by the matching unit 23, that is, the matching unit 23 matches the first output result with the matching rules in the matching rule set to obtain the matching result; the second processing unit 22 is used to process the first output result and the set second prompt word using an open source language model when the matching result obtained by the matching unit 23 is that the first output result does not match any of the matching rules in the matching rule set to obtain the second output data.

[0138] Figure 2 Both the parameter configuration module and the result segmentation module can be implemented through the second processing unit 22. Figure 2 The evaluation score calculation module can be implemented through the evaluation unit 24.

[0139] It should be noted that the model evaluation device provided in the above embodiments is only illustrated by the division of the above program modules when performing model evaluation. In practical applications, the above processing can be assigned to different program modules as needed, that is, the internal structure of the device can be divided into different program modules to complete all or part of the processing described above. In addition, the model evaluation device and the model evaluation method embodiments provided in the above embodiments belong to the same concept, and the specific implementation process can be found in the method embodiments, which will not be repeated here.

[0140] This invention also provides an electronic device. Figure 5 This is a schematic diagram of the hardware composition structure of the electronic device according to an embodiment of the present invention, such as... Figure 5 As shown, the electronic device includes a memory 32, a processor 31, and a computer program stored in the memory 32 and executable on the processor 31. When the processor 31 executes the program, it implements the steps of the model evaluation method of the present invention.

[0141] Optionally, the various components in the electronic device are coupled together via a bus system 33. It is understood that the bus system 33 is used to implement communication between these components. In addition to a data bus, the bus system 33 also includes a power bus, a control bus, and a status signal bus. However, for clarity, in... Figure 5 The general will label all buses as Bus System 33.

[0142] It is understood that memory 32 can be volatile memory or non-volatile memory, or both. Non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), magnetic random access memory (FRAM), flash memory, magnetic surface memory, optical disc, or compact disc read-only memory (CD-ROM); magnetic surface memory can be disk storage or magnetic tape storage. Volatile memory can be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of RAM are available, such as Static Random Access Memory (SRAM), Synchronous Static Random Access Memory (SSRAM), Dynamic Random Access Memory (DRAM), Synchronous Dynamic Random Access Memory (SDRAM), Double Data Rate Synchronous Dynamic Random Access Memory (DDRSDRAM), Enhanced Synchronous Dynamic Random Access Memory (ESDRAM), SyncLink Dynamic Random Access Memory (SLDRAM), and Direct Rambus Random Access Memory (DRRAM).The memory 32 described in the embodiments of the present invention is intended to include, but is not limited to, these and any other suitable types of memory.

[0143] The methods disclosed in the above embodiments of the present invention can be applied to or implemented by processor 31. Processor 31 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the integrated logic circuit of the hardware in processor 31 or by instructions in software form. The processor 31 may be a general-purpose processor, DSP, or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Processor 31 can implement or execute the methods, steps and logic block diagrams disclosed in the embodiments of the present invention. The general-purpose processor may be a microprocessor or any conventional processor, etc. The steps of the methods disclosed in the embodiments of the present invention can be directly manifested as being executed by a hardware decoding processor, or being executed by a combination of hardware and software modules in the decoding processor. The software modules may be located in a storage medium, which is located in memory 32. Processor 31 reads the information in memory 32 and completes the steps of the aforementioned method in combination with its hardware.

[0144] In an exemplary embodiment, the electronic device may be implemented by one or more application-specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), FPGAs, general-purpose processors, controllers, MCUs, microprocessors, or other electronic components to perform the aforementioned method.

[0145] In an exemplary embodiment, the present invention also provides a computer-readable storage medium, such as a memory 32 including a computer program, which can be executed by a processor 31 of an electronic device to perform the steps described in the foregoing method. The computer-readable storage medium may be a memory such as FRAM, ROM, PROM, EPROM, EEPROM, Flash Memory, magnetic surface memory, optical disc, or CD-ROM; or it may be various devices including one or any combination of the above-mentioned memories.

[0146] The computer-readable storage medium provided in the embodiments of the present invention stores a computer program thereon, which, when executed by a processor, implements the steps of the model evaluation method of the embodiments of the present invention.

[0147] This application also provides a computer program product, including a computer program that can be executed by an electronic device (such as a processor 31 of an electronic device) to complete the steps of any of the aforementioned model evaluation methods.

[0148] The methods disclosed in the several method embodiments provided in this application can be arbitrarily combined without conflict to obtain new method embodiments.

[0149] The features disclosed in the several product embodiments provided in this application can be arbitrarily combined without conflict to obtain new product embodiments.

[0150] The features disclosed in the several method or device embodiments provided in this application can be arbitrarily combined without conflict to obtain new method or device embodiments.

[0151] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods, such as: multiple units or components can be combined, or integrated into another system, or some features can be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the various components shown or discussed can be through some interfaces, and the indirect coupling or communication connection between devices or units can be electrical, mechanical, or other forms.

[0152] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the units may be selected to achieve the purpose of this embodiment according to actual needs.

[0153] In addition, in the various embodiments of the present invention, each functional unit can be integrated into one processing unit, or each unit can be a separate unit, or two or more units can be integrated into one unit; the integrated unit can be implemented in hardware or in the form of hardware plus software functional units.

[0154] Those skilled in the art will understand that all or part of the steps of the above method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps of the above method embodiments. The aforementioned storage medium includes various media that can store program code, such as mobile storage devices, ROM, RAM, magnetic disks, or optical disks.

[0155] Alternatively, if the integrated units of this invention are implemented as software functional modules and sold or used as independent products, they can also be stored in a computer-readable storage medium. Based on this understanding, the technical solutions of the embodiments of this invention, or the parts that contribute to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the methods described in the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as mobile storage devices, ROM, RAM, magnetic disks, or optical disks.

[0156] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A model evaluation method, characterized in that, The method includes: The model to be evaluated is used to process the first prompt word and the data in the dataset to obtain the first output data, which includes the first output result corresponding to the data. If the first output result does not match any of the matching rules in the matching rule set, the first output result and the set second prompt word are processed using an open-source language model to obtain second output data. The second output data includes the first matching rule and the second output result corresponding to the data. The first matching rule is added to the matching rule set for matching the next output result of the model to be evaluated, and a score of the model to be evaluated is obtained based on the second output result and the label data corresponding to the data; the label data is the expected output result corresponding to the data. The evaluation result of the model to be evaluated is obtained based on the scores corresponding to each of the multiple data points.

2. The method according to claim 1, characterized in that, The process of using an open-source language model to process the first output result and the set second prompt word to obtain second output data includes: The first output result, all matching rules in the matching rule set, and the set second prompt word are processed using an open-source language model to obtain a second output result corresponding to the first matching rule and the data; wherein, the second output result is obtained by the open-source language model based on all matching rules in the matching rule set and the first output result.

3. The method according to claim 2, characterized in that, Before adding the first matching rule to the matching rule set, the method further includes: The second output data is identified based on preset keywords or preset key phrases, and the first matching rule is extracted.

4. The method according to claim 1, characterized in that, The method further includes: If the first output matches the second matching rule in the matching rule set, the score of the model to be evaluated is obtained based on the first output and the label data corresponding to the data.

5. The method according to claim 4, characterized in that, The step of obtaining the score of the model to be evaluated based on the first output result and the label data corresponding to the data includes: When the first output result matches the label data, the score of the model to be evaluated is determined; when the first output result does not match the label data, the score of the model to be evaluated is determined to be zero. And / or, obtaining the score of the model to be evaluated based on the second output result and the label data corresponding to the data includes: When the second output matches the label data, the score of the model to be evaluated is determined; when the second output does not match the label data, the score of the model to be evaluated is determined to be zero.

6. The method according to claim 1, characterized in that, Before adding the first matching rule to the matching rule set, the method further includes: The second output data is identified based on preset keywords or key phrases, and the first matching rule is identified and extracted.

7. The method according to claim 1, characterized in that, Before processing the first prompt word and the data in the dataset using the model to be evaluated, the method further includes: Obtain the dataset and the first prompt word, wherein the dataset includes multiple data points and the corresponding label data for each data point.

8. A model evaluation device, characterized in that, The device includes: a first processing unit, a second processing unit, a matching unit, and an evaluation unit; wherein... The first processing unit is used to process the first prompt word and the data in the dataset using the model to be evaluated to obtain first output data, wherein the first output data includes the first output result corresponding to the data; The second processing unit is used to process the first output result and the set second prompt word using an open-source language model when the first output result does not match any of the matching rules in the matching rule set, so as to obtain second output data. The second output data includes the first matching rule and the second output result corresponding to the data. The matching unit is used to add the first matching rule to the matching rule set for matching the next output result of the model to be evaluated; The evaluation unit is configured to obtain a score of the model to be evaluated based on the second output result and the label data corresponding to the data; the label data is the expected output result corresponding to the data; and is also configured to obtain an evaluation result of the model to be evaluated based on the scores corresponding to multiple data.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the program implements the steps of the method according to any one of claims 1 to 7.

10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps of the method according to any one of claims 1 to 7.

11. A computer program product, characterized in that, It includes computer program instructions that cause a computer to perform the steps of the method according to any one of claims 1 to 7.