Model evaluation method and device and computing equipment

By generating fine-grained requirements evaluation constraints, adaptive evaluation of model evaluation results is solved, and the problem of inaccurate and flexible model evaluation methods in the prior art is solved, achieving higher evaluation accuracy and adaptability.

CN119939190APending Publication Date: 2025-05-06SHUXING TECH (BEIJING) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510074625.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-17
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

In the prior art, the model evaluation method cannot be adapted to different application scenarios, resulting in inaccurate and flexible evaluation results and lack of adaptability.

Method used

By obtaining evaluation use cases and reference results, fine-grained requirements evaluation constraints are generated, and the evaluation results are evaluated based on these constraints, so as to achieve adaptability and flexibility of model evaluation results.

Benefits of technology

Improve the accuracy and flexibility of model evaluation, so that the evaluation results can better adapt to different types and inference needs evaluation use cases, and are suitable for a variety of application scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119939190A_ABST
    Figure CN119939190A_ABST
Patent Text Reader

Abstract

Embodiments of the invention provide a model evaluation method and apparatus, and a computing device. The model evaluation method comprises the steps of obtaining an evaluation case and a corresponding reference result; according to the evaluation case and the reference result, a demand evaluation constraint corresponding to the reasoning demand of the evaluation case is generated, and the demand evaluation constraint is used for constraining a positive factor and / or a negative factor of the evaluation case; generating a to-be-evaluated result for the evaluation case through the to-be-evaluated model; performing evaluation processing on the evaluation case and the corresponding to-be-evaluated result through an evaluation model based on the demand evaluation constraint to obtain a reasoning quality evaluation result corresponding to the to-be-evaluated result of the evaluation case; and determining an evaluation result of the to-be-evaluated model according to a reasoning quality evaluation result corresponding to a to-be-evaluated result of the evaluation case. According to the method, fine-grained demand evaluation constraints are adaptively generated according to the reasoning demand characteristics of the evaluation cases, and the method is better generalized to reasoning quality evaluation of the evaluation cases of various different types and different reasoning demands.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of this specification relate to the field of computer technology, and more particularly to a model evaluation method, apparatus, and computing device. Background Art

[0002] With the rapid development of computer technology, machine learning, and deep learning technology, machine learning models have been widely used in many fields such as natural language processing, machine translation, speech synthesis, and image generation. Large language models have been born accordingly. Large language models have been proven to have excellent performance in understanding and generating natural language texts. Large language model technology has developed rapidly, and it has shown strong capabilities in instruction-following, problem-solving, and open-ended chatting, which has brought greater challenges to how to evaluate the processing capabilities of such large language models.

[0003] In the prior art, a certain number of evaluation sets are generally created, and each evaluation question in the evaluation set is used to ask a large language model to obtain an answer. The evaluation questions and the answers of the large language model are then input into the evaluation model to obtain the evaluation results of the answer quality of the evaluation questions, so as to evaluate the analysis and processing capabilities of the large language model.

[0004] However, in the above evaluation methods, the evaluation model uses a unified method to evaluate different evaluation questions and large language model answers, which cannot adapt to various application scenarios. The evaluation results of the answer quality are not accurate and flexible, resulting in poor accuracy and flexibility of model evaluation. Therefore, a more accurate and flexible model evaluation solution is urgently needed. Summary of the invention

[0005] In view of this, an embodiment of this specification provides a model evaluation method. One or more embodiments of this specification also relate to a model evaluation device, a computing device, a computer-readable storage medium and a computer program product to solve the technical defects existing in the prior art.

[0006] According to a first aspect of an embodiment of this specification, a model evaluation method is provided, comprising: Obtain evaluation cases and corresponding reference results; Generate, according to the evaluation case and the reference result, a requirement evaluation constraint corresponding to the reasoning requirement of the evaluation case, wherein the requirement evaluation constraint is used to constrain the positive factor and / or negative factor of the evaluation case; Generate an evaluation result for the evaluation case through the model to be evaluated; The evaluation model is used to evaluate the evaluation case and the corresponding result to be evaluated based on the requirement evaluation constraint, so as to obtain the reasoning quality evaluation result corresponding to the result to be evaluated of the evaluation case; The evaluation result of the model to be evaluated is determined according to the reasoning quality evaluation result corresponding to the result to be evaluated of the evaluation case.

[0007] According to a second aspect of an embodiment of this specification, a model evaluation device is provided, comprising: An acquisition module, configured to acquire evaluation cases and corresponding reference results; A first generating module is configured to generate a requirement evaluation constraint corresponding to the reasoning requirement of the evaluation case according to the evaluation case and the reference result, wherein the requirement evaluation constraint is used to constrain the positive factor and / or negative factor of the evaluation case; A second generating module is configured to generate a result to be evaluated for the evaluation case through the model to be evaluated; A first evaluation module is configured to evaluate the evaluation case and the result to be evaluated based on the requirement evaluation constraint through an evaluation model, and obtain a reasoning quality evaluation result corresponding to the result to be evaluated of the evaluation case; The second evaluation module is configured to determine the evaluation result of the model to be evaluated according to the reasoning quality evaluation result corresponding to the result to be evaluated of the evaluation case.

[0008] According to a third aspect of an embodiment of this specification, a computing device is provided, including: Memory and processor; The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions. When the computer-executable instructions are executed by the processor, the steps of the above-mentioned model evaluation method are implemented.

[0009] According to a fourth aspect of the embodiments of this specification, a computer-readable storage medium is provided, which stores computer-executable instructions, and when the instructions are executed by a processor, the steps of the above-mentioned model evaluation method are implemented.

[0010] According to a fifth aspect of the embodiments of this specification, a computer program product is provided, comprising a computer program / instruction, which implements the steps of the above-mentioned model evaluation method when executed by a processor.

[0011] The embodiments of this specification provide a model evaluation method, which realizes the reasoning requirements based on the evaluation case and the reference results, and adaptively generates the requirement evaluation constraints corresponding to the evaluation case, so as to guide the evaluation model to evaluate the evaluation case and the corresponding results to be evaluated according to the corresponding positive factors and / or negative factors based on the requirement evaluation constraints, so as to guide the evaluation model to evaluate the reasoning requirements based on the evaluation case. In this way, each evaluation case can adapt to its own reasoning requirement characteristics, and adaptively generate a set of fine-grained requirement evaluation constraints, which specifically constrain the detailed positive factors and / or negative factors of the evaluation case, and can be better generalized to the reasoning quality evaluation of evaluation cases of various types and different reasoning requirements, and can be adapted to a variety of application scenarios. More accurate and flexible automated reasoning quality evaluation results can be obtained, and then more accurate and flexible model evaluation results can be obtained. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] Figure 1 is a flow chart of a model evaluation method provided by an embodiment of this specification; Figure 2 It is a schematic diagram of a generation process of a requirement evaluation constraint corresponding to an evaluation case provided by an embodiment of this specification; Figure 3 is a process flow chart of a model evaluation method provided by an embodiment of this specification; Figure 4 is a structural schematic diagram of a model evaluation device provided by an embodiment of this specification; Figure 5 It is a structural block diagram of a computing device provided by an embodiment of this specification. DETAILED DESCRIPTION

[0013] Many specific details are described in the following description to facilitate a full understanding of this specification. However, this specification can be implemented in many other ways than those described herein, and those skilled in the art can make similar generalizations without violating the connotation of this specification, so this specification is not limited to the specific implementation disclosed below.

[0014] The terms used in one or more embodiments of this specification are only for the purpose of describing specific embodiments, and are not intended to limit one or more embodiments of this specification. The singular forms of "a" and "the" used in one or more embodiments of this specification and the appended claims are also intended to include plural forms, unless the context clearly indicates other meanings. It should also be understood that the term "and / or" used in one or more embodiments of this specification refers to and includes any or all possible combinations of one or more associated listed items.

[0015] It should be understood that although the terms first, second, etc. may be used to describe various information in one or more embodiments of this specification, this information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of one or more embodiments of this specification, the first may also be referred to as the second, and similarly, the second may also be referred to as the first. Depending on the context, the word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".

[0016] In addition, it should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in one or more embodiments of this specification are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.

[0017] First, the terms involved in one or more embodiments of this specification are explained.

[0018] Large Language Models (LLMS): refers to a type of language processing model with a large number of parameters. These models are trained through deep learning technology and can show excellent performance in a variety of natural language processing tasks. In recent years, with the growth of computing resources and the advancement of algorithms, large language models have become one of the important breakthroughs in the field of natural language processing. Large language models usually contain hundreds of millions, tens of billions, hundreds of billions, trillions or even more than ten trillion model parameters. Large language models can also be called cornerstone models / foundation models. Large language models are pre-trained using large-scale unlabeled corpora to produce pre-trained models with more than 100 million parameters. Such models can adapt to a wide range of downstream tasks and have good generalization capabilities. For example, large-scale language models (LLM) and multi-modal pre-training models have shown strong capabilities in instruction-following, problem-solving and open-ended chatting, and can be applied to many fields such as natural language processing, machine translation, speech synthesis, image generation, etc.

[0019] Question Instance: Instance-Wise refers to every question in the assessment question set.

[0020] Self-Adaptive: Self-Adaptive can be understood as customization.

[0021] Unbiased scoring: Objectively scoring capabilities. In this setting, only the degree to which LLMs address the reasoning requirements in the input evaluation case is considered, without considering other factors, such as the impact of diverse human preferences, which may cause score fluctuations.

[0022] LLM-as-a-judge: refers to using a large language model as an evaluation model to evaluate the quality of content generated by other models.

[0023] Evaluation Model: It can be a large language model used to evaluate the quality of content generated by other models. In other words, the evaluation model can be used to evaluate the analytical processing capabilities of other models.

[0024] It should be noted that the rapid development of large language model technology has led to its widespread use. It has shown strong capabilities in instruction-following, problem-solving, and open-ended chatting, which poses a greater challenge to how to evaluate such open-ended cognitive abilities. The general evaluation method is generally to establish a certain number of question sets for various capabilities of the model, such as mathematics, reasoning, knowledge questions and answers, and creation, and then use these question sets to ask each large language model (LLMs) to judge the ability level of each LLM based on the quality of their answers.

[0025] Evaluating these large language models in open-ended question answering scenarios is a major challenge. Metric-based evaluations provide speed and convenience, but often fall short in accuracy and effectiveness due to the diversity of ground truth. In contrast, human-based evaluations provide reliable evaluations but require a lot of resources.

[0026] Therefore, the LLM-as-a-judge paradigm attempts to strike a balance between automation and manual evaluation. It uses proprietary evaluation LLMs (Judge Model) to evaluate the quality of model answers. It uses predefined principles, such as the 3H principle (usefulness, honesty, and harmlessness), to evaluate the quality of model answers. However, the above methods use unified, question-agnostic evaluation rules. For example, Prometheus uses customized evaluation rules, but each evaluation rule is not directly related to a specific question. These methods ignore the uniqueness of each question, because each question contains its own emphasis, which can be divided into primary and secondary content. Unified evaluation rules cannot well reflect the characteristics of these answer requirements, and unified evaluation rules are not easy to be understood and generalized by the Jusge model to a variety of actual scenarios, which will lead to large evaluation errors.

[0027] Therefore, in order to be able to adaptively align the specific evaluation process, the embodiment of this specification provides a model evaluation scheme. Different from the coarse-grained unified evaluation method, the model evaluation scheme provided by the embodiment of this specification can adaptively and automatically generate a set of fine-grained evaluation constraints for each evaluation case. Under the constraints of the primary and secondary needs of the evaluation case itself, detailed positive factors and negative factors are defined step by step, and negative factors are used to punish content that does not conform to objective facts. At the same time, some background knowledge can also be added to the evaluation constraints, which can guide the evaluation model to better understand the background knowledge and better generalize to the evaluation of various reasoning results.

[0028] In this specification, a model evaluation method is provided. This specification also relates to a model evaluation device, a computing device, a computer-readable storage medium, and a computer program product, which are described in detail one by one in the following embodiments.

[0029] See also Figure 1 , Figure 1 A flow chart of a model evaluation method provided according to an embodiment of the present specification is shown, which specifically includes the following steps.

[0030] Step 102: Obtain evaluation cases and corresponding reference results.

[0031] Among them, the evaluation case refers to the input that needs to be analyzed by the model to be evaluated. The evaluation case is used to evaluate the reasoning quality of the model to be evaluated. For example, the evaluation case can be a question, a problem, a task instruction, a scenario setting, a story creation, a translation, etc. The reference result can refer to the standard reasoning result given for the evaluation case. For example, if the evaluation case is a question, the reference result is the standard answer stored in the question bank or manually given. If the evaluation case is a task instruction, the reference result is the instruction execution result manually marked for the task indicator or the standard execution result of the historical output.

[0032] It should be noted that the model evaluation method provided in the embodiments of this specification can be applied to the client or to the cloud server. If applied to the client, the evaluation case and the corresponding reference result can be obtained through the cloud server or other devices to perform reasoning quality evaluation on the evaluation case on the client; if applied to the cloud server, the evaluation case and the corresponding reference result can be obtained locally from the client, other servers or the cloud server to perform reasoning quality evaluation on the evaluation case on the cloud server. Among them, the cloud server can be connected to one or more clients through a local area network connection, a wide area network connection, an Internet connection, or other types of data networks. The client here can include but is not limited to: smart phones, tablet computers, laptops, PDAs, personal computers, smart home devices, vehicle-mounted devices, etc. The client can interact with the user through a graphical user interface to achieve management such as uploading, querying, downloading, and evaluating the evaluation case, thereby realizing the model evaluation method provided in the embodiments of this specification.

[0033] In actual implementation, a set number of evaluation cases and reference results corresponding to each evaluation case can be obtained from a local database, a third-party platform and / or the Internet, and each evaluation case and the corresponding reference result can be used as an evaluation data set. When performing reasoning quality assessment, any evaluation case can be obtained from the evaluation data set, and the corresponding reference result can be obtained for subsequent generation of corresponding demand assessment constraints, and reasoning quality assessment can be performed on the evaluation case. For example, the evaluation case is a question, and the evaluation data set is an evaluation question set. Any question instance in the evaluation question set can be used as an evaluation case.

[0034] Among them, different evaluation cases in the evaluation data set can come from different application scenarios, or they can all come from the same application scenario. For example, in the news writing scenario, the evaluation case is the task instruction "Please give the press release corresponding to XXXX"; in the knowledge question and answer scenario, the evaluation case can be the question "Why is Li Bai so much younger than Du Fu?" etc.

[0035] As an example, the evaluation case may be "Why is Li Bai so much younger than Du Fu?", and the corresponding reference result may be "Li Bai (701-762) and Du Fu (712-770) are famous poets in China's Tang Dynasty, and are respectively known as the "Poet Immortal" and the "Poet Sage". Although Li Bai is 11 years older than Du Fu, they both have a high status in literary achievement and influence. The poetry styles and themes of Li Bai and Du Fu are distinctive. Li Bai's poems are known for their boldness, unrestrainedness, freshness, and loftiness. His poems have a strong personality and romanticism. Du Fu's poems are known for their rigor, melancholy, realism, and concern for people's livelihood. His poems reflect social reality and people's suffering. Although Li Bai is older than Du Fu, their poetry achievements and influence are world-famous. In literary history, they are known as "Li Du", representing the pinnacle of Tang poetry."

[0036] Specifically, the evaluation data set can be expressed as the following formula (1): …… (1) in, Represents the evaluation dataset Evaluation use cases; For the The reference results corresponding to the evaluation cases.

[0037] It should be noted that, for each evaluation case in the evaluation data set, a subsequent evaluation process can be performed to obtain the corresponding reasoning quality evaluation result.

[0038] Step 104: Generate a requirement evaluation constraint corresponding to the reasoning requirement of the evaluation case based on the evaluation case and the reference result, wherein the requirement evaluation constraint is used to constrain the positive factor and / or negative factor of the evaluation case.

[0039] The reasoning requirements of the evaluation case are used to indicate the content that the reasoning results of the evaluation case need to include, which can usually include the core content of the reasoning analysis of the evaluation case, that is, the main content, and at the same time, cover as much additional information as possible that can enhance the depth and breadth of the reasoning results, such as secondary content. The demand evaluation constraint refers to the fine-grained evaluation conditions that are adaptively generated according to the reasoning requirements of the evaluation case, and are used to constrain the positive factors and / or negative factors of the evaluation case.

[0040] A positive factor refers to positive feedback given when the reasoning result corresponding to the evaluation case includes correct or expected content. This feedback can be a numerical increase, score improvement, or other forms of optimization. A negative factor refers to negative feedback given when the reasoning result corresponding to the evaluation case includes incorrect or unexpected content. This feedback can be a numerical decrease, score reduction, or other forms of suppression. As an example, for an evaluation case that is a question, a positive factor can be a score item for the answer to the question, and a negative factor can be a deduction item.

[0041] In actual implementation, the requirement evaluation constraints corresponding to the reasoning requirements of the evaluation case are generated based on the evaluation case and the reference results. The requirement evaluation constraints are an unbiased scoring, that is, only the degree to which the reasoning results solve the reasoning requirements of the evaluation case is considered, and other factors, such as the impact of diversified human preferences, are not considered, which may cause evaluation fluctuations.

[0042] It should be noted that, according to the reasoning requirements of the evaluation case and the reference results, the demand evaluation constraints corresponding to the evaluation case can be adaptively generated, so as to facilitate the subsequent evaluation of the evaluation case and the corresponding results to be evaluated by guiding the evaluation model based on the demand evaluation constraints.

[0043] In an optional implementation of this embodiment, the reasoning requirement includes a primary requirement content and a secondary requirement content; according to the evaluation case and the reference result, the requirement evaluation constraint corresponding to the reasoning requirement of the evaluation case is generated, including: Extract the corresponding primary and secondary requirements based on the evaluation cases and reference results; According to the set indicator allocation strategy, corresponding indicators are allocated to the extracted primary and secondary demand contents, and the demand assessment constraints corresponding to the evaluation case are generated.

[0044] Among them, the main content usually refers to the part that directly feeds back the core content of the evaluation case, the core and key part of the evaluation case, that is, the main demand or goal for reasoning the evaluation case to obtain the reasoning result, that is, the information that must be included to ensure the accuracy and completeness of the reasoning result. The main content is the basis for reasoning and analyzing the evaluation case, and directly responds to the key information that the questioner wants to know. Secondary content refers to information that can supplement and support the main content, which may not be necessary, but can provide a more comprehensive understanding and increase the value of the reasoning result. Secondary content is usually used to provide additional and rich information to help gain a deeper understanding.

[0045] In actual implementation, the index allocation strategy is a constraint strategy for allocating indicators to each main demand content and secondary demand content. The index allocation strategy can constrain the number of main demand content and secondary demand content, as well as the indicators corresponding to each main demand content and secondary demand content. Among them, the indicator can refer to the degree of quality of reasoning used to measure the main demand content and secondary demand content. The indicator can be a specific score or weight, such as 2 points, 5 points, 0.8, 0.6, etc., or a grade of excellence, such as excellent, relatively excellent, sub-excellent, etc., or a specific level, such as the first level, the second level, the third level, etc.

[0046] For example, taking the indicator as a score, the indicator allocation strategy can generate up to 5 main demand contents. When the number of main demand contents is equal to 1, the indicator is set to 3 points; when the number of main demand contents is 2, each main demand content indicator is set to 2 points; when the number of main demand contents is greater than 2, each main demand content indicator is set to 1 point. For secondary demand contents, at most two are generated. According to the number of main demand contents and the total score, the score setting of the secondary demand content is adjusted. The minimum score of the secondary demand content is 1 point and the maximum score is 2 points.

[0047] It should be noted that a primary and secondary content extractor can be defined, and the specific expression form of the primary and secondary content extractor can be diverse, which can be a trained large language model, a key content extraction model, or an extraction program configured based on a programming language or other content extraction algorithms. The content extractor can extract the primary and secondary content contained therein based on a given evaluation case and the corresponding reference results.

[0048] As an example, natural language processing (NLP) technology can be used to extract the main and secondary required content contained in a given evaluation case and the corresponding reference results. Specifically, the evaluation case and the corresponding reference results can be preprocessed, such as text cleaning: removing irrelevant characters, punctuation marks, stop words, etc., word segmentation: breaking sentences into words or phrases, part-of-speech tagging: marking each word with its part of speech, such as noun, verb, etc., named entity recognition (NER): identifying entity names in input data, such as names of people, places, organizations, etc. Then, TF-IDF (TermFrequency-Inverse Document Frequency), TextRank or other algorithms can be used to extract keywords, dependency parsing or component parsing can be used to understand sentence structure, determine subject-verb-object relationships, identify content distribution through methods such as LDA (Latent DirichletAllocation), and perform demand content modeling. Afterwards, based on domain knowledge and the characteristics of evaluation cases, you can define a set of strategies to distinguish which information is necessary (primary content) and which is supplementary (secondary content), use regular expressions or custom functions to match these strategies, and mark the corresponding primary and secondary content; alternatively, you can prepare a training set that contains sample content distribution and labeled primary and secondary content samples, construct feature vectors, such as word frequency, position information, sentence length, keyword density, etc., select a suitable classification algorithm (such as SVM, random forest, neural network, etc.) for training, obtain a key content extraction model, use the trained key content extraction model to identify the content distribution of the evaluation case and the corresponding reference results, and output the corresponding primary and secondary content.

[0049] Of course, you can also directly prepare a training set, including multiple sample data, sample results, and corresponding primary and secondary required contents, train to obtain a key content extraction model, input the evaluation case into the trained key content extraction model, and output the corresponding primary and secondary required contents.

[0050] In the embodiments of the present specification, targeted analysis can be performed on the evaluation use cases and reference results to determine their reasoning demand characteristics, extract corresponding primary and secondary demand contents, allocate corresponding indicators to the extracted primary and secondary demand contents according to the set indicator allocation strategy, and generate demand assessment constraints corresponding to the evaluation use cases. Thus, under the constraints of the primary and secondary demand contents of the evaluation use cases themselves, detailed positive factors and / or negative factors for the reasoning requirements of the evaluation use cases are defined step by step, and each evaluation use case can adapt to its own reasoning demand characteristics, and adaptively generate a set of fine-grained demand assessment constraints to obtain more accurate and flexible automated reasoning quality evaluation results.

[0051] In an optional implementation of this embodiment, corresponding indicators are allocated to the extracted primary demand content and secondary demand content according to a set indicator allocation strategy, and demand evaluation constraints corresponding to the evaluation case are generated, including: Based on the set indicator allocation strategy, corresponding indicators are allocated to the extracted primary demand contents and secondary demand contents in turn; Generate a first positive factor and / or a first negative factor corresponding to the evaluation case according to the indicators corresponding to each primary demand content and secondary demand content; According to the first positive factor and / or the first negative factor, a requirement assessment constraint corresponding to the assessment case is generated.

[0052] In actual implementation, the indicators corresponding to each main demand content can be determined according to the set indicator allocation strategy, and then the indicators corresponding to each secondary demand content can be determined based on the total indicator, the indicator occupied by the main demand content and the number of secondary demand content; then, based on whether the reasoning result contains the main demand content, the secondary demand content and the allocated indicator, the corresponding first positive factor and / or first negative factor are generated. Among them, if the reasoning result contains the main demand content and the secondary demand content, the first positive factor of the corresponding indicator can be generated; if it does not contain the main demand content and the secondary demand content, the first negative factor of the corresponding indicator can be generated. By combining the first positive factors and / or first negative factors corresponding to each main demand content and the secondary demand content, the demand assessment constraints for the adaptability of the evaluation use case can be obtained.

[0053] Using the above example, for the evaluation case "Why is Li Bai so much younger than Du Fu", the main requirement of this evaluation case is "Li Bai is older than Du Fu or Du Fu is younger than Li Bai", and the secondary requirements are "Li Bai was born in 701, or the year of Li Bai's birth and death", "Du Fu was born in 712, or the year of Du Fu's birth and death", and "other additional relevant information". According to the above-set indicator allocation strategy, there is 1 main requirement point, and the corresponding indicator of "Li Bai is older than Du Fu or Du Fu is younger than Li Bai" is 3 points. Assuming the total indicator is 6 points, since there are 3 secondary requirements and the main requirements have been allocated 3 points, 1 point is allocated to each of the 3 secondary requirements, that is, "Li Bai was born in 701, or the year of Li Bai's birth and death" corresponds to 1 point, "Du Fu was born in 712, or the year of Du Fu's birth and death" corresponds to 1 point, and "other additional relevant information" corresponds to 1 point.

[0054] At this point, the adaptive demand assessment constraint for the assessment case "Why is Li Bai so much younger than Du Fu" can be generated: "Answer that Li Bai is older than Du Fu or Du Fu is younger than Li Bai - get 3 points; answer that Li Bai was born in 701 or Li Bai's birth and death years - get 1 point; answer that Du Fu was born in 712 or Du Fu's birth and death years - get 1 point; answer other additional relevant information - get 1 point" (first positive factor). Or, "Failed to answer that Li Bai is older than Du Fu or Du Fu is younger than Li Bai - deduct 3 points; failed to answer that Li Bai was born in 701 or Li Bai's birth and death years - deduct 1 point; failed to answer that Du Fu was born in 712 or Du Fu's birth and death years - deduct 1 point; failed to answer other additional relevant information - deduct 1 point" (first negative factor). Or, "Answering that Li Bai is older than Du Fu or Du Fu is younger than Li Bai - get 3 points, not answering that Li Bai is older than Du Fu or Du Fu is younger than Li Bai - deduct 3 points; answering that Li Bai was born in 701 or the year of his birth and death - get 1 point, not answering that Li Bai was born in 701 or the year of his birth and death - deduct 1 point; answering that Du Fu was born in 712 or the year of his birth and death - get 1 point, not answering that Du Fu was born in 712 or the year of his birth and death - deduct 1 point; answering other additional relevant information - get 1 point, not answering other additional relevant information - deduct 1 point" (the first positive factor and the first negative factor).

[0055] In the embodiments of the present specification, corresponding indicators can be allocated to the extracted main demand contents and secondary demand contents in turn according to the set indicator allocation strategy, and the first positive factor and / or the first negative factor corresponding to the generated evaluation use case can be flexibly selected to generate the corresponding demand evaluation constraints. The generation method of the demand evaluation constraints is more flexible and can be adapted to a variety of application scenarios with different needs, with higher flexibility and adaptability.

[0056] In an optional implementation of this embodiment, the reasoning requirement also includes factual content; and the method further includes: Determine whether there is factual content in the reference results corresponding to the evaluation case; If there is factual content, the second positive factor or the second negative factor corresponding to the evaluation case is configured according to the set indicator item; Accordingly, based on the evaluation case and the reference results, the requirement evaluation constraints corresponding to the reasoning requirements of the evaluation case are generated, including: According to the first positive factor, the first negative factor, the second positive factor and / or the second negative factor, a requirement assessment constraint corresponding to the assessment case is generated.

[0057] Among them, factual content refers to the specific, objective and verifiable information contained in the reasoning analysis of evaluation cases. Factual content exists objectively and does not rely on personal opinions or interpretations. This type of content is usually based on known facts, data, statistics, historical events, scientific principles or other verified knowledge. Factual content provides a solid foundation for the reasoning analysis of evaluation cases and ensures the accuracy and reliability of information.

[0058] The indicator items are set to reward content that is consistent with objective facts or punish content that is inconsistent with objective facts. Suppose the indicator item is set to correspond to 5 points for factual content. At this time, the corresponding second positive factor "the answer does not include information that is inconsistent with the factual content - get 5 points" can be generated, or the corresponding second negative factor "the answer includes information that is inconsistent with the factual content - deduct 5 points" can be generated.

[0059] It should be noted that in addition to generating the demand assessment constraints corresponding to the evaluation use case based on the main demand content and the secondary demand content as mentioned above, it is also possible to further determine whether there is factual content in the reference results corresponding to the evaluation use case, so as to set the corresponding second positive factor or second negative factor, thereby combining the first positive factor, the first negative factor, the second positive factor and / or the second negative factor to generate the demand assessment constraints corresponding to the evaluation use case.

[0060] In actual implementation, a factual content judger can be defined. The specific expression form of the factual content judger can be diverse, as long as it can judge whether there is factual content in the reference result based on the evaluation case and the corresponding reference result.

[0061] As an example, natural language processing (NLP) technology can be used to determine whether the reference results include factual content. Specifically, the reference results can be preprocessed first, such as text cleaning: removing irrelevant characters, punctuation marks, stop words, etc., word segmentation: breaking sentences into words or phrases, part-of-speech tagging: marking each word with its part of speech, such as noun, verb, etc., named entity recognition (NER): identifying entity names in input data, such as names of people, places, organizations, etc. Then, TF-IDF (Term Frequency-Inverse Document Frequency), TextRank or other algorithms can be used to extract keywords, dependency parsing or component parsing can be used to understand sentence structure, determine subject-verb-object relationships, and identify the distribution of various types of content through methods such as LDA (Latent Dirichlet Allocation). After that, prepare a labeled dataset containing question-answer pairs, and each answer is labeled as whether it contains factual content. The labeling can be binary classification (containing / not containing factual content) or multi-classification (different types of factual content). Construct feature vectors, such as keyword density, number of entities, frequency of specific parts of speech, sentence length, etc., and use word embeddings (such as Word2Vec, GloVe) to represent text features. Then select a suitable classification model, such as logistic regression, support vector machine (SVM), random forest, gradient boosting tree (such as XGBoost) or deep learning model (such as convolutional neural network CNN, recurrent neural network RNN), etc., divide the labeled dataset into a training set and a validation set, train and select the classification model to be tested on the training set, and evaluate and tune it on the validation set. After obtaining the trained classification model, use the trained classification model to identify the distribution of each type of content in the reference results corresponding to the evaluation case obtained based on the above steps, and output whether it includes factual content.

[0062] Of course, you can also directly prepare a training set, including multiple sample data, sample results, and whether the annotation includes factual content, train to obtain a classification model, input the reference results corresponding to the evaluation case into the trained classification model, and output whether it includes factual content.

[0063] Continuing with the above example, the reference result corresponding to the evaluation case includes the factual content "Li Bai is older than Du Fu". Assuming that the indicator item is set to correspond to 5 points for factual content, the corresponding second positive factor "There is no relevant description in the answer that Li Bai is younger than Du Fu - get 5 points" can be generated, or the corresponding second negative factor "There is a relevant description in the answer that Li Bai is younger than Du Fu - deduct 5 points" can be generated.

[0064] As an example, the first positive factor and the second negative factor can be combined to generate the demand assessment constraint corresponding to the assessment use case: "Answer that Li Bai is older than Du Fu or Du Fu is younger than Li Bai - get 3 points; answer that Li Bai was born in 701 or the year of Li Bai's birth and death - get 1 point; answer that Du Fu was born in 712 or the year of Du Fu's birth and death - get 1 point; answer other additional relevant information - get 1 point; if there is a description that Li Bai is younger than Du Fu - deduct 5 points".

[0065] In the embodiments of the present specification, in addition to defining detailed positive factors and negative factors step by step under the constraints of the primary and secondary needs of the evaluation case itself, detailed positive factors and / or negative factors can also be further defined in combination with factual content. Through the positive factors and / or negative factors corresponding to the factual content, content that is consistent with objective facts is rewarded, or content that is inconsistent with objective facts is punished, so as to obtain the demand evaluation constraints corresponding to the evaluation case, adaptively generate a set of fine-grained demand evaluation constraints, and flexibly select the positive factors and / or negative factors corresponding to the evaluation case. The generation method of demand evaluation constraints is more flexible and can be adapted to a variety of application scenarios with different needs, with higher flexibility and adaptability.

[0066] In an optional implementation of this embodiment, after generating the requirement evaluation constraints corresponding to the reasoning requirements of the evaluation case according to the evaluation case and the reference result, the following is further included: Find the background content corresponding to the evaluation case; Add the background content to the requirement evaluation constraint to obtain the updated evaluation constraint.

[0067] It should be noted that different evaluation cases may have different background content, which can help the evaluation model better understand the evaluation case and the corresponding results to be evaluated. Therefore, in actual implementation, after generating the demand evaluation constraints, you can also search for the background content corresponding to the evaluation case through local databases, cloud servers, third-party platforms, search engines, etc., add the background content to the demand evaluation constraints, and obtain updated evaluation constraints. Subsequently, the evaluation model can be guided to analyze the evaluation case and the corresponding results to be evaluated based on the updated evaluation constraints, helping to guide the evaluation model to better understand the background knowledge and better generalize to the evaluation of various reasoning results, thereby improving the accuracy of the evaluation results.

[0068] Continuing with the above example, you can use a search engine to search for relevant background content for "Why is Li Bai so much younger than Du Fu?": "Both Li Bai and Du Fu lived in the Tang Dynasty, a period in history when culture was prosperous and economy was developed. The open policy and cultural exchanges of the Tang Dynasty provided a good environment for the development of literature and art. The Anshi Rebellion had a profound impact on the Tang Dynasty. The social unrest of this period was also reflected in the works of Li Bai and Du Fu. Li Bai and Du Fu met in Luoyang in the third year of Tianbao (744) and forged a deep friendship. Although their poetry styles were different, they respected each other." Add this background content to the above demand evaluation constraint to obtain the updated evaluation constraint "Both Li Bai and Du Fu lived in the Tang Dynasty, a period in history when culture was prosperous and economy was developed. During the period of economic development, the Tang Dynasty's open policy and cultural exchanges provided a good environment for the development of literature and art. The Anshi Rebellion had a profound impact on the Tang Dynasty. The social unrest during this period was also reflected in the works of Li Bai and Du Fu. Li Bai and Du Fu met in Luoyang in the third year of Tianbao (744) and formed a deep friendship. Although their poetry styles were different, they respected each other. Answer that Li Bai is older than Du Fu or Du Fu is younger than Li Bai - get 3 points; answer that Li Bai was born in 701, or the year of Li Bai's birth and death - get 1 point; answer that Du Fu was born in 712, or the year of Du Fu's birth and death - get 1 point; answer other additional relevant information - get 1 point; if there is a description that Li Bai is younger than Du Fu - deduct 5 points.

[0069] In an optional implementation of this embodiment, after generating the requirement evaluation constraints corresponding to the reasoning requirements of the evaluation case according to the evaluation case and the reference result, the following is further included: Generate the control results corresponding to the evaluation case according to the demand evaluation constraints; Based on the reference results and control results, determine the target evaluation constraints of the evaluation case.

[0070] The control result refers to the reasoning result generated for the evaluation case according to the positive factors and / or negative factors corresponding to the demand evaluation constraints, so as to measure the advantages and disadvantages of the generated demand evaluation constraints through the reference results and the control results.

[0071] It should be noted that after generating the requirement evaluation constraints corresponding to the reasoning requirements of the evaluation case based on the evaluation case and the reference results, the evaluation model can be directly guided to evaluate the results to be evaluated corresponding to the evaluation case based on the requirement evaluation constraints; or, the control results corresponding to the evaluation case can be generated based on the requirement evaluation constraints, and the final target evaluation constraints of the evaluation case can be determined based on the reference results and the control results, and then the evaluation model can be guided to evaluate the results to be evaluated corresponding to the evaluation case based on the target evaluation constraints.

[0072] In the embodiments of the present specification, the control result is a control group generated according to the detailed evaluation rules of the demand evaluation constraints adaptively generated for the evaluation case. Comparing the reference result with the control result can reflect the accuracy of the demand evaluation constraints adaptively generated for the evaluation case, thereby determining the final target evaluation constraints of the evaluation case, improving the accuracy of the demand evaluation constraints adaptively generated for the evaluation case, and further ensuring the accuracy of the evaluation results of the evaluation model.

[0073] In an optional implementation of this embodiment, determining the target evaluation constraint of the evaluation case according to the reference result and the comparison result includes: Determining the similarity between the reference results and the control results; If the similarity is greater than the similarity threshold, the requirement evaluation constraint is determined as the target evaluation constraint; If the similarity is less than or equal to the similarity threshold, return to execute the operation steps of generating the requirement evaluation constraints corresponding to the reasoning requirements of the evaluation case based on the evaluation case and the reference results, regenerate the requirement evaluation constraints corresponding to the evaluation case, until the iteration stop condition is met, and determine the requirement evaluation constraints of the current iteration round as the target evaluation constraints.

[0074] In actual implementation, a similarity judger can be defined to judge the similarity between two input data. The similarity judger can be diverse. For example, the similarity judger can output a value between 0 and 1. The closer the value is to 1, the more similar the two input data are.

[0075] In one possible implementation, the feature vectors of the reference result and the control result may be extracted, and then the cosine value of the angle between the two feature vectors may be calculated, and a value between [-1, 1] may be output, where 1 indicates that they are exactly the same, and -1 indicates that they are completely opposite.

[0076] In another possible implementation, a training set may be prepared, including multiple inference result pairs, wherein one inference result pair includes two inference results for the same sample data, and each inference result pair is annotated with a corresponding similarity label, and a similarity model is trained based on the training set. The reference result and the control result are input into the trained similarity model, and the corresponding similarity is output. If the similarity label is a specific similarity value, then the trained similarity model can output the similarity value corresponding to the reference result and the control result, such as a value between 0 and 1. The closer the value is to 1, the more similar the two input data are; if the similarity label is a similarity classification, such as dissimilar, relatively similar, similar, etc., then the trained similarity model can output the similarity type corresponding to the reference result and the control result.

[0077] It should be noted that if the similarity is greater than the similarity threshold, it means that the currently generated demand evaluation constraint is relatively accurate, and the currently generated demand evaluation constraint is determined as the final target evaluation constraint. If the similarity is less than or equal to the similarity threshold, it means that the error of the currently generated demand evaluation constraint is large and the reasoning quality cannot be accurately evaluated. At this time, you can return to continue to execute the operation steps of generating the demand evaluation constraints corresponding to the reasoning requirements of the evaluation case based on the evaluation case and the reference results, and regenerate the demand evaluation constraints corresponding to the evaluation case until the iteration stop condition is met, and the demand evaluation constraints of the current iteration round are determined as the target evaluation constraints. Among them, the iteration stop condition refers to the condition for stopping the iterative optimization of the demand evaluation constraint, such as the similarity is less than or equal to the similarity threshold, or the similarity is greater than the similarity threshold but the number of iterations reaches the number threshold, assuming that it can be iterated 10 times.

[0078] In actual implementation, if the demand assessment constraints are updated based on the background content, the above-mentioned iterative optimization process can be performed on the updated assessment constraints. The specific implementation is similar to the above-mentioned iterative optimization process for the demand assessment constraints, and the implementation of this specification will not be repeated here.

[0079] In the embodiments of the present specification, after the initial requirement evaluation constraints of the evaluation case are determined, a self-loop verification can be performed to iteratively optimize the requirement evaluation constraints until a relatively accurate target evaluation constraint is obtained, thereby greatly improving the accuracy of the fine-grained evaluation constraints generated for the evaluation case, thereby ensuring the accuracy of the evaluation results of the evaluation model.

[0080] For example, Figure 2 is a schematic diagram of a generation process of a demand assessment constraint corresponding to an assessment case provided by an embodiment of this specification, such as Figure 2 As shown, get the evaluation dataset .

[0081] Assumptions For "Why is Li Bai so much younger than Du Fu?" - "Li Bai (701-762) and Du Fu (712-770) are famous poets in China's Tang Dynasty, who are respectively known as the "Poet Immortal" and the "Poet Sage". Although Li Bai is 11 years older than Du Fu, they both have a high status in literary achievement and influence. The poetry styles and themes of Li Bai and Du Fu each have their own characteristics. Li Bai's poetry is known for its boldness, unrestrainedness, freshness, and loftiness, and his poetry has a strong personality and romanticism. Du Fu's poetry is known for its rigor, melancholy, realism, and concern for people's livelihood, and his poetry reflects social reality and people's suffering. Although Li Bai is older than Du Fu, their poetry achievements and influence are world-famous. In literary history, they are known as "Li Du", representing the pinnacle of Tang poetry." Among them, for - Any value in .

[0082] against ,Will Input to the main demand and secondary demand content extractor to obtain the corresponding main demand content and secondary demand content; and Input to the factual content judgement device to obtain Whether it includes factual content.

[0083] Generate the content based on the main content, secondary content and whether it includes factual content. Corresponding requirement assessment constraints. Find assessment use cases Corresponding background content, add the background content to the demand assessment constraints to obtain Corresponding update evaluation constraints .

[0084] Evaluate constraints based on updates Generate evaluation cases The corresponding control results , the comparison results And reference results Input to the similarity judgement to obtain the similarity. Determine whether the similarity is greater than the similarity threshold. If so, update the current evaluation constraint As an evaluation case The final target evaluation constraint. If not, determine whether the number of iterations has been reached. If not, then Return to the above steps and continue to generate The corresponding next update evaluation constraint ; If so, update the evaluation constraint of the current iteration round Identify as an evaluation case The final target evaluation constraint. Among them, The number of iterations to stop.

[0085] like Figure 2 As shown, the target evaluation constraint can be "4 scoring points in total, 1 deduction point, and the total score is 0-6 points. Scoring points: answering that Li Bai is older than Du Fu or Du Fu is younger than Li Bai [3 points]; answering that Li Bai was born in 701 or the year of Li Bai's birth and death [1 point]; answering that Du Fu was born in 712 or the year of Du Fu's birth and death [1 point]; answering other additional relevant information [1 point]; deduction points: if there is a description that Li Bai is younger than Du Fu [deduction 5 points]".

[0086] Step 106: Generate an evaluation result for the evaluation case using the model to be evaluated.

[0087] Specifically, the model to be evaluated refers to a large model whose reasoning ability is waiting to be evaluated, and the result to be evaluated is the reasoning result output by the model to be evaluated after reasoning and analyzing the evaluation case.

[0088] It should be noted that the model to be evaluated can be a large language model in various application scenarios. For example, in the application scenario of news writing, the model to be evaluated is a large language model that can generate coherent and accurate news releases; in the application scenario of customer service chatbots, the model to be evaluated is a large language model that provides accurate and timely customer service; in the application scenario of educational tutoring, the model to be evaluated is a large language model that provides personalized learning suggestions and feedback.

[0089] In actual implementation, the evaluation case can be input into the model to be evaluated to obtain the evaluation result corresponding to the evaluation case. Subsequently, the evaluation result output by the model to be evaluated can be evaluated to obtain the reasoning quality evaluation result to reflect the processing and analysis capability of the model to be evaluated, thereby realizing the evaluation of the model to be evaluated.

[0090] Step 108: Based on the requirement evaluation constraints, the evaluation model is used to evaluate the evaluation case and the result to be evaluated, and the reasoning quality evaluation result corresponding to the result to be evaluated of the evaluation case is obtained.

[0091] The evaluation model may be a large language model used to evaluate the quality of the reasoning results of the evaluation case.

[0092] In an optional implementation, after the demand evaluation constraints corresponding to the evaluation case are generated based on the primary demand content, secondary demand content, and whether there is factual content corresponding to the evaluation case and the reference result, the evaluation case, the result to be evaluated, and the demand evaluation constraints can be directly input into the evaluation model, and the evaluation model can be guided to perform evaluation processing based on the demand evaluation constraints to obtain the reasoning quality evaluation results corresponding to the result to be evaluated of the evaluation case. In another optional implementation, if the demand evaluation constraints are updated based on the background content corresponding to the evaluation case, and an updated evaluation constraint is obtained, the evaluation case, the result to be evaluated, and the updated evaluation constraints can be input into the evaluation model, and the evaluation model can be guided to perform evaluation processing based on the updated evaluation constraints to obtain the reasoning quality evaluation results corresponding to the result to be evaluated of the evaluation case.

[0093] In another optional implementation, if the demand evaluation constraint or the updated evaluation constraint is iteratively optimized to obtain the target evaluation constraint, the evaluation case, the result to be evaluated, and the target evaluation constraint can be input into the evaluation model, and the evaluation model can be guided to perform evaluation processing based on the target evaluation constraint to obtain the reasoning quality evaluation result corresponding to the result to be evaluated of the evaluation case, that is, the evaluation case and the corresponding result to be evaluated are evaluated and processed based on the demand evaluation constraint by the evaluation model, and the reasoning quality evaluation result corresponding to the result to be evaluated is obtained, including: the evaluation case, the result to be evaluated, and the target evaluation constraint are input into the evaluation model for evaluation processing to obtain the reasoning quality evaluation result corresponding to the result to be evaluated of the evaluation case. The final target evaluation constraint of the evaluation case is determined by iterative optimization, and then the evaluation model is guided to perform evaluation based on the target evaluation constraint, which greatly improves the accuracy of the reasoning quality evaluation result.

[0094] Step 110: Determine the evaluation result of the model to be evaluated according to the inference quality evaluation result corresponding to the evaluation result to be evaluated of the evaluation case.

[0095] It should be noted that the evaluation case is any one in the evaluation data set. After evaluating the results to be evaluated of the evaluation case and obtaining the corresponding reasoning quality evaluation results, the model to be evaluated can be directly evaluated based on the reasoning quality evaluation results corresponding to the results to be evaluated of the evaluation case to obtain the evaluation results of the model to be evaluated; or, the model to be evaluated can be comprehensively evaluated based on the reasoning quality evaluation results corresponding to the results to be evaluated of each evaluation case in the evaluation data set to obtain the evaluation result of the model to be evaluated.

[0096] The evaluation result of the model to be evaluated may be a specific quality score, such as 95 points, 90 points, 88 points, etc., or a quality level, such as excellent, relatively excellent, good, poor, etc.

[0097] In actual implementation, the result to be evaluated is the inference result of the model to be evaluated performing reasoning analysis on the evaluation case, and the evaluation case is any data in the evaluation data set. The evaluation model evaluates the evaluation case and the corresponding result to be evaluated based on the requirement evaluation constraints, and after obtaining the inference quality evaluation result corresponding to the result to be evaluated of the evaluation case, the model evaluation result of the model to be evaluated can also be determined based on the inference quality evaluation result of the result to be evaluated corresponding to each evaluation case in the evaluation data set.

[0098] It should be noted that the model evaluation method provided in the embodiments of this specification can be divided into two parts. One part is the process of generating adaptive demand evaluation constraints for evaluation use cases, and the other part is the process of evaluating the results to be evaluated corresponding to the evaluation use cases based on the adaptive demand evaluation constraints corresponding to the evaluation use cases.

[0099] In actual implementation, for any evaluation case in the evaluation data set, the corresponding target evaluation constraint can be generated based on the above-mentioned generation process of the adaptive demand evaluation constraint, and then the model reasoning quality evaluation is performed on the result to be evaluated corresponding to the evaluation case (that is, the reasoning result output by the model to be evaluated) based on the target evaluation constraint to obtain the reasoning quality evaluation result of the evaluation case. Based on the reasoning quality evaluation results of the result to be evaluated corresponding to each evaluation case in the evaluation data set, the model evaluation result of the model to be evaluated can be determined.

[0100] Specifically, the model score of the model to be evaluated can be obtained according to the reasoning quality evaluation results of each evaluation case in the evaluation data set; the model capability level of the model to be evaluated can be determined according to the reasoning quality evaluation results of each evaluation case in the evaluation data set; the model score and / or model capability level can be used as the model evaluation result of the model to be evaluated. Among them, the model score of the model to be evaluated can be obtained by summing the reasoning quality evaluation results of each evaluation case; or, the geometric mean of the reasoning quality evaluation results of each evaluation case can be calculated as the model score of the model to be evaluated. Of course, other data calculation methods can also be used to process the reasoning quality evaluation results of each evaluation case to obtain the model score of the model to be evaluated, such as the harmonic mean, which is not limited in the embodiments of this specification.

[0101] In addition, the model capability level of the model to be evaluated can also be determined based on the reasoning quality evaluation results of each evaluation case, where the model capability level is used to indicate the quality of the processing capability of the model to be evaluated. For example, the model capability level can be excellent, medium, poor, or level one, level two, level three, etc. The higher the level, the better the processing capability.

[0102] Specifically, different scoring ranges and the capability levels corresponding to the scoring ranges can be pre-configured, and the inference quality evaluation results of each evaluation case can be summed, the geometric mean or other data calculation methods can be used to obtain the comprehensive score of the inference quality evaluation results of each evaluation case. The calculation method of the comprehensive score can be the same as or different from the calculation method of the model score of the model to be evaluated. Then, the target scoring range corresponding to the comprehensive score is determined, and the target capability level corresponding to the target scoring range is used as the model capability level of the model to be evaluated.

[0103] It should be noted that the model score of the model to be evaluated can be used as the model evaluation result, or the model capability level of the model to be evaluated can be used as the model evaluation result, or the model score and the model capability level can be used together as the model evaluation result. The model score and / or model capability level of the model to be evaluated can be determined as the model evaluation result based on the inference quality evaluation results of the results to be evaluated corresponding to each evaluation case in the evaluation data set, which ensures the accuracy of the model evaluation results, and the evaluation method is richer and more flexible, which can adapt to a variety of different application scenarios and improve the model evaluation effect.

[0104] The embodiments of this specification provide a model evaluation method, which can adaptively generate demand evaluation constraints corresponding to the evaluation case based on the primary demand content, secondary demand content, and whether factual content is included in the evaluation case and the reference results, and can generate corresponding control results based on the demand evaluation constraints, perform iterative optimization of the demand evaluation constraints, and obtain the final target evaluation constraints of the evaluation case. The target evaluation constraints contain positive factors and / or negative factors, comprehensively consider the reasoning situation, and guide the evaluation model to evaluate the evaluation case and the corresponding results to be evaluated based on the target evaluation constraints, so as to obtain the reasoning quality evaluation results corresponding to the results to be evaluated of the evaluation case, and the model capability of the model to be evaluated can be determined based on the reasoning quality evaluation results corresponding to the results to be evaluated of each evaluation case in the evaluation data set.

[0105] In this way, each evaluation case can adapt to its own reasoning demand characteristics and adaptively generate a set of fine-grained demand evaluation constraints, which can be better generalized to the reasoning quality evaluation of evaluation cases of various types and different reasoning requirements, and can be adapted to a variety of application scenarios, so as to obtain more accurate and flexible automated reasoning quality evaluation results. Moreover, the entire evaluation process simulates the process in the manual evaluation process, which can help the evaluation model better approach the manual evaluation process and improve the consistency rate with the manual evaluation. The above-mentioned detailed demand evaluation constraints can save the evaluation model from the task planning and decomposition process, so that the evaluation model only needs to follow the demand evaluation constraints for evaluation, which greatly reduces the reliance on the inherent knowledge and capabilities of the evaluation model itself, and can bring more accurate automated evaluation results.

[0106] The following combination Figure 3 , taking the application of the model evaluation method provided in this specification in the question-answering model scenario as an example, the model evaluation method is further explained. Figure 3 A processing flow chart of a model evaluation method provided by an embodiment of the present specification is shown, which specifically includes the following steps.

[0107] Step 302: Obtain an evaluation question set based on the question-answering platform, where the evaluation question set includes at least one question and a corresponding answer.

[0108] Step 304: Obtain any question from the evaluation question set as a question to be evaluated, and obtain the corresponding reference answer.

[0109] Step 306: Input the question to be evaluated into the model to be evaluated to generate the answer to be evaluated.

[0110] Step 308: Determine the corresponding primary content, secondary content and whether factual content is included according to the question to be evaluated and the reference answer, and generate the requirement evaluation constraint corresponding to the question to be evaluated according to the primary content, secondary content and whether factual content is included.

[0111] Step 310: Find the background content corresponding to the topic to be evaluated, add the background content to the requirement evaluation constraint, and obtain the updated evaluation constraint.

[0112] Step 312: Generate a reference answer corresponding to the question to be evaluated based on the required evaluation constraints, and determine the similarity between the reference answer and the reference answer.

[0113] Step 314: If the similarity is greater than the similarity threshold, the requirement evaluation constraint is determined as the target evaluation constraint.

[0114] Step 316: If the similarity is less than or equal to the similarity threshold, return to execute the above step 308 until the iteration stop condition is met, and determine the demand evaluation constraint of the current iteration round as the target evaluation constraint.

[0115] Step 318: The question to be evaluated, the answer to be evaluated, and the target evaluation constraint are input into the evaluation model for evaluation processing to obtain the answer quality evaluation result corresponding to the evaluation result of the question to be evaluated.

[0116] Return to execute the above step 304 until the last question in the evaluation question set, and obtain the answer quality evaluation results of the answers to be evaluated corresponding to each question.

[0117] Step 320: Determine the model evaluation result of the model to be evaluated according to the answer quality evaluation results of the answers to be evaluated corresponding to each question in the evaluation question set.

[0118] The embodiments of this specification provide a model evaluation method. For a question-and-answer platform, each question to be evaluated can adapt to its own answering requirements and adaptively generate a set of fine-grained requirement evaluation constraints, which can be better generalized to the answer quality evaluation of questions of various types and with different reasoning requirements, and can be adapted to a variety of application scenarios, so as to obtain more accurate and flexible evaluation results of the quality of automated question answers; moreover, the entire evaluation process simulates the process in the manual evaluation process, which can help the evaluation model to better approach the manual evaluation process and improve the consistency rate with the manual evaluation. The above-mentioned detailed requirement evaluation constraints can save the evaluation model from the task planning and decomposition process, so that the evaluation model only needs to follow the requirement evaluation constraints for evaluation, which greatly reduces the reliance on the inherent knowledge and capabilities of the evaluation model itself, and can bring more accurate automated evaluation results.

[0119] Corresponding to the above method embodiment, this specification also provides a model evaluation device embodiment, Figure 4 FIG. 1 is a schematic diagram showing the structure of a model evaluation device provided by an embodiment of the present specification. Figure 4 As shown, the device comprises: An acquisition module 402 is configured to acquire evaluation cases and corresponding reference results; The first generating module 404 is configured to generate a requirement evaluation constraint corresponding to the reasoning requirement of the evaluation case according to the evaluation case and the reference result, wherein the requirement evaluation constraint is used to constrain the positive factor and / or negative factor of the evaluation case; The second generating module 406 is configured to generate a result to be evaluated for the evaluation case through the model to be evaluated; The first evaluation module 408 is configured to evaluate the evaluation case and the result to be evaluated based on the demand evaluation constraint through the evaluation model, and obtain the reasoning quality evaluation result corresponding to the result to be evaluated of the evaluation case; The second evaluation model 410 is configured to determine the evaluation result of the model to be evaluated according to the reasoning quality evaluation result corresponding to the result to be evaluated of the evaluation case.

[0120] Optionally, the inference requirement includes primary requirement content and secondary requirement content; the first generating module 404 is further configured to: Extract the corresponding primary and secondary requirements based on the evaluation cases and reference results; According to the set indicator allocation strategy, corresponding indicators are allocated to the extracted primary and secondary demand contents, and the demand assessment constraints corresponding to the evaluation case are generated.

[0121] Optionally, the first generating module 404 is further configured to: Based on the set indicator allocation strategy, corresponding indicators are allocated to the extracted primary demand contents and secondary demand contents in turn; Generate a first positive factor and / or a first negative factor corresponding to the evaluation case according to the indicators corresponding to each primary demand content and secondary demand content; According to the first positive factor and / or the first negative factor, a requirement assessment constraint corresponding to the assessment case is generated.

[0122] Optionally, the reasoning requirement further includes factual content; the device further includes a first determining module configured to: Determine whether there is factual content in the reference results corresponding to the evaluation case; If there is factual content, the second positive factor or the second negative factor corresponding to the evaluation case is configured according to the set indicator item; Accordingly, the first generating module 404 is further configured to: According to the first positive factor, the first negative factor, the second positive factor and / or the second negative factor, a requirement assessment constraint corresponding to the assessment case is generated.

[0123] Optionally, the device further includes an updating module configured to: Find the background content corresponding to the evaluation case; Add the background content to the requirement evaluation constraint to obtain the updated evaluation constraint.

[0124] Optionally, the device further includes a second determining module, configured to: Generate the control results corresponding to the evaluation case according to the demand evaluation constraints; Determine the target evaluation constraints of the evaluation case based on the reference results and control results; Accordingly, the first evaluation module 408 is further configured to: The evaluation case, the result to be evaluated, and the target evaluation constraint are input into the evaluation model for evaluation processing to obtain the reasoning quality evaluation result corresponding to the result to be evaluated of the evaluation case.

[0125] Optionally, the second determining module is further configured to: Determining the similarity between the reference results and the control results; If the similarity is greater than the similarity threshold, the requirement evaluation constraint is determined as the target evaluation constraint; If the similarity is less than or equal to the similarity threshold, return to execute the operation steps of generating the requirement evaluation constraints corresponding to the reasoning requirements of the evaluation case based on the evaluation case and the reference results, regenerate the requirement evaluation constraints corresponding to the evaluation case, until the iteration stop condition is met, and determine the requirement evaluation constraints of the current iteration round as the target evaluation constraints.

[0126] The embodiment of this specification provides a model evaluation device, including an acquisition module, a first generation module, a second generation module, a first evaluation module and a second evaluation module. The modules interact with each other to realize the reasoning requirements based on the evaluation case and the reference results, and adaptively generate the demand evaluation constraints corresponding to the evaluation case, so as to guide the evaluation model to evaluate the evaluation case and the corresponding results to be evaluated according to the corresponding positive factors and / or negative factors based on the demand evaluation constraints, so as to guide the evaluation model to evaluate the reasoning requirements based on the evaluation case. In this way, each evaluation case can adapt to its own reasoning requirement characteristics, adaptively generate a set of fine-grained demand evaluation constraints, and specifically constrain the detailed positive factors and / or negative factors of the evaluation case, which can be better generalized to the reasoning quality evaluation of evaluation cases of various types and different reasoning requirements, and adapt to a variety of application scenarios. More accurate and flexible automated reasoning quality evaluation results can be obtained, and then more accurate and flexible model evaluation results can be obtained.

[0127] The above is a schematic scheme of a model evaluation device of this embodiment. It should be noted that the technical scheme of the model evaluation device and the technical scheme of the above-mentioned model evaluation method belong to the same concept, and the details not described in detail in the technical scheme of the model evaluation device can be referred to the description of the technical scheme of the above-mentioned model evaluation method.

[0128] Figure 5 The structure block diagram of a computing device 500 provided according to an embodiment of the present specification is shown. The components of the computing device 500 include but are not limited to a memory 510 and a processor 520. The processor 520 is connected to the memory 510 via a bus 530, and the database 550 is used to store data.

[0129] The computing device 500 also includes an access device 540 that enables the computing device 500 to communicate via one or more networks 560. Examples of these networks include a public switched telephone network (PSTN), a local area network (LAN), a wide area network (WAN), a personal area network (PAN), or a combination of communication networks such as the Internet. The access device 540 may include one or more of any type of network interface (e.g., a network interface card (NIC)) that is wired or wireless, such as an IEEE 802.11 wireless local area network (WLAN) wireless interface, a world-wide interoperability for microwave access (Wi-MAX) interface, an Ethernet interface, a universal serial bus (USB) interface, a cellular network interface, a Bluetooth interface, and a near field communication (NFC).

[0130] In one embodiment of the present specification, the above components of the computing device 500 and Figure 5 Other components not shown in the figure may also be connected to each other, for example, via a bus. It should be understood that Figure 5 The computing device structure block diagram shown is only for the purpose of illustration, and is not intended to limit the scope of this specification. Those skilled in the art can add or replace other components as needed.

[0131] The computing device 500 may be any type of stationary or mobile computing device, including a mobile computer or mobile computing device (e.g., a tablet computer, a personal digital assistant, a laptop computer, a notebook computer, a netbook, etc.), a mobile phone (e.g., a smart phone), a wearable computing device (e.g., a smart watch, smart glasses, etc.), or other types of mobile devices, or a stationary computing device such as a desktop computer or a personal computer (PC). The computing device 500 may also be a mobile or stationary server.

[0132] The processor 520 is used to execute the following computer executable instructions, which, when executed by the processor, implement the steps of the above-mentioned model evaluation method.

[0133] The above is a schematic scheme of a computing device of this embodiment. It should be noted that the technical scheme of the computing device and the technical scheme of the above-mentioned model evaluation method belong to the same concept, and the details not described in detail in the technical scheme of the computing device can be referred to the description of the technical scheme of the above-mentioned model evaluation method.

[0134] An embodiment of the present specification also provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the steps of the above-mentioned model evaluation method.

[0135] The above is a schematic scheme of a computer-readable storage medium of this embodiment. It should be noted that the technical scheme of the storage medium and the technical scheme of the above-mentioned model evaluation method belong to the same concept, and the details not described in detail in the technical scheme of the storage medium can be referred to the description of the technical scheme of the above-mentioned model evaluation method.

[0136] An embodiment of the present specification also provides a computer program, wherein when the computer program is executed in a computer, the computer is caused to execute the steps of the above-mentioned model evaluation method.

[0137] The above is a schematic scheme of a computer program of this embodiment. It should be noted that the technical scheme of the computer program and the technical scheme of the above-mentioned model evaluation method belong to the same concept, and the details not described in detail in the technical scheme of the computer program can be referred to the description of the technical scheme of the above-mentioned model evaluation method.

[0138] The above is a description of a specific embodiment of the specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in an order different from that in the embodiments and still achieve the desired results. In addition, the processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0139] Computer instructions include computer program codes, which may be in source code form, object code form, executable files or some intermediate form, etc. Computer readable media may include: any entity or device capable of carrying computer program codes, recording media, USB flash drives, mobile hard disks, magnetic disks, optical disks, computer memories, read-only memories (ROM), random access memories (RAM), electric carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the contents of computer readable media may be appropriately increased or decreased according to the requirements of patent practice. For example, in some regions, according to patent practice, computer readable media do not include electric carrier signals and telecommunication signals.

[0140] It should be noted that, for the above-mentioned method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations, but those skilled in the art should be aware that the embodiments of this specification are not limited by the order of the actions described, because according to the embodiments of this specification, some steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the embodiments of this specification.

[0141] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0142] The preferred embodiments of this specification disclosed above are only used to help explain this specification. The optional embodiments do not describe all the details in detail, nor do they limit the invention to only the specific implementation methods described. Obviously, many modifications and changes can be made according to the content of the embodiments of this specification. This specification selects and specifically describes these embodiments in order to better explain the principles and practical applications of the embodiments of this specification, so that technicians in the relevant technical field can well understand and use this specification. This specification is only limited by the claims and their full scope and equivalents.

Claims

1. A model evaluation method, characterized in that: include: Obtain evaluation cases and corresponding reference results; Generate, according to the evaluation case and the reference result, a requirement evaluation constraint corresponding to the reasoning requirement of the evaluation case, wherein the requirement evaluation constraint is used to constrain the positive factor and / or negative factor of the evaluation case; Generate an evaluation result for the evaluation case through the model to be evaluated; The evaluation model is used to evaluate the evaluation case and the result to be evaluated based on the requirement evaluation constraint to obtain a reasoning quality evaluation result corresponding to the result to be evaluated of the evaluation case; The evaluation result of the model to be evaluated is determined according to the reasoning quality evaluation result corresponding to the result to be evaluated of the evaluation case.

2. The method according to claim 1, characterized in that The reasoning requirement includes a primary requirement content and a secondary requirement content; the generating, according to the evaluation case and the reference result, a requirement evaluation constraint corresponding to the reasoning requirement of the evaluation case includes: Extracting corresponding primary and secondary required contents according to the evaluation case and the reference result; Corresponding indicators are allocated to the extracted primary demand content and secondary demand content according to the set indicator allocation strategy, and demand evaluation constraints corresponding to the evaluation case are generated.

3. The method according to claim 2, characterized in that The step of allocating corresponding indicators to the extracted primary demand content and secondary demand content according to the set indicator allocation strategy, and generating demand assessment constraints corresponding to the assessment case, includes: Based on the set indicator allocation strategy, corresponding indicators are allocated to the extracted primary demand contents and secondary demand contents in turn; Generate a first positive factor and / or a first negative factor corresponding to the evaluation case according to the indicators corresponding to the primary and secondary requirements; A requirement assessment constraint corresponding to the assessment use case is generated according to the first positive factor and / or the first negative factor.

4. The method according to claim 3, characterized in that The reasoning requirement also includes factual content; the method also includes: Determining whether there is factual content in the reference result corresponding to the evaluation case; If the factual content exists, configuring the second positive factor or the second negative factor corresponding to the evaluation case according to the set indicator item; Accordingly, generating a requirement evaluation constraint corresponding to the reasoning requirement of the evaluation case according to the evaluation case and the reference result includes: A requirement assessment constraint corresponding to the assessment case is generated according to the first positive factor, the first negative factor, the second positive factor and / or the second negative factor.

5. The method according to claim 1, characterized in that After generating the requirement evaluation constraint corresponding to the reasoning requirement of the evaluation case according to the evaluation case and the reference result, the method further includes: Find the background content corresponding to the evaluation case; The background content is added to the demand evaluation constraint to obtain an updated evaluation constraint.

6. The method according to any one of claims 1 to 5, characterized in that: After generating the requirement evaluation constraint corresponding to the reasoning requirement of the evaluation case according to the evaluation case and the reference result, the method further includes: Generating a comparison result corresponding to the evaluation case according to the demand evaluation constraint; Determining a target evaluation constraint of the evaluation case according to the reference result and the comparison result; Accordingly, the evaluation model performs evaluation processing on the evaluation case and the result to be evaluated based on the requirement evaluation constraint to obtain the reasoning quality evaluation result corresponding to the result to be evaluated of the evaluation case, including: The evaluation case, the result to be evaluated and the target evaluation constraint are input into the evaluation model for evaluation processing to obtain the reasoning quality evaluation result corresponding to the result to be evaluated of the evaluation case.

7. The method according to claim 6, characterized in that Determining the target evaluation constraint of the evaluation case according to the reference result and the comparison result includes: Determining a similarity between the reference result and the control result; If the similarity is greater than a similarity threshold, determining the demand evaluation constraint as the target evaluation constraint; If the similarity is less than or equal to the similarity threshold, return to execute the operation step of generating the requirement evaluation constraints corresponding to the reasoning requirements of the evaluation case based on the evaluation case and the reference result, regenerate the requirement evaluation constraints corresponding to the evaluation case, until the iteration stop condition is met, and determine the requirement evaluation constraints of the current iteration round as the target evaluation constraints.

8. A model evaluation device, characterized in that: include: An acquisition module, configured to acquire evaluation cases and corresponding reference results; A first generating module is configured to generate a requirement evaluation constraint corresponding to the reasoning requirement of the evaluation case according to the evaluation case and the reference result, wherein the requirement evaluation constraint is used to constrain the positive factor and / or negative factor of the evaluation case; A second generating module is configured to generate a result to be evaluated for the evaluation case through the model to be evaluated; A first evaluation module is configured to evaluate the evaluation case and the result to be evaluated based on the requirement evaluation constraint through an evaluation model, and obtain a reasoning quality evaluation result corresponding to the result to be evaluated of the evaluation case; The second evaluation module is configured to determine the evaluation result of the model to be evaluated according to the reasoning quality evaluation result corresponding to the result to be evaluated of the evaluation case.

9. A computing device, characterized in that include: Memory and processor; The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions. When the computer-executable instructions are executed by the processor, the steps of the model evaluation method according to any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium, characterized in that: It stores computer executable instructions, which, when executed by a processor, implement the steps of the model evaluation method described in any one of claims 1 to 7.

11. A computer program product, characterized in that The method comprises a computer program / instruction, which, when executed by a processor, implements the steps of the model evaluation method according to any one of claims 1 to 7.