Model Evaluation Method, Device, Electronic Device and Storage Medium

By adopting unified scoring standards and processes in large model evaluation, quantifying the external link knowledge base data and building standard evaluation sets, the problem of low accuracy of large model evaluation is solved, and the efficiency and accuracy of model evaluation and training are improved.

CN117829294BActive Publication Date: 2025-05-30BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311767414.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-12-20
Publication Date
2025-05-30
Estimated Expiration
2043-12-20

AI Technical Summary

Technical Problem

The accuracy of large-scale model evaluation is poor, which affects the speed of large-scale model tuning and research and development.

Method used

By providing a unified scoring standard and scoring process, the data in the external link knowledge base is quantified, a standard evaluation set is constructed, the large model is evaluated, the target score is determined, and the evaluation results of the model to be evaluated are finally obtained.

Benefits of technology

It improves the effectiveness and accuracy of model evaluation, makes the evaluation and training of large models more quantitative, promotes model fine-tuning and evaluation, and facilitates the research and development efficiency and performance of large models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117829294B_ABST
    Figure CN117829294B_ABST
Patent Text Reader

Abstract

The present disclosure provides a model evaluation method, apparatus, electronic device, and storage medium. The present disclosure relates to the field of computer technologies, and particularly to technical fields such as artificial intelligence, deep learning, data processing, and large models. The specific implementation solution is as follows: determining an evaluation set based on multiple evaluation tasks, where the evaluation set includes test documents corresponding to each evaluation task, as well as test questions and standard answers corresponding to the test documents; testing the model to be evaluated based on the test documents corresponding to each evaluation task in the evaluation set to obtain test results output by the model to be evaluated; determining a target score corresponding to each evaluation task based on the difference between the intelligent answer and the standard answer corresponding to each evaluation task; and determining an evaluation result of the model to be evaluated based on the target score corresponding to each evaluation task. According to the solution of the present disclosure, quantitative analysis can be performed on the evaluation and training of the model, and the accuracy of model evaluation can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technologies, and particularly to technologies such as artificial intelligence, deep learning, data processing, and large models. Background Art

[0002] Over time, large model services based on external link knowledge bases have been increasingly widely applied in the medical field. At the same time, large model tuning strategies and large model training and evaluation methods based on external link knowledge bases have also become an indispensable part of the daily service provision process. In related technologies, the accuracy of large model evaluation is poor. The accuracy of large model evaluation will seriously affect the speed of large model tuning and research and development. Summary of the Invention

[0003] The present disclosure provides a model evaluation method, apparatus, electronic device, and storage medium.

[0004] According to a first aspect of the present disclosure, there is provided a model evaluation method, including:

[0005] Determining an evaluation set based on multiple evaluation tasks, where the evaluation set includes test documents corresponding to each evaluation task, as well as test questions and standard answers corresponding to the test documents;

[0006] Testing the model to be evaluated based on the test documents corresponding to each evaluation task in the evaluation set, obtaining test results output by the model to be evaluated, where the test results include intelligent answers output based on the test questions included in the test documents;

[0007] Determining a target score corresponding to each evaluation task based on the difference between the intelligent answer and the standard answer corresponding to each evaluation task; wherein, each evaluation task corresponds to at least two sub-evaluation results from two target objects, and the target score is determined based on at least two sub-evaluation results corresponding to each evaluation task;

[0008] Determining an evaluation result of the model to be evaluated based on the target score corresponding to each evaluation task.

[0009] According to a second aspect of the present disclosure, there is provided a model evaluation apparatus, including:

[0010] A first determination module, configured to determine an evaluation set based on multiple evaluation tasks, where the evaluation set includes test documents corresponding to each evaluation task, as well as test questions and standard answers corresponding to the test documents;

[0011] A testing module, configured to test the model to be evaluated based on the test documents corresponding to each evaluation task in the evaluation set, obtaining test results output by the model to be evaluated, where the test results include intelligent answers output based on the test questions included in the test documents;

[0012] A second determination module, configured to determine a target score corresponding to each evaluation task based on the difference between the intelligent answer corresponding to each evaluation task and the standard answer; wherein, each evaluation task corresponds to at least two sub-evaluation results from two target objects, and the target score is determined based on at least two sub-evaluation results corresponding to each evaluation task;

[0013] A third determination module, configured to determine an evaluation result of the model to be evaluated based on the target score corresponding to each evaluation task.

[0014] According to a third aspect of the present disclosure, there is provided an electronic device, including:

[0015] At least one processor;

[0016] A memory communicatively connected to the at least one processor;

[0017] The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method of any embodiment in the present disclosure.

[0018] According to a fourth aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to execute the method of any embodiment in the present disclosure.

[0019] According to a fifth aspect of the present disclosure, there is provided a computer program product, including a computer program stored on a storage medium, where the computer program, when executed by a processor, implements the method of any embodiment in the present disclosure.

[0020] According to the solution of the present disclosure, by adopting a unified scoring standard and scoring process, the data in the external link knowledge base is quantified, so as to perform quantitative analysis on the evaluation and training of the large model, and improve the effectiveness of model evaluation; by using a standard evaluation set to evaluate the large model, the performance of the large model can be evaluated more accurately, which is convenient for better model fine-tuning and evaluation, thereby improving the R & D efficiency of the large model and enhancing the performance of the large model.

[0021] The above summary is only for the purpose of the specification and is not intended to be limiting in any way. In addition to the above-described illustrative aspects, embodiments, and features, further aspects, embodiments, and features of the present application will be readily apparent by reference to the drawings and the following detailed description. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] In the accompanying drawings, unless otherwise specified, the same reference numerals throughout the several views denote the same or similar components or elements. These drawings are not necessarily drawn to scale. It should be understood that these drawings depict only some embodiments disclosed in accordance with the present application and should not be regarded as limiting the scope of the present application.

[0023] Figure 1 is a schematic flowchart of a model evaluation method according to an embodiment of the present disclosure;

[0024] Figure 2 is a flowchart of large model pre-training evaluation based on an external link knowledge base according to an embodiment of the present disclosure;

[0025] Figure 3 is a framework diagram of large model pre-training task classification and question formulation based on an external link knowledge base according to an embodiment of the present disclosure;

[0026] Figure 4 is a framework diagram of scoring criteria and task statistical dimensions for large model pre-training based on an external link knowledge base according to an embodiment of the present disclosure;

[0027] Figure 5 is a scoring flowchart of large model pre-training based on an external link knowledge base according to an embodiment of the present disclosure;

[0028] Figure 6 is a schematic diagram of scoring metric dimensions for large model pre-training based on an external link knowledge base according to an embodiment of the present disclosure;

[0029] Figure 7 is a schematic structural diagram of a model evaluation device according to an embodiment of the present disclosure;

[0030] Figure 8 is a schematic diagram of a scenario of a model evaluation method according to an embodiment of the present disclosure;

[0031] Figure 9 is a schematic structural diagram of an electronic device for implementing the model evaluation method according to an embodiment of the present disclosure. Detailed Embodiments

[0032] The following describes exemplary embodiments of the present disclosure in conjunction with the accompanying drawings. Various details of the embodiments of the present disclosure are included to assist in understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted below.

[0033] In the embodiments of the specification, claims and the above-mentioned drawings of the present disclosure, terms such as "first", "second" and "third" are used to distinguish similar objects and do not necessarily describe a specific order or sequence. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a method, system, product or device comprising a series of steps or units does not necessarily limit to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0034] In the related art, with the continuous development of big data and artificial intelligence technologies, the application of large models based on external link knowledge bases has become increasingly widespread. However, there are many problems in the evaluation and training of such models. The most prominent problems are the inconsistency of scoring criteria and the non-standardization of scoring processes, which lead to the divergence and difficulty in quantification of the evaluation and training of large models. Due to the lack of a unified standard, it is also difficult to effectively optimize the fine-tuning process of large models. During the service provision process, it is difficult to explain the degree and direction of optimization of the fine-tuning process, which brings great inconvenience to the fine-tuning and evaluation of large models. Therefore, the efficiency and quality of the optimization and research and development work of large models are severely restricted.

[0035] To at least partially address one or more of the above problems and other potential problems, the present disclosure proposes a model evaluation method. By providing a unified evaluation standard and process, the data in the external link knowledge base is quantified, so as to conduct quantitative analysis on the evaluation and training of large models, providing a unified standard and direction for the fine-tuning and evaluation of large models, and effectively solving the problem of low efficiency in the optimization and research and development work of large models. At the same time, since the present disclosure uses a standard evaluation set to evaluate large models, it can conduct quantitative analysis on the evaluation and training of models, can more accurately evaluate the performance of large models, is convenient for better model fine-tuning and evaluation, and thus improves the research and development efficiency of large models, enhances the performance of large models, and provides strong support for the research and development work of large models. Moreover, since the present disclosure provides a standardized model evaluation system, using the method of the present disclosure to evaluate models can improve the effectiveness, fairness and universality of model evaluation.

[0036] An embodiment of the present disclosure provides a model evaluation method. Figure 1It is a schematic flowchart of a model evaluation method according to an embodiment of the present disclosure. The model evaluation method can be applied to a model evaluation device. The model evaluation device is located in an electronic device. The electronic device includes, but is not limited to, a fixed device and / or a mobile device. For example, the fixed device includes, but is not limited to, a server, and the server can be a cloud server or a general server. For example, the mobile device includes, but is not limited to, a mobile phone, a tablet computer, a vehicle-mounted terminal, etc. In some possible implementation manners, the model evaluation method can also be implemented by a processor calling computer-readable instructions stored in a memory. As Figure 1 shown, the model evaluation method includes:

[0037] S101: Determine an evaluation set based on multiple evaluation tasks. The evaluation set includes test documents corresponding to each evaluation task, as well as test questions and standard answers corresponding to the test documents;

[0038] S102: Based on the test documents corresponding to each evaluation task in the evaluation set, test the model to be evaluated, and obtain test results output by the model to be evaluated. The test results include intelligent answers output based on the test questions included in the test documents;

[0039] S103: Determine the target score corresponding to each evaluation task based on the difference between the intelligent answer and the standard answer corresponding to each evaluation task; wherein, each evaluation task corresponds to at least two sub-evaluation results from two target objects, and the target score is determined based on at least two sub-evaluation results corresponding to each evaluation task;

[0040] S104: Determine the evaluation result of the model to be evaluated based on the target score corresponding to each evaluation task.

[0041] In the embodiment of the present disclosure, the model to be evaluated is a large model based on an external link knowledge base. For example, the model to be evaluated is a medical large model based on an external link knowledge base. The large model of the present disclosure can be a trained large model, an improved large model, or a pre-training model of a large model.

[0042] In the embodiment of the present disclosure, the external link knowledge base refers to a knowledge base containing a large number of external links. It is a database that aggregates knowledge and information from different sources. These links can guide users to jump to other websites or resources to obtain more in-depth information. The purpose of the external link knowledge base is to provide comprehensive and extensive information, enabling users to conduct in-depth research and understand relevant topics. These external links can be citations, references, supports, or extensions of the current topic being discussed.

[0043] In some embodiments, for each evaluation task, multiple test documents may be included. For each test document, multiple test questions may be included. It should be noted that the test questions correspond one-to-one with the standard answers. For example, if N test questions are determined based on Test Document 1, then for each of the N test questions in Test Document 1, there is a unique standard answer, that is, the 1st test question corresponds to the 1st standard answer, the 2nd test question corresponds to the 2nd standard answer, and the Nth test question corresponds to the Nth standard answer.

[0044] In some embodiments, for different models, the number of test documents corresponding to different types of evaluation tasks may be the same or different. For example, Evaluation Task 1 has k1 test documents, and Evaluation Task 2 has k2 test documents. The values of k1 and k2 may be equal or unequal.

[0045] In some embodiments, for different models, the number of test questions included in the test documents for the same evaluation task may be the same or different. For example, Evaluation Task 1 has k1 test documents. The 1st test document includes q1 test questions, and the 2nd test document includes q2 test questions. The values of q1 and q2 may be equal or unequal.

[0046] In some embodiments, the evaluation set is a data set determined based on M types of evaluation tasks, where M is an integer greater than or equal to 2. The value of M can be set or adjusted according to user needs. For example, it can be extracted from the five dimensions of Population Intervention Comparison Outcome Study design (PICOS), adverse reaction monitoring, single medical literature summary, multiple medical literature summary, and document-level relationship extraction. Set 5 types of evaluation tasks, and provide no less than 5 document contents as test documents for each evaluation task. Each test document provides 5 to 10 test questions.

[0047] In some embodiments, multiple evaluation tasks include at least two of the following:

[0048] Evaluation tasks related to PICOS extraction;

[0049] Evaluation tasks related to adverse reaction monitoring;

[0050] Evaluation tasks related to single medical literature summary;

[0051] Evaluation tasks related to multiple medical literature summary;

[0052] Evaluation tasks related to document-level relationship extraction.

[0053] Among them, P in PICOS represents the research object, for example, a specific group of people suffering from a certain disease. Among them, I in PICOS represents the intervention measure, for example, the treatment plan or exposure factor of the intervention group. Among them, C in PICOS represents the control measure, for example, the treatment plan or exposure factor of the control group. Among them, O in PICOS represents the important clinical outcome, for example, effectiveness, survival rate. Among them, S in PICOS represents the research type, for example, what is the research design, randomized controlled study, cohort study, case-control study, etc.

[0054] In some embodiments, test documents can be selected according to the type of evaluation task. For example, in the process of selecting test documents, 8 test documents are formulated respectively for the five types of evaluation tasks of PICOS extraction, adverse reaction monitoring, single medical literature summary, multiple medical literature summary, and document detailed extraction. 10 test questions and 10 standard answers are formulated for each test document, and finally an evaluation set is generated.

[0055] In some embodiments, test questions for various evaluation tasks can be formulated according to the type of evaluation task. Each type of evaluation task includes test questions that can reflect the characteristics of this type of evaluation task as completely as possible. In practical applications, test question examples can be set for various evaluation tasks, so as to facilitate other users to select test documents for the large model from the external link knowledge base they enjoy, and refer to the test question examples to set test questions for the selected test documents.

[0056] In some embodiments, for the evaluation tasks related to PICOS extraction, the corresponding test questions can include the following 5 categories:

[0057] (1) Who is the research object of this study? / what′s participant in this study?

[0058] (2) What scientific problem does this experiment address? / what′s background in this study?

[0059] (3) What is the intervention measure of this trial? / what′s intervention in this study?

[0060] (4) What are the research results of this trial? / what research methods were employed inthis study?

[0061] (5) What is the research conclusion of this trial? / what′s the conclusion of this study?

[0062] In some embodiments, for the assessment tasks related to adverse reaction monitoring, the corresponding test questions may include the following four categories:

[0063] (1) What are the classification or grouping names of the patients, the grouping method, and the number of patients?

[0064] (2) What are the treatment regimens and drug names used for patients in different groups?

[0065] (3) What are the adverse reactions occurred in patients in different groups and the specific number of patients? Please explain the reasons.

[0066] (4) What is the severity of the adverse reactions occurred in patients in different groups?

[0067] In some embodiments, for the assessment tasks related to the summary of a single medical literature, the corresponding test questions may include the following two categories:

[0068] (1) Please summarize the main viewpoints of this article;

[0069] (2) Please summarize the important information of this article.

[0070] In some embodiments, for the assessment tasks related to the summary of multiple medical literatures, the corresponding test questions may include the following two categories:

[0071] (1) Please summarize the main viewpoints of these articles;

[0072] (2) Please summarize the important information of these articles.

[0073] In some embodiments, for the assessment tasks related to the extraction of document-level detailed information, the corresponding test questions may include the following two categories:

[0074] (1) Please summarize the valid numerical values in this article;

[0075] (2) Please summarize the detailed content of this article.

[0076] It should be noted that the types included in the test questions corresponding to the above different types of assessment tasks are only for exemplary illustration, and do not limit all possible type numbers and contents included in the test questions corresponding to different types of assessment tasks. Only exhaustive listing is not done here.

[0077] Examples of assessment tasks, their corresponding test questions, and examples of test document contents can be referred to Table 1.

[0078]

[0079]

[0080] Table 1

[0081] It should be noted that the example test questions and the example test document content corresponding to each evaluation task in Table 1 above are only for illustrative purposes, and do not limit all possible contents included in the example test questions corresponding to different types of evaluation tasks, nor do they limit all possible contents included in the test document content corresponding to different types of evaluation tasks. Here, exhaustive listing is not done.

[0082] In the embodiments of the present disclosure, the target object can be an electronic product with an evaluation function, a robot with an evaluation function, or a real evaluation personnel. The target object can give a score for each evaluation task based on the difference between the intelligent answer and the standard answer corresponding to each evaluation task.

[0083] In the embodiments of the present disclosure, the evaluation result is used to reflect the performance of the model to be evaluated. For example, the evaluation result can be a summary of the performance of the model to be evaluated. The evaluation result can be represented by a total score value. The higher the total score value, the better the model performance. Here, the maximum value of the total score value can be determined according to requirements. If the maximum value of the total score value is 100, the value range of the total score value can be specifically represented by 0 to 100. If the maximum value of the total score value is 3, the value range of the total score value can be specifically represented by 0 to 3. The evaluation result can also be represented by a grade value. The value range of the grade value includes excellent, good, medium, pass, and fail; among them, the model performance corresponding to excellent, good, medium, pass, and fail decreases in turn. For example, if the grade value is fail, it indicates that the model performance is very poor. If the grade value is pass or medium, it indicates that the model performance is average. If the grade value is good or excellent, it indicates that the model performance is very good. The above is only for illustrative purposes and does not limit all possible representation methods included in the evaluation result. Here, exhaustive listing is not done.

[0084] The technical solution described in the embodiments of the present disclosure determines an evaluation set based on multiple evaluation tasks, which can improve the coverage breadth of the evaluation set, contribute to improving the evaluation breadth of the model, and thus improve the accuracy of the evaluation results of the model. By formulating test documents and standard answers for each evaluation task according to a unified standard, and evaluating the model to be evaluated based on the test documents and standard answers to obtain the evaluation results of the model to be evaluated, the data in the external link knowledge base can be quantified, so that the evaluation and training of the large model can be quantitatively analyzed, the performance of the large model can be evaluated more accurately, it is convenient to better perform model fine-tuning and evaluation, thereby improving the R & D efficiency of the large model, enhancing the performance of the large model, and providing strong support for the R & D work of the large model. At the same time, the present disclosure provides a unified evaluation standard and evaluation process, quantifies the data in the external link knowledge base, can quantitatively analyze the evaluation and training of the large model, provides a unified standard and direction for the fine-tuning and evaluation of the large model, and effectively solves the problem of low efficiency in the fine-tuning and R & D work of the large model. In addition, by evaluating the model through the standardized model evaluation system provided by the present disclosure, the effectiveness, fairness and universality of model evaluation can be improved.

[0085] In the embodiments of the present disclosure, each sub-evaluation result is obtained by each target object evaluating the difference between the intelligent answer corresponding to the test question and the standard answer; wherein, each test question corresponds to at least two target objects.

[0086] Here, the sub-evaluation result of each evaluation task is relative to the evaluation result of the model to be evaluated.

[0087] Here, each test question corresponds to at least two target objects, which can be understood as each test question is evaluated by at least two target objects. Since each test question corresponds to at least two target objects, it follows that each test document has at least two target objects, which in turn leads to each evaluation task having at least two target objects, and further leads to the model to be evaluated having at least two target objects.

[0088] In some embodiments, obtain the i-th sub-evaluation result determined by the i-th target object. The i-th sub-evaluation result is determined by the i-th target object based on the difference i between the intelligent answer yi of the test question y and the standard answer yi. The i-th target object obtains the difference i after comparing the intelligent answer yi and the standard answer yi; obtain the j-th sub-evaluation result determined by the j-th target object. The j-th sub-evaluation result is determined by the j-th target object based on the difference j between the intelligent answer yj and the standard answer yj. The j-th target object obtains the difference j after comparing the intelligent answer yj and the standard answer yj; determine the sub-evaluation result of the test question y based on the i-th sub-evaluation result and the j-th sub-evaluation result. Wherein, i ≠ j.

[0089] Exemplarily, Q target objects evaluate based on the differences between the intelligent answers of the model to be evaluated and the standard answers. Specifically, for each evaluation task, at least two target objects are involved in the evaluation. For example, the first target object can obtain a first sub-evaluation result from the first dimension, that is, obtain the first sub-evaluation result based on the first evaluation task; similarly, the second target object can obtain a second sub-evaluation result from the first dimension, that is, obtain the second sub-evaluation result based on the first evaluation task. The sub-evaluation result of the first evaluation task is obtained based on the first sub-evaluation result of the first target object and the second sub-evaluation result of the second target object. For example, at least two target objects evaluate the second evaluation task. Specifically, the third target object can obtain a third sub-evaluation result from the second dimension, that is, obtain the third sub-evaluation result based on the second evaluation task; similarly, the fourth target object can obtain a fourth sub-evaluation result from the second dimension, that is, obtain the fourth sub-evaluation result based on the second evaluation task. The sub-evaluation result of the second evaluation task is obtained based on the third sub-evaluation result of the third target object and the fourth sub-evaluation result of the fourth target object. For example, at least two target objects evaluate the third evaluation task. Specifically, the fifth target object can obtain a fifth sub-evaluation result from the third dimension, that is, obtain the fifth sub-evaluation result based on the third evaluation task; similarly, the sixth target object can obtain a sixth sub-evaluation result from the third dimension, that is, obtain the sixth sub-evaluation result based on the third evaluation task. The sub-evaluation result of the third evaluation task is obtained based on the fifth sub-evaluation result of the fifth target object and the sixth sub-evaluation result of the sixth target object. For example, at least two target objects evaluate the fourth evaluation task. Specifically, the seventh target object can obtain a seventh sub-evaluation result from the fourth dimension, that is, obtain the seventh sub-evaluation result based on the fourth evaluation task; similarly, the eighth target object can obtain an eighth sub-evaluation result from the fourth dimension, that is, obtain the eighth sub-evaluation result based on the fourth evaluation task. The sub-evaluation result of the fourth evaluation task is obtained based on the seventh sub-evaluation result of the seventh target object and the eighth sub-evaluation result of the eighth target object. For example, at least two target objects evaluate the fifth evaluation task. Specifically, the ninth target object can obtain a ninth sub-evaluation result from the fifth dimension, that is, obtain the ninth sub-evaluation result based on the fifth evaluation task; similarly, the tenth target object can obtain a tenth sub-evaluation result from the fifth dimension, that is, obtain the tenth sub-evaluation result based on the fifth evaluation task. The sub-evaluation result of the fifth evaluation task is obtained based on the ninth sub-evaluation result of the ninth target object and the tenth sub-evaluation result of the tenth target object.

[0090] In practical applications, at least two target objects evaluate the first evaluation task, the second evaluation task, the third evaluation task, the fourth evaluation task, and the fifth evaluation task. The above-mentioned first target object, third target object, fifth target object, seventh target object, and ninth target object may be the same target object. The above-mentioned second target object, fourth target object, sixth target object, eighth target object, and tenth target object may be the same target object.

[0091] In this way, since each test question corresponds to at least two target objects, it helps to improve the fairness and effectiveness of each sub-evaluation result, thereby improving the accuracy of the target score of each evaluation task and the accuracy of the evaluation result of the model to be evaluated.

[0092] In the embodiments of the present disclosure, determining the target score corresponding to each evaluation task includes: determining at least two sub-evaluation results corresponding to the test documents included in each evaluation task based on at least two sub-evaluation results corresponding to the test questions included in each evaluation task; determining at least two sub-evaluation results corresponding to each evaluation task based on at least two sub-evaluation results corresponding to the test documents included in each evaluation task; and determining the target score corresponding to each evaluation task based on at least two sub-evaluation results corresponding to each evaluation task.

[0093] Exemplarily, evaluation task 1 corresponds to n test documents, and each test document corresponds to m test questions. Specifically, at least two sub-evaluation results corresponding to each of the m test questions in test document 1 are used to determine at least two sub-evaluation results corresponding to test document 1; at least two sub-evaluation results corresponding to each of the n test documents are used to determine at least two sub-evaluation results corresponding to evaluation task 1; and the target score corresponding to evaluation task 1 is determined based on at least two sub-evaluation results corresponding to evaluation task 1.

[0094] In this way, determining the target score corresponding to each evaluation task based on at least two sub-evaluation results corresponding to each evaluation task can improve the accuracy of the target score of each evaluation task and the accuracy of the evaluation result of the model to be evaluated.

[0095] In the embodiments of the present disclosure, the sub-evaluation result includes a specific score value. Determining the target score corresponding to each evaluation task based on at least two sub-evaluation results corresponding to each evaluation task may include: determining the target score corresponding to each evaluation task based on whether the specific score values included in at least two sub-evaluation results corresponding to each evaluation task are in the same score range.

[0096] In the embodiments of the present disclosure, the evaluation criteria can be divided into multiple levels as needed. For example, the evaluation criteria are divided into 5 levels, namely: the large model fails to recall effective content, the large model recalls answers that do not meet expectations, the large model recalls some answers that meet expectations, the large model recalls answers that basically meet expectations, and the large model recalls answers that completely meet expectations.

[0097] In the embodiments of the present disclosure, the detailed definitions of the 5-level evaluation criteria are shown in Table 2.

[0098]

[0099]

[0100] Table 2

[0101] In practical applications, the above 5 levels can be set as independent levels, or any two adjacent levels can be combined. For example, the specific score values 0 and 1 are divided into the same scoring level (denoted as the first scoring level), the specific score values 2 and 3 are divided into the same scoring level (denoted as the second scoring level), and the specific score value -1 is divided into an independent scoring level (denoted as the third scoring level). Since the specific score values of the third level are all -1 and belong to the situation of not recalling answers, when the specific score values included in at least two sub-evaluation results corresponding to the evaluation task are all in the third scoring level, the target score corresponding to the evaluation task is -1.

[0102] In this way, based on whether the specific score values included in at least two sub-evaluation results corresponding to each evaluation task are in the same scoring level, determining the target score corresponding to each evaluation task can further improve the accuracy of the target score.

[0103] In the embodiments of the present disclosure, based on whether the specific score values included in at least two sub-evaluation results corresponding to each evaluation task are in the same scoring level, determining the target score corresponding to each evaluation task includes: when the specific score values included in at least two sub-evaluation results corresponding to the first evaluation task are in the same scoring level and are not equal to 0, calculating the average of the specific score values included in at least two sub-evaluation results corresponding to the first evaluation task to obtain the target score corresponding to the first evaluation task, where the first evaluation task is any one of multiple evaluation tasks.

[0104] Exemplarily, the first evaluation task corresponds to two sub-evaluation results (sub-evaluation result 1 and sub-evaluation result 2), where the specific score value in sub-evaluation result 1 is 2 and the specific score value in sub-evaluation result 2 is 3. Calculating the average of the specific score values of sub-evaluation result 1 and sub-evaluation result 2, the target score corresponding to the first evaluation task is (2 + 3) / 2 = 2.5.

[0105] Exemplarily, if there are three sub - evaluation results (sub - evaluation result 1, sub - evaluation result 2, and sub - evaluation result 3) corresponding to the first evaluation task, where the specific score value in sub - evaluation result 1 is 2, the specific score value in sub - evaluation result 2 is 3, and the specific score value in sub - evaluation result 3 is 3, then the average value of the specific score values of sub - evaluation result 1, sub - evaluation result 2, and sub - evaluation result 3 is calculated, and the target score of the first evaluation task is obtained as (2 + 3 + 3) / 3 = 2.67.

[0106] Among them, the average value calculation can adopt methods such as simple arithmetic mean or weighted mean, etc. The present disclosure does not limit the calculation method of the average value.

[0107] In this way, when the specific score values included in at least two sub - evaluation results corresponding to the first evaluation task are in the same score range and are not equal to 0, calculating the average value of the specific score values included in at least two sub - evaluation results corresponding to the first evaluation task to obtain the target score corresponding to the first evaluation task can further improve the accuracy of the target score.

[0108] In the embodiments of the present disclosure, based on whether the specific score values included in at least two sub - evaluation results corresponding to each evaluation task are in the same score range, determining the target score corresponding to each evaluation task includes: when the specific score values included in at least two sub - evaluation results corresponding to the first evaluation task are not in the same score range and are not equal to 0, obtaining the target score corresponding to the first evaluation task by re - evaluating the first evaluation task, where the first evaluation task is any one of multiple evaluation tasks.

[0109] Exemplarily, if there are two sub - evaluation results (sub - evaluation result 1 and sub - evaluation result 2) corresponding to the first evaluation task, where the specific score value in sub - evaluation result 1 is 1 and the specific score value in sub - evaluation result 2 is 3, then the first evaluation task is re - evaluated.

[0110] In this way, when the specific score values included in at least two sub - evaluation results corresponding to the first evaluation task are not in the same score range and are not equal to 0, obtaining the target score corresponding to the first evaluation task by re - evaluating the first evaluation task can further improve the accuracy of the target score.

[0111] In the embodiments of the present disclosure, based on whether the specific score values included in at least two sub - evaluation results corresponding to each evaluation task are in the same score range, determining the target score corresponding to each evaluation task includes: when there is a specific score value equal to 0 among the specific score values included in at least two sub - evaluation results corresponding to the first evaluation task, obtaining the target score corresponding to the first evaluation task by re - evaluating the first evaluation task, where the first evaluation task is any one of multiple evaluation tasks.

[0112] Exemplarily, if there are two sub-evaluation results (sub-evaluation result 1 and sub-evaluation result 2) corresponding to the first evaluation task, where the specific score value in sub-evaluation result 1 is 0 and the specific score value in sub-evaluation result 2 is 1, then the first evaluation task is re-evaluated.

[0113] In this way, when there is a specific score value equal to 0 among the specific score values included in at least two sub-evaluation results corresponding to the first evaluation task, by re-evaluating the first evaluation task to obtain the target score corresponding to the first evaluation task, the accuracy of the target score can be further improved.

[0114] In the embodiments of the present disclosure, obtaining the target score corresponding to the first evaluation task by re-evaluating the first evaluation task includes: obtaining multiple score data of the test questions included in the first evaluation task evaluated from multiple evaluation factors respectively, where at least two target objects evaluate the multiple score data; determining the score result corresponding to the test question based on the multiple score data under multiple evaluation factors corresponding to the test question; and obtaining the target score corresponding to the first evaluation task according to the score results corresponding to the test questions included in the first evaluation task.

[0115] In some embodiments, the multiple evaluation factors include at least two of factors such as question difficulty, knowledge point mastery level, mis-evaluation times, etc.

[0116] Exemplarily, the first evaluation task includes n test documents, and each test document corresponds to m test questions. Specifically, for each test question, the first target object scores the test question from evaluation factor 1 and evaluation factor 2 respectively, obtaining score data 11 under evaluation factor 1 and score data 21 under evaluation factor 2; the second target object scores the test question from evaluation factor 1 and evaluation factor 2 respectively, obtaining score data 12 under evaluation factor 1 and score data 22 under evaluation factor 2; and the score result of the test question is determined according to score data 11 under evaluation factor 1, score data 12 under evaluation factor 1, score data 21 under evaluation factor 2, and score data 22 under evaluation factor 2. By using the above method, the score results of each test question included in each test document are obtained; and then the target score corresponding to the first evaluation task is obtained according to the score results corresponding to the test questions included in the first evaluation task.

[0117] In this way, if the sub-evaluation results of different target objects for the same evaluation task have a large gap, further re-evaluation can be carried out to ensure the fairness and reasonableness of the final score result, and the accuracy of the target score can be further improved.

[0118] In the embodiments of the present disclosure, the model evaluation method further includes: in response to the target score corresponding to the first evaluation task being lower than the preset threshold corresponding to the first evaluation task, retrieving the test document corresponding to the first evaluation task; determining an optimization plan for the model to be evaluated based on the test document, test questions, and standard answers corresponding to the first evaluation task; wherein the first evaluation task is any one of multiple evaluation tasks.

[0119] Among them, different types of evaluation tasks may have different preset thresholds.

[0120] In some embodiments, the optimization plan includes, but is not limited to, optimization directions, optimization suggestions, etc.

[0121] Here, the optimization direction is the improvement direction proposed based on the test document, test questions, and standard answers of the first evaluation task.

[0122] Among them, the optimization direction includes, but is not limited to, the following aspects:

[0123] Algorithm optimization: Improve the model's algorithm or introduce a new algorithm to improve the accuracy, efficiency, and stability of the model.

[0124] Data optimization: Improve the quality of the model's input data through data preprocessing, feature engineering, etc., to improve the generalization ability of the model.

[0125] Architecture optimization: Optimize the network structure and layers of the model to improve the training speed and performance of the model.

[0126] Parameter optimization: Optimize the parameter settings of the model through hyperparameter tuning, regularization, etc., to improve the generalization ability and robustness of the model.

[0127] Deployment optimization: Optimize the performance and efficiency of the model in the actual deployment environment to better meet the actual needs.

[0128] Here, the optimization suggestion is the improvement opinion proposed based on the test document, test questions, and standard answers of the first evaluation task.

[0129] Among them, the optimization suggestion includes, but is not limited to, the following aspects:

[0130] Data quality assurance: Ensure data quality, including cleaning data, handling outliers and missing values, etc., to improve the accuracy and stability of the model.

[0131] Feature engineering optimization: Improve the input features of the model through feature selection, feature extraction, and feature transformation, etc., to improve the performance of the model.

[0132] Model Selection and Tuning: Consider trying different types of models and tuning the hyperparameters of the models to find the model that best fits the data.

[0133] Ensemble Learning: Consider adopting ensemble learning methods to combine the prediction results of multiple models and improve the generalization ability of the models.

[0134] Model Compression: For deep learning models, consider model compression techniques such as pruning, quantization, etc. to reduce the model size and computational cost and improve the deployment efficiency of the models.

[0135] Automated Hyperparameter Tuning: Try using automated hyperparameter tuning tools such as grid search, Bayesian optimization, etc. to find the optimal combination of model hyperparameters.

[0136] Deployment Optimization: In the model deployment stage, consider techniques such as model compression and accelerated computing to improve the performance and efficiency of the models in practical applications.

[0137] Thus, in response to the target score of the first evaluation task being lower than the preset threshold corresponding to the first evaluation task, retrieve the test document corresponding to the first evaluation task; based on the test document, test questions, and standard answers corresponding to the first evaluation task, determine the optimization plan for the model to be evaluated, and use the evaluation results as a reference for subsequent model research and application, providing more accurate and reliable support for the development and application of large models.

[0138] In the embodiments of the present disclosure, the evaluation results include the total score value. Based on the target scores corresponding to each evaluation task, determine the evaluation results of the model to be evaluated, including: determining the total score value of the model to be evaluated according to the weight corresponding to each evaluation task and the target score corresponding to each evaluation task.

[0139] Among them, each evaluation task has a corresponding weight, and each evaluation task has a target score.

[0140] Exemplarily, assume there are n evaluation tasks, each task is represented by i (i = 1, 2,..., n), the weight of each task is wi, and the target score is pi. Then the total score value = ∑(wi × (mi / pi)), where ∑ represents summation, mi is the actual score of the model on the i-th task, and pi is the target score of the i-th task.

[0141] It can be understood that the above formula for calculating the total score value is illustrative rather than restrictive. In practical applications, the formula can be adjusted or changed according to requirements. For example, only consider some evaluation tasks, or adjust the weights of some evaluation tasks, etc.

[0142] Thus, by comprehensively considering the target scores of multiple evaluation tasks, the accuracy, effectiveness, and fairness of the total score value of the model to be evaluated can be improved.

[0143] In the embodiments of the present disclosure, the multiple evaluation tasks include at least two of the following:

[0144] Evaluation tasks related to PICOS extraction;

[0145] Evaluation tasks related to adverse reaction monitoring;

[0146] Evaluation tasks related to the summary of a single medical document;

[0147] Evaluation tasks related to the summary of multiple medical documents;

[0148] Evaluation tasks related to document-level relation extraction.

[0149] It should be noted that the above multiple evaluation tasks belong to different test dimensions.

[0150] In this way, by first setting the task dimension and then constructing an evaluation set for the evaluation model, it helps to promote the evaluation tasks of the model and improve the unity and accuracy of model evaluation.

[0151] In the embodiments of the present disclosure, an evaluation set is determined based on multiple evaluation tasks, including: selecting test documents related to each evaluation task from an external link knowledge base according to the characteristics of each evaluation task; generating test questions and corresponding standard answers for the test documents included in each evaluation task according to the test question examples of each evaluation task; and determining the evaluation set based on the test documents, test questions, and standard answers corresponding to each evaluation task.

[0152] In some embodiments, the evaluation set includes test documents, test questions, and standard answers related to each evaluation task.

[0153] In some embodiments, selecting test documents related to each evaluation task from an external link knowledge base according to the characteristics of each evaluation task includes: determining the number of test documents included in each evaluation task and the number of test questions included in each test document according to the characteristics of each evaluation task. For example, 10 test documents are formulated for each task, and 15 test questions are formulated for each test document.

[0154] In practical applications, corresponding test documents are selected from a publicly available medical literature database according to the characteristics of each evaluation task, test questions and standard answers are determined based on the test documents, and an evaluation set is constructed for the pre-training task of a medical large model based on an external link knowledge base based on the test documents, test questions, and standard answers.

[0155] Thus, when constructing the evaluation set, the characteristics and requirements of each evaluation task are fully considered, and representative test documents and test questions are selected to fully cover the knowledge points and skills involved in the evaluation task. At the same time, by setting the number of test documents included in each evaluation task and the number of test questions included in each test document, it helps to ensure the diversity and reliability of the evaluation set, and avoid problems such as overfitting and insufficient generalization ability.

[0156] In the embodiments of the present disclosure, according to the characteristics of each evaluation task, test documents related to each evaluation task are selected from the external link knowledge base, including: determining document features related to each evaluation task according to the characteristics of each evaluation task; and selecting test documents related to each evaluation task from the external link knowledge base according to the document features related to each evaluation task.

[0157] Thus, according to the document features related to each evaluation task, test documents related to each evaluation task are selected from the external link knowledge base. Using this test document to formulate the evaluation set can more accurately evaluate the performance of the large model and improve the R & D efficiency and performance of the large model.

[0158] In the embodiments of the present disclosure, the evaluation result includes an overall level metric value. Determining the evaluation result of the model to be evaluated may include: determining the overall level metric value of the model to be evaluated according to multiple first evaluation dimensions; where the multiple first evaluation dimensions at least include the following two dimensions: accuracy rate, weighted average score, and recall rate.

[0159] In some embodiments, the formula for calculating the accuracy rate is: (the number of questions scored 2 + the number of questions scored 3) / (the total number of questions - the number of questions not recalled).

[0160] In some embodiments, the formula for calculating the weighted average score is: (0 * the number of questions with a final score of 0 + 1 * the number of questions with a final score of 1 + 2 * the number of questions with a final score of 2 + 3 * the number of questions with a final score of 3) / (the total number of questions - the number of questions not recalled).

[0161] In some embodiments, the formula for calculating the recall rate is: (the number of questions scored 0 + the number of questions scored 1 + the number of questions scored 2 + the number of questions scored 3) / the total number of questions.

[0162] Thus, determining the overall level metric value of the model to be evaluated according to multiple first evaluation dimensions can improve the accuracy of the evaluation result, more accurately evaluate the performance of the large model, and thus improve the R & D efficiency and performance of the large model.

[0163] In the embodiments of the present disclosure, the evaluation result includes local-level metric values. Determining the evaluation result of the model to be evaluated includes: determining the local-level metric values of the model to be evaluated according to multiple second evaluation dimensions; where the multiple second evaluation dimensions at least include the following two dimensions: the test effect of the number of test documents, the test effect of the language of the test documents, and the test effect of the types of evaluation tasks.

[0164] Among them, the number of test documents can be single or multiple.

[0165] Among them, the language of the test documents can be Chinese or foreign language.

[0166] Among them, the types of evaluation tasks include but are not limited to five categories: PICOS extraction, adverse reaction monitoring, single medical literature summary, multiple medical literature summary, and detailed document extraction.

[0167] In this way, determining the local-level metric values of the model to be evaluated according to multiple second evaluation dimensions helps to quickly determine the performance optimization direction of the large model, thereby improving the R & D efficiency and performance of the large model.

[0168] Figure 2 shows a flowchart of pre-training evaluation of a large model based on an external link knowledge base according to an embodiment of the present disclosure, as Figure 2 shown, the pre-training evaluation process includes:

[0169] S201: Determine the task dimensions of the pre-training of the large model based on the external link knowledge base, and then execute S202;

[0170] S202: Obtain pre-training data and promote the pre-training task, and then execute S203;

[0171] S203: Determine the evaluation set of the pre-training, and test the large model based on the external link knowledge base based on the evaluation set, and then execute S204;

[0172] S204: Cross-evaluate tasks of multiple target objects, and then execute S205;

[0173] S205: Calculate and combine the scores provided by multiple target objects, and then execute S206;

[0174] S206: Statistically analyze the scores according to the task metric dimensions.

[0175] Here, after obtaining the score of the large model based on the external link knowledge base according to S206, return to S202. Based on S202 to S206, periodically promote the pre-training task of the large model based on the external link knowledge base.

[0176] In practical applications, corresponding documents can be selected from publicly available medical literature databases to construct an evaluation set. When constructing the evaluation set, it is necessary to consider the characteristics and requirements of each type of task, and select representative test documents and test questions to fully cover the knowledge points and skills involved in that type of task. At the same time, it is also necessary to ensure the diversity and reliability of the evaluation set to avoid problems such as overfitting and insufficient generalization ability.

[0177] When pre-training a medical large model, use a pre-training task dataset to pre-train the medical large model to improve the performance and accuracy of the medical large model in handling medical domain tasks. During the pre-training process, it is necessary to select an appropriate model architecture and optimization algorithm, and adjust the hyperparameters according to the task requirements. For example, supervised learning or unsupervised learning methods can be used for pre-training, and the choice of specific method depends on the task type and data characteristics.

[0178] In this way, a simple and comprehensive standard evaluation process is provided for the pre-training of medical large models. Through the implementation of this solution, the performance of the large model can be evaluated more accurately, and the R & D efficiency and performance of the large model can be improved. At the same time, this solution also has high scalability and flexibility, and can easily adapt to different application scenarios and requirements.

[0179] Figure 3 Shows the classification of pre-training tasks of the large model based on the external link knowledge base and the framework diagram for formulating questions, as Figure 3 shown, the classification of pre-training tasks of the medical large model based on the external link knowledge base is divided into Task 1, Task 2, Task 3, Task 4, and Task 5. This Task 1 is PICOS extraction, this Task 2 is adverse reaction monitoring, this Task 3 is single-document summary, this Task 4 is multi-document summary, and this Task 5 is detailed information extraction from documents. This PICOS extraction task can include: selection of target PICOS documents, drafting of target PICOS questions, and obtaining answers from the medical large model service. This adverse reaction monitoring task can include: selection of target adverse reaction documents, drafting of target adverse reaction questions, and obtaining answers from the medical large model service. This single-document summary task can include: selection of target single-summary documents, drafting of target single-summary document questions, and obtaining answers from the medical large model service. This multi-document summary task can include: selection of target multi-summary documents, drafting of target multi-summary document questions, and obtaining answers from the medical large model service. This detailed information extraction task from documents can include: selection of target detailed information extraction documents, drafting of target detailed information extraction questions, and obtaining answers from the medical large model service.

[0180] In this way, a simple and comprehensive evaluation set formulation solution is provided for the medical large model. Through the implementation of this solution, the performance of the large model can be evaluated more accurately, and the R & D efficiency and performance of the large model can be improved.

[0181] Figure 4 shows the scoring criteria and task statistic dimension framework diagram for large model pre-training based on an external link knowledge base, as Figure 4 shown, the measurement dimensions of the medical large model document understanding scoring device based on the external link knowledge base can include Dimension 1, Dimension 2, and Dimension 3. Dimension 1 is the evaluation of the number of external link knowledge articles (single article / multiple articles), Dimension 2 is the evaluation of the language type of external link knowledge (Chinese / English), and Dimension 3 is the evaluation of the task category of external link knowledge. The scoring criteria can include a score of 3 (i.e., fully compliant), and the answer accuracy can be based on the document content to correctly answer the question, including almost all key information points; when the score is 2 (basically compliant), the answer accuracy can be to partially answer the question based on the document content, with a small part of the key points missing and a small amount of divergence in the answer, but there are errors that can be found by reading the document content; when the score is 1 (partially compliant), the answer accuracy can be to partially answer the question based on the document content, but there are some obvious errors or omissions; when the score is 0 (non-compliant), the answer accuracy can be that there are multiple errors or omissions in the answer and the answer is not relevant to the instruction, answering off-topic; when the score is -1 (not recalled), the answer accuracy can output the recalled content, which can be: "Sorry, the question you asked cannot be answered from the document. You can ask the question again."

[0182] In this way, it provides a standard scoring range for the medical large model and a paradigm for formulating evaluation sets for other industries in the future.

[0183] Figure 5 is the scoring flow chart for large model pre-training based on an external link knowledge base according to an embodiment of the present disclosure; as Figure 5 shown, the process can include:

[0184] S501: For the same test question, summarize the scores of multiple target objects, and then execute S502;

[0185] S502: Determine whether the scores of multiple target objects are the same value. If so, execute S505; if not, execute S503;

[0186] S503: Determine whether the scores of multiple target objects are in the same scoring range; if so, execute S506; if not, execute S504;

[0187] For example, the first scoring range includes score values of 0 and 1, and the second scoring range includes score values of 2 and 3.

[0188] S504: If the scores of multiple target objects are not in the same scoring range, re-scoring is required;

[0189] S505: Take this value as the score value of this test question;

[0190] S506: If there is no score of 0 and the items are in the same scoring range, take the average value as the scoring value of this test item; if there is a score of 0, re-score.

[0191] Here, taking the average value as the scoring value of this test item may further include: taking the closest integer value according to the rounding method as the scoring value of this test item.

[0192] In this way, a standard process for combined scoring is provided for the medical large model, providing a paradigm for the scoring process of the large model.

[0193] Figure 6 It is a schematic diagram of the scoring metric dimensions for large model pre-training based on an external link knowledge base according to an embodiment of the present disclosure. As Figure 6 shown, the final scoring calculation of the medical large model document understanding based on the external link knowledge base may include Dimension 1, Dimension 2, and Dimension 3; Dimension 1 is the accuracy rate, and the calculation formula for Dimension 1 is (the number of questions scored 2 + the number of questions scored 3) / (the total number of questions - the number of questions not recalled); Dimension 2 is the weighted average score, and the calculation formula for Dimension 2 is (0 * the number of questions with a final score of 0 + 1 * the number of questions with a final score of 1 + 2 * the number of questions with a final score of 2 + 3 * the number of questions with a final score of 3) / (the total number of questions - the number of questions not recalled); Dimension 3 is the recall rate, and the calculation formula for Dimension 3 is (the number of questions scored 0 + the number of questions scored 1 + the number of questions scored 2 + the number of questions scored 3 / the total number of questions).

[0194] In this way, a metric index system of the external link knowledge base is provided for the medical large model, providing a paradigm for the scoring standards of other subsequent industries.

[0195] An embodiment of the present disclosure provides a model evaluation device. As Figure 7 shown, the model evaluation device may include: a first determination module 701, configured to determine an evaluation set based on multiple evaluation tasks, where the evaluation set includes a test document corresponding to each evaluation task, as well as test questions and standard answers corresponding to the test document; a test module 702, configured to test the model to be evaluated based on the test document corresponding to each evaluation task in the evaluation set, and obtain a test result output by the model to be evaluated, where the test result includes an intelligent answer output based on the test questions included in the test document; a second determination module 703, configured to determine a target score corresponding to each evaluation task based on the difference between the intelligent answer and the standard answer corresponding to each evaluation task; where each evaluation task corresponds to at least two sub-evaluation results from two target objects, and the target score is determined based on at least two sub-evaluation results corresponding to each evaluation task; a third determination module 704, configured to determine an evaluation result of the model to be evaluated based on the target score corresponding to each evaluation task.

[0196] In some embodiments, each sub-evaluation result is obtained after each target object evaluates the difference between the intelligent answer corresponding to the test question and the standard answer; wherein, there are at least two target objects corresponding to the test question.

[0197] In some embodiments, the second determination module 703 includes: a first determination sub-module, configured to determine at least two sub-evaluation results corresponding to the test document included in each evaluation task based on at least two sub-evaluation results corresponding to the test questions included in each evaluation task; a second determination sub-module, configured to determine at least two sub-evaluation results corresponding to each evaluation task based on at least two sub-evaluation results corresponding to the test document included in each evaluation task; a third determination sub-module, configured to determine the target score corresponding to each evaluation task based on at least two sub-evaluation results corresponding to each evaluation task.

[0198] In some embodiments, the sub-evaluation result includes a specific score value. The third determination sub-module is configured to: determine the target score corresponding to each evaluation task based on whether the specific score values included in at least two sub-evaluation results corresponding to each evaluation task are in the same score range.

[0199] In some embodiments, the first evaluation task is any one of multiple evaluation tasks. The third determination sub-module is configured to: when the specific score values included in at least two sub-evaluation results corresponding to the first evaluation task are in the same score range and are not equal to 0, calculate the average of the specific score values included in at least two sub-evaluation results corresponding to the first evaluation task to obtain the target score corresponding to the first evaluation task.

[0200] In some embodiments, the first evaluation task is any one of multiple evaluation tasks. The third determination sub-module is configured to: when the specific score values included in at least two sub-evaluation results corresponding to the first evaluation task are not in the same score range and are not equal to 0, obtain the target score corresponding to the first evaluation task by re-evaluating the first evaluation task.

[0201] In some embodiments, the first evaluation task is any one of multiple evaluation tasks. The third determination sub-module is configured to: when there is a specific score value equal to 0 among the specific score values included in at least two sub-evaluation results corresponding to the first evaluation task, obtain the target score corresponding to the first evaluation task by re-evaluating the first evaluation task.

[0202] In some embodiments, the third determination sub-module is configured to: obtain multiple scoring data of the test questions included in the first assessment task evaluated from multiple assessment factors respectively, where the multiple scoring data is evaluated by at least two target objects; determine the scoring result corresponding to the test question based on the multiple scoring data under the multiple assessment factors corresponding to the test question; and obtain the target score corresponding to the first assessment task according to the scoring result corresponding to the test questions included in the first assessment task.

[0203] In some embodiments, the first assessment task is any one of multiple assessment tasks. The model assessment device further includes: a retrieval module ( Figure 7 not shown in the figure) configured to retrieve the test document corresponding to the first assessment task in response to the target score corresponding to the first assessment task being lower than the preset threshold corresponding to the first assessment task; a fourth determination module ( Figure 7 not shown in the figure) configured to determine an optimization solution for the model to be assessed based on the test document, test questions, and standard answers corresponding to the first assessment task.

[0204] In some embodiments, the assessment result includes a total score value. The third determination module includes: a fourth determination sub-module configured to determine the total score value of the model to be assessed according to the weight corresponding to each assessment task and the target score corresponding to each assessment task.

[0205] In some embodiments, the multiple assessment tasks include at least two of the following: an assessment task related to PICOS extraction; an assessment task related to adverse reaction monitoring; an assessment task related to single medical literature summarization; an assessment task related to multi-medical literature summarization; an assessment task related to document-level relation extraction.

[0206] In some embodiments, the first determination module 701 includes: a selection sub-module configured to select a test document related to each assessment task from the external link knowledge base according to the characteristics of each assessment task; a generation sub-module configured to generate test questions and the standard answers corresponding to the test questions for the test document included in each assessment task according to the test question examples of each assessment task; and a fifth determination sub-module configured to determine an assessment set based on the test document, test questions, and standard answers corresponding to each assessment task.

[0207] In some embodiments, the selection sub-module is configured to: determine the document features related to each assessment task according to the characteristics of each assessment task; and select the test document related to each assessment task from the external link knowledge base according to the document features related to each assessment task.

[0208] In some embodiments, the evaluation result includes an overall level metric value. The third determination module includes: a sixth determination sub-module, configured to determine the overall level metric value of the model to be evaluated according to a plurality of first evaluation dimensions; wherein, the plurality of first evaluation dimensions include at least two of the following dimensions: accuracy rate, weighted average score, and recall rate.

[0209] In some embodiments, the evaluation result includes a local level metric value. The third determination module includes: a seventh determination sub-module, configured to determine the local level metric value of the model to be evaluated according to a plurality of second evaluation dimensions; wherein, the plurality of second evaluation dimensions include at least two of the following dimensions: test effect of the number of test documents, test effect of the language of test documents, and test effect of the types of evaluation tasks.

[0210] Those skilled in the art should understand that the functions of the various processing modules in the model evaluation device of the embodiments of the present disclosure can be understood with reference to the relevant descriptions of the foregoing model evaluation method. The various processing modules in the model evaluation device of the embodiments of the present disclosure can be implemented by a simulation circuit that implements the functions of the embodiments of the present disclosure, or can be implemented by the operation of software that executes the functions of the embodiments of the present disclosure on an electronic device.

[0211] The model evaluation device of the embodiments of the present disclosure can quantify the data in the external link knowledge base, so as to perform quantitative analysis on the evaluation and training of the model, and improve the effectiveness of model evaluation.

[0212] The embodiments of the present disclosure provide a schematic diagram of a model evaluation scenario, as Figure 8 shown.

[0213] As mentioned above, the model evaluation method provided by the embodiments of the present disclosure is applied to an electronic device. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smart phones, wearable devices, and other similar computing devices.

[0214] Specifically, the electronic device can specifically perform the following operations:

[0215] Determine an evaluation set based on a variety of evaluation tasks, where the evaluation set includes test documents corresponding to each evaluation task, as well as test questions and standard answers corresponding to the test documents;

[0216] Based on the test documents corresponding to each evaluation task in the evaluation set, test the model to be evaluated to obtain the test results output by the model to be evaluated. The test results include intelligent answers output based on the test questions included in the test documents;

[0217] Based on the differences between the intelligent answers corresponding to each evaluation task and the standard answers, determine the target score corresponding to each evaluation task; wherein, each evaluation task corresponds to at least two sub-evaluation results from two target objects, and the target score is determined based on at least two sub-evaluation results corresponding to each evaluation task;

[0218] Based on the target score corresponding to each evaluation task, determine the evaluation result of the model to be evaluated.

[0219] Among them, various evaluation tasks and evaluation sets include the test documents corresponding to each evaluation task, and the test questions and standard answers corresponding to the test documents can be obtained from the data source. The data source can be various forms of data storage devices, such as, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The data source can also represent various forms of mobile devices, such as, personal digital assistants, cellular phones, smart phones, wearable devices, and other similar computing devices.

[0220] It should be understood that Figure 8 the illustrated scenario diagrams are merely illustrative and not restrictive, and those skilled in the art can make various obvious changes and / or substitutions based on Figure 8 the examples, and the obtained technical solutions still fall within the scope of the disclosure of the embodiments of the present disclosure.

[0221] In the technical solution of the present disclosure, the acquisition, storage, and application of the user's personal information involved all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.

[0222] According to the embodiments of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0223] Figure 9 FIG. shows a schematic block diagram of an exemplary electronic device 900 that can be used to implement the embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, personal digital assistants, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described herein and / or claimed.

[0224] As Figure 9As shown, device 900 includes a computing unit 901, which can perform various appropriate actions and processes according to computer programs stored in a read-only memory (ROM) 902 or computer programs loaded from a storage unit 908 into a random access memory (RAM) 903. In the RAM 903, various programs and data required for the operation of device 900 can also be stored. The computing unit 901, the ROM 902, and the RAM 903 are connected to each other via a bus 904. An input / output (I / O) interface 905 is also connected to the bus 904.

[0225] Multiple components in device 900 are connected to the I / O interface 905, including: an input unit 906, such as a keyboard, a mouse, etc.; an output unit 907, such as various types of displays, speakers, etc.; a storage unit 908, such as a magnetic disk, an optical disc, etc.; and a communication unit 909, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 909 allows device 900 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0226] The computing unit 901 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 901 include but are not limited to a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The computing unit 901 executes the various methods and processes described above, such as the model evaluation method. For example, in some embodiments, the model evaluation method can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as the storage unit 908. In some embodiments, part or all of the computer program can be loaded and / or installed onto device 900 via the ROM 902 and / or the communication unit 909. When the computer program is loaded into the RAM 903 and executed by the computing unit 901, one or more steps of the model evaluation method described above can be executed. Alternatively, in other embodiments, the computing unit 901 can be configured to execute the model evaluation method by any other appropriate means (e.g., by means of firmware).

[0227] The various embodiments of the systems and techniques described above in this document may be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application-specific standard products (ASSPs), system on chip (SOC) systems, complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include: being implemented in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which may be a special-purpose or general-purpose programmable processor that receives data and instructions from, and transmits data and instructions to, a storage system, at least one input device, and at least one output device.

[0228] The program code for implementing the methods of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general purpose computer, special purpose computer, or other programmable data processing apparatus, such that the program codes, when executed by the processor or controller, cause the functions / operations specified in the flowchart and / or block diagram to be implemented. The program code may execute entirely on the machine, partly on the machine, as a stand-alone software package partly on the machine and partly on a remote machine, or entirely on the remote machine or server.

[0229] In the context of this disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory, a read-only memory, an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0230] For purposes of providing an interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a cathode ray tube (CRT) or a liquid crystal display (LCD) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide for interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic, speech, or tactile input).

[0231] The systems and techniques described herein can be implemented in a computing system that includes a back-end component (e.g., as a data server), or a computing system that includes a middleware component (e.g., an application server), or a computing system that includes a front-end component (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such back-end, middleware, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of a communication network include: a local area network (LAN), a wide area network (WAN), and the Internet.

[0232] A computer system may include a client and a server. The client and the server are generally far away from each other and usually interact through a communication network. The client-server relationship is generated by computer programs running on respective computers and having a client-server relationship with each other. The server may be a cloud server, a server of a distributed system, or a server combined with a blockchain.

[0233] It should be understood that various forms of the processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. There is no limitation herein.

[0234] The above specific embodiments do not constitute a limitation on the protection scope of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent replacements, and improvements made within the principles of this disclosure shall be included within the protection scope of this disclosure.

Claims

1. A model evaluation method, including: determining an evaluation set based on multiple evaluation tasks, where the evaluation set includes test documents corresponding to each evaluation task, as well as test questions and standard answers corresponding to the test documents; testing the model to be evaluated based on the test documents corresponding to each evaluation task in the evaluation set, obtaining test results output by the model to be evaluated, where the test results include intelligent answers output based on the test questions included in the test documents; determining a target score corresponding to each evaluation task based on the differences between the intelligent answers and the standard answers corresponding to each evaluation task; where each evaluation task corresponds to at least two sub-evaluation results from two target objects, and the target score is determined based on at least two sub-evaluation results corresponding to each evaluation task; determining an evaluation result of the model to be evaluated based on the target scores corresponding to each evaluation task; the determining of the target score corresponding to each evaluation task includes: determining at least two sub-evaluation results corresponding to the test documents included in each evaluation task based on at least two sub-evaluation results corresponding to the test questions included in each evaluation task; determining at least two sub-evaluation results corresponding to each evaluation task based on at least two sub-evaluation results corresponding to the test documents included in each evaluation task; determining the target score corresponding to each evaluation task based on at least two sub-evaluation results corresponding to each evaluation task; the sub-evaluation results include specific score values, and the determining of the target score corresponding to each evaluation task based on at least two sub-evaluation results corresponding to each evaluation task includes: determining the target score corresponding to each evaluation task based on whether the specific score values included in at least two sub-evaluation results corresponding to each evaluation task are in the same score range.

2. The method according to claim 1, wherein, each sub-evaluation result is obtained by each target object evaluating the difference between the intelligent answer corresponding to the test question and the standard answer; where there are at least two target objects corresponding to the test question.

3. The method according to claim 1, wherein, the determining of the target score corresponding to each evaluation task based on whether the specific score values included in at least two sub-evaluation results corresponding to each evaluation task are in the same score range includes: in the case where the specific score values included in at least two sub-evaluation results corresponding to the first evaluation task are in the same score range and are not equal to 0, calculating the average of the specific score values included in at least two sub-evaluation results corresponding to the first evaluation task to obtain the target score corresponding to the first evaluation task, where the first evaluation task is any one of the multiple evaluation tasks.

4. The method according to claim 1, wherein, the determining of the target score corresponding to each evaluation task based on whether the specific score values included in at least two sub-evaluation results corresponding to each evaluation task are in the same score range includes: When the specific score values included in at least two sub - evaluation results corresponding to the first evaluation task are not in the same score range and are not equal to 0, the target score corresponding to the first evaluation task is obtained by re - evaluating the first evaluation task, where the first evaluation task is any one of the multiple evaluation tasks.

5. The method according to claim 1, wherein, determining the target score corresponding to each evaluation task based on whether the specific score values included in at least two sub - evaluation results corresponding to each evaluation task are in the same score range includes: When there is a specific score value equal to 0 among the specific score values included in at least two sub - evaluation results corresponding to the first evaluation task, the target score corresponding to the first evaluation task is obtained by re - evaluating the first evaluation task, where the first evaluation task is any one of the multiple evaluation tasks.

6. The method according to claim 4 or 5, wherein, obtaining the target score corresponding to the first evaluation task by re - evaluating the first evaluation task includes: Obtaining multiple score data of the test questions included in the first evaluation task evaluated from multiple evaluation factors, where at least two target objects evaluate the multiple score data; Based on the multiple score data under multiple evaluation factors corresponding to the test questions, determining the score result corresponding to the test questions; According to the score result corresponding to the test questions included in the first evaluation task, obtaining the target score corresponding to the first evaluation task.

7. The method according to claim 1, further including: In response to the target score corresponding to the first evaluation task being lower than the preset threshold corresponding to the first evaluation task, retrieving the test document corresponding to the first evaluation task; Based on the test document, the test questions, and the standard answers corresponding to the first evaluation task, determining an optimization scheme for the model to be evaluated, where the first evaluation task is any one of the multiple evaluation tasks.

8. The method according to claim 1, wherein, the evaluation result includes a total score value, and determining the evaluation result of the model to be evaluated based on the target score corresponding to each evaluation task includes: Determining the total score value of the model to be evaluated according to the weight corresponding to each evaluation task and the target score corresponding to each evaluation task.

9. The method according to claim 1, wherein, the evaluation result includes an overall level metric value, and determining the evaluation result of the model to be evaluated includes: Determining the overall level metric value of the model to be evaluated according to multiple first evaluation dimensions; where the multiple first evaluation dimensions at least include the following two dimensions: accuracy rate, weighted average score, and recall rate.

10. The method according to claim 1, wherein, the evaluation result includes a local level metric value, and determining the evaluation result of the model to be evaluated includes: Determining the local level metric value of the model to be evaluated according to multiple second evaluation dimensions; Among them, the multiple second evaluation dimensions at least include the following two dimensions: the test effect of the number of test documents, the test effect of the languages of test documents, and the test effect of the types of evaluation tasks.

11. The method according to claim 1, wherein, the determining the evaluation set based on multiple evaluation tasks includes: selecting the test documents related to each evaluation task from the external link knowledge base according to the characteristics of each evaluation task; generating the test questions and the corresponding standard answers for the test documents included in each evaluation task according to the test question examples of each evaluation task; determining the evaluation set based on the test documents, the test questions and the standard answers corresponding to each evaluation task.

12. The method according to claim 11, wherein, the selecting the test documents related to each evaluation task from the external link knowledge base according to the characteristics of each evaluation task includes: determining the document features related to each evaluation task according to the characteristics of each evaluation task; selecting the test documents related to each evaluation task from the external link knowledge base according to the document features related to each evaluation task.

13. The method according to claim 1, wherein, the multiple evaluation tasks include at least the following two: evaluation tasks related to the extraction of PICOS including research object, intervention measure, control measure, outcome, and research type; evaluation tasks related to adverse reaction monitoring; evaluation tasks related to the summary of a single medical literature; evaluation tasks related to the summary of multiple medical literatures; evaluation tasks related to document-level relation extraction.

14. A model evaluation device, including: a first determination module, configured to determine an evaluation set based on multiple evaluation tasks, where the evaluation set includes the test documents corresponding to each evaluation task and the test questions and standard answers corresponding to the test documents; a test module, configured to test the model to be evaluated based on the test documents corresponding to each evaluation task in the evaluation set, and obtain the test results output by the model to be evaluated, where the test results include the intelligent answers output based on the test questions included in the test documents; a second determination module, configured to determine the target score corresponding to each evaluation task based on the difference between the intelligent answers and the standard answers corresponding to each evaluation task; wherein, each evaluation task corresponds to at least two sub-evaluation results from two target objects, and the target score is determined based on at least two sub-evaluation results corresponding to each evaluation task; a third determination module, configured to determine the evaluation result of the model to be evaluated based on the target scores corresponding to each evaluation task; the second determination module includes: a first determination sub-module, configured to determine at least two sub-evaluation results corresponding to the test documents included in each evaluation task based on at least two sub-evaluation results corresponding to the test questions included in each evaluation task; a second determination sub-module, configured to determine at least two sub-evaluation results corresponding to each evaluation task based on at least two sub-evaluation results corresponding to the test documents included in each evaluation task; A third determination sub-module, configured to determine the target score corresponding to each assessment task based on at least two sub-assessment results corresponding to each assessment task; The sub-assessment results include specific score values, and the third determination sub-module is configured to: Determine the target score corresponding to each assessment task based on whether the specific score values included in at least two sub-assessment results corresponding to each assessment task are in the same score range.

15. The apparatus according to claim 14, wherein, Each sub-assessment result is obtained by each target object evaluating the difference between the intelligent answer corresponding to the test question and the standard answer; wherein, there are at least two target objects corresponding to the test question.

16. The apparatus according to claim 14, wherein, The third determination sub-module is configured to: When the specific score values included in at least two sub-assessment results corresponding to the first assessment task are in the same score range and are not equal to 0, calculate the average of the specific score values included in at least two sub-assessment results corresponding to the first assessment task to obtain the target score corresponding to the first assessment task, and the first assessment task is any one of the multiple assessment tasks.

17. The apparatus according to claim 14, wherein, The third determination sub-module is configured to: When the specific score values included in at least two sub-assessment results corresponding to the first assessment task are not in the same score range and are not equal to 0, obtain the target score corresponding to the first assessment task by re-evaluating the first assessment task, and the first assessment task is any one of the multiple assessment tasks.

18. The apparatus according to claim 14, wherein, The third determination sub-module is configured to: When there is a specific score value equal to 0 among the specific score values included in at least two sub-assessment results corresponding to the first assessment task, obtain the target score corresponding to the first assessment task by re-evaluating the first assessment task, and the first assessment task is any one of the multiple assessment tasks.

19. The apparatus according to claim 17 or 18, wherein, The third determination sub-module is configured to: Obtain multiple score data of the test questions included in the first assessment task evaluated from multiple assessment factors respectively, wherein at least two target objects evaluate the multiple score data; Determine the score result corresponding to the test question based on the multiple score data under multiple assessment factors corresponding to the test question; Obtain the target score corresponding to the first assessment task according to the score result corresponding to the test question included in the first assessment task.

20. The apparatus according to claim 14, wherein, The apparatus further includes: An extraction module, configured to extract the test document corresponding to the first assessment task in response to the target score corresponding to the first assessment task being lower than the preset threshold corresponding to the first assessment task, and the first assessment task is any one of the multiple assessment tasks; A fourth determination module, configured to determine an optimization solution for the model to be evaluated based on the test document, the test questions, and the standard answers corresponding to the first evaluation task.

21. The apparatus according to claim 14, wherein, the evaluation result includes a total score value, and the third determination module includes: A fourth determination sub-module, configured to determine the total score value of the model to be evaluated according to the weight corresponding to each evaluation task and the target score corresponding to each evaluation task.

22. The apparatus according to claim 14, wherein, the evaluation result includes an overall level metric value, and the third determination module includes: A fifth determination sub-module, configured to determine the overall level metric value of the model to be evaluated according to a plurality of first evaluation dimensions; wherein, the plurality of first evaluation dimensions include at least two of the following dimensions: accuracy rate, weighted average score, and recall rate.

23. The apparatus according to claim 14, wherein, the evaluation result includes a local level metric value, and the third determination module includes: A sixth determination sub-module, configured to determine the local level metric value of the model to be evaluated according to a plurality of second evaluation dimensions; wherein, the plurality of second evaluation dimensions include at least two of the following dimensions: test effect of the number of test documents, test effect of the language of test documents, and test effect of the types of evaluation tasks.

24. The apparatus according to claim 14, wherein, the first determination module includes: A selection sub-module, configured to select the test documents related to each evaluation task from the external link knowledge base according to the characteristics of each evaluation task; A generation sub-module, configured to generate the test questions and the standard answers corresponding to the test questions for the test documents included in each evaluation task according to the test question examples of each evaluation task; A seventh determination sub-module, configured to determine the evaluation set based on the test documents, the test questions, and the standard answers corresponding to each evaluation task.

25. The apparatus according to claim 24, wherein, the selection sub-module is configured to: Determine the document features related to each evaluation task according to the characteristics of each evaluation task; Select the test documents related to each evaluation task from the external link knowledge base according to the document features related to each evaluation task.

26. The apparatus according to claim 14, wherein, the multiple evaluation tasks include at least two of the following: Evaluation tasks related to the extraction of PICOS of research objects, intervention measures, control measures, outcomes, and research types; Evaluation tasks related to adverse reaction monitoring; Evaluation tasks related to the summary of a single medical document; Evaluation tasks related to the summary of multiple medical documents; Evaluation tasks related to document-level relationship extraction.

27. An electronic device, including: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method according to any one of claims 1-13.

28. A non-transitory computer-readable storage medium storing computer instructions, wherein, the computer instructions are for causing the computer to execute the method according to any one of claims 1-13.

29. A computer program product, comprising a computer program stored on a storage medium, the computer program, when executed by a processor, implementing the method according to any one of claims 1-13.

Citation Information

Patent Citations

  • Question and answer scoring method, question and answer scoring device, electronic equipment and storage medium

    CN116561538A

  • Model evaluation method and device, electronic equipment and storage medium

    CN116737881A