Model evaluation method and device, computer equipment and storage medium

By publishing evaluation tasks including model response methods, test sets and case parameters on the task release platform, and visually displaying them, the problem of difficulty in centralized management of model evaluation data is solved, and intuitive understanding of model effects and low-cost acquisition are achieved.

CN120256891APending Publication Date: 2025-07-04CHONGQING SELIS PHOENIX INTELLIGENT INNOVATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510292500.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-12
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

The model evaluation methods in the prior art are independent and difficult to centrally manage, making it difficult for R&D and testers to intuitively understand the model effects.

Method used

The evaluation task is published on the task release platform, including model reply method, model test set, model testing field and evaluation case parameters. The evaluation personnel conduct the evaluation and obtain the evaluation parameters for visual display.

Benefits of technology

Centralized management of model evaluation parameters is realized, which reduces the time cost of R&D personnel for obtaining data, and allows R&D personnel to intuitively understand the model effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120256891A_ABST
    Figure CN120256891A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of model analysis, in particular to a model evaluation method and device, computer equipment and a storage medium. The method comprises the steps that an evaluation task for a to-be-evaluated model is issued on a task issuing platform, and the evaluation task comprises a model reply mode, a model test set, a model test field and evaluation case parameters of the to-be-evaluated model; enabling the evaluation personnel to receive the evaluation task in the task issuing platform, and evaluating the evaluation task according to the model reply mode, the model test set, the model test field and the evaluation case parameters; obtaining evaluation parameters obtained after evaluation of the evaluation task is finished; and visually displaying the evaluation parameters of the to-be-evaluated model on the task publishing platform. According to the method, research and development personnel can intuitively know the model effect, centralized management of the evaluation parameters of the to-be-evaluated model is realized, the research and development personnel do not need to consume a large amount of time to perform data extraction and data processing, and the acquisition cost of the evaluation parameters is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of model analysis, and particularly to a model evaluation method, device, computer device, and storage medium. Background Art

[0002] With the continuous development and progress of model technology, more and more models are built and put into use. To ensure a good user experience during the application stage of the model, the model is evaluated during the R & D and update stage to obtain the model effect.

[0003] However, there are many existing evaluation methods, and these methods are relatively independent of each other. Therefore, the data generated by the evaluation is difficult to manage centrally, which is not conducive to R & D and testing personnel intuitively understanding the model effect. Summary of the Invention

[0004] Based on this, in view of the above technical problems, it is necessary to provide a model evaluation method, device, computer device, and storage medium that can facilitate R & D and testing personnel to intuitively understand the model effect.

[0005] In a first aspect, this application provides a model evaluation method. The method includes:

[0006] Publish an evaluation task for the model to be evaluated on a task publishing platform. The evaluation task includes the model reply method, model test set, model test field, and evaluation case parameters of the model to be evaluated, so that the evaluator can receive the evaluation task on the task publishing platform and evaluate the evaluation task according to the model reply method, model test set, model test field, and evaluation case parameters.

[0007] Obtain the evaluation parameters obtained after the evaluation of the evaluation task is completed.

[0008] Visually display the evaluation parameters of the model to be evaluated on the task publishing platform.

[0009] In one embodiment, the visually displaying the evaluation parameters of the model to be evaluated on the task publishing platform includes:

[0010] Determine the model distribution in the visual area on the task publishing platform according to the number of models of the model to be evaluated.

[0011] Visually display the evaluation results of the model to be evaluated in the visual area according to the model distribution.

[0012] In one embodiment, the evaluation parameters include the evaluation result parameters of the evaluation task and the evaluation score parameters for the evaluation result parameters.

[0013] The evaluation scoring parameters include the case evaluation scoring parameters of the evaluation cases; the field evaluation scoring parameters of the model test fields; and the model evaluation scoring parameters of each model to be evaluated in the case where there are models to be evaluated.

[0014] In one embodiment, the case evaluation scoring parameters include the case evaluation scores of different evaluation cases and the distribution of the case evaluation scores with different values; the field evaluation scoring parameters include the field evaluation scores of different model test fields and at least one of the total score, average score, variance, and median corresponding to each field evaluation score; the model evaluation scoring parameters include the model evaluation scores of different models to be evaluated and at least one of the total score, average score, variance, and median corresponding to each model evaluation score.

[0015] In one embodiment, releasing an evaluation task for a model to be evaluated on the task publishing platform includes:

[0016] Releasing an evaluation task for a model to be evaluated on the task publishing platform and setting an end time for the evaluation task on the task publishing platform; the end time is used to represent the time when the evaluation task cannot be received on the task publishing platform.

[0017] In one embodiment, the method further includes:

[0018] If the current time exceeds the end time, or the task status of the evaluation task is adjusted from the receivable status to the end status, then the evaluator is prohibited from receiving the evaluation task on the task publishing platform.

[0019] In one embodiment, visualizing the evaluation parameters of the model to be evaluated on the task publishing platform includes:

[0020] Obtaining the historical parameters of the historical model;

[0021] Comparing the historical parameters with the evaluation parameters to obtain a comparison result;

[0022] Visualizing the evaluation parameters, the historical parameters, and the comparison result on the task publishing platform.

[0023] In a second aspect, the present application further provides a model evaluation device. The device includes:

[0024] A publishing module, configured to publish an evaluation task for a model to be evaluated on a task publishing platform, where the evaluation task includes a model response method, a model test set, a model test field, and evaluation case parameters of the model to be evaluated, so that evaluators can receive the evaluation task on the task publishing platform and evaluate the evaluation task according to the model response method, the model test set, the model test field, and the evaluation case parameters;

[0025] An acquisition module, configured to acquire evaluation parameters obtained after the evaluation of the evaluation task;

[0026] A display module, configured to visually display the evaluation parameters of the model to be evaluated on the task publishing platform.

[0027] In a third aspect, the present application further provides a computer device. The computer device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:

[0028] Publish an evaluation task for a model to be evaluated on a task publishing platform, where the evaluation task includes a model response method, a model test set, a model test field, and evaluation case parameters of the model to be evaluated, so that evaluators can receive the evaluation task on the task publishing platform and evaluate the evaluation task according to the model response method, the model test set, the model test field, and the evaluation case parameters;

[0029] Acquire the evaluation parameters obtained after the evaluation of the evaluation task;

[0030] Visually display the evaluation parameters of the model to be evaluated on the task publishing platform.

[0031] In a fourth aspect, the present application further provides a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the following steps are implemented:

[0032] Publish an evaluation task for a model to be evaluated on a task publishing platform, where the evaluation task includes a model response method, a model test set, a model test field, and evaluation case parameters of the model to be evaluated, so that evaluators can receive the evaluation task on the task publishing platform and evaluate the evaluation task according to the model response method, the model test set, the model test field, and the evaluation case parameters;

[0033] Acquire the evaluation parameters obtained after the evaluation of the evaluation task;

[0034] Visualize the evaluation parameters of the to-be-evaluated model on the task publishing platform.

[0035] In a fifth aspect, the present application also provides a computer program product. The computer program product includes a computer program, and when the computer program is executed by a processor, the following steps are implemented:

[0036] Publish an evaluation task for the to-be-evaluated model on the task publishing platform, where the evaluation task includes the model response mode, model test set, model test field, and evaluation case parameters of the to-be-evaluated model; so that the evaluator can receive the evaluation task on the task publishing platform and evaluate the evaluation task according to the model response mode, the model test set, the model test field, and the evaluation case parameters;

[0037] Obtain the evaluation parameters obtained after the evaluation of the evaluation task;

[0038] Visualize the evaluation parameters of the to-be-evaluated model on the task publishing platform.

[0039] The above model evaluation method, device, computer device, and storage medium publish an evaluation task for the to-be-evaluated model on the task publishing platform, so that the evaluator can receive the evaluation task on the task publishing platform and evaluate the evaluation task according to the model response mode, model test set, model test field, and evaluation case parameters; furthermore, obtain the evaluation parameters obtained after the evaluation of the evaluation task to realize visual display of the evaluation parameters of the to-be-evaluated model on the task publishing platform. According to the above content, it can be known that in the process of evaluating the model in the present application, an evaluation task for the to-be-evaluated model will be pre-constructed, where the evaluation task includes the model response mode, model test set, model test field, and evaluation case parameters of the to-be-evaluated model to ensure that the evaluator can evaluate the model according to the actual evaluation needs of the to-be-evaluated model, and by obtaining the evaluation parameters and visualizing the evaluation parameters of the to-be-evaluated model on the task publishing platform, the R & D personnel can intuitively know the model effect, realizing centralized management of the evaluation parameters of the to-be-evaluated model, without the R & D personnel consuming a large amount of time for data extraction and data processing, reducing the acquisition cost of the evaluation parameters. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] Figure 1 It is an application environment diagram of a model evaluation method provided by an embodiment of the present application;

[0041] Figure 2 It is a flowchart of the first model evaluation method provided by an embodiment of the present application;

[0042] Figure 3Schematic flowchart of the second model evaluation method provided by the embodiments of the present application;

[0043] Figure 4 Schematic diagram of the first visualization area in a task publishing platform provided by the embodiments of the present application;

[0044] Figure 5 Schematic diagram of the second visualization area in a task publishing platform provided by the embodiments of the present application;

[0045] Figure 6 Schematic diagram of the third visualization area in a task publishing platform provided by the embodiments of the present application;

[0046] Figure 7 Schematic diagram of the fourth visualization area in a task publishing platform provided by the embodiments of the present application;

[0047] Figure 8 Schematic diagram of the fifth visualization area in a task publishing platform provided by the embodiments of the present application;

[0048] Figure 9 Schematic diagram of the sixth visualization area in a task publishing platform provided by the embodiments of the present application;

[0049] Figure 10 Schematic diagram of the framework structure of model evaluation provided by the embodiments of the present application;

[0050] Figure 11 Schematic flowchart of the processing flow of model evaluation provided by the embodiments of the present application;

[0051] Figure 12 Block diagram of the structure of a model evaluation device provided by the embodiments of the present application;

[0052] Figure 13 Internal structure diagram of a computer device in an embodiment. Detailed implementation manners

[0053] In order to make the objectives, technical solutions and advantages of the present application more clear and understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0054] The model evaluation method provided by the embodiments of the present application can be applied to, for example Figure 1In the application environment shown. Among them, the terminal 102 communicates with the server 104 through the network. The data storage system can store the data that the server 104 needs to process. The data storage system can be integrated on the server 104, or can be placed on the cloud or other network servers. By publishing an evaluation task for the model to be evaluated on the task publishing platform, so that the evaluators can receive the evaluation task on the task publishing platform, and evaluate the evaluation task according to the model response method, model test set, model test field and evaluation case parameters; furthermore, obtain the evaluation parameters obtained after the evaluation of the evaluation task is completed, so as to realize the visual display of the evaluation parameters of the model to be evaluated on the task publishing platform. Among them, the terminal 102 can be but is not limited to various personal computers, laptop computers, smart phones, tablet computers, Internet of Things devices and portable wearable devices. The Internet of Things devices can be smart speakers, smart TVs, smart air conditioners, smart vehicle-mounted devices, etc. The portable wearable devices can be smart watches, smart bracelets, head-mounted devices, etc. The server 104 can be implemented by an independent server or a server cluster composed of multiple servers.

[0055] In one embodiment, as Figure 2 shown, a model evaluation method is provided. Taking the server 104 in Figure 1 as an example for illustration, the method includes the following steps:

[0056] S201, publish an evaluation task for the model to be evaluated on the task publishing platform.

[0057] Among them, the evaluation task includes the model response method, model test set, model test field and evaluation case parameters of the model to be evaluated; so that the evaluators can receive the evaluation task on the task publishing platform, and evaluate the evaluation task according to the model response method, model test set, model test field and evaluation case parameters.

[0058] It should be noted that in order to ensure that the model to be evaluated can fully meet the evaluation needs of the R & D personnel in the subsequent evaluation process, therefore, it is necessary to pre-determine the model response method, model test set, model test field and evaluation case parameters of the model to be evaluated, and then, according to the model response method, model test set, model test field and evaluation case parameters of the model to be evaluated, construct an evaluation task for the model to be evaluated.

[0059] Specifically, the model reply method can include a streaming reply method and a non-streaming reply method. The model test set can include the test set of the corresponding database of the task publishing platform and a custom data set. Among them, the custom data set needs to be uploaded in a fixed format before it can be used, and a template format of the custom data set needs to be provided. The model test fields include a dedicated field and a general field. If the model test field is the general field, the model test set can be randomly selected from the data sets in all fields; if the model test field is the dedicated field, the model test set can be randomly selected from the data sets in a specific field. The evaluation case parameters include the number of cases in different fields and the total number of cases. Among them, the total number of cases and the number of cases in different fields can be calculated from each other, that is: when setting the number of cases in different fields, the total number of cases is automatically updated; when modifying the total number of cases, the number of cases in different fields is automatically updated. At this time, it should be noted that the number of cases in different fields is at least 1, and the total number of cases is evenly distributed.

[0060] Furthermore, when creating an evaluation task for the model to be evaluated, the execution times of each case can be configured. After configuration, the total number of cases called is N*M; N is the number of cases in all fields, and M is the execution times of each case.

[0061] It should be noted that the corresponding database of the task publishing platform can include model parameter configuration and test set configuration; among them, the model parameter configuration includes fields for URL (Uniform Resource Locator), header (the header part in the data stream), body (the main part of the code block or data structure), and the interface return. The test set configuration includes field configuration, sub-field configuration, question configuration, and reference reply configuration.

[0062] In an embodiment of the present application, an evaluator can receive a published and unfinished evaluation task on the task publishing platform. After receiving the task, the evaluation can be carried out at any time before the task ends.

[0063] Among them, the evaluation methods adopted by the evaluator when evaluating the evaluation task can include but are not limited to: centralized evaluation and individual evaluation. Centralized evaluation means waiting until all test set cases of the model to be evaluated are executed, and then conducting the evaluation; individual evaluation means randomly selecting the test set cases that need to be executed, and after the corresponding model to be evaluated for the current case is executed, conducting the evaluation, and after the evaluation is completed, randomly selecting the next test set case that needs to be executed until all test set cases are executed.

[0064] When evaluating a test task, the evaluator can display the evaluation progress in real time on the task release platform; during the process of the evaluator evaluating the model to be evaluated through the centralized evaluation method or the individual evaluation method, the evaluation process for the model to be evaluated can be displayed on the task release platform in the form of a percentage.

[0065] Furthermore, during the process of the evaluator evaluating the test task, the basic information of the test task can be displayed. Among them, the basic information can include but is not limited to: task name, evaluation model, evaluation field, evaluation rules, evaluation progress, etc.; the evaluation cases, the reference answers of the evaluation cases, and the similarity between the reference answers and the model answers can also be displayed.

[0066] To prevent the subjective awareness of the evaluator from affecting the fairness of the evaluation of the model to be evaluated, the evaluation order of the model to be evaluated can be randomly disrupted, and the model name of the model to be evaluated can be hidden, so that the evaluator does not know which model the answer they are evaluating corresponds to.

[0067] In an embodiment of the present application, during the process of the evaluator evaluating the test task, the corresponding test set can be selected according to the model test set and the model test field corresponding to the test task; that is, after the field evaluation task by the evaluator, the domain range corresponding to the model test field corresponding to the test task and the test set source corresponding to the model test set are both used for evaluation; and during the process of multiple evaluators conducting evaluations, the evaluation questions are random and non-repetitive, that is, the test sets of different evaluators are not exactly the same, but the test questions of the test sets are non-repetitive, unless it is a case with an execution count greater than one.

[0068] After the evaluation is completed, the evaluator can only see the evaluation results of their current evaluation, that is, an evaluation of the model to be evaluated by themselves.

[0069] S202, obtain the evaluation parameters obtained after the evaluation of the test task ends.

[0070] It should be noted that when it is necessary to obtain the evaluation parameters obtained after the evaluation of the test task ends, the test task after the evaluation ends can be initially evaluated to obtain the initial parameters; then, the initial parameters are adjusted by integrating the historical evaluation experience of the evaluators to obtain the adjusted evaluation parameters.

[0071] In an embodiment of the present application, an initial scoring model can be pre-trained so that after the model to be evaluated completes all evaluation cases, an initial evaluation of the model to be evaluated is performed based on the execution results of the model to be evaluated and the standard responses of each evaluation case to obtain initial parameters; the evaluator adjusts the initial parameters adaptively according to the actual situation of the model to be evaluated so that the adjusted initial parameters better conform to the actual situation of the model to be evaluated, and the evaluation parameters obtained after the evaluation task is completed are obtained.

[0072] S203. Visualize the evaluation parameters of the model to be evaluated on the task publishing platform.

[0073] It should be noted that the visualization area can be determined in advance on the task publishing platform; furthermore, the evaluation parameters of the model to be evaluated are visualized in the visualization area to achieve visualization display in a specific area of the task publishing platform, which is convenient for R & D personnel to intuitively understand the model effect.

[0074] Furthermore, when visualizing the evaluation parameters of the model to be evaluated on the task publishing platform, if there are multiple models to be evaluated, the display order of each model to be evaluated can be determined according to the evaluation end time of each model to be evaluated; furthermore, according to the display order, the evaluation results of each model to be evaluated are visualized in the visualization area.

[0075] In the above model evaluation method, an evaluation task for the model to be evaluated is published on the task publishing platform so that the evaluator can receive the evaluation task on the task publishing platform and evaluate the evaluation task according to the model response method, model test set, model test field, and evaluation case parameters; furthermore, the evaluation parameters obtained after the evaluation task is completed are obtained to achieve visualization display of the evaluation parameters of the model to be evaluated on the task publishing platform. According to the above content, it can be seen that in the process of evaluating a model in the present application, an evaluation task for the model to be evaluated will be pre-constructed, where the evaluation task includes the model response method, model test set, model test field, and evaluation case parameters of the model to be evaluated to ensure that the evaluator can evaluate the model according to the actual evaluation requirements of the model to be evaluated. And by obtaining the evaluation parameters and visualizing the evaluation parameters of the model to be evaluated on the task publishing platform, R & D personnel can intuitively understand the model effect, realizing centralized management of the evaluation parameters of the model to be evaluated, without the need for R & D personnel to consume a large amount of time for data extraction and data processing, and reducing the acquisition cost of the evaluation parameters.

[0076] In an embodiment, as Figure 3 shown, when it is necessary to visualize the evaluation parameters of the model to be evaluated on the task publishing platform, it specifically includes the following content:

[0077] S301. Determine the model distribution in the visualization area of the task publishing platform according to the number of models of the model to be evaluated.

[0078] Among them, the model distribution is used to represent the maximum number of models displayed horizontally and / or the maximum number of models displayed vertically in the visualization area.

[0079] In an embodiment of the present application, if the maximum number of models displayed horizontally recorded in the model distribution is three, when visualizing the evaluation results of the model to be evaluated in the visualization area, the visualization area is as Figure 4 shown. Among them, the evaluation method used by the evaluator for the evaluation task is centralized evaluation, and the visualization area may also include the specific content of the evaluation rules, that is, case questions, and reference answers to the case questions.

[0080] It should be noted that when it is necessary to determine the model distribution in the visualization area of the task publishing platform, the mapping relationship between different candidate numbers and different candidate distribution situations can be specified in advance; then, after determining the number of models, determine the reference data with the same number as the number of models from the candidate numbers; and use the candidate distribution situation corresponding to the reference number in the mapping relationship as the model distribution in the visualization area of the task publishing platform.

[0081] Among them, the mapping relationship between different candidate numbers and different candidate distribution situations can be set or adjusted according to the actual situation, and the specific content of the mapping relationship is not limited here.

[0082] S302. Visualize the evaluation results of the model to be evaluated in the visualization area according to the model distribution.

[0083] Among them, during the process of visualizing the evaluation results of the model to be evaluated, the display form can also be limited. Specifically, the display form can include but is not limited to: table form, bar chart form, etc.

[0084] It should be noted that the evaluation parameters include the evaluation result parameters of the evaluation task and the evaluation scoring parameters for the evaluation result parameters;

[0085] The evaluation scoring parameters include the case evaluation scoring parameters of the evaluation case; the field evaluation scoring parameters of the model test field; the model evaluation scoring parameters of each model to be evaluated in the case where there are models to be evaluated.

[0086] Further explanation: The case evaluation scoring parameters include the case evaluation scores of different evaluation cases and the distribution of case evaluation scores with different values; the domain evaluation scoring parameters include the domain evaluation scores of different model test domains and at least one of the total score, average score, variance, and median corresponding to each domain evaluation score; the model evaluation scoring parameters include the model evaluation scores of different models to be evaluated and at least one of the total score, average score, variance, and median corresponding to each model evaluation score.

[0087] In one embodiment, when visually displaying the evaluation parameters of the model to be evaluated, it may also include the task name, evaluation model, model test domain, evaluation case parameters, case execution status, that is, the number of completed tasks and the number of uncompleted tasks, task status, task start time, and task end time. When visually displaying the evaluation parameters of the model to be evaluated in tabular form, as Figure 5 shown. When visually displaying the evaluation parameters of the model to be evaluated in the form of a bar chart, as Figure 6 shown; when displaying the total score, average score, variance, and median corresponding to each model evaluation score in tabular form, as Figure 7 shown, when displaying the total score, average score, variance, and median corresponding to each model evaluation score in the form of a bar chart, as Figure 8 shown.

[0088] In one embodiment of the present application, when visually displaying the evaluation parameters of the model to be evaluated on the task publishing platform, the following content may be included: obtaining the historical parameters of the historical model; comparing the historical parameters with the evaluation parameters to obtain a comparison result; visually displaying the evaluation parameters, historical parameters, and comparison result on the task publishing platform.

[0089] Further explanation, as Figure 9 shown, when visually displaying the evaluation parameters, historical parameters, and comparison result on the task publishing platform, a bar chart form may be used to arrange the comparison results in order; among them, the arrangement order can be set or adjusted according to the actual situation, for example, it can be arranged according to the task execution order, etc.

[0090] The above model evaluation method determines the model distribution in the visual area of the task publishing platform, so as to visually display the evaluation results of the model to be evaluated in the visual area according to the model distribution, enabling R & D personnel to intuitively know the model effect, realizing centralized management of the evaluation parameters of the model to be evaluated, without R & D personnel consuming a large amount of time for data extraction and data processing, and reducing the acquisition cost of evaluation parameters.

[0091] In one embodiment, when issuing an evaluation task for a model to be evaluated on a task publishing platform, the following content may also be included: issuing an evaluation task for the model to be evaluated on the task publishing platform, and setting a task end time for the evaluation task on the task publishing platform; the task end time is used to represent the time when the evaluation task cannot be claimed on the task publishing platform.

[0092] Furthermore, if the current time exceeds the task end time, or the task status of the evaluation task is adjusted from the claimable status to the ended status, then the evaluators are prohibited from claiming the evaluation task on the task publishing platform.

[0093] The above model evaluation method realizes flexible control of the evaluation task by stipulating the task end time, so as to ensure that the claim process of the evaluation task can be standardized according to the actual situation. Moreover, by setting the task status of the evaluation task with the task end time, it is prevented that the evaluation task is claimed when there is no need to execute the evaluation task, resulting in waste of resources.

[0094] In one embodiment, when it is necessary to evaluate the model to be evaluated, the model can be pre-configured in the database, which specifically includes the following content: a) Name: mode1-gpt; b) URL: http: / / 127.0.0..1 / xxxxx; c) Call parameters: {prompt: xxxxx}; d) Call header: {"Content-Type": "application / json"……}; e) Interface return field path: result.data.content, result.data[-1].content (depending on the return of different model interfaces). If it is a streaming interface, the processing method of the streaming interface is used. And different test sets are configured in the database, which need to include: field, question, reference answer, etc. Select test models: demo1-gpt, demo2-gpt; When selecting the built-in test set, the data mentioned above is used for evaluation; When customizing the test set, the data uploaded by the publisher is used for evaluation. Select the model test field: general field and specific field; When selecting the general field, directly set the number of cases and the number of executions for each case; When selecting the specific field, a specific field needs to be selected again, such as: automobile, airplane, etc. After selection, the number of cases needs to be set for each field, and the setting will be linked to the total number of cases. For example, if the number of cases in the automobile field is 2 and the number of cases in the airplane field is 3, then the total number of cases will be automatically displayed as 5. When not setting the number of cases for each field and directly setting the total number of cases, each field will score the number of cases. For example, when the total number of cases is set to 10, the number of cases allocated to the automobile field and the airplane field is 5 each; When the total number of cases is set to 9, 4 cases will be evenly allocated to the automobile field first, 4 cases to the airplane field, and the remaining 1 case will be allocated to the automobile field in order. Set the task end time, such as: 2024-12-05 23:59:59; After setting, create the task, and at this time, the task can be edited and modified again; After the task is created, the task needs to be manually published. Only the published task can be obtained by the evaluator.

[0095] The evaluators receive evaluation tasks on the task publishing platform, which may include the following: When the evaluators receive tasks, if there are three evaluation tasks on the task publishing platform at this time, namely: taskA: not published; taskB: published, not reaching the task end time; taskC: ended, exceeding the task end time; therefore, the evaluators can only select taskB for evaluation. When the task is not ended, the receiver can evaluate taskB. During the evaluation, the cases of different evaluators are not completely the same, that is, a set number of data are randomly selected from the test set. During the evaluation, it can be interrupted at any time and the evaluator can leave the evaluation. When evaluating again later, the previous evaluation progress can be continued. After the evaluation is completed, the evaluator can view his own evaluation results, as follows: (1) In the form of a list, by score or number of times, viewed by field, model viewing, etc.; (2) In the form of a bar chart, by score or number of times, viewed by field, model viewing, etc.

[0096] Among them, the framework of the model evaluation is as Figure 10 shown; the processing flow of the model evaluation is as Figure 11 shown.

[0097] It should be understood that although the steps in the flowcharts involved in the above embodiments are shown in sequence according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear description in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or alternately with at least a part of other steps or steps in other steps.

[0098] Based on the same inventive concept, the embodiment of the present application also provides a model evaluation device for implementing the above-mentioned model evaluation method. The implementation solutions provided by this device to solve problems are similar to the implementation solutions described in the above method. Therefore, the specific limitations in one or more embodiments of the model evaluation device provided below can refer to the limitations on the model evaluation method in the above text, and will not be repeated here.

[0099] In one embodiment, as Figure 12 shown, a model evaluation device is provided, including: a publishing module 10, an acquisition module 20, and a display module 30, where:

[0100] A publishing module 10 is used to publish an evaluation task for a model to be evaluated on a task publishing platform. The evaluation task includes the model response method, model test set, model test field, and evaluation case parameters of the model to be evaluated, so that evaluators can receive the evaluation task on the task publishing platform and evaluate the evaluation task according to the model response method, model test set, model test field, and evaluation case parameters.

[0101] An acquisition module 20 is used to acquire the evaluation parameters obtained after the evaluation of the evaluation task.

[0102] A display module 30 is used to visually display the evaluation parameters of the model to be evaluated on the task publishing platform.

[0103] In one embodiment, according to the number of models of the model to be evaluated, the model distribution in the visual area on the task publishing platform is determined.

[0104] According to the model distribution, the evaluation results of the model to be evaluated are visually displayed in the visual area.

[0105] In one embodiment, the evaluation scoring parameters include the case evaluation scoring parameters of the evaluation case, the field evaluation scoring parameters of the model test field, and the model evaluation scoring parameters of each model to be evaluated in the case where there are models to be evaluated.

[0106] In one embodiment, the case evaluation scoring parameters include the case evaluation scores of different evaluation cases and the distribution of case evaluation scores with different values. The field evaluation scoring parameters include the field evaluation scores of different model test fields and at least one of the total score, average score, variance, and median corresponding to each field evaluation score. The model evaluation scoring parameters include the model evaluation scores of different models to be evaluated and at least one of the total score, average score, variance, and median corresponding to each model evaluation score.

[0107] In one embodiment, an evaluation task for a model to be evaluated is published on the task publishing platform, and a task end time for the evaluation task is set on the task publishing platform. The task end time is used to represent the time when the evaluation task cannot be received on the task publishing platform.

[0108] In one embodiment, if the current time exceeds the task end time, or the task status of the evaluation task is adjusted from the receivable status to the end status, evaluators are prohibited from receiving the evaluation task on the task publishing platform.

[0109] In one embodiment, historical parameters of a historical model are acquired.

[0110] The historical parameters are compared with the evaluation parameters to obtain a comparison result.

[0111] Visualize the evaluation parameters, historical parameters, and comparison results on the task publishing platform.

[0112] In the above model evaluation method, an evaluation task for the model to be evaluated is published on the task publishing platform, so that the evaluator can receive the evaluation task on the task publishing platform and evaluate the evaluation task according to the model response method, model test set, model test field, and evaluation case parameters. Furthermore, the evaluation parameters obtained after the evaluation of the evaluation task are obtained to realize the visualization display of the evaluation parameters of the model to be evaluated on the task publishing platform. According to the above content, in the process of evaluating the model, this application will pre-construct an evaluation task for the model to be evaluated, where the evaluation task includes the model response method, model test set, model test field, and evaluation case parameters of the model to be evaluated, so as to ensure that the evaluator can evaluate the model according to the actual evaluation needs of the model to be evaluated. And by obtaining the evaluation parameters and visualizing the evaluation parameters of the model to be evaluated on the task publishing platform, the R & D personnel can intuitively know the model effect, realizing the centralized management of the evaluation parameters of the model to be evaluated, without the need for R & D personnel to consume a lot of time for data extraction and data processing, reducing the acquisition cost of the evaluation parameters.

[0113] Each module in the above model evaluation device can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in the processor of the computer device in hardware form or independent of it, or stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to the above modules.

[0114] In one embodiment, a computer device is provided. The computer device can be a terminal, and its internal structure diagram can be as Figure 13As shown in the figure. The computer device includes a processor, a memory, an input / output interface, a communication interface, a display unit, and an input device. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface, the display unit, and the input device are connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals in a wired or wireless manner, and the wireless manner can be implemented through WIFI, a mobile cellular network, NFC (Near Field Communication), or other technologies. The computer program, when executed by the processor, implements a model evaluation method. The display unit of the computer device is used to form a visually visible picture, which can be a display screen, a projection device, or a virtual reality imaging device. The display screen can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer covering the display screen, or a button, a trackball, or a touchpad provided on the computer device housing, or an external keyboard, touchpad, or mouse, etc.

[0115] Those skilled in the art can understand that Figure 13 the structure shown in the figure is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0116] In one embodiment, a computer device is provided, including a memory and a processor. A computer program is stored in the memory. When the processor executes the computer program, the following steps are implemented:

[0117] Publish an evaluation task for the model to be evaluated on a task publishing platform. Among them, the evaluation task includes the model reply method, the model test set, the model test field, and the evaluation case parameters of the model to be evaluated; so that the evaluator can receive the evaluation task on the task publishing platform and evaluate the evaluation task according to the model reply method, the model test set, the model test field, and the evaluation case parameters;

[0118] Obtain the evaluation parameters obtained after the evaluation of the evaluation task;

[0119] Visually display the evaluation parameters of the model to be evaluated on the task publishing platform.

[0120] In one embodiment, when the processor executes the computer program, the following steps are further implemented:

[0121] Determine the model distribution in the visualization area of the task publishing platform according to the number of models of the model to be evaluated;

[0122] Visualize the evaluation results of the model to be evaluated in the visualization area according to the model distribution.

[0123] In one embodiment, when the processor executes the computer program, the following steps are further implemented:

[0124] The evaluation scoring parameters include the case evaluation scoring parameters of the evaluation cases; the domain evaluation scoring parameters of the model test domain; and the model evaluation scoring parameters of each model to be evaluated in the case where there are models to be evaluated.

[0125] In one embodiment, when the processor executes the computer program, the following steps are further implemented:

[0126] The case evaluation scoring parameters include the case evaluation scores of different evaluation cases and the distribution of the case evaluation scores with different values; the domain evaluation scoring parameters include the domain evaluation scores of different model test domains and at least one of the total score, average score, variance, and median corresponding to each domain evaluation score; the model evaluation scoring parameters include the model evaluation scores of different models to be evaluated and at least one of the total score, average score, variance, and median corresponding to each model evaluation score.

[0127] In one embodiment, when the processor executes the computer program, the following steps are further implemented:

[0128] Publish an evaluation task for the model to be evaluated on the task publishing platform and set an end time for the evaluation task on the task publishing platform; the end time is used to represent the time when the evaluation task cannot be retrieved on the task publishing platform.

[0129] In one embodiment, when the processor executes the computer program, the following steps are further implemented:

[0130] If the current time exceeds the end time of the task, or the task status of the evaluation task is adjusted from the retrievable state to the end state, then it is prohibited for the evaluators to retrieve the evaluation task on the task publishing platform.

[0131] In one embodiment, when the processor executes the computer program, the following steps are further implemented:

[0132] Obtain the historical parameters of the historical model;

[0133] Compare the historical parameters with the evaluation parameters to obtain a comparison result;

[0134] Visualize the evaluation parameters, historical parameters, and comparison results on the task publishing platform.

[0135] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:

[0136] Publish an evaluation task for the model to be evaluated on the task publishing platform, where the evaluation task includes the model response method, model test set, model test field, and evaluation case parameters of the model to be evaluated; so that the evaluator can receive the evaluation task on the task publishing platform and evaluate the evaluation task according to the model response method, model test set, model test field, and evaluation case parameters.

[0137] Obtain the evaluation parameters obtained after the evaluation of the evaluation task is completed.

[0138] Visualize the evaluation parameters of the model to be evaluated on the task publishing platform.

[0139] In one embodiment, when the computer program is executed by a processor, the following steps are also implemented:

[0140] Determine the model distribution in the visualization area on the task publishing platform according to the number of models of the model to be evaluated.

[0141] Visualize the evaluation results of the model to be evaluated in the visualization area according to the model distribution.

[0142] In one embodiment, when the computer program is executed by a processor, the following steps are also implemented:

[0143] The evaluation scoring parameters include the case evaluation scoring parameters of the evaluation cases; the field evaluation scoring parameters of the model test fields; the model evaluation scoring parameters of each model to be evaluated in the case where there are models to be evaluated.

[0144] In one embodiment, when the computer program is executed by a processor, the following steps are also implemented:

[0145] The case evaluation scoring parameters include the case evaluation scores of different evaluation cases and the distribution of case evaluation scores with different values; the field evaluation scoring parameters include the field evaluation scores of different model test fields and at least one of the total score, average score, variance, and median corresponding to each field evaluation score; the model evaluation scoring parameters include the model evaluation scores of different models to be evaluated and at least one of the total score, average score, variance, and median corresponding to each model evaluation score.

[0146] In one embodiment, when the computer program is executed by a processor, the following steps are also implemented:

[0147] Publish an evaluation task for the model to be evaluated on the task publishing platform, and set an end time for the evaluation task on the task publishing platform; the end time is used to represent the time when the evaluation task cannot be claimed on the task publishing platform.

[0148] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented:

[0149] If the current time exceeds the end time of the task, or the task status of the evaluation task is adjusted from the claimable state to the ended state, then the evaluator is prohibited from claiming the evaluation task on the task publishing platform.

[0150] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented:

[0151] Obtain the historical parameters of the historical model;

[0152] Compare the historical parameters with the evaluation parameters to obtain a comparison result;

[0153] Visually display the evaluation parameters, historical parameters, and comparison result on the task publishing platform.

[0154] In one embodiment, a computer program product is provided, including a computer program, which when executed by a processor, implements the following steps:

[0155] Publish an evaluation task for the model to be evaluated on the task publishing platform, where the evaluation task includes the model reply method, model test set, model test field, and evaluation case parameters of the model to be evaluated; so that the evaluator can claim the evaluation task on the task publishing platform and evaluate the evaluation task according to the model reply method, model test set, model test field, and evaluation case parameters;

[0156] Obtain the evaluation parameters obtained after the evaluation of the evaluation task;

[0157] Visually display the evaluation parameters of the model to be evaluated on the task publishing platform.

[0158] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented:

[0159] Determine the model distribution in the visualization area on the task publishing platform according to the number of models of the model to be evaluated;

[0160] Visually display the evaluation results of the model to be evaluated in the visualization area according to the model distribution.

[0161] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented:

[0162] The evaluation scoring parameters include the case evaluation scoring parameters of the evaluation cases; the domain evaluation scoring parameters of the model test fields; and the model evaluation scoring parameters of each model to be evaluated in the case where there are models to be evaluated.

[0163] In one embodiment, when the computer program is executed by the processor, the following steps are further implemented:

[0164] The case evaluation scoring parameters include the case evaluation scores of different evaluation cases and the distribution of the case evaluation scores with different values; the domain evaluation scoring parameters include the domain evaluation scores of different model test fields and at least one of the total score, average score, variance, and median corresponding to each domain evaluation score; the model evaluation scoring parameters include the model evaluation scores of different models to be evaluated and at least one of the total score, average score, variance, and median corresponding to each model evaluation score.

[0165] In one embodiment, when the computer program is executed by the processor, the following steps are further implemented:

[0166] Publish an evaluation task for the model to be evaluated on the task publishing platform and set an end time for the evaluation task on the task publishing platform; the end time is used to represent the time when the evaluation task cannot be received on the task publishing platform.

[0167] In one embodiment, when the computer program is executed by the processor, the following steps are further implemented:

[0168] If the current time exceeds the end time of the task, or the task status of the evaluation task is adjusted from the receivable status to the end status, then it is prohibited for the evaluators to receive the evaluation task on the task publishing platform.

[0169] In one embodiment, when the computer program is executed by the processor, the following steps are further implemented:

[0170] Obtain the historical parameters of the historical model;

[0171] Compare the historical parameters with the evaluation parameters to obtain a comparison result;

[0172] Visually display the evaluation parameters, historical parameters, and comparison results on the task publishing platform.

[0173] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the users or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions.

[0174] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in the present application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in the present application can be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, data processing logics based on quantum computing, etc., without limitation.

[0175] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.

[0176] The above embodiments only represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the appended claims.

Claims

1. A model evaluation method, characterized in that, The method includes: Posting an evaluation task for the model to be evaluated on a task publishing platform, where the evaluation task includes the model response method, model test set, model test field, and evaluation case parameters of the model to be evaluated; so that evaluators can receive the evaluation task on the task publishing platform and evaluate the evaluation task according to the model response method, the model test set, the model test field, and the evaluation case parameters. Obtaining the evaluation parameters obtained after the evaluation of the evaluation task. Visualizing the evaluation parameters of the model to be evaluated on the task publishing platform.

2. The method according to claim 1, characterized in that, The visualizing the evaluation parameters of the model to be evaluated on the task publishing platform includes: Determining the model distribution in the visualization area on the task publishing platform according to the number of models of the model to be evaluated. Visualizing the evaluation results of the model to be evaluated in the visualization area according to the model distribution.

3. The method according to claim 2, wherein The evaluation parameters include the evaluation result parameters of the evaluation task and the evaluation scoring parameters for the evaluation result parameters. The evaluation scoring parameters include the case evaluation scoring parameters of the evaluation case. The field evaluation scoring parameters of the model test field. The model evaluation scoring parameters of each model to be evaluated in the presence of models to be evaluated.

4. The method according to claim 3, wherein The case evaluation scoring parameters include the case evaluation scores of different evaluation cases and the distribution of case evaluation scores with different values; the field evaluation scoring parameters include the field evaluation scores of different model test fields and at least one of the total score, average score, variance, and median corresponding to each field evaluation score; the model evaluation scoring parameters include the model evaluation scores of different models to be evaluated and at least one of the total score, average score, variance, and median corresponding to each model evaluation score.

5. The method according to claim 1, wherein The posting an evaluation task for the model to be evaluated on the task publishing platform includes: Posting an evaluation task for the model to be evaluated on the task publishing platform and setting an end time for the evaluation task on the task publishing platform; the end time is used to represent the time when the evaluation task cannot be received on the task publishing platform.

6. The method according to claim 5, characterized in that, The method further includes: If the current time exceeds the end time, or the task status of the evaluation task is adjusted from the receivable status to the end status, then the evaluator is prohibited from receiving the evaluation task on the task publishing platform.

7. The method according to claim 1, wherein The visualizing the evaluation parameters of the model to be evaluated on the task publishing platform includes: Obtaining the historical parameters of the historical model. Comparing the historical parameters with the evaluation parameters to obtain a comparison result. Visualizing the evaluation parameters, the historical parameters, and the comparison result on the task publishing platform.

8. A model evaluation device, characterized in that The device includes: A publishing module, configured to publish an evaluation task for a model to be evaluated on a task publishing platform, wherein the evaluation task includes a model response mode, a model test set, a model test field, and evaluation case parameters of the model to be evaluated; so that evaluators can receive the evaluation task on the task publishing platform and evaluate the evaluation task according to the model response mode, the model test set, the model test field, and the evaluation case parameters; An acquisition module, configured to acquire evaluation parameters obtained after the evaluation of the evaluation task; A display module, configured to visually display the evaluation parameters of the model to be evaluated on the task publishing platform.

9. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method described in any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method described in any one of claims 1 to 7.