Model evaluation method and device, equipment, storage medium and product

By distributing and displaying the results of large-scale model processing, we ensure that the reviewers can only determine the evaluation parameters based on content comparison, solving the problem of biased evaluation results and achieving more accurate performance evaluation.

CN119938616APending Publication Date: 2025-05-06BEIJING 360 INTELLIGENT TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411998311.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-31
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

In the prior art, when evaluating the results of large-scale model processing, the reviewers are susceptible to the model version number, resulting in biased evaluation results and the accuracy of performance evaluation results cannot be guaranteed.

Method used

By determining the large model to be evaluated, multiple samples are processed separately for each large model to obtain multiple result files. Then, based on the file name of the result file, multiple result files are distributed to multiple evaluation folders to ensure that multiple result files in each evaluation folder do not belong to the same model. After displaying the contents of multiple evaluation folders, the reviewer can only compare the contents of multiple result files whose file names contain the same sample identification to determine the content evaluation parameters of each result file to avoid being affected by the model or folder.

Benefits of technology

Through this method, the objectivity of the evaluation results is ensured, evaluation bias is avoided, and the accuracy of the performance evaluation results of the large model is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119938616A_ABST
    Figure CN119938616A_ABST
Patent Text Reader

Abstract

The invention discloses a model evaluation method and device, equipment, a storage medium and a product, relates to the technical field of artificial intelligence, and discloses a method for determining a plurality of to-be-evaluated large models, and processing a plurality of samples based on each to-be-evaluated large model to obtain a plurality of result files; the file name of each result file is determined, the file name comprises a sample identifier corresponding to the result file and a model encryption identifier, and the model encryption identifier is obtained by encrypting a model identifier corresponding to the result file; based on the file names, the plurality of result files are distributed to a plurality of evaluation folders, each evaluation folder comprises a result file corresponding to each sample in the plurality of samples, and the plurality of result files in each evaluation folder do not belong to the same model; displaying the contents of the plurality of evaluation folders; and determining a performance evaluation result of each large model based on content evaluation parameters of each result file set by the displayed content. The method can improve the accuracy of the performance evaluation result of the large model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence technology, and in particular to a model evaluation method, device, equipment, storage medium and product. Background Art

[0002] The performance evaluation of large models is a complex process that involves multiple aspects of evaluation to ensure a comprehensive measurement of the model's performance. In addition to some indirect quantitative evaluation indicators, it is often necessary for evaluators to carefully compare the processing results of multiple large models to determine which large model performs better.

[0003] In related technologies, the processing results of each large model are provided to evaluators for evaluation. However, when evaluating the processing results, the evaluators are often influenced by the model to which the processing results belong. For example, evaluators tend to believe that the processing results of large models with higher version numbers are better, and thus give higher evaluations to large models with higher version numbers. This leads to biased evaluation results and makes it impossible to guarantee the accuracy of the performance evaluation results of large models.

[0004] The above contents are only used to assist in understanding the technical solution of the present application and do not constitute an admission that the above contents are prior art. Summary of the invention

[0005] The main purpose of this application is to provide a model evaluation method, device, equipment, storage medium and product that can improve the accuracy of performance evaluation results of large models.

[0006] To achieve the above objectives, the present application proposes a model evaluation method, which includes:

[0007] Determine multiple large models to be evaluated, and process multiple samples based on each large model to be evaluated to obtain multiple result files;

[0008] Determine a file name for each result file, wherein the file name includes a sample identifier and a model encryption identifier corresponding to the result file, and the model encryption identifier is obtained by encrypting the model identifier corresponding to the result file;

[0009] Distributing the multiple result files to multiple evaluation folders based on the file names, each evaluation folder containing a result file corresponding to each sample in the multiple samples, and the multiple result files in each evaluation folder do not belong to the same model;

[0010] Displaying the contents of the multiple evaluation folders;

[0011] The performance evaluation results of each model are determined based on the content evaluation parameters of each result file set based on the displayed content.

[0012] Optionally, distributing the plurality of result files to a plurality of evaluation folders based on the file names comprises:

[0013] The multiple result files are grouped based on the sample identifier in the file name to obtain multiple groups of files, wherein the file names of the multiple result files in each group of files contain the same sample identifier and different model encryption identifiers;

[0014] Randomly sort multiple result files in each group of files;

[0015] Each result file in each group of files is distributed to an evaluation folder corresponding to the sorting position of each result file.

[0016] Optionally, distributing the plurality of result files to a plurality of evaluation folders based on the file names comprises:

[0017] Based on the file names, the multiple result files are stored in a result record table, wherein the file names of the multiple result files in the same column of the result record table contain the same model encryption identifier, and the file names of the multiple result files in the same row contain the same sample identifier;

[0018] After randomly swapping the positions of multiple result files in the same row, multiple result files in the same column are distributed to the same evaluation folder.

[0019] Optionally, determining the file name of each result file includes:

[0020] Perform hash operation based on the sample identifier and model identifier corresponding to each result file to obtain a hash value corresponding to each result file;

[0021] Determine the hash value corresponding to each result file as the model encryption identifier corresponding to each result file;

[0022] The sample identifier and the model encryption identifier corresponding to each result file are determined as the file name of each result file.

[0023] Optionally, the content evaluation parameters of each result file set based on the displayed content to determine the performance evaluation results of each model include:

[0024] Determine content evaluation parameters of multiple result files corresponding to each major model based on content evaluation parameters of each result file set by the displayed content;

[0025] The content evaluation parameters of multiple result files corresponding to each major model are weighted and summed to obtain the performance evaluation results of each major model.

[0026] Optionally, after determining the file name of each result file, the method further includes:

[0027] Store the mapping relationship between the file name of each result file and the model identifier corresponding to each result file;

[0028] The content evaluation parameters of each result file set based on the displayed content determine the content evaluation parameters of multiple result files corresponding to each major model, including:

[0029] Determine the file name of the result file corresponding to each set content evaluation parameter, and map the file name to a model identifier based on the mapping relationship;

[0030] The content evaluation parameters of the multiple result files mapped to the same model identifier are determined as the content evaluation parameters of the multiple result files corresponding to the large model to which the model identifier belongs.

[0031] Optionally, displaying the contents of the multiple evaluation folders includes:

[0032] In the model evaluation interface, the contents of multiple result files whose file names contain the same sample identifier in the multiple evaluation folders are compared and displayed, wherein the file names of the multiple result files contain different model encryption identifiers.

[0033] Optionally, the model evaluation interface includes the multiple evaluation folders and content display controls corresponding to each sample identifier;

[0034] In the model evaluation interface, the contents of multiple result files whose file names contain the same sample identifier in the multiple evaluation folders are compared and displayed, including:

[0035] In response to a triggering operation on a content display control corresponding to any sample identifier in the model evaluation interface, determining a plurality of result files whose file names contain the sample identifier from the plurality of evaluation folders;

[0036] In the model evaluation interface, the contents of the multiple determined result files are compared and displayed.

[0037] Optionally, the method further comprises:

[0038] In response to a triggering operation on a content display control corresponding to any sample identifier in the model evaluation interface, the content of the sample to which the sample identifier belongs and the functional information of the multiple large models are displayed in the model evaluation interface.

[0039] Optionally, the large model is any one of the following models:

[0040] Large language models, image processing models, video processing models, and audio processing models.

[0041] In addition, to achieve the above purpose, the present application also proposes a model evaluation device, which includes:

[0042] A file acquisition module, used to determine multiple large models to be evaluated, and process multiple samples based on each large model to be evaluated to obtain multiple result files;

[0043] A name determination module, used to determine the file name of each result file, wherein the file name includes a sample identifier and a model encryption identifier corresponding to the result file, and the model encryption identifier is obtained by encrypting the model identifier corresponding to the result file;

[0044] A file distribution module, used to distribute the multiple result files to multiple evaluation folders based on the file names, each evaluation folder contains a result file corresponding to each sample in the multiple samples, and the multiple result files in each evaluation folder do not belong to the same model;

[0045] A content display module, used to display the contents of the multiple evaluation folders;

[0046] The performance evaluation module is used to determine the performance evaluation results of each model based on the content evaluation parameters of each result file set based on the displayed content.

[0047] Optionally, the file distribution module is used to group the multiple result files based on the sample identifier in the file name to obtain multiple groups of files, and the file names of the multiple result files in each group of files contain the same sample identifier and different model encryption identifiers; randomly sort the multiple result files in each group of files; and distribute each result file in each group of files to an evaluation folder corresponding to the sorting position of each result file.

[0048] Optionally, the file distribution module is used to store the multiple result files in a result record table based on the file names, the file names of the multiple result files located in the same column of the result record table contain the same model encryption identifier, and the file names of the multiple result files located in the same row contain the same sample identifier; after randomly swapping the positions of the multiple result files located in the same row, the multiple result files located in the same column are distributed to the same evaluation folder.

[0049] Optionally, the name determination module is used to perform hash operations based on the sample identifier and model identifier corresponding to each result file to obtain a hash value corresponding to each result file; determine the hash value corresponding to each result file as the model encryption identifier corresponding to each result file; and determine the sample identifier and model encryption identifier corresponding to each result file as the file name of each result file.

[0050] Optionally, the performance evaluation module includes:

[0051] A parameter determination unit, configured to determine content evaluation parameters of multiple result files corresponding to each major model based on content evaluation parameters of each result file set by the displayed content;

[0052] The performance evaluation unit is used to perform weighted summation on the content evaluation parameters of multiple result files corresponding to each major model to obtain the performance evaluation results of each major model.

[0053] Optionally, the name determination module is further used to store a mapping relationship between the file name of each result file and the model identifier corresponding to each result file;

[0054] The parameter determination unit is used to determine the file name of the result file corresponding to each set content evaluation parameter, and map the file name to a model identifier based on the mapping relationship; and determine the content evaluation parameters of multiple result files mapped to the same model identifier as the content evaluation parameters of multiple result files corresponding to the large model to which the model identifier belongs.

[0055] Optionally, the content display module is used to compare and display the contents of multiple result files whose file names contain the same sample identifier in the multiple evaluation folders on the model evaluation interface, wherein the file names of the multiple result files contain different model encryption identifiers.

[0056] Optionally, the model evaluation interface includes the multiple evaluation folders and content display controls corresponding to each sample identifier;

[0057] The content display module is used to respond to the triggering operation of the content display control corresponding to any sample identifier in the model evaluation interface, determine multiple result files whose file names contain the sample identifier from the multiple evaluation folders; in the model evaluation interface, compare and display the contents of the multiple determined result files.

[0058] Optionally, the device further comprises:

[0059] The content display module is also used to respond to the triggering operation of the content display control corresponding to any sample identifier in the model evaluation interface, and display the content of the sample to which the sample identifier belongs and the functional information of the multiple large models in the model evaluation interface.

[0060] Optionally, the large model is any one of the following models:

[0061] Large language models, image processing models, video processing models, and audio processing models.

[0062] In addition, to achieve the above objectives, the present application also proposes a model evaluation device, which includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program is configured to implement the steps of the model evaluation method described above.

[0063] In addition, to achieve the above-mentioned purpose, the present application also proposes a storage medium, which is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, the steps of the model evaluation method described above are implemented.

[0064] In addition, to achieve the above-mentioned purpose, the present application also provides a computer program product, which includes a computer program, and when the computer program is executed by a processor, the steps of the model evaluation method described above are implemented.

[0065] One or more technical solutions proposed in this application have at least the following technical effects:

[0066] The model evaluation scheme provided in the present application determines multiple large models to be evaluated, and processes multiple samples respectively based on each large model to be evaluated to obtain multiple result files. Then determine the file name of each result file, the file name contains the sample identification corresponding to the result file and the model encryption identification, and the model encryption identification is encrypted by the model identification corresponding to the result file. Therefore, the evaluator cannot judge which large model the result file comes from based on the file name of the result file. Then, based on the file name, multiple result files are distributed to multiple evaluation folders, each evaluation folder contains the result file corresponding to each sample in multiple samples, and the multiple result files in each evaluation folder do not belong to the same model. In other words, the result files generated by each large model are scattered in each evaluation folder. Therefore, after displaying the contents of multiple evaluation folders, the evaluator can only compare the contents of multiple result files whose file names contain the same sample identification in each evaluation folder to determine the content evaluation parameters of each result file without being disturbed by the model to which the result file belongs or the folder where the result file is located. This ensures the objectivity of the content evaluation parameters set for each result file. Therefore, the performance evaluation results of each major model are determined based on the content evaluation parameters of each result file set based on the displayed content, making the performance evaluation results unbiased and improving the accuracy of the determined performance evaluation results. BRIEF DESCRIPTION OF THE DRAWINGS

[0067] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.

[0068] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.

[0069] Figure 1 A schematic diagram of an implementation environment of the model evaluation method of this application;

[0070] Figure 2 A flowchart of the first embodiment of the model evaluation method of the present application is provided;

[0071] Figure 3 A flowchart of the second embodiment of the model evaluation method of the present application is provided;

[0072] Figure 4 A flowchart of the third embodiment of the model evaluation method of the present application is provided;

[0073] Figure 5 A flowchart of the fourth embodiment of the model evaluation method of the present application is provided;

[0074] Figure 6 A flowchart of the fifth embodiment of the model evaluation method of the present application is provided;

[0075] Figure 7 A schematic diagram of a model evaluation process provided in an embodiment of the present application;

[0076] Figure 8 This is a schematic diagram of the module structure of the model evaluation device of the embodiment of the present application;

[0077] Fig. 9 Schematic diagram of the device structure of the hardware operating environment involved in the model evaluation method in the embodiment of the present application.

[0078] The purpose, features and advantages of this application will be further described in conjunction with the embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION

[0079] It should be understood that the specific embodiments described herein are only used to explain the technical solutions of the present application and are not used to limit the present application.

[0080] In order to better understand the technical solution of the present application, a detailed description will be given below in conjunction with the accompanying drawings and specific implementation methods.

[0081] Figure 1 is a schematic diagram of an implementation environment provided by an embodiment of the present disclosure. Figure 1, the implementation environment includes a first terminal 101 and a second terminal 102. The first terminal 101 is used to provide data for reference in evaluating the large model. The second terminal 102 is used to determine the performance evaluation results of each large model based on the data. The first terminal 101 and the second terminal 102 are connected via a wireless or wired network. Exemplarily, the first terminal 101 and the second terminal 102 are computers, mobile phones, tablet computers or other terminals.

[0082] In the present application, the first terminal 101 is used to determine multiple large models to be evaluated. Based on each large model to be evaluated, multiple samples are processed separately to obtain multiple result files. The file name of each result file is determined, and the file name includes a sample identifier corresponding to the result file and a model encryption identifier, and the model encryption identifier is encrypted by the model identifier corresponding to the result file. Based on the file name, the multiple result files are distributed to multiple evaluation folders, each evaluation folder contains a result file corresponding to each sample in the multiple samples, and the multiple result files in each evaluation folder do not belong to the same model. Then the first terminal 101 sends multiple evaluation folders to the second terminal 102. The second terminal 102 receives multiple evaluation folders and displays the contents of the multiple evaluation folders. Based on the content evaluation parameters of each result file set by the displayed content, the performance evaluation results of each major model are determined.

[0083] Alternatively, the above model evaluation process can also be completed by any terminal alone. This embodiment of the present application does not limit this.

[0084] The model evaluation method provided in this application is applicable to a variety of scenarios. Exemplarily, a scenario in which multiple large language models are evaluated. For example, multiple sample texts are translated respectively based on multiple large language models to be evaluated to obtain multiple result files. Distribute the multiple result files to multiple evaluation folders according to the method provided in this application. Provide multiple evaluation folders to evaluators, who set content evaluation parameters for each result file based on the contents of the multiple evaluation folders displayed. Determine the performance evaluation results of each major model based on the content evaluation parameters of each result file. Exemplarily, the model evaluation method can also be applied to the evaluation of image processing models, video processing models, and audio processing models.

[0085] Figure 2 This is a flow chart of the first embodiment of the model evaluation method of this application. Figure 2 Taking the execution subject as the terminal as an example, the model evaluation method includes the following steps S10 to S50:

[0086] Step S10, determining a plurality of large models to be evaluated, and processing a plurality of samples respectively based on each large model to be evaluated to obtain a plurality of result files.

[0087] The large model to be evaluated refers to the large model that needs to be evaluated for performance. Exemplarily, the large model is a large language model, which is based on deep learning technology and has powerful natural language processing capabilities, and can perform tasks such as text generation, question answering, and translation. Alternatively, the large model is an image processing model, which is based on deep learning technology and has powerful image processing capabilities, and can perform tasks such as image generation, image restoration, and image recognition. Alternatively, the large model is a video processing model, which is based on deep learning technology and has powerful video processing capabilities, and can perform tasks such as video generation, video background replacement, and video object recognition. Alternatively, the large model is an audio processing model, which is based on deep learning technology and has powerful audio processing capabilities, and can perform tasks such as audio generation, audio-to-text conversion, and audio object recognition.

[0088] Samples are used to perform performance tests on large models. Multiple samples form a test data set, which is used to comprehensively evaluate the processing capabilities of large models under different input conditions. The sample format will change accordingly depending on the performance type of the large model being tested. For example, multiple large models are large language models, and this performance evaluation aims to test the translation performance of each language model. Accordingly, samples can be news reports, academic papers, etc. to be translated. Alternatively, this performance evaluation aims to test the answering performance of each language model. Accordingly, samples can be various questions. For another example, multiple large models are image generation models, and samples can be text descriptions of images, etc.

[0089] The result files are the files output by the big model after processing each sample. These files contain the processing results given by the big model for the sample, such as the translated text, the answer content, the generated image, etc. Each result file corresponds to a specific sample and a specific big model.

[0090] Step S20, determining the file name of each result file, the file name includes the sample identifier and the model encryption identifier corresponding to the result file, and the model encryption identifier is obtained by encrypting the model identifier corresponding to the result file.

[0091] Each sample has a specific sample ID. The sample ID is an identifier used to uniquely identify each sample. It can be a text, number, letter, number or other unique string combination. Through the sample ID, each sample and its corresponding result file can be accurately tracked and managed.

[0092] Each large model has a specific model identifier. The model identifier is an identifier used to uniquely identify each large model. It can be the name, version number, R&D organization and other information of the model, which is used to clearly distinguish different large models. The model encryption identifier is an identifier obtained by encrypting the model identifier. The purpose of encryption is to protect the identity information of the model and prevent unauthorized access and identification. For example, the model identifier is encrypted using a specific encryption algorithm to generate a seemingly irregular string as the model encryption identifier. The file name of the result file contains the model encryption identifier, which can not only protect the model information to a certain extent, but also be used for subsequent file management and identification.

[0093] The file name is the name of the result file. The file name contains the sample ID and model encryption ID corresponding to the result file. This naming method can clearly indicate which model processed which sample to obtain the result file, which is convenient for subsequent file management, distribution and search.

[0094] Step S30: Distribute the multiple result files to multiple evaluation folders based on the file names, each evaluation folder contains a result file corresponding to each sample in the multiple samples, and the multiple result files in each evaluation folder do not belong to the same model.

[0095] The evaluation folder is used to store result files. Each evaluation folder contains result files corresponding to multiple samples, but these result files come from different large models.

[0096] Step S40, displaying the contents of multiple evaluation folders.

[0097] Displaying the contents of multiple evaluation folders includes displaying each result file in each evaluation folder, for example, displaying the name of each result file, the content of each result file, etc.

[0098] Optionally, displaying the contents of multiple evaluation folders includes: in the model evaluation interface, comparing and displaying the contents of multiple result files whose file names contain the same sample identifier in multiple evaluation folders, wherein the file names of the multiple result files contain different model encryption identifiers.

[0099] The model evaluation interface is a user interface used to display and operate model evaluation related information. It can be in the form of a web page, application window, etc., providing users with a visual platform to facilitate users to view, compare and analyze the evaluation results of different models. On this interface, users can perform various operations related to model evaluation, such as viewing the content of result files, setting evaluation parameters, generating evaluation reports, etc.

[0100] It can be understood that in the model evaluation interface, the contents of multiple result files containing the same sample identifier and different model encryption identifiers are compared and displayed, that is, the processing results of different models on the same sample are compared and displayed, which can facilitate the evaluator to compare and evaluate the performance of different models on the same sample. Exemplarily, the comparative display can be in the form of side-by-side display, column display, list display, etc., so that users can easily compare and evaluate.

[0101] In the embodiment of the present application, on the model evaluation interface, the contents of multiple result files whose file names contain the same sample identifier in multiple evaluation folders are compared and displayed, and the file names of these multiple result files contain different model encryption identifiers, so that users can intuitively see the differences between the results generated by different models for the same sample. This display method can help users compare and analyze the performance of different models to evaluate the pros and cons of each model.

[0102] Optionally, the model evaluation interface includes multiple evaluation folders and content display controls corresponding to each sample identifier. Accordingly, in the model evaluation interface, the contents of multiple result files whose file names contain the same sample identifier in multiple evaluation folders are compared and displayed, including: the terminal responds to the triggering operation of the content display control corresponding to any sample identifier in the model evaluation interface, and determines multiple result files whose file names contain the sample identifier from multiple evaluation folders; in the model evaluation interface, the contents of the determined multiple result files are compared and displayed.

[0103] In an embodiment of the present application, the user only needs to trigger the content display control corresponding to a certain sample identifier, and the terminal can automatically and accurately locate all result files containing the sample identifier from multiple evaluation folders, thereby centrally displaying the result file contents of multiple models for the same sample. This avoids the user from manually searching and screening in a large number of files, saves a lot of time and energy, and significantly improves evaluation efficiency.

[0104] Optionally, the terminal responds to the triggering operation of the content display control corresponding to any sample identifier in the model evaluation interface, and also displays the content of the sample to which the sample identifier belongs and the functional information of multiple large models in the model evaluation interface. Among them, displaying the content of the sample to which the sample identifier belongs allows the evaluator to closely combine the original sample when comparing the result file, and understand the object and task background of the model processing. This helps to more accurately evaluate the rationality, relevance and accuracy of the model output results, avoid viewing the results in isolation, and improve the scientificity and objectivity of the evaluation. Moreover, presenting the functional information of multiple large models allows the evaluator to understand the functional characteristics of each model while comparing the results. This provides more dimensional references for evaluators to evaluate model results. For example, evaluators can judge whether the output results of the model meet expectations based on the functional characteristics of the model, so as to analyze the performance of the model in more depth.

[0105] Step S50, determining the performance evaluation results of each major model based on the content evaluation parameters of each result file set according to the displayed content.

[0106] Content evaluation parameters are a series of indicators for evaluating the content of the result file. These parameters can include accuracy, completeness, relevance, fluency, logic, etc. Depending on the type of large model and the type of task, corresponding content evaluation parameters will be set. By setting these parameters, the result file corresponding to the large model can be quantitatively evaluated to determine the performance of each large model.

[0107] Exemplarily, the large model is a large language model, and the task type of this test is text translation. Accordingly, the content evaluation parameters include accuracy parameters, completeness parameters, and fluency parameters. Exemplarily, the large model is an image generation model, and the task type of this test is to generate a corresponding image based on a text description. Accordingly, the content evaluation parameters include relevance parameters, completeness parameters, and picture aesthetic parameters.

[0108] The performance evaluation results are the evaluation results of the performance of each large model after evaluating the content of each result file according to the set content evaluation parameters. The performance evaluation results can be a comprehensive score, ranking, or a specific performance analysis of each large model on different content evaluation parameters. Through the performance evaluation results, you can intuitively understand the advantages and disadvantages of different large models when processing specific tasks, and provide a basis for model selection and improvement.

[0109] The model evaluation scheme provided in the present application determines multiple large models to be evaluated, and processes multiple samples respectively based on each large model to be evaluated to obtain multiple result files. Then determine the file name of each result file, the file name contains the sample identification corresponding to the result file and the model encryption identification, and the model encryption identification is encrypted by the model identification corresponding to the result file. Therefore, the evaluator cannot judge which large model the result file comes from based on the file name of the result file. Then, based on the file name, multiple result files are distributed to multiple evaluation folders, each evaluation folder contains the result file corresponding to each sample in multiple samples, and the multiple result files in each evaluation folder do not belong to the same model. In other words, the result files generated by each large model are scattered in each evaluation folder. Therefore, after displaying the contents of multiple evaluation folders, the evaluator can only compare the contents of multiple result files whose file names contain the same sample identification in each evaluation folder to determine the content evaluation parameters of each result file without being disturbed by the model to which the result file belongs or the folder where the result file is located. This ensures the objectivity of the content evaluation parameters set for each result file. Therefore, the performance evaluation results of each major model are determined based on the content evaluation parameters of each result file set based on the displayed content, making the performance evaluation results unbiased and improving the accuracy of the determined performance evaluation results.

[0110] Based on the above first embodiment, a second embodiment of the present application is proposed. For the same or similar contents as the first embodiment, please refer to the above introduction and will not be described in detail later. Figure 3 In the second embodiment, the above step S30 includes steps S301 to S303:

[0111] Step S301 , grouping multiple result files based on the sample identifier in the file name to obtain multiple groups of files, wherein the file names of multiple result files in each group of files contain the same sample identifier and different model encryption identifiers.

[0112] Exemplarily, the terminal classifies multiple result files whose file names contain the same sample identifier into the same group, thereby obtaining multiple groups of files. Each group of files contains result files of multiple large models for the same sample.

[0113] Step S302: randomly sort the multiple result files in each group of files.

[0114] The multiple result files in each group of files are randomly sorted so that the sorting positions of multiple result files of the same model in each group of files are not uniform. For example, the result file A of model A is ranked first in the file group corresponding to sample identifier A. The result file B of model A is ranked third in the file group corresponding to sample identifier B.

[0115] Step S303: distribute each result file in each group of files to an evaluation folder corresponding to the sorting position of each result file.

[0116] It can be understood that the number of result files contained in each group of files is consistent with the number of large models. Therefore, the number of sorting positions in each group of files is consistent with the number of large models. Among them, each sorting position corresponds to an evaluation folder, so each result file in each group of files is distributed to the evaluation folder corresponding to the sorting position of each result file, and the number of evaluation folders obtained is consistent with the number of large models.

[0117] Take 4 samples and 5 large models as an example. After grouping by samples, 4 groups of files will be obtained. Each group of files contains 5 result files, corresponding to 5 large models. After randomly sorting multiple result files in each group of files, the result file ranked first in each group of files is distributed to the evaluation folder corresponding to sorting position 1, and the result file ranked second in each group of files is distributed to the evaluation folder corresponding to sorting position 2, and so on, to obtain 5 evaluation folders.

[0118] The solution provided by the present application, since the file names of multiple result files in each group of files after grouping contain the same sample identifier and different model encryption identifiers. The multiple result files in each group of files are randomly sorted, and each result file in each group of files is distributed to the evaluation folder corresponding to the sorting position of each result file. It can be ensured that each evaluation folder obtained contains the result file corresponding to each sample in multiple samples, and the multiple result files in each evaluation folder do not belong to the same model. In other words, the result files generated by each large model are scattered in each evaluation folder. This ensures that the result files of each large model have the same display environment during the evaluation process, avoiding the subjective bias that may be caused by the result files of the model being fixed to a certain evaluation folder. Thereby ensuring that the evaluator can evaluate the results of each large model more objectively, so that the model evaluation results can more truly reflect the performance of each model.

[0119] Based on the above first embodiment, a third embodiment of the present application is proposed. For the same or similar contents as the first embodiment, please refer to the above introduction and will not be described in detail later. Figure 4 In the third embodiment, the above step S30 includes steps S304 to S305:

[0120] Step S304, storing multiple result files in a result record table based on file names, wherein the file names of multiple result files in the same column of the result record table contain the same model encryption identifier, and the file names of multiple result files in the same row contain the same sample identifier.

[0121] Take 4 large models and 4 samples as an example. The model encryption identifiers corresponding to the 4 large models are X1, X2, X3, and X4 respectively. The sample identifiers corresponding to the 4 samples are F1, F2, F3, and F4 respectively. Multiple result files are stored in the result record table based on the file name, and the result record table obtained is shown in Table 1:

[0122] F1_X1 F1_X2 F1_X3 F1_X4 F2_X1 F2_X2 F2_X3 F2_X4 F3_X1 F3_X2 F3_X3 F3_X4 F4_X1 F4_X2 F4_X3 F4_X4

[0123] Table 1

[0124] Step S305 , after randomly swapping the positions of the multiple result files in the same row, the multiple result files in the same column are distributed to the same evaluation folder.

[0125] Still taking the above 4 large models and 4 samples as an example, after randomly swapping the positions of multiple result files in the same row in the result record table, the result record table obtained is shown in Table 2:

[0126] F1_X1 F1_X4 F1_X2 F1_X3 F2_X4 F2_X3 F2_X2 F2_X1 F3_X3 F3_X1 F3_X4 F3_X2 F4_X2 F4_X1 F4_X3 F4_X4

[0127] Table 2

[0128] Then, multiple result files in the same column are distributed to the same evaluation folder, ensuring that each evaluation folder contains the result files corresponding to each sample in multiple samples, and the multiple result files in each evaluation folder do not belong to the same model. In other words, the result files generated by each large model are scattered in each evaluation folder. The result files in the same evaluation folder come from different models, so there will be no situation where the result files in a certain evaluation folder are significantly better than those in other evaluation folders, thus avoiding evaluation bias. Ensure that evaluators can evaluate the results of each large model more objectively, so that the model evaluation results can more truly reflect the performance of each model.

[0129] Based on the above-mentioned first embodiment of the present application, a fourth embodiment of the present application is proposed. The same or similar contents as those of the first embodiment can be referred to the above introduction, and will not be repeated in the following. Figure 5 In the fourth embodiment, the above step S20 includes step S201 and step S203.

[0130] Step S201 , performing a hash operation based on the sample identifier and the model identifier corresponding to each result file to obtain a hash value corresponding to each result file.

[0131] A hash operation is an operation that converts an input of any length into an output of a fixed length through a specific hash function. In the embodiment of the present application, the input is a sample identifier and a model identifier, and the output is a hash value. The hash function is unidirectional, so it is difficult to reversely deduce the original input, i.e., the sample identifier and the model identifier, based on the hash value, and different inputs will obtain different hash values. The hash value is an output value of a fixed length obtained by the hash operation, which is a characteristic representation of the input data, i.e., the sample identifier and the model identifier.

[0132] Step S202: determine the hash value corresponding to each result file as the model encryption identifier corresponding to each result file.

[0133] In an embodiment of the present application, the hash value corresponding to each result file is calculated based on its sample identifier and model identifier, and is used as a model encryption identifier.

[0134] Step S203: determine the sample identifier and the model encryption identifier corresponding to each result file as the file name of each result file.

[0135] By performing a hash operation on the sample identifier and the model identifier to generate a model encryption identifier, the real identification information of the model is hidden, making it impossible for evaluators to determine the model to which the result file belongs based on the model encryption identifier, thus avoiding the subjective bias of the evaluators caused by the exposure of the model information to which the result file belongs. In addition, each result file has a unique hash value generated based on its sample identifier and model identifier as the model encryption identifier, which together with the sample identifier constitutes the file name. This ensures that each result file has a unique identifier in the entire evaluation system, which facilitates the accurate classification, management, and retrieval of a large number of result files. And this unified file naming rule makes the file organization in the evaluation process more standardized and orderly, which can improve the automation and efficiency of the evaluation process.

[0136] Based on the above first embodiment of the present application, a fifth embodiment of the present application is proposed. For the same or similar contents as the first embodiment, please refer to the above introduction, and no further description will be given later. Figure 6 In the fifth embodiment, the above step S50 includes step S501 and step S502.

[0137] Step S501, based on the content evaluation parameters of each result file set by the displayed content, content evaluation parameters of multiple result files corresponding to each major model are determined.

[0138] Exemplarily, the model evaluation interface includes an evaluation parameter input box corresponding to each result file. After viewing the content of the result file, the evaluator can input the content evaluation parameter corresponding to the result file in the evaluation parameter input box.

[0139] In order to perform performance evaluation on each large model, it is necessary to first determine multiple content evaluation parameters corresponding to each large model. That is, the content evaluation parameters of multiple result files corresponding to each large model. Optionally, after determining the file name of each result file, the terminal stores the mapping relationship between the file name of each result file and the model identifier corresponding to each result file.

[0140] Correspondingly, after the terminal obtains the content evaluation parameters of each result file set based on the displayed content, it determines the file name of the result file corresponding to each content evaluation parameter set, and maps the file name to the model identifier based on the mapping relationship; the content evaluation parameters of multiple result files mapped to the same model identifier are determined as the content evaluation parameters of multiple result files corresponding to the large model to which the model identifier belongs. This solution utilizes the mapping relationship between the file name of each result file and the model identifier corresponding to each result file, and can quickly determine the large model to which each result file belongs, thereby determining the content evaluation parameters of multiple result files corresponding to each large model, thereby improving the model evaluation efficiency.

[0141] Step S502, performing weighted summation on the content evaluation parameters of the multiple result files corresponding to each major model to obtain the performance evaluation results of each major model.

[0142] Exemplarily, the weights of the content evaluation parameters of each result file are the same.

[0143] By weighted summing the content evaluation parameters of multiple result files corresponding to each large model, a specific value is generated as the performance evaluation result, making the performance comparison between different models more intuitive and objective. The pros and cons of different models can be known based on the numerical value of the performance evaluation result, providing an objective basis for model launch or model training corpus construction.

[0144] Figure 7 A schematic diagram of a model evaluation process provided in an embodiment of the present application. Figure 7 , the terminal generates result files in batches through multiple large models, that is, based on each large model to be evaluated, multiple samples are processed separately to obtain multiple result files. Then the file name is updated. That is, the file name of the result file is renamed to the sample identifier plus the model encryption identifier. Then a mapping relationship is generated and the mapping relationship is persisted. The mapping relationship is the mapping relationship between the new file name and the corresponding model identifier. At the same time, the result files are distributed to obtain multiple evaluation folders. The contents of the evaluation folder are displayed so that the evaluators can set the content evaluation parameters of each result file. Then, based on the mapping relationship, the multiple result files corresponding to each major model are determined, and based on the content evaluation parameters of the multiple result files corresponding to each major model, the content evaluation results of each major model are determined.

[0145] This application proposes an unbiased evaluation method for large models, which helps evaluators to conduct objective and unbiased comparative evaluation of large model performance. On the one hand, it can determine the model version with the best performance for product services, which can provide users with a better product experience. On the other hand, the accurate and objective content evaluation parameters given by the evaluators are important human preference signals for the iterative upgrade of the model version, which helps to provide the correct training corpus for the iterative upgrade of the model.

[0146] Another point that needs to be explained is that the above examples are only used to understand the present application and do not constitute a limitation on the model evaluation method of the present application. More simple transformations based on this technical concept are all within the scope of protection of the present application.

[0147] This application also provides a model evaluation device, please refer to Figure 8 , the model evaluation device includes:

[0148] The file acquisition module 10 is used to determine multiple large models to be evaluated, and process multiple samples based on each large model to be evaluated to obtain multiple result files;

[0149] A name determination module 20 is used to determine the file name of each result file, where the file name includes a sample identifier and a model encryption identifier corresponding to the result file, where the model encryption identifier is encrypted from the model identifier corresponding to the result file;

[0150] A file distribution module 30, for distributing the plurality of result files to a plurality of evaluation folders based on the file names, each evaluation folder containing a result file corresponding to each sample in the plurality of samples, and the plurality of result files in each evaluation folder do not belong to the same model;

[0151] A content display module 40, used to display the contents of multiple evaluation folders;

[0152] The performance evaluation module 50 is used to determine the performance evaluation results of each major model based on the content evaluation parameters of each result file set according to the displayed content.

[0153] Optionally, the file distribution module 30 is used to group multiple result files based on the sample identifier in the file name to obtain multiple groups of files, and the file names of the multiple result files in each group of files contain the same sample identifier and different model encryption identifiers; randomly sort the multiple result files in each group of files; and distribute each result file in each group of files to an evaluation folder corresponding to the sorting position of each result file.

[0154] Optionally, the file distribution module 30 is used to store multiple result files in a result record table based on file names, wherein the file names of the multiple result files located in the same column in the result record table contain the same model encryption identifier, and the file names of the multiple result files located in the same row contain the same sample identifier; after randomly swapping the positions of the multiple result files located in the same row, the multiple result files located in the same column are distributed to the same evaluation folder.

[0155] Optionally, the name determination module 20 is used to perform hash operations based on the sample identifier and model identifier corresponding to each result file to obtain a hash value corresponding to each result file; determine the hash value corresponding to each result file as the model encryption identifier corresponding to each result file; and determine the sample identifier and model encryption identifier corresponding to each result file as the file name of each result file.

[0156] Optionally, the performance evaluation module 50 includes:

[0157] A parameter determination unit, configured to determine content evaluation parameters of multiple result files corresponding to each major model based on content evaluation parameters of each result file set by the displayed content;

[0158] The performance evaluation unit is used to perform weighted summation on the content evaluation parameters of multiple result files corresponding to each major model to obtain the performance evaluation results of each major model.

[0159] Optionally, the name determination module 20 is further used to store a mapping relationship between the file name of each result file and the model identifier corresponding to each result file;

[0160] A parameter determination unit is used to determine the file name of the result file corresponding to each set content evaluation parameter, and map the file name to a model identifier based on a mapping relationship; content evaluation parameters of multiple result files mapped to the same model identifier are determined as content evaluation parameters of multiple result files corresponding to the large model to which the model identifier belongs.

[0161] Optionally, the content display module 40 is used to compare and display the contents of multiple result files whose file names contain the same sample identifier in multiple evaluation folders on the model evaluation interface, wherein the file names of the multiple result files contain different model encryption identifiers.

[0162] Optionally, the model evaluation interface includes multiple evaluation folders and content display controls corresponding to each sample identifier;

[0163] The content display module 40 is used to respond to the triggering operation of the content display control corresponding to any sample identifier in the model evaluation interface, determine multiple result files whose file names contain the sample identifier from multiple evaluation folders; in the model evaluation interface, compare and display the contents of the multiple determined result files.

[0164] Optionally, the device further comprises:

[0165] The content display module 40 is also used to respond to the triggering operation of the content display control corresponding to any sample identifier in the model evaluation interface, and display the content of the sample to which the sample identifier belongs and the functional information of multiple large models in the model evaluation interface.

[0166] Optionally, the large model is any of the following models:

[0167] Large language models, image processing models, video processing models, and audio processing models.

[0168] The model evaluation device provided by the present application adopts the model evaluation method in the above embodiment, which can solve the technical problem that the performance evaluation results determined by the relevant technology are biased, resulting in inaccurate performance evaluation results. Compared with the prior art, the beneficial effects of the model evaluation device provided by the present application are the same as the beneficial effects of the model evaluation method provided by the above embodiment, and the other technical features in the model evaluation device are the same as the features disclosed in the above embodiment method, which will not be repeated here.

[0169] The present application provides a model evaluation device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the model evaluation method in the above-mentioned embodiment one.

[0170] Reference below Fig. 9 , which shows a schematic diagram of the structure of a model evaluation device suitable for implementing an embodiment of the present application. The model evaluation device in the embodiment of the present application may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Descriptions), PMPs (Portable Media Players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Fig. 9 The model evaluation device shown is merely an example and should not bring any limitation to the functions and scope of use of the embodiments of the present application.

[0171] like Fig. 9 As shown, the model evaluation device may include a processing device 1001 (such as a central processing unit, a graphics processor, etc.), which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM: Read Only Memory) 1002 or a program loaded from a storage device 1003 to a random access memory (RAM: Random Access Memory) 1004. In RAM1004, various programs and data required for the operation of the model evaluation device are also stored. The processing device 1001, ROM1002 and RAM1004 are connected to each other through a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Generally, the following systems can be connected to the I / O interface 1006: an input device 1007 including, for example, a touch screen, a touch pad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; an output device 1008 including, for example, a liquid crystal display (LCD: Liquid Crystal Display), a speaker, a vibrator, etc.; a storage device 1003 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1009. The communication device 1009 can allow the model evaluation device to communicate with other devices wirelessly or by wire to exchange data. Although the figure shows a model evaluation device with various systems, it should be understood that it is not required to implement or have all the systems shown. More or fewer systems can be implemented or provided instead.

[0172] In particular, according to the embodiments disclosed in the present application, the process described above with reference to the flowchart can be implemented as a computer software program. For example, the embodiments disclosed in the present application include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network through a communication device, or installed from a storage device 1003, or installed from a ROM 1002. When the computer program is executed by the processing device 1001, the above-mentioned functions defined in the method of the embodiment disclosed in the present application are executed.

[0173] The model evaluation device provided by the present application adopts the model evaluation method in the above embodiment, which can solve the technical problem that the performance evaluation results determined by the relevant technology are biased, resulting in inaccurate performance evaluation results. Compared with the prior art, the beneficial effects of the model evaluation device provided by the present application are the same as the beneficial effects of the model evaluation method provided by the above embodiment, and the other technical features in the model evaluation device are the same as the features disclosed in the previous embodiment method, which will not be repeated here.

[0174] It should be understood that the various parts disclosed in this application can be implemented by hardware, software, firmware or a combination thereof. In the description of the above embodiments, specific features, structures, materials or characteristics can be combined in any one or more embodiments or examples in a suitable manner.

[0175] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art who is familiar with the present technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.

[0176] The present application provides a computer-readable storage medium having computer-readable program instructions (ie, computer programs) stored thereon, and the computer-readable program instructions are used to execute the model evaluation method in the above-mentioned embodiment.

[0177] The computer-readable storage medium provided in the present application may be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, systems or devices, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in combination with an instruction execution system, system or device. The program code contained on the computer-readable storage medium may be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination of the above.

[0178] The computer-readable storage medium may be included in the model evaluation device; or may exist independently without being assembled into the model evaluation device.

[0179] The above-mentioned computer-readable storage medium carries one or more programs. When the above-mentioned one or more programs are executed by the model evaluation device, the model evaluation device: determines multiple large models to be evaluated, and processes multiple samples respectively based on each large model to be evaluated to obtain multiple result files; determines the file name of each result file, and the file name includes a sample identifier and a model encryption identifier corresponding to the result file, and the model encryption identifier is encrypted by the model identifier corresponding to the result file; distributes the multiple result files to multiple evaluation folders based on the file name, each evaluation folder contains a result file corresponding to each sample in the multiple samples, and the multiple result files in each evaluation folder do not belong to the same model; displays the contents of multiple evaluation folders; and determines the performance evaluation results of each large model based on the content evaluation parameters of each result file set based on the displayed content.

[0180] Computer program code for performing the operations of the present application may be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, C++, and conventional procedural programming languages ​​such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0181] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the system, method and computer program product according to various embodiments of the present application. In this regard, each square box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and the module, the program segment or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two square boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each square box in the block diagram and / or flow chart, and the combination of the square boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0182] The modules involved in the embodiments described in this application may be implemented by software or hardware, wherein the name of the module does not constitute a limitation on the unit itself in some cases.

[0183] The readable storage medium provided in this application is a computer-readable storage medium, which stores computer-readable program instructions (i.e., computer programs) for executing the above-mentioned model evaluation method, and can solve the technical problem that the performance evaluation results determined by the relevant technology are biased, resulting in inaccurate performance evaluation results. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in this application are the same as the beneficial effects of the model evaluation method provided in the above-mentioned embodiment, and will not be repeated here.

[0184] The present application also provides a computer program product, including a computer program, which implements the steps of the above-mentioned model evaluation method when executed by a processor.

[0185] The computer program product provided by this application can solve the technical problem that the performance evaluation results determined by the related technology are biased, resulting in inaccurate performance evaluation results. Compared with the prior art, the beneficial effects of the computer program product provided by this application are the same as the beneficial effects of the model evaluation method provided by the above embodiment, which will not be repeated here.

[0186] The above descriptions are only some embodiments of the present application, and are not intended to limit the patent scope of the present application. All equivalent structural changes made using the contents of the present application specification and drawings under the technical concept of the present application, or direct / indirect applications in other related technical fields are included in the patent protection scope of the present application.

Claims

1. A model evaluation method, characterized in that: The method comprises: Determine multiple large models to be evaluated, and process multiple samples based on each large model to be evaluated to obtain multiple result files; Determine a file name for each result file, wherein the file name includes a sample identifier and a model encryption identifier corresponding to the result file, and the model encryption identifier is obtained by encrypting the model identifier corresponding to the result file; Distributing the multiple result files to multiple evaluation folders based on the file names, each evaluation folder containing a result file corresponding to each sample in the multiple samples, and the multiple result files in each evaluation folder do not belong to the same model; Displaying the contents of the multiple evaluation folders; The performance evaluation results of each model are determined based on the content evaluation parameters of each result file set based on the displayed content.

2. The method according to claim 1, characterized in that The distributing the plurality of result files to a plurality of evaluation folders based on the file names comprises: The multiple result files are grouped based on the sample identifier in the file name to obtain multiple groups of files, wherein the file names of the multiple result files in each group of files contain the same sample identifier and different model encryption identifiers; Randomly sort multiple result files in each group of files; Each result file in each group of files is distributed to an evaluation folder corresponding to the sorting position of each result file.

3. The method according to claim 1, characterized in that The distributing the plurality of result files to a plurality of evaluation folders based on the file names comprises: Based on the file names, the multiple result files are stored in a result record table, wherein the file names of the multiple result files in the same column of the result record table contain the same model encryption identifier, and the file names of the multiple result files in the same row contain the same sample identifier; After randomly swapping the positions of multiple result files in the same row, multiple result files in the same column are distributed to the same evaluation folder.

4. The method according to claim 1, characterized in that Determining the file name of each result file includes: Perform hash operation based on the sample identifier and model identifier corresponding to each result file to obtain a hash value corresponding to each result file; Determine the hash value corresponding to each result file as the model encryption identifier corresponding to each result file; The sample identifier and the model encryption identifier corresponding to each result file are determined as the file name of each result file.

5. The method according to claim 1, characterized in that The content evaluation parameters of each result file set based on the displayed content determine the performance evaluation results of each model, including: Determine content evaluation parameters of multiple result files corresponding to each major model based on content evaluation parameters of each result file set by the displayed content; The content evaluation parameters of multiple result files corresponding to each major model are weighted and summed to obtain the performance evaluation results of each major model.

6. The method according to claim 5, characterized in that After determining the file name of each result file, the method further includes: Store the mapping relationship between the file name of each result file and the model identifier corresponding to each result file; The content evaluation parameters of each result file set based on the displayed content determine the content evaluation parameters of multiple result files corresponding to each major model, including: Determine the file name of the result file corresponding to each set content evaluation parameter, and map the file name to a model identifier based on the mapping relationship; The content evaluation parameters of the multiple result files mapped to the same model identifier are determined as the content evaluation parameters of the multiple result files corresponding to the large model to which the model identifier belongs.

7. A model evaluation device, characterized in that: The device comprises: A file acquisition module, used to determine multiple large models to be evaluated, and process multiple samples based on each large model to be evaluated to obtain multiple result files; A name determination module, used to determine the file name of each result file, wherein the file name includes a sample identifier and a model encryption identifier corresponding to the result file, and the model encryption identifier is obtained by encrypting the model identifier corresponding to the result file; A file distribution module, used to distribute the multiple result files to multiple evaluation folders based on the file names, each evaluation folder contains a result file corresponding to each sample in the multiple samples, and the multiple result files in each evaluation folder do not belong to the same model; A content display module, used to display the contents of the multiple evaluation folders; The performance evaluation module is used to determine the performance evaluation results of each model based on the content evaluation parameters of each result file set based on the displayed content.

8. A model evaluation device, characterized in that: The device comprises: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program is configured to implement the steps of the model evaluation method according to any one of claims 1 to 6.

9. A storage medium, characterized in that: The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, the steps of the model evaluation method according to any one of claims 1 to 6 are implemented.

10. A computer program product, characterized in that The computer program product comprises a computer program, and when the computer program is executed by a processor, the steps of the model evaluation method according to any one of claims 1 to 6 are implemented.