Model evaluation method, system, equipment, medium and product

By configuring evaluation benchmarks and generating lists in the model evaluation system, the difficulties of machine learning model evaluation and comparison in different fields are solved, and automated and effective model performance evaluation is achieved.

CN120144979APending Publication Date: 2025-06-13BEIJING VOLCANO ENGINE TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510220745.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-26
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

How to effectively evaluate and compare machine learning models in different fields, taking into account the differences in the characteristics of model structure, base module, parameter quantity, training methods and training data.

Method used

Provide a model evaluation method and system to evaluate machine learning models through pre-configured evaluation benchmarks (including evaluation data sets, evaluation indicators and evaluation methods). The system is able to respond to evaluation requests, acquire multiple machine learning models and corresponding evaluation benchmarks, and generate lists to describe the model's ranking in performance.

Benefits of technology

It realizes automated model evaluation and comparison, overcomes the shortcomings of users' manual participation in the evaluation process, and can effectively evaluate and compare the performance of different machine learning models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120144979A_ABST
    Figure CN120144979A_ABST
Patent Text Reader

Abstract

The invention discloses a model evaluation method, system, device, medium and product, at least one pre-configured evaluation benchmark is stored in the system, and the working principle of the system comprises the following steps: responding to an evaluation request triggered by the system; acquiring a plurality of machine learning models indicated by the request and acquiring an evaluation reference indicated by the request so as to subsequently evaluate the performance of each model by using the evaluation reference indicated by the request; according to the method, a request is sent to each model, a list is generated according to the evaluation result of each model under the evaluation reference indicated by the request, so that the list is used for describing the ranking of each model in performance, the list can better represent the relative quality degree of the models, the quality of the models can be automatically evaluated and compared, and the user experience is improved. In this way, the defects caused by the fact that a user manually participates in the evaluation and comparison process of the models can be effectively overcome.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of data processing, and in particular, to a method, system, device, medium, and product for model evaluation. Background Art

[0002] With the rapid development of machine learning models, various machine learning models are widely applied in multiple fields to meet the data processing requirements in each field, such as video generation requirements or code rewriting requirements, etc.

[0003] In addition, in order to better adapt to the requirements in different fields, machine learning models in different fields have different characteristics, such as different model structures, different base modules, different numbers of parameters, different training methods, different training data, etc., making it an urgent technical problem to evaluate and compare the advantages and disadvantages of these models. Summary of the Invention

[0004] To solve the above technical problems, this application provides a method, system, device, medium, and product for model evaluation.

[0005] To achieve the above object, the technical solutions provided by this application are as follows:

[0006] This application provides a method for model evaluation. The method is applied to a model evaluation system in which at least one pre-configured evaluation benchmark is stored. The evaluation benchmark includes an evaluation data set, evaluation metrics, and an evaluation method. The evaluation benchmark is used to evaluate the performance of a machine learning model. The method includes: in response to an evaluation request triggered through the system, obtaining a plurality of machine learning models indicated by the evaluation request, and obtaining the evaluation benchmark indicated by the evaluation request. The at least one evaluation benchmark includes the evaluation benchmark indicated by the evaluation request; generating a leaderboard based on the evaluation results of each model in the plurality of machine learning models under the evaluation benchmark indicated by the evaluation request. The leaderboard is used to describe the ranking of each model in the plurality of machine learning models in terms of performance.

[0007] In a possible implementation, the system includes evaluation benchmarks in multiple dimensions. The evaluation benchmarks in different dimensions are used to evaluate the performance of machine learning models in different dimensions. For any dimension, the evaluation benchmarks in this dimension include evaluation benchmarks corresponding to multiple version numbers, and there are differences in the versions of some or all of the content between the evaluation benchmarks corresponding to different version numbers in the same dimension. The part or all of the content includes at least one of an evaluation data set, evaluation metrics, and an evaluation method; the process of obtaining the evaluation benchmark indicated by the evaluation request includes: determining a target dimension and a target version number according to the evaluation request, the multiple dimensions including the target dimension, and the multiple version numbers in the target dimension including the target version number; searching for the evaluation benchmark indicated by the evaluation request in the storage space of the system according to the target dimension and the target version number, the version number corresponding to the evaluation benchmark indicated by the evaluation request being the target version number, and the evaluation benchmark indicated by the evaluation request being used to evaluate the performance of each model in the multiple machine learning models in the target dimension.

[0008] In a possible implementation, if the evaluation request is automatically triggered by the system at a preset period, the target version number is the version number with the latest creation time in the target dimension; if the evaluation request is manually triggered by a user, the evaluation request carries a version number specified by the user, and the target version number is the version number carried by the evaluation request.

[0009] In a possible implementation, the system includes evaluation benchmarks in multiple dimensions. For any dimension, the evaluation benchmarks in this dimension include evaluation benchmarks corresponding to multiple version numbers; the method further includes: for any dimension, in response to an update operation triggered for at least one piece of content in the evaluation benchmark corresponding to the version number with the latest creation time in this dimension, creating a new version number in this dimension, determining the evaluation benchmark corresponding to the new version number according to the update operation, and determining the creation reason corresponding to the new version number according to the update operation, the creation reason including the content change indicated by the update operation, and the at least one piece of content including some or all of an evaluation data set, evaluation metrics, and an evaluation method.

[0010] In a possible implementation, the evaluation benchmark indicated by the evaluation request includes at least two evaluation benchmarks; the process of generating the leaderboard includes: for any model in the multiple machine learning models, analyzing the leaderboard score of the model according to the evaluation results of the model under each evaluation benchmark in the at least two evaluation benchmarks; sorting the multiple machine learning models according to the leaderboard score to obtain the leaderboard.

[0011] In a possible implementation manner, for any one of the multiple machine learning models, the process of determining the list score of this model includes: calculating the average value of the evaluation results of this model under each evaluation benchmark among the at least two evaluation benchmarks to obtain the list score of this model.

[0012] In a possible implementation manner, for any one of the multiple machine learning models, the process of determining the list score of this model includes: determining the evaluation results of this model in each scenario based on the scenario classification information of each evaluation benchmark among the at least two evaluation benchmarks and the evaluation results of this model under each evaluation benchmark among the at least two evaluation benchmarks, and performing a weighted summation process on the evaluation results of this model in each scenario according to the weights corresponding to each scenario to obtain the list score of this model.

[0013] In a possible implementation manner, for any one of the multiple machine learning models, the process of determining the list score of this model includes: performing a weighted summation process on the evaluation results of this model under each evaluation benchmark among the at least two evaluation benchmarks according to the weights configured for each evaluation benchmark among the at least two evaluation benchmarks to obtain the list score of this model.

[0014] In a possible implementation manner, for any one of the multiple machine learning models, the process of determining the list score of this model includes: determining the ranking of this model under each evaluation benchmark among the at least two evaluation benchmarks based on the evaluation results of this model under each evaluation benchmark among the at least two evaluation benchmarks, and calculating the average value of the rankings of this model under each evaluation benchmark among the at least two evaluation benchmarks to obtain the list score of this model.

[0015] In a possible implementation manner, after obtaining the multiple machine learning models indicated by the evaluation request and the evaluation benchmarks indicated by the evaluation request, the method further includes: for any one of the multiple machine learning models, checking whether there is an execution record of the evaluation task of this model under the evaluation benchmark indicated by the evaluation request in the historical execution record of the system; if it exists, querying the execution result of the evaluation task from the storage space of the system to obtain the evaluation result of this model under the evaluation benchmark indicated by the evaluation request; if it does not exist, creating and executing the evaluation task to obtain the evaluation result of this model under the evaluation benchmark indicated by the evaluation request.

[0016] In a possible implementation, the method further includes: in response to an update of the model, creating and executing the evaluation task to obtain an evaluation result of the model under the evaluation benchmark indicated by the evaluation request; the step of looking up in the historical execution records of the system whether there is an execution record of an evaluation task of the model under the evaluation benchmark indicated by the evaluation request includes: in response to the model not being updated, looking up in the historical execution records of the system whether there is an execution record of an evaluation task of the model under the evaluation benchmark indicated by the evaluation request.

[0017] In a possible implementation, the system includes an import tool for importing files from different storage systems into the evaluation dataset stored in the system; the import process includes: in response to a data import request triggered by the import tool, determining meta-information according to the file or file storage path carried in the data import request, where the meta-information is used to describe the file indicated by the data import request; creating a data table in the system according to the meta-information, where the data table is used to describe the file indicated by the data import request according to the data structure of each data in the evaluation dataset; and updating the evaluation dataset stored in the system by using the data table.

[0018] In a possible implementation, for any one of the evaluation benchmarks, the evaluation metrics of the evaluation benchmark include a main metric and an auxiliary metric, where the main metric is used to affect the ranking, and the auxiliary metric is used to provide auxiliary information for the list.

[0019] The present application provides a model evaluation system, in which at least one pre-configured evaluation benchmark is stored, the evaluation benchmark includes an evaluation dataset, evaluation metrics and an evaluation method, and the evaluation benchmark is used to evaluate the performance of a machine learning model. The system includes: an acquisition unit for, in response to an evaluation request triggered by the system, acquiring a plurality of machine learning models indicated by the evaluation request and an evaluation benchmark indicated by the evaluation request, where the at least one evaluation benchmark includes the evaluation benchmark indicated by the evaluation request; and a generation unit for generating a list according to the evaluation results of each model in the plurality of machine learning models under the evaluation benchmark indicated by the evaluation request, where the list is used to describe the ranking of each model in the plurality of machine learning models in terms of performance.

[0020] The present application provides an electronic device, which includes: a processor and a memory; the memory is used to store instructions or computer programs; the processor is used to execute the instructions or computer programs in the memory so that the electronic device executes the model evaluation method provided by the present application.

[0021] The present application provides a computer-readable medium, characterized in that instructions or a computer program are stored in the computer-readable medium, and when the instructions or the computer program run on a device, the device is caused to execute the model evaluation method provided by the present application.

[0022] The present application provides a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program includes program code for executing the model evaluation method provided by the present application.

[0023] Compared with the related art, the present application has at least the following advantages:

[0024] For the model evaluation system with an automatic evaluation function provided by the present application, at least one pre-configured evaluation benchmark (such as the evaluation benchmark under content understanding, the evaluation benchmark under logical reasoning, etc.) is stored in the system, and because the evaluation benchmark includes an evaluation data set (such as some data tables, some images or some videos, etc.), evaluation metrics (such as recall rate, accuracy rate, etc.) and an evaluation method, so that the evaluation benchmark is used to evaluate the performance of any machine learning model. In addition, the working principle of the system includes: in response to an evaluation request triggered through the system, first obtain a plurality of machine learning models indicated by the request and obtain the evaluation benchmark indicated by the request, so as to be able to use the evaluation benchmark indicated by the request to evaluate the performance of each model later; then generate a list according to the evaluation results of each model under the evaluation benchmark indicated by the request, so that the list is used to describe the ranking of each model in terms of performance, so that the list can better represent the relative superiority and inferiority degree among these models, and further be able to realize the automatic evaluation and comparison of the superiority and inferiority among these models, so that the defects caused by the user manually participating in the evaluation and comparison process of these models can be effectively overcome. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the related art, the following will briefly introduce the drawings required for use in the description of the embodiments or the related art. Obviously, the drawings in the following description are only some embodiments recorded in the present application. For those of ordinary skill in the art, other drawings can be obtained according to these drawings without creative efforts.

[0026] Figure 1 It is a flowchart of a model evaluation method provided by an embodiment of the present application;

[0027] Figure 2 It is a schematic structural diagram of a model evaluation system provided by an embodiment of the present application;

[0028] Figure 3 It is a schematic diagram of a list generation process with a cache mechanism provided by an embodiment of the present application;

[0029] Figure 4 A structural schematic diagram of a model evaluation system provided by an embodiment of the present application;

[0030] Figure 5 A structural schematic diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners

[0031] In order to enable those skilled in the art to better understand the solution of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0032] To better understand the technical solution provided by the present application, the model evaluation method provided by the present application will be described below with reference to some accompanying drawings. As Figure 1 shown, the model evaluation method provided by the embodiment of the present application is applied to a model evaluation system (such as the model evaluation system shown in Figure 2 ). At least one pre-configured evaluation benchmark (Benchmark) is stored in the system. The evaluation benchmark includes an evaluation data set, evaluation metrics, and an evaluation method. The evaluation benchmark is used to evaluate the performance of a machine learning model, and the model evaluation method includes S1-S2 below.

[0033] S1: In response to an evaluation request triggered by the model evaluation system, obtain multiple machine learning models indicated by the evaluation request, and obtain the evaluation benchmark indicated by the evaluation request. At least one evaluation benchmark stored in the system includes the evaluation benchmark indicated by the evaluation request.

[0034] Among them, the evaluation benchmark is a pre-constructed benchmark evaluation system for performance evaluation of machine learning models, such as the Benchmark shown in Figure 2 . And the evaluation benchmark includes an evaluation data set, evaluation metrics, and an evaluation method, so that the evaluation standard can perform performance evaluation processing for machine learning models based on the evaluation data set, evaluation metrics, and evaluation method.

[0035] For the above-mentioned evaluation datasets (such as the evaluation set), the evaluation dataset refers to the set of data (such as data tables, images, videos, audio, etc.) required for performance evaluation of machine learning models, so that the evaluation process implemented based on this set can evaluate the advantages and disadvantages of the model under different difficulties, different scenarios, and different dimensions. In addition, the evaluation dataset can be divided into a training set, a validation set, and a test set according to a certain ratio. Furthermore, this application does not limit the acquisition method of the evaluation dataset.

[0036] For the above-mentioned evaluation metrics, the evaluation metric refers to the evaluation metrics required for performance evaluation of machine learning models, such as accuracy, recall rate, and other evaluation metrics. It should be noted that this application does not limit the acquisition method of the evaluation metric.

[0037] In addition, for the same evaluation benchmark, the evaluation metrics of the evaluation benchmark can include multiple metrics, so that the evaluation benchmark can better evaluate the performance of the machine learning model in a certain dimension (such as the dimension of content understanding) with the help of these metrics.

[0038] Furthermore, in order to better improve the evaluation effect, for the same evaluation benchmark, when the evaluation benchmark includes multiple metrics, the multiple metrics can include a main metric and at least one auxiliary metric. The main metric refers to the metric that is mainly concerned under the evaluation benchmark, so that the main metric becomes the metric that must be used when determining the list ranking; the auxiliary metric refers to the other metrics used by the evaluation benchmark except the main metric, so that the auxiliary metric does not participate in the list ranking, but needs to be provided to the user as auxiliary information of the list to enrich the information displayed by the list.

[0039] It can be seen that for any evaluation benchmark, the evaluation metrics of the evaluation benchmark can include a main metric and an auxiliary metric. The main metric is used to affect the list ranking, and the auxiliary metric is used to provide auxiliary information for the list. It should be noted that in some scenarios, in order to better improve the evaluation effect, a Benchmark has and only has one main metric; and when determining the auxiliary information of the list based on the auxiliary metric, information such as the metric name, metric description, and metric control parameters of the auxiliary metric can be provided.

[0040] For the above-mentioned evaluation method, the evaluation method is used to describe how to use the evaluation dataset and evaluation metrics to achieve the performance evaluation of machine learning models, so that the evaluation method can not only describe the performance evaluation process of the machine learning model, but also describe the use process of the evaluation dataset and the use process of the evaluation metrics, so that the evaluation method can describe as detailed as possible the execution process of the evaluation task under the Benchmark including the evaluation method. It should be noted that this application does not limit the acquisition method of the evaluation method.

[0041] For the above model evaluation system, for a platform including Online Analytical Processing (OLAP), at least one pre-configured evaluation benchmark (such as Figure 3 the Benchmark shown 1 , Benchmark 2 , Benchmark 3 ) is stored in the system, so that some or all of the Benchmarks can be automatically or manually selected from these evaluation benchmarks for evaluating the performance of machine learning models later.

[0042] The evaluation request is used to request the evaluation of multiple machine learning models according to one or more evaluation benchmarks, so that the evaluation request can not only indicate the machine learning models to be evaluated (such as Figure 3 the model 1, model 2,... shown Figure 3 ), but also indicate the evaluation benchmarks required to be used when evaluating the performance of the machine learning models (such as 1 the Benchmark shown 2 , Benchmark 3 ).

[0043] In addition, the present application does not limit the triggering method of the evaluation request. For example, in some scenarios, such as the automatic evaluation scenario, in order to better reduce the defects caused by manual participation, the evaluation request can be automatically triggered by the model evaluation system according to a preset period. Wherein, the preset period refers to the scheduled execution time pre-configured for the automatic evaluation task, such as executed weekly, bi-weekly, or monthly.

[0044] For another example, in some scenarios, such as the manual evaluation scenario, in order to better meet the user's needs, the above evaluation request can be manually triggered by the user, so that the evaluation request can describe the user's evaluation needs. It should be noted that the present application does not limit the implementation manner of this triggering.

[0045] Based on the relevant content of S1 above, for the model evaluation system, after detecting an evaluation request triggered through the system (such as an automatically triggered evaluation request or a manually triggered evaluation request), obtain the multiple machine learning models to be evaluated indicated by the request, and obtain the evaluation benchmarks required to be used when evaluating these machine learning models indicated by the request, so that the performance evaluation of these machine learning models can be realized based on these evaluation benchmarks later.

[0046] S2: Generate a leaderboard based on the evaluation results of each model among multiple machine learning models under the evaluation criteria indicated by the evaluation request. This leaderboard is used to describe the performance rankings of the models among the multiple machine learning models.

[0047] Among them, the evaluation result of the j-th machine learning model under the evaluation criteria indicated by the evaluation request refers to the result (such as a score) obtained by evaluating the performance of the j-th machine learning model using this evaluation criteria, so that this result can represent the performance presented by the j-th machine learning model under this evaluation criteria, such as content understanding performance, etc. j is a positive integer, j ≤ J, and J is a positive integer representing the number of models among the above multiple machine learning models.

[0048] In addition, the present application does not limit the implementation manner of the above S2. For example, specifically, it can be: first, sort the multiple machine learning models according to the evaluation results of each model among the multiple machine learning models under the evaluation criteria indicated by the evaluation request to obtain a sorting result, so that this sorting result can represent the performance rankings of the models; then draw and display the leaderboard based on this sorting result.

[0049] Based on the relevant content of the above S1 to S2, for the model evaluation system with an automatic evaluation function provided by the present application, the working principle of this system includes: in response to an evaluation request triggered through this system, first obtain the multiple machine learning models indicated by this request and the evaluation criteria indicated by this request, so as to be able to evaluate the performance of each model using the evaluation criteria indicated by this request subsequently; then generate a leaderboard based on the evaluation results of each model under the evaluation criteria indicated by this request, so that this leaderboard is used to describe the performance rankings of the models, thereby enabling this leaderboard to better represent the relative superiority and inferiority degrees among these models, and further enabling the automatic evaluation and comparison of the superiority and inferiority of these models, so that the defects caused by the user's manual participation in the evaluation and comparison process of these models can be effectively overcome.

[0050] In addition, the present application does not limit the execution entity of the above model evaluation method. For example, this method can be applied to a terminal device or a server. Another example is that this method can also be implemented by means of the data interaction process between the terminal device and the server. Among them, the terminal device can be a smart phone, a computer, a personal digital assistant (Personal Digital Assistant, PDA), a tablet computer, etc. The server can be an independent server, a cluster server, or a cloud server.

[0051] It has been found through research that a single evaluation criterion may only be used to evaluate the performance of a machine learning model in a certain dimension (such as content understanding), so that the model evaluation system implemented based on this evaluation criterion can only evaluate the performance presented by the model in a single aspect, thus limiting the application scope of this system.

[0052] Based on the above research, in order to overcome the above problems, the above model evaluation system may include evaluation benchmarks in multiple dimensions, and the evaluation benchmarks in different dimensions are used to evaluate the performance of the machine learning model in different dimensions. In this way, it is beneficial to use this system to comprehensively evaluate the machine learning model, so that this system is applicable to perform model evaluation processing in various situations. It should be noted that the embodiments of the multiple dimensions are not limited in this application. For example, the multiple dimensions can be set according to actual application requirements. Again, in some cases, the multiple dimensions may include some or all of content understanding, logical reasoning, code writing, mathematical calculation, and video generation.

[0053] It can be seen that in a possible embodiment, the above model evaluation system may include evaluation benchmarks in multiple dimensions, and the multiple dimensions may include content understanding, logical reasoning, code writing, mathematical calculation, and video generation. Among them, the evaluation benchmark under content understanding is used to evaluate the content understanding performance of the machine learning model; the evaluation benchmark under logical reasoning is used to evaluate the logical reasoning performance of the machine learning model; the evaluation benchmark under code writing is used to evaluate the code writing performance of the machine learning model; the evaluation benchmark under mathematical calculation is used to evaluate the mathematical calculation performance of the machine learning model; the evaluation benchmark under video generation is used to evaluate the video generation performance of the machine learning model.

[0054] In addition, in order to better improve the evaluation effect, when the above model evaluation system includes evaluation benchmarks in multiple dimensions, for any dimension (such as content understanding), the evaluation benchmarks under this dimension may include evaluation benchmarks corresponding to multiple version numbers (such as Figure 3 the version number V 1 or the version number V 3 shown). And the evaluation benchmarks corresponding to different version numbers under the same dimension are different in the version of some or all of the content, and there are differences between the contents of different versions. The some or all of the content includes at least one of the evaluation data set, evaluation index, and evaluation method. For the convenience of understanding, the following will be described with examples.

[0055] As an example, when the above model evaluation system includes the evaluation benchmark Benchmark 1 under the first dimension, the evaluation benchmark Benchmark 2 under the second dimension,... (and so on), and the evaluation benchmark Benchmark N under the Nth dimension, where N is a positive integer, the evaluation benchmark Benchmark n under the nth dimension may include the evaluation benchmark Benchmark n .V1 、The benchmark corresponding to the second version number n .V 2 、……(and so on), and the benchmark corresponding to the Mth version number n .V M ; moreover, the benchmark corresponding to the mth version number n .V m differs from the benchmark corresponding to the (m + 1)th version number n .V m+1 in at least one content item (such as the evaluation dataset). For example, the evaluation dataset in this Benchmark n .V m+1 is obtained by making a certain change (such as data update, split ratio update, etc.) to the evaluation dataset in this Benchmark n .V m so that the version of the evaluation dataset in this Benchmark n .V m+1 is different from the version of the evaluation dataset in this Benchmark n .V m . Here, m is a positive integer, m + 1 ≤ M, M is a positive integer, n is a positive integer, and n ≤ N.

[0056] It can be seen that for the model evaluation system, the system can at least meet the following constraints: The system stores benchmarks in multiple dimensions; and the benchmarks in each dimension include benchmarks corresponding to at least one version number, so that the benchmarks stored in the system can not only meet the dimensional requirements of model evaluation processing, but also meet the version requirements of model evaluation processing, which is conducive to better improving the evaluation effect.

[0057] In addition, based on the model evaluation system shown in the previous paragraph, this application also provides a way to determine the "benchmark indicated by the evaluation request" above. In this way, when the model evaluation system includes benchmarks in multiple dimensions, and for any dimension, the benchmarks in this dimension include benchmarks corresponding to multiple version numbers, the determination process of the "benchmark indicated by the evaluation request" may include the following steps 11 - step 12.

[0058] Step 11: Determine the target dimension and the target version number according to the above evaluation request. The above multiple dimensions include the target dimension, and the multiple version numbers in the target dimension include the target version number, so that the target dimension can represent which aspect (such as content understanding) of the machine learning model needs to be evaluated, and the target version number can represent which version of the benchmark is selected in the target dimension.

[0059] It should be noted that the number of target dimensions in this application is not limited. For example, it can be multiple. Additionally, the method for determining the target dimension in this application is not limited. For example, the target dimension can be extracted from the evaluation request.

[0060] It should also be noted that the method for determining the target version number in this application is not limited. For example, to improve flexibility, if the evaluation request is automatically triggered by the model evaluation system at a preset cycle, then the target version number can be the version number with the latest creation time under the target dimension (that is, the latest version number); however, if the evaluation request is manually triggered by the user, then the evaluation request carries a version number specified by the user (such as the version number specified by the user for the target dimension), and the target version number can be the version number carried by the evaluation request (such as the aforementioned specified version number).

[0061] Step 12: According to the target dimension and the target version number, search for the evaluation benchmark indicated by the evaluation request in the storage space of the model evaluation system, so that the version number corresponding to the evaluation benchmark is the target version number, and the evaluation benchmark is used to evaluate the performance of each model among multiple machine learning models under the target dimension, so that the evaluation benchmark indicated by the evaluation request refers to the evaluation benchmark stored in the system that successfully matches both the target dimension and the target version number.

[0062] Based on the relevant content of Steps 11 to 12 above, it can be seen that when the model evaluation system includes evaluation benchmarks under multiple dimensions, and there is at least one version number for the evaluation benchmarks under each dimension, the dimension + version number can be selected according to actual needs, so that the evaluation benchmark determined based on the selection result is more suitable for participating in the model evaluation process under the actual needs.

[0063] In addition, to better improve the evaluation effect, this application also provides a possible implementation manner of the model evaluation system. In this manner, when the system includes evaluation benchmarks under multiple dimensions, and there is at least one version number for the evaluation benchmarks under each dimension, the system can also be used for: for any dimension, in response to an update operation triggered for at least one piece of content in the evaluation benchmark corresponding to the version number with the latest creation time (such as V 3 ), create a new version number (such as V 4 ) under this dimension, determine the evaluation benchmark corresponding to the new version number according to the update operation, and determine the creation reason corresponding to the new version number according to the update operation, so that the creation reason includes the content change indicated by the update operation. The at least one piece of content includes some or all of the evaluation dataset, evaluation metrics, and evaluation methods, so that automatic version control of Benchmark can be implemented in the system.

[0064] It should be noted that the present application does not limit the implementation manner of the update operation. For example, the update operation may include operations such as updating the name of the evaluation metric, modifying the metric caliber of the evaluation metric, updating the data in the evaluation dataset, and updating the splitting ratio in the evaluation dataset. Among them, the metric caliber is used to describe the unified specification of the metric definition, calculation method, and data source. The splitting ratio is used to indicate the proportion occupied by the training set, validation set, and test set in the evaluation dataset.

[0065] It should also be noted that the present application does not limit the manner of obtaining the above new version number. For example, the new version number may be determined based on the "version number with the latest creation time in this dimension" described above, so that the new version number is greater than the "version number with the latest creation time in this dimension", which can meet the requirement that the version number increases in chronological order.

[0066] Based on the above three paragraphs, it can be seen that as the evaluation dataset, evaluation metric, or evaluation method is updated and changed, the version of the evaluation benchmark in the model evaluation system can be automatically updated, so that the following relationships exist among different entities involved in the system in terms of version: If it is detected that the evaluation metric involved in updating the evaluation benchmark (such as a change in the metric name or a modification of the metric caliber, etc.), a new version of the evaluation metric and a new version of the evaluation benchmark including the evaluation metric are automatically added, so that the evaluation benchmark corresponding to the new version includes the updated evaluation metric; If it is detected that the evaluation dataset involved in updating the evaluation benchmark (such as a modification of the splitting ratio, etc.), a new version of the evaluation dataset and a new version of the evaluation benchmark including the evaluation dataset are automatically added, so that the evaluation benchmark corresponding to the new version includes the updated evaluation dataset; If it is detected that the evaluation method involved in updating the evaluation benchmark, a new version of the evaluation method and a new version of the evaluation benchmark including the evaluation dataset are automatically added, so that the evaluation benchmark corresponding to the new version includes the updated evaluation method. It can be seen that the Benchmark stored in the system will automatically add a new version when a new version of the evaluation metric, evaluation dataset, or evaluation method it includes is added, without manually adding a new version to the Benchmark, which is beneficial to improving the update efficiency.

[0067] It can be seen that the model evaluation system has the following characteristics for the automatic version control of Benchmark: for any dimension, a unique version number is assigned to each Benchmark under that dimension, and the version number increases in chronological order; the reason for each version update is recorded (such as data update in the evaluation dataset, modification of metric caliber, etc.); the relevant data of each version number is saved (such as evaluation dataset, evaluation metrics, evaluation methods, etc.) for subsequent backtracking and comparison. In addition, the system also provides a version comparison function to enable users to clearly see the differences between Benchmarks corresponding to different version numbers in a visual manner, so as to better understand the evolution process of Benchmarks under a certain dimension. In addition, when generating the list, it is necessary to clearly specify the version number of the Benchmark required for the list to ensure the accuracy and comparability of the evaluation results. It can be seen that through automatic version control, the system can better manage and track the changes of Benchmarks under each dimension, providing strong support for the evaluation and optimization of the model, which is conducive to better meeting various evaluation needs.

[0068] It is found through research that in order to better improve the evaluation effect, the list can be generated by means of evaluation benchmarks under at least two dimensions, so that the list can describe the advantages and disadvantages of different models as comprehensively as possible, thus effectively overcoming the defects caused when generating the list based on the evaluation benchmark under a single dimension.

[0069] Based on the above research, the present application also provides a generation method for the list. In this method, when the evaluation benchmarks indicated by the evaluation request include at least two evaluation benchmarks, the generation process of the list may include the following steps 21 - step 22.

[0070] Step 21: For any one of the above multiple machine learning models, analyze the list score of the model based on the evaluation results of the model under each evaluation benchmark in at least two evaluation benchmarks.

[0071] Among them, the list score of the j-th machine learning model is used to characterize the comprehensive performance of the j-th machine learning model in some dimensions (such as dimensions of content understanding, logical reasoning, code writing, mathematical calculation, and video generation, etc.) concerned by the list, where j is a positive integer and j ≤ J.

[0072] In addition, the present application does not limit the implementation manner of the above step 21. For example, it can be implemented by means of any data statistics method (such as the average value calculation method).

[0073] It can be seen that in a possible implementation manner, the above step 21 can specifically be: for any one of the above multiple machine learning models, calculate the average value of the evaluation results of the model under each evaluation benchmark in at least two evaluation benchmarks to obtain the list score of the model, so that the list constructed based on the list score can not only display the list score, but also display the evaluation results of the model under each evaluation benchmark, so as to better meet the information viewing needs of users. For the convenience of understanding, the following will be described with examples.

[0074] As an example, when the evaluation benchmarks indicated by the evaluation request include the evaluation benchmark Benchmark under content understanding 1 .V 1 , the evaluation benchmark Benchmark under logical reasoning 2 .V 1 , the evaluation benchmark Benchmark under code writing 3 .V 3 , the evaluation benchmark Benchmark under mathematical calculation 4 .V 2 , and the evaluation benchmark Benchmark under video generation 5 .V 4 , the average value between the evaluation result of the j-th machine learning model under Benchmark 1 .V 1 , the evaluation result of the j-th machine learning model under Benchmark 2 .V 1 , the evaluation result of the j-th machine learning model under Benchmark 3 .V 3 , the evaluation result of the j-th machine learning model under Benchmark 4 .V 2 , and the evaluation result of the j-th machine learning model under Benchmark 5 .V 4 can be used as the list score of the j-th machine learning model, where j is a positive integer and j ≤ J.

[0075] It has been found through research that since the performance to be concerned about may vary in different scenarios, different evaluation benchmarks in different dimensions need to be used to implement model evaluation processing in different scenarios. For example, since Scenario 1 (such as the video restoration scenario) needs to focus on two types of performance: logical reasoning + video generation, in this Scenario 1, the evaluation benchmark under logical reasoning + the evaluation benchmark under video generation need to be used to implement model evaluation processing. Another example is that since Scenario 2 (such as the code rewriting scenario) needs to focus on three types of model performance: mathematical calculation + logical reasoning + code writing, in this Scenario 2, the evaluation benchmark under mathematical calculation + the evaluation benchmark under logical reasoning + the evaluation benchmark under code writing need to be used to implement model evaluation processing. It can be seen that in order to better meet the foregoing requirements, the evaluation benchmarks stored in the model evaluation system can be pre-classified by scenario to obtain the scenario classification information of each evaluation benchmark, so that the scenario classification information can indicate which model evaluation tasks each evaluation benchmark is applicable to participate in, so as to be able to implement model evaluation processing in each scenario based on this scenario classification information subsequently.

[0076] It has also been found through research that when constructing a list by referring to the evaluation benchmarks in multiple scenarios at the same time, since different scenarios may show different influences on this list, in order to better meet this demand, the weights corresponding to each scenario can be pre-configured, so that the weight can indicate the influence degree presented by each scenario on this list.

[0077] Based on the above two paragraphs of research, in some cases (such as when comprehensively evaluating machine learning models in multiple scenarios), the above step 21 may include: for any one of the above multiple machine learning models, based on the scenario classification information of each evaluation benchmark in the above at least two evaluation benchmarks and the evaluation results of this model under each evaluation benchmark in the above at least two evaluation benchmarks, determine the evaluation results of this model in each scenario, and perform weighted summation processing on the evaluation results of this model in each scenario according to the weights corresponding to each scenario to obtain the list score of this model, so that the list constructed based on this list score can not only display this list score, but also display the evaluation results of this model in each scenario, in order to better meet the information viewing needs of users. For the convenience of understanding, an example is given below for illustration.

[0078] As an example, when the evaluation benchmark indicated by the evaluation request includes the evaluation benchmark Benchmark under content understanding 1 .V 1 , the evaluation benchmark Benchmark under logical reasoning 2 .V 1 , the evaluation benchmark Benchmark under code writing 3 .V 3 , the evaluation benchmark Benchmark under mathematical calculation 4 .V2 and the Benchmark for evaluation under video generation 5 .V 4 , Benchmark 1 .V 1 and Benchmark 5 .V 4 all belong to Scenario 3 (such as general scenarios), and Benchmark 2 .V 1 , Benchmark 3 .V 3 and Benchmark 4 .V 2 when all belong to Scenario 4 (such as code-related scenarios), the process for determining the leaderboard score of the j-th machine learning model is as follows:

[0079] First, based on the scenario classification information of each Benchmark in the Benchmark indicated by the evaluation request, determine the classification results, that is, Benchmark 1 .V 1 and Benchmark 5 .V 4 all belong to Scenario 3, and Benchmark 2 .V 1 , Benchmark 3 .V 3 and Benchmark 4 .V 2 all belong to Scenario 4 for these classification results;

[0080] Then, take the average between the evaluation result of the j-th machine learning model under Benchmark 1 .V 1 and the evaluation result of this j-th machine learning model under Benchmark 5 .V 4 as the evaluation result of the j-th machine learning model under Scenario 3, so that this evaluation result can represent the performance presented by the j-th machine learning model under Scenario 3, and the evaluation result of the j-th machine learning model under Benchmark 2 .V 1 , the evaluation result of this j-th machine learning model under Benchmark 3 .V 3 and the evaluation result of this j-th machine learning model under Benchmark 4 .V 2The average value among the evaluation results under the following is used as the evaluation result of the j-th machine learning model in Scenario 4, so that the evaluation result can represent the performance presented by the j-th machine learning model in Scenario 4;

[0081] Finally, according to the weights corresponding to Scenario 3 and the weights corresponding to Scenario 4, a weighted sum processing is performed on the evaluation result of the j-th machine learning model in Scenario 3 and the evaluation result of the j-th machine learning model in Scenario 4 to obtain the list score of the j-th machine learning model, so that the list score can represent the performance comprehensively presented by the j-th machine learning model in multiple scenarios. j is a positive integer, j ≤ J. In this way, the list constructed based on the list score can more accurately represent which machine learning model presents better performance in the case of multi-scenario fusion, so as to better meet the performance requirements in the case of multi-scenario fusion.

[0082] It has been found through research that in the case of constructing a list based on at least one evaluation benchmark, in order to better meet the requirement that different evaluation benchmarks have different influence strengths on the list, the weights corresponding to each evaluation benchmark stored in the model evaluation system can be pre-configured, so that the weight can represent the influence strength presented by each evaluation benchmark during the construction of the list.

[0083] Based on the above research, in a possible implementation manner, step 21 above may include: for any one of the above-mentioned multiple machine learning models, a weighted sum processing is performed on the evaluation results of the model under each of the at least two evaluation benchmarks according to the weights configured for each evaluation benchmark in the at least two evaluation benchmarks, to obtain the list score of the model, so that the list constructed based on the list score can not only display the list score, but also display the weights configured for each evaluation benchmark, so as to better meet the user's information viewing requirements. For ease of understanding, the following is illustrated with examples.

[0084] As an example, when the evaluation benchmarks indicated by the evaluation request include the evaluation benchmark Benchmark under content understanding 1 .V 1 , the evaluation benchmark Benchmark under logical reasoning 2 .V 1 , the evaluation benchmark Benchmark under code writing 3 .V 3 , the evaluation benchmark Benchmark under mathematical calculation 4 .V 2 , and the evaluation benchmark Benchmark under video generation 5 .V 4 , it is possible to follow Benchmark 1 .V 1Configured weights, Benchmark 2 .V 1 Configured weights, Benchmark 3 .V 3 Configured weights, Benchmark 4 .V 2 Configured weights, and Benchmark 5 .V 4 The configured weights, for the j-th machine learning model in Benchmark 1 .V 1 The evaluation results under it, the evaluation results of the j-th machine learning model in Benchmark 2 .V 1 The evaluation results under it, the evaluation results of the j-th machine learning model in Benchmark 3 .V 3 The evaluation results under it, the evaluation results of the j-th machine learning model in Benchmark 4 .V 2 The evaluation results under it, and the evaluation results of the j-th machine learning model in Benchmark 5 .V 4 The evaluation results under it are weighted and summed to obtain the list score of the j-th machine learning model, where j is a positive integer and j ≤ J.

[0085] It has been found through research that due to the data distributions presented by the evaluation results of multiple machine learning models determined based on evaluation benchmarks in different dimensions, there may be significant differences (such as large differences between peaks and valleys), which may interfere with the list. Therefore, to overcome this interference, the rankings of each model under each evaluation benchmark can be used instead of the evaluation results of each model under each evaluation benchmark to participate in the construction of the list.

[0086] Based on the above research, in a possible implementation manner, step 21 above can specifically be: for any one of the above multiple machine learning models, based on the evaluation results of the model under each evaluation benchmark in at least two evaluation benchmarks, determine the rankings of the model under each evaluation benchmark in at least two evaluation benchmarks, and calculate the average value of the rankings of the model under each evaluation benchmark in at least two evaluation benchmarks to obtain the list score of the model, so that the list constructed based on this list score can not only display the list score but also display the rankings of the model under each evaluation benchmark, to better meet the information viewing needs of users. For ease of understanding, an example is given below for illustration.

[0087] As an example, when the evaluation benchmark indicated by the evaluation request includes the evaluation benchmark Benchmark under content understanding 1 .V 1, Benchmark under logical reasoning 2 .V 1 , Benchmark under code writing 3 .V 3 , Benchmark under mathematical calculation 4 .V 2 , and Benchmark under video generation 5 .V 4 When, the determination process of the list scores of the above multiple machine learning models includes:

[0088] First, sort the evaluation results of each machine learning model under Benchmark 1 .V 1 to obtain the rankings of each machine learning model under Benchmark 1 .V 1 ; sort the evaluation results of each machine learning model under Benchmark 2 .V 1 to obtain the rankings of each machine learning model under Benchmark 2 .V 1 ; sort the evaluation results of each machine learning model under Benchmark 3 .V 3 to obtain the rankings of each machine learning model under Benchmark 3 .V 3 ; sort the evaluation results of each machine learning model under Benchmark 4 .V 2 to obtain the rankings of each machine learning model under Benchmark 4 .V 2 ; sort the evaluation results of each machine learning model under Benchmark 5 .V 4 to obtain the rankings of each machine learning model under Benchmark 5 .V 4 ;

[0089] Then, for any machine learning model, the ranking of this model under Benchmark 1 .V 1 , the ranking of this model under Benchmark 2 .V 1 , the ranking of this model under Benchmark 3 .V 3 , the ranking of this model under Benchmark 4 .V2 rank under and the ranking of the model in Benchmark 5 .V 4 The average value between the rankings is used as the leaderboard score of the model.

[0090] Based on the relevant content of step 21 above, for any one of multiple machine learning models, after obtaining the evaluation results of the model under each evaluation benchmark in at least two evaluation benchmarks, these evaluation results can be aggregated in a certain way to obtain the leaderboard score of the model, so that the leaderboard constructed based on the leaderboard score can better meet the performance evaluation requirements of the current leaderboard.

[0091] Step 22: Sort the multiple machine learning models according to the above leaderboard scores to obtain a leaderboard.

[0092] Based on the relevant content of steps 21 to 22 above, for the model evaluation system provided in this application, the system can provide multiple calculation methods for leaderboard rankings, so that the system can better meet the model evaluation processing in different situations, so that relevant personnel can then select the calculation method of the leaderboard ranking according to actual needs, which is conducive to better improving the automation effect of the evaluation.

[0093] It has been found through research that for the current evaluation process, some or all of the tasks involved in this evaluation process may have been executed in historical time. Therefore, in order to better reduce the overhead, the historical execution results can be directly reused.

[0094] Based on the above research, in order to better reduce the resource overhead, this application also provides a possible implementation manner of the model evaluation method. In this manner, the model evaluation method may include steps 31 to 33 below.

[0095] Step 31: In response to an evaluation request triggered by the model evaluation system, obtain multiple machine learning models indicated by the evaluation request, and obtain the evaluation benchmarks indicated by the evaluation request. At least one evaluation benchmark stored in the system includes the evaluation benchmark indicated by the evaluation request.

[0096] It should be noted that for the relevant content of step 31, please refer to the relevant content of S1 above.

[0097] Step 32: For any one of the multiple machine learning models, check the historical execution records of the model evaluation system to see if there is an execution record of the evaluation task of this model under the evaluation benchmark indicated by the evaluation request; if so, query the execution result of this evaluation task from the storage space of this system to obtain the evaluation result of this model under the evaluation benchmark indicated by the evaluation request; however, if not, create and execute this evaluation task to obtain the evaluation result of this model under the evaluation benchmark indicated by the evaluation request.

[0098] In this application, for the current evaluation process, if it is determined that the performance of the j-th machine learning model needs to be evaluated, it is possible to first check the historical execution records of the model evaluation system to see if there is an execution record of the evaluation task of this j-th machine learning model under the evaluation benchmark indicated by the evaluation request. If so, it can be determined that this system has already executed and cached this evaluation task, so it can be determined that this evaluation task belongs to the tasks cached within this system. Therefore, to save resources, the execution result of this evaluation task can be directly queried from the storage space (such as the cache space) of this system; however, if not, it can be determined that this system has not yet executed this evaluation task, so it can be determined that this evaluation task does not belong to the tasks cached within this system. Therefore, it is necessary to create, execute, and cache this evaluation task to obtain and cache the evaluation result of this model under the evaluation benchmark indicated by the evaluation request for subsequent use. In this way, a list generation process with a caching mechanism can be implemented in this system, effectively reducing the overhead of the list generation process, such as time overhead + resource overhead, etc. It should be noted that this application does not limit the acquisition method of the above historical execution records. For example, it can be read from the storage space (such as the cache space) of the model evaluation system.

[0099] Step 33: Generate a list based on the evaluation results of each model among the multiple machine learning models under the evaluation benchmark indicated by the evaluation request. This list is used to describe the ranking of each model among the multiple machine learning models in terms of performance.

[0100] It should be noted that for the relevant content of Step 33, please refer to the relevant content of S2 above.

[0101] Based on the relevant content from Step 31 to Step 33, for the model evaluation system, when it is used to automatically generate a leaderboard on schedule, if it is detected that the latest version numbers of the evaluation benchmarks involved in the leaderboard have not been updated, it can be determined that the evaluation benchmarks used to construct the leaderboard have not changed. Therefore, it is possible to check in the cache whether there are evaluation tasks involved in the leaderboard to ensure that tasks that have been cached do not initiate new tasks, but tasks that are not cached need to initiate new tasks, so that after their execution, their relevant data can be recorded in the cache, so that after it is determined that all evaluation tasks involved in the leaderboard are completed, the leaderboard is constructed and displayed according to the leaderboard ranking calculation method.

[0102] In addition, when the model evaluation system is used to generate a leaderboard according to a manually triggered request, it is possible to check in the cache whether there are evaluation tasks involved in the leaderboard to ensure that tasks that have been cached do not initiate new tasks, but tasks that are not cached need to initiate new tasks, so that after their execution, their relevant data can be recorded in the cache, so that after it is determined that all evaluation tasks involved in the leaderboard are completed, the leaderboard is constructed and displayed according to the leaderboard ranking calculation method.

[0103] It has been found through research that in some cases, users may adjust the machine learning models involved in the leaderboard. Therefore, in order to better improve the evaluation effect, the present application also provides a possible implementation manner of the above Step 32, which can specifically be: for any one of the above multiple machine learning models, in response to an update of the model, create and execute an evaluation task of the model under the evaluation benchmark indicated by the evaluation request to obtain the evaluation result of the model under the evaluation benchmark indicated by the evaluation request; in response to the model not being updated, check in the historical execution record of the model evaluation system whether there is an execution record of the evaluation task; if there is, query the execution result of the evaluation task from the storage space of the system to obtain the evaluation result of the model under the evaluation benchmark indicated by the evaluation request; if not, create and execute the evaluation task to obtain the evaluation result of the model under the evaluation benchmark indicated by the evaluation request.

[0104] It can be seen that for the current list construction process, first determine whether the j-th machine learning model involved in the list has been updated. If it has been updated, it can be determined that the evaluation task of the j-th machine learning model under the evaluation benchmark indicated by the evaluation request needs to be re-executed. Therefore, it is necessary to create, execute, and cache the evaluation task, and obtain and cache the evaluation result of the model under the evaluation benchmark indicated by the evaluation request for subsequent use. However, if there is no update, check whether the evaluation task exists in the cache. If it exists, no new task needs to be initiated. If it does not exist, a new task needs to be initiated so that the relevant data can be recorded in the cache after its execution. j is a positive integer and j ≤ J. In this way, the overhead can be reduced on the premise of meeting the evaluation requirements of the model iteration process, which is conducive to improving the evaluation effect.

[0105] It has been found through research that in some cases, the source of data in the evaluation dataset may not be unique (for example, the data can come from a local file storage system, a cloud document storage system, or a distributed file storage system, etc.). This makes it a difficult problem to construct an evaluation dataset using data from different sources.

[0106] Based on the above research, in order to overcome the above problems, the present application also provides a possible implementation manner of the model evaluation system. In this manner, the system includes an import tool, which is used to import files from different storage systems into the evaluation dataset stored in the system. The working principle of the import tool is as follows: in response to a data import request triggered through the import tool, first determine the meta-information (such as column names, column types, etc.) based on the file or file storage path carried by the data import request, so that the meta-information is used to describe the file indicated by the data import request, thereby enabling the meta-information to represent the characteristics (such as source, format, structure, content, etc.) of the file indicated by the request; then create a data table (such as an OLAP table) in the system based on the meta-information, so that the data table is used to describe the file indicated by the data import request according to the data structure of each data in the evaluation dataset (such as the data structure of the OLAP table); then, use the data table to update the evaluation dataset stored in the system, so that the updated evaluation dataset includes the data table. In this way, it is possible to use the import tool to import data from different sources into the same evaluation dataset, so that subsequent processing such as visual display, statistical analysis, and list construction can be performed on the evaluation dataset through the system.

[0107] It should be noted that the present application does not limit the implementation manner of "determining meta-information based on the file or file storage path carried by the data import request". For example, specifically, it may be: first, parse the data import request to obtain a parsing result, so that the parsing result includes the file or file storage path; then, perform meta-information extraction processing based on the parsing result to obtain the meta-information of the file indicated by the data import request.

[0108] Based on the above two paragraphs, for the evaluation data set stored in the model evaluation system, the output and synchronization of the evaluation data set can be realized by means of the import tool in the system, so that the tool can import files from different storage systems into the evaluation data set. In addition, the tool can create an offline import task for the data import request to implement different import operations according to the storage system types of different files, so as to ensure that the data in the file or path provided by the user is successfully saved to the OLAP of the system. In addition, the import process realized by means of the tool can be: first, parse the file or file storage path provided by the user through the request, and extract meta-information based on the parsing result and create an OLAP table in the system based on the meta-information; then, generate an offline import task based on the OLAP table, so that the task is used to read the file or path provided by the user, and save the read information to the system through a series of conversions and cleanings, so that subsequent statistical analysis, visual display, and building of a leaderboard and other processes can be performed based on the data saved in the system. In addition, the import tool is also used to inform the user of the data import situation through a certain notification method after detecting that the offline import task is completed, so as to achieve real-time performance.

[0109] Based on the relevant content of the model evaluation method applied to the model evaluation system, the present application has realized a complete model evaluation leaderboard system in the big data scenario, optimized some components of the leaderboard and Benchmark, and provided a better process system for model evaluation. In addition, the present application has also realized a solution for automatic version control of Benchmark, and based on this solution, has realized a leaderboard output process with a caching mechanism, effectively reducing the evaluation overhead (such as time overhead + resource overhead) in the leaderboard output process. In addition, the present application has proposed some leaderboard ranking calculation methods (such as average ranking by scenario, weighted summation ranking, average ranking, etc.), so as to effectively meet different leaderboard requirements. It can be seen that the model evaluation system provided by the present application can provide a clear process system for model evaluation, provide more comprehensive support for model iteration comparison, effectively improve the automation degree of evaluation, reduce human input, shorten the leaderboard output time, and reduce resource overhead.

[0110] Based on the model evaluation method provided in the embodiments of the present application, the embodiments of the present application also provide a model evaluation system, which will be explained and described below in conjunction with Figure 4 wherein Figure 4 is a schematic structural diagram of a model evaluation system provided in the embodiments of the present application. It should be noted that for the technical details of the model evaluation system provided in the embodiments of the present application, please refer to the relevant content of the above model evaluation method.

[0111] As Figure 4 shown, the model evaluation system 400 provided in the embodiments of the present application, at least one pre-configured evaluation benchmark is stored in the system 400, and the evaluation benchmark includes an evaluation data set, an evaluation index, and an evaluation method. The evaluation benchmark is used to evaluate the performance of a machine learning model. The system 400 includes:

[0112] An acquisition unit 401, configured to, in response to an evaluation request triggered by the system 400, acquire a plurality of machine learning models indicated by the evaluation request, and acquire an evaluation benchmark indicated by the evaluation request. The at least one evaluation benchmark includes the evaluation benchmark indicated by the evaluation request;

[0113] A generation unit 402, configured to generate a list according to the evaluation results of each model in the plurality of machine learning models under the evaluation benchmark indicated by the evaluation request. The list is used to describe the performance ranking of each model in the plurality of machine learning models.

[0114] In a possible implementation manner, the system 400 includes evaluation benchmarks in multiple dimensions. The evaluation benchmarks in different dimensions are used to evaluate the performance of the machine learning model in different dimensions. For any dimension, the evaluation benchmarks in this dimension include evaluation benchmarks corresponding to multiple version numbers. There are differences in the versions of some or all of the content between the evaluation benchmarks corresponding to different version numbers in the same dimension. The part or all of the content includes at least one of an evaluation data set, an evaluation index, and an evaluation method. The acquisition unit 401 is specifically configured to: determine a target dimension and a target version number according to the evaluation request. The multiple dimensions include the target dimension, and the multiple version numbers under the target dimension include the target version number; according to the target dimension and the target version number, search for the evaluation benchmark indicated by the evaluation request in the storage space of the system. The version number corresponding to the evaluation benchmark indicated by the evaluation request is the target version number, and the evaluation benchmark indicated by the evaluation request is used to evaluate the performance of each model in the plurality of machine learning models in the target dimension.

[0115] In a possible implementation, if the evaluation request is automatically triggered by the system at a preset period, the target version number is the version number with the latest creation time in the target dimension; if the evaluation request is manually triggered by the user, the evaluation request carries the version number specified by the user, and the target version number is the version number carried by the evaluation request.

[0116] In a possible implementation, the system 400 includes evaluation benchmarks in multiple dimensions. For any dimension, the evaluation benchmarks in this dimension include evaluation benchmarks corresponding to multiple version numbers; the system 400 further includes: a creation unit, for any dimension, in response to an update operation triggered for at least one content in the evaluation benchmark corresponding to the version number with the latest creation time in this dimension, create a new version number in this dimension, determine the evaluation benchmark corresponding to the new version number according to the update operation, and determine the creation reason corresponding to the new version number according to the update operation, where the creation reason includes the content change indicated by the update operation, and the at least one content includes some or all of the evaluation data set, evaluation index, and evaluation method.

[0117] In a possible implementation, the evaluation benchmarks indicated by the evaluation request include at least two evaluation benchmarks; the generating unit 402 is specifically configured to: for any one of the multiple machine learning models, analyze the list score of the model according to the evaluation results of the model under each evaluation benchmark in the at least two evaluation benchmarks; sort the multiple machine learning models according to the list score to obtain the list.

[0118] In a possible implementation, the generating unit 402 is specifically configured to: for any one of the multiple machine learning models, calculate the average value of the evaluation results of the model under each evaluation benchmark in the at least two evaluation benchmarks to obtain the list score of the model.

[0119] In a possible implementation, the generating unit 402 is specifically configured to: for any one of the multiple machine learning models, determine the evaluation results of the model in each scenario according to the scenario classification information of each evaluation benchmark in the at least two evaluation benchmarks and the evaluation results of the model under each evaluation benchmark in the at least two evaluation benchmarks, and perform a weighted summation process on the evaluation results of the model in each scenario according to the weights corresponding to each scenario to obtain the list score of the model.

[0120] In a possible implementation manner, the generating unit 402 is specifically configured to: for any one of the multiple machine learning models, perform a weighted summation process on the evaluation results of the model under each of the at least two evaluation benchmarks according to the weights configured for each of the at least two evaluation benchmarks, to obtain the list score of the model.

[0121] In a possible implementation manner, the generating unit 402 is specifically configured to: for any one of the multiple machine learning models, determine the ranking of the model under each of the at least two evaluation benchmarks according to the evaluation results of the model under each of the at least two evaluation benchmarks, and calculate the average value of the rankings of the model under each of the at least two evaluation benchmarks, to obtain the list score of the model.

[0122] In a possible implementation manner, the system 400 further includes: a determining unit, configured to, for any one of the multiple machine learning models, check whether there is an execution record of an evaluation task of the model under the evaluation benchmark indicated by the evaluation request in the historical execution record of the system; if there is, query the execution result of the evaluation task from the storage space of the system, to obtain the evaluation result of the model under the evaluation benchmark indicated by the evaluation request; if not, create and execute the evaluation task, to obtain the evaluation result of the model under the evaluation benchmark indicated by the evaluation request.

[0123] In a possible implementation manner, the determining unit is specifically configured to: in response to an update of the model, create and execute the evaluation task, to obtain the evaluation result of the model under the evaluation benchmark indicated by the evaluation request; in response to no update of the model, check whether there is an execution record of an evaluation task of the model under the evaluation benchmark indicated by the evaluation request in the historical execution record of the system.

[0124] In a possible implementation manner, the system includes an import tool, and the import tool is configured to import files of different storage systems into the evaluation data set stored in the system; the import process includes: in response to a data import request triggered by the import tool, determine meta information according to the file or file storage path carried by the data import request, where the meta information is used to describe the file indicated by the data import request; create a data table in the system according to the meta information, where the data table is used to describe the file indicated by the data import request according to the data structure of each data in the evaluation data set; and update the evaluation data set stored in the system by using the data table.

[0125] In a possible implementation, for any one of the above evaluation benchmarks, the evaluation metrics of the evaluation benchmark include a main metric and an auxiliary metric. The main metric is used to affect the ranking, and the auxiliary metric is used to provide auxiliary information for the list.

[0126] Based on the relevant content of the above model evaluation system 400, it can be known that at least one pre-configured evaluation benchmark is stored in the system 400, and the working principle of the system includes: in response to an evaluation request triggered by the system, first obtain multiple machine learning models indicated by the request and obtain the evaluation benchmark indicated by the request, so as to be able to use the evaluation benchmark indicated by the request to evaluate the performance of each model later; then generate a list according to the evaluation results of each model under the evaluation benchmark indicated by the request, so that the list is used to describe the ranking of each model in terms of performance, so that the list can better represent the relative advantages and disadvantages among these models, and then be able to automatically evaluate and compare the advantages and disadvantages among these models, so as to effectively overcome the defects caused by users manually participating in the evaluation and comparison process of these models.

[0127] In addition, an embodiment of the present application also provides an electronic device, where the device includes a processor and a memory: the memory is used to store instructions or computer programs; the processor is used to execute the instructions or computer programs in the memory so that the electronic device executes any implementation manner of the model evaluation method provided by the embodiment of the present application.

[0128] See Figure 5 , which shows a schematic structural diagram of an electronic device 500 suitable for implementing the embodiments of the present disclosure. The terminal devices in the embodiments of the present disclosure may include, but are not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Tablet Computers), PMPs (Portable Multimedia Players), in-vehicle terminals (such as in-vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 5 The electronic device shown is only an example and should not bring any limitation to the functions and usage scopes of the embodiments of the present disclosure.

[0129] As Figure 5As shown, the electronic device 500 may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 501, which may perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 502 or a program loaded from a storage device 508 into a random access memory (RAM) 503. In the RAM 503, various programs and data required for the operation of the electronic device 500 are also stored. The processing device 501, the ROM 502, and the RAM 503 are connected to each other through a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.

[0130] Generally, the following devices may be connected to the I / O interface 505: an input device 506 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 507 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 508 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 509. The communication device 509 may allow the electronic device 500 to communicate with other devices wirelessly or wiredly to exchange data. Although Figure 5 an electronic device 500 with various devices is shown, it should be understood that it is not required to implement or have all the shown devices. Instead, more or fewer devices may be implemented or had.

[0131] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart may be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program includes program codes for performing the method shown in the flowchart. In such an embodiment, the computer program may be downloaded and installed from a network through the communication device 509, or installed from the storage device 508, or installed from the ROM 502. When the computer program is executed by the processing device 501, the above-mentioned functions defined in the method of the embodiment of the present disclosure are executed.

[0132] The electronic device provided by the embodiment of the present disclosure and the method provided by the above embodiment belong to the same inventive concept. Technical details not described in detail in this embodiment may be referred to the above embodiment, and this embodiment has the same beneficial effects as the above embodiment.

[0133] An embodiment of the present application also provides a computer-readable medium, in which instructions or a computer program are stored. When the instructions or the computer program run on a device, the device is caused to execute any implementation manner of the model evaluation method provided by the embodiment of the present application.

[0134] It should be noted that the computer-readable medium described above can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, the computer-readable storage medium can be any tangible medium that contains or stores a program, which can be used by or in conjunction with an instruction execution system, apparatus, or device. In the present disclosure, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.

[0135] In some embodiments, the client and the server can communicate using any currently known or future-developed network protocol such as HTTP (Hyper Text Transfer Protocol), and can be interconnected with digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet), and end-to-end networks (e.g., ad hoc end-to-end networks), as well as any currently known or future-developed networks.

[0136] The above computer-readable medium can be included in the above electronic device; or it can exist separately and not be assembled into the electronic device.

[0137] The above computer-readable medium carries one or more programs, and when the one or more programs are executed by the electronic device, the electronic device can execute the above method.

[0138] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages or combinations thereof. The programming languages include, but are not limited to, object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).

[0139] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a part of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system for performing the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.

[0140] The units involved in the embodiments described in the present disclosure may be implemented in software or in hardware. Among them, the name of the unit / module does not, in some cases, constitute a limitation on the unit itself.

[0141] The functions described above herein may be performed, at least in part, by one or more hardware logic components. For example, by way of non-limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGA), application specific integrated circuits (ASIC), application specific standard products (ASSP), system on a chip (SOC), complex programmable logic devices (CPLD), and so on.

[0142] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. The machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium would include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0143] It should be noted that the various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts among the various embodiments can be referred to each other. For the systems or apparatuses disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the description of the method part.

[0144] It should be understood that in this application, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects and indicates that three relationships can exist. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist at the same time. Among them, A and B can be singular or plural. The character " / " generally indicates that the associated objects before and after are in an "or" relationship. "At least one (one) of the following" or its similar expression refers to any combination of these items, including any combination of single item (one) or plural items (ones). For example, at least one (one) of a, b, or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.

[0145] It should also be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprising", "including" or any other variant thereof are intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements not only includes those elements but also includes other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising said element.

[0146] The steps of the methods or algorithms described in connection with the embodiments disclosed herein may be implemented directly in hardware, in software modules executed by a processor, or in a combination thereof. The software modules may be placed in a random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium well-known in the art.

[0147] The foregoing description of the disclosed embodiments enables those skilled in the art to implement or use the present application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present application. Thus, the present application is not intended to be limited to the embodiments shown herein but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A model evaluation method, characterized in that: The method is applied to a model evaluation system, wherein the system stores at least one pre-configured evaluation benchmark, wherein the evaluation benchmark includes an evaluation data set, an evaluation index, and an evaluation method, and wherein the evaluation benchmark is used to evaluate the performance of a machine learning model. The method includes: In response to an evaluation request triggered by the system, obtaining a plurality of machine learning models indicated by the evaluation request, and obtaining an evaluation benchmark indicated by the evaluation request, wherein the at least one evaluation benchmark includes the evaluation benchmark indicated by the evaluation request; Based on the evaluation results of each model in the multiple machine learning models under the evaluation benchmark indicated by the evaluation request, a list is generated, where the list is used to describe the performance ranking of each model in the multiple machine learning models.

2. The method according to claim 1, characterized in that The system includes evaluation benchmarks in multiple dimensions, and the evaluation benchmarks in different dimensions are used to evaluate the performance of the machine learning model in different dimensions. For any dimension, the evaluation benchmark in the dimension includes evaluation benchmarks corresponding to multiple version numbers, and the evaluation benchmarks corresponding to different version numbers in the same dimension have differences in the versions of part or all of the content, and the part or all of the content includes at least one of the evaluation data set, evaluation indicators and evaluation methods; The process of obtaining the evaluation benchmark indicated by the evaluation request includes: Determining a target dimension and a target version number according to the evaluation request, wherein the multiple dimensions include the target dimension, and the multiple version numbers under the target dimension include the target version number; According to the target dimension and the target version number, the evaluation benchmark indicated by the evaluation request is searched from the storage space of the system, the version number corresponding to the evaluation benchmark indicated by the evaluation request is the target version number, and the evaluation benchmark indicated by the evaluation request is used to evaluate the performance of each of the multiple machine learning models under the target dimension.

3. The method according to claim 2, characterized in that If the evaluation request is automatically triggered by the system according to a preset period, the target version number is the version number with the latest creation time under the target dimension; If the evaluation request is manually triggered by a user, the evaluation request carries a version number specified by the user, and the target version number is the version number carried by the evaluation request.

4. The method according to claim 1, characterized in that: The system includes evaluation benchmarks under multiple dimensions. For any dimension, the evaluation benchmark under the dimension includes evaluation benchmarks corresponding to multiple version numbers; The method further comprises: For any dimension, in response to an update operation triggered by at least one content in an evaluation benchmark corresponding to a version number with the latest creation time under the dimension, a new version number is created under the dimension, the evaluation benchmark corresponding to the new version number is determined based on the update operation, and the creation reason corresponding to the new version number is determined based on the update operation, the creation reason includes the content change indicated by the update operation, and the at least one content includes part or all of the evaluation data set, evaluation indicators and evaluation methods.

5. The method according to claim 1, characterized in that The evaluation benchmark indicated by the evaluation request includes at least two evaluation benchmarks; The process of generating the list includes: For any model among the multiple machine learning models, analyzing the list score of the model according to the evaluation results of the model under each of the at least two evaluation benchmarks; The multiple machine learning models are sorted according to the list scores to obtain the list.

6. The method according to claim 5, characterized in that For any model among the multiple machine learning models, a process of determining a list score of the model includes: Calculate the average value of the evaluation results of the model under each of the at least two evaluation benchmarks to obtain the list score of the model; or, Determine the evaluation results of the model in each scenario according to the scenario classification information of each of the at least two evaluation benchmarks and the evaluation results of the model in each of the at least two evaluation benchmarks, and perform weighted sum processing on the evaluation results of the model in each scenario according to the weight corresponding to each scenario to obtain the list score of the model; or, According to the weights of each evaluation benchmark configuration in the at least two evaluation benchmarks, weighted sum processing is performed on the evaluation results of the model under each evaluation benchmark in the at least two evaluation benchmarks to obtain the list score of the model; or, According to the evaluation results of the model under each of the at least two evaluation benchmarks, the ranking of the model under each of the at least two evaluation benchmarks is determined, and the average of the rankings of the model under each of the at least two evaluation benchmarks is calculated to obtain the list score of the model.

7. The method according to claim 1, characterized in that After obtaining the plurality of machine learning models indicated by the evaluation request and obtaining the evaluation benchmark indicated by the evaluation request, the method further includes: For any model among the multiple machine learning models, searching from the historical execution records of the system whether there is an execution record of the evaluation task of the model under the evaluation benchmark indicated by the evaluation request; If so, query the execution result of the evaluation task from the storage space of the system to obtain the evaluation result of the model under the evaluation benchmark indicated by the evaluation request; If it does not exist, then create and execute the evaluation task to obtain the evaluation result of the model under the evaluation benchmark indicated by the evaluation request.

8. The method according to claim 7, characterized in that The method further comprises: In response to the model being updated, creating and executing the evaluation task to obtain an evaluation result of the model under the evaluation benchmark indicated by the evaluation request; The searching from the historical execution records of the system whether there is an execution record of the evaluation task of the model under the evaluation benchmark indicated by the evaluation request includes: In response to the model not being updated, searching the historical execution records of the system to see whether there is an execution record of the evaluation task of the model under the evaluation benchmark indicated by the evaluation request.

9. The method according to claim 1, characterized in that: The system includes an import tool, which is used to import files from different storage systems into the evaluation data set stored in the system; The import process includes: In response to a data import request triggered by the import tool, determining meta information according to a file or a file storage path carried by the data import request, the meta information being used to describe the file indicated by the data import request; Creating a data table in the system according to the meta-information, the data table being used to describe the file indicated by the data import request according to the data structure of each data in the evaluation data set; The data table is used to update the evaluation data set stored in the system.

10. The method according to any one of claims 1 to 9, characterized in that: For any of the evaluation benchmarks, the evaluation indicators of the evaluation benchmark include main indicators and auxiliary indicators, the main indicators are used to influence the ranking, and the auxiliary indicators are used to provide auxiliary information for the list.

11. A model evaluation system, characterized in that: The system stores at least one pre-configured evaluation benchmark, which includes an evaluation data set, an evaluation index, and an evaluation method. The evaluation benchmark is used to evaluate the performance of a machine learning model. The system includes: an acquisition unit, configured to, in response to an evaluation request triggered by the system, acquire a plurality of machine learning models indicated by the evaluation request, and acquire an evaluation benchmark indicated by the evaluation request, wherein the at least one evaluation benchmark includes the evaluation benchmark indicated by the evaluation request; A generating unit is used to generate a list based on the evaluation results of each model in the multiple machine learning models under the evaluation benchmark indicated by the evaluation request, wherein the list is used to describe the performance ranking of each model in the multiple machine learning models.

12. An electronic device, characterized in that: The device comprises: a processor and a memory; The memory is used to store instructions or computer programs; The processor is used to execute the instructions or computer programs in the memory so that the electronic device executes the method according to any one of claims 1 to 10.

13. A computer readable medium, characterized in that The computer-readable medium stores instructions or computer programs, and when the instructions or computer programs are executed on a device, the device executes the method according to any one of claims 1 to 10.

14. A computer program product, characterized in that It comprises a computer program carried on a non-transitory computer-readable medium, the computer program comprising a program code for executing the method according to any one of claims 1 to 10.