Data processing method and device based on artificial intelligence model and related equipment

By dividing the use cases to be evaluated into multiple sharded tasks, processing and testing in parallel, the problem of long time consumption in the process of artificial intelligence model evaluation is solved, and more efficient model evaluation is achieved.

CN120386714APending Publication Date: 2025-07-29TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410104513.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-24
Publication Date
2025-07-29

AI Technical Summary

Technical Problem

In the model evaluation process of artificial intelligence models in the prior art, the use case processing in the evaluation data set is carried out serially, resulting in a long time consumption, and the test of data processing results needs to wait for all use cases to complete, which reduces the efficiency of model evaluation.

Method used

By dividing the use cases to be evaluated into multiple shard tasks, processing each shard task in parallel, using multiple models to process the use case data in parallel, and testing is carried out after the processing results of a single use case data is released, reducing the waiting time.

Benefits of technology

The model evaluation efficiency of artificial intelligence models has been improved, and the overall processing time has been reduced and the efficiency of data processing has been improved through parallel processing and timely testing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120386714A_ABST
    Figure CN120386714A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a data processing method and device based on an artificial intelligence model and related equipment, and can be applied to the technical field of data processing. The method comprises the steps of receiving a model evaluation task; determining K to-be-evaluated cases based on the model evaluation task; creating M model processing threads associated with the M fragmentation tasks based on the parallel processing quantity, and dividing the K to-be-evaluated cases into the M fragmentation tasks to obtain fragmentation task processing queues of the M fragmentation tasks; determining a first to-be-assessed case in to-be-assessed cases contained in a fragmentation task processing queue of the fragmentation task i, inputting the first to-be-assessed case into a to-be-assessed model through a model processing thread j, and performing data processing on the first to-be-assessed case by the to-be-assessed model to obtain a data processing result; and performing test processing on the data processing result based on the test processing thread to obtain a use case test result. By adopting the embodiment of the invention, the model evaluation efficiency of the artificial intelligence model can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of data processing, and particularly to data processing methods, devices, and related equipment. Background Art

[0002] Currently, when evaluating an artificial intelligence model (such as a code generation model), it is usually necessary to input all the use cases in the selected evaluation dataset into the artificial intelligence model in sequence to obtain the data processing results corresponding to all the use cases in the evaluation dataset, and then test the data processing results corresponding to all the use cases in sequence to obtain the evaluation result of the artificial intelligence model.

[0003] However, the inventors found in the practical process that when processing the use cases in the evaluation dataset through the artificial intelligence model, each use case is processed serially, that is, the artificial intelligence model is used to process one use case at a time, and then the next use case is processed after one use case is processed. Once there are many use cases in the evaluation dataset, it takes a long data processing time to obtain the data processing results of all the use cases in the evaluation dataset. In addition, when testing the data processing results of the use cases, it is necessary to test after determining the data processing results of all the use cases in the evaluation dataset, resulting in a long test waiting time, thus causing a certain waste of time in the process of model evaluation, and reducing the model evaluation efficiency of the artificial intelligence model. Summary of the Invention

[0004] Embodiments of this application provide a data processing method, device, and related equipment based on an artificial intelligence model, which helps to improve the model evaluation efficiency of the artificial intelligence model.

[0005] On the one hand, embodiments of this application provide a data processing method based on an artificial intelligence model, which is executed by a service processing device; the method includes:

[0006] Receiving a model evaluation task submitted by a service object; the model evaluation task is determined by the service object based on an evaluation dataset and a model to be evaluated configured on an evaluation task configuration page; the evaluation dataset includes N use case data; N is a positive integer; the concurrent processing capacity of the model to be evaluated is used to indicate M shard tasks associated with the model evaluation task; M is a positive integer;

[0007] Determining K use cases to be evaluated associated with the N use case data based on the model evaluation task; K is a positive integer multiple of N;

[0008] Create M model processing threads associated with M shard tasks based on the number of parallel processes, divide K test cases to be evaluated into M shard tasks, and obtain shard task processing queues for the M shard tasks; one shard task corresponds to one model processing thread; among the M shard tasks, there is a shard task i, and the model processing thread associated with the shard task i is the model processing thread j, where both i and j are positive integers less than or equal to M.

[0009] Among the test cases to be evaluated included in the shard task processing queue of the shard task i, determine the first test case to be evaluated, input the first test case to be evaluated into the model to be evaluated through the model processing thread j, and have the model to be evaluated perform data processing on the first test case to be evaluated to obtain a data processing result for the first test case to be evaluated.

[0010] Obtain a test processing thread associated with the model evaluation task, and perform test processing on the data processing result based on the test processing thread to obtain a test result for the first test case to be evaluated.

[0011] On the one hand, an embodiment of the present application provides a data processing method based on an artificial intelligence model. The method is executed by a service client, and the method includes:

[0012] Display an evaluation task configuration page; the evaluation task configuration page includes a model configuration area and a data set configuration area;

[0013] In response to a model configuration operation in the model configuration area, display the model to be evaluated in the model configuration area; the concurrent processing volume of the model to be evaluated is used to indicate M shard tasks associated with the model evaluation task; M is a positive integer;

[0014] In response to a data set configuration operation in the data set configuration area, display an evaluation data set in the data set configuration area; the evaluation data set includes N case data; N is a positive integer;

[0015] In response to a task confirmation operation on the task confirmation control in the evaluation task configuration page, generate a model evaluation task based on the model to be evaluated and the evaluation data set, and send the model evaluation task to a service processing device, so that the service processing device determines K test cases to be evaluated associated with the N case data based on the model evaluation task, divide the K test cases to be evaluated into M shard tasks, and obtain shard task processing queues for the M shard tasks; the shard task processing queue is used to indicate that when the first test case to be evaluated is determined from the shard task processing queue, input the first test case to be evaluated into the model to be evaluated through the model processing thread, and have the model to be evaluated perform data processing on the first test case to be evaluated to obtain a data processing result for the first test case to be evaluated.

[0016] Receive the use case test result of the first use case to be evaluated and display the use case test result; the use case test result is obtained by the service processing device testing the data processing result based on the test processing thread.

[0017] On the one hand, an embodiment of the present application provides a data processing device based on an artificial intelligence model. The device is run by a service processing device. The device includes:

[0018] A task receiving module, configured to receive a model evaluation task submitted by a service object; the model evaluation task is determined by the service object based on the evaluation data set and the model to be evaluated configured on the evaluation task configuration page; the evaluation data set includes N use case data; N is a positive integer; the concurrent processing capacity of the model to be evaluated is used to indicate M shard tasks associated with the model evaluation task; M is a positive integer;

[0019] A use case determination module, configured to determine K use cases to be evaluated associated with the N use case data based on the model evaluation task; K is a positive integer multiple of N;

[0020] A task sharding module, configured to create M model processing threads associated with the M shard tasks based on the parallel processing quantity, divide the K use cases to be evaluated into the M shard tasks, and obtain a shard task processing queue for each of the M shard tasks; one shard task corresponds to one model processing thread; among the M shard tasks, there is a shard task i, and the model processing thread associated with the shard task i is the model processing thread j, and both i and j are positive integers less than or equal to M;

[0021] A model processing module, configured to determine a first use case to be evaluated among the use cases to be evaluated included in the shard task processing queue of the shard task i, input the first use case to be evaluated into the model to be evaluated through the model processing thread j, and the model to be evaluated processes the data of the first use case to be evaluated to obtain a data processing result for the first use case to be evaluated;

[0022] A testing module, configured to obtain a test processing thread associated with the model evaluation task, and test the data processing result based on the test processing thread to obtain a use case test result for the first use case to be evaluated.

[0023] On the one hand, an embodiment of the present application provides a data processing device based on an artificial intelligence model. The device is run by a service client. The device includes:

[0024] A configuration page display module, configured to display an evaluation task configuration page; the evaluation task configuration page includes a model configuration area and a data set configuration area;

[0025] A model configuration module, configured to display a model to be evaluated in a model configuration area in response to a model configuration operation in the model configuration area; the concurrent processing volume of the model to be evaluated is used to indicate M shard tasks associated with a model evaluation task; M is a positive integer;

[0026] A dataset configuration module, configured to display an evaluation dataset in a dataset configuration area in response to a dataset configuration operation for the dataset configuration area; the evaluation dataset includes N case data; N is a positive integer;

[0027] A task generation module, configured to generate a model evaluation task based on the model to be evaluated and the evaluation dataset in response to a task confirmation operation on a task confirmation control on an evaluation task configuration page, and send the model evaluation task to a business processing device, so that the business processing device determines K cases to be evaluated associated with the N case data based on the model evaluation task, and divides the K cases to be evaluated into M shard tasks to obtain a shard task processing queue for the M shard tasks; the shard task processing queue is used to indicate that when a first case to be evaluated is determined from the shard task processing queue, the first case to be evaluated is input into the model to be evaluated through a model processing thread, and the model to be evaluated processes the data of the first case to be evaluated to obtain a data processing result for the first case to be evaluated;

[0028] A result receiving module, configured to receive a case test result of the first case to be evaluated and display the case test result; the case test result is obtained by the business processing device performing test processing on the data processing result based on a test processing thread.

[0029] On the one hand, an embodiment of the present application provides a computer-readable storage medium storing a computer program, which is adapted to be loaded and executed by a processor so that a computer device having the processor executes the method provided by the embodiment of the present application.

[0030] On the one hand, an embodiment of the present application provides a computer program product, characterized in that it includes a computer program / instructions, and when the computer program / instructions are executed by a processor, they are used to execute the method provided by the embodiment of the present application.

[0031] In an embodiment of the present application, when a model evaluation task is obtained, K test cases to be evaluated can be determined based on the case data in the dataset, where K is a positive integer. Then, the K test cases to be evaluated can be divided into M sharding tasks, and each sharding task corresponds to a model processing thread. Thus, the test cases to be evaluated in different sharding tasks can be input into the model to be evaluated for data processing in parallel through M model processing threads to obtain corresponding data processing results. Therefore, multiple test cases to be evaluated can be processed in parallel at the same time, thereby improving the efficiency of processing the test cases to be evaluated. In addition, when the data processing result of a single test case to be evaluated (such as the data processing result of the first test case to be evaluated) is obtained, the obtained data processing result of the single case can be tested, without waiting for all the test cases to be evaluated to determine the corresponding data processing results before testing. This helps to improve the efficiency of data testing and further improve the model evaluation efficiency of the artificial intelligence model. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] To more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for the description of the embodiments will be briefly introduced below. Obviously, the drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0033] Figure 1 is a schematic structural diagram of a data processing system provided by an embodiment of the present application;

[0034] Figure 2 is a schematic scenario diagram of a data processing method based on an artificial intelligence model provided by an embodiment of the present application;

[0035] Figure 3 is a schematic flowchart of a data processing method based on an artificial intelligence model provided by an embodiment of the present application;

[0036] Figure 4 is a schematic flowchart of a sharding task processing process provided by an embodiment of the present application;

[0037] Figure 5 is a schematic architecture diagram of a model evaluation system provided by an embodiment of the present application;

[0038] Figure 6 is a schematic flowchart of a data processing method based on an artificial intelligence model provided by an embodiment of the present application;

[0039] Figure 7 is a schematic diagram of a request processing process provided by an embodiment of the present application;

[0040] Figure 8It is a schematic flowchart of a test task processing provided by an embodiment of the present application;

[0041] Figure 9 It is a system architecture diagram of another model evaluation system provided by an embodiment of the present application;

[0042] Figure 10 It is a schematic flowchart of a data processing method based on an artificial intelligence model provided by an embodiment of the present application;

[0043] Figure 11 It is a schematic diagram of the effect of a task configuration page provided by an embodiment of the present application;

[0044] Figure 12 It is a schematic diagram of the effect of an evaluation task execution page provided by an embodiment of the present application;

[0045] Figure 13 It is a schematic diagram of the effect of a historical task viewing page provided by an embodiment of the present application;

[0046] Figure 14 It is a schematic diagram of the processing time sequence of a model evaluation process provided by an embodiment of the present application;

[0047] Figure 15 It is a schematic diagram of a task processing process provided by an embodiment of the present application;

[0048] Figure 16 It is a schematic structural diagram of a data processing device based on an artificial intelligence model provided by an embodiment of the present application;

[0049] Figure 17 It is a schematic structural diagram of a data processing device based on an artificial intelligence model provided by an embodiment of the present application;

[0050] Figure 18 It is a schematic structural diagram of a computer device provided by an embodiment of the present application. Detailed implementation manners

[0051] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts shall fall within the protection scope of the present application.

[0052] Please refer to Figure 1 , Figure 1 It is a schematic structural diagram of a data processing system provided by an embodiment of the present application. As Figure 1As shown, the data processing system may include terminal devices (such as device 11a, device 12a, device 13a) and a server 200a. It can be understood that Figure 1 the numbers of the terminal devices and the server in

[0053] are merely illustrative. According to the implementation requirements, there can be any number of terminal devices and servers. The terminal devices (such as device 11a, device 12a, device 13a) can communicate with the server through a network (i.e., a medium that provides a communication link through wired, wireless communication links, fiber optic cables, etc.), and then data can be transmitted.

[0054] It can be understood that a client can run on the terminal device (such as device 12a), and the client can be a program that provides local services for users (also called service objects, operation objects). The server 200a can be the server corresponding to the client, and a program for providing resources, service data, and other services can run in the server 200a. It can be understood that the client running on the terminal device can also be called an application client, a service client, etc. For example, the client running on the terminal device can be a client for performing a configuration model evaluation task, and then the user can configure the model evaluation task on the terminal device and can display the evaluation results of the model evaluation task on the terminal device.

[0055] It can be understood that the embodiments of the present application can be applied to computer devices. For example, the computer device can be the Figure 1 terminal device in Figure 1 above, or can be the Figure 1 server in Figure 1When in the server, the user can configure a model evaluation task through a terminal device and send the model evaluation task to the server, so that after the server obtains the model evaluation task, it determines the evaluation result of the model evaluation task and returns the evaluation result to the terminal device for display.

[0056] Please refer to Figure 2 , Figure 2 FIG. is a schematic diagram of a scenario of a data processing method based on an artificial intelligence model provided by an embodiment of the present application. As Figure 2 shown, the terminal device 20a can be the terminal device corresponding to the service object A. A task configuration page can be displayed on the terminal device 20a, and the service object A can configure a model evaluation task through the task configuration page. When configuring the model evaluation task, the artificial intelligence model to be evaluated and the data set used to evaluate the configured artificial intelligence model can be configured. For example, as Figure 2 shown, the page 201a is a task configuration page. In this task configuration page, it can be displayed that the configured model is the artificial intelligence model A and the configured data set is the data set S. It should be understood that when the service object clicks the confirmation control for the task configured for the task configuration page (such as Figure 2 the control 201b shown in FIG.), the terminal device 20a sends a model evaluation task to the service processing device 21a (i.e., step S21).

[0057] Further, the service processing device 21a can obtain the model evaluation task sent by the terminal device 20a. Then, the service processing device 21a can determine a plurality of test cases to be evaluated based on the data set S (such as Figure 2 202a shown in FIG.). This process can also be called test case disassembling. The test cases to be evaluated refer to the test cases used for model evaluation and can include information such as input data and expected results required for evaluating the artificial intelligence model A. For example, as Figure 2 shown in 203a in FIG., test case 1 to be evaluated, test case 2 to be evaluated, test case 3 to be evaluated,..., test case n to be evaluated, etc. can be determined. Further, the determined plurality of test cases to be evaluated can be divided into M task shards, and the M task shards can be determined based on the concurrent processing quantity of the artificial intelligence model A. For example, as Figure 2 shown in FIG., the plurality of test cases to be evaluated can be divided into 3 task shards, specifically task shard 204a, task shard 205a, and task shard 206a. Among them, the number of test cases to be evaluated in each task shard should be as average as possible.

[0058] Further, the model processing thread corresponding to each task shard can be determined, and one task shard corresponds to one model processing thread. For example, as Figure 2As shown, the model processing thread of task shard 204a is model processing thread 1 (as shown by 207a in Figure 2 ), the model processing thread of task shard 205a is model processing thread 2 (as shown by 208a in Figure 2 ), and the model processing thread of task shard 206a is model processing thread 3 (as shown by 209a in Figure 2 ). It should be understood that these 3 model processing threads can process the test cases to be evaluated in the corresponding task shards in parallel, and a model processing thread can process the test cases to be evaluated in the corresponding task shards sequentially. For example, model processing thread 1 can sequentially process the test cases to be evaluated in task shard 204a.

[0059] Here, taking test case 1 to be evaluated in shard task 204a as an example, the model processing thread can notify artificial intelligence model A (as shown by 210a in Figure ) to perform data processing on test case 1 to be evaluated, obtaining the data processing result corresponding to test case 1 to be evaluated (as shown by 211a in ​ ), and then can process the data processing result through the test processing thread to obtain the test result of test case 1 to be evaluated (as shown by 212a in ​ ).

[0060] Furthermore, when the service processing device obtains the test result of test case 1 to be evaluated, it can return the test result of the test case to terminal device 20a, so that terminal device 20a can display the test result of test case 1 to be evaluated. In addition, after the service processing device obtains the test results of all test cases to be evaluated, it can determine the task evaluation result of the model evaluation task based on the test results of all test cases to be evaluated, and then return the task evaluation result to terminal device 20a, so that terminal device 20a can display the task evaluation result.

[0061] In an embodiment of the present application, when a model evaluation task is obtained, K test cases to be evaluated can be determined based on the case data in the dataset, where K is a positive integer. Then, the K test cases to be evaluated can be divided into M sharding tasks, and each sharding task corresponds to a model processing thread. Thus, the test cases to be evaluated in different sharding tasks can be input into the model to be evaluated for data processing in parallel by M model processing threads, and corresponding data processing results can be obtained. Therefore, multiple test cases to be evaluated can be processed in parallel at the same time, thereby improving the efficiency of processing the test cases to be evaluated. In addition, when the data processing result of a single test case to be evaluated (such as the data processing result of the first test case to be evaluated) is obtained, the obtained data processing result of the single case can be tested, without waiting for all the test cases to be evaluated to determine the corresponding data processing results before testing. This helps to improve the efficiency of data testing and further improve the model evaluation efficiency of the artificial intelligence model.

[0062] It can be understood that the embodiments of the present application can be applied to the field of artificial intelligence technology. For example, it is used to evaluate the model of an artificial intelligence model trained based on artificial intelligence technology. Artificial Intelligence (AI) uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results in theory, methods, technologies, and application systems. In other words, artificial intelligence is a comprehensive technology in computer science. It attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines, enabling the machines to have the functions of perception, reasoning, and decision-making. Artificial intelligence technology is an interdisciplinary subject, involving a wide range of fields, including both hardware-level technologies and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0063] It should be noted that before and during the collection of relevant data of the user, this application can display a prompt interface, a pop-up window, or output a voice prompt message. The prompt interface, pop-up window, or voice prompt message is used to prompt the user that their relevant data is currently being collected, so that this application only starts to execute the relevant steps of obtaining the user's relevant data after obtaining the confirmation operation of the user on the prompt interface or pop-up window. Otherwise (that is, when the confirmation operation of the user on the prompt interface or pop-up window is not obtained), the relevant steps of obtaining the user's relevant data are ended, that is, the relevant data of the user is not obtained. In other words, all user data collected by this application is collected with the consent and authorization of the user, and the collection, use, and processing of relevant user data need to comply with relevant laws, regulations, and standards in the relevant region.

[0064] It can be understood that the above scenarios are only examples and do not constitute a limitation on the application scenarios of the technical solutions provided by the embodiments of this application. The technical solutions of this application can also be applied to other scenarios. For example, as known to those of ordinary skill in the art, with the evolution of the system architecture and the emergence of new business scenarios, the technical solutions provided by the embodiments of this application are equally applicable to similar technical problems.

[0065] Further, please refer to ​ , ​ is a schematic flowchart of a data processing method based on an artificial intelligence model provided by an embodiment of this application. This method can be executed by a service processing device, such as the service processing device 21a in the above ​ . This method can at least include the following steps S101-step S105.

[0066] S101. Receive a model evaluation task submitted by a service object; the model evaluation task is determined by the service object based on an evaluation data set and a model to be evaluated configured on an evaluation task configuration page; the evaluation data set includes N case data; N is a positive integer; the concurrent processing capacity of the model to be evaluated is used to indicate M shard tasks associated with the model evaluation task; M is a positive integer.

[0067] Among them, the model evaluation task can be a task for evaluating an artificial intelligence model. Among them, the model evaluation task is determined by the service object based on an evaluation data set and a model to be evaluated configured on an evaluation task configuration page. The service object can be the object that submits the model evaluation task. The task configuration page can be a page for configuring the model evaluation task, and the task configuration page can be used to configure parameters required for the model evaluation task, such as the evaluation data set and the model to be evaluated.

[0068] Among them, the model to be evaluated can be an artificial intelligence model to be evaluated. The model to be evaluated can be a code generation model (i.e., program statement generation model) for generating code (i.e., program statements), also known as a program statement generation model, or a model for code annotation, a model for code completion, which is not limited here. The model to be evaluated can also be a model for text processing, such as an artificial intelligence model for text translation, a model for intelligent question answering, etc. The model to be evaluated can also be a model for image processing, such as a model for object detection, a model for classifying and identifying objects in an image, etc. It can be understood that the model to be evaluated can also be other types of artificial intelligence models, such as multimodal models, etc., which is not limited here. For example, the model to be evaluated can be chatGPT (an artificial intelligence model), Wenxin Yiyan (an artificial intelligence model), Github Copilot (an artificial intelligence model), etc., which is not limited here.

[0069] The evaluation dataset can be a dataset for testing the model to be evaluated. It can be understood that the evaluation dataset can include N case data. The case data can be the data of the cases in the evaluation dataset for evaluating the model to be evaluated. For example, the evaluation dataset can be datasets such as HumanEval (a dataset), MBPP (a dataset), etc., which is not limited here. In the case data, it can include the input data (also known as case description information) required for evaluating the model to be evaluated, case test information, etc. Among them, the case description information can be the data input to the model to be evaluated for data processing, and the case test information can be the information for testing the data processing result. For example, in the evaluation dataset for evaluating the code generation model, the case description information of each case data can be the prompt for prompting the code generation model to generate what code segment, and the case test information can be the code segment for testing the code generated by the model. Another example is that in the dataset for evaluating the image classification model, the case description information of each case can be the image to be classified, and the case test information can be used to indicate the actual category corresponding to the image. Then, when evaluating the model to be evaluated, the input data in the case data (such as case data A) can be input into the model to be evaluated, and through the data processing of the model to be evaluated, a data processing result (i.e., the result obtained through the data processing of the model to be evaluated) can be obtained. Then, by testing the data processing result with the case test information, the case test result for this case data (such as case data A) can be obtained. Furthermore, the model evaluation result of the model to be evaluated can be determined through the test results of multiple cases.

[0070] It can be understood that the model to be evaluated has a corresponding concurrent processing capacity. The concurrent processing capacity can refer to the number of use cases that the model to be evaluated can process within a unit time. Simply put, it is the number of use cases that the model to be evaluated can process in parallel. It can be understood that the concurrent processing capacity of the model to be evaluated can be determined according to the structure of the model to be evaluated itself and the processing capacity of the business processing device. If the structure of the model to be evaluated is smaller and the configuration of the business processing device is better (i.e., the processing capacity of the business processing device is better), then the concurrent processing capacity of the model is larger.

[0071] The concurrent processing capacity of the model to be evaluated is used to indicate M shard tasks associated with the model evaluation task; M is a positive integer. The shard task can be a sub-task obtained by sharding the model evaluation task. It should be understood that the number of M shard tasks should be less than or equal to the concurrent processing capacity of the model to be evaluated. For example, if the concurrent processing capacity of the model to be evaluated is 3, then the maximum number of determined shard tasks can be 3. Of course, the number of shard tasks can also be 2 or 1, which is not limited here and is specifically determined according to the actual situation. To improve the efficiency of model evaluation, a number less than or equal to the concurrent processing capacity and the largest value can be selected as the number of shard tasks.

[0072] S102. Determine K use cases to be evaluated associated with N use case data based on the model evaluation task; K is a positive integer multiple of N.

[0073] Among them, the use case to be evaluated can refer to the use case used for model evaluation. It should be understood that the use case to be evaluated is determined based on the N use case data in the evaluation dataset. It can be understood that K is a positive integer multiple of N, such as 1 time, 2 times, etc. It should be understood that the method for determining K use cases to be evaluated can be determined according to actual needs.

[0074] Optionally, the model to be evaluated is a program statement generation model; the model evaluation task carries the program language type associated with the program statement generation model; the number of types of the program language type is L, and L is a positive integer; then, determining K use cases to be evaluated associated with N use case data based on the model evaluation task can include the following steps: determining K use cases to be evaluated associated with N use case data based on the number of types of the program language type and N use case data; K = L * N; a use case to be evaluated is obtained based on one program language type and one use case data.

[0075] Among them, the program statement generation model can be an artificial intelligence model for generating program statements (i.e., code), and can also be called a code generation model. It can be understood that the program language type can be used to indicate the language type of the program statements generated by the model to be evaluated. For example, the program language type can be python, Java, go, etc., which is not limited here.

[0076] It can be understood that the number of types of the program language type is L, and L is a positive integer. In other words, the number of types of the program language type is greater than or equal to 1. It can be understood that the program language type can be configured by the business object through the above task configuration page.

[0077] It can be understood that a use case to be evaluated can be obtained based on a program language type and a use case data. The number of determined use cases to be evaluated is equal to the product of the number of use case data in the evaluation dataset and the number of types of the program language type, that is, K = L * N.

[0078] For example, the program language types carried in the model evaluation task include language type 1 and language type 2, that is, the number of types of the program language types carried is 2. The evaluation dataset includes 164 pieces of use case data. Then, 164 use cases to be evaluated can be determined based on language type 1 and 164 pieces of use case data, so as to indicate that the model to be evaluated determines the program statements corresponding to language type 1 based on these 164 use cases to be evaluated; in addition, 164 use cases to be evaluated can be determined based on language type 2 and 164 pieces of use case data, so as to indicate that the model to be evaluated determines the program statements corresponding to language type 2 based on these 164 use cases to be evaluated. In other words, for each language type (i.e., language type 1 and language type 2), 164 corresponding use cases to be evaluated are determined, and then 164 * 2 use cases to be evaluated can be obtained. Another example is that if the program language type carried in the model evaluation task includes language type 1, that is, the number of types of the program language type carried is 1, and the evaluation dataset includes 164 pieces of use case data, then the number of determined use cases to be evaluated is 164.

[0079] It can be understood that different artificial intelligence models have differences in request methods, parameters, and return values, and different programming languages also have different coding specifications. Therefore, in order to enable the artificial intelligence model to output more standardized data during model evaluation, the data input into the model (i.e., the input data in the use case to be evaluated) can be character-adjusted so that the artificial intelligence model can input more standardized results.

[0080] Specifically, each of the K use case data includes use case description information. Then, determining K to-be-evaluated use cases associated with N use case data based on the model evaluation task may include the following steps: obtaining a character processing strategy associated with the to-be-evaluated model, adjusting the use case description information in the N use case data based on the character processing strategy to obtain N updated use case data corresponding to the N use case data; one use case data corresponds to one updated use case data; determining K to-be-evaluated use cases associated with the N use case data based on the model evaluation task and the N updated use case data.

[0081] Among them, the use case description information may be information used to instruct the to-be-evaluated model to process the to-be-evaluated use case. For example, when the to-be-evaluated model is a code generation model, the use case description information may be a prompt for instructing the to-be-evaluated model to generate program code. For example, the to-be-evaluated model is instructed to generate a piece of code for sorting through the use case description information (i.e., the prompt). Another example is that when the to-be-evaluated model is a text translation model, the use case description information may be the text used to instruct the to-be-evaluated model to perform translation.

[0082] It should be understood that the character processing strategy may be a strategy for processing the use case description information, so that the processed use case description information can be input into the to-be-evaluated model for processing subsequently, which helps to standardize the results output by the to-be-evaluated model. Customized logic for the evaluation model and evaluation language. On the evaluation system interface, provide the "plugin development" function to the evaluation user (such as the above-mentioned business object). This character processing strategy can be configured by the evaluation object (such as the above-mentioned business object) on the evaluation system interface (for example, the character processing strategy can be developed using the python language). The evaluation system will provide the intermediate data in the evaluation process (such as the use case description information of the to-be-evaluated use case, the data processing result generated by the to-be-evaluated model) to the developer, and the developer can develop personalized character processing logic based on the intermediate data provided in the evaluation process. The evaluation system will load the code of the user-defined character processing strategy during operation to implement the user-defined character processing, so as to achieve the plug-in development of the character processing logic.

[0083] Among them, the updated use case data may refer to the use case data obtained by adjusting the use case description information based on the character processing strategy. One use case data may correspond to one updated use case data. By adjusting the use case description information in N use case data based on the character processing strategy, N updated use case data corresponding to the N use case data are obtained, that is, the use case description information in the use case data is adjusted based on the code of the character processing strategy. The process of determining K to-be-evaluated use cases associated with the N use case data based on the model evaluation task and the N updated use case data can refer to the relevant description of determining K to-be-evaluated use cases based on the model evaluation task and the N use case data, which will not be elaborated here.

[0084] For example, the code logic of the character processing strategy is as follows:

[0085] {user_code = “‘

[0086] prompt = ${prompt}+‘, it is necessary to add comments to each line of code’

[0087] ”’

[0088] }

[0089] Based on the character processing strategy indicated by the above code segment, the description of ‘it is necessary to add comments to each line of code’ can be added after the prompt indicated by the use case description information in the use case data to obtain the updated use case description information. Furthermore, when the updated use case description information is input into the to-be-evaluated model, the to-be-evaluated use cases can add comments to each line of code according to the indication of the added description, so that the to-be-evaluated use cases can perform code commenting more standardly.

[0090] Another example, the code logic of the character processing strategy is as follows:

[0091] {user_code = “‘

[0092] prompt = ${prompt}+‘, please complete the code’

[0093] ”’

[0094] }

[0095] Based on the character processing strategy indicated by the above code segment, the description of ‘please complete the code’ can be added after the prompt indicated by the use case description information in the use case data to obtain the updated use case description information. Furthermore, when the updated use case description information is input into the to-be-evaluated model, the to-be-evaluated use cases can complete the code according to the indication of the added description, so that the to-be-evaluated model can perform code completion more standardly.

[0096] S103. Create M model processing threads associated with M shard tasks based on the number of parallel processes, divide K test cases to be evaluated into M shard tasks, and obtain a shard task processing queue for each of the M shard tasks; one shard task corresponds to one model processing thread; among the M shard tasks, there is a shard task i, and the model processing thread associated with the shard task i is the model processing thread j, where both i and j are positive integers less than or equal to M.

[0097] Among them, the model processing thread can be a thread used to notify the model to be evaluated to perform data processing on the test case to be evaluated. A thread is the smallest unit that the operating system can perform operation scheduling on, and multiple threads can run simultaneously. It can be understood that one shard task corresponds to one model processing thread.

[0098] It can be understood that dividing K test cases to be evaluated into M shard tasks means determining the shard task corresponding to each test case to be evaluated, and then forming a shard task processing queue based on the test tasks divided into each shard task. It can be understood that when dividing K test cases to be evaluated into M shard tasks, the number of test cases to be evaluated divided into each shard task should be as average as possible. For example, if the number of test cases to be evaluated is 164 (i.e., K = 164) and the number of shard tasks is 2 (i.e., M = 2), then the data of the test cases to be evaluated divided into the 2 shard tasks are both 82. Another example, if the number of test cases to be evaluated is 164 (i.e., K = 164) and the number of shard tasks is 3 (i.e., M = 3), then among the 3 shard tasks, 2 shard tasks are divided into 55 test cases to be evaluated, and 1 shard task is divided into 54 test cases to be evaluated.

[0099] Among them, the shard task processing queue can be a queue composed of the test cases to be evaluated divided into the shard task. It should be understood that each shard task processing queue can include at least one test case to be evaluated. It can be understood that the test tasks in the shard task queue can be arranged in a certain order. The order of the test cases to be evaluated in the shard task queue can be randomly determined or determined according to the order of the case data corresponding to the test cases to be evaluated in the evaluation dataset, which is not limited here. Subsequently, when determining the test cases to be processed from the shard task queue, it is carried out sequentially based on the order of the test cases to be evaluated.

[0100] Among them, the sharding task i can be any one of the M sharding tasks. The model processing thread j can be the model processing thread associated with the sharding task i. Both i and j are positive integers less than or equal to M. For example, if M is 2, the 2 sharding tasks can include sharding task 1 and sharding task 2, and the model processing threads can include model processing thread 1 and model processing thread 2. The sharding task i can be sharding task 1 or sharding task 2, and the model processing thread j can be model processing thread 1 or model processing thread 2. It should be understood that here, taking the sharding task i as an example, the processing process of the use cases to be evaluated in the sharding task i is described, and the processing processes of other sharding tasks refer to the relevant descriptions of the sharding task i. In addition, the M sharding tasks can be processed in parallel. Thus, by sharding the model evaluation task into subtasks for parallel processing, it helps to improve the processing efficiency of the model evaluation task, that is, to improve the model evaluation efficiency of the model to be evaluated.

[0101] S104. Among the use cases to be evaluated included in the sharding task processing queue of the sharding task i, determine the first use case to be evaluated, input the first use case to be evaluated into the model to be evaluated through the model processing thread j, and have the model to be evaluated perform data processing on the first use case to be evaluated to obtain a data processing result for the first use case to be evaluated.

[0102] It can be understood that the first use case to be evaluated can be the use case to be evaluated with the first order in the sharding task processing queue of the sharding task i. After the processing of the first use case to be evaluated is completed (that is, the corresponding data processing result is generated by the model to be evaluated), the first use case to be evaluated can be removed from the sharding task queue, so as to use the next use case to be evaluated of the first use case to be evaluated as the new first use case to be evaluated (that is, update the first use case to be evaluated based on the next use case to be evaluated of the first use case to be evaluated) to process the new first use case to be evaluated. For example, in the sharding task queue i, the arrangement of the use cases to be evaluated is: use case Y1, use case Y2, use case Y3,..., use case Yn. Then, use case Y1 can be first used as the first use case to be evaluated for processing. After the processing of use case Y1 is completed, use case Y1 can be removed from the sharding task queue i. The arrangement of the use cases to be evaluated in the sharding task queue i after removing use case Y1 is: use case Y2, use case Y3,..., use case Yn. Furthermore, use case Y2 can be used as the new first use case to be evaluated for processing, and so on, until all the use cases to be evaluated in the sharding task queue i are processed as the first use case to be evaluated.

[0103] It can be understood that the data processing result refers to the result obtained by the model to be evaluated in processing the first use case to be evaluated. When the model to be evaluated processes the first use case to be evaluated to obtain the data processing result for the first use case to be evaluated, it can be processed according to the actual model result of the model to be evaluated, which is not limited here. For example, when the model to be evaluated is a code generation model, the data processing result can be a piece of code determined based on the prompt in the first use case to be evaluated (i.e., the use case description information). Another example is that when the model to be evaluated is a text translation model, the data processing result can be the translated text determined based on the text in the first use case to be evaluated (i.e., the use case description information).

[0104] Specifically, a business processing component and a model processing component are deployed in the business processing device. Then, inputting the first use case to be evaluated into the model to be evaluated through the model processing thread j may include the following steps: calling the model processing thread j in the business processing component, generating a model processing request corresponding to the first use case to be evaluated based on the model processing thread j, and sending the model processing request to the model processing component; the model processing request carries the first use case to be evaluated; inputting the first use case to be evaluated into the model to be evaluated through the model processing component.

[0105] Among them, the business processing component may refer to the component in the business processing device used for model evaluation task processing. It can be understood that the business processing component can be used to execute the above steps S102 - S103, that is, it can be used to determine K use cases to be evaluated based on the use case data in the evaluation dataset when obtaining a model evaluation task, and divide the determined K use cases to be evaluated into M task slices. Moreover, the business processing component can also be used to call the model processing thread to determine the model processing request for the use case to be evaluated.

[0106] Among them, the model processing component can be a component used to call the model to be processed to process the use case to be evaluated. It can be understood that due to the inconsistent usage methods of artificial intelligence models, the model processing component can encapsulate the service interfaces of different artificial intelligence models (such as chatGPT, Wenxin Yiyan, Github Copilot, etc.) and provide a unified standard HTTP API (an interface form) interface for the business processing component to call.

[0107] It can be understood that the model processing request can be a request for instructing the model processing component to perform data processing on the first use case to be evaluated. It can be understood that the first use case to be evaluated can be carried in the model processing request, specifically, it can carry the use case description information and use case identifier of the first use case to be evaluated. The use case identifier can be used to uniquely identify a use case among the K use cases to be evaluated. Generating a model processing request based on model processing thread j can be generating a model processing request by model processing thread j based on the use case description information of the first use case to be evaluated. Further, the model processing request is sent to the model processing component through the model processing thread. For example, the model processing request can be sent based on the URL (Uniform Resource Locator) request address. For instance, the request address is http: / / {baseUrl} / api / generation / code?model_name=hy, where model_name=hy is used to indicate that the model identifier of the model to be evaluated is hy.

[0108] Further, the model processing component can input the first use case to be evaluated into the model to be evaluated, so that the model to be evaluated can perform data processing on the first use case to be evaluated (specifically, it can be the use case description information of the first use case to be evaluated) to obtain a data processing result. For example, when the model processing component is a code generation component, the use case description information of the first use case to be evaluated (i.e., the prompt for indicating what kind of code to generate) is used to indicate generating a sorting algorithm code, then the model processing component can generate a sorting algorithm code (i.e., the data processing result) based on the use case description information.

[0109] For example, when the model to be evaluated is a code generation model, the data structure of the model processing request can be:

[0110] {task_id; / / The identifier of the use case to be evaluated;

[0111] Prompt; / / The prompt of the use case to be evaluated;

[0112] .......}

[0113] Among them, in this model processing request, the task_id field can be used to indicate the identifier of the use case to be evaluated. Prompt can be used to indicate the prompt of the use case to be evaluated (i.e., the use case description information), and this prompt is used to prompt what kind of code the code generation should generate. The model processing request can also include some other field information, which is determined according to actual needs and is not limited here.

[0114] Further, when the model to be evaluated is a code generation model (also simply referred to as the model), the data structure of the result obtained after the code generation model processes the use case to be evaluated can be:

[0115] {task_id; / / Identifier of the test case to be evaluated;

[0116] generation; / / Code result generated by the model;

[0117] }

[0118] Among them, in this result, the task_id field represents the identifier of the test case to be evaluated; the generation field is used to indicate the code result generated by the code generation model, also known as the model-generated code.

[0119] Here, the processing process of the sharding task is described in combination with the illustration. Please refer to ​ , ​ which is a schematic flowchart of a sharding task processing process provided by an embodiment of this application. As ​ shown, multiple test cases to be evaluated can be obtained based on the evaluation dataset (such as ​ shown as 401a in), for example, the multiple test cases obtained can include: Test Case 1 to be evaluated, Test Case 2 to be evaluated, Test Case 3 to be evaluated,..., Test Case n to be evaluated, and so on. Furthermore, the multiple test cases obtained can be divided into multiple sharding tasks. As ​ shown, sharding task 402a, sharding task 403a, and sharding task 404a can be obtained. Among them, sharding task 402a can include test cases to be evaluated such as Test Case 1, Test Case 2, and Test Case n1, sharding task 403a can include test cases to be evaluated such as Test Case n2, Test Case n3, and Test Case n4, and sharding task 404a can include test cases to be evaluated such as Test Case n5, Test Case n6, and Test Case n. Furthermore, the test cases in each sharding task can be processed in parallel through the model processing threads corresponding to each sharding task. Specifically, Test Case 1 is processed through model processing thread 1 (such as ​ shown as 405a in), that is, Test Case 1 is input into the model to be evaluated (such as ​ shown as 408a in), and data processing result 1 (such as ​ shown as 409a in) is obtained. At the same time, Test Case n2 can be processed through model processing thread 2 (such as ​ shown as 406a in), that is, Test Case n2 is input into the model to be evaluated (such as ​ shown as 408a in), and data processing result 2 (such as ​ shown as 410a in) is obtained; at the same time, Test Case n5 can be processed through model processing thread 3 (such as ​ shown as 407a in), that is, Test Case n5 is input into the model to be evaluated (such as ​as shown by 408a in [reference], to obtain a data processing result 3 (such as ​ as shown by 411a in [reference]). Thus, multiple shard tasks can be processed in parallel by multiple threads for the test cases to be evaluated, thereby improving the efficiency of data processing by the model and further improving the efficiency of model evaluation.

[0120] S105: Obtain a test processing thread associated with the model evaluation task, and perform test processing on the data processing result based on the test processing thread to obtain a test result for the first test case to be evaluated.

[0121] It can be understood that the test processing thread can be a thread used to test the data processing result of the test case to be evaluated (such as the first test case to be evaluated). The test result of the case can be the result obtained by testing the data processing result of the test case to be evaluated (such as the first test case to be evaluated), and this test result of the case can be used to indicate whether the test case to be evaluated passes the test. It can be understood that the specific method of testing the data processing result can be determined according to the actual situation and will not be elaborated here. For example, when testing the code generated by the code generation model (i.e., the model-generated code), executable code segments can be determined based on the generated code, and then during the test, the executable code segments are run to obtain the test result. If the run is successful, it indicates that the test passes; if the run is not successful, it indicates that the test fails. Another example is when testing the classification result identified by the image classification and recognition model for an image. The identified classification result can be compared with the expected result in the test case to be evaluated. If the comparison result indicates that the two are consistent, it indicates that the test passes; if they are inconsistent, it indicates that the test fails.

[0122] It should be understood that in the embodiments of the present application, when obtaining the data processing result of a single test case to be evaluated, the obtained data processing result of the single case can be immediately tested, without waiting for all the test cases to be evaluated to determine their corresponding data processing results before testing. This helps to improve the efficiency of data testing and further improve the efficiency of model evaluation.

[0123] It can be understood that a test processing component can be deployed in the service processing device, and the test processing component can be a component used to test the data processing result of the test case to be evaluated. Furthermore, when the service processing device obtains the data processing result of the first test case to be evaluated, a test processing request related to the data processing result can be generated through the test processing thread. Then, the test processing thread can send the test processing request to the test processing component, and the test processing component tests the data processing result.

[0124] It should be understood that the use case test result of the first use case to be evaluated can be used to determine the task evaluation result of the model evaluation task. In the embodiments of the present application, the use case test results of K use cases to be evaluated can be determined, and then the task evaluation result of the entire model evaluation task can be determined through the use case test result of each use case to be evaluated.

[0125] Specifically, when the use case test results for K use cases to be evaluated are obtained, based on the use case test results for K use cases to be evaluated, the task evaluation result for the model evaluation task is determined.

[0126] Among them, the task evaluation result may refer to the evaluation result of the model evaluation task. For example, the task evaluation result can be to determine a test pass rate based on the use case test results of each use case to be evaluated, and then determine the task evaluation result based on the test pass rate. The test pass rate can be the ratio of the number of use cases to be evaluated with a passed use case test result to all use cases to be evaluated. Another example is that the index value of a certain task index can be determined for the use case test results of each use case to be evaluated, and then the task evaluation result can be determined based on the index values of each use case to be evaluated. For example, the average value of the index values of each use case to be evaluated is determined as the task evaluation result. The task index here can be configured by the business object on the task configuration page. For example, the task index is Pass@1, that is, the ratio of the test pass for the result obtained by inputting the model to be evaluated only once for a use case to be evaluated. Another example is that the task index is Pass@10, that is, among the 10 results (one result is obtained for each processing) obtained by repeatedly inputting a use case to be evaluated into the model to be evaluated 10 times, the ratio of at least one test pass.

[0127] It can be understood that based on the above description, in the embodiments of the present application, during the process of the business processing device performing model evaluation, it mainly involves three data processing components, namely the business processing component, the model processing component, and the test processing component. The specific introduction of each component can refer to the above relevant description and will not be elaborated here.

[0128] Here, the system architecture for model evaluation in the evaluation scenario of the code generation model is described in conjunction with the drawings. Please refer to ​ , ​ is the schematic diagram of the architecture of a model evaluation system provided by the embodiments of the present application. As ​As shown in the figure, the business processing device 51a may include three components, namely, a business processing component 501a, a model processing component 502a, and a test processing component 503a. Among them, the business processing component 501a may include sub-components such as a dataset processing sub-component 511a and an evaluation processing sub-component 512a. Among them, the evaluation processing sub-component 512a is the core function of the web background service (i.e., the business processing component). Its main function is to receive the user's evaluation request (i.e., the request associated with the model evaluation task), and process all the test cases to be evaluated according to the sharding strategy (i.e., the strategy for determining the sharding task processing queue of M sharding tasks mentioned above). And it parallelly calls the code generation service (i.e., the model processing component) and the code testing service (the test processing component) through multiple threads, and finally calculates the task evaluation result to return the task evaluation result to the user. Among them, other sub-components except the evaluation processing sub-component 512a can be called auxiliary function components. For example, the dataset processing sub-component 511a here, also known as the evaluation dataset management sub-component, can be used to manage the test dataset (i.e., the dataset) and the test cases (i.e., the case data). Specifically, for datasets such as HumanEval (a dataset) and MBPP (a dataset), and then it can provide the function of selecting the dataset for model evaluation to the business object when the business object configures the model evaluation task. It can be understood that the auxiliary function components may also include an evaluation model management sub-component (also known as the model management sub-component), an evaluation language management sub-component, an evaluation scenario management sub-component, and so on. Among them, the function of the evaluation model management sub-component is to manage artificial intelligence models, such as chatGPT, Wenxin Yiyan, Github Copilot, etc., and then it can provide the function of selecting the model to be evaluated to the business object when the business object configures the model evaluation task. The function of the evaluation language management sub-component is to manage the types of programming languages, such as python, Java, go, etc., and then it can provide the function of selecting the type of programming language to the business object when the business object configures the model evaluation task. The evaluation scenario management sub-component can be used to manage the evaluation scenarios, such as code completion, code generation, unit test (i.e., the smallest testable unit in the software) generation, comment generation, etc., and then it can provide the function of selecting the evaluation scenario to the business object when the business object configures the model evaluation task.

[0129] Among them, the model processing component 502a can provide data processing functions for models such as artificial intelligence model A, artificial intelligence model B, and artificial intelligence model C. The test processing component 503a can provide tests for the code of programming language types such as programming language type 1, programming language type 2, and programming language type 3.

[0130] Among them, the service processing device 51a can receive the model evaluation task 500a submitted by the service object. Furthermore, the service processing device can execute step S51 through the evaluation processing sub-component 512a in the service processing component 501a: determine K to-be-evaluated use cases based on the case data in the evaluation dataset, and divide the K to-be-evaluated use cases into M shard tasks. Furthermore, the to-be-evaluated use cases in the shard tasks can be processed through the model processing threads corresponding to each shard task. For example, the model processing threads in the service processing device can include Thread 1 (as shown by 513a in ​ ), Thread 2 (as shown by 514a in ​ ), and Thread 3 (as shown by 515a in ​ ), and each thread is used to process the corresponding shard task. Taking Thread 2 as an example here, the service processing device executes step S52 through Thread 2: perform data processing on a single to-be-evaluated use case (such as the to-be-evaluated use case 1) for the notification service processing component. Furthermore, the model processing component can call the artificial intelligence model (such as the artificial intelligence model A) indicated by the model evaluation task to process the to-be-evaluated use case 1 to obtain a data processing result (i.e., model-generated code), and then execute step S53 through the model processing component 502a: return the determined data processing result (i.e., model-generated code) to the service processing component. Further, the service processing device can execute step S54: notify the test processing component 503a to test the data processing result of the to-be-evaluated use case 1.

[0131] Furthermore, please refer to ​ . ​ is a schematic flowchart of a data processing method based on an artificial intelligence model provided by an embodiment of the present application. This method can be executed by a service processing device, such as the service processing device 21a in the above ​ . This method can at least include the following steps S201 - step S206.

[0132] S201. Receive the model evaluation task submitted by the service object; the model evaluation task is determined by the service object based on the evaluation dataset and the to-be-evaluated model configured on the evaluation task configuration page; the evaluation dataset includes N case data; N is a positive integer; the concurrent processing volume of the to-be-evaluated model is used to indicate M shard tasks associated with the model evaluation task; M is a positive integer.

[0133] S202. Determine K to-be-evaluated use cases associated with the N case data based on the model evaluation task; K is a positive integer multiple of N.

[0134] S203. Create M model processing threads associated with M shard tasks based on the number of parallel processes, divide K test cases to be evaluated into M shard tasks, and obtain shard task processing queues for the M shard tasks; one shard task corresponds to one model processing thread; among the M shard tasks, there is a shard task i, and the model processing thread associated with the shard task i is the model processing thread j, where both i and j are positive integers less than or equal to M.

[0135] S204. Among the test cases to be evaluated included in the shard task processing queue of the shard task i, determine the first test case to be evaluated, input the first test case to be evaluated into the model to be evaluated through the model processing thread j, and have the model to be evaluated perform data processing on the first test case to be evaluated to obtain a data processing result for the first test case to be evaluated.

[0136] Among them, for the processing procedures of steps S201 - S204, the processing logic of the above steps S201 - S204 can be referred to, and details are not elaborated here.

[0137] It can be understood that the model processing component is restricted by the access quantity, that is, there is an access bottleneck in the model processing component. For example, if accessing an external model to be evaluated (i.e., the model to be evaluated and the evaluation system belong to different systems), it is easily rejected by the external model when the concurrent request access volume is large. If accessing an internally privately deployed model service, the privately deployed model is easily crashed when the concurrent request access volume is large. In the embodiment of the present application, the model processing component is an independent microservice architecture. Therefore, before a model processing request reaches the model processing component, it will first pass through the gateway to determine whether it reaches the maximum concurrent number that the model to be evaluated can bear. If it exceeds the maximum concurrent number of the model to be evaluated, the access will be rejected. Currently, there are already various gateway components provided in the industry, such as apisix (a gateway component), and the specific type of the gateway component is not elaborated here in detail.

[0138] Specifically, the business processing device includes a gateway component; then, generating a model processing request corresponding to the first test case to be evaluated based on the model processing thread j and sending the model processing request to the model processing component may include the following steps: generating a model processing request through the model processing thread j and sending the model processing request to the gateway component; the gateway component counts the number of requests received for the model to be evaluated, and when the number of requests is less than the request quantity threshold, the gateway component forwards the model processing request to the model processing component.

[0139] It can be understood that the gateway component can be a component responsible for the transmission of network data, and here it can be used to control the number of requests sent to the model processing component. It can be understood that in the embodiments of the present application, since there can be multiple model processing threads processing the use cases to be evaluated in the corresponding shard task processing queue in parallel (that is, generating model processing requests corresponding to the use cases to be evaluated), the gateway component can receive the requests for the model to be evaluated sent by each model processing thread (that is, requests for calling the model to be evaluated for data processing), and thus the number of requests for the model to be evaluated received can be counted. It should be understood that the number of requests for the model to be evaluated received here refers to the number of unprocessed requests, and the number of processed requests is not counted.

[0140] Among them, the request quantity threshold can be the maximum number of requests when the gateway component rejects requests. This request quantity threshold can be determined according to the maximum request concurrency of the model to be evaluated. This request quantity threshold can be pre-configured in the gateway component or can be dynamically adjusted according to the actual request concurrency situation. Further, when the request quantity does not reach the request quantity threshold, the gateway component forwards the model processing request to the model processing component. When the request quantity reaches the request quantity threshold, the gateway component rejects the model processing request. For example, if the request quantity threshold is 10, then when the gateway component obtains that the number of requests for the model to be evaluated reaches 10, it rejects the received model processing requests. When the number of requests for the model to be evaluated obtained does not reach 10, the gateway component can forward the model processing request to the model processing component.

[0141] It can be understood that when the gateway component rejects a model processing request, it can return a prompt message for prompting the rejection of the model processing request to the service processing component. Further, the service processing component can perform subsequent retry processing based on the prompt message returned by the gateway component.

[0142] For example, model processing thread j sends a model processing request. In the configuration of the gateway component, the maximum QPS (Queries Per Second, that is, the number of transactions processed per unit time) that the configuration of the model to be evaluated indicated by the model processing request can bear is 10. After the requests received by the gateway component reach the maximum QPS of 10, the gateway component rejects the subsequent received requests, that is, returns a prompt message to the request initiator. When the requests received by the gateway component do not reach the maximum QPS of 10, the gateway component can forward the model processing request to the model processing component.

[0143] For example, please refer to ​ , ​ is a schematic diagram of a request processing process provided by an embodiment of the present application. As ​As shown, the model processing thread (such as thread 2) can determine a model processing request Q1 for the use case to be evaluated (such as the use case to be evaluated 1). Then, thread 2 can send this model processing request to the gateway component 704a (i.e., step S71a). Then, the gateway component S704a can count the number of requests received for the model to be evaluated (such as the artificial intelligence model A). When the number of requests is less than the request quantity threshold, the gateway component 704a forwards the model processing request to the model processing component 702a (i.e., step S71b). Further, the service processing component 702a can call the artificial intelligence model 1 to process the data of the use case to be evaluated 1, obtain a data processing result, and return the data processing result to the gateway component 704a (i.e., step S72a). Then, the gateway component 704a can return the data processing result to the service processing component 701a (i.e., step S72b) so that the data processing result can be sent to the test processing component 703a for testing subsequently.

[0144] It can be understood that based on the above description, in order to improve the accuracy of the evaluation of the artificial intelligence model, during model evaluation, for a use case to be evaluated, it can be executed and tested multiple times. Then, based on the test results of the data processing results of multiple executions, the final use case test result for the use case to be evaluated is determined.

[0145] Specifically, the model evaluation task carries a use case processing times threshold. Then, the embodiments of the present application may further include: counting the number of times the first use case to be evaluated is input into the model to be evaluated; when the number of use case processing times is less than the use case processing times threshold, input the first use case to be evaluated into the model to be evaluated through the model processing thread j.

[0146] Among them, the use case processing times threshold can be the maximum number of times that the use case to be evaluated needs to be repeatedly processed. It should be understood that this use case processing times threshold can be configured by the business object on the task configuration page. The number of use case processing times can be the number of times the data of the first use case to be evaluated has been processed. It should be understood that each time the first use case to be evaluated is input into the model to be evaluated for data processing, the number of use case processing times of the first use case to be evaluated can be incremented by one. When the number of use case processing times has not reached the use case processing times threshold (i.e., the number of use case processing times is less than the use case processing times threshold), the first use case to be evaluated will be input into the model to be evaluated through the model processing thread j to obtain the data processing result of the first use case to be evaluated again until the number of use case processing times reaches the use case processing times threshold. Only then is the next use case to be evaluated in the shard task processing queue of the shard task i determined as the first use case to be evaluated. It can be understood that once the data processing result of the first use case to be evaluated is processed, the data processing result obtained from this data processing will be immediately tested, thereby improving the efficiency of model testing.

[0147] S205. Obtain a test processing thread associated with the model evaluation task, call the test processing thread through the service processing component, generate a test processing request corresponding to the data processing result based on the test processing thread, and send the test processing request to the test processing component; the test processing request carries a first test case to be evaluated and the data processing result.

[0148] Among them, the test processing component can be a component for testing the data processing result generated by the model to be processed. It can be understood that this test processing component can provide a unified standard HTTP API (a form of interface) interface for the service processing component to call.

[0149] It can be understood that the test processing request can be a request for instructing the test processing component to test the data processing result. It can be understood that this test processing request carries the first test case to be evaluated and the data processing result. It can be understood that the test processing request can specifically carry the case identifier of the first test case to be evaluated. The data processing result carried in the test processing request can be written into the test processing request after being processed according to actual requirements.

[0150] For example, when the model to be evaluated is a code generation model, the data processing result generated by the model to be evaluated can be a piece of code, such as the model-generated code. The service processing component can assemble the programming language header file (i.e., the header file corresponding to the program language type in the test case to be evaluated), the main function code output by the code generation model (i.e., the data processing result), and the test code (i.e., the code for instructing the test of the data processing result, i.e., the above-mentioned case test information) into a runnable code snippet (such as the code to be tested), write the assembled runnable code snippet (such as the code to be tested) into the test processing request, and send the test processing request to the test processing component. Moreover, when the model to be evaluated is a code generation model, the code test component can include the running environments and language dependency packages of all program language types to be evaluated. Therefore, when the test processing component obtains the test processing request, it can parse the assembled code snippet from it, and the test processing component compiles the assembled code snippet to obtain the result based on the compilation execution, and determines the case test result (i.e., pass or fail the test) based on the result of the compilation execution, and then can return the case test result to the service processing component.

[0151] For example, when the model to be evaluated is a code generation model, the code generation model can generate a piece of code as the data processing result. Furthermore, the data structure of the test processing request for this data processing result can be:

[0152] {task_id; / / Identifier of the use case to be evaluated;

[0153] Test_code; / / Code executed during testing;

[0154] generation; / / Code result generated by the model;

[0155] passed; / / Whether the code execution was successful;

[0156] .......}

[0157] Among them, in the data structure of the above test processing request, task_id represents the identifier of the use case to be evaluated, the Test_code field represents the code executed during testing, that is, the above-mentioned programming language header file (i.e., the header file corresponding to the program language type in the use case to be evaluated), the main function code output by the code generation model (i.e., the data processing result), and the test code (i.e., the code used to indicate the test of the data processing result) are assembled into a runnable code snippet (such as called the code to be tested). The generation field represents the code result generated by the model (such as called the model-generated code). The passed field indicates whether the code execution was successful. If the execution is successful, the passed field returns true; otherwise, it returns false to the business processing component.

[0158] It can be understood that different artificial intelligence models have differences in request methods, parameters, and return values, and different programming languages also have different coding specifications. Therefore, in order to make the artificial intelligence model output more standardized data during model evaluation, the data input to the model (i.e., the input data in the use case to be evaluated) can be character-adjusted so that the artificial intelligence model can input more standardized results.

[0159] Specifically, each of the K use case data includes use case description information; then, based on the model evaluation task, determining the K use cases to be evaluated associated with the N use case data can include the following steps: obtaining the character processing strategy associated with the model to be evaluated, adjusting the use case description information in the N use case data based on the character processing strategy to obtain N updated use case data corresponding to the N use case data; one use case data corresponds to one updated use case data; based on the model evaluation task and the N updated use case data, determining the K use cases to be evaluated associated with the N use case data.

[0160] Among them, the use case description information can be information used to instruct the model to be evaluated to process the use case to be evaluated. For example, when the model to be evaluated is a code generation model, the use case description information can be a prompt used to instruct the model to be evaluated to generate program code. For example, through the use case description information (i.e., the prompt), the model to be evaluated is instructed to generate a piece of code for sorting. Another example is that when the model to be evaluated is a text translation model, the use case description information can be the text used to instruct the model to be evaluated to perform translation.

[0161] It should be understood that the character processing strategy can be a strategy for processing the use case description information, so that the processed use case description information can be input into the model to be evaluated for processing subsequently, which helps to standardize the results output by the model to be evaluated. Customized logic for the evaluation model and evaluation language. On the evaluation system interface, a "plugin development" function is provided to the evaluation users (such as the above-mentioned business objects). This character processing strategy can be that the evaluation object (such as the above-mentioned business object) can configure the character processing strategy on the evaluation system interface (for example, the development of the character processing strategy can be carried out using the Python language). The evaluation system will provide the intermediate data in the evaluation process (such as the use case description information of the use case to be evaluated, the data processing results generated by the model to be evaluated) to the developers. The developers can develop personalized character processing logic based on the intermediate data provided in the evaluation process. When the evaluation system runs, it will load the code of the user-defined character processing strategy to implement the user-defined character processing, so as to achieve the plugin development of the character processing logic.

[0162] Among them, the updated use case data can refer to the use case data after adjusting the use case description information based on the character processing strategy. One use case data can correspond to one updated use case data. Adjust the use case description information in N use case data based on the character processing strategy to obtain N updated use case data corresponding to the N use case data, that is, adjust the use case description information in the use case data based on the code of the character processing strategy. The process of determining K use cases to be evaluated associated with N use case data based on the model evaluation task and N updated use case data can refer to the relevant description of determining K use cases to be evaluated based on the model evaluation task and N use case data above, and will not be elaborated here.

[0163] For example, the code logic of the character processing strategy is:

[0164] {user_code = “‘

[0165] prompt = ${prompt}+‘, need to add comments to each line of code’

[0166] ”’

[0167] }

[0168] Based on the character processing strategy indicated by the above code segment, the description of "Need to add comments to each line of code" can be added after the prompt indicated by the use case description information in the use case data to obtain the updated use case description information. Furthermore, when the updated use case description information is input into the model to be evaluated, the use case to be evaluated can add comments to each line of code according to the indication of the added description, so that the use case to be evaluated can perform code commenting more standardly.

[0169] Another example is that the code logic of the character processing strategy is as follows:

[0170] {user_code = "‘

[0171] prompt = ${prompt}+‘, please complete the code’

[0172] ”’

[0173] }

[0174] Based on the character processing strategy indicated by the above code segment, the description of "Please complete the code" can be added after the prompt indicated by the use case description information in the use case data to obtain the updated use case description information. Furthermore, when the updated use case description information is input into the model to be evaluated, the use case to be evaluated can complete the code according to the indication of the added description, so that the model to be evaluated can perform code completion more standardly.

[0175] In addition, when the model to be evaluated is a code generation model, the data processing result generated by the model to be evaluated can be a piece of code, such as the model-generated code. Since different evaluation models have differences in request methods, parameters, and return values, and different programming languages also have different coding specifications, the code generated by the code generation model may contain some information that does not conform to the compilation specifications, such as some unnecessary spaces and line breaks. Therefore, when generating a test processing request for the model-generated code, it is necessary to screen out the correct code from the model return value (i.e., the model-generated code) to determine the code to be tested based on the correct code, that is, assemble the programming language header file (i.e., the header file corresponding to the program language type in the use case to be evaluated), the code screened out from the code output by the code generation model, and the test code (i.e., the code used to indicate the test of the data processing result) into the code to be tested. It can be understood that screening out the correct code from the model return value (i.e., the model-generated code) can be processed based on a pre-configured character processing strategy, and its specific processing logic can be determined according to the actual situation. This can improve the process automation degree of the entire model evaluation task, and the method of customizing the character processing strategy can help improve the standardization of intermediate data (i.e., the input data of the model and the input data of the test component, such as the code to be tested), and solve the inefficiency problem in the manual evaluation process.

[0176] It can be understood that in the business processing device, the number of test processing threads can be multiple, and then multiple test processing threads can test the data processing result of the use case to be evaluated in parallel, thereby improving the efficiency of testing the use case to be evaluated. Here, the process of testing the data processing result of the first use case to be evaluated is described.

[0177] Specifically, the number of test processing threads is P; P is a positive integer; then, calling the test processing thread through the business processing component to generate a test processing request corresponding to the data processing result can include the following steps: determining a target test processing thread for testing the data processing result from P test processing threads through the business processing component; generating a test processing request corresponding to the data processing result based on the target test processing thread.

[0178] Among them, P is the number of test processing threads, and P is a positive integer. In other words, the number of test processing threads is one or more. It can be understood that the number of test processing threads and the number of model processing threads can be the same or different, which is not limited here. Generally speaking, in order to improve the processing efficiency of the model evaluation task, the number of test processing threads can be greater than or equal to the number of model processing threads, thereby improving the efficiency of testing the use case to be evaluated and further improving the evaluation efficiency of the model evaluation task.

[0179] It can be understood that the target test processing thread can be a test processing thread used to test the data processing result of the first test case to be evaluated. It should be understood that after the above model processing component generates the data processing result of the first test case to be evaluated, it can return the data processing result to the service processing component. Therefore, the service processing component can generate a test processing request corresponding to the data processing result based on the target test processing thread.

[0180] Specifically, a model processing component is deployed in the service processing device; the data processing result is determined by the model processing component; then, the service processing component determines the target test processing thread for testing the data processing result from P test processing threads, which can include the following steps: when the service processing component obtains the data processing result, the service processing component determines the test task associated with the first test case to be evaluated and adds the test task to the test task processing queue; the test task is determined based on the data processing result; when the test task meets the task execution condition associated with the test task processing queue, the service processing component determines the target test processing thread for testing the data processing result from P test processing threads.

[0181] Among them, the test task refers to a task used to test the data processing result of the test case to be evaluated (such as the first test case to be evaluated). It should be understood that the test task is determined based on the data processing result of the first test case to be evaluated, and the test task can include the data processing result and the case identifier of the first test case to be evaluated.

[0182] Among them, the test task processing queue can be a queue storing tasks for testing the data processing results of the use cases to be evaluated. It should be understood that for each use case to be evaluated among the K use cases to be evaluated, when obtaining the test task associated with the data processing result of the use case to be evaluated, the test task can be added to the test task processing queue. It should be understood that the service processing device can process the tasks in the test task processing queue in sequence based on the order in which the test tasks are added to the test task processing queue. Moreover, after processing the test tasks in the test task processing queue, the processed test tasks (i.e., the test tasks that have completed the testing of the data processing results) can be removed from the task processing queue, so as to process the next test task. In other words, the test task processing queue processes the test tasks according to the principle of first in, first out, that is, the test tasks that are first added to the test task processing queue are processed first. For example, if there is only one test processing thread, the service processing device can process the first test task A in the test task processing queue, and after processing the test task A (i.e., completing the test through the test processing component), remove the test task A from the test task queue, and then process the next test task of the test task A (such as test task B).

[0183] Among them, the task execution condition can be the condition that needs to be met for generating the book number processing request corresponding to the test task through the test processing thread. For example, the task execution condition can be: there is an idle test processing thread, and the test task is the first task to be processed in the test task processing queue. The task to be processed here can be a task for which no test processing thread has been assigned for processing. Furthermore, when the test task meets the task execution condition, a target test processing thread is determined to process the test task, specifically, a target test processing thread can be determined from the idle test processing threads. The idle test processing threads here can refer to the test processing threads that are not used for transaction processing.

[0184] It should be understood that when a test task (such as test task A) is added to the test task processing queue, if there is no unprocessed test task in the test task queue and there are idle test processing threads, then after test task A is added to the test task processing queue, a test processing thread for this test task A can be determined immediately from the idle test processing threads to process test task A. When a test task (such as test task A) is added to the test task processing queue, if there are unprocessed test tasks in the test task queue, then after test task A is added to the test task processing queue, it is necessary to wait until test task A becomes the first task to be processed and there are idle test processing threads before test task A can be processed. Generally speaking, in order to improve the processing efficiency of model evaluation tasks, the efficiency of the test processing thread in processing test tasks can be greater than or equal to the efficiency of the data processing results generated by the model to be evaluated. Thus, it is possible to immediately perform a test after generating a data processing result, improving the efficiency of data processing.

[0185] Please refer to ​ , ​ which is a schematic flowchart of a test task processing provided by an embodiment of the present application. As ​ described, when determining the data processing result of the use case to be evaluated, the test task corresponding to the data processing result can be added to the test task processing queue (such as ​ shown as 801a in ​ ), and then the thread for processing the test task can be scheduled by the scheduling thread 802a. For example, test task 1 is processed by test processing thread 1 (such as ​ shown as 803a in ​ ), test task 2 is processed by test processing thread 2 (such as ​ shown as 804a in

[0186] S206. Perform test processing on the data processing result through the test processing component to obtain the use case test result for the first use case to be evaluated.

[0187] It can be understood that in some scenarios, testing the data processing results may have a certain negative impact on the business processing device. For example, in the code generation scenario, the model to be evaluated is a code generation model, and the data processing result determined by the code generation model can be a piece of code, which may be some code that poses a threat to the device. Therefore, when testing this piece of code (i.e., the data processing result), running this piece of code is likely to have a negative impact on the business processing device. Therefore, the test processing component can be deployed in a secure test environment to ensure the security of the test process and thus the security of model evaluation.

[0188] Specifically, the business processing device includes a first business processing device and a second business processing device; the business processing component is deployed in the first business processing device; the test processing component running in the secure test environment is deployed in the second business processing device; the second business processing device is independent of the first business processing device; then, testing the data processing result through the test processing component to obtain the use case test result for the first use case to be evaluated may include the following steps: when the second business processing device receives the test processing request sent by the first business processing device, in the secure test environment, the test processing component tests the data processing result indicated by the test processing request to obtain the use case test result for the first use case to be evaluated.

[0189] Among them, the first business processing device is the business processing device with the business processing component deployed, and the second business processing device is the business processing device with the test processing component deployed. At the architecture level, the business processing component and the test processing component are independent microservice architectures, so the business processing component and the test disposal component can be deployed on different devices.

[0190] It should be understood that the test processing component runs in the secure test environment. It can be understood that this secure test environment can be an environment for securely testing the data processing results. In other words, when testing the data processing results, in the secure test environment of the second business processing device, the test processing component tests the data processing result indicated by the test processing request to obtain the use case test result for the first use case to be evaluated. This can avoid potential security hazards in the model evaluation process and improve the security during the model evaluation process.

[0191] For example, the test processing component will be deployed to a Kubernetes (a containerized application) container cluster (i.e., the second service processing device) with separate "sandbox" (a security mechanism) capabilities to provide standard HTTP services. Specifically, only limited network access is provided in the second service processing device. Secondly, the Kubernetes container cluster of the sandbox (i.e., the second service processing device) and the normal business Kubernetes container cluster (i.e., the first service processing device) are completely independent and will not affect each other. The process of testing the data processing results is set to log in with a low-privilege user. Even if the test code goes wrong, it will not be able to affect functions that require a high-privilege threshold due to the low user privilege of the logged-in user. In the code generation scenario, when testing the code generated by the model (i.e., the data processing result), the generated code needs to be run. When the generated code is running, the execution time of the code will be detected. If it exceeds a certain time (e.g., 5 seconds), it will be killed by the management process. This can avoid the code running in an infinite loop and causing the second service processing device to crash.

[0192] For example, here, taking the model evaluation task for the code generation model as an example, the deployment process of the security test environment will be described. In the model evaluation task for the code generation model, the model processing component can also be called the code generation service, and the test processing component can also be called the code test service. The deployment of the security test environment here can include the following steps: Deploy the code test service (i.e., the test processing component) to a separately built Kubernetes container cluster (i.e., the second service processing device); Close the access permission of the external network (i.e., the wide area network, public network) of the "sandbox" Kubernetes container cluster (i.e., the second service processing device); Deploy the code test service (i.e., the test processing component) in the "sandbox" Kubernetes container cluster (i.e., the second service processing device) and start the code test service (i.e., the test processing component); The code test service (i.e., the test processing component) receives the test processing requests sent in parallel by the web background service (i.e., the service processing component); The code test service (i.e., the test processing component) parses the code to be tested and writes it into a temporary file in a fixed directory, and assigns it to a low-privilege system user, for example, writing it into the / tmp / test_code / abc.py directory; The code test service (i.e., the test processing component) calls the system command and uses the low-privilege user to execute the generated temporary file, that is, to test the code to be tested, for example, using the system command python abc.py to execute the generated temporary file; The code test service (i.e., the test processing component) obtains the execution duration of the temporary file (i.e., the duration of executing the temporary file). If it exceeds a certain time threshold, such as exceeding 5 seconds, stop the execution of the temporary file.

[0193] For example, please refer to ​ , ​This is the system architecture diagram of another model evaluation system provided by the embodiments of the present application. As ​ shown, in the service processing device 900a, it may include a service processing device 91a (i.e., the above-mentioned first service processing device) and a service processing device 92a (i.e., the above-mentioned second service processing device). In the service processing device 91a, it may include a service processing component 901a and a model processing component 902a. In the service processing device 92a, it may include a test data component 903a running in a security test environment. Furthermore, when it is necessary to test the data processing result of the use case to be evaluated, the test processing thread in the service processing device 91a needs to send a test processing request to the service processing device 92a, so that the service processing device 92a can perform tests in the security test environment.

[0194] It can be understood that the model processing component and the test processing component are two independent microservice architectures that support elastic scaling. When the capabilities of the model processing component or the test processing component are insufficient, the system processing capacity can be improved by scaling out, so as to improve the throughput of the evaluation process as much as possible, thereby making the evaluation system operate more efficiently.

[0195] Optionally, a model processing component is deployed in the service processing device; the M sharding tasks include the sharding task u; the model processing thread associated with the sharding task u is the model processing thread v, and both u and v are positive integers less than or equal to M; then, the embodiments of the present application may further include: obtaining the resource occupancy information of the service processing device; the resource occupancy information is the information of the occupied data processing resources in the data processing resources of the service processing device; when the resource occupancy information meets the component backup condition, notifying the first backup processing device to enable the backup model processing component; the backup model processing component has the same business function as the model processing component; when the second use case to be evaluated is determined from the sharding task processing queue of the sharding task u, notifying the backup model processing component to perform data processing on the second use case to be evaluated through the model processing thread v.

[0196] Among them, the sharding task u is a sharding task different from the sharding task i; the model processing thread v is the model processing thread associated with the sharding task u, and both u and v are positive integers less than or equal to M.

[0197] Among them, the resource occupancy information may be the occupancy situation of the data processing resources in the current service processing device. For example, the occupancy rate of the CPU (core processor) of the service processing device. The component backup condition may be a condition for enabling the backup model processing component in the backup device. For example, the component backup condition may be that the resource occupancy information is greater than a certain threshold, such as the CPU occupancy rate being greater than 90%. The first backup processing device may be a device deployed with the backup model processing component. The first backup processing device may be independent of the service processing device, or may be a device composed of other processing resources in the service processing device except for the data processing resources originally configured to process the model evaluation task. There is no limitation here. The backup model processing component may be a component having the same service function as the model processing component, but the backup model processing component is deployed in the first backup processing device. It can be understood that the number of the first backup processing devices may also be one or more. One or more backup model processing components may be deployed on one first backup processing device. In other words, the number of the backup model processing components may be one or more. There is no limitation here.

[0198] Among them, the second test case to be evaluated may be a test case to be processed in the shard task processing queue of the shard task u. The process of notifying the backup model processing component to perform data processing on the second test case to be evaluated through the model processing thread v may refer to the relevant description of notifying the model processing component to perform data processing on the first test case to be evaluated through the model processing thread j above. There is no need to repeat it here.

[0199] Optionally, a test processing component is deployed in the service processing device; among the K test cases to be evaluated, there is a third test case to be evaluated; the third test case to be evaluated is different from the first test case to be evaluated; then, the embodiments of the present application may further include: obtaining the resource occupancy information of the service processing device; the resource occupancy information is the information of the occupied data processing resources in the data processing resources of the service processing device; when the resource occupancy information meets the component backup condition, notifying the backup test processing component in the second backup processing device to be enabled; the backup test processing component has the same service function as the test processing component; when obtaining the data processing result for the third test case to be evaluated, notifying the backup test processing component to perform test processing on the data processing result for the third test case to be evaluated through the test processing thread, and obtaining the test case test result for the third test case to be evaluated.

[0200] Among them, the introduction of the resource occupancy information and the component backup condition may refer to the relevant description above. There is no need to repeat it here.

[0201] The second backup processing device may be a device on which a backup test processing component is deployed. The second backup processing device may be independent of the service processing device or may be a device composed of other processing resources in the service processing device except for the data processing resources originally configured to process model evaluation tasks, which is not limited herein. The second backup processing device may be the same as or different from the first backup processing device, which is not limited herein. The backup test processing component may be a component having the same service function as the test processing component, but the test model processing component is deployed in the second backup processing device.

[0202] Among them, the third test case to be evaluated may be a test case to be evaluated different from the first test case to be evaluated. The backup test processing component may be notified by the test processing thread to perform test processing on the data processing result for the third test case to be evaluated, and the relevant description of notifying the test processing component by the test processing thread to perform test processing on the data processing result for the first test case to be evaluated may be referred to, which is not elaborated herein. It can be understood that the number of second backup processing devices may also be one or more. One or more backup test processing components may be deployed on one second backup processing device. In other words, the number of backup test processing components may be one or more, which is not limited herein.

[0203] For example, in the model evaluation task of the code generation model, the model processing component can also be called the code generation service, and the test processing component can also be called the code testing service. The code generation service and the code testing service can be stateless services deployed in Kubernetes (a containerized application) (a service that does not need to save any data state and does not depend on other requests for the processing of a single request). They will accept requests for single test cases sent in parallel by multiple threads from the business processing component. The code generation service and the code testing service are located in two pods of Kubernetes (i.e., the smallest deployable computing units created and managed in Kubernetes). During actual operation, by pre-configuring the elastic scaling policy of the pod (i.e., the elastic scaling policy for the code generation service and the code testing service in the pod), the ability to automatically scale the pod horizontally is achieved. For example, when the CPU occupancy rate of the code testing service continuously reaches 90%, by increasing the number of replicas of the code testing service, that is, enabling the backup test processing component mentioned above, the throughput capacity of the code testing service is improved. The specific implementation steps are as follows: Configure the elastic scaling policies of the code generation service (i.e., a model processing component) and the code testing service (i.e., the test processing component) (i.e., the policies for enabling the backup test processing component and the backup model processing component under what conditions); The business object can submit a model evaluation task in the web background service (i.e., the business processing component) through the application client; The web background service (i.e., the business processing component) disassembles the submitted model evaluation task into test cases to be evaluated, that is, determines K test cases to be evaluated based on the case data in the evaluation dataset carried by the model evaluation task; The web background service (i.e., the business processing component) slices the test cases to be evaluated based on the concurrency ability of the model to be evaluated (i.e., the concurrent processing quantity of the model to be evaluated), that is, divides the K test cases to be evaluated into M sharding tasks; The web background service (i.e., the business processing component) starts multiple worker threads (i.e., M model processing threads) according to the number of shards (i.e., M) to send model processing requests corresponding to the test cases to be evaluated to the code generation service in parallel, so that the model to be evaluated generates the corresponding code; The web background service (i.e., the business processing component) receives the code generation data (i.e., the code generated by the code generation model), and the web background service (i.e., the business processing component) starts multiple worker threads (i.e., test processing threads) according to the number of shards (i.e., M) to send test processing requests corresponding to the code to be tested (i.e., the code assembled based on information such as the code generated by the model and the programming language header file) to the code testing service in parallel, so as to obtain the code test result (i.e., the use case test result) through the test processing component; The web background service (i.e., the business processing component) accepts the code test data and calculates the task evaluation result.

[0204] Further, please refer to ​ , ​It is a schematic flowchart of a data processing method based on an artificial intelligence model provided by an embodiment of the present application. This method can be executed by a business client. This method can at least include the following steps S301 - step S305.

[0205] S301. Display an evaluation task configuration page; the evaluation task configuration page includes a model configuration area and a dataset configuration area.

[0206] Among them, the task configuration page can refer to a page used to configure a model evaluation task. The model configuration area can be an area used to configure the model to be evaluated, and the dataset configuration area can be an area used to configure the evaluation dataset.

[0207] S302. In response to a model configuration operation in the model configuration area, display the model to be evaluated in the model configuration area; the concurrent processing volume of the model to be evaluated is used to indicate M shard tasks associated with the model evaluation task; M is a positive integer.

[0208] Among them, the model configuration operation in the model configuration area can be an operation to select the model to be evaluated in the model configuration area, such as clicking on a control for selecting a model in the model configuration area to display multiple models to be selected, and then clicking on the model that needs to be selected. Furthermore, after a business object selects a certain model, the selected model, that is, the model to be evaluated, can be displayed in the model configuration area.

[0209] S303. In response to a dataset configuration operation in the dataset configuration area, display the evaluation dataset in the dataset configuration area; the evaluation dataset includes N case data; N is a positive integer.

[0210] Among them, the dataset configuration operation in the dataset configuration area can be an operation to select the evaluation dataset in the dataset configuration area, such as clicking on a control for selecting a dataset in the dataset configuration area to display multiple datasets to be selected, and then clicking on the dataset that needs to be selected. Furthermore, after a business object selects a certain dataset, the selected dataset, that is, the evaluation dataset, can be displayed in the dataset configuration area.

[0211] S304. In response to a task confirmation operation on the task confirmation control in the evaluation task configuration page, generate a model evaluation task based on the model to be evaluated and the evaluation data set, and send the model evaluation task to the service processing device, so that the service processing device determines K test cases to be evaluated associated with N use case data based on the model evaluation task, and divides the K test cases to be evaluated into M shard tasks to obtain a shard task processing queue for the M shard tasks; the shard task processing queue is used to indicate that when a first test case to be evaluated is determined from the shard task processing queue, the first test case to be evaluated is input into the model to be evaluated through a model processing thread, and the model to be evaluated processes the data of the first test case to be evaluated to obtain a data processing result for the first test case to be evaluated.

[0212] Among them, the task confirmation control can be a control used to indicate confirmation of the task configured in the evaluation task configuration page. The task confirmation operation can be a touch operation on the task confirmation control, such as clicking the task confirmation control. The processing process of the service processing device for the model evaluation task will not be elaborated here.

[0213] For example, please refer to ​ , ​ is a schematic diagram of the effect of a task configuration page provided by an embodiment of the present application. As ​ shown, page 1101a is a task configuration page. The task configuration page may include an area for configuring the model to be evaluated (i.e., as shown in 1101a), and then select the model to be evaluated in the area shown in 1101a. The task configuration page may also include an area for selecting an evaluation scenario (i.e., as shown in 1102a), and then select the scenario to be evaluated in the area shown in 1102a, such as a code generation scenario, a code completion scenario, a code annotation scenario, etc. The task configuration page may include an area for configuring the data set of the evaluation model (i.e., as shown in 1103a), and then select the data set for evaluating the model in the area shown in 1101a. The task configuration page may include an area for configuring the programming language type of the program statements generated by the model (i.e., as shown in 1104a), and then select the programming language type of the program statements generated by the model in the area shown in 1104a.

[0214] The task configuration page may include an area for evaluation result metrics (i.e., as shown in 1105a), and then select evaluation result metrics (also referred to as evaluation metrics) in the area shown in 1105a. For example, as ​As shown, the result indicators to be evaluated can be selected from Indicator 1 and Indicator 2. For Indicator 1, specific parameters can be configured in the area shown in 1106a, such as Parameter 1, Parameter 2, Parameter 3, Parameter 4, etc.; for Indicator 2, specific parameters can be configured in the area shown in 1107a, such as Parameter 1, Parameter 2, Parameter 3, Parameter 4, etc. For example, Parameter 1 here can be the use case processing times threshold mentioned above. Parameter 2 can be Top-k. Top-k is a parameter used to control the number of context words considered by the model when generating an answer (i.e., the data processing result). When the model predicts the next word, it will consider the previous k words. This parameter helps limit the number of words the model needs to consider, thus accelerating the calculation and reducing the length of the generated answer. Usually, the value of Top-k ranges from 10 to 5, depending on the specific requirements of the task. Parameter 3 can be Top-p. Top-p is another commonly used ChatGLM parameter. It controls which words the model should prioritize when generating an answer (i.e., the data processing result). The value of the Top-p parameter ranges between 0 and 1. A high Top-p value means the model will prioritize high-quality words, while a low Top-p value means the model will consider generating different words more. By adjusting the Top-p parameter, the diversity and quality of the model can be balanced. Parameter 4 can be Temperature. Temperature (i.e., the temperature parameter) is an important factor controlling the way the model generates an answer (i.e., the data processing result). The temperature parameter can affect the diversity and certainty of the generated answer. A high temperature value means the model will generate less common words and more creative answers with a higher probability. On the contrary, a low temperature value means the model will be more conservative and tend to generate more common and more certain answers. By adjusting the temperature parameter, the diversity and accuracy of the model can be balanced.

[0215] On this task configuration page, there can be an area (shown in 1108a) for describing the evaluation plan (i.e., the model evaluation task), and then information for describing the configured model evaluation task can be entered in area 1105a. On this task configuration page, there can be a control (shown in 1109a) for confirming the information configured on the page, that is, the above-mentioned task confirmation control. Then, when the business object clicks on control 1109a, a model evaluation task can be generated based on the information configured on the task configuration page, so that the business client can send the model evaluation task to the business processing device, such as determining the model evaluation request corresponding to the model evaluation task and sending the model evaluation request to the business processing device.

[0216] S305. Receive the use case test result of the first use case to be evaluated and display the use case test result; the use case test result is obtained by the business processing device through testing the data processing result based on the test processing thread.

[0217] Among them, the determination method of the use case test result can refer to the above description and will not be elaborated here.

[0218] It can be understood that displaying the use case test result means that it can show whether the first use case to be evaluated passes the test. In addition, displaying the use case test result can show the quantity of the obtained use case test results, that is, the quantity of the use cases to be evaluated for which the data processing and the test process have been completed currently.

[0219] Optionally, display the evaluation task execution page; the evaluation task progress is displayed in the evaluation task execution page; the evaluation task progress is determined based on the quantity of K use cases to be evaluated and the quantity of the use cases to be evaluated for which the use case test results have been obtained.

[0220] Among them, the evaluation task execution page can be a page for displaying the execution status of the model evaluation task. The evaluation task progress can be the progress of the model evaluation task execution. The evaluation task progress is determined based on the quantity of K use cases to be evaluated and the quantity of the use cases to be evaluated for which the use case test results have been obtained. The quantity of the use cases to be evaluated for which the use case test results have been obtained can be the use cases to be evaluated for which the test has been completed currently. For example, the total quantity of the use cases to be evaluated is 164. If the quantity of the use cases to be evaluated for which the test has been completed currently is 32, the current evaluation task progress is 32 / 164, that is, 19.5%.

[0221] Optionally, in the evaluation task execution page, the status of each use case to be evaluated corresponding to the model evaluation task can also be displayed. For example, the use case to be evaluated is in the to-be-processed state (that is, the state of not having undergone data processing and testing), the state of undergoing data processing through the to-be-evaluated model (also called the model processing state), the state of undergoing testing (also called the testing state), the state of having completed the test (also called the use case completion state), and so on. And when in the use case completion state, the use case test result of the corresponding use case to be evaluated can be displayed.

[0222] Please refer to ​ , ​It is a schematic diagram showing the effect of an evaluation task execution page provided by an embodiment of the present application. Page 1200a can be an evaluation task execution page. In this evaluation task execution page, there can be a region for displaying the progress of the evaluation task (as shown in 1201a). In the region shown in 1201a, the current evaluation task progress can be displayed in a visual manner. For example, the evaluation task progress in the region shown in 1201a is 32 / 164. In this evaluation task execution page, there can be a region for displaying the execution status of the current model evaluation task (as shown in 1202a). For example, the execution status can be in progress, completed, paused, etc. In this evaluation task execution page, there can be a region for displaying the task identifier of the current model evaluation task (as shown in 1203a). In this evaluation task execution page, there can be a control for pausing the model evaluation task (as shown in 1204a). When the business object clicks on this control 1204, the execution of model evaluation task 1 can be paused. In this evaluation task execution page, there can be a control for exiting this evaluation task execution page (as shown in 1205a). When the business object clicks on this control 1204, this evaluation task execution page can be exited.

[0223] In addition, in this evaluation task execution page, there can be a region for displaying the current model evaluation task information (as shown in 1206a). For example, it can display the model name, evaluation scenario, data set, task creation time, completion time, language type of the program statements used for generation, task description information, etc. of the current model evaluation task.

[0224] Optionally, in this evaluation task execution page, there can also be a region for displaying the recently executed tasks (as shown in 1207a). For example, it can display the task identifier, task initiator, task creation time, task execution progress, etc. of the model evaluation tasks executed within the most recent week. In this evaluation task execution page, there can also be a region for displaying the task evaluation results of the most recently executed task (as shown in 1208a). For example, it can display the results of the most recently executed model evaluation task on metrics 1 and 2.

[0225] Optionally, in this evaluation task execution page, there can also be a control for indicating the page for displaying more historically executed model evaluation tasks (as shown in 1209a). When the business object clicks on this control 1209, a historical task viewing page can be displayed, that is, a page for displaying more historically executed model evaluation tasks. This historical task viewing page can be used to view the historically executed model evaluation tasks. Optionally, in this historical task viewing page, some eligible model evaluation tasks can also be filtered and displayed based on some options.

[0226] For example, please refer to​ , ​ is a schematic diagram showing the effect of a historical task viewing page provided by an embodiment of the present application. Page 1300a can be a historical task viewing page. In this historical task viewing page, controls for filtering model evaluation tasks can be displayed. For example, a control 1301a for filtering data sets, a control 1302a for filtering language types, a control 1303a for filtering the creators of tasks, a control 1304a for filtering the time when the evaluation task is completed, and so on. Further, when corresponding filtering conditions are selected, the historically executed model evaluation tasks obtained by filtering can be displayed in the task display area (such as the area shown as 1305a in ​ ). In this task display area, information about each historically executed model evaluation task can be displayed. For example, model name, evaluation scenario, data set, creation time of the task, completion time, language type of the program statements used for generation, description information of the task, etc., which are not limited here. Optionally, in the historical task viewing page, a control for viewing the task evaluation result (also called an evaluation report) of each historical model evaluation task (i.e., the historically executed model evaluation task) can be displayed. For example, the control shown as 1306a. Furthermore, when a business object clicks on this control 1306a, the task evaluation result of the corresponding model evaluation task can be viewed.

[0227] Please refer to ​ , ​ is a schematic diagram showing the processing time sequence of a model evaluation process provided by an embodiment of the present application. This model evaluation process can involve a front-end evaluation processing component, a data set processing sub-component, a model management sub-component, an evaluation processing sub-component, a thread pool, a model to be evaluated, a test processing component, a database, and so on. Among them, the front-end evaluation processing component is the above-mentioned application client, and the data set processing sub-component, the model management sub-component, the evaluation processing sub-component, the thread pool, etc. correspond to the above-mentioned business processing components. The model to be evaluated is called by the above-mentioned model processing component. The test processing component is the above-mentioned test processing component, which will not be elaborated here. The database can be used to store data during the model evaluation process. This database can be deployed in the business processing device or in other devices independent of the business processing device, which is not limited here.

[0228] Among them, the business object A can select a data set through the evaluation processing front end (i.e., the application client) (i.e., step S1.1). The evaluation processing front end pulls the data set list from the data set processing sub-component (i.e., step S1.2). Then, the data set processing sub-component returns the data set list to the evaluation processing front end (i.e., step S1.3). That is to say, when the business object clicks the control for selecting the data set through the application client, the application client pulls the data set list for display so that the business object can select the evaluation data set from the data set list. Among them, the data set list can be a list of data sets available for model evaluation. In addition, the business object A selects a model through the evaluation processing front end (i.e., the application client) (i.e., step S2.1). The evaluation processing front end pulls the model list and model parameters from the model management sub-component (i.e., step S2.2). Then, the model management sub-component returns the model list and model parameters to the evaluation processing front end (i.e., step S2.3). That is to say, when the business object clicks the control for selecting the model through the application client, the application client pulls the model list for display so that the business object can select the model to be evaluated from the model list. Among them, the model list can be a list of models available for selection for evaluation, and the model parameters can refer to the parameters for evaluating the model (such as parameter 1, parameter 2, parameter 3, etc. in the above ​ . Further, the business object A submits an evaluation plan through the evaluation processing front end (i.e., step S3.1). This evaluation plan can be the above-mentioned model evaluation task. For example, the business object A clicks the control for indicating the confirmation of the model evaluation task (such as the control 1109a in the above ​ ). The evaluation processing front end sends a model evaluation request to the evaluation processing sub-component (i.e., step S3.2). This model evaluation request can carry the model evaluation task (i.e., the evaluation plan). The evaluation processing sub-component can obtain the detailed data set information from the data set processing sub-component (step S3.2.1). Then, the evaluation processing sub-component can assemble the data to be evaluated (i.e., step S3.2.2), that is, the evaluation processing sub-component can determine multiple evaluation cases based on the evaluation data set. In addition, the evaluation processing sub-component can save the evaluation plan in the database (step S3.3) so that the subsequent business line can query the evaluation plan.

[0229] Further, the evaluation processing sub-component can apply for a thread pool, submit and shard model evaluation tasks (step S3.4), that is, create model processing threads corresponding to M shard tasks, divide the test cases to be evaluated into M task shards, and the thread pool requests model generation code from the model to be evaluated through the model processing threads (such as the above-mentioned model processing thread j) (that is, step S3.5.1), that is, generate a model processing request through the model processing thread, and send the model processing request to the model processing component to call the model to be evaluated for data processing to generate corresponding code (that is, the data processing result). Then, the model to be evaluated returns the code to the thread (that is, step S3.5.2). In the thread pool, a thread used to process the generated code can process the generated code based on a character processing strategy (that is, step S3.6), that is, determine the useful part from the generated code to facilitate subsequent testing of the determined useful part of the code. In addition, the business processing component can save the model code and status to the database (that is, step S3.7). Here, the status can refer to the task status of the model evaluation task, and the task status of the model evaluation task can include the processing status of each test case associated with the model evaluation task, such as the data processing of the test case to be evaluated by the model (that is, generating the code corresponding to the test case to be evaluated), and the completion of the test of the data processing result of the test case to be evaluated, etc. Further, the test processing thread in the thread pool can request evaluation test code from the test processing component (that is, step S3.8), that is, generate a test processing request through the test processing thread, and send the test processing request to the test processing component to test the generated code corresponding to the test case to be evaluated. Then, the test processing component can return the test result after testing the code (that is, step S3.9). In addition, the test result and status can be saved to the database (that is, step S3.10). Here, the status can be the above-mentioned task status of the model evaluation task. Further, the threads in the thread pool (such as the model processing thread and the test processing thread) can determine whether all shard tasks are completed (that is, step S3.11), that is, determine whether the test cases in all shard tasks have generated corresponding data processing results, and the test of the data processing results is completed. When the evaluation in the thread pool is completed, it notifies the evaluation processing front end (that is, step S3.12), that is, prompts the business object that the model evaluation task has been completed.

[0230] In addition, the business object A can initiate a request to view the evaluation task through the evaluation processing front end (i.e., step S4.1). The evaluation processing front end sends a view request to the evaluation processing sub-component (i.e., step S4.2). The evaluation processing front end views the task status from the database (i.e., step S4.3). Subsequently, the evaluation processing front end can assemble the evaluation result (i.e., step S4.4) and return the evaluation result to the evaluation processing front end (i.e., step S4.5). The evaluation result here is the above-mentioned task evaluation result, which may include the index values of specified task metrics, the use case test results for each use case to be evaluated, and so on.

[0231] Please refer to ​ , ​ which is a schematic diagram of a task processing process provided by an embodiment of the present application. As ​ shown, during the entire model evaluation process, it may include model evaluation task submission 1501a, model evaluation task execution 1502a, and model evaluation task viewing 1503a. The processing processes of model evaluation task submission and model evaluation task viewing can refer to the relevant descriptions in ​ the embodiment. The model evaluation task execution can refer to the relevant descriptions in the above ​ and ​ embodiments, and will not be elaborated here.

[0232] Please refer to ​ , ​ which is a schematic structural diagram of a data processing device based on an artificial intelligence model provided by an embodiment of the present application. As ​ shown, the data processing device 1 based on the artificial intelligence model can be a computer program (including program code) running on a business processing device (for example, the business processing device 21a in the above ​ ). For example, the data processing device 1 based on the artificial intelligence model is an application software. It can be understood that the data processing device 1 based on the artificial intelligence model can be used to execute the corresponding steps in the data processing method provided by the embodiments of the present application. As ​ shown, the data processing device 1 based on the artificial intelligence model may include: a task receiving module 11, a use case determination module 12, a task sharding module 13, a model processing module 14, and a testing module 15;

[0233] The task receiving module 11 is configured to receive a model evaluation task submitted by a business object. The model evaluation task is determined by the business object based on the evaluation data set and the model to be evaluated configured on the evaluation task configuration page. The evaluation data set includes N use case data. N is a positive integer. The concurrent processing capacity of the model to be evaluated is used to indicate M sharding tasks associated with the model evaluation task. M is a positive integer;

[0234] A use case determination module 12, configured to determine K to-be-evaluated use cases associated with N use case data based on a model evaluation task; K is a positive integer multiple of N;

[0235] A task sharding module 13, configured to create M model processing threads associated with M sharding tasks based on the number of parallel processes, divide the K to-be-evaluated use cases into M sharding tasks, and obtain sharding task processing queues for the M sharding tasks; one sharding task corresponds to one model processing thread; among the M sharding tasks, there is a sharding task i, and the model processing thread associated with the sharding task i is the model processing thread j, where both i and j are positive integers less than or equal to M;

[0236] A model processing module 14, configured to determine a first to-be-evaluated use case among the to-be-evaluated use cases included in the sharding task processing queue of the sharding task i, input the first to-be-evaluated use case into the to-be-evaluated model through the model processing thread j, and have the to-be-evaluated model perform data processing on the first to-be-evaluated use case to obtain a data processing result for the first to-be-evaluated use case;

[0237] A testing module 15, configured to obtain a test processing thread associated with the model evaluation task, and perform test processing on the data processing result based on the test processing thread to obtain a use case test result for the first to-be-evaluated use case.

[0238] Among them, a business processing component and a model processing component are deployed in the business processing device;

[0239] The model processing module 14 is specifically configured to:

[0240] Call the model processing thread j in the business processing component, generate a model processing request corresponding to the first to-be-evaluated use case based on the model processing thread j, and send the model processing request to the model processing component; the model processing request carries the first to-be-evaluated use case;

[0241] Input the first to-be-evaluated use case into the to-be-evaluated model through the model processing component.

[0242] Among them, a gateway component is included in the business processing device;

[0243] The model processing module 14 is specifically configured to:

[0244] Generate a model processing request through the model processing thread j, and send the model processing request to the gateway component;

[0245] The gateway component counts the number of requests received for the to-be-evaluated model, and when the number of requests is less than the request quantity threshold, the gateway component forwards the model processing request to the model processing component.

[0246] Among them, the model evaluation task carries a use case processing times threshold;

[0247] Among them, the data processing device based on the artificial intelligence model further includes: a use case input module 16;

[0248] The use case input module 16 is specifically used for:

[0249] Count the number of use case processing times when the first use case to be evaluated is input into the model to be evaluated;

[0250] When the number of use case processing times is less than the use case processing times threshold, input the first use case to be evaluated into the model to be evaluated through the model processing thread j.

[0251] Among them, a business processing component and a test processing component are deployed in the business processing device;

[0252] The test module 15 is specifically used for:

[0253] Call the test processing thread through the business processing component, generate a test processing request corresponding to the data processing result based on the test processing thread, and send the test processing request to the test processing component; the test processing request carries the first use case to be evaluated and the data processing result;

[0254] Perform test processing on the data processing result through the test processing component to obtain a use case test result for the first use case to be evaluated.

[0255] Among them, the number of test processing threads is P; P is a positive integer;

[0256] The test module 15 is specifically used for:

[0257] Determine a target test processing thread for testing the data processing result from P test processing threads through the business processing component;

[0258] Generate a test processing request corresponding to the data processing result based on the target test processing thread.

[0259] Among them, a model processing component is deployed in the business processing device; the data processing result is determined by the model processing component;

[0260] The test module 15 is specifically used for:

[0261] When obtaining the data processing result through the business processing component, determine a test task associated with the first use case to be evaluated through the business processing component, and add the test task to the test task processing queue; the test task is determined based on the data processing result;

[0262] When the test task meets the task execution condition associated with the test task processing queue, the business processing component determines a target test processing thread for the test data processing result from P test processing threads.

[0263] Among them, the business processing device includes a first business processing device and a second business processing device; the business processing component is deployed in the first business processing device; the test processing component running in the security test environment is deployed in the second business processing device; the second business processing device is independent of the first business processing device;

[0264] The test module 15 is specifically used for:

[0265] When the second business processing device obtains the test processing request sent by the first business processing device, in the security test environment, the test processing component performs test processing on the data processing result indicated by the test processing request to obtain the use case test result for the first to-be-evaluated use case.

[0266] Among them, the model processing component is deployed in the business processing device; the shard task u is included in M shard tasks; the model processing thread associated with the shard task u is the model processing thread v, and both u and v are positive integers less than or equal to M;

[0267] Among them, the data processing device 1 based on the artificial intelligence model further includes: a first backup component module 17;

[0268] The first backup component module 17 is specifically used for:

[0269] Obtain the resource occupancy information of the business processing device; the resource occupancy information is the information of the occupied data processing resources in the data processing resources of the business processing device;

[0270] When the resource occupancy information meets the component backup condition, notify the first backup processing device to enable the backup model processing component; the backup model processing component has the same business function as the model processing component;

[0271] When the second to-be-evaluated use case is determined from the shard task processing queue of the shard task u, notify the backup model processing component to perform data processing on the second to-be-evaluated use case through the model processing thread v.

[0272] Among them, the test processing component is deployed in the business processing device; the third to-be-evaluated use case is included in K to-be-evaluated use cases; the third to-be-evaluated use case is different from the first to-be-evaluated use case;

[0273] Among them, the data processing device 1 based on the artificial intelligence model further includes: a second backup component module 18;

[0274] The second backup component module 18 is specifically used for:

[0275] Obtain the resource occupancy information of the service processing device; the resource occupancy information is the information of the occupied data processing resources in the data processing resources of the service processing device;

[0276] When the resource occupancy information meets the component backup condition, notify the backup test processing component in the second backup processing device to be enabled; the backup test processing component has the same service function as the test processing component;

[0277] When the data processing result for the third to-be-evaluated use case is obtained, notify the backup test processing component to perform test processing on the data processing result for the third to-be-evaluated use case through the test processing thread, and obtain the use case test result for the third to-be-evaluated use case.

[0278] Among them, the to-be-evaluated model is a program statement generation model; the model evaluation task carries the program language type associated with the program statement generation model; the number of types of the program language type is L, and L is a positive integer;

[0279] The use case determination module 12 is specifically used for:

[0280] Based on the number of types of the program language type and N use case data, determine K to-be-evaluated use cases associated with the N use case data; K = L * N; a to-be-evaluated use case is obtained based on one program language type and one use case data.

[0281] Among them, each of the K use case data includes use case description information;

[0282] The use case determination module 12 is specifically used for:

[0283] Obtain the character processing strategy associated with the to-be-evaluated model, and adjust the use case description information in the N use case data based on the character processing strategy to obtain N updated use case data corresponding to the N use case data; one use case data corresponds to one updated use case data;

[0284] Based on the model evaluation task and the N updated use case data, determine K to-be-evaluated use cases associated with the N use case data.

[0285] Among them, the data processing device 1 based on the artificial intelligence model further includes: a task evaluation result determination module 19;

[0286] The task evaluation result determination module 19 is specifically used for:

[0287] When the use case test results for the K to-be-evaluated use cases are obtained, based on the use case test results for the K to-be-evaluated use cases, determine the task evaluation result for the model evaluation task.

[0288] Please refer to​ , ​ is a schematic structural diagram of a data processing device based on an artificial intelligence model provided by an embodiment of the present application. As ​ shown, the data processing device 2 based on the artificial intelligence model may be a computer program (including program code) running on an application client (for example, the application client in the terminal device 20a described above ​ ), for example, the data processing device 2 based on the artificial intelligence model is an application software; it can be understood that the data processing device 2 based on the artificial intelligence model can be used to execute the corresponding steps in the data processing method provided by the embodiment of the present application. As ​ shown, the data processing device 2 based on the artificial intelligence model may include: a configuration page display module 21, a model configuration module 22, a data set configuration module 23, a task generation module 24, and a result receiving module 25;

[0289] The configuration page display module 21 is configured to display an evaluation task configuration page; the evaluation task configuration page includes a model configuration area and a data set configuration area;

[0290] The model configuration module 22 is configured to, in response to a model configuration operation on the model configuration area, display a model to be evaluated in the model configuration area; the concurrent processing amount of the model to be evaluated is used to indicate M shard tasks associated with the model evaluation task; M is a positive integer;

[0291] The data set configuration module 23 is configured to, in response to a data set configuration operation on the data set configuration area, display an evaluation data set in the data set configuration area; the evaluation data set includes N case data; N is a positive integer;

[0292] The task generation module 24 is configured to, in response to a task confirmation operation on the task confirmation control in the evaluation task configuration page, generate a model evaluation task based on the model to be evaluated and the evaluation data set, and send the model evaluation task to the service processing device, so that the service processing device determines K cases to be evaluated associated with the N case data based on the model evaluation task, divides the K cases to be evaluated into M shard tasks, and obtains a shard task processing queue of the M shard tasks; the shard task processing queue is used to indicate that when a first case to be evaluated is determined from the shard task processing queue, the first case to be evaluated is input into the model to be evaluated through a model processing thread, and the model to be evaluated processes the data of the first case to be evaluated to obtain a data processing result for the first case to be evaluated;

[0293] The result receiving module 25 is configured to receive the case test result of the first case to be evaluated and display the case test result; the case test result is obtained by the service processing device performing a test process on the data processing result based on a test processing thread.

[0294] Please refer to ​ , ​ which is a schematic structural diagram of a computer device provided by an embodiment of the present application. As ​ shown, the computer device 1000 may include: a processor 1001, a network interface 1004, and a memory 1005. In addition, the above computer device 1000 may further include: a user interface 1003, and at least one communication bus 1002. Among them, the communication bus 1002 is used to realize the connection and communication between these components. Among them, the user interface 1003 may include a display screen (Display) and a keyboard (Keyboard). Optionally, the user interface 1003 may further include a standard wired interface and a wireless interface. The network interface 1004 may optionally include a standard wired interface and a wireless interface (such as a WI-FI interface). The memory 1005 may be a high-speed RAM memory or a non-volatile memory, such as at least one disk memory. Optionally, the memory 1005 may also be at least one storage device located far from the aforementioned processor 1001. As ​ shown, the memory 1005, as a computer-readable storage medium, may include an operating system, a network communication module, a user interface module, and a device control application program.

[0295] In the computer device 1000 as ​ shown, the network interface 1004 can provide network communication functions; while the user interface 1003 is mainly used to provide an input interface for users; and the processor 1001 can be used to call the device control application program stored in the memory 1005 to execute the description of the data processing method in any one of the foregoing corresponding embodiments, which will not be elaborated here. In addition, the description of the beneficial effects of adopting the same method will not be elaborated either.

[0296] In addition, it should be pointed out here that: The embodiment of the present application also provides a computer-readable storage medium, and the computer-readable storage medium stores computer programs executed by the data processing device 1 and the data processing device 2 based on the artificial intelligence model mentioned above, and the computer program includes program instructions. When the processor executes the program instructions, it can execute the description of the data processing method in the foregoing embodiments, so it will not be elaborated here. In addition, the description of the beneficial effects of adopting the same method will not be elaborated either. For the technical details not disclosed in the embodiment of the computer-readable storage medium involved in the present application, please refer to the description of the method embodiment of the present application.

[0297] The above computer-readable storage medium may be the data processing device provided in any of the foregoing embodiments or the internal storage unit of the above computer device, such as the hard disk or memory of the computer device. The computer-readable storage medium may also be an external storage device of the computer device, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the computer device. Further, the computer-readable storage medium may also include both the internal storage unit and the external storage device of the computer device. The computer-readable storage medium is used to store the computer program and other programs and data required by the computer device. The computer-readable storage medium may also be used to temporarily store the data that has been output or is to be output.

[0298] In addition, it should be noted here that: The embodiments of the present application also provide a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the method provided in any of the foregoing corresponding embodiments. In addition, the description of the beneficial effects of using the same method will not be repeated. For the technical details not disclosed in the computer program product or the computer program embodiments involved in the present application, please refer to the description of the method embodiments of the present application.

[0299] In the embodiments of the present application, the term "module" or "unit" refers to a computer program with a predetermined function or a part of a computer program, which works together with other related parts to achieve a predetermined goal, and can be fully or partially implemented by using software, hardware (such as a processing circuit or a memory), or a combination thereof. Similarly, one processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of the overall module or unit that includes the function of the module or unit.

[0300] In the description of the embodiments of the present application, the terms "first", "second", etc. in the specification, claims and drawings are used to distinguish different objects, rather than to describe a specific order. In addition, the term "comprising" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, device, product or equipment that includes a series of steps or units is not limited to the listed steps or modules, but optionally further includes steps or modules not listed, or optionally further includes other step units inherent to these processes, methods, devices, products or equipment.

[0301] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.

[0302] The foregoing disclosure is only for the preferred embodiments of this application, and of course it cannot be used to limit the scope of rights of this application. Therefore, equivalent changes made according to the claims of this application still fall within the scope covered by this application.

Claims

1. A data processing method based on an artificial intelligence model, characterized in that, The method is executed by a service processing device; the method includes: Receiving a model evaluation task submitted by a service object; the model evaluation task is determined by the service object based on an evaluation data set and a model to be evaluated configured on an evaluation task configuration page; the evaluation data set includes N case data; N is a positive integer; the concurrent processing volume of the model to be evaluated is used to indicate M shard tasks associated with the model evaluation task; M is a positive integer; Determining K cases to be evaluated associated with the N case data based on the model evaluation task; K is a positive integer multiple of N; Creating M model processing threads associated with the M shard tasks based on the parallel processing quantity, dividing the K cases to be evaluated into the M shard tasks to obtain shard task processing queues of the M shard tasks; one shard task corresponds to one model processing thread; the M shard tasks include a shard task i, and the model processing thread associated with the shard task i is a model processing thread j, and both i and j are positive integers less than or equal to M; Determining a first case to be evaluated among the cases to be evaluated included in the shard task processing queue of the shard task i, inputting the first case to be evaluated into the model to be evaluated through the model processing thread j, and performing data processing on the first case to be evaluated by the model to be evaluated to obtain a data processing result for the first case to be evaluated; Obtaining a test processing thread associated with the model evaluation task, and performing test processing on the data processing result based on the test processing thread to obtain a case test result for the first case to be evaluated.

2. The method according to claim 1, characterized in that A service processing component and a model processing component are deployed in the service processing device; The step of inputting the first case to be evaluated into the model to be evaluated through the model processing thread j includes: Invoking the model processing thread j through the service processing component, generating a model processing request corresponding to the first case to be evaluated based on the model processing thread j, and sending the model processing request to the model processing component; the model processing request carries the first case to be evaluated; Inputting the first case to be evaluated into the model to be evaluated through the model processing component.

3. The method according to claim 2, wherein A gateway component is included in the service processing device; The step of generating a model processing request corresponding to the first case to be evaluated based on the model processing thread j and sending the model processing request to the model processing component includes: Generating a model processing request through the model processing thread j and sending the model processing request to the gateway component; The gateway component counts the number of requests received for the model to be evaluated, and when the number of requests is less than a request quantity threshold, the gateway component forwards the model processing request to the model processing component.

4. The method according to claim 1, characterized in that, The model evaluation task carries a case processing times threshold; The method further includes: Counting the number of times the first case to be evaluated is input into the model to be evaluated. When the number of use case processing times is less than the use case processing times threshold, input the first use case to be evaluated into the model to be evaluated through the model processing thread j.

5. The method according to claim 1, characterized in that A service processing component and a test processing component are deployed in the service processing device; Performing test processing on the data processing result based on the test processing thread to obtain a use case test result for the first use case to be evaluated, including: Invoking the test processing thread through the service processing component, generating a test processing request corresponding to the data processing result based on the test processing thread, and sending the test processing request to the test processing component; the test processing request carries the first use case to be evaluated and the data processing result; Performing test processing on the data processing result through the test processing component to obtain a use case test result for the first use case to be evaluated.

6. The method according to claim 5, wherein The number of the test processing threads is P; P is a positive integer; The step of invoking the test processing thread through the service processing component and generating a test processing request corresponding to the data processing result based on the test processing thread includes: Determining a target test processing thread for testing the data processing result from the P test processing threads through the service processing component; Generating a test processing request corresponding to the data processing result based on the target test processing thread.

7. The method according to claim 6, wherein A model processing component is deployed in the service processing device; the data processing result is determined by the model processing component; The step of determining a target test processing thread for testing the data processing result from the P test processing threads through the service processing component includes: When the data processing result is obtained through the service processing component, determining a test task associated with the first use case to be evaluated through the service processing component, and adding the test task to a test task processing queue; the test task is determined based on the data processing result; When the test task meets the task execution condition associated with the test task processing queue, determining a target test processing thread for testing the data processing result from the P test processing threads through the service processing component.

8. The method according to claim 5, characterized in that, The service processing device includes a first service processing device and a second service processing device; the service processing component is deployed in the first service processing device; the test processing component running in a security test environment is deployed in the second service processing device; the second service processing device is independent of the first service processing device; The step of performing test processing on the data processing result through the test processing component to obtain a use case test result for the first use case to be evaluated includes: When the second service processing device obtains a test processing request sent by the first service processing device, in the security test environment, performing test processing on the data processing result indicated by the test processing request through the test processing component to obtain a use case test result for the first use case to be evaluated.

9. The method according to claim 1, characterized in that, A model processing component is deployed in the service processing device; the M shard tasks include shard task u; the model processing thread associated with shard task u is model processing thread v, and both u and v are positive integers less than or equal to M; The method further includes: Obtaining resource occupancy information of the service processing device; the resource occupancy information is information on the occupied data processing resources in the data processing resources of the service processing device; When the resource occupancy information meets the component backup condition, notifying the first backup processing device to enable a backup model processing component; the backup model processing component has the same service function as the model processing component; When a second test case to be evaluated is determined from the shard task processing queue of shard task u, notifying the backup model processing component to perform data processing on the second test case to be evaluated through model processing thread v.

10. The method according to claim 1, characterized in that, A test processing component is deployed in the service processing device; the K test cases to be evaluated include a third test case to be evaluated; the third test case to be evaluated is different from the first test case to be evaluated; The method further includes: Obtaining resource occupancy information of the service processing device; the resource occupancy information is information on the occupied data processing resources in the data processing resources of the service processing device; When the resource occupancy information meets the component backup condition, notifying the second backup processing device to enable a backup test processing component; the backup test processing component has the same service function as the test processing component; When the data processing result for the third test case to be evaluated is obtained, notifying the backup test processing component to perform test processing on the data processing result for the third test case to be evaluated through the test processing thread to obtain a test result for the third test case to be evaluated.

11. The method according to claim 1, characterized in that, The model to be evaluated is a program statement generation model; the model evaluation task carries a program language type associated with the program statement generation model; the number of types of the program language type is L, and L is a positive integer; Determining the K test cases to be evaluated associated with the N case data based on the model evaluation task includes: Determining the K test cases to be evaluated associated with the N case data based on the number of types of the program language type and the N case data; K = L * N; one test case to be evaluated is obtained based on one program language type and one case data.

12. The method according to claim 1, characterized in that, Each of the K case data includes case description information; Determining the K test cases to be evaluated associated with the N case data based on the model evaluation task includes: Obtaining a character processing strategy associated with the model to be evaluated, and adjusting the case description information in the N case data based on the character processing strategy to obtain N updated case data corresponding to the N case data; one case data corresponds to one updated case data; Determining the K test cases to be evaluated associated with the N case data based on the model evaluation task and the N updated case data.

13. The method according to claim 1, characterized in that, The method further includes: When obtaining the test results of the use cases for the K to-be-evaluated use cases, based on the test results of the use cases for the K to-be-evaluated use cases, determine the task evaluation result for the model evaluation task.

14. A data processing method based on an artificial intelligence model, characterized in that, The method is executed by a business client, and the method includes: Display an evaluation task configuration page; the evaluation task configuration page includes a model configuration area and a dataset configuration area; In response to a model configuration operation in the model configuration area, display a to-be-evaluated model in the model configuration area; the concurrent processing volume of the to-be-evaluated model is used to indicate M shard tasks associated with the model evaluation task; M is a positive integer; In response to a dataset configuration operation in the dataset configuration area, display an evaluation dataset in the dataset configuration area; the evaluation dataset includes N use case data; N is a positive integer; In response to a task confirmation operation on the task confirmation control in the evaluation task configuration page, generate a model evaluation task based on the to-be-evaluated model and the evaluation dataset, and send the model evaluation task to the business processing device, so that the business processing device determines K to-be-evaluated use cases associated with the N use case data based on the model evaluation task, and divides the K to-be-evaluated use cases into the M shard tasks to obtain a shard task processing queue of the M shard tasks; the shard task processing queue is used to indicate that when a first to-be-evaluated use case is determined from the shard task processing queue, input the first to-be-evaluated use case into the to-be-evaluated model through a model processing thread, and the to-be-evaluated model performs data processing on the first to-be-evaluated use case to obtain a data processing result for the first to-be-evaluated use case; Receive the test result of the first to-be-evaluated use case and display the test result; the test result of the use case is obtained by the business processing device performing test processing on the data processing result based on a test processing thread.

15. A data processing device based on an artificial intelligence model, characterized in that, The device runs on a business processing device; the device includes: A task receiving module, configured to receive a model evaluation task submitted by a business object; the model evaluation task is determined by the business object based on an evaluation dataset and a to-be-evaluated model configured on an evaluation task configuration page; the evaluation dataset includes N use case data; N is a positive integer; the concurrent processing volume of the to-be-evaluated model is used to indicate M shard tasks associated with the model evaluation task; M is a positive integer; A use case determination module, configured to determine K to-be-evaluated use cases associated with the N use case data based on the model evaluation task; K is a positive integer multiple of N; A task sharding module, configured to create M model processing threads associated with the M shard tasks based on the parallel processing quantity, divide the K to-be-evaluated use cases into the M shard tasks to obtain a shard task processing queue of the M shard tasks; one shard task corresponds to one model processing thread; among the M shard tasks, there is a shard task i, and the model processing thread associated with the shard task i is model processing thread j, where both i and j are positive integers less than or equal to M; A model processing module, configured to determine a first case to be evaluated from the cases to be evaluated included in the shard task processing queue of the shard task i, input the first case to be evaluated into the model to be evaluated through the model processing thread j, and perform data processing on the first case to be evaluated by the model to be evaluated to obtain a data processing result for the first case to be evaluated; A testing module, configured to obtain a test processing thread associated with the model evaluation task, and perform test processing on the data processing result based on the test processing thread to obtain a case test result for the first case to be evaluated.

16. A data processing device based on an artificial intelligence model, characterized in that, The apparatus is run by a service client, and the apparatus includes: A configuration page display module, configured to display an evaluation task configuration page; the evaluation task configuration page includes a model configuration area and a data set configuration area; A model configuration module, configured to, in response to a model configuration operation for the model configuration area, display a model to be evaluated in the model configuration area; the concurrent processing capacity of the model to be evaluated is used to indicate M shard tasks associated with the model evaluation task; M is a positive integer; A data set configuration module, configured to, in response to a data set configuration operation for the data set configuration area, display an evaluation data set in the data set configuration area; the evaluation data set includes N case data; N is a positive integer; A task generation module, configured to, in response to a task confirmation operation for a task confirmation control in the evaluation task configuration page, generate a model evaluation task based on the model to be evaluated and the evaluation data set, and send the model evaluation task to the service processing device, so that the service processing device determines K cases to be evaluated associated with the N case data based on the model evaluation task, divide the K cases to be evaluated into the M shard tasks to obtain a shard task processing queue of the M shard tasks; the shard task processing queue is used to indicate that when a first case to be evaluated is determined from the shard task processing queue, input the first case to be evaluated into the model to be evaluated through a model processing thread, and perform data processing on the first case to be evaluated by the model to be evaluated to obtain a data processing result for the first case to be evaluated; A result receiving module, configured to receive the case test result of the first case to be evaluated and display the case test result; the case test result is obtained by the service processing device performing test processing on the data processing result based on a test processing thread.

17. A computer device, characterized in that, including a memory and a processor; The memory is connected to the processor, the memory is used to store a computer program, and the processor is used to call the computer program so that the computer device executes the method according to any one of claims 1-14.

18. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium, and the computer program is suitable for being loaded and executed by a processor so that a computer device having the processor executes the method according to any one of claims 1-14.

19. A computer program product, characterized in that, Comprising a computer program / instructions which, when executed by a processor, implement the method according to any one of claims 1 - 14.