Apparatus and method for training and evaluating ai model
By building process task templates and adopting stage splitting and product escape analysis on the containerized application management platform, the training and evaluation of AI models are automated and parallelized, solving the problem of high manual participation and improving efficiency.
Patent Information
- Application Number
- PCT/CN2024/143372
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-29
- Filing Date
- 2024-12-27
- Publication Date
- 2025-07-03
AI Technical Summary
There are problems such as high manual participation, low efficiency and high labor costs in the training and evaluation process of existing AI models, which makes it impossible to achieve parallel progress of training and evaluation.
On the containerized application management platform, by building process task templates, the training and evaluation tasks of AI models are automated and parallelized, and phase splitting and product escape analysis are adopted, allowing the evaluation tasks to be started concurrently when the node reaches the training stage, reducing manual intervention.
The continuous training and evaluation process of AI models is automated and parallelized, development efficiency is improved, manual participation is reduced, and training and evaluation cannot be carried out in parallel.
Smart Images

Figure CN2024143372_03072025_PF_FP_ABST
Abstract
Description
AI model training and evaluation device and method
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims priority to Chinese patent application No. 202311869522.0, filed on December 29, 2023, entitled “Training and Evaluation Device and Method for AI Model”. The entire contents of the Chinese patent application are incorporated herein by reference. Technical Field
[0003] The present application generally relates to the field of artificial intelligence (AI), and more specifically to a training and evaluation device and method for an AI model. Background Art
[0004] In recent years, with the rapid development of artificial intelligence technology, more and more AI models have been applied to various fields. The generation of an AI model usually goes through two stages: model training and model evaluation.
[0005] Model training involves training and fine-tuning model parameters using a large amount of training data. Model evaluation, on the other hand, involves testing the AI models generated during training using pre-prepared and designed test data. After the model evaluation process is complete, test results, such as images and metrics, are typically generated for the evaluated model on the test data. These metrics reflect the effectiveness of the AI model, helping algorithm developers select the most effective models for production use.
[0006] Currently, AI model training and evaluation are typically performed manually by algorithm developers. For example, after training a model, algorithm developers temporarily pause training and manually retrieve the model to run the evaluation process. After the evaluation is complete, they decide whether to continue training based on the evaluation results. Alternatively, after training is complete, they retrieve multiple models generated by training and run the evaluation process on each model in turn. After all evaluation processes are complete, they manually select the optimal model.
[0007] This fully manual training and evaluation process involves too many manual operations, resulting in problems such as low evaluation efficiency and high labor costs. Summary of the Invention
[0008] In view of the above-mentioned problems, according to one aspect of an embodiment of the present application, a device for training and evaluation of an AI model is provided, comprising: a memory for storing program instructions; and a processor, which is coupled to the memory and configured to execute program instructions to: create a process task template applied to a containerized application management platform, the process task template comprising: process configuration information associated with the training and evaluation task of the AI model, and stage flags corresponding to each step in the training and evaluation task; and in response to receiving a start command for a continuous training and evaluation task, perform the following operations: query the process task template, and according to the process configuration information and stage flags in the process task template, split the training and evaluation task into stages to construct a training stage task, perform product escape analysis on the training stage task to generate a product escape analysis result associated with the training stage task, start the training stage task based on the product escape result associated with the training stage task, and when the training stage task reaches a predetermined training stage node, start the concurrent running of one or more evaluation stage tasks without interrupting the training stage task.
[0009] According to another aspect of an embodiment of the present application, a method for training and evaluation of AI models is provided, including: creating a process task template applied to a containerized application management platform, the process task template including: process configuration information associated with the training and evaluation task of the AI model, and stage flags corresponding to each step in the training and evaluation task; and in response to receiving a start command for a continuous training and evaluation task: querying the process task template, performing phase splitting on the training and evaluation task to construct a training phase task according to the process configuration information and phase flags in the process task template, performing product escape analysis on the training phase task to generate a product escape analysis result associated with the training phase task, starting the training phase task based on the product escape result associated with the training phase task, and when the training phase task reaches a predetermined training phase node, starting the concurrent running of one or more evaluation phase tasks without interrupting the training phase task.
[0010] According to another aspect of an embodiment of the present application, a machine-readable medium is provided, on which program instructions are stored. When the program instructions are executed by a processor, the processor executes the method for training and evaluating an AI model as described above.
[0011] According to another aspect of an embodiment of the present application, a device for training and evaluating an AI model is provided, comprising a device for executing the method for training and evaluating an AI model as described above.
[0012] According to the apparatus, method, device and machine-readable medium for training and evaluation of AI models in the embodiments of the present application, the training and evaluation process configuration of the AI model is abstracted and solidified, continuous training and evaluation of the AI model is realized, manual participation in the training and evaluation process is reduced, and the problem that model training and model evaluation cannot be carried out in parallel is solved, thereby improving the efficiency of AI model development.
[0013] In addition, according to some embodiments of the present application, a unified AI model single / continuous training evaluation method can also be implemented. In these embodiments, the platform can respond to a single training evaluation start command or a continuous training evaluation start command input by the user, and accordingly start a single training evaluation or a continuous training evaluation program. That is, after a unified training evaluation configuration, the user can choose to run a single training evaluation task or a continuous training evaluation task without modifying the configuration. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] The features, advantages and technical effects of exemplary embodiments of the present application will be described below with reference to the accompanying drawings.
[0015] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0016] FIG1 shows a schematic overall flow chart of the training and evaluation process of an AI model according to some embodiments of the present application;
[0017] FIG2 shows an example process of process configuration in the training and evaluation process of an AI model according to some embodiments of the present application;
[0018] FIG3 shows a single training and evaluation task process of an AI model according to some embodiments of the present application;
[0019] FIG4 shows an overall flow chart of the continuous training and evaluation task process of the AI model according to some embodiments of the present application;
[0020] FIG5 is a schematic diagram showing a stage splitting process in a continuous training and evaluation task flow of an AI model according to some embodiments of the present application;
[0021] FIG6 is a schematic diagram showing a process of using a sidecar agent to assist in delivering product files in a continuous training and evaluation task process of an AI model according to some embodiments of the present application;
[0022] FIG7 shows a schematic diagram of non-invasive Sidecar proxy container injection in the continuous training and evaluation task process of the AI model according to some embodiments of the present application;
[0023] FIG8 is a schematic diagram showing the concurrent execution of multiple model evaluation tasks in the continuous training and evaluation task process of an AI model according to some embodiments of the present application;
[0024] FIG9 shows a detailed flowchart of the continuous training and evaluation task process of the AI model according to some embodiments of the present application;
[0025] Figure 10 shows a schematic diagram of the hardware structure of an AI model training and evaluation device according to some embodiments of the present application. DETAILED DESCRIPTION
[0026] The features and exemplary embodiments of various aspects of the present application will be described in detail below. In the detailed description below, many specific details are proposed to provide a comprehensive understanding of the application. However, it will be apparent to those skilled in the art that the application can be implemented without the need for some of these specific details. The following description of the embodiments is merely to provide a better understanding of the application by illustrating the examples of the present application. The application is in no way limited to any specific configuration proposed below, but covers any modification, replacement and improvement of elements, parts and algorithms without departing from the spirit of the application. In the accompanying drawings and the following description, known structures and technologies are not shown to avoid causing unnecessary ambiguity to the application.
[0027] In order to solve the problems of high manual participation, low evaluation efficiency, and high labor costs caused by the use of full manual training and evaluation workflow in artificial intelligence model training and evaluation, this application proposes a continuous training and evaluation process for AI models that can be implemented on a containerized application management platform (for example, a Kubernetes-based platform). This process abstracts and solidifies the training and evaluation process configuration of the AI model, solving the problem that model training and model evaluation cannot be carried out in parallel. In addition, the platform for implementing the continuous training and evaluation process of the AI model according to the present application can also be easily compatible with the single training and evaluation process of the AI model. Therefore, according to some embodiments of the present application, a unified AI model single or continuous training and evaluation process can be implemented. The overall process of this unified training and evaluation process will be described below with reference to Figure 1, and then the process of abstracting and solidifying the training and evaluation process configuration of the AI model, as well as the specific processes of the single training and evaluation process and the continuous training and evaluation process will be described in detail with reference to Figures 2 to 9, especially the implementation process of the continuous training and evaluation process will be described in detail.
[0028] Figure 1 shows a schematic overall flow chart of the training and evaluation process of an AI model according to some embodiments of the present application. As shown in Figure 1, the overall process of AI model training and evaluation may include a process configuration process, a training start process, and a single training and evaluation task process or a continuous training and evaluation task process initiated according to user selection.
[0029] In general, the training and evaluation process can be implemented on a containerized application management platform. First, the platform can construct a task step composition of the training and evaluation task of the AI model based on the knowledge of the relevant fields of artificial intelligence. The task step composition includes the various steps required to complete the training and evaluation task of the AI model. For example, for the field of image detection, a task step composition including main steps such as image preprocessing steps, model training steps, model evaluation steps, and post-processing steps can be constructed. The platform can solidify these task steps into a process task template through software methods to facilitate the subsequent running of training and evaluation tasks based on the process task template. Then, based on the solidified process task template, the platform can start a single training and evaluation task or a continuous training and evaluation task of the AI model in response to the training and evaluation start command input by the user, thereby realizing the automation and parallelization of the training and evaluation functions of the AI model, and can provide a unified storage and display method for models and reports.
[0030] In terms of specific implementation methods, for example, by using containerized application management engines (for example, Kubernetes), sidecar mechanisms, object storage, process engines, software development kits (SDKs) and other related technical means, the training and evaluation task processes of AI models are configured in a templated manner. At the same time, through mechanisms such as stage splitting and product escape analysis, it is possible to choose to conduct a single training and evaluation task process or a continuous training and evaluation task process according to user needs. If the user chooses to start a single training and evaluation task process, a complete process including the training and evaluation phases of a single AI model will be run, which is similar to the fully manual method, but each step in the training and evaluation process will be completed automatically and sequentially without the need for excessive manual intervention. If the user selects the continuous training and evaluation task process, the process task will be automatically split into stages and product escape analysis will be performed. Based on the results of the stage splitting and product escape analysis, the training stage task (also referred to as the training task in this article) will be started, and when the training stage task reaches the predetermined training stage node, without interrupting the training stage task, one or more evaluation stage tasks (also referred to as the evaluation task in this article) will be started for concurrent operation, thereby realizing the parallel operation of the training stage task and the evaluation stage task.
[0031] The following will describe in detail the process of configuring the training and evaluation process of the AI model, as well as the specific processes of the single training and evaluation process and the continuous training and evaluation process, with reference to Figures 2 to 9. In particular, the specific implementation process of the continuous training and evaluation process will be described in detail with reference to Figures 4 to 9.
[0032] Figure 2 shows an example process of process configuration in the training and evaluation process of the AI model according to some embodiments of the present application. According to an embodiment of the present application, process configuration refers to the process of using software methods to correspond the relevant running codes, configurations, running environments, etc. in the workflow of algorithm experts in AI-related fields to the abstracted running steps (for example, for the field of image detection, it may include image preprocessing steps, model training steps, model evaluation steps, post-processing steps, etc.), that is, solidifying the workflow of algorithm experts in AI-related fields into a process task template.
[0033] For example, the technical solution according to the embodiment of the present application can be implemented on a platform based on a containerized application management engine (for example, Kubernetes). All training and evaluation tasks are ultimately scheduled and run in the form of containers (Pods). Therefore, in the specific implementation, the relevant configurations of the algorithm expert workflow in the AI field can be mapped to parameters such as container startup commands, container images, stage products, configuration files, and then process task templates are generated based on these parameters at the software level and solidified into the training and evaluation process of the AI model.
[0034] Specifically, in the process configuration stage, the training and evaluation process of the AI model can be mapped and parameterized to generate a process task template associated with the training and evaluation task of the AI model for running on the containerized application management platform. The process task template may include process configuration information associated with the training and evaluation task of the AI model. As shown in Figure 2, for example, for an image detection model, the process configuration information may include the image detection training set, training set preprocessing configuration, image detection test set, and test set preprocessing configuration associated with the image preprocessing step; the image configuration, startup command configuration, resource configuration, product configuration, and input file configuration associated with the model training step; the image configuration, startup command configuration, resource configuration, product configuration, and input file configuration associated with the model evaluation step; and the model evaluation report analysis and model post-processing configuration associated with the post-processing step.
[0035] In addition, as described above, according to some embodiments of the present application, when the user chooses to start a continuous training and evaluation task process, the platform needs to automatically perform phase splitting and product escape analysis on the process task, start the training phase task based on the results of the phase splitting and product escape analysis, and when the training phase task reaches a predetermined training phase node, start the concurrent running of one or more evaluation phase tasks without interrupting the training phase task.
[0036] In these embodiments, when creating a process task template, a task step composition for the AI model training and evaluation task is constructed based on domain knowledge associated with the AI model. The task step composition includes the various steps required to complete the AI model training and evaluation task; and each step in the task step composition is assigned a stage to which the step belongs to obtain a stage marker corresponding to the step. The stage marker corresponding to each step can be included in the process task template to provide a basis for constructing subsequent stage tasks.
[0037] For example, for the training and evaluation tasks of image detection models, the task steps usually include image preprocessing steps, model training steps, model evaluation steps and post-processing steps. These steps can be divided into stages based on the knowledge of the image detection field, that is, each step is assigned to its own stage, providing a basis for the construction of subsequent stage tasks. For example, the training and evaluation tasks of image detection models can be divided into two stages: the model training stage and the model evaluation stage. The two stages are assigned stage marks 1 and 2 respectively, and then the corresponding stage marks are assigned to each step in the task step composition. For example, since both the training stage tasks and the evaluation stage tasks require image preprocessing steps, the image preprocessing step should be equipped with stage marks 1 and 2, while the model training step is only equipped with stage mark 1, and the model evaluation step is only equipped with stage mark 2.
[0038] Therefore, the process task template created by the platform includes not only the process configuration information associated with the training and evaluation tasks of the AI model, but also the stage marks corresponding to the various steps in the training and evaluation tasks. In this way, when the platform needs to start a stage task corresponding to a certain stage mark, it can construct the corresponding stage task to be started according to the stage marks corresponding to the various steps in the process task template. For example, when it is necessary to start the model training stage task (corresponding to stage mark 1), the model training stage task can be constructed based on the stage marks configured in the task step composition, including the various steps of stage mark 1, so that the training stage task can be started later.
[0039] As shown in Figure 1, after configuring the process task template for the AI model training and evaluation process, you can start the training and evaluation task with a single click. When starting the training and evaluation task, users can choose to perform a single or continuous training and evaluation task as needed. After the training and evaluation task is started, the platform will run the model training and evaluation containers in Kubernetes containers based on the process task template created during the process configuration phase and monitor their running status.
[0040] According to an embodiment of the present application, if the user chooses to start a single training and evaluation task when starting training, the platform will directly run the single training and evaluation task of the AI model. That is, based on the process task template, the single training phase task and the single evaluation phase task of the AI model are sequentially executed to complete the single training and evaluation task. Figure 3 shows the process of a single training and evaluation task of an AI model according to some embodiments of the present application.
[0041] As shown in Figure 3, taking the image detection model as an example, the single training and evaluation task process of the image detection model can include a preprocessing stage, a training stage, and an evaluation stage. Specifically, in the preprocessing stage of the training and evaluation of the image detection model, the following operations are mainly performed: automatically pulling the image data required for training and testing from the object storage on demand according to the image pulling conditions configured by the image detection field experts, and at the same time, automatically preprocessing the pulled image data according to the image preprocessing algorithm configured by the algorithm experts to generate preprocessed products (image data). After completing the pulling of the image data required for the training and evaluation tasks of the image detection model and the preprocessing of the image data, the training stage container will be automatically run. In the training stage, the training code of the image detection model will be run, and the image detection model will be trained using the preprocessed image data. The training stage will generate some product files. The most important product files include model files. The software level will automatically complete operations such as archiving these product files. After the training phase is complete, the image detection model's evaluation phase will be executed sequentially. The base model files used in the training phase and the model files generated after training will be automatically transferred to the evaluation phase container. At the same time, the relevant image test set selected by algorithm experts in the field of image detection will also be automatically transferred to the evaluation phase. The evaluation phase container generates some model report files by running the evaluation code. For image detection models, the report typically includes indicators such as image segmentation result images, object detection images, object detection positive detection rate, object detection false alarm rate, and object detection omission rate. Also in this phase, the software automatically completes operations such as archiving these product files.
[0042] According to some embodiments of the present application, if the user chooses to start a continuous training and evaluation task when starting training, the platform will automatically perform phase splitting and product escape analysis on the training and evaluation task, and start the training phase task based on the results of the phase splitting and product escape analysis. When the training phase task reaches a predetermined training phase node, one or more evaluation phase tasks can be started concurrently without interrupting the training phase task to perform continuous training and evaluation. Figure 4 shows an overall flow chart of the continuous training and evaluation task process of the AI model according to some embodiments of the present application.
[0043] Table 1 below shows some key operations involved in the continuous training and evaluation task process of the AI model in an example embodiment of the present application. The serial numbers in Table 1 correspond to the serial numbers of these operations shown in Figure 4.
[0044] Table 1
[0045] As described above with reference to Figure 1, the continuous training and evaluation task process involves multiple aspects, including phase splitting, product escape analysis, product storage and transfer, and the parallel execution of model training and model evaluation tasks. The operations shown in Table 1 are actually operations related to these aspects. These operations are described in more detail below with reference to Figures 5 to 9.
[0046] Figure 5 shows a schematic diagram of the stage splitting process in the continuous training and evaluation task process of the AI model according to some embodiments of the present application. As shown in Figure 5, stage splitting refers to automatically splitting the evaluation stage task in the training and evaluation task of the AI model from the process, that is, only starting the training stage task first. Specifically, when the platform receives the start command of the continuous training and evaluation task, the platform will query the process task template, and according to the process configuration information and stage flags in the process task template, the training and evaluation task will be stage split to construct the training stage task of the AI model.
[0047] As mentioned above, according to an embodiment of the present application, it is proposed that when creating a process task template, based on the domain knowledge associated with the AI model, the task step composition of the training and evaluation task of the AI model is constructed, and each step in the task step composition is assigned the stage to which the step belongs to obtain the stage mark corresponding to the step. The stage mark corresponding to each step can be included in the process task template to provide a basis for the construction of subsequent stage tasks. In this way, when the platform needs to start the stage task corresponding to a certain stage mark, the corresponding stage task to be started can be constructed according to the stage mark corresponding to each step in the process task template.
[0048] Based on the phase flags configured in the process task template, when it is necessary to start the training phase task corresponding to phase flag 1, the training phase task can be constructed during the phase splitting process based on the phase flags configured in the task step composition, including the various steps of phase flag 1.
[0049] In addition, according to an embodiment of the present application, after the platform performs stage splitting and constructs the training stage task, it is also necessary to perform product escape analysis. Specifically, the algorithm expert can configure corresponding products (for example, including input products and output products) for each step involved in the training and evaluation task. These products need to be transferred and used between multiple steps. However, due to the splitting of the stage tasks, the product transfer method originally applied to the various steps in a single task may no longer be applicable. Therefore, it is necessary to analyze and specially handle the situation of cross-stage transfer of products, which is also referred to as product escape analysis in this article.
[0050] Specifically, when starting a task at a certain stage, first obtain the process task template, which contains the input products and output products of each step configured by the algorithm expert. Traverse the input products and output products of these steps and analyze the transfer relationship between these products in each step (for example, product 1 is generated by step 1 and is used as input in step 2). After analyzing the transfer relationship between the products in each step, screen out products whose transfer relationship spans different stages in the training and evaluation tasks (for example, the product needs to be transferred between step 1 and step 2, and step 1 belongs to the training stage with stage mark 1, and step 2 belongs to the evaluation stage with stage mark 2). These products are processed as escaped products. When there are escaped products, the platform first cancels the original processing logic for these escaped products in the original single task, and then establishes the transfer processing logic of the escaped products based on the stage marks corresponding to each step in the stage task as the product escape analysis result associated with the stage task.
[0051] It should be noted that when starting the training phase task and the evaluation phase task, cross-phase product transfer may occur, and therefore, corresponding product escape analysis is required. Specifically, the product escape analysis performed when starting the training phase task includes: based on the process task template, analyzing the transfer relationship between the products of each step in the training phase task; screening out the products whose transfer relationship spans across different stages in the training and evaluation tasks, and processing them as escaped products; and based on the stage marks corresponding to each step in the training phase task, establishing the transfer processing logic of the escaped products as the product escape analysis result associated with the training phase task. Correspondingly, the product escape analysis performed when starting the evaluation phase task includes: based on the process task template, analyzing the transfer relationship between the products of each step in the evaluation phase task; screening out the products whose transfer relationship spans across different stages in the training and evaluation tasks, and processing them as escaped products; and based on the stage marks corresponding to each step in the evaluation phase task, establishing the transfer processing logic of the escaped products as the product escape analysis result associated with the evaluation phase task.
[0052] In addition, regarding the transmission of product escape analysis results, according to an embodiment of the present application, it is proposed that the transmission of product escape analysis results can be carried out in the following manner: configuring an associated proxy container for the stage task, using the proxy container to store product files according to the product escape analysis results, and assisting in the transmission of product files between training stage tasks and evaluation stage tasks.
[0053] For example, the product escape analysis results associated with the training phase tasks can be passed to the sidecar proxy container in the form of environment variables. At the same time, the sidecar proxy container shares storage with the training container by mounting the same storage, so that the products generated by the training phase tasks can be obtained; after the training task achieves phased results, the training container notifies the sidecar proxy container through a dedicated software development tool kit SDK to start the model evaluation task. The sidecar proxy container processes all escaped products (such as packaging, archiving, etc.) according to the product escape analysis results obtained by the environment variables, and finally starts the model evaluation task through the platform service interface, and passes the archive information of the escaped products to the platform service through the interface.
[0054] Figure 6 shows a schematic diagram of the process of using a sidecar proxy to assist in delivering product files in the continuous training and evaluation task process of an AI model according to some embodiments of the present application. As shown in Figure 6, when the product delivery spans different stage tasks, the product needs to be marked as an escaped product, and when an independent evaluation stage task is subsequently initiated, the sidecar proxy container can assist in the storage and delivery of the product files.
[0055] According to some embodiments of the present application, in order to autonomously complete operations such as updating the preprocessing products of the training dataset, storing the training phase products (models), and launching concurrent model evaluation tasks during the operation of the model training container, a non-invasive proxy container injection method is adopted in the technical implementation. This proxy container can be, for example, a sidecar proxy container based on the sidecar architecture pattern for a containerized application management platform.
[0056] Figure 7 shows a schematic diagram of the non-invasive Sidecar proxy container injection in the continuous training and evaluation task process of the AI model according to some embodiments of the present application. Since the actions of initiating storage of training phase products, updating data set preprocessing products, etc. in the training phase tasks involve complex and time-consuming operations, and starting concurrent evaluation tasks requires calling the platform service interface, it is inappropriate to implement these operations directly by the training container. To solve this problem, in the embodiments of the present application, it is proposed to adopt the method of combining SDK with Sidecar proxy container, and the Sidecar proxy container completes the complex and time-consuming operations and calls the platform service interface, and the SDK layer only encapsulates a simple interface layer.
[0057] As shown in Figure 7, the training phase task is configured with an associated sidecar proxy container. This sidecar proxy container can interact with the training and evaluation phase tasks through a dedicated SDK to store and transfer product artifacts between the sidecar proxy container and the training and evaluation phase tasks. Specifically, by automatically injecting a self-developed sidecar proxy container into the container belonging to the training task on Kubernetes, the training container only needs to communicate with the sidecar proxy container over the local network (this communication is also encapsulated by the Python SDK). At the same time, the sidecar proxy container and the training container can share part of the file system, that is, they can complete the required file storage or transfer operations synchronously or asynchronously on demand, or call remote platform service interfaces to start the concurrent execution of multiple model evaluation tasks.
[0058] Figure 8 shows a schematic diagram of the concurrent operation of multiple model evaluation tasks in the continuous training and evaluation task process of the AI model according to some embodiments of the present application. As shown in Figure 8, in an embodiment according to the present application, the training task and the evaluation task can run in parallel without affecting each other, and multiple evaluation tasks can also run concurrently. For example, when the training phase task reaches a predetermined training phase node, the concurrent operation of one or more evaluation phase tasks can be started without interrupting the training phase task. In an embodiment of the present application, the predetermined training phase node may include: a node at which the training phase task reaches the phase model training target, or a node at which the training phase task reaches a predefined number of training iterations. For example, when the training phase task reaches a predetermined first training phase node, the first evaluation phase task can be started for the training model obtained at the first training phase node; when the training phase task reaches a predetermined second training phase node, the second evaluation phase task can be started for the training model obtained at the second training phase node; and so on. While starting the concurrent operation of multiple evaluation phase tasks, the training phase tasks can continue. In addition, as shown in Table 1, the Sidecar proxy container can be used to call the platform service interface to start the concurrent running of multiple model evaluation tasks.
[0059] According to the embodiments of the present application, the use of SDK combined with Sidecar proxy to implement the storage and delivery of product files and the calling of remote platform service interfaces will have the following advantages: after separating this part of the logic from the training code of the AI model, the difficulty of algorithm implementation and the complexity of the code are reduced; the specific implementation is completed by the Sidecar proxy, which realizes the separation of logic and algorithm and increases the flexibility of subsequent modification of the implementation logic; and avoids the training code directly calling the platform service interface, thereby avoiding the exposure of the platform service interface.
[0060] The various possible operations and steps involved in the continuous training and evaluation process of AI models based on the containerized application management platform have been described above in conjunction with Figures 4 to 8. To more clearly illustrate the overall flow of this continuous training and evaluation process, the following will further describe this continuous training and evaluation process with reference to Figure 9.
[0061] Figure 9 shows a detailed flowchart of the continuous training and evaluation task process of the AI model according to some embodiments of the present application. The continuous training and evaluation task process can be implemented on a containerized application management platform. It should be noted here that in the previous description, the training and evaluation task process (including a single training and evaluation task process and a continuous evaluation task process) provided by this application is mainly described using the image detection model in the image processing field as an example, but it can be understood that the training and evaluation task process can be applied to the training and evaluation of different AI models in various fields. For different AI models in different fields, the platform can create a process task template applied to the containerized application management platform based on the process configuration information associated with the training and evaluation task of the AI model input by algorithm experts in the relevant field. The process task template may include: process configuration information associated with the training and evaluation task of the AI model, and stage marks corresponding to each step in the training and evaluation task. Therefore, there may be multiple process task templates associated with different AI models in various fields on the platform. Accordingly, when a training or evaluation task is to be started, it is necessary to query the process task template to which the training or evaluation task belongs.
[0062] As shown in Figure 9, in response to receiving the start command of the continuous training and evaluation task, the platform can perform the following operations: query the process task template, split the training and evaluation task into stages to construct a training stage task based on the process configuration information and stage flags in the process task template, perform product escape analysis based on the training stage task to generate a product escape analysis result associated with the training stage task, and start the training stage task based on the product escape result associated with the training stage task.
[0063] During the execution of the training phase tasks, the following operations can be performed: perform model training iterations; during the iteration process, notify the sidecar agent through the SDK to update the data set preprocessing products as needed; when the training phase task reaches the predetermined training phase node (for example, the phase model training target or the predefined number of training iterations is reached), without interrupting the training phase task, notify the sidecar agent through the SDK to start the evaluation phase task; use the sidecar agent to store the escaped products according to the phase operation configuration and product escape analysis results; the sidecar agent calls the platform service interface to start the evaluation phase task; and continue subsequent model training.
[0064] As shown in Figure 9, the platform can initiate a corresponding evaluation phase task each time a training phase task reaches a predetermined training phase node, without interrupting the training phase task. This allows for the concurrent execution of multiple evaluation phase tasks (evaluation tasks 1 through n) on the platform. Furthermore, for each evaluation phase task, the following operations can be performed: querying the process task template; constructing the evaluation phase task based on the process configuration information and phase flags in the process task template; performing product escape analysis on the evaluation phase task to generate product escape analysis results associated with the evaluation phase task; and initiating the evaluation phase task based on the product escape results associated with the evaluation phase task.
[0065] In summary, the continuous training and evaluation process according to this application can automate and parallelize the training and evaluation functions of various AI models in different fields. In addition, the SDK is combined with the Sidecar proxy to implement the storage and delivery of product files and the calling of platform service interfaces, so that the training task container does not need to directly call the platform service interface, avoiding the exposure of the platform service interface.
[0066] Figure 10 shows a schematic diagram of the hardware structure of the AI model training and evaluation device according to an embodiment of the present application.
[0067] The apparatus may include a processor 1001 and a memory 1002 storing program instructions.
[0068] Specifically, the processor 1001 may include a central processing unit (CPU), or an application-specific integrated circuit (ASIC), or may be configured to implement one or more integrated circuits of the embodiments of the present application.
[0069] The memory 1002 may include a large capacity memory for data or instructions. By way of example and not limitation, the memory 1002 may include a hard disk drive (HDD), a floppy disk drive, a flash memory, an optical disk, a magneto-optical disk, a magnetic tape, or a universal serial bus (USB) drive, or a combination of two or more of these. Where appropriate, the memory 1002 may include removable or non-removable (or fixed) media. Where appropriate, the memory 1002 may be inside or outside the integrated gateway disaster recovery device. In a specific embodiment, the memory 1002 is a non-volatile solid-state memory.
[0070] In certain embodiments, memory 1002 includes read-only memory (ROM). The ROM may be mask-programmed ROM, programmable ROM (PROM), erasable PROM (EPROM), electrically erasable PROM (EEPROM), electrically alterable ROM (EAROM), or flash memory, or a combination of two or more thereof, where appropriate.
[0071] The memory may include a read-only memory (ROM), a random access memory (RAM), a magnetic disk storage medium device, an optical storage medium device, a flash memory device, an electrical, optical or other physical / tangible storage device. Therefore, typically, the memory includes one or more tangible (non-transitory) readable storage media (e.g., a memory device) encoded with software including computer-executable instructions, and when the software is executed (e.g., by one or more processors), it is operable to perform the operations described with reference to the training and evaluation method of the AI model according to the embodiment of the present application.
[0072] The processor 1001 implements the training and evaluation method of the AI model described in the above embodiment by reading and executing the program instructions stored in the memory 1002.
[0073] In one example, the training and evaluation device of the AI model may further include a communication interface 1003 and a bus 1010. As shown in FIG10 , the processor 1001, the memory 1002, and the communication interface 1003 are connected via the bus 1010 and communicate with each other.
[0074] The communication interface 1003 is mainly used to implement communication between various modules, devices, units and / or equipment in the embodiments of the present application.
[0075] Bus 1010 includes hardware, software or both, couples the parts of device to each other.For example, and not limitation, bus may include accelerated graphics port (AGP) or other graphics bus, enhanced industry standard architecture (EISA) bus, front side bus (FSB), hypertransport (HT) interconnection, industry standard architecture (ISA) bus, infinite bandwidth interconnection, low pin count (LPC) bus, memory bus, micro channel architecture (MCA) bus, peripheral component interconnection (PCI) bus, PCI-Express (PCI-X) bus, serial advanced technology attachment (SATA) bus, video electronics standard association local (VLB) bus or other suitable bus or two or more of these combinations. In appropriate cases, bus 1010 may include one or more buses. Although the present application embodiment describes and shows specific bus, the application considers any suitable bus or interconnection.
[0076] In conjunction with the AI model training and evaluation method described in the above embodiments, the present embodiments can be implemented by a machine-readable medium having program instructions stored thereon; when the program instructions are executed by a processor, the processor executes the AI model training and evaluation method described in the above embodiments.
[0077] It should be understood that the present application is not limited to the specific configurations and processes described above and illustrated in the figures. For the sake of brevity, a detailed description of known methods is omitted here. In the above embodiments, several specific steps are described and illustrated as examples. However, the method process of the present application is not limited to the specific steps described and illustrated. Those skilled in the art can make various changes, modifications, and additions, or change the order of the steps after understanding the spirit of the present application.
[0078] The functional blocks shown in the above-described block diagram can be implemented as hardware, software, firmware or a combination thereof. When implemented in hardware, they can be, for example, electronic circuits, application specific integrated circuits (ASICs), appropriate firmware, plug-ins, function cards, etc. When implemented in software, the elements of the present application are programs or code segments that are used to perform the required tasks. The program or code segment can be stored in a machine-readable medium, or transmitted on a transmission medium or a communication link by a data signal carried in a carrier wave. "Machine-readable medium" can include any medium that can store or transmit information. Examples of machine-readable media include electronic circuits, semiconductor memory devices, ROMs, flash memories, erasable ROMs (EROMs), floppy disks, CD-ROMs, optical disks, hard disks, optical fiber media, radio frequency (RF) links, etc. The code segment can be downloaded via a computer network such as the Internet, an intranet, etc.
[0079] It should also be noted that the exemplary embodiments mentioned in this application describe some methods based on a series of steps. However, this application is not limited to the order of the above steps. In other words, the steps can be performed in the order mentioned in the embodiments, or in a different order, or several steps can be performed simultaneously.
[0080] The above reference is according to the flow chart and / or block diagram of the method and apparatus of the embodiment of the application, described various aspects of the application.It should be understood that each box in the flow chart and / or block diagram and the combination of each box in the flow chart and / or block diagram can be realized by program instructions.These program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer or other programmable data processing device, to produce a kind of machine, so that these instructions executed via the processor of the computer or other programmable data processing device enable the realization of the function / action specified in one or more boxes of the flow chart and / or block diagram.Such a processor can be, but is not limited to, a general-purpose processor, a special-purpose processor, a special application processor or a field programmable logic circuit.It is also understood that each box in the block diagram and / or flow chart and the combination of the boxes in the block diagram and / or flow chart can also be realized by the dedicated hardware that performs the specified function or action, or can be realized by the combination of dedicated hardware and computer instructions.
[0081] The above description is only a specific embodiment of the present application. Those skilled in the art will clearly understand that for the convenience and brevity of description, the specific working process of the device described above can refer to the corresponding process in the aforementioned method embodiment, and will not be repeated here. It should be understood that the scope of protection of the present application is not limited to this. Any person skilled in the art can easily think of various equivalent modifications or replacements within the technical scope disclosed in this application, and these modifications or replacements should be included in the scope of protection of this application.
Claims
1. A device for training and evaluation of an artificial intelligence (AI) model, comprising: A memory for storing program instructions; And A processor coupled to the memory and configured to execute the program instructions to: Create a process task template applied to a containerized application management platform, the process task template including: process configuration information associated with the training and evaluation task of the AI model, and stage flags corresponding to each step in the training and evaluation task; And In response to receiving a start command for a continuous training and evaluation task, perform the following operations: Query the process task template, According to the process configuration information and the stage flags in the process task template, split the training and evaluation task into stages to construct a training stage task, Perform product escape analysis based on the training stage task to generate a product escape analysis result associated with the training stage task, Based on the product escape result associated with the training stage task, start the training stage task, and When the training stage task reaches a predetermined training stage node, start the concurrent execution of one or more evaluation stage tasks without interrupting the training stage task.
2. The device according to claim 1, wherein, Starting the concurrent execution of one or more evaluation stage tasks includes performing the following operations for each evaluation stage task in the one or more evaluation stage tasks: Query the process task template; According to the process configuration information and the stage flags in the process task template, construct the evaluation stage task, Perform product escape analysis based on the evaluation stage task to generate a product escape analysis result associated with the evaluation stage task, and Based on the product escape result associated with the evaluation stage task, start the evaluation stage task.
3. The device according to claim 1, wherein The predetermined training stage node includes: a node where the training stage task reaches a stage model training target, or a node where the training stage task reaches a predefined number of training iterations.
4. The device according to claim 1, wherein, Performing product escape analysis based on the training stage task includes: Based on the process task template, analyze the transfer relationship of the products of each step in the training stage task, where the products include input products and output products; Screen out the products whose transfer relationship crosses different stages in the training and evaluation task and process them as escape products; and Based on the stage flags corresponding to each step in the training stage task, establish the transfer processing logic of the escape products as the product escape analysis result associated with the training stage task.
5. The device according to claim 2, wherein, Performing product escape analysis based on the evaluation stage task includes: Based on the process task template, analyze the transfer relationship of the products of each step in the evaluation stage task, where the products include input products and output products; Screen out the products whose transfer relationship crosses different stages in the training and evaluation task and process them as escape products; and Based on the phase flags corresponding to each step in the evaluation phase task, establish the transfer processing logic of the escaped product as the product escape analysis result associated with the evaluation phase task.
6. The device according to any one of claims 1 to 5, wherein, The training phase task is configured with an associated proxy container, and starting the evaluation phase task in the training and evaluation task includes: using the proxy container to start the concurrent execution of the one or more evaluation phase tasks by invoking the service interface of the containerized application management platform.
7. The apparatus according to claim 6, wherein, The operation further includes: using the proxy container to store the product file according to the product escape analysis result and assisting in realizing the transfer of the product file between the training phase task and the one or more evaluation phase tasks.
8. The apparatus according to claim 6, wherein, The training phase task interacts with the proxy container through a dedicated software development kit (SDK) to notify the proxy container to update the product associated with the training phase task and to start the concurrent execution of the one or more evaluation phase tasks.
9. The apparatus according to claim 6, wherein The proxy container is a Sidecar proxy container based on the Sidecar architecture mode for the containerized application management platform.
10. The device according to claim 1, wherein, The processor is further configured to execute the program instructions to: In response to receiving a start command for a single training and evaluation task, sequentially execute the single training phase task and the single evaluation phase task of the AI model based on the process task template.
11. The device according to claim 1, wherein, Creating a process task template applied to the containerized application management platform includes: Based on the domain knowledge associated with the AI model, constructing the composition of the task steps of the training and evaluation task of the AI model, which composition includes all the steps required to complete the training and evaluation task of the AI model; and Assigning the phase to which each step in the composition of the task steps belongs to obtain the phase flag corresponding to the step.
12. The device according to claim 11, wherein, The AI model includes an image detection model, and the composition of the task steps includes an image preprocessing step, a model training step, a model evaluation step, and a post-processing step.
13. A method for training and evaluating an artificial intelligence (AI) model, including: Creating a process task template applied to the containerized application management platform, which process task template includes: process configuration information associated with the training and evaluation task of the AI model, and phase flags corresponding to each step in the training and evaluation task; and In response to receiving a start command for a continuous training and evaluation task: Query the process task template, According to the process configuration information and the phase flags in the process task template, split the training and evaluation task into phases to construct a training phase task, Perform product escape analysis based on the training phase task to generate a product escape analysis result associated with the training phase task, Based on the product escape result associated with the training phase task, start the training phase task, and When the training phase task reaches a predetermined training phase node, start the concurrent execution of one or more evaluation phase tasks without interrupting the training phase task.
14. The method according to claim 13, wherein, Initiating the concurrent execution of one or more evaluation phase tasks includes performing the following operations for each of the one or more evaluation phase tasks: Querying the process task template; Constructing the evaluation phase task according to the process configuration information and the phase flag in the process task template, Performing product escape analysis based on the evaluation phase task to generate a product escape analysis result associated with the evaluation phase task, and Based on the product escape result associated with the evaluation phase task, initiating the evaluation phase task.
15. The method according to claim 13, wherein, The predetermined training phase node includes: a node at which the training phase task reaches the stage model training target, or a node at which the training phase task reaches the predefined number of training iterations.
16. The method according to claim 13, wherein Performing product escape analysis based on the training phase task includes: Based on the process task template, analyzing the transfer relationship of the products of each step in the training phase task between each step, where the products include input products and output products; Screening out the products whose transfer relationship spans different phases in the training and evaluation task as escape products for processing; and Based on the phase flags corresponding to each step in the training phase task, establishing the transfer processing logic of the escape products as the product escape analysis result associated with the training phase task.
17. The method according to claim 14, wherein Performing product escape analysis based on the evaluation phase task includes: Based on the process task template, analyzing the transfer relationship of the products of each step in the evaluation phase task between each step, where the products include input products and output products; Screening out the products whose transfer relationship spans different phases in the training and evaluation task as escape products for processing; and Based on the phase flags corresponding to each step in the evaluation phase task, establishing the transfer processing logic of the escape products as the product escape analysis result associated with the evaluation phase task.
18. The method according to any one of claims 13 to 17, wherein The training phase task is configured with an associated proxy container, and initiating the evaluation phase task in the training and evaluation task includes: using the proxy container to initiate the concurrent execution of the one or more evaluation phase tasks by calling the service interface of the containerized application management platform.
19. The method according to claim 18, further comprising: Using the proxy container to store the product file according to the product escape analysis result and assisting in realizing the transfer of the product file between the training phase task and the one or more evaluation phase tasks.
20. The method according to claim 18, wherein The training phase task interacts with the proxy container through a dedicated software development kit (SDK) to notify the proxy container to update the products associated with the training phase task and initiate the concurrent execution of the one or more evaluation phase tasks.
21. The method according to claim 18, wherein, The proxy container is a Sidecar proxy container based on the Sidecar architecture mode for the containerized application management platform.
22. The method according to claim 13, further comprising: In response to receiving a start command for a single training and evaluation task, sequentially executing the single training phase task and the single evaluation phase task of the AI model based on the process task template.
23. The method according to claim 13, wherein Creating a process task template for a containerized application management platform includes: Based on the domain knowledge associated with the AI model, constructing the composition of task steps for the training and evaluation task of the AI model, where the composition of task steps includes each step required to complete the training and evaluation task of the AI model; and Assigning the phase to which each step in the composition of task steps belongs to obtain the phase flag corresponding to the step.
24. The method according to claim 23, wherein, The AI model includes an image detection model, and the composition of task steps includes an image preprocessing step, a model training step, a model evaluation step, and a post-processing step.
25. A machine-readable medium storing program instructions that, when executed by a processor, cause the processor to perform the method for training and evaluating an artificial intelligence (AI) model as described in any one of claims 13 to 24.
26. A device for training and evaluating an artificial intelligence (AI) model, including means for performing the method for training and evaluating an AI model as described in any one of claims 13 to 24.
Citation Information
Patent Citations
Model training test optimization and deployment method and device based on container
CN112463301A
Automatic training system and method for package recommendation machine learning model
CN112685457A
Model training visualization method and device and cloud platform
CN114861773A
AI-driven project automation method and system based on stage retraining
CN116468131A
Training system and method based on cognitive models
US20110097697A1
Cited By
Language model service performance evaluation method, device, equipment, system and medium
CN120780574A