AI model training evaluation device and method

By building process task templates on the containerized application management platform and performing phase splitting and product escape analysis, the problem of AI model training and evaluation cannot be parallelized, and efficient automated training and evaluation of AI models is realized.

CN120236160APending Publication Date: 2025-07-01NUCTECH JIANGSU CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311869522.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-29
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

There are problems such as high manual participation, low efficiency and high labor costs in the training and evaluation of existing AI models, and the training and evaluation cannot be carried out in parallel.

Method used

Implement the continuous training and evaluation process of AI models on the containerized application management platform. By building process task templates, performing phase splitting and product escape analysis, the concurrent operation of training phase tasks and evaluation phase tasks is realized.

Benefits of technology

It improves the training and evaluation efficiency of AI models, reduces manual participation, realizes the parallelization of training and evaluation, and improves development efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120236160A_ABST
    Figure CN120236160A_ABST
Patent Text Reader

Abstract

The invention provides a training evaluation device and method for an AI model. The method comprises the steps that a process task template applied to a containerized application management platform is created, wherein the process task template comprises process configuration information associated with a training evaluation task of an AI model and stage marks corresponding to all steps in the training evaluation task; and in response to a received starting command of the continuous training evaluation task, querying a process task template, performing stage splitting on the training evaluation task according to the process configuration information and the stage mark to construct a training stage task, performing product escape analysis according to the training stage task, and performing product escape analysis according to the product escape analysis. In one embodiment, the method includes receiving a training stage task, generating a product escape analysis result associated with the training stage task, initiating the training stage task based on the product escape result, and initiating concurrent operation of one or more evaluation stage tasks without interrupting the training stage task when the training stage task reaches a predetermined training stage node.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application generally relates to the field of artificial intelligence (AI), and more specifically to an apparatus and method for training and evaluating an AI model. Background Art

[0002] In recent years, with the rapid development of artificial intelligence technology, more and more AI models have been applied to various fields. The generation of an AI model usually goes through two stages: model training and model evaluation.

[0003] Model training refers to the process of training and optimizing model parameters through a large amount of training data; while model evaluation refers to the process of testing the AI model generated during the training process through some pre-prepared and designed test data. After the model evaluation process, some test results of the evaluated model in the test data are usually generated, such as images, metrics, etc. These metrics can reflect the effectiveness of the AI model, thus helping algorithm personnel select better-performing models for production applications.

[0004] Currently, the training and evaluation of AI models are usually completed entirely manually by algorithm personnel. For example, after an algorithm personnel trains a model, they will temporarily interrupt the training, manually take out the model and run the evaluation process to evaluate the model. After the evaluation is completed, they will decide whether to continue the training based on the results of the model evaluation; or after all the training is completed, they will take out multiple models generated during the training at the same time, run the evaluation process for each model in turn. After all the evaluation processes are completed, they will then manually select the optimal model.

[0005] This fully manual training and evaluation process has problems such as low evaluation efficiency and high labor costs due to the inclusion of too many manual operations. Summary of the Invention

[0006] In view of the above problems, according to one aspect of the embodiments of the present application, there is provided an apparatus for training and evaluation of an AI model, including: a memory for storing program instructions; and a processor coupled to the memory and configured to execute the program instructions to: create a process task template applied to a containerized application management platform, the process task template including: process configuration information associated with the training and evaluation task of the AI model, and stage flags corresponding to each step in the training and evaluation task; and in response to receiving a start command for a continuous training and evaluation task, perform the following operations: query the process task template, split the training and evaluation task into stages according to the process configuration information and stage flags in the process task template to construct a training stage task, perform product escape analysis based on the training stage task to generate a product escape analysis result associated with the training stage task, start the training stage task based on the product escape result associated with the training stage task, and when the training stage task reaches a predetermined training stage node, start the concurrent execution of one or more evaluation stage tasks without interrupting the training stage task.

[0007] According to another aspect of the embodiments of the present application, there is provided a method for training and evaluation of an AI model, including: creating a process task template applied to a containerized application management platform, the process task template including: process configuration information associated with the training and evaluation task of the AI model, and stage flags corresponding to each step in the training and evaluation task; and in response to receiving a start command for a continuous training and evaluation task: query the process task template, split the training and evaluation task into stages according to the process configuration information and stage flags in the process task template to construct a training stage task, perform product escape analysis based on the training stage task to generate a product escape analysis result associated with the training stage task, start the training stage task based on the product escape result associated with the training stage task, and when the training stage task reaches a predetermined training stage node, start the concurrent execution of one or more evaluation stage tasks without interrupting the training stage task.

[0008] According to yet another aspect of the embodiments of the present application, there is provided a machine-readable medium having program instructions stored thereon, and when the program instructions are executed by a processor, the processor is caused to execute the method for training and evaluation of an AI model as described above.

[0009] According to yet another aspect of the embodiments of the present application, there is provided a device for training and evaluation of an AI model, including an apparatus for executing the method for training and evaluation of an AI model as described above.

[0010] The device, method, equipment, and machine-readable medium for training and evaluation of an AI model according to embodiments of the present application abstract and solidify the process configuration of training and evaluation of the AI model, realize continuous training and evaluation of the AI model, reduce manual participation in the training and evaluation process, and solve the problem that model training and model evaluation cannot be carried out in parallel, thereby improving the efficiency of AI model development.

[0011] In addition, according to some embodiments of the present application, a unified method for single / continuous training and evaluation of an AI model can also be realized. In these embodiments, the platform can respond to a single training and evaluation start command or a continuous training and evaluation start command input by the user, and start the single training and evaluation or continuous training and evaluation program accordingly. That is, after unified training and evaluation configuration, the user can either choose to run a single training and evaluation task or, without modifying the configuration, choose to run a continuous training and evaluation task. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] The features, advantages, and technical effects of exemplary embodiments of the present application will be described below with reference to the accompanying drawings.

[0013] To more clearly illustrate the technical solutions of the embodiments of the present application, the accompanying drawings required for use in the embodiments of the present application will be briefly introduced below. For those of ordinary skill in the art, other drawings can also be obtained based on these drawings without creative efforts.

[0014] Figure 1 FIG. shows a schematic overall flowchart of the training and evaluation process of an AI model according to some embodiments of the present application;

[0015] Figure 2 FIG. shows an example process of the process configuration in the training and evaluation process of an AI model according to some embodiments of the present application;

[0016] Figure 3 FIG. shows the single training and evaluation task process of an AI model according to some embodiments of the present application;

[0017] Figure 4 FIG. shows the overall flowchart of the continuous training and evaluation task process of an AI model according to some embodiments of the present application;

[0018] Figure 5 FIG. shows a schematic diagram of the stage splitting process in the continuous training and evaluation task process of an AI model according to some embodiments of the present application;

[0019] Figure 6 FIG. shows a schematic diagram of the process of using a sidecar proxy to assist in the transfer of product files in the continuous training and evaluation task process of an AI model according to some embodiments of the present application;

[0020] Figure 7 Shows a schematic diagram of non-invasive Sidecar proxy container injection in the continuous training and evaluation task process of an AI model according to some embodiments of the present application;

[0021] Figure 8 Shows a schematic diagram of concurrent execution of multiple model evaluation tasks in the continuous training and evaluation task process of an AI model according to some embodiments of the present application;

[0022] Figure 9 Shows a detailed flowchart of the continuous training and evaluation task process of an AI model according to some embodiments of the present application;

[0023] Figure 10 Shows a schematic diagram of the hardware structure of a training and evaluation device for an AI model according to some embodiments of the present application. Detailed implementation manners

[0024] The features and exemplary embodiments of various aspects of the present application will be described in detail below. In the following detailed description, many specific details are set forth in order to provide a thorough understanding of the present application. However, it will be apparent to those skilled in the art that the present application may be practiced without some of these specific details. The following description of the embodiments is only intended to provide a better understanding of the present application by showing examples of the present application. The present application is in no way limited to any specific configuration set forth below, but covers any modification, replacement, and improvement of elements, components, and algorithms without departing from the spirit of the present application. Well-known structures and technologies are not shown in the drawings and the following description in order to avoid unnecessarily obscuring the present application.

[0025] To solve the problems of high manual participation, low evaluation efficiency, and high labor costs in the full-manual training and evaluation workflow in artificial intelligence model training and evaluation, a continuous training and evaluation process of an AI model that can be implemented on a containerized application management platform (for example, a Kubernetes-based platform) is proposed in the present application. This process abstracts and solidifies the training and evaluation process configuration of the AI model, and solves the problem that model training and model evaluation cannot be performed in parallel. In addition, the platform for implementing the continuous training and evaluation process of the AI model according to the present application can also easily be compatible with the single training and evaluation process of the AI model. Therefore, according to some embodiments of the present application, a unified single or continuous training and evaluation process of the AI model can be realized. The overall process of this unified training and evaluation process will be described below with reference to Figure 1 and then with reference to Figures 2 to 9Describe in detail the process of abstracting and solidifying the training and evaluation process configuration of the AI model, as well as the specific processes of a single training and evaluation process and a continuous training and evaluation process. In particular, the implementation process of the continuous training and evaluation process will be described in detail.

[0026] Figure 1 FIG. shows a schematic overall flowchart of the training and evaluation process of an AI model according to some embodiments of the present application. As Figure 1 shown, the overall process of training and evaluating the AI model may include a process configuration process, a start training process, and a single training and evaluation task process or a continuous training and evaluation task process started according to user selection.

[0027] Generally speaking, this training and evaluation process can be implemented on a containerized application management platform. First, the platform can build the task step composition of the training and evaluation task of the AI model based on knowledge in the field of artificial intelligence. The task step composition includes each step required to complete the training and evaluation task of the AI model. For example, for the field of image detection, a task step composition including main steps such as image preprocessing steps, model training steps, model evaluation steps, and post-processing steps can be built. The platform can solidify these task steps into a process task template by software methods for subsequent running of the training and evaluation task based on this process task template. Then, based on the solidified process task template, the platform can start a single training and evaluation task or a continuous training and evaluation task of the AI model in response to a training and evaluation start command input by the user, thereby realizing the automation and parallelization of the training and evaluation function of the AI model and providing a unified storage and display method for models and reports.

[0028] In specific implementation manners, for example, by using relevant technical means such as a containerized application management engine (e.g., Kubernetes), sidecar mechanism, object storage, process engine, software development kit (SDK), etc., the training and evaluation task process of the AI model is configured in a templated manner. At the same time, through mechanisms such as stage splitting and product escape analysis, it is possible to select to perform a single training and evaluation task process or a continuous training and evaluation task process according to user needs. If the user selects to start a single training and evaluation task process, a complete process including the training and evaluation stages of the AI model will be run, which is similar to the fully manual method, but each step in the training and evaluation process will be automatically completed in sequence without much manual intervention. If the user selects the continuous training and evaluation task process, the process tasks will be automatically split into stages and product escape analysis will be performed. Based on the results of the stage splitting and product escape analysis, the training stage task (also simply referred to as the training task in this article) will be started, and when the training stage task reaches a predetermined training stage node, one or more evaluation stage tasks (also simply referred to as the evaluation tasks in this article) will be started to run concurrently without interrupting the training stage task, so as to achieve the parallel operation of the training stage task and the evaluation stage task.

[0029] The following will combine Figures 2 to 9 Describe in detail the process of configuring the training and evaluation process of the AI model, as well as the specific processes of the single training and evaluation process and the continuous training and evaluation process. In particular, reference will be made to Figures 4 to 9 Focus on describing the specific implementation process of the continuous training and evaluation process.

[0030] Figure 2 FIG. shows an example process of process configuration in the training and evaluation process of an AI model according to some embodiments of the present application. According to the embodiments of the present application, process configuration refers to the process of mapping relevant running codes, configurations, running environments, etc. in the algorithm expert workflow in the AI-related field to the abstracted running steps (for example, in the field of image detection, it may include image preprocessing steps, model training steps, model evaluation steps, post-processing steps, etc.) through software methods, that is, the process of solidifying the workflow of algorithm experts in the AI-related field into a process task template.

[0031] For example, the technical solution according to the embodiments of the present application can be implemented on a platform based on a containerized application management engine (e.g., Kubernetes). All training and evaluation tasks are ultimately scheduled and run in the form of containers (Pods). Therefore, in specific implementation, the relevant configurations of the algorithm expert workflow in the AI field can be mapped to parameters such as container startup commands, container images, stage products, configuration files, etc., and then a process task template is generated at the software level according to these parameters and solidified into the training and evaluation process of the AI model.

[0032] Specifically, in the process configuration stage, the training and evaluation process of the AI model can be mapped and parameter-configured to generate a process task template associated with the training and evaluation task of the AI model for running on the containerized application management platform. This process task template may include process configuration information associated with the training and evaluation task of the AI model. As Figure 2 shown, for example, for an image detection model, the process configuration information may include an image detection training set associated with the image preprocessing step, training set preprocessing configuration, an image detection test set, and test set preprocessing configuration; mirror configuration, startup command configuration, resource configuration, product configuration, and input file configuration associated with the model training step; mirror configuration, startup command configuration, resource configuration, product configuration, and input file configuration associated with the model evaluation step; and model evaluation report analysis and model post-processing configuration associated with the post-processing step.

[0033] In addition, as described above, according to some embodiments of the present application, when the user selects to start the continuous training and evaluation task process, the platform needs to automatically perform stage splitting and product escape analysis on the process task, start the training stage task based on the results of the stage splitting and product escape analysis, and when the training stage task reaches a predetermined training stage node, start the concurrent execution of one or more evaluation stage tasks without interrupting the training stage task.

[0034] In these embodiments, it is proposed that when creating the process task template, based on the domain knowledge associated with the AI model, the composition of the task steps of the training and evaluation task of the AI model is constructed. This composition of task steps includes each step required to complete the training and evaluation task of the AI model; and a stage to which each step in the composition of task steps belongs is assigned to obtain a stage flag corresponding to this step. The stage flags corresponding to each step may be included in the process task template to provide a basis for the construction of subsequent stage tasks.

[0035] For example, for the training and evaluation task of an image detection model, the composition of task steps usually includes an image preprocessing step, a model training step, a model evaluation step, and a post-processing step. These steps can be divided into stages based on the domain knowledge of image detection, that is, a stage to which each step belongs is assigned to provide a basis for the construction of subsequent stage tasks. For example, the training and evaluation task of an image detection model can be divided into two stages: the model training stage and the model evaluation stage. Stage flags 1 and 2 are assigned to these two stages respectively, and then corresponding stage flags are assigned to each step in the composition of task steps. For example, since the image preprocessing step is required in both the training stage task and the evaluation stage task, stage flags 1 and 2 should be assigned to the image preprocessing step, while only stage flag 1 is assigned to the model training step and only stage flag 2 is assigned to the model evaluation step.

[0036] Therefore, in the process task template created on the platform, it not only includes the process configuration information associated with the training and evaluation tasks of the AI model, but also includes the stage flags corresponding to each step in the training and evaluation tasks. In this way, when the platform needs to start the stage task corresponding to a certain stage flag, it can construct the corresponding stage task to be started according to the stage flags corresponding to each step in the process task template. For example, when it is necessary to start the model training stage task (corresponding to stage flag 1), the model training stage task can be constructed based on the steps in which the stage flag configured in the task step composition includes stage flag 1, so as to start the training stage task subsequently.

[0037] As Figure 1 shown, after the configuration of the process task template for the training and evaluation process of the AI model is completed, the training and evaluation task can be started with one key. When starting the training and evaluation task, the user can choose to perform a single training and evaluation task or a continuous training and evaluation task according to needs. After the training and evaluation task is started, the platform will, according to the process task template created in the process configuration stage, run the model training container and the model evaluation container based on the Kubernetes container operation model and monitor their running status.

[0038] According to an embodiment of the present application, if the user selects to start a single training and evaluation task when starting the training, the platform will directly run the single training and evaluation task of the AI model, that is, sequentially execute the single training stage task and the single evaluation stage task of the AI model based on the process task template to complete the single training and evaluation task. Figure 3 shows the single training and evaluation task process of the AI model according to some embodiments of the present application.

[0039] As Figure 3As shown below, taking an image detection model as an example for illustration, the single training and evaluation task process of the image detection model may include a preprocessing stage, a training stage, and an evaluation stage. Specifically, in the preprocessing stage of the training and evaluation of the image detection model, the following operations are mainly carried out: Automatically pull the required image data for training and testing from the object storage according to the image pulling conditions configured by experts in the field of image detection, and at the same time, according to the image preprocessing algorithm configured by algorithm experts, automatically preprocess the pulled image data to generate preprocessing products (image data). After pulling the image data required for the training and evaluation task of the image detection model and preprocessing the image data, the training stage container will be automatically run. In the training stage, the training code of the image detection model will be run, and the preprocessed image data will be used to train the image detection model. Some product files will be generated in the training stage, and the most important product file contains the model file. At the software level, operations such as archiving these product files will be automatically completed. After the training stage is completed, the evaluation stage of the image detection model will be sequentially executed. The basic model file used in the training stage and the model file generated after training will be automatically transmitted to the evaluation stage container. At the same time, the relevant image test sets selected by algorithm experts in the field of image detection will also be automatically transmitted to the evaluation stage. The evaluation stage container generates some model report files by running the evaluation code. For the image detection model, the report usually contains indicators such as image segmentation result diagrams, object detection diagrams, positive detection rates of object detection, false alarm rates of object detection, and missed detection rates of object detection. Similarly, at this stage, operations such as archiving these product files will be automatically completed at the software level.

[0040] According to some embodiments of the present application, if the user selects to start a continuous training and evaluation task when starting the training, the platform will automatically perform stage splitting and product escape analysis on the training and evaluation task, start the training stage task based on the results of the stage splitting and product escape analysis, and when the training stage task reaches a predetermined training stage node, one or more evaluation stage tasks can be started to run concurrently without interrupting the training stage task to perform continuous training and evaluation. Figure 4 Fig. shows the overall flowchart of the continuous training and evaluation task process of the AI model according to some embodiments of the present application.

[0041] Table 1 below shows some key operations involved in the continuous training and evaluation task process of the AI model in the exemplary embodiments according to the present application. The serial numbers in Table 1 correspond to the serial numbers of these operations shown in Figure 4 shown above.

[0042]

[0043] Table 1

[0044] As combined aboveFigure 1 As described, the continuous training and evaluation task process involves multiple aspects such as stage splitting, product escape analysis, product storage and transfer, and parallel execution of model training tasks and model evaluation tasks. The operations shown in Table 1 are actually operations related to these aspects. The following will combine Figures 5 to 9 to describe the operations related to these aspects in more detail.

[0045] Figure 5 FIG. shows a schematic diagram of the stage splitting process in the continuous training and evaluation task process of an AI model according to some embodiments of the present application. As Figure 5 shown, stage splitting refers to automatically splitting the evaluation stage task in the training and evaluation task of the AI model from the process, that is, only starting the training stage task first. Specifically, when the platform receives a start command for the continuous training and evaluation task, the platform will query the process task template and, based on the process configuration information and stage flags in the process task template, perform stage splitting on the training and evaluation task to construct the training stage task of the AI model.

[0046] As mentioned above, according to the embodiments of the present application, when creating the process task template, based on the domain knowledge associated with the AI model, the composition of the task steps of the training and evaluation task of the AI model is constructed, and the stage to which each step in the task step composition belongs is assigned to obtain the corresponding stage flag for the step. The stage flags corresponding to each step can be included in the process task template to provide a basis for the construction of subsequent stage tasks. In this way, when the platform needs to start the stage task corresponding to a certain stage flag, the corresponding stage task to be started can be constructed according to the stage flags corresponding to each step in the process task template.

[0047] Based on the stage flags configured in the process task template, when it is necessary to start the training stage task corresponding to stage flag 1, in the stage splitting process, the training stage task can be constructed based on each step in the task step composition whose configured stage flag includes stage flag 1.

[0048] In addition, according to the embodiments of the present application, after the platform performs stage splitting to construct the training stage task, it is also necessary to perform product escape analysis. Specifically, the algorithm expert can configure corresponding products (for example, including input products and output products) for each step involved in the training and evaluation task. These products need to be transferred and used between multiple steps. However, due to the splitting of the stage tasks, the original product transfer method between the steps in a single task may no longer be applicable. Therefore, it is necessary to analyze and specially process the situation of product cross-stage transfer, which is also referred to as performing product escape analysis in this article.

[0049] Specifically, when starting a task for a certain phase, first obtain the process task template. The process task template contains the input and output products of each step configured by algorithm experts. Traverse the input and output products of these steps and analyze to obtain the transfer relationships between these products among the steps (for example, product 1 is generated by step 1 and used as input in step 2). After analyzing the transfer relationships between the products among the steps, filter out the products whose transfer relationships span different phases in the training and evaluation tasks (for example, the product needs to be transferred between step 1 and step 2, where step 1 belongs to the training phase with a phase flag of 1 and step 2 belongs to the evaluation phase with a phase flag of 2). These products are then processed as escape products. When there are escape products, the platform first cancels the original processing logic for these escape products in the original single task, and then based on the phase flags corresponding to each step in this phase task, establishes the transfer processing logic for the escape products as the product escape analysis result associated with this phase task.

[0050] It should be noted that when starting the training phase task and the evaluation phase task, the situation of product cross-phase transfer may occur. Therefore, corresponding product escape analysis needs to be carried out. Specifically, the product escape analysis carried out when starting the training phase task includes: based on the process task template, analyzing the transfer relationships between the products of each step in the training phase task among the steps; filtering out the products whose transfer relationships span different phases in the training and evaluation tasks and processing them as escape products; and based on the phase flags corresponding to each step in the training phase task, establishing the transfer processing logic for the escape products as the product escape analysis result associated with the training phase task. Correspondingly, the product escape analysis carried out when starting the evaluation phase task includes: based on the process task template, analyzing the transfer relationships between the products of each step in the evaluation phase task among the steps; filtering out the products whose transfer relationships span different phases in the training and evaluation tasks and processing them as escape products; and based on the phase flags corresponding to each step in the evaluation phase task, establishing the transfer processing logic for the escape products as the product escape analysis result associated with the evaluation phase task.

[0051] In addition, regarding the transfer of the product escape analysis result, according to the embodiments of the present application, it is proposed that the product escape analysis result can be transferred in the following manner: configure an associated proxy container for the phase task, and use this proxy container to store the product file according to the product escape analysis result and assist in realizing the transfer of the product file between the training phase task and the evaluation phase task.

[0052] For example, the product escape analysis result associated with the training phase task can be passed to the sidecar proxy container in the form of environment variables. At the same time, the sidecar proxy container shares the storage with the training container by mounting the same storage, that is, it can obtain the products generated by the training phase task. After the training task reaches a phased result, the training container notifies the sidecar proxy container to start the model evaluation task through a dedicated software development kit (SDK). The sidecar proxy container processes all the escaped products (such as packaging, archiving, etc.) according to the product escape analysis result obtained from the environment variables, and finally starts the model evaluation task through the platform service interface, and passes the archived information of the current escaped products to the platform service through the interface.

[0053] Figure 6 FIG. shows a schematic diagram of the process of using a sidecar proxy to assist in the transfer of product files in the continuous training and evaluation task process of an AI model according to some embodiments of the present application. As Figure 6 shown, when the product transfer spans different stage tasks, the product needs to be marked as an escaped product, and when initiating an independent evaluation stage task subsequently, the sidecar proxy container can assist in the storage and transfer of the product files.

[0054] According to some embodiments of the present application, in order to autonomously complete operations such as updating the preprocessed products of the dataset used for training, storing the products (models) in the training phase, and starting concurrent model evaluation tasks during the operation of the model training container, a non-invasive proxy container injection method is adopted in the technical implementation. This proxy container can be, for example, a Sidecar proxy container based on the Sidecar architecture mode for the containerized application management platform.

[0055] Figure 7 FIG. shows a schematic diagram of non-invasive Sidecar proxy container injection in the continuous training and evaluation task process of an AI model according to some embodiments of the present application. Since actions such as initiating the storage of training phase products and updating preprocessed dataset products in the training phase task involve complex and time-consuming operations, and starting concurrent evaluation tasks requires calling the platform service interface, it is not appropriate for the training container to directly implement these operations. To solve this problem, in the embodiments of the present application, a method combining an SDK and a Sidecar proxy container is proposed. The Sidecar proxy container completes complex and time-consuming operations and calls the platform service interface, and only a simple interface layer is encapsulated at the SDK level.

[0056] As Figure 7As shown, the training phase task is configured with an associated Sidecar proxy container, which can interact with the training phase task and the evaluation phase task through a dedicated SDK to store products and transfer products between the Sidecar proxy container, the training phase task, and the evaluation phase task. Specifically, by automatically injecting a self-developed Sidecar proxy container into the container to which the training task belongs on Kubernetes, the training container only needs to communicate with the Sidecar proxy container through the local network (this communication will also be encapsulated by the python SDK). At the same time, the Sidecar proxy container and the training container can share part of the file system, that is, synchronize / asynchronize the required file storage or transfer operations as needed, or call the remote platform service interface to start the concurrent operation of multiple model evaluation tasks.

[0057] Figure 8 The figure shows a schematic diagram of the concurrent operation of multiple model evaluation tasks in the continuous training and evaluation task process of an AI model according to some embodiments of the present application. As Figure 8 shown, in the embodiments according to the present application, the training task and the evaluation task can run in parallel without affecting each other, and at the same time, multiple evaluation tasks can also run concurrently. For example, when the training phase task reaches a predetermined training phase node, one or more evaluation phase tasks can be started to run concurrently without interrupting the training phase task. In the embodiments of the present application, the predetermined training phase node may include: a node where the training phase task reaches the stage model training target, or a node where the training phase task reaches the predefined number of training iterations. For example, when the training phase task reaches a predetermined first training phase node, a first evaluation phase task can be started for the training model obtained at the first training phase node; when the training phase task reaches a predetermined second training phase node, a second evaluation phase task can be started for the training model obtained at the second training phase node; and so on. While starting the concurrent operation of multiple evaluation phase tasks, the training phase task can continue. In addition, as shown in Table 1, the Sidecar proxy container can be used to call the platform service interface to start the concurrent operation of multiple model evaluation tasks.

[0058] According to the embodiments of the present application, adopting the method of combining the SDK with the Sidecar proxy to implement the storage and transfer of product files and the call of the remote platform service interface will have the following advantages: After separating this part of the logic from the training code of the AI model, the difficulty of algorithm implementation and the complexity of the code are reduced; the specific implementation is completed by the Sidecar proxy, realizing the separation of logic and algorithm, and increasing the flexibility of modifying the implementation logic in the future; avoiding the training code from directly calling the platform service interface and preventing the exposure of the platform service interface.

[0059] As described above in conjunction with Figures 4 to 8 the various possible operations and steps involved in the continuous training and evaluation process of an AI model based on a containerized application management platform have been described. To more clearly illustrate the overall process of this continuous training and evaluation process, the following will refer to Figure 9 to further describe this continuous training and evaluation process.

[0060] Figure 9 FIG. shows a detailed flowchart of the continuous training and evaluation task process of an AI model according to some embodiments of the present application. This continuous training and evaluation task process can be implemented on a containerized application management platform. Here, it should be noted that in the previous description, the training and evaluation task process provided by the present application (including the single training and evaluation task process and the continuous evaluation task process) was mainly described by taking the image detection model in the field of image processing as an example. However, it can be understood that this training and evaluation task process can be applied to the training and evaluation of different AI models in various fields. For different AI models in different fields, the platform can create a process task template applied to the containerized application management platform according to the process configuration information associated with the training and evaluation tasks of the AI model input by algorithm experts in the relevant field. This process task template may include: process configuration information associated with the training and evaluation tasks of the AI model, and phase flags corresponding to each step in the training and evaluation tasks. Therefore, there may be multiple process task templates associated with different AI models in various fields on the platform. Correspondingly, when starting a certain training or evaluation task, it is necessary to query the process task template to which the training or evaluation task belongs.

[0061] As Figure 9 shown, in response to receiving a start command for a continuous training and evaluation task, the platform can perform the following operations: query the process task template, split the training and evaluation task into phases according to the process configuration information and phase flags in the process task template to construct training phase tasks, perform product escape analysis based on the training phase tasks to generate a product escape analysis result associated with the training phase tasks, and start the training phase tasks based on the product escape result associated with the training phase tasks.

[0062] During the execution of the training phase task, the following operations can be performed: perform model training iterations; during the iteration process, notify the sidecar proxy to update the preprocessed product of the dataset as needed through the SDK; when the training phase task reaches a predetermined training stage node (for example, reaches the stage model training goal or the predefined number of training iterations), without interrupting the training phase task, notify the sidecar proxy to start the evaluation phase task through the SDK; use the sidecar proxy to store the escaped products according to the stage running configuration and the product escape analysis result; the sidecar proxy calls the platform service interface to start the evaluation phase task; continue with subsequent model training.

[0063] As Figure 9 shown, the platform can start a corresponding evaluation phase task each time the training phase task reaches a predetermined training stage node without interrupting the training phase task, so that multiple evaluation phase tasks (evaluation task 1 to evaluation task n) can be concurrently run on the platform. Moreover, for each evaluation phase task, the following operations can be performed: query the process task template; construct the evaluation phase task according to the process configuration information and the stage flag in the process task template, perform product escape analysis according to the evaluation phase task to generate a product escape analysis result associated with the evaluation phase task, and based on the product escape result associated with the evaluation phase task, start the evaluation phase task.

[0064] In summary, by using the continuous training and evaluation process according to the present application, the automation and parallelization of the training and evaluation functions of various AI models in different fields can be realized. In addition, the SDK combined with the Sidecar proxy is used to implement the storage and transfer of product files and the invocation of the platform service interface, so that the training task container does not need to directly call the platform service interface, avoiding the exposure of the platform service interface.

[0065] Figure 10 Fig. shows a schematic hardware structure diagram of an AI model training and evaluation device according to an embodiment of the present application.

[0066] The device may include a processor 1001 and a memory 1002 storing program instructions.

[0067] Specifically, the above-mentioned processor 1001 may include a central processing unit (CPU), or an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present application.

[0068] The memory 1002 may include a mass storage for data or instructions. By way of example and not limitation, the memory 1002 may include a hard disk drive (HDD), a floppy disk drive, a flash memory, an optical disk, a magneto-optical disk, a magnetic tape, or a universal serial bus (USB) drive, or a combination of two or more of these. In a suitable case, the memory 1002 may include removable or non-removable (or fixed) media. In a suitable case, the memory 1002 may be inside or outside the integrated gateway disaster recovery device. In a particular embodiment, the memory 1002 is a non-volatile solid-state memory.

[0069] In a particular embodiment, the memory 1002 includes a read-only memory (ROM). In a suitable case, the ROM may be a mask-programmed ROM, a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), an electrically rewritable ROM (EAROM), or a flash memory, or a combination of two or more of these.

[0070] The memory may include a read-only memory (ROM), a random access memory (RAM), a magnetic disk storage media device, an optical storage media device, a flash memory device, an electrical, optical, or other physical / tangible storage device. Thus, generally, the memory includes one or more tangible (non-transitory) readable storage media (e.g., memory devices) encoded with software including computer-executable instructions, and when the software is executed (e.g., by one or more processors), it is operable to perform the operations described with reference to the training and evaluation method of the AI model according to embodiments of the present application.

[0071] The processor 1001 reads and executes the program instructions stored in the memory 1002 to implement the training and evaluation method of the AI model described in the above embodiments.

[0072] In one example, the training and evaluation device of the AI model may further include a communication interface 1003 and a bus 1010. Among them, as Figure 10 shown, the processor 1001, the memory 1002, and the communication interface 1003 are connected through the bus 1010 and communicate with each other.

[0073] The communication interface 1003 is mainly used to implement the communication between the various modules, devices, units, and / or devices in the embodiments of the present application.

[0074] Bus 1010 includes hardware, software, or both, and couples components of the device together. By way of example and not limitation, the bus may include an Accelerated Graphics Port (AGP) or other graphics bus, an Enhanced Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), a HyperTransport (HT) interconnect, an Industry Standard Architecture (ISA) bus, an InfiniBand interconnect, a Low Pin Count (LPC) bus, a memory bus, a MicroChannel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCI-X) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local (VLB) bus, or other suitable bus or a combination of two or more of these. Where appropriate, bus 1010 may include one or more buses. Although embodiments of the present application describe and illustrate specific buses, the present application contemplates any suitable bus or interconnect.

[0075] In combination with the AI model training and evaluation method in the above embodiments, the embodiments of the present application can be implemented by a machine-readable medium. Program instructions are stored on the machine-readable medium; when the program instructions are executed by a processor, the processor executes the AI model training and evaluation method described in the above embodiments.

[0076] It should be clear that the present application is not limited to the specific configurations and processes described above and illustrated in the figures. For the sake of brevity, detailed descriptions of known methods are omitted here. In the above embodiments, several specific steps are described and illustrated as examples. However, the method process of the present application is not limited to the specific steps described and illustrated, and those skilled in the art can make various changes, modifications, and additions, or change the order between steps after understanding the spirit of the present application.

[0077] The functional blocks shown in the above-described block diagrams can be implemented as hardware, software, firmware, or a combination thereof. When implemented in hardware, it can be, for example, an electronic circuit, an Application Specific Integrated Circuit (ASIC), appropriate firmware, a plug-in, a functional card, and so on. When implemented in software, the elements of the present application are programs or code segments for performing the required tasks. The program or code segment can be stored in a machine-readable medium or transmitted via a data signal carried in a carrier wave on a transmission medium or communication link. "Machine-readable medium" can include any medium capable of storing or transmitting information. Examples of machine-readable media include electronic circuits, semiconductor memory devices, ROMs, flash memories, Erasable ROMs (EROMs), floppy disks, CD-ROMs, optical discs, hard disks, fiber optic media, Radio Frequency (RF) links, and so on. The code segment can be downloaded via a computer network such as the Internet, an intranet, and so on.

[0078] It should also be noted that in the exemplary embodiments mentioned in this application, some methods are described based on a series of steps. However, this application is not limited to the order of the above steps. That is to say, the steps can be executed in the order mentioned in the embodiments, can be executed in an order different from that in the embodiments, or several steps can be executed simultaneously.

[0079] The above has described various aspects of this application with reference to the flowcharts and / or block diagrams of the methods and apparatuses according to the embodiments of this application. It should be understood that each block in the flowchart and / or block diagram, as well as the combinations of blocks in the flowchart and / or block diagram, can be implemented by program instructions. These program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, or other programmable data processing devices to generate a machine, such that these instructions executed by the processor of the computer or other programmable data processing devices enable the implementation of the functions / actions specified in one or more blocks of the flowchart and / or block diagram. Such a processor can be, but is not limited to, a general-purpose processor, a special-purpose processor, a special application processor, or a field-programmable logic circuit. It can also be understood that each block in the block diagram and / or flowchart, as well as the combinations of blocks in the block diagram and / or flowchart, can also be implemented by dedicated hardware that executes the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.

[0080] The above description is only the specific implementation manner of this application. Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working process of the above-described device can refer to the corresponding process in the foregoing method embodiments and will not be elaborated herein. It should be understood that the protection scope of this application is not limited thereto. Any person skilled in the art within the technical scope disclosed in this application can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should all be covered within the protection scope of this application.

Claims

1. A device for training and evaluation of an artificial intelligence (AI) model, comprising: a memory for storing program instructions; and a processor coupled to the memory and configured to execute the program instructions to: create a process task template applied to a containerized application management platform, the process task template including: process configuration information associated with the training and evaluation task of the AI model, and phase flags corresponding to each step in the training and evaluation task; and in response to receiving a start command for a continuous training and evaluation task, perform the following operations: query the process task template, split the training and evaluation task into phases according to the process configuration information and the phase flags in the process task template to construct a training phase task, perform product escape analysis based on the training phase task to generate a product escape analysis result associated with the training phase task, start the training phase task based on the product escape result associated with the training phase task, and when the training phase task reaches a predetermined training phase node, start the concurrent execution of one or more evaluation phase tasks without interrupting the training phase task.

2. The device according to claim 1, wherein Starting the concurrent execution of one or more evaluation phase tasks includes performing the following operations for each evaluation phase task among the one or more evaluation phase tasks: query the process task template; construct the evaluation phase task according to the process configuration information and the phase flags in the process task template, perform product escape analysis based on the evaluation phase task to generate a product escape analysis result associated with the evaluation phase task, and start the evaluation phase task based on the product escape result associated with the evaluation phase task.

3. The device according to claim 1, wherein The predetermined training phase node includes: a node where the training phase task reaches a phased model training target, or a node where the training phase task reaches a predefined number of training iterations.

4. The device according to claim 1, wherein Performing product escape analysis based on the training phase task includes: analyzing the transfer relationship of products between each step in the training phase task based on the process task template, where the products include input products and output products; screening out products whose transfer relationship spans different phases in the training and evaluation task as escape products for processing; and establishing the transfer processing logic of the escape products based on the phase flags corresponding to each step in the training phase task as the product escape analysis result associated with the training phase task.

5. The apparatus according to claim 2, wherein Performing product escape analysis based on the evaluation phase task includes: analyzing the transfer relationship of products between each step in the evaluation phase task based on the process task template, where the products include input products and output products; screening out products whose transfer relationship spans different phases in the training and evaluation task as escape products for processing; and Based on the stage flags corresponding to each step in the evaluation stage task, establish the transfer processing logic of the escape product as the product escape analysis result associated with the evaluation stage task.

6. The device according to any one of claims 1 to 5, wherein The training stage task is configured with an associated proxy container, and starting the evaluation stage task in the training and evaluation task includes: using the proxy container to start the concurrent execution of the one or more evaluation stage tasks by calling the service interface of the containerized application management platform.

7. The device according to claim 6, wherein The operation further includes: using the proxy container to store the product file according to the product escape analysis result and assisting in realizing the transfer of the product file between the training stage task and the one or more evaluation stage tasks.

8. The apparatus according to claim 6, wherein, The training stage task interacts with the proxy container through a dedicated software development kit (SDK) to notify the proxy container to update the product associated with the training stage task and to start the concurrent execution of the one or more evaluation stage tasks.

9. The apparatus according to claim 6, wherein, The proxy container is a Sidecar proxy container based on the Sidecar architecture mode for the containerized application management platform.

10. The device according to claim 1, wherein The processor is further configured to execute the program instructions to: In response to receiving a start command for a single training and evaluation task, sequentially execute the single training stage task and the single evaluation stage task of the AI model based on the process task template.

11. The device according to claim 1, wherein Creating a process task template applied to a containerized application management platform includes: Based on the domain knowledge associated with the AI model, constructing the task step composition of the training and evaluation task of the AI model, which includes all the steps required to complete the training and evaluation task of the AI model; and Assigning the stage to which each step in the task step composition belongs to obtain the stage flag corresponding to the step.

12. The device according to claim 11, wherein, The AI model includes an image detection model, and the task step composition includes an image preprocessing step, a model training step, a model evaluation step, and a postprocessing step.

13. A method for training and evaluating an artificial intelligence (AI) model, including: Creating a process task template applied to a containerized application management platform, the process task template including: process configuration information associated with the training and evaluation task of the AI model, and the stage flags corresponding to each step in the training and evaluation task; and In response to receiving a start command for a continuous training and evaluation task: Query the process task template, According to the process configuration information and the stage flags in the process task template, split the training and evaluation task into stages to construct a training stage task, Perform product escape analysis based on the training stage task to generate a product escape analysis result associated with the training stage task, Based on the product escape result associated with the training stage task, start the training stage task, and When the training stage task reaches a predetermined training stage node, start the concurrent execution of one or more evaluation stage tasks without interrupting the training stage task.

14. The method according to claim 13, wherein, Initiating concurrent execution of one or more evaluation phase tasks includes performing the following operations for each of the one or more evaluation phase tasks: Querying the process task template; Constructing the evaluation phase task based on the process configuration information and the phase flag in the process task template; Performing product escape analysis based on the evaluation phase task to generate a product escape analysis result associated with the evaluation phase task, and Initiating the evaluation phase task based on the product escape result associated with the evaluation phase task.

15. The method according to claim 13, wherein The predetermined training stage nodes include: the node where the training phase task reaches the stage model training target, or the node where the training phase task reaches the predefined number of training iterations.

16. The method according to claim 13, wherein, Performing product escape analysis based on the training phase task includes: Analyzing the transfer relationship of products between each step in the training phase task based on the process task template, where the products include input products and output products; Filtering out the products whose transfer relationship spans different stages in the training and evaluation task as escape products for processing; and Establishing the transfer processing logic of the escape products based on the phase flags corresponding to each step in the training phase task as the product escape analysis result associated with the training phase task.

17. The method according to claim 14, wherein, Performing product escape analysis based on the evaluation phase task includes: Analyzing the transfer relationship of products between each step in the evaluation phase task based on the process task template, where the products include input products and output products; Filtering out the products whose transfer relationship spans different stages in the training and evaluation task as escape products for processing; and Establishing the transfer processing logic of the escape products based on the phase flags corresponding to each step in the evaluation phase task as the product escape analysis result associated with the evaluation phase task.

18. The method according to any one of claims 13 to 17, wherein The training phase task is configured with an associated proxy container, and initiating the evaluation phase task in the training and evaluation task includes: using the proxy container to initiate concurrent execution of the one or more evaluation phase tasks by calling the service interface of the containerized application management platform.

19. The method according to claim 18, further comprising: Using the proxy container to store product files according to the product escape analysis result and assisting in realizing the transfer of the product files between the training phase task and the one or more evaluation phase tasks.

20. The method according to claim 18, wherein The training phase task interacts with the proxy container through a dedicated software development kit (SDK) to notify the proxy container to update the products associated with the training phase task and initiate concurrent execution of the one or more evaluation phase tasks.

21. The method according to claim 18, wherein The proxy container is a Sidecar proxy container based on the Sidecar architecture mode for the containerized application management platform.

22. The method according to claim 13, further comprising: In response to receiving a start command for a single training and evaluation task, sequentially executing the single training phase task and the single evaluation phase task of the AI model based on the process task template.

23. The method according to claim 13, wherein, Creating a process task template for a containerized application management platform includes: Based on the domain knowledge associated with the AI model, constructing the composition of the task steps for the training and evaluation task of the AI model, where the composition of the task steps includes each step required to complete the training and evaluation task of the AI model; and Assigning the phase to which each step in the composition of the task steps belongs to obtain the phase flag corresponding to the step.

24. The method according to claim 23, wherein, The AI model includes an image detection model, and the composition of the task steps includes an image preprocessing step, a model training step, a model evaluation step, and a post-processing step.

25. A machine-readable medium having program instructions stored thereon, which when executed by a processor cause the processor to execute the method for training and evaluating an artificial intelligence (AI) model as recited in any one of claims 13 to 24.

26. An apparatus for training and evaluating an artificial intelligence (AI) model, including means for executing the method for training and evaluating an AI model as recited in any one of claims 13 to 24.