Multi-stage collaborative algorithm modeling method, device, system and equipment and storage medium
Through the multi-level collaborative algorithm modeling method, the disconnection between model training and application feedback in the existing technology is solved, and the algorithm closed loop from training to application is realized, which improves the adaptability and applicability of the model.
Patent Information
- Application Number
- CN202311453240.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-02
- Publication Date
- 2025-05-06
AI Technical Summary
The existing algorithm modeling platform lacks complete model application, data collection, and continuous optimization, which leads to a disconnect between model training and application feedback, making it difficult to form an end-to-end complete design solution.
It provides a multi-level collaborative algorithm modeling method, which obtains the terminal's business data for preprocessing, allocates configuration information for scheduling, performs business training, performs algorithm modeling, and distributes the target model to the terminal to realize the algorithm closed loop from training to application.
The algorithm closed loop from training to application is realized, ensuring the effective application and continuous optimization of the model in the operating environment, and improving the adaptability and applicability of the model.
Smart Images

Figure CN119938288A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a multi-level collaborative algorithm modeling method, device, system, equipment and storage medium. Background Art
[0002] In the related technologies, common algorithm modeling platforms lack complete model application, data collection, and continuous optimization, and fail to effectively explain the closed-loop concept of algorithm modeling. There is no systematic planning and design for how to apply the trained model to the operating environment, how to effectively collect real-time data to supplement the training library during execution, and how to continuously optimize the model to meet the high adaptability requirements under environmental changes. It is difficult to form an end-to-end complete design solution, that is, the disconnection between model training and application feedback. There is currently no effective solution to this problem. Summary of the invention
[0003] To solve related technical problems, the embodiments of the present application provide a multi-level collaborative algorithm modeling method, apparatus, system, device and storage medium.
[0004] To achieve the above purpose, the technical solution of the embodiment of the present application is implemented as follows:
[0005] The present application provides a multi-level collaborative algorithm modeling method, the method comprising:
[0006] Acquiring service data of at least one terminal;
[0007] Preprocessing the business data to obtain training data and test data corresponding to the business data;
[0008] Perform scheduling processing according to the configuration information allocated to the training data and the test data to obtain scheduling information;
[0009] Performing business training on the training data and the test data based on the scheduling information to obtain target parameters of the training;
[0010] Algorithmic modeling is performed according to the target parameters to obtain a target model; the target model is used to distribute to the terminal.
[0011] In the above solution, the preprocessing of the business data to obtain training data and test data corresponding to the business data includes:
[0012] Classify the business data to obtain initial training data and initial test data corresponding to the business data;
[0013] In the case that the initial training data and the initial test data are abnormal, adding and / or deleting the initial training data and the initial test data to obtain the training data and the test data;
[0014] In the case that the initial training data and the initial test data have no abnormality, the initial training data is used as the training data and the initial test data is used as the test data.
[0015] In the above solution, the configuration information includes at least one of the following:
[0016] Resource information available for the test data;
[0017] Training parameter information corresponding to the test data;
[0018] The training target information corresponding to the test data.
[0019] In the above solution, the scheduling process is performed according to the configuration information allocated to the training data and the test data to obtain the scheduling information, including:
[0020] Obtaining a first parameter of each of the at least one model to be trained;
[0021] Performing a first preset algorithm processing on each of the models to be trained to obtain matching information between each of the models to be trained and the central cloud or the edge cloud;
[0022] The matching information, the first parameter and the configuration information are processed by a second preset algorithm to obtain the scheduling information.
[0023] In the above scheme, the first preset algorithm is performed on each of the models to be trained to obtain matching information between each of the models to be trained and the central cloud or the edge cloud, including:
[0024] Obtaining a priority parameter corresponding to each of the models to be trained and the central cloud or the edge cloud;
[0025] The first preset algorithm is performed on each of the training models according to the priority parameters to obtain the matching information.
[0026] In the above solution, performing service training on the training data and the test data based on the scheduling information to obtain target parameters of the training includes:
[0027] Performing service training on the training data and the test data based on the scheduling information to obtain initial parameters;
[0028] Performing business training evaluation on the initial parameters to obtain the target parameters.
[0029] In the above solution, the performing of business training evaluation on the initial parameters to obtain the target parameters includes:
[0030] Performing business training evaluation on the initial parameters according to a preset evaluation expert database to obtain an evaluation result;
[0031] The target parameter is determined based on the evaluation result.
[0032] The embodiment of the present application also provides a multi-level collaborative algorithm modeling system, which is applied to a distributed training application integrated architecture;
[0033] The multi-level collaborative algorithm modeling system is used to obtain business data of at least one terminal; pre-process the business data to obtain training data and test data corresponding to the business data; perform scheduling processing according to the configuration information allocated to the training data and the test data to obtain scheduling information; perform business training on the training data and the test data based on the scheduling information to obtain target parameters for training; perform algorithm modeling according to the target parameters to obtain a target model; and the target model is used to distribute to the terminal.
[0034] The embodiment of the present application also provides a multi-level collaborative algorithm modeling device, the device includes an acquisition unit, a preprocessing unit, a scheduling processing unit, a training unit and an algorithm modeling unit, wherein:
[0035] The acquisition unit is used to acquire service data of at least one terminal;
[0036] The preprocessing unit is used to preprocess the business data to obtain training data and test data corresponding to the business data;
[0037] The scheduling processing unit is used to perform scheduling processing according to the configuration information allocated to the training data and the test data to obtain scheduling information;
[0038] The training unit is used to perform service training on the training data and the test data based on the scheduling information to obtain target parameters for training;
[0039] The algorithm modeling unit is used to perform algorithm modeling according to the target parameters to obtain a target model; the target model is used to distribute to the terminal.
[0040] An embodiment of the present application further provides a storage medium having a computer program stored thereon; when the computer program is executed by a processor, the steps of any of the above methods are implemented.
[0041] An embodiment of the present application also provides a multi-level collaborative algorithm modeling device, which includes: a processor and a memory for storing a computer program that can be run on the processor, wherein the processor executes the steps of the above-mentioned method when running the computer program.
[0042] The embodiments of the present application provide a multi-level collaborative algorithm modeling method, device, equipment, system and storage medium, wherein the method includes: obtaining business data of at least one terminal; preprocessing the business data to obtain training data and test data corresponding to the business data; scheduling and processing according to the configuration information allocated to the training data and the test data to obtain scheduling information; performing business training on the training data and the test data based on the scheduling information to obtain the target parameters of the training; performing algorithm modeling according to the target parameters to obtain the target model; the target model is used to distribute to the terminal. The technical solution of the embodiments of the present application is adopted, by receiving the business data of the terminal, scheduling and processing the business data and performing business training, obtaining the target model, and distributing the target model to the terminal, that is, the model is trained and applied according to the data received from the terminal, realizing the algorithm closed loop from training to application. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] Figure 1 A flowchart of a multi-level collaborative algorithm modeling method provided in an embodiment of the present application;
[0044] Figure 2 A schematic diagram of a multi-level cloud collaborative training process provided in an embodiment of the present application;
[0045] Figure 3 A schematic diagram of the structure of a multi-level collaborative algorithm modeling system application provided by an embodiment of the present invention;
[0046] Figure 4 A schematic diagram of an edge-cloud collaborative distributed training framework provided in an embodiment of the present application;
[0047] Figure 5 A schematic diagram of a multi-level cloud collaborative execution process provided in an embodiment of the present application;
[0048] Figure 6 A schematic diagram of a multi-level collaborative algorithm modeling device provided in an embodiment of the present application;
[0049] Figure 7 A schematic diagram of the hardware entity structure of a multi-level collaborative algorithm modeling device in an embodiment of the present application. DETAILED DESCRIPTION
[0050] The present application is further described in detail below with reference to the accompanying drawings and specific embodiments.
[0051] At present, most of the related technologies of AI collaborative computing based on cloud computing have considered the impact of sample size and quantity, cloud computing network latency, and the computing resources that can be called by containers in cloud computing on modeling and training tasks, but have not effectively combined the modeling, training and evaluation optimization of AI algorithms with cloud network resources from the perspectives of data collection, model construction, training scheduling, adversarial simulation, intelligent evaluation, application feedback and modeling update. In addition, some cases are strongly bound to the industry (such as industrial big data, smart factories, etc.), and the versatility and cross-domain capabilities of their methods are not strong, mainly including the following situations:
[0052] 1. A deep learning model training acceleration method for end-edge-cloud collaboration was built, focusing on model segmentation and computing task allocation in the end-edge-cloud environment, but without considering application data feedback, continuous modeling optimization, adversarial models, and intelligent scheduling of training tasks for computing nodes.
[0053] 2. A product quality prediction model under the end-edge-cloud system is proposed. However, due to its focus on industrial scenarios, its cross-domain applicability and adaptability are limited. The solution does not involve research in related directions such as distributed training task scheduling and adversarial model design.
[0054] 3. A federal defense method for the security of the Artificial Intelligence of Things (AIoT) was proposed, and a distributed adversarial training model was designed. However, due to the focus on the AIoT security field and the lack of attention to the continuous improvement of model training, its application scenarios and the adaptability of the adversarial model were affected to a certain extent.
[0055] The AI modeling training platform needs to make continuous breakthroughs in the fields of intelligence, cross-domain, layering, and closed loop. Introduce adversarial models to intelligently judge training results; intelligently schedule distributed training tasks, and monitor and feedback training results in real time; implement wide-area cross-data center training application models, reasonably schedule cloud network resources, and establish a full-cycle algorithm modeling closed-loop platform from model training to application time, and then to feedback upgrades. For the above scenarios, the following problems need to be solved:
[0056] 1. The disconnect between model training and application feedback in related technologies: Currently common algorithm modeling platforms lack complete model application, data collection, and continuous optimization content, and have not effectively explained the closed-loop concept of algorithm modeling. There is no systematic planning and design for how to apply the trained model to the operating environment, how to effectively collect real-time data during execution to supplement the training library, and how to continuously optimize the model to meet the high adaptability requirements under environmental changes. It is difficult to form an end-to-end design plan.
[0057] 2. Model training in related technologies lacks continuously evolving adversarial simulation: Currently, common adversarial models mostly focus on special fields (such as security defense), lack universal supporting scenarios and continuous adversarial evolution (that is, there is a lack of algorithm modeling ideas, and the need to simultaneously consider sustainably evolving adversarial models or quality evaluation models when building models). This limits the level of unmanned and intelligent modeling training, and is not conducive to the rapid improvement of algorithm training efficiency and cross-domain promotion and practice.
[0058] 3. Model training in related technologies lacks a hierarchical task management engine: Currently, most common algorithm training models are single-layer centralized models, lacking a hierarchical training system and a plan and design for the effective coordination of cross-domain computing resources, which is not conducive to promoting the coordination of computing resources under wide-area conditions.
[0059] Based on this, the embodiment of the present application provides an integrated design method for artificial intelligence algorithm modeling, adversarial training and application feedback based on a multi-level cloud collaboration scenario to solve the integration problem of the distributed training system and the model application system, and proposes end-to-end closed-loop management measurement of artificial intelligence from training to application to optimization, so as to avoid the problem of insufficient continuous evolution and upgrading capabilities of artificial intelligence application systems, and the in-depth bundling of algorithm training and application models, which leads to increased power consumption of hardware terminals.
[0060] Figure 1 A flowchart of a multi-level collaborative algorithm modeling method provided in an embodiment of the present application; Figure 1 As shown, the method includes:
[0061] S101: Acquire service data of at least one terminal;
[0062] S102: Preprocess the business data to obtain training data and test data corresponding to the business data;
[0063] S103: Perform scheduling processing according to the configuration information allocated to the training data and the test data to obtain scheduling information;
[0064] S104: Performing service training on the training data and the test data based on the scheduling information to obtain target parameters of the training;
[0065] S105: Perform algorithm modeling according to the target parameters to obtain a target model; the target model is used to distribute to the terminal.
[0066] It should be noted that the multi-level collaboration in the embodiments of the present application can be understood as collaboration among the three layers of cloud, edge and end.
[0067] In S101, the acquisition of business data of at least one terminal can be understood as the edge cloud side collecting real-time data generated on the terminal, and the real-time data generates local situation data after the algorithm execution engine on the edge cloud side executes the algorithm business; the central cloud side collects the local situation data on the edge cloud side and the real-time data on the terminal. The business data can be understood as the real-time data generated by the terminal.
[0068] In S102, it should be noted that the training data can be understood as a training set; the test data can be understood as a test set. The proportion of the training data and the test data in the business data can be determined according to actual conditions and is not limited here. In practical applications, the training set can account for 70%; the test set can account for 30%.
[0069] The preprocessing of the business data to obtain training data and test data corresponding to the business data can be understood as: classifying the business data to obtain initial training data and initial test data corresponding to the business data; in the event that the initial training data and the initial test data are abnormal, adding and / or deleting the initial training data and the initial test data to obtain the training data and the test data; in the event that the initial training data and the initial test data are not abnormal, using the initial training data as the training data and using the initial test data as the test data.
[0070] In S103, the scheduling processing is performed according to the configuration information allocated according to the training data and the test data to obtain the scheduling information, which can be understood as obtaining the first parameter of each of the at least one model to be trained; performing the first preset algorithm processing on each of the model to be trained to obtain the matching information between each of the model to be trained and the central cloud or the edge cloud; performing the second preset algorithm processing on the matching information, the first parameter and the configuration information to obtain the scheduling information.
[0071] In S104, the business training is performed on the training data and the test data based on the scheduling information to obtain the target parameters of the training. This can be understood as performing business training on the training data and the test data based on the scheduling information to obtain initial parameters; and performing business training evaluation on the initial parameters to obtain the target parameters.
[0072] In S105, the algorithm modeling is performed according to the target parameters to obtain the target model; the target model is used to distribute to the terminal, which can be understood as the training engine generating the algorithm model data according to the target parameters, and then distributing the trained target model to the terminal. It should be noted that after obtaining the target model, the cloud platforms at all levels first store the model data in the local cloud.
[0073] In the embodiment of the present application, business data of the terminal is received and the business data is trained to obtain a target model, and the target model is distributed to the terminal, that is, the model is trained and the trained model is applied, thereby realizing an algorithm closed loop from training to application.
[0074] In one embodiment, the preprocessing of the business data to obtain training data and test data corresponding to the business data includes:
[0075] Classify the business data to obtain initial training data and initial test data corresponding to the business data;
[0076] In the case that the initial training data and the initial test data are abnormal, adding and / or deleting the initial training data and the initial test data to obtain the training data and the test data;
[0077] In the case that the initial training data and the initial test data have no abnormality, the initial training data is used as the training data and the initial test data is used as the test data.
[0078] In this embodiment, the method for classifying and processing the business data can be determined according to actual conditions and is not limited here.
[0079] In the case where anomalies occur in the initial training data and the initial test data, the initial training data and the initial test data are added and / or deleted to obtain the training data and the test data; it can be understood that in the case where the initial training data and the initial test data are missing or noise data exists, the initial training data and the initial test data are added and / or deleted to obtain the training data and the test data. In practical applications, in the case where the initial training data and the initial test data are missing, the tuple attributes used for training sample data can be supplemented. The method of supplementation can be determined according to actual conditions and is not limited here. As an example, Bayesian or decision tree can be used to fill in the most likely value; in the case where noise data exists in the initial training data and the initial test data, the noise data needs to be removed or generated.
[0080] In one embodiment, the configuration information includes at least one of the following:
[0081] Resource information available for the test data;
[0082] Training parameter information corresponding to the test data;
[0083] The training target information corresponding to the test data.
[0084] In this embodiment, the resource information can be understood as system resources distributed by a multi-level collaborative algorithm modeling system; the training parameter information can be understood as training parameters configured for the model to be trained; and the training target information can be understood as training targets configured for the model to be trained.
[0085] In one embodiment, the scheduling process is performed according to the configuration information allocated to the training data and the test data to obtain the scheduling information, including:
[0086] Obtaining a first parameter of each of the at least one model to be trained;
[0087] Performing a first preset algorithm processing on each of the models to be trained to obtain matching information between each of the models to be trained and the central cloud or the edge cloud;
[0088] The matching information, the first parameter and the configuration information are processed by a second preset algorithm to obtain the scheduling information.
[0089] In this embodiment, it should be noted that the first parameter can be understood as the task request corresponding to the model to be trained. In actual applications, the first parameter includes model complexity, training data volume, and resource conditions of the training server.
[0090] The first preset algorithm is performed on each of the models to be trained to obtain matching information between each of the models to be trained and the central cloud or the edge cloud; it can be understood as obtaining the corresponding priority parameters between each of the models to be trained and the central cloud or the edge cloud; and the first preset algorithm is performed on each of the training models according to the priority parameters to obtain the matching information.
[0091] The second preset algorithm may be determined according to actual conditions. As an example, the preset algorithm may be an adaptive scheduling algorithm based on a greedy idea.
[0092] The matching information, the first parameter and the configuration information are processed by a second preset algorithm to obtain the scheduling information, which can be illustrated as follows: according to the underlying perception information reported by each server, based on the "greedy" idea, each training task maximizes the allocation of server resources to achieve the purpose of completing the training as soon as possible, and adaptively adjusts the hyperparameters of the model training and the required number of GPUs.
[0093] In one embodiment, the performing of a first preset algorithm processing on each of the models to be trained to obtain matching information between each of the models to be trained and the central cloud or the edge cloud includes:
[0094] Obtaining a priority parameter corresponding to each of the models to be trained and the central cloud or the edge cloud;
[0095] The first preset algorithm is performed on each of the training models according to the priority parameters to obtain the matching information.
[0096] In this embodiment, the matching information is the matching information between the training model and the server; the first preset algorithm can be determined according to actual conditions. As an example, the first preset algorithm can be a fitness function of a genetic algorithm.
[0097] The obtaining of the corresponding priority parameters between each model to be trained and the central cloud or the edge cloud can be understood in practical applications as follows: the task sets between levels are allocated to the optimal training server in a cascading manner, and the level with the highest priority is matched first.
[0098] The first preset algorithm is performed on each of the training models according to the priority parameters to obtain the matching information. In practical applications, it can be understood that for each level of task set, a genetic algorithm is used to search for the optimal solution to obtain the optimal match between this batch of training tasks and the server.
[0099] For ease of understanding, the scheduling training process in actual application is described in detail here.
[0100] Artificial Intelligence (AI) model training not only requires processing massive amounts of data, but also requires a large amount of Graphic Processing Unit (GPU) resources for computing. Each edge server and cloud server is equipped with GPU computing resources. The types, numbers, and computing capabilities of GPU cards vary between servers, and the requirements for GPU resources for each AI training task are also different. How to reasonably schedule multiple AI training tasks and achieve high-performance AI training and computing directly affects the efficiency of AI production and R&D, and is crucial to the iteration efficiency of AI products.
[0101] This application proposes an adaptive model training scheduling method based on the greedy idea. For a batch of AI model training task requests, according to the model complexity, training deadline time, training data volume, and training server resource conditions, the method aims to maximize the allocation of server resources to complete the training as soon as possible. The method adaptively adjusts the model training hyperparameters and reasonably schedules the edge cloud servers to meet the multi-objective optimization requirements of minimizing the overall training delay and deadline violation rate under the constraints of graphics card resources.
[0102] Mathematical modeling of AI training task scheduling: Assume that there are n model training tasks W = {w1,w2,…,w n}, m servers E = {e1, e2, …, e m}. The video memory resources occupied by model training task i are shown in formula (1),
[0103] M=M model +M layer (1)
[0104] In formula (1), M is the video memory resource occupied by model training task i, M model The number of bytes occupied by the floating point number of the model parameter; M layer The number of bytes occupied by the floating-point numbers of the output layer parameters.
[0105] The overall delay of model training task i on server j is shown in formula (2):
[0106]
[0107] The formula for deadline violation rate is shown in formula (3):
[0108]
[0109] The objective optimization function to minimize the overall training delay and deadline violation rate is shown in formula (4):
[0110]
[0111] In formulas (2), (3) and (4), T i,j is the completion time of training task i running on server j, T Wait is the task waiting delay, T dt is the data transmission delay, T train is the model training delay, T mt The model distribution delay; for a complete AI training task, the task completion time includes the task waiting delay T Wait, data transmission delay T dt , model training delay T train and model distribution delay T mt ; M is the resource required by the currently running task server; For the current server k available resources; n' is the number of tasks that complete training within the deadline; ω is the weight of the latency indicator in the overall optimization goal, and its value can be determined according to actual conditions and is not limited here; T is the benchmark value of the training task time.
[0112] Among them, T Wait Related to the remaining resources of all current servers; T dt Related to the amount of training data and the communication bandwidth between the data feedback engine and the training server; T mt Related to the size of model parameters and the communication bandwidth between the training server and the application execution engine; T train It is related to the complexity of the network model, the amount of training data, the computing power of the GPU and central processing unit (CPU), and the training hyperparameters, as shown in formulas (5) and (6):
[0113]
[0114]
[0115] In formulas (5) and (6), c j The CPU resource capacity that server j can provide; g i is the GPU floating point computing capability of server j; f i r The CPU computing power required to perform input data preprocessing and data copying between memory and video memory during model training. In image processing models, the input image is preprocessed in the CPU, and this delay is a non-negligible part; i l is the time complexity of network model i, using the number of floating point operations FLOPs of the model i to measure; deep is the number of convolutional layers of the network; layer is the current convolutional layer; Fm layer K is the size of the feature map output by the current convolutional layer; layer is the convolution kernel size of the current layer; C layer-1 is the number of input channels of the current layer, that is, the number of output channels of the previous layer; C layeris the number of output channels of the current layer, that is, the number of convolution kernels; Data_set is the amount of training data; batch_size is the size of a batch of data in the training hyperparameters, and one batch_size completes one iteration; epoch i The number of generations that model i completes a full training of all data. All data are fully trained once, and one epoch includes Iterations; N is the number of GPU cards for multi-GPU training, and multiple GPU cards are trained in parallel; δ is the attenuation coefficient of multi-GPU training. The more GPUs are used, the more complicated the communication between devices is, especially in the batch normalization (BatchNormalization, BN) layer of the network model. If the BN layer is to be synchronized, the training time will be extended; this value is generally taken as 0.8.
[0116] M model is the number of bytes occupied by the floating point number of the model parameter, which can be expressed by formula (7) as follows:
[0117]
[0118] In formula (7), The number of model parameters, the number of bytes occupied by their floating point numbers, should be multiplied by 4. During the training process, it is necessary to save the model parameters, gradient parameters, momentum parameters and adaptive moment estimation (Adaptive moment estimation, Adam) optimization parameters at the same time, a total of 4 parameters, so the total number of bytes occupied should be multiplied by 4 again.
[0119] M layer is the number of bytes occupied by the floating point number of the output layer parameter, which can be expressed as formula (8):
[0120]
[0121] AI training task scheduling algorithm process:
[0122] Step 1: First, prioritize the models according to their application scenarios.
[0123] Each AI model has different application scenarios. Some models are used for global situation prediction, and the data in each local area has an impact on them. Such models are more suitable for training on cloud servers; some models are only related to local area data, and the amount of data is extremely large, and the cost of remote transmission is high, so they are more suitable for training on edge servers. We will first classify all models according to the application characteristics of the models: edge server training class, hybrid training class, and cloud server training class. The edge server training class has a higher priority, which often means that the model iteration needs to be faster, and the number of task requests for this class is the largest; the hybrid training class is second; the cloud server training class has the lowest priority and the number of task requests is the least.
[0124] Step 2: Based on the "greedy" idea, adaptively adjust the hyperparameters of model training and the number of GPU cards occupied.
[0125] Each server contains multiple GPU cards, which are connected by NVLINK topology. The edge server reports its own underlying topology perception information to the central cloud server, including GPU type, GPU computing power, communication bandwidth between GPUs, the number of currently idle GPU cards, and the memory capacity of each GPU card; each model to be trained provides an adjustable hyperparameter range.
[0126] According to the underlying perception information reported by each server, based on the "greedy" idea, each training task maximizes the allocation of server resources to achieve the goal of completing training as soon as possible, and adaptively adjusts the hyperparameters of model training and the number of GPUs required. The adjustment rules are as follows: According to the remaining video memory resources of the current server, within the available batch_size range (generally {16, 32, 64, 128, 256, 512}), select a suitable batch_size, calculate the video memory capacity required for model training according to the above formulas (1), (7) and (8), and obtain the required number of GPU cards N; according to the batch_size, use the linear scaling rule to adjust the learning rate accordingly to ensure the training accuracy of the model; according to the amount of data for this training Data_set, adjust the number of epochs of model training accordingly, so that the model achieves a good training loss without overfitting. Generally, the larger the amount of data, the larger the number of iterations required. According to the amount of training data, Data_set, batch_size, epoch and N value, the estimated training delay T is obtained according to formulas (2), (5) and (6): train .
[0127] Step 3: Perform hierarchical cascade matching between training tasks and servers.
[0128] Considering the high time cost of cross-domain communication between multiple servers, a training task can only be assigned to one server. A server can receive multiple training tasks within the resource limit. The training task allocation method is:
[0129] (1) According to the priority stratification in step 1, the task sets between the levels are allocated the optimal training server in a cascading manner, and the level with the highest priority is matched first;
[0130] (2) For each level of task set, use the genetic algorithm to search for the optimal solution to obtain the optimal match between this batch of training tasks and the server. The above formula (4) is used as the fitness function of the genetic algorithm;
[0131] (3) In each iterative search in the genetic algorithm, the underlying perception information of each server is updated according to the solution set of the previous generation. In the next crossover and mutation, Constraints of
[0132] (4) In the genetic algorithm, each combination crossover and mutation requires the encoding of the change to be re-adjusted according to step 2 to adaptively adjust the model parameters.
[0133] Step 4: For tasks in training, provide real-time feedback of the remaining completion time to the central cloud.
[0134] The central cloud obtains the remaining completion time of the tasks running on each server and serves as a reference for the waiting time of the next round of scheduled tasks.
[0135] In one embodiment, performing service training on the training data and the test data based on the scheduling information to obtain target parameters of the training includes:
[0136] Performing service training on the training data and the test data based on the scheduling information to obtain initial parameters;
[0137] Performing business training evaluation on the initial parameters to obtain the target parameters.
[0138] In this embodiment, the business training evaluation of the initial parameters to obtain the target parameters can be understood as performing business training evaluation on the initial parameters according to a preset evaluation expert database to obtain evaluation results; and determining the target parameters based on the evaluation results. The initial parameters are parameters obtained after business training.
[0139] In one embodiment, performing business training evaluation on the initial parameters to obtain the target parameters includes:
[0140] Performing business training evaluation on the initial parameters according to a preset evaluation expert database to obtain an evaluation result;
[0141] The target parameter is determined based on the evaluation result.
[0142] In this embodiment, the preset evaluation expert library is selected according to the actual situation and is not limited here. In practical applications, it may include accuracy (Accurcy), false negative rate / missed diagnosis rate (FNR), false positive rate / misdiagnosis rate (FPR), ROC curve area (AUC), average accuracy (AP), F1 score (F1-score), and root mean square error (RMSE).
[0143] Determining the target parameter based on the evaluation result can be understood as, when the evaluation result indicates that the initial parameter is better than the preset parameter, updating the preset parameter to the initial parameter, that is, using the initial parameter as the target parameter; when the evaluation result indicates that the initial parameter is not better than the preset parameter, no updating is performed.
[0144] For ease of understanding, an example from a practical application is given here.
[0145] Different network models are suitable for different evaluation methods. In order to more comprehensively evaluate the model performance, multiple dimensions and multiple indicators are often combined for comprehensive evaluation. Establish a model evaluation expert database, including but not limited to accuracy, false negative rate / missed diagnosis rate, false positive rate / misdiagnosis rate, ROC curve area, average accuracy, F1 score, root mean square error, and select one or more evaluation indicators according to the actual business model to evaluate the model. For example, for target detection models, AP is used as the evaluation indicator, and AP is also combined. @0.5 (AP value when IOU threshold is 0.5), AP @0.75 (AP value when IOU threshold is 0.75), AP s (AP value of small-sized targets), AP m (AP value for medium-sized targets), AP L (AP value of large-size targets) to assist in comprehensive evaluation; for lane models, the F1-score evaluation indicator is used; for classification models with an imbalance in the number of positive and negative samples, the AUC evaluation indicator is used; for situation prediction models, the RMSE evaluation indicator is used; in the model training engine, a certain proportion of the data is divided into test sets using manually annotated or actually collected data sets to evaluate the quality of the model.
[0146] In the model execution engine, actual operation data is collected, and model quality assessment is performed manually, automatically or semi-automatically according to the model type, and the model operation status is promptly fed back to the central cloud to generate a model update request. For detection business scenarios, user satisfaction feedback can be received, and the collected data can be manually annotated regularly to feedback quality assessment results; for situation prediction business, the model quality is automatically updated according to the actual situation detected, and when the quality is lower than the preset threshold, a model update request is automatically initiated to the central cloud.
[0147] In order to understand the embodiments of the present application, the embodiments of the present application illustrate a specific application scenario of a multi-level collaborative algorithm modeling method.
[0148] In the embodiment of the present application, the algorithm model training is carried out based on the collaborative division of labor of the cloud, edge and end layers:
[0149] Central cloud (first-level management node): Training on the central cloud is aimed at global situation prediction or statistical algorithms; the training management service of the central cloud uniformly manages and schedules the algorithm model, and plans, schedules, and dispatches training tasks that can be run on the central cloud and edge cloud according to the resource availability and environmental factors of the cloud computing environment; manages and controls the executing training tasks, supports their suspension, termination, and re-running after model parameter adjustment; the execution status of the training tasks on the central cloud is periodically (k seconds) reported to the training management service of the central cloud so that the training management service can monitor and schedule the training service; after the training task is completed, the model files and parameters in the training results are collected and stored in the model database, and it is decided whether to distribute them to the algorithm execution environment on the central cloud and edge cloud as needed, and whether the algorithm model needs to be updated is determined based on manual or preset thresholds, supporting the continuous optimization of the algorithm on the actual execution environment.
[0150] Regional / edge cloud or terminal side (secondary execution node): Training on regional / edge cloud or computing terminal is targeted at local areas and algorithms with strong real-time performance; the training scheduling model of the central cloud determines whether the regional / edge side or terminal side training tasks need to be executed. The relevant training tasks are uniformly managed by the central side, and the execution status is periodically (k seconds) reported to the training management service of the central cloud for the training management service to monitor and schedule the training service. After the training task on the regional / edge cloud is completed, the model files and parameters in the training results are collected and synchronized to the training model database of the central cloud for storage. The training management service of the central cloud decides whether to distribute it to the algorithm execution environment on the regional / edge cloud with the same business, and decides whether the algorithm model needs to be updated based on manual or preset thresholds, supporting the continuous optimization of the algorithm on the actual execution environment.
[0151] Training samples: The samples of training and test data come from regional / edge cloud collection data sets, local situation data sets gathered on the edge side of the central cloud, and generator networks in the Generative Adversarial Network (GAN) on the regional / edge / central cloud.
[0152] Figure 2 A schematic diagram of a multi-level cloud collaborative training process provided in an embodiment of the present application is shown in FIG. Figure 2 As shown, the training is divided into 8 steps:
[0153] (1) Collecting sample data: The cloud platforms at all levels running on the edge / region and core continuously run business systems. The edge application system collects real-time data generated on the terminal. The real-time data is sent to the algorithm training engine on the edge as the sample and test data set of the training engine for online training or offline training; the real-time data is sent to the algorithm execution engine on the edge to execute the local algorithm business on the edge and generate local situation data; the central cloud side deploys data collection and distribution tools on the edge to collect local situation data on the edge and real-time data (training samples, verification and test data sets are outside the edge data set and are used for global situation prediction) for subsequent online training or offline training on the central cloud.
[0154] (2) Data preprocessing: classify the data into training sets (accounting for 70%) and test sets (accounting for 30%); complete the tuple attributes used for training sample data (fill in the most likely values for Bayesian, decision trees, etc.); remove or generate noise data; detect deviations and data transformations.
[0155] (3) Determine the training task configuration: Determine the configuration for each training task, such as the allocatable system resources, training parameters, training objectives, etc.
[0156] (4) Scheduling training tasks: Determine the execution priority and target execution environment of the training task, and allocate the execution resources of the target environment to it; the training server schedules the resources and training timing of the training task according to the training priority.
[0157] (5) Execute training tasks: The edge cloud performs algorithm training on the training set obtained in the regional scope to obtain a local training model; the central cloud obtains the training set, test set and trained model on multiple edge clouds and performs global training on the core area. The algorithm training engines on the cloud platforms at all levels execute the training tasks according to the preconditions and complete the training tasks after the training objectives are achieved or manual intervention stops.
[0158] (6) Evaluate training quality: The training engine obtains the feature values of the training set and the test set, and integrates multiple evaluation models and evaluation indicators to obtain the quality evaluation results of the training model. The evaluation results are used to determine whether the algorithm training models on subsequent cloud platforms at all levels need to be updated and retrained. At the same time, the training model is sent to the algorithm execution engine on the edge cloud and the central cloud through the edge cloud collaboration component, and the execution engine decides whether it is necessary to replace the algorithm model or parameters used in the business system.
[0159] (7) Generate algorithm model: After the training task is completed, the training engine generates algorithm model data, and the cloud platforms at all levels first store the model data in the local cloud.
[0160] (8) Distributing algorithm models: The edge cloud aggregates the local algorithm models on the edge side and sends them to the central cloud through the edge cloud collaboration component to perform global training to avoid local data overfitting, or to perform global situation prediction training. The central cloud dynamically sends the global training model or global situation prediction to the target edge cloud based on the training status, algorithm execution status, dynamic allocation / idle status of cloud resources collected from the edge cloud, and the set training goals.
[0161] Based on the same inventive concept as above, the embodiment of the present invention also provides a multi-level collaborative algorithm modeling system. Figure 3 A schematic diagram of a multi-level collaborative algorithm modeling system application provided by an embodiment of the present invention is shown in FIG. Figure 3 As shown, the system is applied to the integrated architecture of distributed training applications;
[0162] The multi-level collaborative algorithm modeling system is used to obtain business data of at least one terminal; pre-process the business data to obtain training data and test data corresponding to the business data; perform business training based on the configuration information allocated to the training data and the test data to obtain target parameters for training; the business training includes scheduling business training and executing business training; perform algorithm modeling according to the target parameters to obtain a target model; and the target model is used to distribute to the terminal.
[0163] Figure 3 The core model consists of three parts:
[0164] 1. AI training management: As a unified management unit, it is generally located at the center and is responsible for the configuration management of algorithm models and training resources and the coordinated scheduling of training tasks. Based on model characteristics and training scenario requirements, this component manages the training plan and the progress of each branch task, synchronizes and integrates training results, distributes (phased) model results to the operating environment, and collects business feedback and application data, similar to a management node (manager_node).
[0165] 2. Distributed training engine: As a training unit, it cooperates with the management unit to build a distributed training environment, executes the training plan, keeps status synchronized with the management node and feeds back the training results; distributes the model to the running environment on demand, similar to a worker node (worker_node).
[0166] 3. AI operating environment execution engine: As an execution unit, it keeps in sync with the management unit, updates training results and loads them into the application system to support application requirements and business scenarios. It also collects operating status and quality feedback and provides them to the algorithm and adversarial model for optimization.
[0167] For ease of understanding, here is an example of a specific multi-level cloud collaborative training and execution. In the overall edge-cloud collaborative distributed training framework, the process can be explained from two aspects: how to start AI training and how to apply the model. Figure 4 A schematic diagram of an edge-cloud collaborative distributed training framework provided in an embodiment of the present application, such as Figure 4 As shown in the figure, this system is divided into the following five modules according to the responsibilities:
[0168] Algorithm Management Center (AI Task Management Unit): manages the currently included algorithm information, including version, field, score, etc., and also provides a visual interface or interface for users to trigger training tasks.
[0169] Algorithm scheduling engine (AI task management unit): responsible for pulling AI models from the algorithm model warehouse on demand and sending them to the training node for training. It is also responsible for collecting trained models and storing them in the algorithm model warehouse, and sending them to the required execution nodes.
[0170] Algorithm model warehouse (part of the algorithm model library in the AI management unit and part of it in distributed training): used to store various versions of algorithm models.
[0171] Training node (training engine in distributed training): a physical service node used to calculate and train AI models.
[0172] Execution node (execution engine in the application environment): a physical service node that uses the algorithm model.
[0173] Nodes need to register themselves to join the framework system. A node can be both a training and execution node. The properties of the node are determined by its autonomous selection and environmental computing resources when it is registered in this system.
[0174] Figure 5 A schematic diagram of a multi-level cloud collaborative execution process provided in an embodiment of the present application is shown in FIG. Figure 5 As shown in the figure, the multi-level cloud collaborative execution process can be divided into the following steps:
[0175] (1) Training task triggering: Training tasks can be divided into three categories: manual triggering, timed triggering, and rule triggering. Manual triggering refers to autonomously triggering the model, time, target, and other information of the specified training task. Timed triggering refers to periodically and autonomously issuing training tasks based on the configuration on the model. Rule triggering refers to defining the quality score target of the model. Before the model score target is reached, training is iterated continuously until the quality score is achieved.
[0176] (2) Algorithm scheduling engine loading: The algorithm scheduling engine pulls the static algorithm model from the algorithm model library according to the task content.
[0177] (3) The algorithm scheduling engine issues training tasks: It issues training tasks based on computing resources, node types, and service areas.
[0178] (4) The algorithm scheduling engine collects training results: After the training is completed, the training node actively requests to upload the training results. After receiving the training results, the scheduling engine scores them.
[0179] (5) The algorithm scheduling engine saves the training results: If the score is better than the current version of the same field and model, the model warehouse is called to save it.
[0180] (6) The algorithm scheduling engine sends the training results: Before sending, it is necessary to compare the model version and domain; then send the training results to the execution node based on the training task information.
[0181] The distributed architecture in the embodiment of the present application is a "vertical-horizontal" expansion mode. Vertically: Through the hierarchical training management architecture, the traditional centralized big data training resources are broken down into parts, computing resources are allocated on demand, distributed scheduling and collaborative training tasks are carried out, and training results are integrated to determine the results. This makes it possible for collaborative AI training in multiple locations to achieve "four crosses" (across Internet data centers (International Data Corporation, IDC), across data centers, across regions, and across fields). Horizontally: By building an "execution engine" to connect the training environment and the application environment, (phased) distribution of training results and collection of feedback information are achieved, thereby promoting the continuous rolling of the "training-application-feedback-optimization" model. Based on the "vertical-horizontal" model, an "integrated" architecture for AI distributed training applications is realized, a grid-based computing power collaboration and application feedback mechanism is constructed, cloud-number fusion capabilities are strengthened, resource efficiency is improved, and green sharing is promoted
[0182] In order to implement the method of the embodiment of the present application, the embodiment of the present application also provides a multi-level collaborative algorithm modeling device, Figure 6 A schematic diagram of a multi-level collaborative algorithm modeling device provided in an embodiment of the present application, such as Figure 6As shown, the device 600 includes an acquisition unit 601, a preprocessing unit 602, a scheduling processing unit 603, a training unit 604 and an algorithm modeling unit 605, wherein:
[0183] The acquisition unit 601 is used to acquire service data of at least one terminal;
[0184] The preprocessing unit 602 is used to preprocess the business data to obtain training data and test data corresponding to the business data;
[0185] The scheduling processing unit 603 is used to perform scheduling processing according to the configuration information allocated to the training data and the test data to obtain scheduling information;
[0186] The training unit 604 is used to perform service training on the training data and the test data based on the scheduling information to obtain target parameters for training;
[0187] The algorithm modeling unit 605 is used to perform algorithm modeling according to the target parameters to obtain a target model; the target model is used to distribute to the terminal.
[0188] In one embodiment, the preprocessing unit 602 is further used to classify the business data to obtain initial training data and initial test data corresponding to the business data; when an abnormality occurs in the initial training data and the initial test data, add and / or delete the initial training data and the initial test data to obtain the training data and the test data; when no abnormality occurs in the initial training data and the initial test data, use the initial training data as the training data and use the initial test data as the test data.
[0189] In one embodiment, the configuration information includes at least one of the following: resource information available for the test data; training parameter information corresponding to the test data; and training target information corresponding to the test data.
[0190] In one embodiment, the scheduling processing unit 603 is also used to obtain a first parameter of each of the at least one model to be trained; perform a first preset algorithm processing on each of the models to be trained to obtain matching information between each of the models to be trained and the central cloud or the edge cloud; perform a second preset algorithm processing on the matching information, the first parameter and the configuration information to obtain the scheduling information.
[0191] In one embodiment, the scheduling processing unit 603 is also used to obtain the priority parameters corresponding to each of the models to be trained and the central cloud or the edge cloud; and perform the first preset algorithm processing on each of the training models according to the priority parameters to obtain the matching information.
[0192] In one embodiment, the training unit 604 is further used to perform business training on the training data and the test data based on the scheduling information to obtain initial parameters; and perform business training evaluation on the initial parameters to obtain the target parameters.
[0193] In one embodiment, the training unit 604 is further used to perform business training evaluation on the initial parameters according to a preset evaluation expert library to obtain evaluation results; and determine the target parameters based on the evaluation results.
[0194] It should be noted that the multi-level collaborative algorithm modeling device provided in the embodiment of the present application and the multi-level collaborative algorithm modeling method provided in the aforementioned embodiment of the present application belong to the same inventive concept. The meaning of the terms appearing here have been explained in detail above and will not be repeated here.
[0195] An embodiment of the present application also provides a storage medium on which a computer program is stored. When the computer program processor is executed by a processor, the steps of the above-mentioned method embodiment are implemented. The aforementioned storage medium includes: a mobile storage device, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and other media that can store program codes.
[0196] An embodiment of the present application also provides a multi-level collaborative algorithm modeling device, which includes: a processor and a memory for storing a computer program that can be run on the processor, wherein when the processor is used to run the computer program, it executes the steps of the above-mentioned method embodiment stored in the memory.
[0197] Figure 7 A schematic diagram of a hardware entity structure of a multi-level collaborative algorithm modeling device in an embodiment of the present application is shown in FIG. Figure 7 As shown, the hardware entity of the multi-level collaborative algorithm modeling device 700 includes: a processor 701 and a memory 702 . Optionally, the multi-level collaborative algorithm modeling device 700 may also include a communication interface 703 .
[0198] It can be understood that the memory 702 can be a volatile memory or a non-volatile memory, and can also include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a magnetic random access memory (FRAM), a flash memory, a magnetic surface memory, an optical disk, or a compact disc read-only memory (CD-ROM); the magnetic surface memory can be a disk memory or a tape memory. The volatile memory can be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of RAM are available, such as static random access memory (SRAM), synchronous static random access memory (SSRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM, SyncLink Dynamic Random Access Memory), and direct RAM bus random access memory (DRRAM, Direct Rambus Random Access Memory).The memory 702 described in the embodiments of the present application is intended to include but is not limited to these and any other suitable types of memories.
[0199] The method disclosed in the above embodiment of the present application can be applied to the processor 701, or implemented by the processor 701. The processor 701 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the hardware integrated logic circuit in the processor 701 or the instruction in the form of software. The above processor 701 may be a general processor, a digital signal processor (DSP, Digital Signal Processor), or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc. The processor 701 can implement or execute the methods, steps and logic block diagrams disclosed in the embodiments of the present application. The general processor may be a microprocessor or any conventional processor, etc. In combination with the steps of the method disclosed in the embodiment of the present application, it can be directly embodied as a hardware decoding processor to execute, or it can be executed by a combination of hardware and software modules in the decoding processor. The software module can be located in a storage medium, which is located in the memory 702, and the processor 701 reads the information in the memory 702 and completes the steps of the above method in combination with its hardware.
[0200] In an exemplary embodiment, the device may be implemented by one or more application specific integrated circuits (ASIC), DSP, programmable logic device (PLD), complex programmable logic device (CPLD), field programmable gate array (FPGA), general processor, controller, microcontroller (MCU), microprocessor, or other electronic components to execute the aforementioned method.
[0201] It should be understood that "one embodiment" or "an embodiment" mentioned throughout the specification means that specific features, structures or characteristics related to the embodiment are included in at least one embodiment of the present application. Therefore, "in one embodiment" or "in an embodiment" appearing throughout the specification does not necessarily refer to the same embodiment. In addition, these specific features, structures or characteristics can be combined in one or more embodiments in any suitable manner. It should be understood that in various embodiments of the present application, the size of the sequence number of the above-mentioned processes does not mean the order of execution, and the execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application. The above-mentioned sequence numbers of the embodiments of the present application are only for description and do not represent the advantages and disadvantages of the embodiments.
[0202] It should be noted that, in this article, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the sentence "comprises a ..." does not exclude the existence of other identical elements in the process, method, article or device including the element.
[0203] The methods disclosed in several method embodiments provided in this application can be arbitrarily combined without conflict to obtain new method embodiments.
[0204] The features disclosed in several product embodiments provided in this application can be arbitrarily combined without conflict to obtain new product embodiments.
[0205] The features disclosed in several method or device embodiments provided in this application can be arbitrarily combined without conflict to obtain new method embodiments or device embodiments.
[0206] The above is only an implementation method of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art who is familiar with the present technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.
Claims
1. A multi-level collaborative algorithm modeling method, characterized in that: The method comprises: Acquiring service data of at least one terminal; Preprocessing the business data to obtain training data and test data corresponding to the business data; Perform scheduling processing according to the configuration information allocated to the training data and the test data to obtain scheduling information; Performing business training on the training data and the test data based on the scheduling information to obtain target parameters of the training; Algorithmic modeling is performed according to the target parameters to obtain a target model; the target model is used to distribute to the terminal.
2. The method according to claim 1, characterized in that The preprocessing of the business data to obtain training data and test data corresponding to the business data includes: Classify the business data to obtain initial training data and initial test data corresponding to the business data; In the case that the initial training data and the initial test data are abnormal, adding and / or deleting the initial training data and the initial test data to obtain the training data and the test data; In the case that the initial training data and the initial test data have no abnormality, the initial training data is used as the training data and the initial test data is used as the test data.
3. The method according to claim 1, characterized in that The configuration information includes at least one of the following: Resource information available for the test data; Training parameter information corresponding to the test data; The training target information corresponding to the test data.
4. The method according to claim 1, characterized in that: The step of performing scheduling processing according to the configuration information allocated to the training data and the test data to obtain scheduling information includes: Obtaining a first parameter of each of the at least one model to be trained; Performing a first preset algorithm processing on each of the models to be trained to obtain matching information between each of the models to be trained and the central cloud or the edge cloud; The matching information, the first parameter and the configuration information are processed by a second preset algorithm to obtain the scheduling information.
5. The method according to claim 4, characterized in that The performing of a first preset algorithm processing on each of the models to be trained to obtain matching information between each of the models to be trained and the central cloud or the edge cloud includes: Obtaining a priority parameter corresponding to each of the models to be trained and the central cloud or the edge cloud; The first preset algorithm is performed on each of the training models according to the priority parameters to obtain the matching information.
6. The method according to claim 1, characterized in that The performing service training on the training data and the test data based on the scheduling information to obtain target parameters of the training includes: Performing service training on the training data and the test data based on the scheduling information to obtain initial parameters; Performing business training evaluation on the initial parameters to obtain the target parameters.
7. The method according to claim 6, characterized in that The performing business training evaluation on the initial parameters to obtain the target parameters includes: Performing business training evaluation on the initial parameters according to a preset evaluation expert database to obtain an evaluation result; The target parameter is determined based on the evaluation result.
8. A multi-level collaborative algorithm modeling system, characterized in that: The system is applied to the integrated architecture of distributed training applications; The multi-level collaborative algorithm modeling system is used to obtain business data of at least one terminal; pre-process the business data to obtain training data and test data corresponding to the business data; perform scheduling processing according to the configuration information allocated to the training data and the test data to obtain scheduling information; Performing business training on the training data and the test data based on the scheduling information to obtain target parameters of the training; Perform algorithm modeling according to the target parameters to obtain a target model; The target model is used for distribution to the terminal.
9. A multi-level collaborative algorithm modeling device, characterized in that: The device includes an acquisition unit, a preprocessing unit, a scheduling processing unit, a training unit and an algorithm modeling unit, wherein: The acquisition unit is used to acquire service data of at least one terminal; The preprocessing unit is used to preprocess the business data to obtain training data and test data corresponding to the business data; The scheduling processing unit is used to perform scheduling processing according to the configuration information allocated to the training data and the test data to obtain scheduling information; The training unit is used to perform service training on the training data and the test data based on the scheduling information to obtain target parameters for training; The algorithm modeling unit is used to perform algorithm modeling according to the target parameters to obtain a target model; the target model is used to distribute to the terminal.
10. A storage medium, characterized in that: The storage medium stores a computer program; when the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
11. A multi-level collaborative algorithm modeling device, characterized in that: The multi-level collaborative algorithm modeling device includes: a processor and a memory for storing a computer program that can be run on the processor, wherein the processor executes the steps of the method described in any one of claims 1 to 7 when running the computer program.