Artificial intelligence network optimization training system and method based on deep learning

Through the multi-objective optimization module and enhanced training module, the problems of single training mode, insufficient edge device support and insufficient model robustness in the existing technology are solved, and efficient and robust deep learning training is achieved.

CN120763618APending Publication Date: 2025-10-10THE 44TH INST OF CHINA ELECTRONICS TECH GROUP CORP

Patent Information

Application Number
CN202510909251.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-02
Publication Date
2025-10-10

AI Technical Summary

Technical Problem

Existing deep learning-based AI network training methods are unable to set appropriate training modes based on the AI ​​network's own parameters, ignore training time and energy consumption, lack distributed training support for edge devices, and have insufficient model robustness, resulting in poor performance when facing adversarial samples.

Method used

A multi-objective optimization module is used to dynamically adjust the training mode, support distributed training on edge devices, optimize model hyperparameters through parameter monitoring and multi-objective optimization algorithms, and use an enhanced training module to inject adversarial samples to enhance model robustness.

Benefits of technology

It achieves multi-objective optimization, improves training accuracy and efficiency, expands the scope of application, and enhances the robustness of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120763618A_ABST
    Figure CN120763618A_ABST
Patent Text Reader

Abstract

The invention discloses an artificial intelligence network optimization training system and method based on deep learning, and the system comprises a request analysis module which is used for determining a training task type and an initial model hyper-parameter, and obtaining an available edge device list; the parameter monitoring module is used for determining a model index monitoring threshold value, monitoring model indexes in the training process in real time and adjusting corresponding model hyper-parameters; the multi-target optimization module is used for initializing target weights and calculating a multi-target optimal solution set; the self-adaptive training module is used for generating a training strategy and optimizing and adjusting the training strategy; and the cooperative training module is used for allocating training tasks to the available edge devices and feeding back model indexes in the training process to the parameter detection module. According to the invention, the training mode is dynamically adjusted according to the specific parameters of the artificial intelligence network, multi-objective optimization is realized, distributed training of edge devices is supported, and the robustness of the model is improved, so that the training precision and efficiency are improved, and the application range is expanded.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of artificial intelligence, in particular to an artificial intelligence network optimization training system and method based on deep learning. BACKGROUND

[0002] Deep learning is the learning of the internal laws and representation levels of sample data. The information obtained in these learning processes is very helpful for the interpretation of data such as text, images and sound. The ultimate goal is to enable machines to have analysis and learning capabilities like humans, and to be able to recognize text, images and sound data. The training method based on deep learning in the prior art often sequentially performs simple training, general training and progressive training mode of senior training, and cannot set a training mode that matches and is suitable for the artificial intelligence network according to its own parameters. The patent for invention with publication number CN116415687A discloses an artificial intelligence network optimization training system and method based on deep learning. The patent initiates a required training request through a request initiation module, receives the initiated training request and collects the corresponding parameters of the artificial intelligence network through a parameter collection module, selects a training model matched with the corresponding parameters of the artificial intelligence network and the initiated training request from the deep learning training model according to the received corresponding parameters of the artificial intelligence network and the initiated training request through a model matching module; determines a training module and parameters related to the training module through a training determination module; receives the parameters related to the training module and the training results through a model hyperparameter collection module, and optimizes the training determination module and the deep learning training model according to the parameters related to the training module and the training results.

[0003] The above optimization training method has the following disadvantages:

[0004] 1. The optimization training system only focuses on a single target, such as the accuracy of the model, while ignoring other important factors such as training time and energy consumption;

[0005] 2. With the development of edge computing technology, more and more artificial intelligence applications need to be trained on edge devices. However, the optimization training system lacks effective support for edge device distributed training scenarios, cannot fully utilize the computing resources of edge devices, and also has difficulty in solving the problems of data transmission and collaborative training between edge devices;

[0006] 3. In actual applications, artificial intelligence networks may face various complex environments and attacks, and therefore need to have strong robustness. However, the optimization training method does not fully consider the robustness of the model, resulting in poor performance of the model when facing adversarial samples. SUMMARY

[0007] In view of the deficiencies of the prior art, the technical problems to be solved by the present application are to provide an artificial intelligence network optimization training system and method based on deep learning, which can dynamically adjust the training mode according to the specific parameters of the artificial intelligence network, realize multi-objective optimization, support distributed training of edge devices, and improve the robustness of the model, thereby improving the training accuracy and efficiency and expanding the application range.

[0008] One of the technical solutions adopted by the present application is to provide an artificial intelligence network optimization training system based on deep learning, which comprises:

[0009] A request analysis module is configured to receive and analyze training task requests, determine training task types and initial model hyperparameters, and obtain a list of available edge devices.

[0010] A parameter monitoring module is configured to determine model index monitoring thresholds according to the training task types, and to monitor model indexes in real time during the training process and adjust corresponding model hyperparameters when the model indexes reach preset thresholds.

[0011] A multi-objective optimization module is configured to initialize target weights according to the training task types, the list of available edge devices, and the model hyperparameters, and to calculate a multi-objective optimal solution set using a multi-objective optimization algorithm.

[0012] An adaptive training module is configured to generate training strategies according to the initial model hyperparameters and the multi-objective optimal solution set, and to optimize and adjust the training strategies according to model indexes in the training process.

[0013] A collaborative training module is configured to assign training tasks to available edge devices according to the training strategies, and to feed back model indexes in the training process to the parameter detection module.

[0014] Further, the system further comprises:

[0015] An enhanced training module is configured to generate adversarial samples, inject the adversarial samples into a training set to generate an enhanced training set, perform adversarial enhanced training on the model using the enhanced training set, calculate the adversarial accuracy of the model, and feed back the adversarial accuracy of the model to the parameter detection module.

[0016] Further, the request analysis module comprises:

[0017] A task request sub-module is configured to receive and analyze training task requests, determine training task types and initial model hyperparameters, and obtain a list of available edge devices; wherein the training task types include classification, detection, and segmentation; and the initial model hyperparameters include learning rate and iteration number.

[0018] The strategy reference submodule is configured to call a historical training task similar to the training task type and the initial model hyperparameter from a historical training library according to the training task type and the initial model hyperparameter, and take a historical training strategy corresponding to the historical training task as a reference strategy of the parameter detection module.

[0019] Further, the parameter detection module comprises:

[0020] The monitoring threshold determination submodule is configured to set a model index monitoring threshold interval according to a model index history record in the historical training task.

[0021] The model index monitoring submodule is configured to monitor a real-time value of the model index in the training process.

[0022] The hierarchical response submodule is configured to compare the real-time value of the model index in the training process with the model index monitoring threshold interval, record a log if the real-time value of the model index exceeds a first interval, adjust a model hyperparameter corresponding to the model index if the real-time value of the model index exceeds a second interval, and stop the training if the real-time value of the model index exceeds a third interval.

[0023] Further, the multi-objective optimization module comprises:

[0024] The target weight configuration submodule is configured to set a training target according to the training task type and the initial model hyperparameter, and set a weight for each training target according to user demand, wherein the training target comprises a model accuracy, a training time and an energy consumption.

[0025] The optimization calculation submodule is configured to calculate a multi-objective optimal solution set according to the training task type, an available edge device list and the model hyperparameter, and utilize an NSGA2 multi-objective optimization algorithm, wherein the multi-objective optimal solution set comprises an edge device coordination parameter, a resource allocation parameter, a model structure parameter and a model hyperparameter.

[0026] Further, the adaptive training module comprises:

[0027] The training generation submodule is configured to generate a boost training unit and a downshift training unit according to a training strategy, and determine a number of the boost training unit and the downshift training unit, and a difference value of adjacent boost training units and a difference value of adjacent downshift training units.

[0028] The dynamic training submodule is configured to adjust the number of the boost training unit and the downshift training unit according to the model index.

[0029] The training termination submodule is configured to determine a model performance according to the model index, and terminate the model training when a model performance improvement amplitude of three consecutive boost training units or downshift training units is less than a preset threshold.

[0030] Further, the collaborative training module comprises:

[0031] An edge device submodule is configured to generate a global training model according to a training strategy, distribute the global training model to available edge devices according to a list of available edge devices, and perform individual training configuration for each available edge device;

[0032] A federated learning submodule is configured to aggregate local model hyperparameters to a central server after the edge devices complete local training, and aggregate the local model hyperparameters according to a federated learning algorithm to obtain model hyperparameters;

[0033] A load balancing submodule is configured to monitor the load of each edge device, and automatically distribute training tasks according to the load of the edge devices.

[0034] Further, the enhanced training module comprises:

[0035] A sample generation submodule is configured to generate adversarial samples according to an adversarial sample generation algorithm, inject the adversarial samples into a local training set in an edge device to obtain an enhanced training set, and perform model enhancement training on the corresponding edge device according to the enhanced training set;

[0036] An adversarial evaluation submodule is configured to monitor the model adversarial accuracy of each edge device, and feed back the model adversarial accuracy to the parameter detection module.

[0037] Another technical solution adopted by the present application is an artificial intelligence network optimization training method based on deep learning, which comprises the following steps:

[0038] S1: receiving and analyzing a training task request, determining a training task type and initial model hyperparameters, and obtaining a list of available edge devices;

[0039] S2: determining a model index monitoring threshold according to the training task type, and monitoring the model index in real time during the training process, and adjusting the corresponding model hyperparameters when the model index reaches a preset threshold;

[0040] S3: initializing target weights according to the training task type, the list of available edge devices, and the model hyperparameters, and calculating a multi-objective optimal solution set using a multi-objective optimization algorithm;

[0041] S4: generating a training strategy according to the initial model hyperparameters and the multi-objective optimal solution set, and optimizing and adjusting the training strategy according to the model index during the training process;

[0042] S5: distributing training tasks to available edge devices according to the training strategy, and feeding back the model index during the training process to step S2.

[0043] Further, the method further comprises:

[0044] S6: generate an adversarial sample, inject the adversarial sample into the training set, generate an enhanced training set, use the enhanced training set to perform adversarial enhancement training on the model, calculate the model adversarial accuracy, and feed back the model adversarial accuracy to step S2.

[0045] The deep learning-based artificial intelligence network optimization training system and method have at least the following beneficial effects: 1. The multi-objective optimization module focuses on the model accuracy, training time, and energy consumption in model training, so that model training no longer focuses on only the single target of model accuracy, and the multi-objective optimization algorithm is used to make the model training more in line with the actual application requirements; 2. The collaborative training module supports distributed training on edge devices, fully utilizes the computing resources of edge devices, and also solves the problems of data transmission and collaborative training between edge devices; 3. The enhanced training module injects adversarial samples into the training set, thereby enhancing the robustness of the model. BRIEF DESCRIPTION OF DRAWINGS

[0046] The accompanying drawings, which are included to provide a further understanding of the application and are incorporated in and constitute a part of this application, illustrate embodiments of the application and serve to explain the application without, however, to limit it. In the drawings:

[0047] Figure 1 The structural block diagram of an embodiment of the deep learning-based artificial intelligence network optimization training system of the application.

[0048] Figure 2 The structural block diagram of the request analysis module of the application.

[0049] Figure 3 The structural block diagram of the parameter detection module of the application.

[0050] Figure 4 The structural block diagram of the multi-objective optimization module of the application.

[0051] Figure 5 The structural block diagram of the adaptive training module of the application.

[0052] Figure 6 The structural block diagram of the collaborative training module of the application.

[0053] Figure 7 The structural block diagram of the enhanced training module of the application.

[0054] Figure 8 The flowchart of an embodiment of the deep learning-based artificial intelligence network optimization training method of the application. DETAILED DESCRIPTION

[0055] The application will be further described below with reference to the accompanying drawings.

[0056] Please refer to Figure 1 , the structural block diagram of an embodiment of the artificial intelligence network optimization training system based on deep learning of the present application. The system comprises:

[0057] The request analysis module 100 is used for receiving and analyzing the training task request, determining the training task type and the initial model hyperparameters and obtaining the list of available edge devices. The request analysis module 100 is connected with the artificial intelligence network, can receive the training task request initiated by the artificial intelligence network, and analyzes the request, determines the type of the training task, such as classification, detection, segmentation, etc., and can determine the initial model hyperparameters of the training model according to the determined training task type, including the model structure parameters that cannot be dynamically adjusted, such as the model network architecture, the optimizer type and the initialization method, etc., and also includes the model hyperparameters that can be dynamically adjusted, such as the learning rate, the batch size, the iteration number and the regularization parameter, etc. Since the scheme can also use edge devices for collaborative training, the list of available edge devices will also be determined when the training task request is analyzed, which is convenient for subsequent training task allocation.

[0058] Please refer to Figure 2 , the structural block diagram of the request analysis module 100 of the present application. In some embodiments, the request analysis module 100 comprises:

[0059] The task request submodule 110 is used for receiving and analyzing the training task request, determining the training task type and the initial model hyperparameters and obtaining the list of available edge devices; wherein the training task type includes classification, detection and segmentation; the initial model hyperparameters include the learning rate and the iteration number; the above training task request is initiated by the artificial intelligence network, and the task request submodule 110 can disassemble the features of the training task request according to the data structure thereof, so as to obtain the specific training task type, and then set the general model hyperparameters according to the training task type, and use them as the initial model hyperparameters of the initial training model.

[0060] The strategy reference submodule 120 is configured to call a historical training task similar to the training task type and the initial model hyperparameter from the historical training library, and take the historical training strategy corresponding to the historical training task as the reference strategy of the parameter detection module. In order to improve the analysis efficiency of the training task request and the accuracy of the initial model hyperparameter, the request analysis module 100 further comprises a strategy reference submodule 120, which is configured to call a historical training task similar to the training task type and the initial model hyperparameter from the historical training library, and select a suitable historical training task by using a similarity algorithm, so as to refer to the training strategy of the historical training task.

[0061] The parameter monitoring module 200 is configured to determine a model index monitoring threshold according to the training task type, monitor the model index in real time during the training process, and adjust the corresponding model hyperparameter when the model index reaches the preset threshold. The parameter monitoring module 200 can monitor various model indexes in real time during the training process. The model indexes can include memory occupation, gradient norm, activation value distribution and other parameters that can directly reflect the model state. When the monitored parameters reach the preset threshold, such as when the memory occupation exceeds 80%, the dynamic module adjustment mechanism is triggered to automatically adjust the corresponding model hyperparameter, such as adjusting the batch size, so as to ensure the stability and efficiency of the training.

[0062] Please refer to Figure 3 That is, the parameter detection module 200 comprises:

[0063] The monitoring threshold determination submodule 210 is configured to set a model index monitoring threshold interval according to the model index historical record of the historical training task. The model index threshold monitored by the monitoring threshold determination submodule 210 can be determined according to the model index record data of the historical training task, such as the historical training data of the same type task, the key inflection point data in the training process and the final model index of the successful model. The purpose of interval division of the model index monitoring threshold is to take different dynamic responses according to the interval of the model index when monitoring the model index subsequently.

[0064] The model index monitoring submodule 220 is configured to monitor the real-time value of the model index to be monitored in the training process. The model index monitoring submodule 220 monitors the model index by using a software tool or an algorithm. For example, for the memory occupation, a memory integration tool can be used to directly monitor the memory occupation. For example, for the gradient norm, an algorithm can be used to calculate the gradient norm in the PyTorch framework.

[0065] The hierarchical response submodule 230 is used to compare the real-time value of the model indicator during the training process with the model indicator monitoring threshold interval. If the real-time value of the model indicator exceeds the first interval, a log is recorded; if the real-time value of the model indicator exceeds the second interval, the model hyperparameter corresponding to the model indicator is adjusted; if the real-time value of the model indicator exceeds the third interval, training is stopped. After confirming the model indicator monitoring threshold, this hierarchical response submodule 230 compares the real-time monitored model indicator with the model indicator monitoring threshold interval, and can adjust the model hyperparameter according to the interval of the real-time monitored model indicator. For example, if the first interval of video memory occupancy is 70%-80%, and the second interval is 80%-90%, when the monitored video memory occupancy exceeds 70%, a log is recorded. If it exceeds 80%, the dynamic module adjustment mechanism is triggered to automatically adjust the parameters of the training module to ensure the stability and efficiency of training. If it exceeds 90%, training is stopped.

[0066] The multi-objective optimization module 300 is used to initialize the target weights according to the training task type, the list of available edge devices and the model hyperparameters, and calculate the multi-objective optimal solution set using the multi-objective optimization algorithm. This multi-objective optimization module 300 integrates the NSGA2 multi-objective algorithm, which can generate a multi-objective optimization solution set, and select the multi-objective optimal solution set from the solution set according to the user-defined target weights, and output the corresponding training parameters. Among them, the model training objectives can be selected according to the actual needs of the user. As an example, the model training objectives of this solution may include three aspects: model accuracy, training time and energy consumption, and their weight configurations can be 0.6, 0.2, and 0.2, that is, the three aspects of model accuracy, training efficiency and economy are comprehensively considered. The multi-objective optimal solution set may include model hyperparameters, which provide a basis for the subsequent generation of training strategies.

[0067] See also Figure 4 , i.e., a structural block diagram of the multi-objective optimization module 300 of the present invention, the multi-objective optimization module 300 includes:

[0068] The target weight configuration submodule 310 is used to set the training target according to the training task type and the initial model hyperparameters, and set the weight for each training target according to user needs; wherein, the training targets include model accuracy, training time and energy consumption; the training task type actually determines the training target, so this target weight configuration submodule 310 can obtain the training task type and the initial model hyperparameters, and set the training target according to the training task type and the initial model hyperparameters. At the same time, the user can also set the weight of each training target according to actual needs, such as setting the weights for the three training targets of model accuracy, training time and energy consumption to: 0.6, 0.2, 0.2.

[0069] The optimization calculation sub-module 320 is configured to calculate a multi-objective optimal solution set according to a training task type, a list of available edge devices, and model hyperparameters, and by using an NSGA2 multi-objective optimization algorithm, the multi-objective optimal solution set including edge device cooperation parameters, resource allocation parameters, model structure parameters, and model hyperparameters, and the like. The above variables are input into the NSGA2 multi-objective optimization algorithm, and a target function is constructed by maximizing model accuracy, minimizing training time, and minimizing energy consumption as targets, and by using the steps of initialization of a population, non-dominated sorting, calculation of crowding degree, selection operation, crossover and mutation operation, elite reservation, and termination condition checking of the NSGA2 multi-objective optimization algorithm, a Pareto optimal solution set (which is a set of optimal solutions of variables under different model accuracy, training time, and energy consumption) is obtained. The above is a conventional step of the NSGA2 multi-objective optimization algorithm, and the construction of the target function can also be reasonably set according to actual model variables and training targets, which need not be repeated here.

[0070] The adaptive training module 330 is configured to generate a training strategy according to initial model hyperparameters and the multi-objective optimal solution set, and to optimize and adjust the training strategy according to model indicators in a training process. The adaptive training module 330 can generate a training strategy according to the multi-objective optimal solution set, and the training strategy includes edge device cooperation parameters, resource allocation parameters, model structure parameters, and model hyperparameters, which are set for the model on the basis of the training target. At the same time, the adaptive training module 330 can also obtain real-time model indicators from the parameter detection module 200, and adaptively optimize and adjust the training strategy according to the real-time model indicators.

[0071] Please refer to Figure 5 That is, the structure block diagram of the adaptive training module 400 of the present application, the adaptive training module 400 comprising:

[0072] The training generation submodule 410 is configured to generate the promotion training unit and the downshift training unit according to a training strategy, and determine the number of the promotion training unit and the downshift training unit, and the difference between adjacent promotion training units and the difference between adjacent downshift training units. In the present scheme, the training efficiency and the model accuracy are improved by using a plurality of promotion training units and downshift training units. A plurality of continuously connected promotion training units constitute a promotion training part of the model, which mainly functions in that if the model reaches the performance requirement after being trained by the initial training unit, the model enters the first promotion training unit and the training difficulty is increased until the termination condition is reached. Similarly, a plurality of continuously connected downshift training units constitute a downshift training part of the model, which mainly functions in that if the model does not reach the performance requirement after being trained by the initial training unit, the model enters the first downshift training unit and the training difficulty is reduced. If the model still does not reach the performance requirement after being trained by the first downshift training unit, the model enters the next downshift training unit and the training difficulty is further reduced until the termination condition is reached. The number of the promotion training unit and the downshift training unit can be determined according to the model structure and resource allocation in the training strategy, and the difference between adjacent promotion training units or downshift training units can be preset. For example, the difference between adjacent promotion training units can be set in an exponentially decreasing form (for example, the difference between the first promotion training unit and the second promotion training unit is 10%, the difference between the second promotion training unit and the third promotion training unit is 5%, and the difference between the third promotion training unit and the fourth promotion training unit is 2.5%). Thus, the range of the promotion training gradually decreases as the training proceeds. The exponentially decreasing difference can improve the training efficiency while ensuring the training accuracy. The difference between adjacent downshift training units can be set as a fixed value (for example, the absolute value of the difference between adjacent downshift training units is fixed as the initial value, that is, the difference between the first downshift training unit and the initial training unit is 1 / 3). Thus, when the training result does not reach the performance requirement, the training difficulty is stably reduced, and the excessively low training efficiency caused by excessive downshift is avoided.

[0073] The dynamic training submodule 420 is configured to adjust the number of the promotion training unit and the downshift training unit according to the model index. The dynamic training submodule 420 can dynamically generate the number of the promotion training unit and the downshift training unit according to the loss value, the accuracy, the gradient distribution and the like in the model index, so as to improve the matching degree of the training.

[0074] The training termination submodule 430 is configured to determine the model performance according to the model indicators, and terminate the model training when the model performance improvement range of the continuous three boosting training units or downshift training units is less than a preset threshold. The training termination condition of the boosting training unit and the downshift training unit is set in the training termination submodule 430, which is to terminate the model training when the model performance improvement range of the continuous three boosting training units or downshift training units is less than a preset threshold, and the preset threshold can be set to 0.4%-1%, so as to avoid unnecessary training and save computing resources and time.

[0075] The collaborative training module 500 is configured to allocate training tasks to available edge devices according to the training strategy, and feed back the model indicators in the training process to the parameter detection module. The collaborative training module 500 supports distributed training of edge devices, fully utilizes the computing resources of edge devices, and solves the problems of data transmission and collaborative training among edge devices. The distributed training architecture can include a center server responsible for storing a global model and a training strategy library, and collecting and processing the data uploaded by the edge devices. The edge devices are responsible for executing local training tasks and uploading intermediate training results to the center server. It should be noted that the edge devices can be various devices with computing capabilities, such as intelligent sensors and edge computing gateways.

[0076] Please refer to Figure 6 That is, the structure block diagram of the collaborative training module 500 of the present application, the collaborative training module 500 comprises:

[0077] The edge device submodule 510 is configured to generate a global training model according to the training strategy, distribute the global training model to available edge devices according to the list of available edge devices, and perform individual training configuration for each available edge device.

[0078] The federated learning submodule 520 is configured to aggregate the local model hyperparameters to the center server after the edge devices complete the local training, and aggregate the local model hyperparameters according to the federated learning algorithm to obtain the model hyperparameters.

[0079] The load balancing submodule 530 is configured to monitor the load of each edge device, and automatically allocate training tasks according to the load of the edge devices. The load balancing submodule 530 will automatically allocate training tasks according to the computing power of the edge devices, such as GPU / CPU utilization. When the computing power of a certain edge device is high, more training tasks will be allocated; otherwise, fewer tasks will be allocated to achieve efficient operation of the entire training process.

[0080] In order to improve the robustness of the model, the system can further comprise:

[0081] The enhanced training module 600 is used to generate adversarial samples, inject the adversarial samples into the training set, generate an enhanced training set, use the enhanced training set to perform adversarial enhancement training on the model, calculate the model adversarial accuracy, and feed the model adversarial accuracy back to the parameter detection module.

[0082] See also Figure 7 , i.e., a structural block diagram of the enhanced training module 600 of the present invention, wherein the enhanced training module 600 includes:

[0083] The sample generation submodule 610 is used to generate adversarial samples based on the adversarial sample generation algorithm, inject the adversarial samples into the local training set of the edge device to obtain an enhanced training set, and perform model enhancement training on the corresponding edge device based on the enhanced training set; this sample generation submodule 610 can generate adversarial samples based on algorithms such as FGSM (Fast Gradient Sign Method) and PGD (Projected Gradient Descent). These adversarial samples can simulate attacks that may be encountered in actual applications and are used to enhance the robustness of the model.

[0084] The adversarial evaluation submodule 620 monitors the model adversarial accuracy of each edge device and feeds this accuracy back to the parameter detection module. This adversarial evaluation submodule 620 can use adversarial accuracy as one of the training termination criteria. Adversarial accuracy refers to the model's accuracy on adversarial examples. Improving this accuracy can make the model more stable in the face of attacks.

[0085] The present invention proposes an artificial intelligence network optimization training system and method based on deep learning. It can use a multi-objective optimization module to focus on multiple objectives in model training, such as model accuracy, training time and energy consumption, so that model training no longer focuses only on the single goal of model accuracy, and uses a multi-objective optimization algorithm to make model training more in line with actual application needs; it can also use a collaborative training module to support distributed training on edge devices, fully utilize the computing resources of edge devices, and at the same time solve the data transmission and collaborative training problems between edge devices; and it can use an enhanced training module to inject adversarial samples into the training set, thereby enhancing the robustness of the model.

[0086] On the other hand, see Figure 8 The present invention also provides an artificial intelligence network optimization training method based on deep learning. The implementation of this method is based on the above-mentioned artificial intelligence network optimization training system based on deep learning. The method includes the following steps:

[0087] S1: Receives and parses training task requests, determines the training task type and initial model hyperparameters, and obtains a list of available edge devices;

[0088] S2: Determine the model indicator monitoring threshold based on the training task type, monitor the model indicators in real time during the training process, and adjust the corresponding model hyperparameters when the model indicators reach the preset threshold;

[0089] S3: Initialize the objective weights based on the training task type, the list of available edge devices, and the model hyperparameters, and use the multi-objective optimization algorithm to calculate the multi-objective optimal solution set;

[0090] S4: Generate a training strategy based on the initial model hyperparameters and the multi-objective optimal solution set, and optimize and adjust the training strategy based on the model indicators during the training process;

[0091] S5: Assign training tasks to available edge devices according to the training strategy, and feed back the model indicators during the training process to step S2;

[0092] S6: Generate adversarial samples, inject the adversarial samples into the training set, generate an enhanced training set, use the enhanced training set to perform adversarial enhancement training on the model, calculate the model adversarial accuracy, and feed the model adversarial accuracy back to step S2.

[0093] The above description merely expresses the preferred embodiments of the present invention, and its description is relatively specific and detailed, but it should not be construed as limiting the scope of the patent. It should be noted that a person skilled in the art may make a number of variations and improvements without departing from the concept of the present invention, and these variations and improvements fall within the scope of protection of the present invention. Therefore, the scope of protection of the patent of the present invention shall be based on the appended claims.

Claims

1. An artificial intelligence network optimization training system based on deep learning, the system comprising: The request parsing module is used to receive and parse training task requests, determine the training task type and initial model hyperparameters, and obtain a list of available edge devices; The parameter monitoring module is used to determine the model indicator monitoring threshold according to the training task type, monitor the model indicators in real time during the training process, and adjust the corresponding model hyperparameters when the model indicators reach the preset threshold; The multi-objective optimization module is used to initialize the objective weights based on the training task type, the list of available edge devices, and the model hyperparameters, and calculate the multi-objective optimal solution set using the multi-objective optimization algorithm; The adaptive training module is used to generate a training strategy based on the initial model hyperparameters and the multi-objective optimal solution set, and optimize and adjust the training strategy based on the model indicators during the training process; The collaborative training module is used to assign training tasks to available edge devices according to the training strategy and feed back model indicators during the training process to the parameter detection module.

2. The deep learning-based artificial intelligence network optimization training system according to claim 1, characterized in that: The system also includes: The enhanced training module is used to generate adversarial samples, inject the adversarial samples into the training set, generate an enhanced training set, use the enhanced training set to perform adversarial enhancement training on the model, calculate the model adversarial accuracy, and feed the model adversarial accuracy back to the parameter detection module.

3. The deep learning-based artificial intelligence network optimization training system according to claim 1, characterized in that: The request parsing module includes: The task request submodule is used to receive and parse training task requests, determine the training task type and initial model hyperparameters, and obtain a list of available edge devices. Training task types include classification, detection, and segmentation; initial model hyperparameters include learning rate and number of iterations. The strategy reference submodule is used to call the historical training tasks in the historical training library that are similar to the training task type and initial model hyperparameters according to the training task type and initial model hyperparameters, and use the historical training strategy corresponding to the historical training task as the reference strategy of the parameter detection module.

4. The deep learning-based artificial intelligence network optimization training system according to claim 3, characterized in that: The parameter detection module includes: The monitoring threshold determination submodule is used to set the model indicator monitoring threshold interval based on the historical model indicator history in the historical training task; The model indicator monitoring submodule is used to monitor the real-time values ​​of the model indicators to be monitored during the training process; The hierarchical response submodule is used to compare the real-time value of the model indicator during the training process with the model indicator monitoring threshold interval. If the real-time value of the model indicator exceeds the first interval, a log is recorded; if the real-time value of the model indicator exceeds the second interval, the model hyperparameter corresponding to the model indicator is adjusted; if the real-time value of the model indicator exceeds the third interval, training is stopped.

5. The deep learning-based artificial intelligence network optimization training system according to claim 1, characterized in that: The multi-objective optimization module includes: The target weight configuration submodule is used to set training targets based on the training task type and initial model hyperparameters, and to set weights for each training target according to user needs; wherein the training targets include model accuracy, training time, and energy consumption; The optimization calculation submodule is used to calculate the multi-objective optimal solution set based on the training task type, the list of available edge devices and the model hyperparameters, and use the NSGA2 multi-objective optimization algorithm. The multi-objective optimal solution set includes edge device collaboration parameters, resource allocation parameters, model structure parameters and model hyperparameters.

6. The deep learning-based artificial intelligence network optimization training system according to claim 1, characterized in that: The adaptive training module includes: A training generation submodule is used to generate a boost training unit and a downshift training unit according to a training strategy, and determine the number of the boost training units and the downshift training unit, as well as the difference between adjacent boost training units and the difference between adjacent downshift training units; A dynamic training submodule is used to adjust the number of upshift training units and downshift training units according to model indicators; The training termination submodule is used to determine the model performance based on the model indicators and terminate the model training when the improvement of the model performance by three consecutive upgrading training units or downshifting training units is less than a preset threshold.

7. The deep learning-based artificial intelligence network optimization training system according to claim 1, characterized in that: The collaborative training module includes: The edge device submodule is used to generate a global training model according to the training strategy, distribute the global training model to available edge devices according to the list of available edge devices, and perform personalized training configuration for each available edge device; The federated learning submodule is used to aggregate local model hyperparameters to the central server after the edge device completes local training and aggregate the local model hyperparameters according to the federated learning algorithm to obtain model hyperparameters; The load balancing submodule is used to monitor the load of each edge device and automatically allocate training tasks according to the load of the edge device.

8. The deep learning-based artificial intelligence network optimization training system according to claim 2, characterized in that: The enhanced training module includes: The sample generation submodule is used to generate adversarial samples according to the adversarial sample generation algorithm, inject the adversarial samples into the local training set in the edge device to obtain an enhanced training set, and perform model enhancement training on the corresponding edge device based on the enhanced training set; The adversarial evaluation submodule is used to monitor the model adversarial accuracy in each edge device and feed the model adversarial accuracy back to the parameter detection module.

9. A deep learning-based artificial intelligence network optimization training method, characterized in that: The method comprises the following steps: S1: Receives and parses training task requests, determines the training task type and initial model hyperparameters, and obtains a list of available edge devices; S2: Determine the model indicator monitoring threshold based on the training task type, monitor the model indicators in real time during the training process, and adjust the corresponding model hyperparameters when the model indicators reach the preset threshold; S3: Initialize the objective weights based on the training task type, the list of available edge devices, and the model hyperparameters, and use the multi-objective optimization algorithm to calculate the multi-objective optimal solution set; S4: Generate a training strategy based on the initial model hyperparameters and the multi-objective optimal solution set, and optimize and adjust the training strategy based on the model indicators during training; S5: Assign training tasks to available edge devices according to the training strategy, and feed back the model indicators during the training process to step S2.

10. The deep learning-based artificial intelligence network optimization training method according to claim 9, characterized in that: The method further includes: S6: Generate adversarial samples, inject the adversarial samples into the training set, generate an enhanced training set, use the enhanced training set to perform adversarial enhancement training on the model, calculate the model adversarial accuracy, and feed the model adversarial accuracy back to step S2.

Citation Information

Patent Citations

  • Artificial intelligence network optimization training system and method based on deep learning

    CN116415687A

Cited By

  • Model training control method and electronic equipment

    CN121523978A