Deep learning-based balance coordination training strategy generation method and system

By establishing a frequency domain stability assessment and response energy index system, the resonance risk in the deep learning model training process is evaluated, and adaptive adjustments are made, thus solving the resonance instability problem in the deep learning model training process and improving the reliability and accuracy of training.

CN121503733AInactive Publication Date: 2026-02-10HEBEI INST OF PHYSICAL EDUCATION
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202511832333.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-08
Publication Date
2026-02-10
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In the training process of deep learning models, existing technologies suffer from resonance instability due to the coupling between policy network control instructions and the inherent dynamic characteristics of the model parameter space. This causes the generated balanced and coordinated training strategy to deviate from the optimal solution and may lead to divergence of the parameter space trajectory, resulting in a waste of computational resources.

Method used

By establishing a dual quantization system of frequency domain stability evaluation value and response energy index value, dynamic training data is collected, frequency domain stability and response energy are analyzed, resonance risk is assessed, resonance suppression configuration is performed, training configuration parameters are adjusted, and the dynamic upper limit of gradient clipping is optimized to achieve adaptive adjustment.

Benefits of technology

It effectively reduces the risk of training resonance, improves the reliability and accuracy of model training, ensures that the training process is in a balanced and coordinated state, and improves the overall platform performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121503733A_ABST
    Figure CN121503733A_ABST
Patent Text Reader

Abstract

The invention discloses a balance coordination training strategy generation method and system based on deep learning, and relates to the technical field of data processing. Firstly, in the operation process of a training strategy generation platform, training dynamic data of the training strategy generation platform are collected, a frequency domain stability evaluation value of the training strategy generation platform is obtained through analysis, then response energy parameters of the training strategy generation platform are collected, and response energy index values of the training strategy generation platform are obtained through processing; then, based on the frequency domain stability evaluation value of the training strategy generation platform, in combination with the response energy index value of the training strategy generation platform, a training resonance risk evaluation value of the training strategy generation platform is obtained through processing, finally, a training instruction of the training strategy generation platform is obtained through analysis, and then resonance suppression configuration is carried out. And the reliability of the generated balance coordination training strategy is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing technology, specifically to a method and system for generating balanced and coordinated training strategies based on deep learning. Background Technology

[0002] In the field of intelligent decision-making technology, with the continuous expansion of the scale of deep learning models and the increasing complexity of application scenarios, achieving autonomous coordination and stable optimization of the training process has become a key technical challenge to improve the reliability of artificial intelligence systems. As a core means of achieving complex decision-making tasks, the dynamic stability of the training process of deep reinforcement learning directly determines the performance of the final model.

[0003] Existing technology, such as the invention patent announcement CN118588238B, discloses a method and system for generating balance coordination training strategies based on machine learning. This method includes collecting user exercise data and medical health data; preprocessing the exercise data; formulating a scientific balance coordination scheme using the medical health data; the exercise data including daily exercise data and balance coordination exercise data; obtaining fixed balance coordination values ​​based on the daily exercise data; filtering the balance coordination exercise data based on balance coordination preferences to obtain balance coordination actions; adjusting the scientific balance coordination scheme using the balance coordination actions to obtain an optimized balance coordination scheme; constructing a balance coordination training strategy generation model based on the fixed balance coordination values ​​and the optimized balance coordination scheme; and optimizing the balance coordination training strategy generation model.

[0004] Based on the above solutions, it was found that in the current field of intelligent decision-making technology, existing technologies typically only analyze static motion feature data. However, during model training, there is a problem of command resonance instability caused by the coupling between the policy network's control commands and the inherent dynamic characteristics of the model's parameter space. For example, during deep neural network training, when the dynamic learning rate adjustment frequency of the policy network output forms a harmonic relationship with the inherent period of the model's convolutional layer weight update, the gradient update error of a specific network layer will be periodically amplified, triggering a vicious cycle of continuous oscillation of filter weights. This internal resonance not only causes the generated balanced and coordinated training strategy to deviate significantly from the optimal solution, but may also cause the parameter space trajectory to diverge, resulting in a continuous shift in the distribution of activation values ​​between network layers. Ultimately, this leads to a decrease in the model's accuracy on the test set, resulting in a serious waste of computational resources. Summary of the Invention

[0005] To address the shortcomings of existing technologies, this invention provides a method and system for generating balanced coordination training strategies based on deep learning, which can effectively solve the problems mentioned in the background technology.

[0006] To achieve the above objectives, the first aspect of the present invention is implemented through the following technical solution: a method for generating balanced and coordinated training strategies based on deep learning, comprising collecting training dynamic data of the training strategy generation platform during the operation of the training strategy generation platform, and analyzing the frequency domain stability evaluation value of the training strategy generation platform.

[0007] Collect the response energy parameters of the training strategy generation platform and process them to obtain the response energy index value of the training strategy generation platform.

[0008] Based on the frequency domain stability evaluation value of the training strategy generation platform, combined with the response energy index value of the training strategy generation platform, the training resonance risk evaluation value of the training strategy generation platform is obtained.

[0009] Based on the training resonance risk assessment value of the training strategy generation platform, the training instructions of the training strategy generation platform are analyzed, and then resonance suppression configuration is performed.

[0010] Furthermore, the analysis yields the frequency domain stability evaluation value of the training strategy generation platform. The specific analysis process is as follows: during the operation of the training strategy generation platform, the training dynamic data of the training strategy generation platform is collected. The training dynamic data of the training strategy generation platform includes the periodic oscillation main frequency amplitude of the parameter update amount of the training strategy generation platform, the fluctuation coefficient of the loss function change rate, and the fluctuation coefficient of the weight update amount.

[0011] Based on the training dynamic data of the training strategy generation platform, the frequency domain stability evaluation value of the training strategy generation platform is obtained. The frequency domain stability evaluation value of the training strategy generation platform represents the quantitative result of the training dynamic data jointly reflecting the degree of the inherent dynamic regularity of the training platform.

[0012] Furthermore, the processing yields the response energy index value of the training policy generation platform. Specifically, the response energy parameters of the training policy generation platform include the sum of the absolute values ​​of the parameter updates of the training policy generation platform, the peak fluctuation ratio of the gradient norm, and the dynamic fluctuation coefficient of the activation function output value.

[0013] Based on the response energy parameters of the training strategy generation platform, the response energy index value of the training strategy generation platform is obtained. The response energy index value of the training strategy generation platform represents the quantification result of the intensity of the output fluctuation of the training platform when it responds to instructions, which is jointly expressed by the response energy parameters.

[0014] Furthermore, the processing yields a training resonance risk assessment value for the training strategy generation platform. Specifically, the process involves: based on the frequency domain stability assessment value of the training strategy generation platform and combined with the response energy index value of the training strategy generation platform, a training resonance risk assessment value for the training strategy generation platform is obtained. This training resonance risk assessment value is used to comprehensively and quantitatively assess the degree of coupling anomaly of the training platform.

[0015] Furthermore, the analysis yields the training instructions of the training strategy generation platform. Specifically, the training instructions of the training strategy generation platform include dynamic decoupling instructions, training maintenance instructions, and strategy execution instructions.

[0016] The training resonance risk assessment value of the training strategy generation platform is compared with the set training resonance risk assessment threshold. If the training resonance risk assessment value of the training strategy generation platform is higher than the set training resonance risk assessment threshold, the training instruction of the training strategy generation platform is marked as a dynamic decoupling instruction; otherwise, the training instruction of the training strategy generation platform is marked as a training maintenance instruction.

[0017] Furthermore, the resonance suppression configuration is then performed, specifically as follows: the training instructions from the training strategy generation platform are extracted; if the training instructions from the training strategy generation platform are training maintenance instructions, the current training configuration parameters are used to continue generating subsequent training strategies; if the training instructions from the training strategy generation platform are dynamic decoupling instructions, the training configuration parameters are adjusted to perform resonance suppression configuration.

[0018] Furthermore, the specific configuration process for adjusting the training configuration parameters to perform resonance suppression is as follows: if the training instruction of the training policy generation platform is a dynamic decoupling instruction, then the output control parameters of the policy network of the training policy generation platform are optimized, and after optimization, the dynamic decoupling instruction of the training policy generation platform is converted into a policy execution instruction.

[0019] After receiving the policy execution instruction, the training policy generation platform adjusts the dynamic upper limit of gradient clipping.

[0020] Furthermore, the specific process of adjusting the gradient pruning dynamic upper limit of the training policy generation platform is as follows: after receiving the policy execution instruction, based on the training resonance risk assessment value of the training policy generation platform, the gradient pruning dynamic upper limit adjustment value of the training policy generation platform is obtained. The current gradient pruning dynamic upper limit value of the training policy generation platform is subtracted from the gradient pruning dynamic upper limit adjustment value to obtain the gradient pruning dynamic upper limit target value of the training policy generation platform. Then, the current gradient pruning dynamic upper limit value is adjusted to the gradient pruning dynamic upper limit target value.

[0021] Based on the gradient pruning dynamic upper limit target value of the training strategy generation platform, the balance and coordination training stability information of the training strategy generation platform is obtained.

[0022] Furthermore, the matching process for obtaining the balance and coordination training stability information of the training strategy generation platform is as follows: the gradient clipping dynamic upper limit target value of the training strategy generation platform is matched with the balance and coordination training stability information corresponding to each gradient clipping dynamic upper limit target value interval stored in the strategy generation database, and the balance and coordination training stability information corresponding to the interval where the gradient clipping dynamic upper limit target value is located is statistically analyzed and recorded as the balance and coordination training stability information of the training strategy generation platform.

[0023] A second aspect of the present invention provides a deep learning-based balanced coordination training strategy generation system, comprising: a frequency domain stability analysis module, used to collect training dynamic data of the training strategy generation platform during its operation and analyze it to obtain a frequency domain stability evaluation value of the training strategy generation platform.

[0024] The response energy analysis module is used to collect the response energy parameters of the training strategy generation platform and process them to obtain the response energy index value of the training strategy generation platform.

[0025] The resonance risk assessment module is used to process the frequency domain stability assessment value of the training strategy generation platform and the response energy index value of the training strategy generation platform to obtain the training resonance risk assessment value of the training strategy generation platform.

[0026] The training instruction analysis module is used to analyze the training instructions of the training strategy generation platform based on the training resonance risk assessment value of the training strategy generation platform, and then configure resonance suppression.

[0027] The present invention has the following beneficial effects: (1) By establishing a dual quantization system of frequency domain stability evaluation value and response energy index value, this invention realizes the accurate evaluation of the risk of coupling between the policy network control command and the inherent frequency of the model parameter space during the training process, and then controls it before resonance instability, reduces the training resonance risk of the platform, improves the reliability of model training in the platform, and thus improves the reliability and accuracy of the generated balanced and coordinated training strategy.

[0028] (2) The present invention adopts an adaptive adjustment training strategy. Through resonance suppression configuration, the training configuration parameters can be automatically adjusted according to the training resonance risk assessment value, which can flexibly cope with different training environments and data characteristics, improve the reliability and accuracy of the training process, not only effectively avoid training collapse or model performance degradation caused by resonance instability, but also optimize training parameters according to real-time feedback, ensuring that the training process is always maintained in a balanced and coordinated state, thereby improving the overall application performance of the platform. Attached Figure Description

[0029] Figure 1 This is a schematic diagram of the method flow of the present invention.

[0030] Figure 2 This is a schematic diagram of the system module connections of the present invention. Detailed Implementation

[0031] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0032] Please see Figure 1 As shown, the first aspect of the present invention provides a technical solution: a method for generating balanced and coordinated training strategies based on deep learning, comprising collecting training dynamic data of the training strategy generation platform during the operation of the training strategy generation platform, and analyzing the frequency domain stability evaluation value of the training strategy generation platform.

[0033] Collect the response energy parameters of the training strategy generation platform and process them to obtain the response energy index value of the training strategy generation platform.

[0034] Based on the frequency domain stability evaluation value of the training strategy generation platform, combined with the response energy index value of the training strategy generation platform, the training resonance risk evaluation value of the training strategy generation platform is obtained.

[0035] Based on the training resonance risk assessment value of the training strategy generation platform, the training instructions of the training strategy generation platform are analyzed, and then resonance suppression configuration is performed.

[0036] Specifically, the frequency domain stability evaluation value of the training strategy generation platform is obtained through analysis. The specific analysis process is as follows: during the operation of the training strategy generation platform, the training dynamic data of the training strategy generation platform is collected. The training dynamic data of the training strategy generation platform includes the periodic oscillation main frequency amplitude of the parameter update amount of the training strategy generation platform, the fluctuation coefficient of the loss function change rate, and the fluctuation coefficient of the weight update amount.

[0037] It should be noted that the periodic oscillation amplitude of the parameter update quantity in the training policy generation platform reflects the dominant periodic fluctuation intensity of the deep learning model parameters in the training process. This can be obtained by performing a Fast Fourier Transform on the parameter update quantity sequence of multiple consecutive training cycles to obtain its spectral distribution, and extracting the amplitude value corresponding to the frequency component with the highest energy in the spectrum as the dominant frequency amplitude. The fluctuation coefficient of the loss function change rate reflects the fluctuation stability of the loss value convergence process. This can be obtained by calculating the change of the loss function value between every two adjacent iteration steps within a fixed training iteration window to obtain the loss change rate sequence, then calculating the standard deviation of the sequence and dividing it by the absolute value of its mean to finally obtain the fluctuation coefficient of the loss function change rate. The fluctuation coefficient of the weight update quantity reflects the fluctuation stability of the deep learning model parameter update process in the training policy generation platform. This can be obtained by calculating the standard deviation of the weight matrix update quantity of each layer of the model within a fixed time window, and then averaging these standard deviations.

[0038] It should be noted that the training strategy generation platform is an intelligent decision-making system based on a deep learning architecture. It adaptively generates a balanced and coordinated training strategy by analyzing the multi-dimensional dynamic features generated during the training process in real time. The deep learning model in this platform is the core component, and its role is to establish a mapping relationship between the dynamic characteristics and stability of the training system, thereby generating an optimized training strategy with balanced and coordinated capabilities.

[0039] Based on the training dynamic data of the training strategy generation platform, the frequency domain stability evaluation value of the training strategy generation platform is obtained. The frequency domain stability evaluation value of the training strategy generation platform represents the quantitative result of the training dynamic data jointly reflecting the degree of the inherent dynamic regularity of the training platform.

[0040] In this embodiment, the frequency domain stability evaluation value of the training strategy generation platform can be obtained through the following analysis method, with the specific analysis conditions as follows: ; In the formula, This represents the frequency domain stability evaluation value of the training policy generation platform. This represents the periodic oscillation amplitude of the parameter update amount of the training strategy generation platform. The stability evaluation factor represents the periodic oscillation frequency amplitude corresponding to the set unit parameter update amount. The fluctuation coefficient represents the rate of change of the loss function of the training policy generation platform. The stability assessment factor represents the volatility coefficient corresponding to the rate of change of the set loss function. This represents the fluctuation coefficient of the weight update volume of the training strategy generation platform. The stability evaluation factor corresponding to the volatility coefficient of the set weight update amount is represented by e, which represents the natural constant.

[0041] It should be explained that by utilizing the properties of the exponential function, the numerical values ​​obtained by the linear weighted combination are mapped to the interval (0, 1] to form the final frequency domain stability evaluation value, making the platform status monitoring more intuitive and easier to explain.

[0042] It should be added that, in this embodiment, the stable evaluation factor corresponding to the periodic oscillation frequency amplitude of the unit parameter update, the stable evaluation factor corresponding to the fluctuation coefficient of the loss function change rate, and the stable evaluation factor corresponding to the fluctuation coefficient of the weight update are obtained from the strategy generation database.

[0043] It should be explained that the stability evaluation factors corresponding to the periodic oscillation amplitude of the unit parameter update, the fluctuation coefficient of the loss function rate of change, and the fluctuation coefficient of the weight update are used to adjust the importance of the training dynamic data of the training strategy generation platform in the process of analyzing and obtaining the frequency domain stability evaluation value. For example, there is a pre-defined mapping relationship between the training dynamic data of the training strategy generation platform and the corresponding stability evaluation factors in the strategy generation database. Through the pre-defined mapping relationship, the stability evaluation factors corresponding to the real-time training dynamic data of the training strategy generation platform can be matched. By matching the training dynamic data of the training strategy generation platform with the pre-defined mapping relationship, the stability evaluation factors corresponding to the periodic oscillation amplitude of the unit parameter update, the fluctuation coefficient of the loss function rate of change, and the fluctuation coefficient of the weight update are obtained.

[0044] In this implementation scheme, the periodic oscillation amplitude of the parameter update quantity, the fluctuation coefficient of the loss function rate of change, and the fluctuation coefficient of the weight update quantity are correlated and do not exist independently. For example, an increase in the periodic oscillation amplitude of the parameter update quantity will directly cause synchronous fluctuations in the loss function rate of change, resulting in a corresponding increase in the fluctuation coefficient of the loss function rate of change. At the same time, this periodic oscillation of parameter updates will disrupt the stability of weight updates, causing the fluctuation coefficient of weight updates to rise as well. This chain reaction will form a positive feedback loop, increasing the instability of the training process. The frequency domain stability coefficient of the training strategy generation platform obtained by comprehensive analysis can be used to evaluate the internal dynamic coordination and anti-interference ability of the training platform, thereby helping to evaluate the stability of the platform model performance.

[0045] Specifically, the response energy index value of the training policy generation platform is obtained through processing. The specific processing procedure is as follows: the response energy parameters of the training policy generation platform include the sum of the absolute values ​​of the parameter update amounts of the training policy generation platform, the peak fluctuation ratio of the gradient norm, and the dynamic fluctuation coefficient of the activation function output value.

[0046] It should be noted that the sum of the absolute values ​​of parameter updates in the training policy generation platform represents the overall strength of the deep learning model parameters' response to training instructions within the platform, which can be obtained by accumulating the absolute values ​​of all weight parameter updates within a single training cycle; the peak-to-peak variability ratio of the gradient norm characterizes the suddenness and instability of gradient changes, and is obtained by calculating the ratio of the maximum value to the average value of the gradient norm within a fixed window; the dynamic variability coefficient of the activation function output value measures the magnitude of change in the activation state within the neural network in response to training instructions, and can be obtained by statistically analyzing the difference between the maximum and minimum values ​​of the activation function output values ​​of each layer in the deep learning model within the training policy generation platform, and then dividing by the standard deviation of the activation values ​​of that layer.

[0047] Based on the response energy parameters of the training strategy generation platform, the response energy index value of the training strategy generation platform is obtained. The response energy index value of the training strategy generation platform represents the quantification result of the intensity of the output fluctuation of the training platform when it responds to instructions, which is jointly expressed by the response energy parameters.

[0048] In this embodiment, the response energy index value of the training strategy generation platform can be obtained through the following analysis method, with the specific analysis conditions as follows: ; In the formula, This represents the response energy metric value of the training strategy generation platform. This represents the sum of the absolute values ​​of parameter updates from the training policy generation platform. This represents the weight factor corresponding to the sum of the absolute values ​​of the set unit parameter update amounts. This represents the peak-to-peak variability of the gradient norm of the training policy generation platform. This represents the weight factor corresponding to the peak-to-peak variability of the set gradient norm. This represents the dynamic fluctuation coefficient of the activation function output value of the training policy generation platform. This represents the weighting factor corresponding to the dynamic fluctuation coefficient of the output value of the set activation function.

[0049] It should be added that, in this embodiment, the weight factors corresponding to the sum of the absolute values ​​of the unit parameter update amounts, the weight factors corresponding to the peak fluctuation ratio of the gradient norm, and the weight factors corresponding to the dynamic fluctuation coefficient of the activation function output value are obtained from the policy generation database.

[0050] It should be explained that the weighting factors corresponding to the sum of absolute values ​​of unit parameter updates, the peak fluctuation ratio of the gradient norm, and the dynamic fluctuation coefficient of the activation function output value are used to adjust the importance of the response energy parameters of the training policy generation platform in the process of analyzing and obtaining the response energy index value of the training policy generation platform.

[0051] In this implementation scheme, the sum of absolute values ​​of parameter updates, the peak variability ratio of the gradient norm, and the dynamic variability coefficient of the activation function output value of the training strategy generation platform are correlated and do not exist independently. For example, an increase in the sum of absolute values ​​of parameter updates will directly lead to drastic changes in the activation state within the network, causing a corresponding increase in the dynamic variability coefficient of the activation function output value. At the same time, such large parameter updates often trigger abnormal fluctuations in gradient calculation, resulting in a significant increase in the peak variability ratio of the gradient norm. This interaction will form a chain reaction, exacerbating the energy accumulation in the training process. By comprehensively analyzing the response energy index value of the training strategy generation platform, the sensitivity and response strength of the training model to external control commands can be evaluated, thereby helping to assess the potential risk of resonance instability during the training process.

[0052] Specifically, the training resonance risk assessment value of the training strategy generation platform is obtained through the following process: based on the frequency domain stability assessment value of the training strategy generation platform and combined with the response energy index value of the training strategy generation platform, the training resonance risk assessment value of the training strategy generation platform is obtained. The training resonance risk assessment value of the training strategy generation platform is used to comprehensively and quantitatively assess the degree of coupling anomaly of the training platform.

[0053] In this embodiment, the training resonance risk assessment value of the training strategy generation platform can be obtained through the following analysis method, with the specific analysis conditions as follows: ; In the formula, This represents the training resonance risk assessment value of the training strategy generation platform. This represents the frequency domain stability evaluation value of the training policy generation platform. This represents the influence factor corresponding to the set frequency domain stability evaluation value. This represents the response energy metric value of the training strategy generation platform. This indicates the influence factor corresponding to the set response energy index value.

[0054] It should be added that, in this embodiment, the influence factors corresponding to the preset frequency domain stability evaluation values ​​and the influence factors corresponding to the response energy index values ​​are obtained from the strategy generation database.

[0055] It should be explained that the influence factors corresponding to the frequency domain stability evaluation value and the response energy index value are used to adjust the importance of the frequency domain stability evaluation value and the response energy index value of the training strategy generation platform in the process of analyzing and obtaining the training resonance risk evaluation value of the training strategy generation platform.

[0056] In this embodiment, the frequency domain stability evaluation value and response energy index value of the platform are generated by the training strategy to quantify the inherent dynamic characteristics and command response strength of the platform, respectively. This can accurately assess the resonance risk caused by the coupling of the two. The training dynamic data is used to detect the potential conditions for resonance to occur, and the response energy parameter is used to assess the energy basis for resonance development. The combined effect of the two determines the stability and safety of the training process. This analysis method provides a data basis for the early diagnosis and effective suppression of resonance instability problems.

[0057] Specifically, the training instructions of the training strategy generation platform are analyzed and obtained. The specific analysis process is as follows: the training instructions of the training strategy generation platform include dynamic decoupling instructions, training maintenance instructions, and strategy execution instructions.

[0058] The training resonance risk assessment value of the training strategy generation platform is compared with the set training resonance risk assessment threshold. If the training resonance risk assessment value of the training strategy generation platform is higher than the set training resonance risk assessment threshold, the training instruction of the training strategy generation platform is marked as a dynamic decoupling instruction; otherwise, the training instruction of the training strategy generation platform is marked as a training maintenance instruction.

[0059] Specifically, resonance suppression configuration is then performed. The specific configuration process is as follows: the training instructions from the training strategy generation platform are extracted. If the training instructions from the training strategy generation platform are training maintenance instructions, the current training configuration parameters are used to continue generating subsequent training strategies. If the training instructions from the training strategy generation platform are dynamic decoupling instructions, the training configuration parameters are adjusted to perform resonance suppression configuration.

[0060] It should be added that the training configuration parameters include the output control parameters of the policy network of the training policy generation platform and the dynamic upper limit of gradient clipping. The analysis of the training resonance risk assessment value of the training policy generation platform can reflect the degree of training resonance risk. If the training resonance risk of the training policy generation platform is high, and subsequent configuration is carried out with the original default training configuration parameters, it may lead to violent oscillations in the training process, unstable model parameter updates, or even training process crash. Therefore, it is necessary to adjust the training configuration parameters in conjunction with the training resonance risk assessment value. The larger the training resonance risk assessment value, the greater the training resonance risk, and the gradient clipping dynamic upper limit needs to be reduced to limit the maximum magnitude of parameter updates.

[0061] Specifically, the training configuration parameters are adjusted to configure resonance suppression. The specific configuration process is as follows: if the training instructions of the training policy generation platform are dynamic decoupling instructions, the output control parameters of the policy network of the training policy generation platform are optimized. After optimization, the dynamic decoupling instructions of the training policy generation platform are converted into policy execution instructions.

[0062] It should be explained that the optimization of the policy network output control parameters of the training policy generation platform is specifically performed by performing low-pass filtering and smoothing on the learning rate and momentum coefficient of the policy network output. This reduces the risk of training resonance by disrupting the generation conditions of high-frequency resonance.

[0063] After receiving the policy execution instruction, the training policy generation platform adjusts the dynamic upper limit of gradient clipping.

[0064] Specifically, the dynamic upper limit of gradient clipping for the training policy generation platform is adjusted. The process is as follows: after receiving the policy execution instruction, the dynamic upper limit of gradient clipping for the training policy generation platform is adjusted based on the training resonance risk assessment value of the platform. The current dynamic upper limit of gradient clipping is subtracted from the adjustment value to obtain the target value. The current dynamic upper limit is then adjusted to the target value to limit the instantaneous update amplitude of the model.

[0065] It should be added that the gradient clipping dynamic upper limit adjustment value of the training policy generation platform is obtained through analysis. Specifically, the absolute value of the difference between the training resonance risk assessment value of the training policy generation platform and the set training resonance risk assessment threshold is recorded as the training resonance risk control value of the training policy generation platform. The training resonance risk control value of the training policy generation platform is matched with the gradient clipping dynamic upper limit adjustment value corresponding to each training resonance risk control value stored in the policy generation database. The gradient clipping dynamic upper limit adjustment value corresponding to the training resonance risk control value is statistically analyzed and recorded as the gradient clipping dynamic upper limit adjustment value of the training policy generation platform. The larger the training resonance risk control value, the greater the resonance risk. The larger the matched gradient clipping dynamic upper limit adjustment value, the smaller the target value of the gradient clipping dynamic upper limit.

[0066] Based on the gradient pruning dynamic upper limit target value of the training strategy generation platform, the balance and coordination training stability information of the training strategy generation platform is obtained.

[0067] Specifically, the process of matching the balance and coordination training stability information of the training strategy generation platform is as follows: the target value of the gradient clipping dynamic upper limit of the training strategy generation platform is matched with the balance and coordination training stability information corresponding to each interval of the gradient clipping dynamic upper limit target value stored in the strategy generation database. The balance and coordination training stability information corresponding to the interval in which the gradient clipping dynamic upper limit target value is located is statistically analyzed and recorded as the balance and coordination training stability information of the training strategy generation platform. The balance and coordination training stability information of the training strategy generation platform includes danger and warning. Then, the balance and coordination training stability information is transmitted to the staff's mobile device for warning prompts. The larger the target value of the gradient clipping dynamic upper limit, the more necessary it is to adjust the gradient clipping dynamic upper limit value to update the constraint, and the greater the possibility that the matched balance and coordination training stability information is dangerous.

[0068] A second aspect of the present invention provides a deep learning-based balanced coordination training strategy generation system, comprising: a frequency domain stability analysis module, used to collect training dynamic data of the training strategy generation platform during its operation and analyze it to obtain a frequency domain stability evaluation value of the training strategy generation platform.

[0069] The response energy analysis module is used to collect the response energy parameters of the training strategy generation platform and process them to obtain the response energy index value of the training strategy generation platform.

[0070] The resonance risk assessment module is used to process the frequency domain stability assessment value of the training strategy generation platform and the response energy index value of the training strategy generation platform to obtain the training resonance risk assessment value of the training strategy generation platform.

[0071] The training instruction analysis module is used to analyze the training instructions of the training strategy generation platform based on the training resonance risk assessment value of the training strategy generation platform, and then configure resonance suppression.

[0072] It should be noted that a method and system for generating balanced coordination training strategies based on deep learning also includes a strategy generation database, which stores the first parameter set, the second parameter set, the third parameter set, and the fourth parameter set obtained by analyzing historical data.

[0073] The first parameter set includes the stability evaluation factor corresponding to the periodic oscillation frequency amplitude of the unit parameter update, the stability evaluation factor corresponding to the volatility coefficient of the loss function change rate, and the stability evaluation factor corresponding to the volatility coefficient of the weight update.

[0074] The second parameter set includes the weight factor corresponding to the sum of the absolute values ​​of unit parameter updates, the weight factor corresponding to the peak variability ratio of the gradient norm, and the weight factor corresponding to the dynamic variability coefficient of the activation function output value.

[0075] The third parameter set includes the influence factors corresponding to the frequency domain stability assessment value and the influence factors corresponding to the response energy index value.

[0076] The fourth parameter set includes the training resonance risk assessment threshold, the gradient clipping dynamic upper limit adjustment value corresponding to each training resonance risk control value, and the balance and coordination training stability information corresponding to the target value range of each gradient clipping dynamic upper limit.

[0077] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.

[0078] The preferred embodiments of the present invention disclosed above are merely illustrative of the invention. These preferred embodiments do not exhaustively describe all details, nor do they limit the invention to the specific implementations described. Clearly, many modifications and variations can be made based on the content of this specification. This specification selects and specifically describes these embodiments to better explain the principles and practical applications of the invention, thereby enabling those skilled in the art to better understand and utilize the invention.

Claims

1. A method for generating a balanced coordination training strategy based on deep learning, characterized in that, include: During the operation of the training strategy generation platform, the training dynamic data of the training strategy generation platform is collected and analyzed to obtain the frequency domain stability evaluation value of the training strategy generation platform. Collect the response energy parameters of the training strategy generation platform and process them to obtain the response energy index value of the training strategy generation platform; Based on the frequency domain stability evaluation value of the training strategy generation platform, combined with the response energy index value of the training strategy generation platform, the training resonance risk evaluation value of the training strategy generation platform is obtained. Based on the training resonance risk assessment value of the training strategy generation platform, the training instructions of the training strategy generation platform are analyzed, and then resonance suppression configuration is performed.

2. The method for generating a balanced coordination training strategy based on deep learning according to claim 1, characterized in that: The analysis yielded the frequency domain stability evaluation value of the training strategy generation platform. The specific analysis process is as follows: During the operation of the training strategy generation platform, training dynamic data of the training strategy generation platform is collected. The training dynamic data of the training strategy generation platform includes the periodic oscillation frequency amplitude of the parameter update amount of the training strategy generation platform, the fluctuation coefficient of the loss function change rate, and the fluctuation coefficient of the weight update amount. Based on the training dynamic data of the training strategy generation platform, the frequency domain stability evaluation value of the training strategy generation platform is obtained. The frequency domain stability evaluation value of the training strategy generation platform represents the quantitative result of the training dynamic data jointly reflecting the degree of the inherent dynamic regularity of the training platform.

3. The method for generating a balanced coordination training strategy based on deep learning according to claim 1, characterized in that: The processing yields the response energy index value of the training policy generation platform. The specific processing procedure is as follows: The response energy parameters of the training strategy generation platform include the sum of the absolute values ​​of the parameter updates of the training strategy generation platform, the peak fluctuation ratio of the gradient norm, and the dynamic fluctuation coefficient of the activation function output value. Based on the response energy parameters of the training strategy generation platform, the response energy index value of the training strategy generation platform is obtained. The response energy index value of the training strategy generation platform represents the quantification result of the intensity of the output fluctuation of the training platform when it responds to instructions, which is jointly expressed by the response energy parameters.

4. The method for generating a balanced coordination training strategy based on deep learning according to claim 1, characterized in that: The processing yields the training resonance risk assessment value for the training strategy generation platform. The specific process is as follows: Based on the frequency domain stability evaluation value of the training strategy generation platform and combined with the response energy index value of the training strategy generation platform, the training resonance risk evaluation value of the training strategy generation platform is obtained. The training resonance risk evaluation value of the training strategy generation platform is used to comprehensively and quantitatively evaluate the coupling anomaly degree of the training platform.

5. The method for generating a balanced coordination training strategy based on deep learning according to claim 1, characterized in that: The analysis yields the training instructions for the training strategy generation platform. The specific analysis process is as follows: The training instructions of the training strategy generation platform include dynamic decoupling instructions, training maintenance instructions, and strategy execution instructions; The training resonance risk assessment value of the training strategy generation platform is compared with the set training resonance risk assessment threshold. If the training resonance risk assessment value of the training strategy generation platform is higher than the set training resonance risk assessment threshold, the training instruction of the training strategy generation platform is marked as a dynamic decoupling instruction; otherwise, the training instruction of the training strategy generation platform is marked as a training maintenance instruction.

6. The method for generating a balanced coordination training strategy based on deep learning according to claim 1, characterized in that: The resonance suppression configuration is then performed, and the specific configuration process is as follows: Extract the training instructions from the training strategy generation platform. If the training instructions from the training strategy generation platform are training maintenance instructions, continue to use the current training configuration parameters for subsequent training strategy generation. If the training instructions from the training strategy generation platform are dynamic decoupling instructions, adjust the training configuration parameters to configure resonance suppression.

7. The method for generating a balanced coordination training strategy based on deep learning according to claim 6, characterized in that: The specific configuration process for adjusting training configuration parameters to configure resonance suppression is as follows: If the training instructions of the training policy generation platform are dynamic decoupling instructions, then the output control parameters of the policy network of the training policy generation platform are optimized. After optimization, the dynamic decoupling instructions of the training policy generation platform are converted into policy execution instructions. After receiving the policy execution instruction, the training policy generation platform adjusts the dynamic upper limit of gradient clipping.

8. The method for generating a balanced coordination training strategy based on deep learning according to claim 7, characterized in that: The process of adjusting the dynamic upper limit of gradient clipping on the training policy generation platform is as follows: After receiving the policy execution instruction, the gradient clipping dynamic upper limit adjustment value of the training policy generation platform is obtained by analyzing the training resonance risk assessment value of the training policy generation platform. The current gradient clipping dynamic upper limit value of the training policy generation platform is subtracted from the gradient clipping dynamic upper limit adjustment value to obtain the gradient clipping dynamic upper limit target value of the training policy generation platform. Then, the current gradient clipping dynamic upper limit value is adjusted to the gradient clipping dynamic upper limit target value. Based on the gradient pruning dynamic upper limit target value of the training strategy generation platform, the balance and coordination training stability information of the training strategy generation platform is obtained.

9. The method for generating a balanced coordination training strategy based on deep learning according to claim 8, characterized in that: The matching process obtains the balance and coordination training stability information of the training strategy generation platform. The specific process is as follows: The dynamic upper limit target value of gradient clipping in the training policy generation platform is matched with the balance and coordination training stability information corresponding to each interval of dynamic upper limit target value of gradient clipping stored in the policy generation database. The balance and coordination training stability information corresponding to the interval of the dynamic upper limit target value of gradient clipping is statistically analyzed and recorded as the balance and coordination training stability information of the training policy generation platform.

10. A deep learning-based balanced coordination training strategy generation system, characterized in that, include: The frequency domain stability analysis module is used to collect training dynamic data of the training strategy generation platform during its operation and analyze it to obtain the frequency domain stability evaluation value of the training strategy generation platform. The response energy analysis module is used to collect the response energy parameters of the training strategy generation platform and process them to obtain the response energy index value of the training strategy generation platform. The resonance risk assessment module is used to process the frequency domain stability assessment value of the training strategy generation platform and the response energy index value of the training strategy generation platform to obtain the training resonance risk assessment value of the training strategy generation platform. The training instruction analysis module is used to analyze the training instructions of the training strategy generation platform based on the training resonance risk assessment value of the training strategy generation platform, and then configure resonance suppression.

Citation Information

Patent Citations

  • A method and system for generating balance and coordination training strategies based on machine learning

    CN118588238B