A cloud edge collaboration-based model dynamic training method, device and management and control system

By analyzing model attributes and computational complexity, dynamically allocating training locations and making adaptive corrections, the problem of insufficient utilization of cloud-edge collaborative resources is solved, thereby improving model training efficiency and resource utilization.

CN120509463BActive Publication Date: 2025-11-07STATE GRID JIANGSU ECONOMIC RES INST +1
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202511000734.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-21
Publication Date
2025-11-07
Estimated Expiration
2045-07-21

AI Technical Summary

Technical Problem

In existing technologies, the resource scheduling and operation and maintenance systems of cloud and edge computing are not yet perfect, which makes it difficult for cloud-edge collaboration to make full use of their respective computing resources and effectively improve model training efficiency.

Method used

By analyzing the attributes and computational complexity of the model to be trained, training locations are dynamically allocated, and adaptive corrections are made in conjunction with model training metrics to achieve cloud-edge collaborative training.

Benefits of technology

This improved the efficiency of model training and resource utilization, and enhanced the adaptability and stability of the integrated energy system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120509463B_ABST
    Figure CN120509463B_ABST
Patent Text Reader

Abstract

The application discloses a kind of based on cloud edge coordination's model dynamic training method, device and management and control system, method includes: obtaining to be trained model;According to the model attribute of to be trained model, the computing complexity of to be trained model is considered, the computing power demand of each to be trained model is analyzed, and to be trained model is distributed, and initial training position is given;Based on initial training position, to be trained model is trained, and to be trained model is adaptively corrected in combination with model training index, and the training of to be trained model is completed.Through the analysis of various model attributes of to be trained model, and to be trained model is distributed, initial training position is determined and is trained, then adaptive correction is integrated in the process of model training, improve the training efficiency of to be trained model.Meanwhile, the training result of to be trained model can be dynamically adjusted, and the model with poor training effect is updated in time, to improve the adaptability and stability of comprehensive energy system.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of edge collaboration of industrial internet platform, and particularly relates to a model dynamic training method and device based on cloud-edge collaboration and a management and control system. BACKGROUND

[0002] Cloud model training relies on powerful computing resources and a large amount of data, and can train large models with excellent performance. However, these models are usually large in size and require a large amount of computing and communication resources to support their operation. In addition, centralized model training faces increasing pressure on storage, model parameter caching and computing costs.

[0003] Edge computing can significantly reduce latency, reduce bandwidth usage, and improve system efficiency and user experience by migrating data processing to edge nodes closer to data sources. Edge model training allows data processing and model updating to be performed locally, reducing the risk of privacy leakage and enabling more personalized services. However, the computing and storage capabilities of edge devices are limited, making it difficult to independently complete large-scale model training.

[0004] Currently, cloud and edge collaboration cannot fully utilize the computing resources of each other. The root cause is that the lack of unified standards and platforms leads to fragmentation of the cloud-edge ecosystem, and the resource scheduling and operation system is not perfect, making it difficult to allocate loads flexibly, and the resources of both ends cannot be well utilized, and the training efficiency of the model is not considered.

[0005] In order to overcome the limitations of single cloud or edge training, researchers have proposed an architecture for training cloud and edge models together. This architecture migrates some tasks from the cloud to the edge, achieving more private, real-time and personalized user experience.

[0006] The patent application CN117997904A discloses a cloud edge collaborative edge device expansion deployment method and system, analyzes the task processing logs of all edge devices under the cloud platform, identifies the target edge device that has a task processing overflow event, and determines the overflow task information of the target edge device; based on the location information of the target edge device within the corresponding cloud network of the cloud platform, determines the edge device cluster that can handle the overflow task of the target edge device, selects an auxiliary edge device to handle the overflow task, realizes the expansion deployment processing of the overflow task of the target edge device, fully utilizes the edge operation computing power of the cloud platform; the normal processing result of the overflow task is also returned to the target edge device, and the non-overflow task processing result of the target edge device and the normal processing result of the received overflow task are integrated to obtain a complete task processing result, and the overflow task of the edge device is effectively and timely processed. The prior art is to train all tasks on the edge side, which is a side-to-side collaborative training, and the cloud end only plays a command role. It collects the running conditions of each edge device and overflow data, and distributes the overflow data to idle edge devices for training, and does not fully utilize the computing power resources of the cloud end.

[0007] How to fully utilize the computing power resources of the cloud and the edge, shorten the model training time, and improve the model training efficiency is a problem to be solved at present. SUMMARY

[0008] In view of the defects in the prior art, the present application provides a cloud edge collaborative model dynamic training method, device and management and control system. The method comprises: obtaining a to-be-trained model; analyzing the computing power requirements of each to-be-trained model according to the model attributes of the to-be-trained model and taking into account the computing complexity of the to-be-trained model, and distributing the to-be-trained model to give an initial training position; training the to-be-trained model based on the initial training position, and adaptively correcting the to-be-trained model in combination with the model training index to complete the training of the to-be-trained model. By analyzing various model attributes of the to-be-trained model and distributing the to-be-trained model, the initial training position is determined and trained, and then adaptive correction is integrated into the model training process, thereby improving the training efficiency of the to-be-trained model. At the same time, the training results of the to-be-trained model can be dynamically adjusted, and the models with poor training effects can be updated in a timely manner, thereby improving the adaptability and stability of the comprehensive energy system.

[0009] In a first aspect, the present application provides a cloud edge collaborative model dynamic training method, which specifically comprises the following steps:

[0010] Obtaining a to-be-trained model;

[0011] According to the model attribute of the to-be-trained model, the computing complexity of the to-be-trained model is considered, the computing power demand of each to-be-trained model is analyzed, and the to-be-trained model is allocated to give an initial training position, wherein the training position is cloud and / or edge;

[0012] Based on the initial training position, the to-be-trained model is trained, the to-be-trained model is adaptively corrected in combination with the model training index, and the training of the to-be-trained model is completed.

[0013] Further, the model attribute includes at least one of training complexity, neuron quantity, memory occupation, and computing intensity;

[0014] According to the model attribute of the to-be-trained model, the computing complexity of the to-be-trained model is considered, the computing power demand of each to-be-trained model is analyzed, and the to-be-trained model is allocated to give an initial training position, and specifically includes:

[0015] Based on the model attribute of the to-be-trained model, the computing complexity of each to-be-trained model is analyzed to give a training position score of each to-be-trained model;

[0016] In combination with the training position score of the to-be-trained model, the computing power of the training position is analyzed to give an initial training position of each to-be-trained model.

[0017] Further, based on the model attribute of the to-be-trained model, the computing complexity of each to-be-trained model is analyzed to give a training position score of each to-be-trained model, and specifically includes:

[0018] Based on the model attribute of the to-be-trained model, the to-be-trained model is analyzed layer by layer, the floating point operation number of each layer in the to-be-trained model is analyzed and accumulated, and the model operation total number of the to-be-trained model is given;

[0019] In combination with the actual processing efficiency of the cloud or the edge, the model operation total number and the ratio of the actual processing efficiency of the cloud or the edge are analyzed to give a corresponding processing time;

[0020] The memory occupation of the to-be-trained model is taken as an example, the network link bandwidth and the network link congestion coefficient in the training process are fused to give a transmission time measurement value;

[0021] Through the processing time, the transmission time measurement value, the fitness of the to-be-trained model is fused, and based on the pre-constructed position score function, a training position score of each to-be-trained model is given.

[0022] Further, in combination with the training position score of the to-be-trained model, the computing power of the training position is analyzed to give an initial training position of each to-be-trained model, and specifically represented as:

[0023] According to the order of the training position scores, in combination with the preset training threshold, the training position scores are sequentially judged;

[0024] If the training position score reaches the training threshold, the initial training position of the to-be-trained model is the cloud end;

[0025] If the training position score is less than the training threshold, the initial training position of the to-be-trained model is the edge end.

[0026] Further, based on the initial training position, the to-be-trained model is trained, and the to-be-trained model is adaptively corrected in combination with the model training index, and the training of the to-be-trained model is completed, specifically including:

[0027] Based on the initial training position, the to-be-trained model is trained, and the model training index value is given, wherein the model training index value is data reflecting the training situation of each to-be-trained model;

[0028] The model training index value is compared with the corresponding training index range, and a comparison result is given;

[0029] According to the comparison result, the to-be-trained model is adaptively corrected, and the training of the to-be-trained model is completed.

[0030] Further, the model training index value includes at least one of the training round value, the loss change rate and the gradient stability;

[0031] Based on the initial training position, the to-be-trained model is trained, and the model training index value is given, specifically including:

[0032] Based on the initial training position, the first moment estimation variable and the second moment estimation variable of the to-be-trained model are initialized, and the to-be-trained model is trained;

[0033] The change of the loss function of the to-be-trained model in the training process is analyzed, and the current loss function gradient is given;

[0034] In combination with the current loss function gradient, the current first moment estimation variable and the current second moment estimation variable are given, and the current first moment estimation variable and the current second moment estimation variable are corrected, and the first moment estimation variable correction value and the second moment estimation variable correction value are given;

[0035] According to the first moment estimation variable correction value and the second moment estimation variable correction value, the to-be-trained model is corrected, and at least one of the training round value, the loss change rate and the gradient stability is given.

[0036] Further, the loss change rate is specifically determined by the following steps:

[0037] The loss function values of each training round of the to-be-trained model in the training process are obtained;

[0038] Based on the loss function values ​​of each training epoch, the difference between the loss function values ​​of adjacent training epochs is given;

[0039] By combining the differences in the loss function values ​​of each adjacent training round, the average value of each difference is given, and the rate of change of loss is obtained.

[0040] Furthermore, the rate of change of loss is specifically expressed as:

[0041] ;

[0042] in, Let be the rate of change of loss over the most recent N training epochs, and t be the number of training epochs for the model to be trained. This represents the loss function value corresponding to the nth training epoch. This is the loss function value corresponding to the (n-1)th training round.

[0043] Furthermore, gradient stability is determined through the following steps:

[0044] Obtain the gradient values ​​of each sample in the training process of the model to be trained;

[0045] Based on the gradient values ​​of each sample, and combined with the mean gradient of all samples, the squared difference between the gradient value and the mean gradient of each sample is given.

[0046] The average value of each squared difference between the gradient value and the mean gradient of each sample is given, thus obtaining the gradient stability.

[0047] Furthermore, gradient stability is specifically expressed as:

[0048] ;

[0049] in, For the gradient stability of the model to be trained, For the sample size, For the first The gradient value of each sample. The gradient mean of all samples.

[0050] Furthermore, based on the comparison results, the model to be trained is adaptively adjusted to complete the training of the model, specifically including:

[0051] If the comparison result is within the corresponding training metric range, then based on the current model parameters in the model to be trained, stochastic gradient descent is used to optimize the model to be trained, thus completing the training of the model to be trained.

[0052] If the comparison result is not within the corresponding training index range, the adaptive matrix estimation is used to optimize the to-be-trained model, and the training of the to-be-trained model is completed.

[0053] Further, the method further comprises:

[0054] Obtaining the model training situation of the edge device and the cloud device;

[0055] According to the model training situation of the edge device and the cloud device, the training waiting time of the new model in the edge device and the cloud device is given.

[0056] Combining the model training situation of the edge device and the cloud device and the training waiting time of the new model in the edge device and the cloud device, a target function is constructed.

[0057] Based on the target function, the training position of the new model is allocated to complete the training of the new model.

[0058] Further, the target function is specifically represented as:

[0059]

[0060] Among them, is the training waiting time of the new model in the edge device, is the training time of the new model in the edge device, is the training waiting time of the new model in the cloud device, is the training time of the new model in the cloud device.

[0061] In the second aspect, the application also provides a model dynamic training device based on cloud-edge collaboration, which adopts the model dynamic training method based on cloud-edge collaboration according to any one of the above, comprising:

[0062] A data acquisition module is configured to acquire a to-be-trained model.

[0063] A position allocation module is configured to analyze the computing power demand of each to-be-trained model according to the model attribute of the to-be-trained model and taking into account the calculation complexity of the to-be-trained model, and to allocate the to-be-trained model to give an initial training position, wherein the training position is the cloud and / or the edge.

[0064] A model training module is configured to train the to-be-trained model based on the initial training position, and to perform adaptive correction on the to-be-trained model in combination with the model training index, so as to complete the training of the to-be-trained model.

[0065] In the third aspect, the application also provides a comprehensive energy management and control system, which comprises a memory, a processor, and a computer program stored in the memory, and when the computer program is run by the processor, the instructions according to the above method are executed.​

[0066] The application provides a cloud edge collaborative model dynamic training method, device and management and control system, which at least has the following beneficial effects:

[0067] (1) The training efficiency of the to-be-trained model is improved by analyzing various model attributes of the to-be-trained model, distributing the to-be-trained model, determining the initial training position and training, and then integrating adaptive correction in the model training process. The above method can dynamically adjust the training result of the to-be-trained model, update the model with poor training effect in time, and thus improve the adaptability and stability of the comprehensive energy system.

[0068] (2) The training position score is given by analyzing the computing complexity of the to-be-trained model, analyzing the computing power demand of each to-be-trained model, judging the matching relationship between the to-be-trained model and the edge end / cloud end, ensuring the matching degree of the to-be-trained model and the training position, and thus improving the utilization rate of the training resources.

[0069] (3) The training waiting time of the new model in the edge end device and the cloud end device is obtained by analyzing the model training situation of the edge end device and the cloud end device, and the training position of the new model is distributed and trained in combination with the objective function. The training ability of the cloud end device and the edge end device is measured by quantifying the training waiting time of the new model in the edge end device and the cloud end device, the training tasks of the cloud end device and the edge end device are reasonably distributed, and the resource utilization rate and operation efficiency of the comprehensive energy system are improved. BRIEF DESCRIPTION OF DRAWINGS

[0070] Figure 1 The flowchart of the cloud edge collaborative model dynamic training method provided by the embodiment of the application is provided.

[0071] Figure 2 The flowchart of giving the training position score of each to-be-trained model provided by the embodiment of the application is provided.

[0072] Figure 3 The flowchart of the model training provided by the embodiment of the application is provided.

[0073] Figure 4 The flowchart of giving the model training index value provided by the embodiment of the application is provided.

[0074] Figure 5 The flowchart of determining the loss change rate provided by the embodiment of the application is provided.

[0075] Figure 6 The flowchart of determining the gradient stability provided by the embodiment of the application is provided.

[0076] Figure 7 The flowchart of judging the to-be-trained model provided by the embodiment of the application is provided.

[0077] Figure 8 A flowchart of a new model position allocation provided for the embodiment of the present application;

[0078] Figure 9 A structural block diagram of a model dynamic training device based on cloud-edge collaboration provided for the embodiment of the present application.

[0079] Among them, 201, data acquisition module; 202, position allocation module; 203, model training module. DETAILED DESCRIPTION

[0080] In order to better understand the above technical solutions, the above technical solutions will be described in detail in the following combined with the description of the drawings and specific embodiments. Obviously, the described embodiments are only part of the embodiments of the present application, not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative labor are within the scope of protection of the present application.

[0081] The terms used in the embodiments of the present application are only for the purpose of describing the specific embodiments, and are not intended to limit the present application. The singular forms "a", "an" and "the" used in the embodiments of the present application and the appended claims are also intended to include the plural forms, unless the context clearly indicates otherwise. "Multiple" generally includes at least two.

[0082] It should also be noted that the terms "include", "contain" or any other variant thereof are intended to cover non-exclusive inclusion, so that the goods or devices including a series of elements not only include those elements, but also include other elements not explicitly listed, or include elements inherent to such goods or devices. Without more limitations, the element defined by the sentence "including a" does not exclude the presence of other identical elements in the goods or devices including the element.

[0083] In various application scenarios of the industrial internet, a large amount of data needs to be collected on site. For example, in a power distribution system, electric parameter data, environmental data, and security data need to be collected. In the field of production and manufacturing, production equipment data, energy consumption data, environmental and monitoring data need to be collected. In these applications, many of the raw data collected are process data, which have little value to users if used directly and need to be processed before they can bring value. If this data is uploaded directly to the platform, it will generate a large amount of "garbage" data, affecting the performance of the platform. In addition, in many industries such as energy and manufacturing, when a device in operation fails or needs maintenance, a timely response may be required. If the fault information is transmitted to the platform through the on-site gateway, the platform judges and processes it, and then issues an instruction to the device through the gateway for execution, the real-time performance cannot be maintained due to the network transmission delay, and the fault cannot be handled in time, which will cause the fault to expand, and even cause fatal risks. In addition, the methods of federated learning, transfer learning, and federated meta-learning have been tried in some scenarios in the industry, but there are still shortcomings such as low data generalization and high communication overhead in actual application. Especially for low-voltage electrical devices in power distribution, which belong to the scene of less feature data, large instantaneous data, and large individual differences, these methods cannot be completely applied.

[0084] Edge intelligence refers to artificial intelligence combined with edge computing, which deploys most of the computing tasks of deep learning to edge cloud rather than central cloud, thus meeting the low-latency demand of deep learning applications and ensuring the quality of service of deep learning applications, thereby realizing a win-win situation of edge computing and artificial intelligence. The development of edge intelligence has a two-way win-win advantage for edge computing and artificial intelligence: on the one hand, edge data can release potential with the help of intelligent algorithms to provide higher availability. On the other hand, edge computing can provide more data and application scenarios for intelligent algorithms.

[0085] Since the training process of a deep neural network model requires a large amount of computing and storage resources, and the computing and storage resources of an edge cloud are relatively limited and cannot be compared with those of a central cloud, in addition, edge data is single, and a model trained with single data usually has poor performance, therefore, model training by an edge cloud alone often cannot obtain a high model accuracy. Deep neural network models include generative adversarial network (GAN), graph neural networks (GNN), convolutional neural networks (CNN), long short-term memory (LSTM), BP neural network, recurrent neural network, multilayer perceptron, and long short-term memory network.

[0086] At present, the update iteration of the equipment is developing towards the direction of automation and intelligence, and many edge side devices are faced with problems such as insufficient intelligence and automation, low resource utilization, and high frequency of manual intervention.

[0087] In order to enhance the automation capability of the edge device and solve the problem of reasonable allocation of cloud edge resources, the application provides a model dynamic training method based on cloud edge cooperation. The model training is realized through edge cloud cooperation to achieve efficient model training. This way can jointly utilize the advantages of central cloud and edge cloud. By considering the training capabilities of the cloud and the edge side, the training speed of the same model by the two methods is calculated, and the shortest training time of the model is found while considering the queuing time, so as to achieve the efficiency optimization of model training. The effect of reasonable utilization of cloud edge resources is achieved, and the problem of too long model training time caused by allocation is solved.

[0088] As shown in Figure 1 The embodiment of the application provides a model dynamic training method based on cloud edge cooperation, and the specific steps are as follows:

[0089] S101: Obtain a to-be-trained model.

[0090] It can be understood that the above to-be-trained model is a model that can be trained on a cloud device or an edge device. According to the different types of models, the to-be-trained model can be a BP neural network, a convolutional neural network, a graph neural network, a long short-term memory network, a generative adversarial network, and other models. In other embodiments, it can be other models, which will not be described here.

[0091] S102: According to the model attribute of the to-be-trained model, the calculation complexity of the to-be-trained model is considered, the computing power demand of each to-be-trained model is analyzed, and the to-be-trained model is distributed to give an initial training position.

[0092] Further, according to the model attribute of the to-be-trained model, the calculation complexity of the to-be-trained model is considered, the computing power demand of each to-be-trained model is analyzed, and the to-be-trained model is distributed to give an initial training position. The training position is the cloud and / or the edge, and specifically includes:

[0093] Based on the model attribute of the to-be-trained model, the calculation complexity of each to-be-trained model is analyzed, and a training position score of each to-be-trained model is given. The model attribute includes at least one of training complexity, neuron number, memory occupation, and calculation intensity.

[0094] In combination with the training position score of the to-be-trained model, the computing power of the training position is analyzed, and an initial training position of each to-be-trained model is given.

[0095] Further, refer to Figure 2, based on the model attribute of the to-be-trained model, analyzing the calculation complexity of each to-be-trained model, and giving a training position score of each to-be-trained model, specifically including:

[0096] Based on the model attribute of the to-be-trained model, the to-be-trained model is analyzed layer by layer, the floating point operations of each layer in the to-be-trained model are analyzed and accumulated, and the total number of model operations of the to-be-trained model is given;

[0097] In combination with the actual processing efficiency of the cloud or the edge, the total number of model operations and the ratio of the actual processing efficiency of the cloud or the edge are analyzed, and the corresponding processing time is given;

[0098] With the memory occupation of the to-be-trained model, the network link bandwidth and the network link congestion coefficient in the training process are fused, and a transmission time metric value is given;

[0099] Through the processing time, the transmission time metric value, the fitness of the to-be-trained model is fused, and based on the pre-constructed position score function, the training position score of each to-be-trained model is given.

[0100] In the embodiments provided by the present application, the model attribute includes at least one of training complexity, neuron number, memory occupation, and calculation intensity. Among them, the training complexity of the model refers to the measure of the calculation resources and time required for the model to complete a complete training in the training process. It is usually used to measure the difficulty and resource consumption degree of training a model. Training complexity can be analyzed from multiple angles, including time complexity, space complexity, and data complexity. The number of neurons of the model refers to the total number of neurons, which are the basic units of the network, in an artificial neural network. Neurons are the basic units of information processing and transmission in neural networks, which are connected together through weights to form a complex network structure. The memory occupation of the model refers to the memory size occupied by the model during the training, inference or deployment of the model. The calculation intensity of the model refers to the number of floating point operations per unit of memory exchanged in the calculation process of the model, which is the ratio of the calculation amount to the memory amount of the model.

[0101] In a specific embodiment, due to the difference in model types, the training complexity of each to-be-trained model has certain differences, and the training complexity of the to-be-trained model can be quantified using floating point operations (Floating Point Operations per Second, FLOPs). FLOPs is used to describe the number of floating point operations completed by a computer, to understand the calculation demand of the model in actual application, and to select appropriate hardware to accelerate the training and inference process of the model.

[0102] Different training models include different layers, such as fully connected layers, convolutional layers, pooling layers, normalization layers, activation layers, recurrent layers, attention layers, etc. Different types of layers have different structures and calculation methods, so their FLOPs are also different. Therefore, the total FLOPs need to be calculated according to the structure of the specific model and the size of the input data to evaluate the computational complexity of the model.

[0103] For a convolutional layer, it is specifically represented as:

[0104] ;;

[0105] wherein, is the FLOPs of the convolutional layer, is the number of input channels of the convolutional layer, K is the size of the convolution kernel, is the height of the output feature map, is the width of the output feature map, is the number of output channels of the convolutional layer.

[0106] For a fully connected layer, it is specifically represented as:

[0107] ;

[0108] wherein, is the FLOPs of the fully connected layer, I is the number of input layer neurons of the fully connected layer, and O is the number of output layer neurons of the fully connected layer.

[0109] In the examples provided by the present application, for a to-be-trained model, the number and size of each fully connected layer and convolutional layer (the remaining layers are ignored) are analyzed, the FLOPs parameters of each layer are accumulated, and the total number of model operations of the to-be-trained model is given. In other examples, the layers to be analyzed are selected according to the actual situation of the to-be-trained model, and the FLOPs of all layers in the to-be-trained model can also be accumulated, and this is not limited.

[0110] The total number of model operations of the to-be-trained model is compared with the actual processing efficiency of the cloud / edge to obtain the processing time of the to-be-trained model on the cloud / edge, which is specifically represented as:

[0111] ;

[0112] wherein, is the processing time of the cloud / edge, is the total number of model operations of the to-be-trained model, is the actual processing efficiency of the cloud / edge.

[0113] It is to be understood that the actual processing efficiency of the cloud can be obtained by consulting the cloud service provider's documentation or actual performance testing, and the actual processing efficiency of the edge can be determined by testing the floating point operation performance of its processor.

[0114] The network link bandwidth and the network link congestion coefficient are fused in the training process, the mutual influence between the network link bandwidth and the network link congestion coefficient is analyzed, the ratio of the input data size to the product of the network link bandwidth and the network link congestion coefficient is calculated, and the transmission time metric value is given, which is specifically represented as:

[0115] ;

[0116] Wherein, is the transmission time metric value, D is the input data size, B is the network link bandwidth, is the network link congestion coefficient.

[0117] The network link bandwidth refers to the amount of data that can be transmitted by the network link in unit time, usually in units of bits per second. The network link bandwidth can be directly read through the management interface of network devices such as routers and switches, or measured using bandwidth testing tools. Network link bandwidth = data amount / time.

[0118] The network link congestion coefficient refers to the data flow on the network link exceeding the bandwidth of the link, causing data packets to queue for transmission, thereby increasing the delay and possible packet loss. The network link congestion coefficient can be observed by measuring the round-trip time of data packets to observe the congestion of the link. If the round-trip time increases significantly, it may indicate that the link is congested. Traffic monitoring tools can also be used to analyze the traffic and congestion on the link. Congestion control algorithms can also be used by TCP and other transport layer protocols to dynamically adjust the sending rate to avoid or alleviate congestion. Network link congestion coefficient = current traffic / network link bandwidth.

[0119] In other embodiments, network traffic monitoring software is used to monitor the traffic on the network link in real time, including upload and download rates, packet transmission delays, packet loss rates, and other key indicators. These indicators can also directly reflect the congestion status of the transmission channel.

[0120] In the embodiments provided by the present application, the training position score of each trained model is given based on the pre-constructed position score function by fusing the processing time of the edge and the transmission time metric value, and the fitness of the trained model, which is specifically represented as:

[0121] ;

[0122] Wherein, S is the training position score, is the weight parameter corresponding to the edge processing time, a weight parameter corresponding to a transmission time metric value, a weight parameter corresponding to a data privacy degree P, 、 、 an adaptability of different to-be-trained models in different scenarios, which is adjusted according to different application scenarios.

[0123] Consider what scenario the to-be-trained model is trained for, and whether the use of training data complies with the requirements of relevant laws and regulations (such as general data protection regulation (GDPR) and other data protection regulations). The data privacy degree can be divided into multiple levels, such as high, medium and low. Highly sensitive data is given a higher score (such as 8-10 out of 1-10), moderately sensitive data is given an intermediate score (4-7), and low-sensitive data is given a lower score (1-3).

[0124] Further, in combination with the training location score of the to-be-trained model, the computing power of the training location is analyzed, and the initial training location of each to-be-trained model is given, which is specifically represented as:

[0125] According to the order of the training location score, in combination with the preset training threshold, the training location score is judged in sequence;

[0126] If the training location score reaches the training threshold, the initial training location of the to-be-trained model is the cloud.

[0127] If the training location score is less than the training threshold, the initial training location of the to-be-trained model is the edge.

[0128] In a specific embodiment, the training threshold is set according to the capabilities of cloud and edge devices and actual application requirements. When the training location score is greater than or equal to the training threshold, the to-be-trained model is placed in cloud training, that is, the initial training location is the cloud, indicating that the to-be-trained model is relatively complex. When the training location score is less than the training threshold, the to-be-trained model is placed in edge training, that is, the initial training location is the edge, indicating that the to-be-trained model is relatively simple and suitable for edge training.

[0129] If there are too many to-be-trained models that need to be trained at the same time, and there is no difference in importance between the to-be-trained models, then the to-be-trained models are distributed according to the shortest time principle, that is, the to-be-trained models are sorted in ascending order according to the training location score, and then trained in the edge or cloud. If the importance of the models has levels, then the to-be-trained models in each level are trained first, and the to-be-trained models in each level are sorted in ascending order according to the training location score and queued for training.

[0130] S103: training the to-be-trained model based on the initial training position, performing self-adaptive correction on the to-be-trained model in combination with the model training index, and completing the training of the to-be-trained model.

[0131] Specifically, referring to Figure 3 , the to-be-trained model is trained based on the initial training position, and a model training index value is given, wherein the model training index value is data reflecting the training situation of each to-be-trained model;

[0132] The model training index value is compared with the corresponding training index range, and a comparison result is given.

[0133] According to the comparison result, the to-be-trained model is adaptively corrected, and the training of the to-be-trained model is completed.

[0134] In the embodiments provided by the present application, the model training index value includes at least one of a training round value, a loss change rate and a gradient stability, wherein the training round value is the current training round in the model training process. For example, there is a training data set containing 1000 samples, and the batch size is 100. Then each training round contains 10 iterations (1000 / 100 = 10). If 10 rounds of training are selected, the model will see the training data set 10 times, a total of 100 iterations (10 rounds x 10 iterations / round), and the range of the training round value is 1-10.

[0135] Further, referring to Figure 4 , the model training index value is given, specifically including:

[0136] Based on the initial training position, the first moment estimation variable and the second moment estimation variable of the to-be-trained model are initialized, and the to-be-trained model is trained;

[0137] The change of the loss function of the to-be-trained model in the training process is analyzed, and the current loss function gradient is given.

[0138] In combination with the current loss function gradient, the current first moment estimation variable and the current second moment estimation variable are given, and the current first moment estimation variable and the current second moment estimation variable are corrected, and the first moment estimation variable correction value and the second moment estimation variable correction value are given.

[0139] According to the first moment estimation variable correction value and the second moment estimation variable correction value, the to-be-trained model is corrected, and at least one of the training round value, the loss change rate and the gradient stability is given.

[0140] It is understandable that the characteristics of a model differ at different stages of model training. Therefore, different strategies are used to optimize the model at different training stages. However, determining which training stage the model is in is a problem that needs to be solved. In the implementation provided by this invention, the strategy to be adopted is determined by judging any one of the training epoch value, loss rate of change, or gradient stability.

[0141] In one specific implementation, during the initial stage of model training, the Adam strategy is used to enable the model to converge quickly, thus rapidly responding to state changes in the integrated energy system. The specific steps are as follows:

[0142] First, initialize the weights of the model to be trained. Bias, first-moment estimators, and second-moment estimators are used to initialize the exponential decay rate in the Adam policy. and and learning rate Among them, the dimensions of the first-order moment estimators and the second-order moment estimators are consistent with the dimensions of the parameters in the model to be trained, while the exponential decay rate... and As a hyperparameter, its value ranges from [0, 1]. For example, =0.9, =0.999.

[0143] Calculate the current gradient of the model to be trained during training. ,in, The gradient of the loss function corresponding to the t-th iteration of the model to be trained is the current gradient of the loss function. For gradient operators, it means that with respect to weights Take the partial derivative, where L is the loss function of the model to be trained. These are the weights corresponding to the (t-1)th iteration of the model to be trained.

[0144] Current first moment estimator variable and the current second moment estimator variables Specifically, it is expressed as:

[0145] ;

[0146] ;

[0147] in, Let be the first moment estimate of the gradient parameter in the t-th iteration, i.e., the current first moment estimate variable. This is the estimate of the first moment of the gradient parameter in the (t-1)th iteration. Let be the second moment estimate of the gradient parameter in the t-th iteration, i.e., the current second moment estimate variable. is the second moment estimation value of the gradient parameter of the t-1th iteration, is the loss function gradient corresponding to the tth iteration of the model to be trained, that is, the current loss function gradient.

[0148] First moment estimation variable correction value and second moment estimation variable correction value , specifically represented as:

[0149] ;

[0150] ;

[0151] Wherein, t is the number of iterations.

[0152] According to the first moment estimation variable correction value and the second moment estimation variable correction value, the weight in the model to be trained is updated:

[0153] ;

[0154] Wherein, is the weight of the model to be trained after the weight corresponding to the tth iteration is updated, is the update parameter, is the weight corresponding to the tth iteration of the model to be trained, is the learning rate. In this example, is a minimum value, which can be .

[0155] Further, referring to Figure 5 , the loss change rate is determined by the following steps:

[0156] Obtain the loss function value of each training round of the model to be trained in the training process;

[0157] Based on the loss function value of each training round, the difference value of the loss function value of adjacent training rounds is given;

[0158] Combined with the difference value of the loss function value of each adjacent training round, the average value of each difference value is given, and the loss change rate is obtained.

[0159] The loss change rate is specifically represented as:

[0160] ;

[0161] Wherein, is the loss change rate corresponding to the last N training rounds, t is the training round value of the model to be trained, is the loss function value corresponding to the nth training round, This is the loss function value corresponding to the (n-1)th training round.

[0162] Furthermore, referring to Figure 6 Gradient stability is determined through the following steps:

[0163] Obtain the gradient values ​​of each sample in the training process of the model to be trained;

[0164] Based on the gradient values ​​of each sample, and combined with the mean gradient of all samples, the squared difference between the gradient value and the mean gradient of each sample is given.

[0165] The average value of each squared difference between the gradient value and the mean gradient of each sample is given, thus obtaining the gradient stability.

[0166] Gradient stability is specifically expressed as:

[0167] ;

[0168] in, For the gradient stability of the model to be trained, For the sample size, For the first The gradient value of each sample. This is the gradient mean of all samples.

[0169] Furthermore, based on the comparison results, the model to be trained is adaptively adjusted to complete the training of the model, referring to... Figure 7 Specifically, it includes:

[0170] If the comparison result is within the corresponding training metric range, then based on the current model parameters in the model to be trained, stochastic gradient descent is used to optimize the model to be trained, thus completing the training of the model to be trained.

[0171] If the comparison result is not within the corresponding training metric range, adaptive moment estimation is used to optimize the model to be trained, thus completing the training of the model to be trained.

[0172] In one specific implementation, during training, it is necessary to determine whether the model has converged and stabilized in order to switch the adaptive adjustment method and thus improve model accuracy. In the first specific example, the training epoch value is used as the model training metric value, and the corresponding training metric range is based on the total number of training epochs. and preset adjustment factors Get the result and determine whether the number of training rounds has reached the total number of training rounds. and regulatory factors the product of the loss change rate and the gradient stability, if the product is reached, the to-be-trained model is optimized by using the stochastic gradient descent strategy, if the product is not reached, the to-be-trained model is continuously optimized by using the adaptive moment estimation strategy, and the training of the to-be-trained model is completed. In a third specific example, the gradient stability is used as the model training index value, if the gradient stability is within the preset corresponding training index range, the to-be-trained model is optimized by using the stochastic gradient descent strategy, if the gradient stability is not within the preset corresponding training index range, the to-be-trained model is continuously optimized by using the adaptive moment estimation strategy, and the training of the to-be-trained model is completed.

[0173] In a specific example, when the adaptive moment estimation strategy is changed to the stochastic gradient descent strategy, in order to make more accurate adjustment to the state change of the integrated energy system, the current weight inherited from the adaptive moment estimation is used to retain the learned features. The learning rate , the initial momentum , the batch size, the maximum training round and other parameters of the stochastic gradient descent strategy are initialized, wherein in the examples provided in the present application, the learning rate is set according to a certain proportion. Finally, the to-be-trained model is optimized by using the stochastic gradient descent strategy, including forward propagation, selection of a loss function according to the task type and calculation of the loss, back propagation, parameter update and termination judgment.

[0174] In another specific embodiment, before the to-be-trained model is optimized by using the adaptive moment estimation (Adam), it is necessary to judge whether the to-be-trained model has the condition of gradient vanishing or gradient explosion. If there is gradient explosion or vanishing, it indicates that the to-be-trained model cannot adapt to the state change of the integrated energy system, and needs to be adaptively adjusted. According to the judgment of the model training index value, the adaptive moment estimation and / or the stochastic gradient descent are selected to adaptively adjust. In this way, the test accuracy of the to-be-trained model can be improved, and the training time of the to-be-trained model can be reduced. By dynamically switching the optimization strategy during the training process, the effect of balancing the convergence speed and the model accuracy is achieved, which better completes the training task of the to-be-trained model and better adapts to the state change of the integrated energy system.

[0175] It can be understood that the above gradient explosion or gradient disappearance judgment is a known technology to those skilled in the art, which will not be repeated here. In one specific example, when the to-be-trained model is a regression task model, one or more of the mean square error, absolute error, and determination coefficient are used to complete the judgment of gradient explosion or gradient disappearance. When the to-be-trained model is a classification task model, the F1 score is used to complete the judgment of gradient explosion or gradient disappearance. In other examples, other ways can be used to judge gradient explosion and gradient disappearance, which are not limited.

[0176] Referring to Figure 8 , the model dynamic training method based on cloud edge collaboration further comprises:

[0177] Obtaining the model training situation of the edge device and the cloud device;

[0178] According to the model training situation of the edge device and the cloud device, the training waiting time of the new model on the edge device and the cloud device is given;

[0179] Combining the model training situation of the edge device and the cloud device and the training waiting time of the new model on the edge device and the cloud device, a target function is constructed;

[0180] Based on the target function, the training position of the new model is allocated to complete the training of the new model.

[0181] It can be understood that when multiple edge devices need to be trained on the cloud device at the same time, it will cause the cloud device to be congested or even overloaded, resulting in the inability to complete the training task well, and also causing the waste of resources of the edge device. On the basis of training of a lightweight model (which can be trained on a cloud device or an edge device), the training tasks of the cloud device and the edge device are reasonably allocated to improve the resource utilization and operation efficiency of the comprehensive energy system.

[0182] Before allocating the training tasks of the cloud device and the edge device, the training capability of the cloud device and the training capability of the edge device need to be determined, and the training capability of the cloud device and the edge device is measured by the training waiting time of the new model on the edge device and the cloud device.

[0183] For the cloud device, parallelism is an important indicator for measuring the training capability of the cloud device. Parallelism refers to the number of models that can be trained simultaneously by the cloud device or the edge device. The parallelism of the cloud device is , the time for the cloud device to train a single model is , the time for uploading data is , the time for downloading the model is , and then the total time consumed by the edge device for uploading the model to the cloud device for training is: When the cloud device is full of queues (i.e. The time that the new model needs to wait (i.e. the training waiting time) is wherein, is the number of models that the new model needs to wait for in the queue of the cloud device.

[0184] For the edge device, the parallelism of the edge device is , and the time that the edge device trains a single model is When the edge device is full of queues (i.e. The time that the new model needs to wait (i.e. the training waiting time) is wherein, is the number of models that the new model needs to wait for in the queue of the edge device. Generally, the edge device can only train one model at a time, so the parallelism of the edge device is 1, and the time that the edge device trains a single model is For different edge devices, the parallelism may be different, and is adjusted according to the actual situation.

[0185] Suppose that there are multiple models that need to be trained at the same time, and the objective function is Without considering the communication delay, a method for finding the shortest time from sending a training request to the end of training is considered for the following two cases:

[0186] Case 1: For each new model, the cloud device and the edge device are all idle, and the selection standard of the training position is that when the training position is arranged on the edge device, and when the training position is arranged on the cloud device.

[0187] Case 2: The cloud device needs to be queued, and there is still a demand for training of new models. In this case, the selection standard of the training position is that when the training position is arranged on the edge device, and when the training position is arranged on the cloud device.

[0188] In the above manner, a suitable training position is arranged for the new model to achieve optimal allocation of training resources.

[0189] In the integrated energy system, the cloud device and the edge device are mutually coordinated in the scenes of energy production and conversion, transmission and distribution, storage and management, and data collection and monitoring. The cloud device provides strong computing power and global optimization capability, and is suitable for processing complex optimization scheduling and data analysis tasks; the edge device focuses more on real-time performance, reliability and local control capability, and is suitable for processing device-level monitoring, control and fault diagnosis tasks.

[0190] For edge devices, configuring a neural network model can better handle tasks such as real-time monitoring and state assessment, fault diagnosis and early warning, demand response and energy management. Through the model trained by the edge device, the integrated energy system can realize intelligent and automated operation in terms of distributed energy access, device management, energy conversion, etc.

[0191] If the configured model can be adaptively trained, the model can dynamically adjust its parameters according to real-time data and environmental changes, reduce errors caused by the stage adaptation of the model, and better adapt to the operation state and demand of the integrated energy system, improving the real-time and accuracy of the task.

[0192] The present application proposes a model dynamic training method based on cloud edge collaboration. First, the model is analyzed by various model attributes, and the model is preliminarily allocated to the initial training location (cloud device or edge device) for training after power matching with the edge device. Then, adaptive adjustment is integrated into the model training process, and the resources of the cloud device and the edge device are optimized to achieve optimal training efficiency. The above method not only dynamically adjusts the training results of the model, updates the model in a timely manner when the training effect is poor, and improves the adaptability and stability of the integrated energy system; but also considers the resource occupation of the cloud device and the edge device, coordinates the allocation of the model, fully utilizes the computing power of the integrated energy system, and improves the system operation efficiency.

[0193] Reference Figure 9 The embodiment of the present application provides a model dynamic training device based on cloud edge collaboration, comprising:

[0194] The data acquisition module 201 is configured to acquire a to-be-trained model.

[0195] The position allocation module 202 is configured to analyze the computing power demand of each to-be-trained model according to the model attributes of the to-be-trained model and the calculation complexity of the to-be-trained model, and allocate the to-be-trained model to give an initial training location, wherein the training location is the cloud and / or the edge.

[0196] The model training module 203 is configured to train the to-be-trained model based on the initial training location, and adaptively correct the to-be-trained model in combination with the model training index, and complete the training of the to-be-trained model.

[0197] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the described module can refer to the corresponding process in the foregoing method embodiments, which will not be repeated here.

[0198] In addition, the application further provides a comprehensive energy management and control system, comprising a memory, a processor, and a computer program stored in the memory, when the computer program is executed by the processor, instructions according to the above method are executed.

[0199] It can be understood that, for the convenience and brevity of description, the specific working process of the described cloud-edge collaboration based model dynamic training device included in the comprehensive energy management and control system can refer to the corresponding process in the foregoing method embodiments, which will not be described here.

[0200] In particular, according to the embodiments of the present application, the above-mentioned process with reference to the flowchart Figure 1 The described process can be implemented as a computer software program. For example, the embodiments of the present application include a computer program product comprising a computer program carried on a machine-readable medium, the computer program containing program code for executing the method shown in the flowchart. In such embodiments, the computer program can be downloaded and installed from the network by the communication part, and / or installed from the detachable medium. When the computer program is executed by the central processing unit, the above-mentioned functions defined in the device of the present application are executed.

[0201] Note that the computer readable medium described above can be a computer readable signal medium or a computer readable storage medium or any combination thereof. The computer readable storage medium can be, for example, but not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus or device, or any suitable combination of the foregoing. More specific examples of the computer readable storage medium can include, but are not limited to, an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing. In the disclosure, the computer readable storage medium can be any tangible medium that contains or stores a program used by or in connection with an instruction execution system, apparatus or device. In the disclosure, the computer readable signal medium can include a computer readable program code propagated on or through a computer readable medium, in baseband or as part of a carrier wave. The computer readable signal medium can take a variety of forms, including but not limited to, electro-magnetic, optical, or any suitable combination of the foregoing. The computer readable signal medium can also be any computer readable medium that is not a computer readable storage medium and that can communicate, propagate or transport a program for use by or in connection with an instruction execution system, apparatus or device. Program code embodied on a computer readable medium can be transmitted using any appropriate medium, including but not limited to wireless, wire line, optical fiber cable, RF, etc., or any suitable combination of the foregoing.

[0202] The computer readable medium described above can be included in the electronic device described above; alternatively, the computer readable medium can exist as a separate entity in which the electronic device is incorporated.

[0203] Computer program code for carrying out operations of the disclosure can be written in any combination of one or more programming languages, including an object oriented programming language such as Java, Smalltalk, C++ or the like and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The program code can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider).

[0204] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0205] The units described in the embodiments of this disclosure can be implemented in software or hardware. The names of the units are not, in some cases, intended to limit the specific unit.

[0206] Although preferred embodiments of the invention have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including both the preferred embodiments and all changes and modifications falling within the scope of the invention. Clearly, those skilled in the art can make various alterations and modifications to the invention without departing from its spirit and scope. Thus, if these modifications and variations of the invention fall within the scope of the claims and their equivalents, the invention is also intended to include these modifications and variations.

Claims

1. A method for dynamic training of a model based on cloud-edge collaboration, characterized in that, The method comprises the following steps: acquiring a to-be-trained model; analyzing the to-be-trained model layer by layer based on the model attribute of the to-be-trained model, accumulating the floating point operation numbers of each layer in the to-be-trained model, and giving the model operation total number of the to-be-trained model; combining the actual processing efficiency of the cloud or the edge, analyzing the ratio of the model operation total number and the actual processing efficiency of the cloud or the edge, and giving the corresponding processing time; giving the transmission time measurement value by fusing the memory occupation of the to-be-trained model, the network link bandwidth in the training process and the network link congestion coefficient; fusing the fitness of the to-be-trained model by the processing time and the transmission time measurement value, and giving the training position score of each to-be-trained model based on the pre-constructed position scoring function, wherein the model attribute comprises at least one of training complexity, neuron quantity, memory occupation and calculation intensity; combining the training position score of the to-be-trained model to analyze the computing power of the training position, and giving the initial training position of each to-be-trained model, wherein the training position is the cloud and / or the edge; training the to-be-trained model based on the initial training position, adaptively correcting the to-be-trained model by combining the model training index, and completing the training of the to-be-trained model; after giving the initial training position of each to-be-trained model, further comprising: acquiring the model training situation of the edge device and the cloud device, wherein the model training situation comprises device parallelism, the number of models waiting for training and the training time of a single model; giving the minimum queuing model number according to the device parallelism and the number of models waiting for training of the edge device; giving the training waiting time of the new model in the edge device based on the minimum queuing model number and the training time of a single model in the edge device; giving the minimum queuing model number according to the device parallelism and the number of models waiting for training of the cloud device; giving the training waiting time of the new model in the cloud device based on the minimum queuing model number and the training time of a single model in the cloud device; constructing a target function by combining the model training situation of the edge device and the cloud device and the training waiting time of the new model in the edge device and the cloud device; and distributing the training position of the new model based on the target function to complete the training of the new model. 2.The cloud edge collaboration based model dynamic training method of claim 1, wherein, combining the training position score of the to-be-trained model to analyze the computing power of the training position, and giving the initial training position of each to-be-trained model, specifically represented as: judging the training position score in turn according to the order of the training position score and combining the preset training threshold; if the training position score reaches the training threshold, the initial training position of the to-be-trained model is the cloud; if the training position score is less than the training threshold, the initial training position of the to-be-trained model is the edge. 3.The cloud edge collaboration based model dynamic training method of claim 1, wherein, training the to-be-trained model based on the initial training position, adaptively correcting the to-be-trained model by combining the model training index, and completing the training of the to-be-trained model, specifically comprising: training the to-be-trained model based on the initial training position, and giving the model training index value, wherein the model training index value is data reflecting the training situation of each to-be-trained model; comparing the model training index value with the corresponding training index range to give a comparison result; According to the comparison result, the to-be-trained model is adaptively corrected, and the training of the to-be-trained model is completed. 4.The cloud-edge collaboration based model dynamic training method of claim 3, wherein, The model training index value includes at least one of a training round value, a loss change rate and a gradient stability. Based on the initial training position, the to-be-trained model is trained to give a model training index value, specifically including: Based on the initial training position, the first moment estimation variable and the second moment estimation variable of the to-be-trained model are initialized, and the to-be-trained model is trained; The change of the loss function of the to-be-trained model in the training process is analyzed to give the current loss function gradient; Combined with the current loss function gradient, the current first moment estimation variable and the current second moment estimation variable are given, and the current first moment estimation variable and the current second moment estimation variable are corrected to give the first moment estimation variable correction value and the second moment estimation variable correction value; According to the first moment estimation variable correction value and the second moment estimation variable correction value, the to-be-trained model is corrected, and at least one of the training round value, the loss change rate and the gradient stability is given. 5.The cloud edge collaboration based model dynamic training method of claim 4, wherein, The loss change rate is determined by the following steps: Obtain the loss function value of each training round of the to-be-trained model in the training process; Based on the loss function value of each training round, the difference value of the loss function value of adjacent training rounds is given; Combined with the difference value of each adjacent training round loss function value, the average value of each difference value is given to obtain the loss change rate. 6.The cloud edge collaboration based model dynamic training method of claim 4, wherein, The gradient stability is determined by the following steps: Obtain the gradient value of each sample of the to-be-trained model in the training process; Based on the gradient value of each sample, combined with the gradient mean value of all samples, the square difference of the gradient value of each sample and the gradient mean value is given; According to the square difference of the gradient value of each sample and the gradient mean value, the average value of each square difference is given to obtain the gradient stability.

7. The cloud edge collaboration based model dynamic training method of claim 3, wherein, According to the comparison result, the to-be-trained model is adaptively corrected, and the training of the to-be-trained model is completed, specifically including: If the comparison result is within the corresponding training index range, the to-be-trained model is optimized based on the current model parameter in the to-be-trained model using stochastic gradient descent to complete the training of the to-be-trained model; If the comparison result is not within the corresponding training index range, the to-be-trained model is optimized using adaptive moment estimation to complete the training of the to-be-trained model. 8.The cloud edge collaboration based model dynamic training method of claim 1, wherein, The objective function is specifically represented as: ; wherein, is the target function, is the training latency of the new model at the edge device, is the training time of the new model at the edge device, is the training latency of the new model at the cloud device, is the training time of the new model at the cloud device. 9.The cloud-edge collaboration based model dynamic training method of claim 8, wherein, Based on the objective function, the training position of the new model is allocated to complete the training of the new model, specifically including: Based on the training waiting time of the edge device and the cloud device in the objective function, the queuing situation of the edge device and the cloud device is given; According to the queuing situation of the edge device and the cloud device, the training time of the edge device is compared with the training time of the cloud device or the total cloud time to determine the training position of the new model and complete the training of the new model, wherein the total cloud time is the sum of the training waiting time and the training time of the new model in the cloud device.

10. A cloud edge collaboration based model dynamic training apparatus, characterized in that, The cloud edge collaborative based model dynamic training method of any one of claims 1-9, comprising: a data acquisition module for acquiring a to-be-trained model; The position distribution module is configured to analyze the to-be-trained models layer by layer based on model attributes of the to-be-trained models, accumulate floating-point operations of each layer in the to-be-trained models, and give model operation total number of the to-be-trained models; analyze the model operation total number and a ratio of the model operation total number to actual processing efficiency of the cloud or the edge, and give corresponding processing time; give a transmission time measurement value by fusing memory occupation of the to-be-trained models, network link bandwidth in a training process, and a network link congestion coefficient; give training position scores of the to-be-trained models based on a pre-constructed position scoring function by fusing the processing time, the transmission time measurement value, and fitness of the to-be-trained models, wherein the model attributes include at least one of training complexity, neuron quantity, memory occupation, and calculation intensity; analyze computing power of the training positions by combining the training position scores of the to-be-trained models, and give initial training positions of the to-be-trained models, wherein the training positions are the cloud and / or the edge. The model training module is configured to train the to-be-trained models based on the initial training positions, adaptively correct the to-be-trained models based on model training indexes, and complete training of the to-be-trained models. After the initial training positions of the to-be-trained models are given, the model training module further includes: acquiring model training conditions of the edge device and the cloud device, wherein the model training conditions include device parallelism, number of models waiting for training, and training time of a single model; giving a minimum queuing model number according to the device parallelism and the number of models waiting for training of the edge device; giving training waiting time of a new model in the edge device based on the minimum queuing model number and the training time of the single model in the edge device; giving a minimum queuing model number according to the device parallelism and the number of models waiting for training of the cloud device; giving training waiting time of a new model in the cloud device based on the minimum queuing model number and the training time of the single model in the cloud device; constructing an objective function based on the model training conditions of the edge device and the cloud device and the training waiting time of the new model in the edge device and the cloud device; and distributing training positions of the new model based on the objective function to complete training of the new model.

11. An integrated energy management system, characterized by, The computer program is run by the processor to execute instructions of the method according to any one of claims 1-9.

Citation Information

Patent Citations

  • Edge device extension deployment method and system based on cloud edge collaboration

    CN117997904A

  • Teaching behavior analysis system and method based on computing power network

    CN117114932A

  • Intelligent computing power distribution method and service system based on cloud edge collaboration

    CN118860675A