Model dynamic training method and device based on cloud edge collaboration and management and control system

By analyzing model attributes and computing complexity, dynamically allocating training locations at the cloud and edge ends, and performing adaptive corrections, the problem of underutilization of cloud and edge device resources is solved, and the model training efficiency and system adaptability are improved.

CN120509463AActive Publication Date: 2025-08-19STATE GRID JIANGSU ECONOMIC RES INST +1

Patent Information

Application Number
CN202511000734.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-21
Publication Date
2025-08-19
Estimated Expiration
2045-07-21

AI Technical Summary

Technical Problem

In the prior art, the computing resources of cloud and edge devices are not fully utilized, resulting in low model training efficiency and ineffective training time.

Method used

By analyzing the attributes and computational complexity of the model to be trained, dynamically allocating the training position, and adaptively correcting it with model training indicators, collaborative training on the cloud and edge side can be achieved.

Benefits of technology

The efficiency and resource utilization of model training are improved, and the adaptability and stability of the integrated energy system are enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120509463A_ABST
    Figure CN120509463A_ABST
Patent Text Reader

Abstract

The invention discloses a dynamic model training method and device based on cloud edge collaboration and a management and control system. The method comprises the steps of obtaining a to-be-trained model; according to the model attributes of the to-be-trained models, calculating complexity of the to-be-trained models is considered, calculating power requirements of the to-be-trained models are analyzed, the to-be-trained models are distributed, and initial training positions are given; and training the to-be-trained model based on the initial training position, and performing adaptive correction on the to-be-trained model in combination with the model training index to complete training of the to-be-trained model. Various model attributes of the to-be-trained model are analyzed, the to-be-trained model is distributed, the initial training position is determined and trained, and then adaptive correction is fused in the model training process, so that the training efficiency of the to-be-trained model is improved. And meanwhile, the training result of the to-be-trained model can be dynamically adjusted, and the model with a poor training effect is updated in time, so that the adaptability and the stability of the comprehensive energy system are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of edge collaboration of industrial Internet platforms, and specifically relates to a model dynamic training method, device and management and control system based on cloud-edge collaboration. Background Art

[0002] Cloud-based model training relies on powerful computing resources and vast amounts of data, enabling the production of large, high-performance models. However, these models are typically large, requiring extensive computing and communication resources to run. Furthermore, centralized model training faces increasing pressure on storage, model parameter caching, and computational costs.

[0003] Edge computing significantly reduces latency and bandwidth usage by migrating data processing to edge nodes closer to the data source, improving system efficiency and user experience. Edge model training allows for local data processing and model updates, reducing the risk of privacy leaks and enabling more personalized services. However, the limited computing and storage capabilities of edge devices make it difficult to independently train large-scale models.

[0004] Currently, it is difficult for the cloud and edge to fully utilize their respective computing resources through collaboration. The fundamental reason is that the lack of unified standards and platforms has led to the fragmentation of the cloud-edge ecosystem, the resource scheduling and operation and maintenance systems are still imperfect and cannot flexibly distribute the load, the resources on both ends cannot be well utilized, and the training efficiency of the model is not taken into consideration.

[0005] To overcome the limitations of single-cloud or edge training, researchers have proposed an architecture that combines cloud and edge model training. This architecture offloads some tasks from the cloud to the edge, enabling a more private, real-time, and personalized user experience.

[0006] Patent application CN117997904A discloses a method and system for expanding and deploying edge devices based on cloud-edge collaboration. The system analyzes the task processing logs of all edge devices under the cloud platform, identifies the target edge device where a task processing overflow event has occurred, and determines the overflow task information of the target edge device. Based on the location information of the target edge device within the cloud platform's corresponding cloud network, the system determines an edge device cluster capable of docking and processing the overflow tasks of the target edge device, selects an auxiliary edge device from the cluster to dock and process the overflow tasks, and implements expanded deployment processing of the overflow tasks of the target edge device, fully utilizing the edge computing power of the cloud platform. The system also returns the normal processing results of the overflow tasks to the target edge device, and then integrates the non-overflow task processing results of the target edge device with the normal processing results of the received overflow tasks to obtain a complete task processing result, effectively and timely processing the overflow tasks of the edge device. The existing technology places all tasks on the edge side for training, which is edge-to-edge collaborative training. The cloud only plays a command role, collecting the operating status and overflow data of each edge device and allocating the overflow data to idle edge devices for training, which does not fully utilize the computing power resources of the cloud.

[0007] How to make full use of cloud and edge computing resources, shorten model training time, and improve model training efficiency is a problem that needs to be solved at present. Summary of the Invention

[0008] In response to the defects existing in the above-mentioned prior art, the present invention provides a model dynamic training method, device and management system based on cloud-edge collaboration, the method comprising: obtaining the model to be trained; according to the model attributes of the model to be trained, taking into account the computational complexity of the model to be trained, analyzing the computing power requirements of each model to be trained, and allocating the model to be trained, and giving an initial training position; based on the initial training position, training the model to be trained, combining the model training indicators, adaptively correcting the model to be trained, and completing the training of the model to be trained. By analyzing the various model attributes of the model to be trained, allocating the model to be trained, determining the initial training position and training, and then incorporating adaptive correction into the model training process, the training efficiency of the model to be trained is improved. At the same time, the training results of the model to be trained can be dynamically adjusted, and the model with poor training effect can be updated in time, thereby improving the adaptability and stability of the integrated energy system.

[0009] In a first aspect, the present invention provides a model dynamic training method based on cloud-edge collaboration, which specifically includes the following steps: Get the model to be trained; Based on the model attributes of the models to be trained and taking into account the computational complexity of the models to be trained, the computing power requirements of each model to be trained are analyzed, and the models to be trained are allocated and the initial training locations are given, where the training locations are the cloud and / or the edge. Based on the initial training position, the model to be trained is trained, and combined with the model training indicators, the model to be trained is adaptively corrected to complete the training of the model to be trained.

[0010] Furthermore, the model attributes include at least one of training complexity, number of neurons, memory usage, and computational intensity; Based on the model attributes of the models to be trained and the computational complexity of the models to be trained, the computing power requirements of each model to be trained are analyzed, and the models to be trained are allocated and the initial training positions are given. Specifically, the following are included: Based on the model properties of the models to be trained, the computational complexity of each model to be trained is analyzed, and a training position score for each model to be trained is given; Combined with the training position scores of the models to be trained, the computing power of the training positions is analyzed to give the initial training positions of each model to be trained.

[0011] Furthermore, based on the model properties of the models to be trained, the computational complexity of each model to be trained is analyzed, and a training position score for each model to be trained is given, specifically including: Based on the model properties of the model to be trained, the model to be trained is analyzed layer by layer, the floating-point operations of each layer in the model to be trained are analyzed and accumulated, and the total number of model operations of the model to be trained is given; Combined with the actual processing efficiency of the cloud or edge, analyze the total number of model operations and the ratio of the actual processing efficiency of the cloud or edge to provide the corresponding processing time; The transmission time measurement value is given by combining the memory usage of the model to be trained, the network link bandwidth during the training process, and the network link congestion coefficient. By combining the processing time and transmission time metrics, the fitness of the model to be trained is integrated, and based on the pre-built position scoring function, the training position score of each model to be trained is given.

[0012] Furthermore, the computing power of the training position is analyzed in combination with the training position score of the model to be trained, and the initial training position of each model to be trained is given, which is specifically expressed as: According to the order of training position scores and the preset training threshold, the training position scores are judged in turn; If the training location score reaches the training threshold, the initial training location of the model to be trained is the cloud; If the training position score is less than the training threshold, the initial training position of the model to be trained is the edge.

[0013] Furthermore, based on the initial training position, the model to be trained is trained, and in combination with the model training indicators, the model to be trained is adaptively corrected to complete the training of the model to be trained, specifically including: Based on the initial training position, the to-be-trained model is trained and a model training index value is given, wherein the model training index value is data reflecting the training status of each to-be-trained model; Compare the model training index value with the corresponding training index range and give the comparison result; According to the comparison results, the model to be trained is adaptively corrected to complete the training of the model to be trained.

[0014] Furthermore, the model training indicator value includes at least one of a training round value, a loss change rate, and a gradient stability; Based on the initial training position, the model to be trained is trained and the model training index values are given, including: Based on the initial training position, the first-order moment estimation variables and the second-order moment estimation variables of the model to be trained are initialized, and the model to be trained is trained; Analyze the changes in the loss function of the model to be trained during the training process and give the current loss function gradient; Combined with the current loss function gradient, the current first-order moment estimation variable and the current second-order moment estimation variable are given, and the current first-order moment estimation variable and the current second-order moment estimation variable are corrected to give the first-order moment estimation variable correction value and the second-order moment estimation variable correction value; The model to be trained is corrected according to the first-order moment estimation variable correction value and the second-order moment estimation variable correction value, and at least one of the training round value, the loss change rate and the gradient stability is given.

[0015] Furthermore, the loss change rate is determined by the following steps: Obtain the loss function value of each training round of the training process of the model to be trained; Based on the loss function value of each training round, the difference between the loss function values of adjacent training rounds is given; Combining the differences in the loss function values of each adjacent training round, the average of each difference is given to obtain the loss change rate.

[0016] Furthermore, the loss change rate is specifically expressed as: ; in, is the loss change rate corresponding to the last N training rounds, t is the training round value of the model to be trained, is the loss function value corresponding to the nth training round, is the loss function value corresponding to the n-1th training round.

[0017] Furthermore, the gradient stability is determined by the following steps: Obtain the gradient value of each sample during the training process of the model to be trained; Based on the gradient value of each sample and the gradient mean of all samples, the square difference between the gradient value of each sample and the gradient mean is given; According to the square difference between the gradient value of each sample and the gradient mean, the average value of each square difference is given to obtain the gradient stability.

[0018] Furthermore, gradient stability is specifically expressed as: ; in, is the gradient stability of the model to be trained, is the sample size, For the The gradient value of the sample, is the mean gradient of all samples.

[0019] Furthermore, based on the comparison results, the model to be trained is adaptively modified to complete the training of the model to be trained, specifically including: If the comparison result is within the corresponding training index range, the model to be trained is optimized using stochastic gradient descent based on the current model parameters in the model to be trained, completing the training of the model to be trained; If the comparison result is not within the corresponding training index range, the adaptive moment estimation is used to optimize the model to be trained to complete the training of the model to be trained.

[0020] Furthermore, the method further comprises: Obtain model training status for edge devices and cloud devices; Based on the model training status of edge devices and cloud devices, the training waiting time of the new model on edge devices and cloud devices is given; Build an objective function based on the model training status of edge devices and cloud devices, as well as the training waiting time of the new model on edge devices and cloud devices. Based on the objective function, the training position of the new model is allocated to complete the training of the new model.

[0021] Furthermore, the objective function is specifically expressed as: ;; in, The waiting time for training the new model on the edge device, is the training time of the new model on the edge device, Waiting time for training new models on cloud devices, The training time of the new model on the cloud device.

[0022] In a second aspect, the present invention further provides a model dynamic training device based on cloud-edge collaboration, which adopts any of the above-mentioned model dynamic training methods based on cloud-edge collaboration, including: Data acquisition module, used to obtain the model to be trained; A location allocation module is used to analyze the computing power requirements of each model to be trained based on the model attributes of the model to be trained and the computational complexity of the model to be trained, and to allocate the models to be trained and provide initial training locations, where the training locations are cloud and / or edge; The model training module is used to train the model to be trained based on the initial training position, and to adaptively correct the model to be trained in combination with the model training indicators to complete the training of the model to be trained.

[0023] In a third aspect, the present invention further provides an integrated energy management and control system, comprising a memory, a processor, and a computer program stored in the memory, wherein when the computer program is run by the processor, the computer program executes instructions according to the above method.

[0024] The present invention provides a model dynamic training method, device, and management system based on cloud-edge collaboration, which have at least the following beneficial effects: (1) By analyzing the various model attributes of the models to be trained, allocating the models to be trained, determining the initial training position and conducting training, and then incorporating adaptive corrections into the model training process, the training efficiency of the models to be trained can be improved. The above method can dynamically adjust the training results of the models to be trained, and timely update the models with poor training results, thereby improving the adaptability and stability of the integrated energy system.

[0025] (2) By analyzing the computational complexity of the model to be trained, the computing power requirements of each model to be trained are analyzed, and a training location score is given. The matching relationship between the model to be trained and the edge / cloud is judged, ensuring the matching degree between the model to be trained and the training location, thereby improving the utilization rate of training resources.

[0026] (3) By analyzing the model training status of edge devices and cloud devices, the training waiting time of the new model on edge devices and cloud devices is obtained, and the training location of the new model is allocated and trained in combination with the objective function. By quantifying the training waiting time of the new model on edge devices and cloud devices, the training capabilities of cloud devices and edge devices are measured, and the reasonable allocation of training tasks for cloud devices and edge devices is achieved, thereby improving the resource utilization and operating efficiency of the integrated energy system. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] Figure 1A flowchart of a model dynamic training method based on cloud-edge collaboration provided by an embodiment of the present invention; Figure 2 A flowchart for providing a training position score for each model to be trained provided in an embodiment of the present invention; Figure 3 A flowchart of the model training provided by an embodiment of the present invention; Figure 4 A flowchart of providing model training index values provided by an embodiment of the present invention; Figure 5 A flow chart for determining the loss change rate provided by an embodiment of the present invention; Figure 6 A flow chart for determining gradient stability provided by an embodiment of the present invention; Figure 7 A flowchart for determining a model to be trained provided by an embodiment of the present invention; Figure 8 A flowchart of new model position allocation provided by an embodiment of the present invention; Figure 9 This is a structural block diagram of the model dynamic training device based on cloud-edge collaboration provided in an embodiment of the present invention.

[0028] Among them, 201 is a data acquisition module; 202 is a location allocation module; 203 is a model training module. DETAILED DESCRIPTION

[0029] To better understand the above technical solution, the following will be described in detail with reference to the accompanying drawings and specific implementation methods. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0030] The terms used in the embodiments of the present invention are for the purpose of describing specific embodiments only and are not intended to limit the present invention. The singular forms "a," "an," "the," and "the" used in the embodiments of the present invention and the appended claims are also intended to include plural forms, and unless the context clearly indicates otherwise, "a plurality" generally includes at least two.

[0031] It should also be noted that the terms "include," "comprises," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a product or device comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such product or device. In the absence of further limitations, an element defined by the phrase "comprises a..." does not exclude the presence of other identical elements in the product or device comprising the element.

[0032] Across various Industrial Internet applications, a vast amount of data needs to be collected on-site. For example, in power distribution systems, electrical parameters, environmental data, and security data are required. In manufacturing, data on production equipment, energy usage, and environmental and monitoring data are collected. Much of the raw data collected in these applications is process data, which offers little value to users directly and requires computational processing. Directly uploading this data to the platform generates a large amount of "junk" data, impacting platform performance. Furthermore, in many industries, such as energy and manufacturing, equipment failures or maintenance requirements may require a prompt response. If fault information is transmitted to the platform via an on-site gateway, where the platform identifies and processes the information and issues instructions, which are then transmitted to the device via the gateway for execution, network transmission latency prevents real-time performance. Failure to address the fault in a timely manner can lead to escalating problems and even fatal risks. In addition, the currently proposed methods such as federated learning, transfer learning, and federated meta-learning have been tried in some scenarios in the industrial sector, but in actual applications, there are still shortcomings such as low data generalization and high communication overhead. This is especially true for equipment such as medium and low voltage electrical appliances in power distribution, which have few feature data, large amounts of instantaneous data, and large individual differences, and cannot be fully applied.

[0033] Edge intelligence refers to artificial intelligence (AI) enabled by edge computing. This approach deploys the majority of deep learning application computing tasks to the edge cloud rather than the central cloud. This not only meets the low-latency requirements of deep learning applications but also ensures the quality of service for these applications, thus achieving a win-win situation for edge computing and AI. The development of edge intelligence offers mutually beneficial benefits for both edge computing and AI: On the one hand, edge data can be leveraged by intelligent algorithms to unlock its potential and provide higher availability. On the other hand, edge computing can provide intelligent algorithms with more data and application scenarios.

[0034] Training deep neural network models requires extensive computing and storage resources, which are relatively limited in edge clouds and cannot match those of central clouds. Furthermore, edge data is often monotonous, and models trained using this monotonous data typically perform poorly. Therefore, training models on the edge cloud alone often cannot achieve high model accuracy. Deep neural network models include generative adversarial networks (GANs), graph neural networks (GNNs), convolutional neural networks (CNNs), long short-term memory networks (LSTMs), backpropagation neural networks, recurrent neural networks, multilayer perceptrons, and long short-term memory networks.

[0035] At this stage, equipment updates and iterations are all moving towards automation and intelligence. However, many edge devices face problems such as insufficient intelligence and automation, low resource utilization, and excessive frequency of manual intervention.

[0036] In order to enhance the automation capability of edge devices and solve the problem of reasonable allocation of cloud-edge resources, the present invention provides a model dynamic training method based on cloud-edge collaboration, which realizes efficient model training through edge-cloud collaboration. This method can jointly utilize the advantages of central cloud and edge cloud. By taking into account the training capabilities of both the cloud and edge sides, the training speed of the same model treated by the two methods is calculated, and the shortest training time of the model is found while considering the queuing time, the efficiency of model training is optimized. The effect of reasonable utilization of cloud-edge resources is achieved, and the problem of excessively long model training time caused by allocation is solved.

[0037] like Figure 1 As shown, an embodiment of the present invention provides a model dynamic training method based on cloud-edge collaboration, and the specific steps are as follows: S101: Obtain the model to be trained.

[0038] It is understood that the above-mentioned model to be trained is a model that can be trained on a cloud device or an edge device. Depending on the type of model, the model to be trained can be a BP neural network, a convolutional neural network, a graph neural network, a long short-term memory network, a generative adversarial network, and other models. In other embodiments, it can be other models, which are not described here.

[0039] S102: Analyze the computing power requirements of each model to be trained based on the model attributes of the model to be trained and the computational complexity of the model to be trained, allocate the models to be trained, and provide an initial training position.

[0040] Furthermore, based on the model attributes of the models to be trained and the computational complexity of the models to be trained, the computing power requirements of each model to be trained are analyzed, and the models to be trained are allocated and the initial training locations are given. The training locations can be cloud-based and / or edge-based, specifically including: Analyzing the computational complexity of each model to be trained based on the model attributes of the model to be trained, and providing a training position score for each model to be trained, wherein the model attributes include at least one of training complexity, number of neurons, memory usage, and computational intensity; Combined with the training position scores of the models to be trained, the computing power of the training positions is analyzed to give the initial training positions of each model to be trained.

[0041] Further, refer to Figure 2 Based on the model properties of the models to be trained, the computational complexity of each model to be trained is analyzed, and the training position score of each model to be trained is given, including: Based on the model properties of the model to be trained, the model to be trained is analyzed layer by layer, the floating-point operations of each layer in the model to be trained are analyzed and accumulated, and the total number of model operations of the model to be trained is given; Combined with the actual processing efficiency of the cloud or edge, analyze the total number of model operations and the ratio of the actual processing efficiency of the cloud or edge to provide the corresponding processing time; The transmission time measurement value is given by combining the memory usage of the model to be trained, the network link bandwidth during the training process, and the network link congestion coefficient. By combining the processing time and transmission time metrics, the fitness of the model to be trained is integrated, and based on the pre-built position scoring function, the training position score of each model to be trained is given.

[0042] In the embodiment provided by the present invention, the model attributes include at least one of training complexity, number of neurons, memory usage, and computational intensity. Among them, the training complexity of the model refers to the measure of the computing resources and time required for the model to complete a complete training during the training process. It is usually used to measure the difficulty and resource consumption of training a model. Training complexity can be analyzed from multiple angles, including time complexity, space complexity, and data complexity. The number of neurons in the model refers to the total number of neurons, the basic units that constitute the network in the artificial neural network. Neurons are the basic units for information processing and transmission in the neural network. They are connected together through weights to form a complex network structure. The memory usage of the model refers to the size of the memory occupied by the model during the training, inference or deployment of the model. The computational intensity of the model refers to the number of floating-point operations corresponding to each unit of memory exchange during the calculation process of the model, which is the ratio of the model's calculation amount to the memory access amount.

[0043] In one specific implementation, the training complexity of each model varies due to different model types. This complexity can be quantified using floating-point operations per second (FLOPs). FLOPs describes the number of floating-point operations a computer performs. This helps understand the computational requirements of a model in practical applications and select appropriate hardware to accelerate model training and inference.

[0044] Different training models include different layers, such as fully connected layers, convolutional layers, pooling layers, normalization layers, activation layers, recurrent layers, and attention layers. Different types of layers have different structures and computational methods, and therefore different FLOPs. The total FLOPs must be calculated based on the specific model structure and input data size to assess the model's computational complexity.

[0045] For the convolutional layer, it is specifically expressed as: ;; in, is the FLOPs of the convolutional layer, is the number of input channels of the convolution layer, K is the convolution kernel size, is the height of the output feature map, is the width of the output feature map, is the number of output channels of the convolutional layer.

[0046] For the fully connected layer, it is specifically expressed as: ; in, is the FLOPs of the fully connected layer, I is the number of neurons in the input layer of the fully connected layer, and O is the number of neurons in the output layer of the fully connected layer.

[0047] In the examples provided herein, the number and size of each fully connected layer and convolutional layer (other layers are ignored) in the model being trained are analyzed, and the FLOPs parameters of each layer are accumulated to provide the total number of model operations for the model being trained. In other examples, the layers to be analyzed can be selected based on the actual conditions of the model being trained, and the FLOPs of all layers in the model being trained can also be accumulated, without limitation.

[0048] The processing time of the model to be trained on the cloud / edge is obtained by calculating the ratio of the total number of model operations of the model to be trained to the actual processing efficiency of the cloud / edge. It can be expressed as: ; in, is the processing time of the cloud / edge, is the total number of model operations for the model to be trained, The actual processing efficiency of the cloud / edge.

[0049] It is important to understand that the actual processing efficiency of the cloud can be obtained by consulting the cloud service provider's documentation or actual performance testing, and the actual processing efficiency of the edge can be determined by testing the floating-point operation performance of its processor, etc.

[0050] During the training process, the network link bandwidth and the network link congestion coefficient are integrated to analyze the mutual influence between the network link bandwidth and the network link congestion coefficient. The ratio of the input data size to the product of the network link bandwidth and the network link congestion coefficient is calculated to give the transmission time measurement value, which is specifically expressed as: ; in, is the transmission time measurement value, D is the input data size, B is the network link bandwidth, is the network link congestion coefficient.

[0051] Network link bandwidth refers to the amount of data a network link can transmit per unit of time, typically measured in bits per second. Network link bandwidth can be read directly from the management interface of network devices (such as routers and switches) or measured using bandwidth testing tools. Network link bandwidth = data volume / time.

[0052] The network link congestion factor indicates that the data traffic on a network link exceeds the link's bandwidth, causing packets to queue for transmission, increasing latency and potentially causing packet loss. The network link congestion factor can be used to monitor link congestion by measuring the round-trip time of packets. A significant increase in round-trip time may indicate link congestion. Traffic monitoring tools can also be used to analyze link traffic and congestion. Transport layer protocols such as TCP can also use congestion control algorithms to dynamically adjust the sending rate to avoid or alleviate congestion. The network link congestion factor = current traffic / network link bandwidth.

[0053] In other implementations, network traffic monitoring software is used to monitor the traffic of the network link in real time, including key indicators such as upload and download rates, data packet transmission delay, packet loss rate, etc. These indicators can also intuitively reflect the congestion status of the transmission channel.

[0054] In the embodiment provided by the present invention, the fitness of the model to be trained is integrated by measuring the processing time and transmission time of the edge, and based on the pre-built position scoring function, a training position score of each model to be trained is given, which is specifically expressed as: ; Among them, S is the training position score, is the weight parameter corresponding to the edge processing time, is the weight parameter corresponding to the transmission time metric, is the weight parameter corresponding to the data privacy degree P, 、 、 To adjust the adaptability of different models to be trained in different scenarios according to different application scenarios.

[0055] Consider the scenarios in which the model will be trained and whether the use of the training data complies with relevant laws and regulations (such as the General Data Protection Regulation (GDPR)). Data privacy can be categorized into multiple levels, such as high, medium, and low. Highly sensitive data can be assigned a higher score (e.g., 8-10 on a scale of 1-10), moderately sensitive data can be assigned an intermediate score (4-7), and low-sensitivity data can be assigned a lower score (1-3).

[0056] Furthermore, the computing power of the training position is analyzed in combination with the training position score of the model to be trained, and the initial training position of each model to be trained is given, which is specifically expressed as: According to the order of training position scores and the preset training threshold, the training position scores are judged in turn; If the training location score reaches the training threshold, the initial training location of the model to be trained is the cloud; If the training position score is less than the training threshold, the initial training position of the model to be trained is the edge.

[0057] In a specific implementation, a training threshold is set according to the capabilities of the cloud-edge devices and actual application requirements. When the training location score is greater than or equal to the training threshold, the model to be trained is placed in the cloud for training, that is, the initial training location is the cloud, indicating that the model to be trained is relatively complex; when the training location score is less than the training threshold, the model to be trained is placed in the edge for training, that is, the initial training location is the edge, indicating that the model to be trained is relatively simple and suitable for edge training.

[0058] If there are too many models to be trained simultaneously, and there's no hierarchy among the models, they're assigned based on the shortest time possible. That is, they're sorted in ascending order by their training location score and sent to the edge or cloud for training. If the models are ranked by importance, they're categorized by importance level, training models in higher-importance categories first. Within each importance level, they're sorted in ascending order by their training location score and queued for training.

[0059] S103: Based on the initial training position, the model to be trained is trained, and the model to be trained is adaptively corrected in combination with the model training index to complete the training of the model to be trained.

[0060] Specifically, refer to Figure 3 , based on the initial training position, train the model to be trained and provide a model training index value, wherein the model training index value is data reflecting the training status of each model to be trained; Compare the model training index value with the corresponding training index range and give the comparison result; According to the comparison results, the model to be trained is adaptively corrected to complete the training of the model to be trained.

[0061] In the embodiments provided herein, model training metrics include at least one of a training round value, a loss change rate, and gradient stability. The training round value represents the current training round in the model training process. For example, consider a training dataset containing 1000 examples with a batch size of 100. Each training round consists of 10 iterations (1000 / 100 = 10). If 10 rounds of training are selected, the model will see the training dataset 10 times, for a total of 100 iterations (10 rounds x 10 iterations / round). The training round value ranges from 1 to 10.

[0062] Further, refer to Figure 4 , giving the model training index values, including: Based on the initial training position, the first-order moment estimation variables and the second-order moment estimation variables of the model to be trained are initialized, and the model to be trained is trained; Analyze the changes in the loss function of the model to be trained during the training process and give the current loss function gradient; Combined with the current loss function gradient, the current first-order moment estimation variable and the current second-order moment estimation variable are given, and the current first-order moment estimation variable and the current second-order moment estimation variable are corrected to give the first-order moment estimation variable correction value and the second-order moment estimation variable correction value; The model to be trained is corrected according to the first-order moment estimation variable correction value and the second-order moment estimation variable correction value, and at least one of the training round value, the loss change rate and the gradient stability is given.

[0063] It is understandable that at different stages of model training, the characteristics of the model are different. Therefore, different strategies are used to optimize the model at different training stages of the model. However, how to determine at which training stage the model is is a problem that needs to be solved. In the embodiment provided by the present invention, the strategy to be adopted is determined by judging any one of the training round value, loss change rate, or gradient stability.

[0064] In a specific implementation, in the initial stage of model training, the Adam strategy is used to enable the trained model to converge quickly so as to quickly respond to changes in the state of the integrated energy system. The specific steps are as follows: First, initialize the weights of the model to be trained , bias, first-order moment estimation variable and second-order moment estimation variable, initialize the exponential decay rate in the Adam strategy and and the learning rate Among them, the dimensions of the first-order moment estimation variables and the second-order moment estimation variables are consistent with the dimensions of the parameters in the model to be trained, and the exponential decay rate and As a hyperparameter, its value range is [0, 1]. For example, =0.9, =0.999.

[0065] Calculate the current gradient of the model to be trained during the training process ,in, is the loss function gradient corresponding to the t-th iteration of the model to be trained, that is, the current loss function gradient, is the gradient operator, which represents the weight Find the partial derivative, L is the loss function of the model to be trained, is the weight corresponding to the t-1th iteration of the model to be trained.

[0066] Current first-order moment estimation variable and the current second moment estimation variable , specifically expressed as: ; ; in, is the first-order moment estimate of the gradient parameter of the t-th iteration, that is, the current first-order moment estimate variable, is the first-order moment estimate of the gradient parameter at the t-1th iteration, is the second-order moment estimate of the gradient parameter of the t-th iteration, that is, the current second-order moment estimate variable, is the second-order moment estimate of the gradient parameter at the t-1th iteration, The gradient of the loss function corresponding to the tth iteration of the model to be trained is the current gradient of the loss function.

[0067] First-order moment estimate variable correction value and the second-order moment estimate variable correction value , specifically expressed as: ; ; Where t is the number of iterations.

[0068] Estimating variable correction values based on first-order moments And the second-order moment estimation variable correction value is used to update the weights in the training model: ; in, is the weight of the model to be trained after the weight corresponding to the tth iteration is updated, To update the parameters, is the weight corresponding to the tth iteration of the model to be trained, is the learning rate. In this example, is a minimum value, which can be .

[0069] Further, refer to Figure 5 , the loss change rate is determined by the following steps: Obtain the loss function value of each training round of the training process of the model to be trained; Based on the loss function value of each training round, the difference between the loss function values of adjacent training rounds is given; Combining the differences in the loss function values of each adjacent training round, the average of each difference is given to obtain the loss change rate.

[0070] The loss change rate is specifically expressed as: ; in, is the loss change rate corresponding to the last N training rounds, t is the training round value of the model to be trained, is the loss function value corresponding to the nth training round, is the loss function value corresponding to the n-1th training round.

[0071] Further, refer to Figure 6 , gradient stability is determined by the following steps: Obtain the gradient value of each sample during the training process of the model to be trained; Based on the gradient value of each sample and the gradient mean of all samples, the square difference between the gradient value of each sample and the gradient mean is given; According to the square difference between the gradient value of each sample and the gradient mean, the average value of each square difference is given to obtain the gradient stability.

[0072] Gradient stability, specifically expressed as: ; in, is the gradient stability of the model to be trained, is the sample size, For the The gradient value of the sample, is the mean gradient of all samples.

[0073] Further, according to the comparison results, the training model is adaptively modified to complete the training of the training model. Figure 7 , specifically including: If the comparison result is within the corresponding training index range, the model to be trained is optimized using stochastic gradient descent based on the current model parameters in the model to be trained, completing the training of the model to be trained; If the comparison result is not within the corresponding training index range, the adaptive moment estimation is used to optimize the model to be trained to complete the training of the model to be trained.

[0074] In a specific implementation, during the training process, it is necessary to determine whether the model has converged and stabilized, so as to switch the adaptive adjustment mode and thus improve the model accuracy. In the first specific example, the training round value is used as the model training index value, and the corresponding training index range is based on the total training rounds. and preset adjustment factors Get, judge whether the training round value reaches the total training round and regulatory factors The product of , if it is reached, the stochastic gradient descent strategy is used to optimize the model to be trained; if it is not reached, the adaptive moment estimation strategy is continued to be used to optimize the model to be trained, and the training of the model to be trained is completed. In the second specific example, the loss change rate is used as the model training index value. If the loss change rate is within the preset corresponding training index range, the stochastic gradient descent strategy is used to optimize the model to be trained; if the loss change rate is not within the preset corresponding training index range, the adaptive moment estimation strategy is continued to be used to optimize the model to be trained, and the training of the model to be trained is completed. In the third specific example, gradient stability is used as the model training index value. If the gradient stability is within the preset corresponding training index range, the stochastic gradient descent strategy is used to optimize the model to be trained; if the gradient stability is not within the preset corresponding training index range, the adaptive moment estimation strategy is continued to be used to optimize the model to be trained, and the training of the model to be trained is completed.

[0075] In a specific example, when changing from the adaptive moment estimation strategy to the stochastic gradient descent strategy, in order to make more accurate adjustments to the state changes of the integrated energy system, the current weights obtained by the adaptive moment estimation are inherited. , retain the learned features. Then initialize the learning rate of the stochastic gradient descent strategy , initial momentum , batch size, maximum training rounds and other parameters, among which, in the example provided by the present invention, the learning rate according to Finally, the stochastic gradient descent strategy is used to optimize the training model, including forward propagation, selecting a loss function based on the task type and calculating the loss, backpropagation, parameter update, and termination judgment.

[0076] In another specific embodiment, before using adaptive moment estimation (Adam) to optimize the model to be trained, it is necessary to first determine whether the model to be trained has gradient vanishing or gradient exploding. If there is gradient explosion or vanishing, it indicates that the model to be trained cannot adapt to the state changes of the integrated energy system and needs to be adaptively adjusted. Based on the judgment of the model training index value, adaptive moment estimation and / or stochastic gradient descent are selected for adaptive adjustment. In this way, the test accuracy of the model to be trained can be improved, and the training time of the model to be trained can be reduced. By dynamically switching the optimization strategy during the training process, the effect of balancing the convergence speed and model accuracy is achieved, and while better completing the training task of the model to be trained, it can also better adapt to the state changes of the integrated energy system.

[0077] It is understood that the above-mentioned determination of gradient explosion or gradient disappearance is a technical content well known to those skilled in the art and will not be elaborated here. In a specific example, when the model to be trained is a regression task model, one or more of the mean square error, absolute error, and determination coefficient are used to complete the determination of gradient explosion or gradient disappearance. When the model to be trained is a classification task model, the F1 score is used to complete the determination of gradient explosion or gradient disappearance. In other examples, other methods can be used to determine gradient explosion and gradient disappearance, which is not limited to this.

[0078] Reference Figure 8 , the model dynamic training method based on cloud-edge collaboration also includes: Obtain model training status for edge devices and cloud devices; Based on the model training status of edge devices and cloud devices, the training waiting time of the new model on edge devices and cloud devices is given; Build an objective function based on the model training status of edge devices and cloud devices, as well as the training waiting time of the new model on edge devices and cloud devices. Based on the objective function, the training position of the new model is allocated to complete the training of the new model.

[0079] Understandably, when multiple edge devices need to train simultaneously on cloud devices, this can cause congestion or even overload on the cloud devices, preventing them from completing training tasks properly and wasting resources on the edge devices. By training lightweight models (which can be trained on both cloud and edge devices), we can rationally allocate training tasks between cloud and edge devices, improving resource utilization and operational efficiency of the integrated energy system.

[0080] Before allocating training tasks to cloud devices and edge devices, it is necessary to first clarify the training capabilities of cloud devices and edge devices, and measure the training capabilities of cloud devices and edge devices by the training waiting time of new models on edge devices and cloud devices.

[0081] For cloud devices, parallelism is an important indicator to measure the training capability of cloud devices. Parallelism refers to the number of models that can be trained simultaneously on cloud devices or edge devices. , the time it takes for a cloud device to train a single model is , the time to upload data is , the time to download the model is , then the total time consumed when the edge device model is uploaded to the cloud device for training is: When the cloud device is fully loaded (i.e. Each team needs to queue up), the time the new model needs to wait (i.e., training waiting time) is ,in, The number of models waiting to be trained on cloud devices for new models.

[0082] For edge devices, the parallelism of edge devices is , the time it takes for an edge device to train a single model is When the edge device is fully loaded (i.e. Each team needs to queue up), the time the new model needs to wait (i.e., training waiting time) is ,in, The number of models waiting to be trained in the queue of edge devices. Generally, an edge device can only train one model at a time, so the parallelism of the edge device is 1, and the time to train a single model is The degree of parallelism may vary for different edge devices and should be adjusted based on actual conditions.

[0083] Assume that there are multiple models to be trained at the same time, and the objective function is Without considering the communication delay, we can find a method to minimize the time from sending a training request to the end of training. Consider the following two cases: Case 1: For each new model, when both the cloud device and the edge device are idle, the training location selection criteria is: When the training position is arranged on the edge device, When training, the training location is arranged on the cloud device.

[0084] Case 2: Cloud devices need to be queued, and there are still new models that need to be trained. In this case, the selection criteria for the training location are: When the training position is arranged on the edge device, When training, the training location is arranged on the cloud device.

[0085] Through the above method, a suitable training location is arranged for the new model to achieve optimal allocation of training resources.

[0086] In integrated energy systems, cloud and edge devices collaborate across scenarios such as energy production and conversion, transmission and distribution, storage and management, and data collection and monitoring. Cloud devices offer powerful computing power and global optimization capabilities, making them suitable for complex optimization scheduling and data analysis tasks. Edge devices, on the other hand, prioritize real-time performance, reliability, and local control, making them suitable for device-level monitoring, control, and fault diagnosis.

[0087] For edge devices, neural network models can better handle tasks such as real-time monitoring and status assessment, fault diagnosis and early warning, demand response, and energy management. Using models trained on edge devices, integrated energy systems can achieve intelligent and automated operation in distributed energy access, device management, and energy conversion.

[0088] If the configured model is capable of adaptive training, the model can dynamically adjust its own parameters according to real-time data and environmental changes, reducing errors caused by the model's stage-by-stage adaptation problems, thereby better adapting to the operating status and needs of the integrated energy system and improving the real-time and accuracy of tasks.

[0089] This paper proposes a dynamic model training method based on cloud-edge collaboration. This method first analyzes the model's various attributes, matches the computing power of edge devices, and then initially allocates the model to an initial training location (either a cloud device or an edge device) for training. Furthermore, adaptive adjustments are incorporated into the model training process, and optimal training efficiency is achieved through rational optimization based on the resources of cloud and edge devices. This method not only dynamically adjusts model training results and promptly updates models with poor training results, thereby improving the adaptability and stability of the integrated energy system, but also comprehensively considers the resource usage of cloud and edge devices, coordinates the allocation of models, fully utilizes the computing power of the integrated energy system, and improves system operational efficiency.

[0090] Reference Figure 9 , an embodiment of the present invention provides a model dynamic training device based on cloud-edge collaboration, comprising: Data acquisition module 201, used to obtain the model to be trained; A location allocation module 202 is configured to analyze the computing power requirements of each model to be trained based on the model attributes of the model to be trained and the computational complexity of the model to be trained, and allocate the models to be trained to provide initial training locations, where the training locations are the cloud and / or the edge. The model training module 203 is used to train the model to be trained based on the initial training position, and to adaptively correct the model to be trained in combination with the model training indicators to complete the training of the model to be trained.

[0091] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the described module can refer to the corresponding process in the aforementioned method embodiment, and will not be repeated here.

[0092] In addition, the present invention also provides an integrated energy management and control system, including a memory, a processor, and a computer program stored in the memory. When the computer program is run by the processor, it executes instructions according to the above method.

[0093] It can be understood that for the convenience and conciseness of description, the specific working process of the model dynamic training device based on cloud-edge collaboration included in the integrated energy management and control system can refer to the corresponding process in the aforementioned method embodiment and will not be repeated here.

[0094] In particular, according to the embodiment of the present application, the above reference flow chart Figure 1 The described processes can be implemented as computer software programs. For example, embodiments of the present application include a computer program product comprising a computer program carried on a machine-readable medium, the computer program containing program code for executing the method shown in the flowchart. In such embodiments, the computer program can be downloaded and installed from a network via a communication component and / or installed from a removable medium. When the computer program is executed by a central processing unit, the above-described functions defined in the apparatus of the present application are performed.

[0095] It should be noted that the computer-readable medium mentioned above in the present disclosure may be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or component, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, device, or component. In the present disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any suitable medium, including but not limited to wires, optical cables, RF (radio frequency), etc., or any suitable combination thereof.

[0096] The computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device.

[0097] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages, or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on the remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).

[0098] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the module, program segment, or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of the boxes in the block diagram and / or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0099] The units described in the embodiments of the present disclosure may be implemented in software or hardware, wherein the name of a unit does not necessarily limit the unit itself.

[0100] Although preferred embodiments of the present invention have been described, those skilled in the art may make additional changes and modifications to these embodiments once they are aware of the basic inventive concepts. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the invention. Obviously, those skilled in the art may make various changes and modifications to the present invention without departing from the spirit and scope of the invention. Thus, the present invention is intended to include such changes and modifications as fall within the scope of the claims and their equivalents.

Claims

1. A model dynamic training method based on cloud-edge collaboration, characterized in that: include: Get the model to be trained; Based on the model attributes of the models to be trained and taking into account the computational complexity of the models to be trained, the computing power requirements of each model to be trained are analyzed, and the models to be trained are allocated and the initial training locations are given, where the training locations are the cloud and / or the edge. Based on the initial training position, the model to be trained is trained, and combined with the model training indicators, the model to be trained is adaptively corrected to complete the training of the model to be trained.

2. The model dynamic training method based on cloud-edge collaboration according to claim 1 is characterized in that: Model attributes include at least one of training complexity, number of neurons, memory usage, and computational intensity; Based on the model attributes of the models to be trained and the computational complexity of the models to be trained, the computing power requirements of each model to be trained are analyzed, and the models to be trained are allocated and the initial training positions are given. Specifically, the following are included: Based on the model properties of the models to be trained, the computational complexity of each model to be trained is analyzed, and a training position score for each model to be trained is given; Combined with the training position scores of the models to be trained, the computing power of the training positions is analyzed to give the initial training positions of each model to be trained.

3. The model dynamic training method based on cloud-edge collaboration according to claim 2 is characterized in that: Based on the model properties of the models to be trained, the computational complexity of each model to be trained is analyzed, and a training position score for each model to be trained is given, specifically including: Based on the model properties of the model to be trained, the model to be trained is analyzed layer by layer, the floating-point operations of each layer in the model to be trained are analyzed and accumulated, and the total number of model operations of the model to be trained is given; Combined with the actual processing efficiency of the cloud or edge, analyze the total number of model operations and the ratio of the actual processing efficiency of the cloud or edge to provide the corresponding processing time; The transmission time measurement value is given by combining the memory usage of the model to be trained, the network link bandwidth during the training process, and the network link congestion coefficient. By combining the processing time and transmission time metrics, the fitness of the model to be trained is integrated, and based on the pre-built position scoring function, the training position score of each model to be trained is given.

4. The model dynamic training method based on cloud-edge collaboration according to claim 2 is characterized in that Combined with the training position scores of the models to be trained, the computing power of the training positions is analyzed to give the initial training positions of each model to be trained, which can be expressed as follows: According to the order of training position scores and the preset training threshold, the training position scores are judged in turn; If the training location score reaches the training threshold, the initial training location of the model to be trained is the cloud; If the training position score is less than the training threshold, the initial training position of the model to be trained is the edge.

5. The model dynamic training method based on cloud-edge collaboration according to claim 1 is characterized in that Based on the initial training position, the model to be trained is trained. Combined with the model training indicators, the model to be trained is adaptively corrected to complete the training of the model to be trained. Specifically, the following steps are performed: Based on the initial training position, the to-be-trained model is trained and a model training index value is given, wherein the model training index value is data reflecting the training status of each to-be-trained model; Compare the model training index value with the corresponding training index range and give the comparison result; According to the comparison results, the model to be trained is adaptively corrected to complete the training of the model to be trained.

6. The model dynamic training method based on cloud-edge collaboration according to claim 5 is characterized in that: The model training indicator value includes at least one of a training round value, a loss change rate, and a gradient stability; Based on the initial training position, the model to be trained is trained and the model training index values are given, including: Based on the initial training position, the first-order moment estimation variables and the second-order moment estimation variables of the model to be trained are initialized, and the model to be trained is trained; Analyze the changes in the loss function of the model to be trained during the training process and give the current loss function gradient; Combined with the current loss function gradient, the current first-order moment estimation variable and the current second-order moment estimation variable are given, and the current first-order moment estimation variable and the current second-order moment estimation variable are corrected to give the first-order moment estimation variable correction value and the second-order moment estimation variable correction value; The model to be trained is corrected according to the first-order moment estimation variable correction value and the second-order moment estimation variable correction value, and at least one of the training round value, the loss change rate and the gradient stability is given.

7. The model dynamic training method based on cloud-edge collaboration according to claim 6 is characterized in that: The loss change rate is determined by the following steps: Obtain the loss function value of each training round of the training process of the model to be trained; Based on the loss function value of each training round, the difference between the loss function values of adjacent training rounds is given; Combining the differences in the loss function values of each adjacent training round, the average of each difference is given to obtain the loss change rate.

8. The model dynamic training method based on cloud-edge collaboration according to claim 6 is characterized in that: Gradient stability is determined by the following steps: Obtain the gradient value of each sample during the training process of the model to be trained; Based on the gradient value of each sample and the gradient mean of all samples, the square difference between the gradient value of each sample and the gradient mean is given; According to the square difference between the gradient value of each sample and the gradient mean, the average value of each square difference is given to obtain the gradient stability.

9. The model dynamic training method based on cloud-edge collaboration according to claim 5, characterized in that: Based on the comparison results, the model to be trained is adaptively modified to complete the training of the model to be trained, specifically including: If the comparison result is within the corresponding training index range, the model to be trained is optimized using stochastic gradient descent based on the current model parameters in the model to be trained, completing the training of the model to be trained; If the comparison result is not within the corresponding training index range, the adaptive moment estimation is used to optimize the model to be trained to complete the training of the model to be trained.

10. The model dynamic training method based on cloud-edge collaboration according to claim 1, characterized in that: After allocating the training model and giving the initial training position, it also includes: Obtain model training status for edge devices and cloud devices; Based on the model training status of edge devices and cloud devices, the training waiting time of the new model on edge devices and cloud devices is given; Build an objective function based on the model training status of edge devices and cloud devices, as well as the training waiting time of the new model on edge devices and cloud devices. Based on the objective function, the training position of the new model is allocated to complete the training of the new model.

11. The model dynamic training method based on cloud-edge collaboration according to claim 10, characterized in that: Model training status, including device parallelism, number of models waiting to be trained, and training time for a single model; Based on the model training status of edge devices and cloud devices, the training waiting time of the new model on edge devices and cloud devices is given, including: Given the minimum number of queued models based on the device parallelism of the edge devices and the number of models waiting to be trained; Based on the minimum number of queued models and the training time of a single model on the edge device, the training waiting time of the new model on the edge device is given; Given the minimum number of queued models based on the device parallelism of the cloud devices and the number of models waiting to be trained; Based on the minimum number of queued models and the training time of a single model on the cloud device, the training waiting time of the new model on the cloud device is given.

12. The model dynamic training method based on cloud-edge collaboration according to claim 10, characterized in that: The objective function is specifically expressed as: ; in, is the objective function, The waiting time for training the new model on the edge device, is the training time of the new model on the edge device, Waiting time for training new models on cloud devices, The training time of the new model on the cloud device.

13. The model dynamic training method based on cloud-edge collaboration according to claim 10, characterized in that: Based on the objective function, the training location of the new model is allocated to complete the training of the new model, specifically including: Based on the training waiting time of edge devices and cloud devices in the objective function, the queue status of edge devices and cloud devices is given; Based on the queuing status of edge devices and cloud devices, the training time on the edge devices is compared with the training time on the cloud devices or the total cloud time to determine the training position of the new model and complete the training of the new model. The total cloud time is the sum of the training waiting time and training time of the new model on the cloud devices.

14. A model dynamic training device based on cloud-edge collaboration, characterized in that: The method for dynamic model training based on cloud-edge collaboration according to any one of claims 1 to 13 is adopted, comprising: Data acquisition module, used to obtain the model to be trained; A location allocation module is used to analyze the computing power requirements of each model to be trained based on the model attributes of the model to be trained and the computational complexity of the model to be trained, and to allocate the models to be trained and provide initial training locations, where the training locations are cloud and / or edge; The model training module is used to train the model to be trained based on the initial training position, and to adaptively correct the model to be trained in combination with the model training indicators to complete the training of the model to be trained.

15. An integrated energy management and control system, characterized in that: The invention comprises a memory, a processor, and a computer program stored in the memory, wherein when the computer program is run by the processor, the computer program executes the instructions of the method according to any one of claims 1 to 13.

Citation Information

Patent Citations

  • Edge device extension deployment method and system based on cloud edge collaboration

    CN117997904A

  • Fault diagnosis optimization method and electronic equipment based on edge-cloud collaborative task offloading

    CN114936708A

  • Teaching behavior analysis system and method based on computing power network

    CN117114932A

  • Intelligent computing power distribution method and service system based on cloud edge collaboration

    CN118860675A

  • Cloud edge cooperative control method

    CN119561945A

Cited By

  • Federal learning-based cross-regional flood risk prediction management system and method

    CN121257865A