Electric power station cloud edge model evolution method and related device

By adopting the cloud-edge model evolution method of the power station in the power station, and using the coordinated work of the module selector and the cloud-side large model, the side equipment is split and deployed, which solves the problems of model update lag, high cost and low performance in the existing technology, and achieves continuous optimization and adaptation of model performance.

CN120163197APending Publication Date: 2025-06-17CHINA ELECTRIC POWER RESEARCH INSTITUTE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510236581.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2025-06-17

AI Technical Summary

Technical Problem

The existing cloud-side learning and side-side learning paradigms cannot update the model in time when dealing with highly variable application environments, resulting in a degradation in model performance, high computing and communication costs, low generalization performance and accuracy, and long delay in model service response.

Method used

A cloud-side model evolution method of power stations is proposed. By splitting and deploying the edge devices based on the trained module selector and cloud-side big model, the edge devices are trained using end-to-end algorithms to dynamically update the module parameters of the cloud-side big model to achieve continuous optimization and adaptation of the model.

Benefits of technology

It effectively improves the model performance on the side equipment, reduces the computing and communication costs, and realizes the timely response to highly variable environmental needs of model performance, while ensuring data privacy and saving communication overhead.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120163197A_ABST
    Figure CN120163197A_ABST
Patent Text Reader

Abstract

The invention belongs to a model evolution method, and provides an electric power station cloud edge model evolution method and a related device for solving the technical problems that the calculation and communication cost of a cloud side learning normal form is high, data features in a new application environment cannot be updated to a model in time, the generalization performance and precision of an edge side learning normal form are low, and the model service response delay is long. And based on the trained module selector and the cloud side large model, carrying out sub-model splitting on given side equipment, issuing and deploying the sub-models for the given side equipment, training the issued and deployed sub-models, aggregating parameters of the sub-models, updating module parameters of the cloud side large model, and then carrying out splitting and issuing and deploying again. According to the method, the cloud side large model can dynamically update knowledge of different sides to ensure that global data distribution covers a side environment with dynamically changed height, so that the model performance can be ensured to timely cope with the environment requirement of height change.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to a model evolution method, and particularly relates to a cloud-edge model evolution method for power stations and related devices. Background Art

[0002] To support artificial intelligence applications in power scenarios driven by deep neural networks, current practices mainly adopt two learning paradigms: cloud-side learning and edge-side learning. However, neither the cloud-side learning paradigm nor the edge-side learning paradigm can effectively handle highly variable application environments and will inevitably be affected by the degradation of model performance. In the cloud-side learning paradigm, when new application environments are encountered at the edge side, the cloud side needs to update the model. However, before the cloud-side model is deployed to the edge side, it is trained based on historical data, so it is impossible to update the data features in the new application environment to the model in a timely manner. At the same time, this paradigm also incurs high computing and communication costs. In the edge-side learning paradigm, although the newly collected data can be used to update the model locally, due to the sparse and biased training data in the edge-side application environment, the generalization performance and accuracy of the model are not satisfactory. In addition, resource competition between edge-side model training and inference may lead to an extended delay in model service response. Summary of the Invention

[0003] This application aims at the technical problems of high computing and communication costs in the cloud-side learning paradigm, inability to update data features in the new application environment to the model in a timely manner, low generalization performance and accuracy in the edge-side learning paradigm, and long delay in model service response, and provides a cloud-edge model evolution method for power stations and related devices.

[0004] To achieve the above object, this application adopts the following technical solutions: In the first aspect, this application proposes a cloud-edge model evolution method for power stations, including: Based on the trained module selector and cloud-side large model, perform sub-model splitting for a given edge device; wherein, the sub-model can minimize the loss of the local data set under resource constraints; the training methods of the module selector and the cloud-side large model include: For all module layers, construct a unified module selector for determining the activated modules of all module layers; determine the activated sub-modules of each reusable module in the cloud-side large model through the module selector, and determine the expression of the cloud-side large model according to the input and output of the activated sub-modules; train the module selector and the cloud-side large model through an end-to-end algorithm; Deploy the sub-model to the given edge device; After training the deployed sub-model, aggregate the parameters of each sub-model and update the module parameters of the cloud-side large model; Based on the module selector and the updated large model on the cloud side, perform sub-model splitting for a given edge device; Deploy the sub-model to the given edge device again.

[0005] Furthermore, the expression of the unified module selector includes:

[0006] wherein, is the layer module selector, is the hyperparameter of the layer module selector, is the feature extracted from the input , is the embedding network for extracting the features of the input , is the hyperparameter of the embedding network, is the total number of module layers; The expression of the large model on the cloud side includes:

[0007] wherein, is the expression of the large model on the cloud side, is the L reusable module, is the L -1 reusable module, is the first reusable module.

[0008] Furthermore, the loss function used when training the module selector and the large model on the cloud side through the end-to-end algorithm includes cross-entropy training loss and module load balancing loss.

[0009] Furthermore, the method for training the module selector and the large model on the cloud side through the end-to-end algorithm includes: Define application-specific subtasks according to the root cause behind the non-independent and identically distributed data distribution on the edge device; Formulate the subtask recognition process as a constrained linear programming problem and maximize the product of the elements of the subtask mapping matrix to determine the model objective of the subtask; the constraints include: making the load less than the maximum number of tasks that the module can withstand; making the maximum number of sub-modules that a subtask can activate less than the maximum number of available modules; Train the module selector and the large model on the cloud side through the end-to-end algorithm based on the model objective of the subtask.

[0010] Furthermore, the method for performing sub-model splitting for a given edge device includes: Define the importance metric of the sub-module based on the output of the module selector, and estimate the resource overhead of the candidate sub-model using the edge-side resource constraints captured by the edge-side resource analyzer of the given edge-side device; Combine the importance metric of the sub-module and the resource overhead of the candidate sub-model, and fit the sub-model for the given edge-side device through a constrained optimization algorithm.

[0011] Further, the method for training the sub-model deployed after distribution includes: Let the sub-model deployed after distribution perform inference runs on different samples; Evaluate the samples according to the information entropy of the probability distribution output after the inference runs, and select the samples with information entropy less than the preset requirements as training samples; Fine-tune the sub-model with the training samples.

[0012] Further, the method for aggregating the parameters of each sub-model includes: .

[0013] Among them, is the parameter of the updated sub-module i, is the set of sub-models containing sub-module i, is the module importance score of the given edge-side device, is the sub-module i of the k th sub-model parameter, is the parameter of sub-module i, is the data collected by the edge-side device.

[0014] In a second aspect, the present application proposes a power station cloud-edge model evolution system, including: A first splitting module, configured to perform sub-model splitting for a given edge-side device based on the trained module selector and the cloud-side large model; wherein, the sub-model can minimize the loss of the local data set under resource constraints; the training methods of the module selector and the cloud-side large model include: Activate all module layers, construct a unified module selector for determining the activated modules of all module layers; determine the activated sub-modules of each reusable module in the cloud-side large model through the module selector, and determine the expression of the cloud-side large model according to the input and output of the activated sub-modules; train the module selector and the cloud-side large model through an end-to-end algorithm; A first deployment module, configured to deploy and distribute the sub-model for the given edge-side device; An update module, configured to aggregate the parameters of each sub-model after training the sub-model deployed after distribution, and update the module parameters of the cloud-side large model; The second splitting module is used to perform sub-model splitting for a given edge device based on the module selector and the updated cloud-side large model; The second deployment module is used to deploy the sub-model to the given edge device again.

[0015] In a third aspect, the present application provides an electronic device, including: a memory and one or more processors; the memory is coupled to the processor; wherein, computer program code is stored in the memory, and the computer program code includes computer instructions, when the computer instructions are executed by the processor, the electronic device executes the steps of the above-mentioned power station cloud-edge model evolution method.

[0016] In a fourth aspect, the present application provides a computer-readable storage medium, in which a computer program is stored, and when the computer program is executed by a processor, the steps of the above-mentioned power station cloud-edge model evolution method are implemented.

[0017] Compared with the prior art, the present application has the following beneficial effects: The present application provides a power station cloud-edge model evolution method. Based on the trained module selector and the cloud-side large model, sub-model splitting is performed for a given edge device. After the sub-model is deployed to the given edge device, the deployed sub-model is trained, the parameters of each sub-model are aggregated, and the module parameters of the cloud-side large model are updated. Then, splitting and deployment are performed again. In the application environment of the present application, the local data distribution is learned by the sub-models deployed on the edge side, and the learned new knowledge can be dynamically aggregated to the cloud side. The cloud-side large model can dynamically update the knowledge of different edge sides, ensuring that the global data distribution covers the highly dynamic edge environment, so as to ensure that the model performance can timely meet the highly variable environmental requirements. Compared with the traditional cloud-edge large model collaboration method, the present application does not need to transmit and collect edge-side data to the cloud side, thus ensuring the privacy of edge-side data and greatly saving the cloud-edge data transmission overhead. Furthermore, compared with distributed machine learning methods such as federated learning, while ensuring data privacy, the present application does not need to transmit and update the training gradients of each node in real time, and can further save communication overhead.

[0018] The present application also provides a power station cloud-edge model evolution system, an electronic device and a computer storage medium, which have all the advantages of the above-mentioned power station cloud-edge model evolution method. Description of the Drawings

[0019] To more clearly illustrate the technical solutions of the embodiments of the present application, the accompanying drawings required for the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present application and should not be regarded as limiting the scope. For those of ordinary skill in the art, without creative efforts, other related drawings can also be obtained based on these drawings.

[0020] Figure 1 It is the first flowchart of the cloud-edge model evolution method for the power station of the present application; Figure 2 It is the second flowchart of the cloud-edge model evolution method for the power station of the present application; Figure 3 It is the schematic diagram of the large model on the cloud side in this embodiment; Figure 4 It is a schematic diagram of the cloud-edge model evolution system for the power station of the present application. Detailed implementation manners

[0021] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. Usually, the components of the embodiments of the present application described and illustrated in the drawings here can be arranged and designed in various different configurations.

[0022] Therefore, the following detailed description of the embodiments of the present application provided in the drawings is not intended to limit the scope of the present application claimed, but merely represents the selected embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application.

[0023] It should be noted that similar reference numerals and letters denote similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings.

[0024] In the description of the embodiments of the present application, it should be noted that if terms such as "upper", "lower", "horizontal", "inner", etc. indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings, or the orientation or positional relationship in which the invention product is usually placed during use, it is only for the convenience of describing the present application and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of the present application. In addition, terms such as "first", "second", etc. are only used for distinguishing descriptions and cannot be understood as indicating or implying relative importance.

[0025] In addition, when the term "horizontal" appears, it does not mean that the component is required to be absolutely horizontal, but it can be slightly inclined. For example, "horizontal" only means that its direction is more horizontal relative to "vertical", and does not mean that the structure must be completely horizontal, but it can be slightly inclined.

[0026] In the description of the embodiments of the present application, it should also be noted that unless otherwise clearly specified and limited, when terms such as "set", "installed", "connected", and "coupled" appear, they should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected, or indirectly connected through an intermediate medium, and it can be the communication inside two components. For those of ordinary skill in the art, the specific meanings of the above terms in the present application can be understood according to specific situations.

[0027] In recent years, with the rapid development of artificial intelligence technology, especially deep learning technology, artificial intelligence applications represented by directions such as computer vision, natural language processing, and speech recognition have been widely applied and developed in power production scenarios.

[0028] To support artificial intelligence applications in power scenarios driven by deep neural networks, current practices mainly adopt two learning paradigms: cloud-side learning and edge-side learning. Among them, cloud-side learning uses rich computing resources on the cloud to provide high-performance services for large models, while edge-side learning runs small models on the user's edge side to achieve fast response and low-cost model services. Although these two paradigms have their own advantages and applicable scenarios, they still have some problems in the face of a highly dynamic application environment. First, the frequent changes in the application-side environment lead to changes in local data distribution and performance requirements. For example, the target objects in video analysis tasks change with the scene, angle, and lighting conditions. Second, the resource requirements of edge-side devices change with different devices and time, and this resource fluctuation requires flexible accuracy-latency trade-offs. These changes require the learning system to quickly adjust the model size and performance to maintain performance that satisfies users.

[0029] When dealing with a highly changing application environment, both cloud-based learning paradigms and edge-side learning paradigms face significant challenges and may lead to a decline in model performance.

[0030] The cloud-based learning paradigm mainly has the following two problems: (1) Model update lag: When the edge side encounters a new application environment, the cloud side needs to update the model. However, since the cloud-side model is trained based on historical data before being deployed to the edge side, it is impossible to incorporate the data features in the new environment into the model in a timely manner, resulting in a lag in model update.

[0031] (2) High computing and communication costs: The data transmission and model update between the cloud side and the edge side consume a large amount of computing and communication resources, increasing the overall cost.

[0032] The edge-side learning paradigm mainly has the following two problems: (1) Data sparsity and bias: The data in the edge-side application environment is often sparse and may be biased, which affects the generalization performance and accuracy of the model. Even though the edge-side learning paradigm allows the use of newly collected data to update the model locally, due to data quality issues, the performance improvement of the model is limited.

[0033] (2) Resource competition and response delay: The edge-side model needs to occupy a certain amount of computing resources during training and inference. When these resources are limited, there may be resource competition between model training and inference, resulting in an extended response delay of the model service.

[0034] To solve the above problems, some researchers have proposed dynamically selecting sub-models from large models, such as nesting multiple deep neural networks in a large model, or using neural network architecture search techniques to search for suitable sub-models from a super-network, aiming to achieve a flexible accuracy-latency trade-off in the face of new edge-side environments. Although this is effective in resisting resource fluctuations, it does not utilize the newly collected data on edge-side devices. Therefore, in a dynamic edge-side environment, the performance of the model using this method will still decline. Some people have also explored collaborative learning between the edge side and the cloud side, such as using knowledge distillation techniques to transfer knowledge between cloud-edge models, or using various model pruning strategies to extract various edge-side models. Although these methods provide flexibility in defining various edge-side models, due to the time-consuming knowledge distillation and model pruning processes, the whole process is not lightweight enough. Therefore, these methods are still difficult to handle frequently changing edge-side environments.

[0035] Based on the above situation, this application proposes a method for evolving cloud-edge models in power stations and related devices. The following will make a detailed description of this application in combination with embodiments and drawings.

[0036] As Figure 1 shown, it is the first process schematic diagram of the method for evolving cloud-edge models in power stations of this application, which may include: S101, based on the trained module selector and the cloud-side large model, perform sub-model splitting for a given edge-side device; wherein, the sub-model can minimize the loss of the local data set under resource constraints; the training methods of the module selector and the cloud-side large model include: For all module layers, construct a unified module selector to determine the activated modules of all module layers; determine the activated sub-modules of each reusable module in the cloud-side large model through the module selector, and determine the expression of the cloud-side large model according to the inputs and outputs of the activated sub-modules; train the module selector and the cloud-side large model through an end-to-end algorithm.

[0037] It should be noted that the main components in the entire neural network architecture are also called the backbone model, which consists of multiple module layers, and each module layer contains multiple sub-modules. These sub-modules can be different layers or components in the neural network, such as convolutional layers, fully connected layers, etc. Each module layer has its own specific function, such as feature extraction, feature transformation, or classification. The module selector is a lightweight neural network used to dynamically select sub-modules in each module layer. It can calculate the importance weights of each sub-module according to the input data and decide which sub-modules should be activated. The activated module refers to the sub-module selected by the module selector. In each module layer, only a part of the sub-modules will be activated to reduce the computational overhead. The outputs of the activated modules will be combined as the final output of the current module layer.

[0038] S102, deploy the sub-model to the given edge device.

[0039] It should be noted that the edge device is the actual environment where the sub-model runs, and it has sufficient computing resources and storage resources to support the operation of the sub-model. At the same time, the edge device also needs to have the ability to communicate with the cloud side so as to update the sub-model or upload data when needed. The process of deployment is to transmit the parameters and structure information of the sub-model to the edge device and configure and initialize it on the device.

[0040] S103, after training the deployed sub-model, aggregate the parameters of each sub-model and update the module parameters of the cloud-side large model.

[0041] The sub-models on the edge devices need to be trained according to the local datasets to improve the performance on local tasks. The training of the sub-models on the edge side needs to consider the computing power and storage capacity of the devices, as well as issues such as energy consumption and heat dissipation during the training process. After training, the parameters of the sub-models on each edge device need to be aggregated to form the updated parameters of the cloud-side large model. The aggregation process needs to consider the data distribution and model performance differences on different devices. Then, update the module parameters of the cloud-side large model according to the aggregated parameters to improve the overall performance and generalization ability of the model.

[0042] S104, based on the module selector and the updated cloud-side large model, perform sub-model splitting for the given edge device.

[0043] The updated module selector can better adapt to different edge devices and task requirements. The module parameters of the updated cloud-side large model have been updated according to the training results of the edge devices, and the overall performance and generalization ability have been improved. Therefore, the sub-models split after the update may be more adaptable to the characteristics of the local dataset.

[0044] S105, deploy the sub-model to the given edge device again.

[0045] In this application, the cloud-edge model evolution method for power stations realizes the processes of sub-model splitting, deployment, edge training, parameter aggregation, and cloud-side update for specific edge devices through the collaborative work of the module selector and the cloud-side large model. It can effectively improve the model performance on edge devices while reducing the computing and communication costs, and has broad application prospects.

[0046] As Figure 2 shown, it is the second flow schematic diagram of the cloud-edge model evolution method for power stations in this application. The cloud-edge model evolution method of this application is mainly divided into two stages. The first stage is the cloud-side large model construction and training stage (steps S201 to S205), and the second stage is the cloud-edge large model collaborative adaptation. The following is a specific implementation method for the two stages, which may include: S201, construct the module selector.

[0047] In practical applications, the module selector is responsible for routing the input of each layer to different subsets of modules in the module layer, which can also be interpreted as the mapping from sub-tasks to activated sub-modules. To learn this mapping, this application uses a lightweight neural network composed of a fully connected layer and a softmax layer. Given the input , the l module selector of the module layer outputs the probability distribution of the sub-modules in the module layer, which can be regarded as the importance weights of each sub-module relative to the input . To reduce the computing overhead on the device, the top-k strategy is adopted in this embodiment, and only k of the available modules are activated for each input

[0048] . To combine the outputs of the activated modules, this application takes their weighted sum as the final output of the current module layer. Therefore, the final output of the module layer can be rewritten as: where is the final output of the module layer, is the input of the entire model, is the parameter of the entire model, is the hyperparameter of the module selector, n is the number of modules in each layer, is the output of the n-th layer module layer, is the parameter of the n-th layer module layer. It should be noted that, Top-k represents an algorithm that arranges the input from large to small and takes the first k values as the output, A is the module number ranked 1 to k in each layer selected by the module selector.

[0049] However, each block layer corresponds to its own module selector, which will make the model selection process less efficient because it is a continuous decision-making process, that is, the module selector takes as the input, and depends on the output of the previous module layer. To speed up this process, this application models the module selectors of all layers as a one-time decision-making process, combining all together to form a unified module selector. Therefore, this application uses an additional embedding network to extract the feature ℎ from the input x. The unified module selector can be formalized as:

[0050] where, is the module selector of the th layer, is the hyperparameter of the module selector of the th layer, is the feature extracted from the input , is the embedding network used to extract the input feature, is the hyperparameter of the embedding network, is the total number of module layers.

[0051] Therefore, the unified module selector can determine the activated modules of all module layers at one time and be decoupled from the backbone model (module), so that it can work independently and help the edge device identify important modules related to its local data distribution locally.

[0052] S202, constructing a modular large model.

[0053] As Figure 3As shown, it is a schematic diagram of the cloud-side large model in this embodiment. The cloud-side large model includes multiple reusable modules, and these reusable modules can be selectively combined to form various sub-models. Among them, the cloud-side large model forms a large feature space and can derive personalized sub-models suitable for different application environments with a very small granularity. Each derived sub-model corresponds to a sub-task. The construction of the reusable modules has a certain degree of flexibility, and the general principle is to be able to complete specific functions in the learning task (such as feature extraction). In order to enable the cloud-side large model to have a large feature space, this embodiment further horizontally expands a number of identical sub-modules for each reusable module.

[0054] In summary, each block in the cloud-side large model is formally defined in this embodiment as:

[0055] Among them, is the input vector of the l th block, is 's parameter, is the l th n nd sub-module in the th block, is l 's parameter. The l th block consists of an activated sub-module set A, and this sub-module set A is specified by the module selector constructed in step S101. The result output by the activated sub-module set of the th block specified by the module selector is the output of the l th block. l th block.

[0056] In turn, the output of the l th block, and as the input of the l +1 th block, the formal definition of the cloud-side large model can be given:

[0057] Among them, the sub-module can be composed of a neural network with any structure as long as its input and output dimensions match the block where it is located.

[0058] S203, Training improvement of the cloud-side large model.

[0059] This application proposes an end-to-end algorithm to pre-train a modular cloud-side large model and a unified module selector. In this process, the module selector decomposes the global task into multiple subtasks and maps each subtask to the sub-modules of the module layer that are correctly activated. These sub-modules are trained under the coordination of the module selector. To train such a model, in addition to the cross-entropy training loss that aligns the model output with the target label, this embodiment also adds a module load balancing loss term to ensure that each sub-module is fully trained. This load balancing technique can route similar data samples to the same activated module. Therefore, a sub-model formed by a subset of activated modules can be trained to handle specific subtasks.

[0060] Taking the classification task as an example, the loss function for training the cloud-side large model can be formalized as:

[0061] where, is the loss function for training the cloud-side large model, is the forward output result of the large model, is the ground truth, is the cross-entropy training loss, is the module load balancing loss, is the output of the unified module selector, is the weight of the module load balancing loss.

[0062] In addition, this embodiment also adopts a noisy top-k technique to achieve end-to-end training with non-differentiable top-k selection ability.

[0063] Although the subtask decomposition and mapping strategy can be automatically learned through the above end-to-end training, when deriving the sub-model for the edge side, a sub-optimal solution may be derived. This is because the sub-model required for a given subtask may be a combination of a large number of sub-modules, which breaks the resource constraints of the edge-side device. Therefore, this embodiment further proposes a model ability enhancement algorithm to learn favorable subtask decomposition and mapping strategies, so that each edge-side subtask can be completed by a sub-model containing as few sub-modules as possible. Specifically, the following method can be adopted: First, define application-specific subtasks. Subtasks can be defined based on the root cause behind the non-IID (Independent and Identically Distributed) data distribution on the edge-side device. For example, label distribution skew is a common type of non-IID data distribution, where each edge has only a small subset of classes instead of all classes in the entire dataset. Therefore, in this case, the classes that usually appear together on this edge can be defined as subtasks.

[0064] Next, determine the sub-task model objective. Based on the current sub-task mapping matrix obtained in the end-to-end training , the goal is to determine the sub-tasks that a given module n is best at, and let the module focus on these sub-tasks, leaving other sub-tasks to other modules. Based on this, this embodiment formulates the sub-task recognition process as a constrained linear programming problem. The first constraint in the algorithm aims to prevent a given module from being overloaded, that is, the load should be less than k1. The second constraint limits the maximum number of sub-modules that a sub-task can activate. To achieve the goal, this application maximizes the product of matrix elements to retain the information of the original matrix, reflecting the strategy learned in the end-to-end training. Retaining this knowledge helps reduce the tuning overhead and improve the convergence speed, because it embeds the internal structure of the global task learned in the end-to-end training stage.

[0065] S204, personalized sub-model splitting.

[0066] To fit a personalized sub-model for heterogeneous edge devices within a huge search space, this embodiment jointly considers local sub-tasks and the system resources available on the edge device to achieve a flexible trade-off between sub-task model performance and resource overhead. The goal of fitting a sub-model for a given edge device is to minimize the loss of its local dataset under resource constraints.

[0067] First, use the output of a unified module selector to define the importance metric of sub-modules, and use the edge resource constraints captured by the edge resource analyzer to estimate the resource overhead of candidate sub-models. Finally, a set of sub-modules can be selected to form a sub-model to achieve the desired performance-cost trade-off.

[0068] To identify the important modules of an edge device, define the module importance score of a given edge device as the average sample score of its local data:

[0069] where is the data collected by the edge device, is the module importance score, is the parameter of module , is the layer and the module selector.

[0070] This importance score embeds the personalized information of the local data distribution and can therefore be used to select sub-modules for edge devices.

[0071] To obtain resource constraints, first use a local resource analyzer to obtain the available resources of edge devices in the dynamic runtime environment, including memory capacity, computing power, and network bandwidth. These results will be used as the resource constraints for deriving submodels. Next, estimate the resource costs of candidate submodels on a given edge device. Since the structure of the modules is determined during the modularization phase, their resource costs can be calculated in advance on the cloud side. The resource cost of a submodel is the sum of the resource costs of all its included submodules.

[0072] After knowing the submodule importance and resource profiles, formulate the personalized submodel derivation process as a constrained optimization algorithm. Meanwhile, the edge device can also locally adjust the submodel to achieve a balance between the required model performance and computing power resources. Each edge device can occupy a set of feasible submodels, which can be dynamically adjusted to adapt to runtime resource fluctuations or changes in data distribution.

[0073] S205, Deploy the submodel.

[0074] After completing the above steps, each edge application environment derives an edge submodel from the cloud-side large model according to on-site requirements. At this time, the edge submodel will be converted, compiled, sent down, and deployed to the edge device by the model automated deployment toolchain according to the edge computing device model.

[0075] S206, Run the submodel inference.

[0076] After the edge submodel is deployed, inference will be invoked according to the business needs of the application scenario. When the submodel runs, sample selection will be based on the information entropy of the prediction results.

[0077] For a multi-classification model, the probability distribution of its output is , then its information entropy is:

[0078] Among them, is the information entropy, is the probability distribution of the i-th category.

[0079] For samples that the submodel is more "certain" about, the probability of a certain category is larger, and the probabilities of the remaining categories are smaller, so the information entropy is smaller; for samples that the model is uncertain about, the probabilities of each category are relatively uniform, and the information entropy is larger. Therefore, the information entropy can be used to evaluate the prediction certainty of the model for samples. The larger the information entropy, the more uncertain the model's prediction of the sample.

[0080] In this embodiment, the submodel performs multiple inferences in batches, calculates the difference in the entropy of the multi-batch inference results, and evaluates the sample value. The sample evaluation function can be:

[0081] 。

[0082] Among them, is the sample evaluation function, is the entropy of the expectation on the distribution of under the condition of a given x and is the expectation of the distribution, is, under the condition of a given x, and the expectation of the joint entropy of

[0083] S207, iterative training of the sub-model.

[0084] After batch selecting sufficient samples through step S201, the sub-model running environment will select the time period with the lowest load of the edge device for fine-tuning training.

[0085] S208, modular sub-model aggregation.

[0086] The purpose of aggregating and updating the sub-model is to transfer the new knowledge learned by the edge device back to the cloud-side large model. To aggregate heterogeneous edge-side sub-models, this embodiment proposes a modular weighted average aggregation method. Its basic principle is that the sub-models are constructed by the same modules, so they can be aggregated in a modular way.

[0087] The specific method is as follows: The weighted average of the parameters of sub-module i of all sub-models within can be calculated to update the parameters of sub-module i. is the set of sub-models containing sub-module i. Considering that each sub-module i can be updated by different sub-models different numbers of times, the (normalized) importance value of sub-module i relative to the sub-model is used as the average weight to balance the contribution of each sub-model. That is, the parameters of sub-module i are updated by the following formula: 。

[0088] Among them, is the parameter of sub-module i after update, is the parameter of sub-module k in sub-model i 。

[0089] This modular aggregation reduces parameter conflicts because each module is trained with data samples of specific sub-tasks and is not interfered by different sub-tasks on other edge devices.

[0090] S209, fine-tuning and enhancement of the large model.

[0091] Based on the obtained target mapping matrix P = H⊙M, where H is the sub-task matrix, M is the large model module matrix, and P is the mapping of the sub-task in the module. There are two goals in the fine-tuning process: one is to use more data in the sub-tasks it focuses on to train each module to further improve its ability on that sub-task, and the other is to let the module selector be updated under the guidance of the new sub-task mapping strategy. For this purpose, each sample of the sub-task is attached with an additional label , indicating the recommended module to be activated.

[0092] S210, Personalized sub-model splitting.

[0093] The method is the same as step S204.

[0094] S211, Deploy the sub-model to the sub-model again.

[0095] The method is the same as step S205.

[0096] This application provides a new modular model decomposition design paradigm. Based on this, sub-models with personalized characteristics can be effectively derived for edge devices in different application environments, and the sub-models can be effectively aggregated and updated to converge the newly learned knowledge on the edge side to the cloud side. Specifically, this application decomposes the cloud-side large model into multiple separable and combinable modules, that is, decomposes the global task (represented by the global data distribution) into multiple sub-tasks (represented by the local data distribution of the edge environment), and each sub-task can be solved by combining an appropriate subset of the cloud-side large model. This design enables the flexible derivation and aggregation of sub-models with different model sizes and personalized capabilities, while avoiding time-consuming model architecture search or knowledge distillation processes.

[0097] As Figure 4 shown, it is a schematic diagram of the cloud-edge model evolution system of this application, which may include: The first splitting module is used to perform sub-model splitting for a given edge device based on the trained module selector and the cloud-side large model; among them, the sub-model can minimize the loss of the local data set under resource constraints; the training methods of the module selector and the cloud-side large model include: Activate a unified module selector for all module layers to determine the activated modules of all module layers; determine the activated sub-modules of each reusable module in the cloud-side large model through the module selector, and determine the expression of the cloud-side large model according to the input and output of the activated sub-modules; train the module selector and the cloud-side large model through an end-to-end algorithm; The first deployment module is used to deploy the sub-model to a given edge device; An update module, which is used to aggregate the parameters of each sub-model after training the sub-models deployed and distributed, and update the module parameters of the large model on the cloud side; A second splitting module, which is used to split the sub-model for a given edge device based on the module selector and the updated large model on the cloud side; A second deployment module, which is used to deploy and distribute the sub-model to the given edge device again.

[0098] It should be noted that in several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the system embodiments described above are only illustrative. For example, the division of each module is only a logical function division. In actual implementation, there may be other division methods. For example, multiple modules can be combined or integrated into another device, or some features can be ignored or not executed. The modules described as separate components may or may not be physically separated. The components shown as modules may be a physical unit or multiple physical units, that is, they can be located in one place or distributed to multiple different places. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0099] In addition, the modules in each embodiment of the present invention can be integrated in a processing unit, or each module can exist physically alone, or two or more modules can be integrated in a unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit.

[0100] The embodiments of the present application further provide an electronic device, which may include one or more processors, a memory, and a communication interface.

[0101] Among them, the memory and the communication interface are coupled to the processor. For example, the memory and the communication interface can be coupled together through a bus.

[0102] Among them, the communication interface is used to transmit data with other devices. The memory stores computer program code. The computer program code includes computer instructions. When the computer instructions are executed by the processor, the electronic device executes the steps of the above-mentioned power station cloud-edge model evolution method.

[0103] Among them, the processor can be a processor or a controller. For example, it can be a Central Processing Unit (CPU), a general-purpose processor, a Digital Signal Processor (DSP), an Application-Specific Integrated Circuit (ASIC), a Field Programmable Gate Array (FPGA), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It can implement or execute various exemplary logic blocks, modules, and circuits described in conjunction with the present disclosure. The processor can also be a combination that realizes computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, and so on. The processor can be used to support the electronic device to execute the method steps provided in the above embodiments.

[0104] Among them, the bus can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. The above buses can be divided into an address bus, a data bus, a control bus, and so on.

[0105] A computer-readable storage medium provided by an embodiment of the present application stores a computer program, and when the computer program is executed by a processor, the steps of the above-mentioned power station cloud-edge model evolution method are realized.

[0106] The computer-readable storage medium involved in the present application includes a Random Access Memory (RAM), a memory, a Read-Only Memory (ROM), an Electrically Programmable ROM, an Electrically Erasable Programmable ROM, a register, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium well-known in the technical field.

[0107] The above are only the preferred embodiments of the present application and are not used to limit the present application. For those skilled in the art, the present application can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the protection scope of the present application.

Claims

1. A method for evolving a cloud edge model of a power station, characterized in that: include: Based on the trained module selector and the cloud-side large model, a sub-model is split for a given edge device; wherein the sub-model can minimize the loss of the local data set under resource constraints; the training method of the module selector and the cloud-side large model includes: For all module layers, a unified module selector is constructed to determine the activation modules of all module layers; the activation submodules of each reusable module in the cloud-side large model are determined by the module selector, and the expression of the cloud-side large model is determined according to the input and output of the activation submodule; the module selector and the cloud-side large model are trained by an end-to-end algorithm; Deploy the sub-model for a given edge device; After training the deployed sub-models, aggregate the parameters of each sub-model and update the module parameters of the cloud-side large model. Based on the module selector and the updated cloud-side large model, split the sub-model for a given edge device; Deploy the sub-model again for the given edge device.

2. The power station cloud edge model evolution method according to claim 1 is characterized in that: The unified module selector expression includes: in, For the layer module selector, For the layer module selector hyperparameters, For input The features extracted from is an embedding network used to extract input feature, is the embedding network hyperparameter, is the total number of module layers; The expression of the cloud-side large model includes: in, is the expression of the large model on the Izumo side, For the L Reusable modules, For the L -1 reusable module, This is the first reusable module.

3. The power station cloud edge model evolution method according to claim 1 is characterized in that: The loss function used when training the module selector and the cloud-side large model through the end-to-end algorithm includes cross entropy training loss and module load balancing loss.

4. The power station cloud edge model evolution method according to claim 1 is characterized in that: The method for training a module selector and a cloud-side large model through an end-to-end algorithm includes: Define application-specific subtasks based on the root causes behind non-IID data distribution on edge devices; The subtask identification process is formulated as a constrained linear programming problem, and the product of the subtask mapping matrix elements is maximized to determine the model objective of the subtask. The constraints include: making the load smaller than the maximum number of tasks that the module can bear; making the maximum number of submodules that can be activated by the subtask smaller than the maximum number of modules that can be used; The subtask-based model objective trains the module selector and the cloud-side large model through an end-to-end algorithm.

5. The power station cloud edge model evolution method according to claim 1 is characterized in that: The method for splitting a given edge device into sub-models includes: Define an importance metric for the sub-module based on the output of the module selector, and estimate the resource overhead of the candidate sub-model using the side resource constraints captured by the side resource analyzer of the given side device; Combining the importance metric of the sub-module and the resource overhead of the candidate sub-model, a constrained optimization algorithm is used to fit the sub-model for a given edge device.

6. The power station cloud edge model evolution method according to claim 1, characterized in that: The method for training the deployed sub-model includes: Enable the deployed sub-model to perform reasoning on different samples; Evaluate the samples according to the information entropy of the probability distribution output after the inference operation, and select samples with information entropy less than the preset requirement as training samples; The sub-model is fine-tuned using training samples.

7. The power station cloud edge model evolution method according to claim 1, characterized in that: The method for aggregating the parameters of each sub-model includes: 。 in, is the updated parameter of submodule i, is the sub-model set containing sub-module i, is the module importance score of a given edge device, For submodules i No. k sub-model parameters, is the parameter of submodule i, Data collected for edge devices.

8. A power station cloud edge model evolution system, characterized in that: include: The first splitting module is used to split the sub-model for a given edge device based on the trained module selector and the cloud-side large model; wherein the sub-model can minimize the loss of the local data set under resource constraints; the training method of the module selector and the cloud-side large model includes: Activation: For all module layers, a unified module selector is constructed to determine the activation modules of all module layers; the activation submodules of each reusable module in the cloud-side large model are determined by the module selector, and the expression of the cloud-side large model is determined according to the input and output of the activation submodule; the module selector and the cloud-side large model are trained through an end-to-end algorithm; The first deployment module is used to send a deployment sub-model to a given edge device; The update module is used to aggregate the parameters of each sub-model after training the deployed sub-models, and update the module parameters of the cloud-side large model; The second splitting module is used to split the sub-model for a given edge device based on the module selector and the updated cloud-side large model; The second deployment module is used to send the deployment sub-model to the given edge device again.

9. An electronic device, characterized in that: include: A memory and one or more processors; the memory is coupled to the processor; wherein the memory stores computer program code, the computer program code includes computer instructions, and when the computer instructions are executed by the processor, the electronic device executes the steps of the power station cloud edge model evolution method as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the power station cloud-edge model evolution method as described in any one of claims 1 to 7 are implemented.