A splitting learning method and system of a power large model

CN122616637APending Publication Date: 2026-08-21GUANGZHOU POWER SUPPLY BUREAU GUANGDONG POWER GRID CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610584400.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-04-29
Publication Date
2026-08-21

AI Technical Summary

Technical Problem

[0003]相关技术中,需要将海量的电力数据传输到部署有电力大模型的云计算中心进行处理,而这不仅导致电力数据经历较长的网络传输延迟,而且将电力数据传输到公网的云计算中心,存在数据泄漏的安全隐患

Benefits of technology

[0041]上述一种电力大模型的拆分学习方法和系统,获取智能电网下发的电力任务;根据电力任务向相应的目标节点下发电力任务处理指令,电力任务处理指令用于指示目标节点根据采集到的电力任务处理数据,基于训练完备的电力大模型进行模型推理计算,以得到推理结果;其中,控制电力大模型进行训练的步骤包括:响应于接收到的电力大模型训练任务,根据多个边缘计算节点的节点计算能力和网络通信延迟,确定电力大模型的拆分映射策略;控制各个算力映射节点和大模型梯度更新控制器基于拆分映射策略执行协同并行模型迭代训练,直至模型收敛,得到训练完备的电力大模型;拆分映射策略表征了将电力大模型拆分为多个拆分段,将不同拆分段分配至边缘计算节点中不同的算力映射节点,基于算力映射节点的计算资源对对应的拆分段进行训练。通过本申请实施例,将电力大模型拆分为多个段并映射到不同的边缘节点,实现了流水线并行的训练和推理,降低了对单个边缘节点的算力和存储要求,减少了数据传输到云端的延时,降低了安全风险。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122616637A_ABST
    Figure CN122616637A_ABST
Patent Text Reader

Abstract

The application relates to a power large model splitting learning method and system, which is applied to a decision controller and comprises the following steps: issuing a power task processing instruction to a target node according to a power task, the power task processing instruction being used for instructing the target node to perform model reasoning based on a power large model according to collected data; and controlling the power large model to perform training: in response to a received power large model training task, determining a splitting mapping strategy of the power large model according to node computing capacity and network communication delay of multiple edge computing nodes; performing collaborative parallel model iterative training based on the splitting mapping strategy to obtain a trained power large model; the splitting mapping strategy represents splitting the power large model into splitting segments, distributing the splitting segments to different computing resource mapping nodes in the edge computing nodes, and training the splitting segments based on the computing resource mapping nodes. The application can efficiently and safely realize training and reasoning of the power large model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of edge computing technology, and in particular to a decomposition learning method and system for a large power model. Background Technology

[0002] Large-scale power models can extract valuable information from massive amounts of multimodal power data to assist in intelligent scheduling and management of the power grid, and are an important approach to the digitalization and intelligentization of the power sector.

[0003] In related technologies, massive amounts of power data need to be transmitted to cloud computing centers that deploy large power models for processing. This not only causes the power data to experience long network transmission delays, but also poses a security risk of data leakage when transmitting power data to cloud computing centers on the public network.

[0004] Therefore, how to efficiently and securely train and infer large-scale power models has become an urgent technical problem to be solved. Summary of the Invention

[0005] Therefore, it is necessary to provide a decomposition learning method and system for large-scale power models to address the aforementioned technical problems.

[0006] Firstly, this application provides a decomposition learning method for a large power model, applied to a decision controller, including:

[0007] Obtain power tasks issued by the smart grid;

[0008] According to the power task, the power task processing instructions are issued to the corresponding target nodes. The power task processing instructions are used to instruct the target nodes to perform model inference calculations based on the collected power task processing data and the fully trained power large model in order to obtain the inference results.

[0009] The steps for controlling the training of the large power model include: responding to the received large power model training task, determining the splitting and mapping strategy of the large power model based on the node computing capabilities and network communication latency of multiple edge computing nodes; controlling each computing power mapping node and the large model gradient update controller to perform collaborative parallel model iterative training based on the splitting and mapping strategy until the model converges and a fully trained large power model is obtained; the splitting and mapping strategy represents splitting the large power model into multiple segments, allocating different segments to different computing power mapping nodes in the edge computing nodes, and training the corresponding segments based on the computing resources of the computing power mapping nodes.

[0010] In one embodiment, the step of establishing a split mapping strategy includes:

[0011] In the edge computing nodes, determine the current computing power mapping node that meets the training requirements of the current model segment; wherein, the storage space of the current computing power mapping node is greater than the preset storage space threshold, and the data transmission delay between the current computing power mapping node and the computing power mapping node of the previous model segment is not higher than the preset communication delay threshold, and the computing resources of the current computing power mapping node are greater than the preset computing resource threshold.

[0012] The number of layers of the large power model included in the current model split segment is determined based on the current computing power mapping node;

[0013] Based on the model segmentation of the large power model and the computing power mapping nodes corresponding to each model segmentation, the splitting and mapping strategy is determined.

[0014] In one embodiment, determining the number of layers of the large power model included in the current model segmentation based on the current computing power mapping node includes:

[0015] The current first layer number is determined based on the maximum number of layers in the storage space that satisfies the current computing power mapping node.

[0016] The current second layer is determined based on the maximum number of layers that meet the preset expected training duration.

[0017] Based on the current first layer number and the current second layer number, determine the number of layers of the large power model included in the current model split segment.

[0018] In one embodiment, the power large-scale model training task includes a set of data sources, which are collected by corresponding data source terminals; the method further includes:

[0019] When the data source terminal has computing power, the data source terminal is used as the first computing power mapping node corresponding to the first model segment in the power big model; the first computing power mapping node is the edge computing node used to perform the first segment model calculation of the power big model.

[0020] In cases where the data source terminal lacks computing power, the edge computing node with the highest transmission rate to the data source terminal is designated as the first computing power mapping node corresponding to the first model segment in the large power model.

[0021] In one embodiment, the power big model includes multiple first computing power mapping nodes, and the method further includes:

[0022] Traverse each first computing power mapping node. If each first computing power mapping node satisfies the preset layer number constraint, determine that the first model splitting segment corresponding to each first computing power mapping node includes the first layer and the second layer of the power big model.

[0023] In the case where the first computing power mapping node does not meet the layer constraint, the first model splitting segment corresponding to each first computing power mapping node is determined to include the first layer of the power big model; the layer constraint is that the total time required for the first computing power mapping node to calculate the first and second layer data of the power big model is less than the preset maximum waiting time, and the memory capacity of the first computing power mapping node is greater than the capacity required for the first and second layer data of the power big model.

[0024] In one embodiment, the step of training the large-scale power model includes:

[0025] The forward propagation parameters of the previous computing power mapping node are received through the first layer of the current model segment of the target computing power mapping node.

[0026] Based on the received forward propagation parameters, the forward propagation training of the current segment of the model is performed according to the forward propagation segment of each micro-batch. The micro-batch is the batch size of each computing power mapping node in each round of training. The micro-batch is determined according to the number of the first computing power mapping nodes and the preset small batch data volume.

[0027] The forward propagation parameters of the target computing power mapping node are passed to the computing power mapping node corresponding to the next model split segment through the last layer of the current model split segment, and the training parameters of each micro-batch of the current split segment are stored in the storage space of the target computing power mapping node.

[0028] Repeat the above steps until the computing power mapping node corresponding to each split segment has been traversed.

[0029] In one embodiment, the method further includes:

[0030] When traversing the computing power mapping nodes corresponding to each model split segment, the backpropagation parameters of the next model split segment are received through the last layer of the current split segment of the target computing power mapping node.

[0031] Based on the received backpropagation parameters, the backpropagation training of each micro-batch in the current model segment is executed sequentially through the current model segment. The backpropagation parameters are then passed to the computing power mapping node corresponding to the previous model segment through the first layer of the current model segment.

[0032] Repeat the above steps until the backpropagation parameters are passed to the first computing power mapping node corresponding to the first model split segment. The first computing power mapping node is used to perform backpropagation training of each layer in the first model split segment in the micro-batch corresponding to the first computing power mapping node when it receives the backpropagation parameters of the next model split segment, to obtain the model weight gradient, and send the model weight gradient to the large model gradient update controller.

[0033] In one embodiment, the method further includes:

[0034] When the model weight gradient is sent to the large model gradient update controller, the target update gradient parameters are calculated by the large model gradient update controller based on the gradient corresponding to each first computing power mapping node.

[0035] The target update gradient is sent to each first computing power mapping node through the large model gradient update controller. The first computing power mapping node is used to update the model parameters in the corresponding first model split segment.

[0036] In one embodiment, the method further includes:

[0037] The training process of the power big data model is monitored based on a preset monitoring period. When the training data of the power big data model meets the preset alarm requirements, a new split mapping strategy is formulated, and the model iterative training is performed according to the new split mapping strategy. The alarm requirements include that at least one computing power mapping stage is in an unreachable state, or that the total training latency of the power big data model is greater than the preset maximum latency threshold.

[0038] Secondly, this application also provides a power large model splitting learning system, including: a decision controller, a large model gradient update controller, an edge computing node, and a power Internet of Things terminal; the power Internet of Things terminal is communicatively connected to the corresponding edge computing node, the decision controller is communicatively connected to the edge computing node through a base station, and the large model gradient update controller is communicatively connected to the edge computing node through a base station.

[0039] The power Internet of Things (IoT) terminal is used to collect corresponding power task processing data based on the power tasks issued by the smart grid.

[0040] A decision controller is used to perform the steps of any of the methods described above.

[0041] The aforementioned method and system for splitting and learning a large-scale power model involves acquiring power tasks issued by a smart grid; issuing power task processing instructions to corresponding target nodes based on the power tasks, which instruct the target nodes to perform model inference calculations based on the collected power task processing data and a fully trained large-scale power model to obtain inference results; wherein, the steps for controlling the training of the large-scale power model include: responding to the received large-scale power model training task, determining a splitting and mapping strategy for the large-scale power model based on the node computing capabilities and network communication latency of multiple edge computing nodes; controlling each computing power mapping node and the large-scale model gradient update controller to perform collaborative parallel model iterative training based on the splitting and mapping strategy until the model converges, resulting in a fully trained large-scale power model; the splitting and mapping strategy represents splitting the large-scale power model into multiple segments, allocating different segments to different computing power mapping nodes in the edge computing nodes, and training the corresponding segments based on the computing resources of the computing power mapping nodes. Through the embodiments of this application, the large power model is divided into multiple segments and mapped to different edge nodes, realizing pipelined parallel training and inference, reducing the computing power and storage requirements of individual edge nodes, reducing the latency of data transmission to the cloud, and reducing security risks. Attached Figure Description

[0042] To more clearly illustrate the technical solutions in the embodiments of this application or related technologies, the drawings used in the description of the embodiments of this application or related technologies will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0043] Figure 1 This is a flowchart illustrating the split learning method in one embodiment;

[0044] Figure 2 This is a flowchart illustrating the process of establishing a split mapping strategy in one embodiment;

[0045] Figure 3 This is a schematic diagram of the process of training a large power model in one embodiment;

[0046] Figure 4 This is a schematic diagram of the process for training a large-scale power model in another embodiment;

[0047] Figure 5 This is a schematic diagram of micro-batch scheduling for training a large model in one embodiment;

[0048] Figure 6 This is a schematic diagram of the process for training a large-scale power model in another embodiment;

[0049] Figure 7 This is a flowchart illustrating the split learning method in a preferred embodiment;

[0050] Figure 8 This is a schematic diagram of the structure of a split learning system in one embodiment. Detailed Implementation

[0051] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0052] The technical background of this application is described below.

[0053] With the mature development of natural language processing and big data technologies, large-scale artificial intelligence models have received widespread attention in recent years. The power Internet of Things (IoT) possesses abundant multimodal data. Utilizing large-scale power models to extract effective information from this massive amount of multimodal power data to assist in intelligent dispatching and management of the power grid is an important approach to the digitalization and intelligentization of the power sector.

[0054] Furthermore, cloud computing centers have extremely strong computing power and can better support the learning requirements of large models. However, transmitting massive amounts of power data to cloud computing centers for model learning not only results in long network transmission delays, which cannot meet the low latency requirements of latency-sensitive production control power businesses, but also easily creates security risks of data leakage when transmitting power data to cloud computing centers on the public network.

[0055] Edge computing is an emerging computing paradigm that has emerged in recent years to extend cloud computing services and improve data privacy and security. Power edge computing reduces network communication latency and mitigates data security risks by deploying 5G power virtual private networks (VPNs) within 5G networks and deploying MEC (Multi-access Edge Computing) servers and edge computing nodes at the edge of these private power networks. However, compared to cloud computing, edge computing nodes have limited computing power and cannot independently support the diverse and large-scale training needs of artificial intelligence models. Therefore, multi-node collaborative computing is required. Furthermore, due to the massive scale of large models, traditional federated learning frameworks—where local terminals train local models and edge servers aggregate and update the global model—are also insufficient to support network edge training of large models. This is because, although the local model only trains on local data, the local terminal trains the entire model. Due to the limitations of terminal computing power and storage capacity, the local terminal cannot independently support the training and storage of the entire large model's massive parameters.

[0056] In summary, the above description illustrates the main technical challenges of the relevant technologies: how to efficiently and securely train and infer large-scale power models.

[0057] Based on this, this application provides a method for splitting and learning a large-scale power model, applied to a decision controller, comprising: splitting a fully trained large-scale power model into multiple computing power mapping nodes for collaborative computation according to a preset splitting and mapping strategy; performing model inference processing on the received power task to obtain inference results; wherein, the training steps of the large-scale power model include: responding to the received large-scale power model training task, determining the splitting and mapping strategy of the large-scale power model based on the node computing capabilities and network communication latency of multiple edge computing nodes; performing collaborative parallel model iterative training based on the splitting and mapping strategy through each computing power mapping node and the large-scale model gradient update controller until the model converges to obtain a fully trained large-scale power model; the splitting and mapping strategy represents splitting the large-scale power model into multiple segments, allocating different segments to different computing power mapping nodes in the edge computing nodes, and training the corresponding segments based on the computing resources of the computing power mapping nodes. See the following embodiments for details.

[0058] In one exemplary embodiment, such as Figure 1 As shown, a decomposition learning method for a large power model is provided and applied to a decision controller, including:

[0059] S110, obtains power tasks issued by the smart grid;

[0060] S120: According to the power task, the power task processing instruction is issued to the corresponding target node. The power task processing instruction is used to instruct the target node to perform model inference calculation based on the collected power task processing data and the fully trained power large model to obtain the inference result.

[0061] The steps for controlling the training of the large power model include: responding to the received large power model training task, determining the splitting and mapping strategy of the large power model based on the node computing capabilities and network communication latency of multiple edge computing nodes; controlling each computing power mapping node and the large model gradient update controller to perform collaborative parallel model iterative training based on the splitting and mapping strategy until the model converges and a fully trained large power model is obtained; the splitting and mapping strategy represents splitting the large power model into multiple segments, allocating different segments to different computing power mapping nodes in the edge computing nodes, and training the corresponding segments based on the computing resources of the computing power mapping nodes.

[0062] The decision controller is connected to the smart grid via a communication network. It receives commands from the smart grid and communicates with multiple computing nodes via base stations. Power IoT terminals communicate with their corresponding computing nodes; the power IoT collects data from various power devices. The large model gradient update controller communicates with edge computing nodes via base stations. The decision controller is a logical control unit deployed in the edge computing network; it can be a standalone server or a control cluster composed of multiple devices.

[0063] The aforementioned power task processing data refers to the data acquired to complete the power tasks issued by the smart grid, such as monitoring images of equipment status, current and voltage time-series data in the equipment, etc. This data can be collected by the power Internet of Things terminal and reported to the decision controller.

[0064] The power large-scale model training task refers to the instruction to train an initial, untrained power large-scale model. This task includes large-scale model information, a data source set, and learning requirement information. The large-scale model information includes the task type, layer partitioning, and mini-batch size. The learning requirement information includes the training epoch threshold and loss rate threshold. The aforementioned data source set is composed of power equipment data collected from at least one power IoT terminal. ,in, This represents the i-th data source, which is the power IoT terminal that carries power data.

[0065] Edge computing nodes are physical devices deployed at the edge of a power grid, possessing certain computing and storage capabilities. A node's computing power characterizes the amount of computation a node can complete per second, determining the computation time required for the node to run one or more layers of a large model. Furthermore, a computing power mapping node is an edge computing node selected by the decision controller to be responsible for the training and inference computation of a specific segment. This node needs sufficient storage space to store all model parameters for that segment, and its computing power must be sufficient to allow the segment to complete forward and backward propagation within a specified latency. A segmentation refers to dividing a complete large power model into multiple consecutive, non-overlapping layers according to the sequence of different layers (such as convolutional layers and fully connected layers in a neural network). For example, the first segment might contain the first and second layers of the model, the second segment might contain layers three through six, and so on, completing the segmentation of all layers within a large power model.

[0066] The collaborative parallel model iterative training mentioned in this application refers to multiple computing power mapping nodes processing different micro-batch data simultaneously in a pipelined parallel manner, with each node responsible for the forward and backward propagation of its own segment.

[0067] The aforementioned target node refers to a computing power mapping node that corresponds to a power task, is used to perform power tasks, and has computing capabilities. It can also be an IoT terminal without computing capabilities.

[0068] In this embodiment, during the model inference phase, the decision controller receives instructions from the smart grid regarding power tasks, such as distribution line fault identification and load forecasting. These power tasks may include determining corresponding power task processing data, such as requiring substation voltage data for the most recent hour. Based on the power task, the decision controller instructs the corresponding power IoT terminal (i.e., the target node) to complete data acquisition and inputs the acquired power task processing data into the computing power mapping node of the first segment of the fully trained power model. After completing the calculation, the computing power mapping node of the first segment transmits the intermediate calculation results to the computing power mapping node of the second segment, and so on, until the node of the last segment outputs the final inference result, which may be: no line faults, load forecast value for the next 15 minutes is 3.2 MW, etc. In summary, the model inference of the power model is completed.

[0069] In the model training phase before the inference phase, a splitting and mapping strategy is determined based on the received power large model training task, the node computing power and network latency of the edge computing nodes. The splitting and mapping strategy is used to indicate that the power large model is split into multiple split segments, each of which includes at least one layer of the model. Different split segments are assigned to different nodes in the edge computing nodes (i.e. computing power mapping nodes). During the training phase, for each micro-batch, each computing power mapping node in the first split segment executes the forward propagation calculation of the first split segment on its own node in parallel. According to the order of model layers, it calculates from the input layer in the first split segment to the last layer of the segment, and passes the output of the last layer to the computing power mapping node of the second split segment, and so on, until it is transmitted to the last split segment. The computing power mapping node corresponding to the last split segment also executes the forward propagation calculation to obtain the final prediction result and calculate the loss function value. In practical applications, the first split segment generally corresponds to at least one computing power mapping node. Each computing power mapping node with the first split segment executes its own calculation in parallel. Except for the first split segment, the other split segments are deployed on one computing power mapping node respectively. For example, if the computing power model is split into five segments, the first segment can be deployed on three computing power mapping nodes, and the other four split segments are each deployed on one computing power mapping node.

[0070] The computational mapping node of the last split segment executes the backpropagation of each micro-batch in reverse order, calculates the first-layer gradient (i.e., error signal) of that segment, and passes this gradient to the computational mapping node of the previous split segment, and so on, until it reaches the computational mapping node of the first split segment. After all the gradients of the computational mapping nodes of the first split segment are received in this round of training, the large model gradient update controller synthesizes these gradients to obtain the target gradient and returns it to each computational mapping node of the first split segment, thus completing the gradient update of the first split segment. For other split segments (except the first split segment), since each segment corresponds to only one computational mapping node, its parameter update is performed locally by that node itself after receiving the backpropagation gradient, without needing to go through the large model gradient update controller.

[0071] Repeat the above steps until the large power model converges. After the training phase is complete, the obtained split mapping strategy is persistently stored. In the subsequent inference phase, the decision controller directly reuses this strategy, that is, according to the split segment division and node allocation determined during training, it performs pipelined parallel computation containing only forward propagation, without re-determining the split mapping decision.

[0072] In summary, it is understandable that the decision controller itself does not perform model training or power task processing calculations. During the inference phase, the decision controller is used to instruct the corresponding nodes to perform model inference calculations based on the fully trained power model. During the training phase, it is used to formulate strategies such as model splitting and split segment allocation, and instruct the corresponding nodes to perform training calculations based on the split mapping strategy. The decision controller itself does not perform model training.

[0073] Through the embodiments of this application, the large power model is divided into multiple segments and mapped to different edge nodes, realizing pipelined parallel training and inference, reducing the computing power and storage requirements of individual edge nodes, reducing the latency of data transmission to the cloud, and reducing security risks.

[0074] In one exemplary embodiment, such as Figure 2 As shown, the steps for establishing a split mapping strategy include:

[0075] S210, determine the current computing power mapping node that meets the training requirements of the current model segment in the edge computing nodes, wherein the storage space of the current computing power mapping node is greater than the preset storage space threshold, the data transmission delay between the current computing power mapping node and the previous model segment is not higher than the preset communication delay threshold, and the computing resources of the current computing power mapping node are greater than the preset computing resource threshold.

[0076] The current model segment is any segment other than the first model segment, and the current computing power mapping node is the computing power mapping node that has the current model segment deployed.

[0077] This application describes a method for determining a computing power mapping node from multiple edge computing nodes. Specifically, a node is randomly selected from the edge computing nodes that meets the following condition: the storage space of the current computing power mapping node is greater than a preset storage space threshold, thereby enabling the node to store the Lth element of a large model. j+1 The training parameter information of at least one layer starting from the first layer; and communication between the current computing power mapping node and the computing power mapping node of the previous model segment is reachable, and the data transmission latency is not higher than a preset communication latency threshold; and the computing resources of the current computing power mapping node are greater than a preset computing resource threshold, so that the current computing power mapping node can run the Lth layer of the large model. j+1 Training must begin at least one layer. In summary, nodes that simultaneously meet the above conditions are designated as the current computing power mapping nodes and labeled as N. k-1 .

[0078] The specific method for determining whether communication is available between the current computing power mapping node and the computing power mapping node of the previous segment and whether the data transmission latency is lower than a preset communication latency threshold is as follows:

[0079] Another N i-1 N i Let represent the computing power mapping nodes of the previous split segment and the current model split segment, respectively. Then, the data transmission rate from the previous split segment to the computing power mapping node of the current segment can be determined by the following formula:

[0080]

[0081] Where B represents the wireless link bandwidth, This indicates the transmit power of the previous stage node. Indicates channel gain. If Gaussian white noise is used, then communication between the previous split segment and the current segment's computing power mapping node is reachable. It is confirmed that, among them, It is the communication rate threshold, let This indicates the amount of data that needs to be transmitted from the previous segment to this segment. Therefore, when the data transmission delay is:

[0082]

[0083] The data transmission delay satisfies At that time, the data transmission delay is lower than the communication delay threshold, where G represents k-1 Stage (i.e., computing power mapping node N) i-1The model parameters (corresponding to the large model splitting stage) are transferred to g. k Stage (i.e., computing power mapping node N) i The maximum tolerable delay for the corresponding large model splitting stage.

[0084] S220, determine the number of layers of the large power model included in the current model splitting segment based on the current computing power mapping node.

[0085] In this embodiment of the application, after determining the current computing power mapping node, the number of layers of the large model included in the current model splitting segment allocated to the current computing power mapping node is determined based on the storage capacity and communication latency of the computing power mapping node.

[0086] In this embodiment of the application, except for the first split segment, all other split segments can use the above method to determine the computing power mapping node corresponding to the split segment, as well as the number of layers of the large model contained in the split segment.

[0087] S230: Determine the splitting and mapping strategy based on the model splitting segments of the power large model and the computing power mapping nodes corresponding to each model splitting segment.

[0088] In this embodiment of the application, the splitting mapping strategy can be determined based on the model splitting method and the computing power mapping nodes corresponding to each splitting segment.

[0089] In an exemplary embodiment, determining the number of layers of the large power model included in the current model split segment based on the current computing power mapping node includes:

[0090] The current first layer number is determined based on the maximum number of layers in the storage space that satisfies the current computing power mapping node.

[0091] The current second layer is determined based on the maximum number of layers that meet the preset expected training duration.

[0092] Based on the current first layer number and the current second layer number, determine the number of layers of the large power model included in the current model split segment.

[0093] In this embodiment of the application, after determining the current computing power mapping node, the number of layers of the large model included in the current model splitting segment corresponding to the current computing power mapping node is determined.

[0094] Determine the maximum number of layers that can meet the storage space requirements of the current computing power mapping nodes, and use this as the current first layer number. Specifically, start from the unsegmented layer, i.e., from L. j+1 Starting with layer L, the algorithm finds the current first layer whose sum of training parameters that needs to be stored is closest to, but not higher than, the storage capacity of the current computing power mapping node by gradually adding layers. For example, adding a split segment ends at layer L. j Layer, that is, the current split segment from the Lth layer. j+1Starting with layer L, first try to include only layer L in the current first layer. j+1 Layer, check if the node memory is sufficient to accommodate the Lth layer. j+1 If there are enough layers, continue trying to add the Lth layer. j+2 Layer, check if the node memory is sufficient to accommodate the Lth layer. j+1 +L j+2 If that's enough, then continue trying to increase the number of Lth digits. j+3 Layer by layer, and so on, until adding another layer would exceed the storage space size of the current computing power mapping node.

[0095] Determine the maximum number of layers that can satisfy the desired training duration, and use it as the current second layer. Specifically, starting from the undivided layers, gradually increase by one layer, finding the layer whose sum of training duration is closest to, but does not exceed, the desired training duration. For example, starting from the Lth layer... j+1 Starting with a layer, calculate the time required for the current computing power mapping node to train only this layer. If it is less than or equal to the preset expected training time (e.g., 50ms), then continue to add the next layer. Check whether the total time for the trainer of the two layers is still less than or equal to the expected training time, and so on, until adding another layer will time out.

[0096] Take the intersection of the current first layer number and the current second layer number as the layer number of the large model included in the current model split segment, and mark this split segment as: , where the natural number t is the number of layers in that segment.

[0097] In summary, through the above steps, for each split segment except the first split segment, the computing power mapping node is first determined, and then the appropriate number of layers contained in the split segment is determined based on the computing power mapping node, until the splitting of all layers in the large computing power model is completed, and the splitting mapping strategy is applied. The code is converted into network instructions that can be recognized by the relevant computing power mapping nodes and sent to the relevant nodes. Here, k∈{1,2,...,k} is the splitting and computing power mapping joint strategy of the kth splitting segment, and the natural number k represents the number of splitting segments in the strategy of the large model.

[0098] In one exemplary embodiment, the power large-scale model training task includes a set of data sources, which are collected by corresponding data source terminals; the method further includes:

[0099] When the data source terminal has computing power, the data source terminal is used as the first computing power mapping node corresponding to the first model segment in the power big model; the first computing power mapping node is the edge computing node used to perform the first segment model calculation of the power big model.

[0100] In cases where the data source terminal lacks computing power, the edge computing node with the highest transmission rate to the data source terminal is designated as the first computing power mapping node corresponding to the first model segment in the large power model.

[0101] Among them, the first computing power mapping node is the node corresponding to the first model splitting segment, the first model splitting segment is the first splitting segment in the power big model, and the first model splitting segment includes at least one layer of model.

[0102] In this embodiment of the application, each data source terminal is determined. If the terminal has computing power, then the terminal is designated as the first computing power mapping node corresponding to the first model splitting segment. This may include multiple first computing power mapping nodes, each of which is equipped with a first model splitting segment.

[0103] If the data source terminal lacks computing power, then the edge computing node with the highest transmission rate to that data source terminal without computing power will be designated as the first computing power mapping node. Specifically, assume data source D... j Lacking data capabilities, another Represents the set of all edge computing nodes, and another This indicates that the edge computing node is connected to the data source D. j The computing node with the highest transmission rate among them is:

[0104]

[0105] in, Indicates data source D j To any computing node N n The transmission rate will then be the data source D. j Data transfer from and will Dataset D in the first split segment j The computing power mapping node.

[0106] In one exemplary embodiment, the power large model includes multiple first computing power mapping nodes, and the method further includes:

[0107] Traverse each first computing power mapping node. If each first computing power mapping node satisfies the preset layer number constraint, determine that the first model splitting segment corresponding to each first computing power mapping node includes the first layer and the second layer of the power big model.

[0108] In the case where the first computing power mapping node does not meet the layer constraint, the first model splitting segment corresponding to each first computing power mapping node is determined to include the first layer of the power big model; the layer constraint is that the total time required for the first computing power mapping node to calculate the first and second layer data of the power big model is less than the preset maximum waiting time, and the memory capacity of the first computing power mapping node is greater than the capacity required for the first and second layer data of the power big model.

[0109] This application provides a method for dividing the number of layers contained in the first model split segment.

[0110] The layer constraint is as follows: when placing the first and second layers into the first split segment, determine whether all first computing power mapping nodes corresponding to the first split segment can meet the computing power operation and model parameter storage requirements of the split segment. Specifically, let W... d Z represents the sum of computational costs of any first-power mapping node in layers L1 and L2. d Let f represent the sum of the number of model parameters required to train the L1 and L2 layers. d MEM d These represent the node's computing power and storage capacity, respectively. This represents the maximum computational delay that the first split segment can tolerate.

[0111] The layer number constraint can be expressed as:

[0112]

[0113] In this embodiment of the application, firstly, let Let L represent the set of layers in a large model. N L1 represents the last layer, and L2 represents the first layer. Traverse each first computing power mapping node and check if each first computing power mapping node satisfies the above layer number constraint. If so, allocate the first and second layers as the first split segment to the first computing power mapping node. If any first computing power mapping node does not satisfy the above layer number constraint, only allocate the first layer as the first split segment to each first computing power mapping node. The first split segment containing the first layer can be represented as: Similarly, when the first segment contains both the first and second layers, it can be represented as: ,in, M represents the first computing power mapping node in the parallel first split segment, and M≥1 represents the number of the first computing power mapping nodes.

[0114] In one exemplary embodiment, such as Figure 3 As shown, the steps for training a large-scale power control model include:

[0115] S310 receives the forward propagation parameters of the previous computing power mapping node through the first layer of the current model segment of the target computing power mapping node.

[0116] The target computing power mapping node is any node other than the first computing power mapping node and the computing power mapping node of the last split segment. The first layer model is the first layer in the current model split segment.

[0117] In this embodiment of the application, each target computing power mapping node receives the forward propagation parameters of the previous split segment, and sequentially executes the forward propagation calculation of each micro-batch at this layer according to the time order of receiving the forward propagation parameters.

[0118] For the first computing power mapping node, since it does not have a previous split segment, each first computing power mapping node corresponding to the first split segment executes the forward propagation training of the first split segment in parallel for its own micro-batch. After the training is completed, the last layer model in the first split segment transmits the forward propagation parameters to the computing power mapping node of the second split segment through the communication network, and stores the training parameters of the segment on the local node.

[0119] S320, through the current model segment, according to the received forward propagation parameters, performs forward propagation training of the current segment of each micro-batch. The micro-batch is the batch size of each computing power mapping node in each round of training. The micro-batch is determined according to the number of the first computing power mapping nodes and the preset small batch data volume.

[0120] The small batch data size is a preset mini-batch, which is obtained by combining the small batch data size with the reciprocal of the number of first computing power mapping nodes (i.e., 1 / M). In this embodiment, the dataset for each first computing power mapping node to perform large model training comes from the local terminal node.

[0121] In this embodiment of the application, for any target computing power mapping node, the forward propagation parameters transmitted by the previous model segment are received through the first layer of the current model segment, thereby performing forward propagation training of each micro-batch through the current model segment.

[0122] S330, through the last layer of the current model split segment, passes the forward propagation parameters of the target computing power mapping node to the computing power mapping node corresponding to the next model split segment, and stores the training parameters of each micro-batch of the current split segment in the storage space of the target computing power mapping node.

[0123] S340, Repeat the above steps until the computing power mapping node corresponding to each split segment is traversed.

[0124] In this embodiment, for the computing power mapping model corresponding to the last split segment, since there is no next model difference segment relative to the last split segment, for the last split segment, it is only necessary to sequentially execute the forward propagation training of each micro-batch at this layer according to the time order of the received forward propagation parameters. This achieves traversal of the computing power mapping node corresponding to each split segment.

[0125] In this embodiment of the application, the above steps are repeated until the last model segment completes the forward propagation calculation.

[0126] In one exemplary embodiment, such as Figure 4 As shown, the method also includes:

[0127] S410, while traversing the computing power mapping nodes corresponding to each model split segment, receives the backpropagation parameters of the next model split segment through the last layer of the current split segment of the target computing power mapping node.

[0128] This application embodiment describes the backpropagation phase during the training process. For the target computing power mapping node, the backpropagation parameters of the next model segment are received through the last layer of the model. It can be understood that for the computing power mapping node corresponding to the last segment, there is no next model segment. Therefore, for the last segment, it is only necessary to transmit the backpropagation parameters calculated by itself to the previous segment.

[0129] S420, through the current model segment, according to the received backpropagation parameters, sequentially executes the backpropagation training of each micro-batch in the current model segment, and through the first layer of the current model segment, passes the backpropagation parameters to the computing power mapping node corresponding to the previous model segment.

[0130] In this embodiment of the application, for the last layer model of the target computing power mapping node, the backpropagation calculation is performed sequentially according to the order in which the backpropagation parameters arrive. After the backpropagation training of each micro-batch is completed, the backpropagation parameters of the first layer in the segment are transmitted to the computing power mapping node corresponding to the previous split segment through the communication network.

[0131] S430, repeat the above steps until the backpropagation parameters are passed to the first computing power mapping node corresponding to the first model split segment. The first computing power mapping node is used to perform backpropagation training of each layer in the first model split segment in the micro-batch corresponding to the first computing power mapping node when it receives the backpropagation parameters of the next model split segment, to obtain the model weight gradient, and send the model weight gradient to the large model gradient update controller.

[0132] In this embodiment of the application, when the first computing power mapping node receives the backpropagation parameters of the next split segment, it performs backpropagation training of the micro-batch in each layer of the segment. Based on the backpropagation parameters and the forward propagation parameters stored by the first computing power mapping node itself, it calculates the model weight gradient of the model weight of the segment and sends the model weight gradient to the large model gradient update controller.

[0133] Figure 5 This is a schematic diagram of micro-batch scheduling for large model training in one embodiment. In the diagram, D1, D2, D3, ..., D... M All of these are computing power mapping nodes for the first split segment. 1-1, 2-1, 3-1, ..., M-1 represent D1, D2, D3, ..., D... respectively. M The executed micro-batches, which are respectively from local batches D1, D2, D3, ..., D... M The data. The computing power mapping nodes D1, D2, D3, ..., D in the first model segment. M After training the first model segment is completed, the last layer model parameters of the local first segment are transmitted to the computing power mapping node N1 of the second model segment via a wireless network. Similarly, assuming that the computing power mapping node N1 of the second model segment g2 receives the forward propagation parameters from D1, D3, and D2 in the following order... M Given the parameters of D2, ..., N1 executes the forward propagation of the second stage in the following order: 1-2, 3-2, M-2, ..., 2-2. In the above representation of XY, X represents the micro-batch number, and Y represents the training stage Y. For example, after N1 completes micro-batch 1-2, it immediately transmits the last layer model parameters of that micro-batch in that stage to the computing power mapping node N2 of the third split segment g3 via the wireless network. Simultaneously, N1 executes the forward propagation training of micro-batch 3-2. Finally, assuming g... k This is the last split segment, and its computing power mapping nodes are N. k-1 Assume the model parameters arrive in the order from D1, D3, D... M Given the parameters D1, ..., D2, the node performs forward propagation training in the order 1-k, 3-k, Mk, ..., 2-k, followed by backward propagation calculations of 2-k, ..., Mk, 3-k, 1-k. Once the micro-batch 2-k backward propagation training is complete, it is immediately passed to the next stage (e.g., g) via the wireless network. i The computing power mapping node N i Send the backpropagation parameters of the first layer in the micro-batch segment, and so on, to complete a full training cycle.

[0134] In the diagram, the horizontal axis represents time, the vertical axis represents different model split segments / computing power mapping nodes, and the green rectangles represent the communication latency required for data transmission between different computing power mapping nodes, because the communication latency varies between different nodes.

[0135] The above diagram can be further abstracted as follows: Figure 6 In this context, solid lines represent forward propagation calculations, and dashed lines represent backward propagation calculations.

[0136] In one exemplary embodiment, the method further includes:

[0137] When the model weight gradient is sent to the large model gradient update controller, the target update gradient parameters are calculated by the large model gradient update controller based on the gradient corresponding to each first computing power mapping node.

[0138] The target update gradient is sent to each first computing power mapping node through the large model gradient update controller. The first computing power mapping node is used to update the model parameters in the corresponding first model split segment.

[0139] In this embodiment of the application, the large model gradient update controller performs gradient calculation after checking that each first computing power mapping node has completed gradient reporting. If gradient averaging is possible, the target update gradient can be obtained.

[0140] The large model gradient update controller sends the target update gradient to each of the first computing power mapping nodes. After receiving the target update gradient parameters from the large model gradient update controller, each first computing power mapping node performs a parameter update operation.

[0141] In one exemplary embodiment, the method further includes:

[0142] The training process of the power big data model is monitored based on a preset monitoring period. When the training data of the power big data model meets the preset alarm requirements, a new split mapping strategy is formulated, and the model iterative training is performed according to the new split mapping strategy. The alarm requirements include that at least one computing power mapping stage is in an unreachable state, or that the total latency of a single round of training of the power big data model is greater than the preset maximum latency threshold.

[0143] In this embodiment, the decision controller periodically monitors the performance of large model training and the status of each computing power mapping node according to a preset monitoring period. When at least one computing power mapping node is in an unreachable state, or the total training latency of a certain round exceeds the maximum latency threshold, the splitting mapping strategy of the large model is redefined, and collaborative parallel pipeline training of the model is executed according to the new splitting mapping strategy. The total training latency is the sum of the training latencies of all stages in a certain round. The training latency further includes computation latency and data transmission latency. The training latency of the first model split segment is the minimum of the computation latency of all computing power mapping nodes in that stage and the data transmission latency experienced when transferring model parameters to the next model split segment. The computation latency of other model split segments is the sum of the computation latencies of each micro-batch in the same round of training in that stage, and the data transmission latency is the data transmission latency experienced when transferring the model parameters of the last micro-batch in that split segment to the next model split segment.

[0144] This application also provides a preferred embodiment of a decomposition learning method for large power models, such as... Figure 7 This is a flowchart illustrating the decomposition learning method for a large power model in a preferred embodiment.

[0145] S710, establish multiple micro-batches. The size of each micro-batch is a preset mini-batch × 1 / M. The micro-batches are determined for each round of training when the computing power mapping nodes in the first model split segment perform large model training. The dataset for each computing power mapping node in the first model split segment when performing large model training is limited to come from the local terminal node.

[0146] S720: Each computing power mapping node in the first model split segment performs forward propagation training of its own micro-batch and passes the model parameters of the last layer in the first model split segment to the second model split segment, and stores the training parameters of the segment on the local node.

[0147] S730, the computing power mapping nodes from the second model split segment to the second-to-last model split segment sequentially perform forward propagation calculations and pass the model parameters of the last layer of the segment to the computing power mapping node of the next split segment. The computing power mapping nodes from the second model split segment to the second-to-last model split segment sequentially perform model training of each layer of each micro-batch in the segment according to the time order of the forward propagation parameters received from the previous split segment, and store the training parameters of each micro-batch of the segment at the computing power mapping node.

[0148] S740, the computing power mapping node of the last model segment executes the forward propagation training of each micro-batch in the order of receiving the forward propagation parameters of the previous model segment.

[0149] S750, after the current round of forward propagation training is completed, the last model segment is backpropagated in reverse order. After the backpropagation training of each micro-batch is completed, the backpropagation parameters of the first layer in the segment are passed to the previous segment.

[0150] S760, the computing power mapping node from the second model split segment to the penultimate model split segment, executes the backward propagation training of each micro-batch in the layer in the order of receiving the candidate propagation parameters of the next model split segment, and passes the backward propagation parameters of the first layer in the segment to the previous split segment.

[0151] S770: After receiving the backpropagation parameters of the next model segment, the first computing power mapping node performs the corresponding micro-batch backpropagation training in this layer and sends the calculated gradient parameters to the large model gradient update controller.

[0152] After receiving all gradient parameters from each of the first model split segments, the S780 large model gradient update controller performs gradient parameter aggregation and update to obtain the target update gradient parameters, and sends the target update gradient parameters to each first computing power mapping node for gradient update.

[0153] S790: Determine whether the large model training has reached the accuracy requirement, or whether the training round threshold has been reached. If yes, end the training; otherwise, jump to S720.

[0154] Similarly, in the inference phase after model training, the system receives power tasks from the smart grid, collects corresponding power task processing data through the terminal based on the power tasks, performs model inference on the power task processing data based on the fully trained smart grid model, and inputs the power task processing data into the computing power mapping node of the first segment. This node performs the calculation of the first segment (forward propagation only), obtains intermediate results, and passes them to the node of the second segment; and so on, until the node of the last segment outputs the final inference result.

[0155] This application provides an embodiment that supports distributed training of large-scale smart grid models with ultra-large parameters and distributed data sources at the network edge, reducing learning latency and meeting the data privacy and security requirements of smart grids. Furthermore, the split-mapping strategy developed in this application ensures the timeliness of large-scale model training at the network edge. Moreover, by aggregating the backpropagation parameters of the last layer in each training round of the large model to the large model gradient update controller for gradient parameter aggregation and updating, the synergy of the distributed parallel training of the large model in micro-batches is improved, meeting the AI ​​accuracy requirements of large-scale model training. In addition, the decision controller periodically monitors the performance and system status of the large model training and updates the strategy when performance deteriorates or computing nodes become unreachable, further providing continuous assurance for the distributed training of the power grid large model at the network edge.

[0156] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.

[0157] Based on the same inventive concept, this application also provides a power model splitting and learning device for implementing the power model splitting and learning method described above. The solution provided by this device is similar to the solution described in the above method; therefore, the specific limitations of the one or more power model splitting and learning device embodiments provided below can be found in the limitations of the power model splitting and learning method described above, and will not be repeated here.

[0158] In one exemplary embodiment, a decomposition learning apparatus for a large power model is provided, comprising:

[0159] The receiving module is used to receive power task processing data corresponding to the power tasks issued by the smart grid.

[0160] The computation module is used to perform model inference processing on power task processing data based on a fully trained power big model to obtain inference results. The training steps of the power big model include: responding to the received power big model training task, determining the splitting and mapping strategy of the power big model based on the node computing capabilities and network communication latency of multiple edge computing nodes; and performing collaborative parallel model iterative training based on the splitting and mapping strategy through each computing power mapping node and the big model gradient update controller until the model converges to obtain a fully trained power big model. The splitting and mapping strategy represents splitting the power big model into multiple segments, allocating different segments to different computing power mapping nodes in the edge computing nodes, and training the corresponding segments based on the computing resources of the computing power mapping nodes.

[0161] The modules in the aforementioned large-scale power model decomposition learning device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the computer device's memory as software, so that the processor can call and execute the corresponding operations of each module.

[0162] In one exemplary embodiment, a decomposition learning system for a large power model is provided, such as... Figure 8 As shown, it includes a decision controller, a large model gradient update controller, edge computing nodes, and a power IoT terminal; the power IoT terminal communicates with the corresponding edge computing node, the decision controller communicates with the edge computing node through a base station, the large model gradient update controller communicates with the edge computing node through a base station, and the edge computing nodes communicate with each other.

[0163] The power Internet of Things (IoT) terminal is used to collect corresponding power task processing data from power equipment based on the power tasks issued by the smart grid.

[0164] A decision controller is used to perform the steps of any of the methods described above.

[0165] In one exemplary embodiment, a computer device is provided, including a memory and a processor, the memory storing a computer program that, when executed by the processor, implements any of the power large model decomposition learning methods described above.

[0166] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements any of the power large model decomposition learning methods described above.

[0167] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements any of the power large model decomposition learning methods described above.

[0168] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, artificial intelligence (AI) processors, etc., and are not limited to these.

[0169] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.

[0170] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.

Claims

1. A decomposition learning method for a large-scale power model, characterized in that, Applied to a decision controller, the method includes: Obtain power tasks issued by the smart grid; According to the power task, the power task processing instruction is issued to the corresponding target node. The power task processing instruction is used to instruct the target node to perform model inference calculation based on the collected power task processing data and the fully trained power large model to obtain the inference result. The steps of controlling the training of the large power model include: responding to the received large power model training task, determining the splitting and mapping strategy of the large power model based on the node computing capabilities and network communication latency of multiple edge computing nodes; controlling each computing power mapping node and the large model gradient update controller to perform collaborative parallel model iterative training based on the splitting and mapping strategy until the model converges and a fully trained large power model is obtained; the splitting and mapping strategy represents splitting the large power model into multiple segments, assigning different segments to different computing power mapping nodes in the edge computing nodes, and training the corresponding segments based on the computing resources of the computing power mapping nodes.

2. The method according to claim 1, characterized in that, The steps for establishing the split mapping strategy include: In the edge computing nodes, determine the current computing power mapping node that meets the training requirements of the current model segment; wherein, the storage space of the current computing power mapping node is greater than a preset storage space threshold, and the data transmission delay between the current computing power mapping node and the computing power mapping node of the previous model segment is not higher than a preset communication delay threshold, and the computing resources of the current computing power mapping node are greater than a preset computing resource threshold. The number of layers of the large power model included in the current model split segment is determined based on the current computing power mapping node; The splitting and mapping strategy is determined based on the model splitting segments of the power big model and the computing power mapping nodes corresponding to each model splitting segment.

3. The method according to claim 2, characterized in that, The step of determining the number of layers of the large power model included in the current model segment based on the current computing power mapping node includes: The current first layer number is determined based on the maximum number of layers in the storage space of the current computing power mapping node. The current second layer number is determined based on the maximum number of layers that meet the preset expected training duration. Based on the current first layer number and the current second layer number, determine the number of layers of the power large model included in the current model split segment.

4. The method according to any one of claims 1 to 3, characterized in that, The power large model training task includes a set of data sources, which are collected by the corresponding data source terminals. The method further includes: When the data source terminal has computing power, the data source terminal is used as the first computing power mapping node corresponding to the first model segment in the power big model; the first computing power mapping node is an edge computing node used to perform the first segment model calculation of the power big model; In the absence of computing power at the data source terminal, the edge computing node with the highest transmission rate to the data source terminal will be used as the first computing power mapping node corresponding to the first model splitting segment in the power big model.

5. The method according to claim 4, characterized in that, The large-scale power model includes multiple first-level computing power mapping nodes, and the method further includes: Traverse each first computing power mapping node. If each first computing power mapping node satisfies the preset layer number constraint, determine that the first model splitting segment corresponding to each first computing power mapping node includes the first layer and the second layer of the power big model. If a first computing power mapping node does not meet the layer number constraint, the first model splitting segment corresponding to each first computing power mapping node is determined to include the first layer of the power big model; the layer number constraint is that the total time required for the first computing power mapping node to calculate the first and second layer data of the power big model is less than the preset maximum waiting time, and the memory capacity of the first computing power mapping node is greater than the capacity required for the first and second layer data of the power big model.

6. The method according to any one of claims 1 to 3, characterized in that, The steps for training the large power model include: The forward propagation parameters of the previous computing power mapping node are received through the first layer of the current model segment of the target computing power mapping node. Based on the received forward propagation parameters, the forward propagation training of the current segment of the model is performed using the current segment. The micro-batch is the batch size of each computing power mapping node in each round of training. The micro-batch is determined based on the number of the first computing power mapping nodes and the preset small batch data volume. The forward propagation parameters of the target computing power mapping node are passed to the computing power mapping node corresponding to the next model split segment through the last layer of the current model split segment, and the training parameters of each micro-batch of the current split segment are stored in the storage space of the target computing power mapping node. Repeat the above steps until the computing power mapping node corresponding to each split segment has been traversed.

7. The method according to claim 6, characterized in that, The method further includes: When traversing the computing power mapping nodes corresponding to each model split segment, the backpropagation parameters of the next model split segment are received through the last layer of the current split segment of the target computing power mapping node. Based on the received backpropagation parameters, the backpropagation training of each micro-batch in the current model segment is executed sequentially through the current model segment. The backpropagation parameters are then passed to the computing power mapping node corresponding to the previous model segment through the first layer of the current model segment. Repeat the above steps until the backpropagation parameters are passed to the first computing power mapping node corresponding to the first model split segment. The first computing power mapping node is used to perform backpropagation training of each layer in the first model split segment in the micro-batch corresponding to the first computing power mapping node when it receives the backpropagation parameters of the next model split segment, to obtain the model weight gradient, and send the model weight gradient to the large model gradient update controller.

8. The method according to claim 7, characterized in that, The method further includes: When the model weight gradient is sent to the large model gradient update controller, the target update gradient parameters are calculated by the large model gradient update controller based on the gradient corresponding to each first computing power mapping node. The target update gradient is sent to each first computing power mapping node through the large model gradient update controller. The first computing power mapping node is used to update the model parameters in the corresponding first model split segment.

9. The method according to any one of claims 1 to 3, characterized in that, The method further includes: The training process of the power big data model is monitored based on a preset monitoring period. When the training data of the power big data model meets the preset alarm requirements, a new split mapping strategy is formulated, and model iterative training is performed according to the new split mapping strategy. The alarm requirements include that at least one computing power mapping stage is in an unreachable state, or that the total latency of a single round of training of the power big data model is greater than the preset maximum latency threshold.

10. A decomposition learning system for a large-scale power model, characterized in that, The system includes a decision controller, a large model gradient update controller, edge computing nodes, and a power IoT terminal; the power IoT terminal is communicatively connected to the corresponding edge computing node, the decision controller is communicatively connected to the edge computing node through a base station, the large model gradient update controller is communicatively connected to the edge computing node through a base station, and the edge computing nodes are communicatively connected to each other. The power Internet of Things terminal is used to collect corresponding power task processing data according to the power tasks issued by the smart grid. The decision controller is configured to perform the steps of the method as described in any one of claims 1 to 9.