Distributed power model updating method and device based on model increment training and electronic equipment

Through singular value decomposition and low-rank decomposition technology, the important parameter matrix of distributed power nodes is extracted and only non-important parameters are updated, which solves the problems of real-time and accuracy in traditional distributed power model training, and achieves efficient model updates and accuracy improvements.

CN120408010AActive Publication Date: 2025-08-01STATE GRID DIGITAL TECHNOLOGY HOLDING CO LTD +1
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510906292.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-02
Publication Date
2025-08-01
Estimated Expiration
2045-07-02

AI Technical Summary

Technical Problem

The traditional distributed power model training method has challenges in real-time, computing resources, energy consumption and communication delay, especially as the amount of power data and model scale continues to increase, it is difficult to effectively update and maintain model accuracy.

Method used

Using a model incremental training method, the local model parameter matrix of distributed power nodes is processed through singular value decomposition and low-rank decomposition technology, the important parameter matrix is extracted and the difference matrix is decomposed. Only non-important parameters are updated, and the model is updated in combination with the global scheduling server.

Benefits of technology

In the case of insufficient incremental data, maintain the original adaptability of the model, avoid drift, improve model update efficiency, and improve model accuracy, and reduce data and calculation amount.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120408010A_ABST
    Figure CN120408010A_ABST
Patent Text Reader

Abstract

The invention provides a distributed power model updating method and device based on model incremental training and electronic equipment. According to the implementation scheme, when local power incremental data of distributed power nodes are insufficient, singular value decomposition is carried out on a local model parameter matrix of a local power model in the distributed power nodes, and a left singular matrix, a diagonal matrix and a right singular matrix are obtained; respectively extracting corresponding column vectors from the left singular matrix and the right singular matrix based on main singular values in the diagonal matrix to obtain an important parameter matrix; and determining a non-important parameter matrix based on a difference matrix between the local model parameter matrix and the important parameter matrix, and performing low-rank decomposition on the non-important parameter matrix to obtain a local first low-rank parameter matrix and a local second low-rank parameter matrix, thereby performing global model updating and updating the local power model of each power node. By adopting the method, the model training speed and precision can be improved, and the model is prevented from drifting.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of computers and power technologies, and in particular, to a distributed power model update method, device, and electronic device based on model incremental training. Background Art

[0002] In some power technology fields, arranging multiple distributed power nodes and a global node can improve the prediction accuracy of power-related indicators or functions, and thus improve the accuracy of distributed power scheduling.

[0003] Among them, in each distributed power node, the local power model is trained using local power data, and then the parameters of the trained local power model are sent to the global node. The global node aggregates the local power model parameters of each distributed power node to obtain global power model parameters, and returns the global power model parameters to each distributed power node to update the local power model of each distributed power node.

[0004] However, with the continuous increase in the amount of power data and the scale of the model, traditional model training methods pose many challenges in terms of real-time performance, computing resources, energy consumption, and communication latency. Summary of the Invention

[0005] The present invention provides a distributed power model update method, device, and electronic device based on model incremental training, which can solve at least one of the above technical problems.

[0006] According to one aspect of the present invention, there is provided a distributed power model update method based on model incremental training, including: When the amount of local power incremental data of the distributed power node is less than a preset threshold, performing singular value decomposition on the local model parameter matrix of the local power model in the distributed power node to obtain a left singular matrix, a diagonal matrix, and a right singular matrix; Based on the main singular values in the diagonal matrix, extracting corresponding column vectors from the left singular matrix and the right singular matrix respectively to obtain an important parameter matrix; Based on the difference matrix between the local model parameter matrix and the important parameter matrix, determining an unimportant parameter matrix, and performing low-rank decomposition on the unimportant parameter matrix to obtain a local first low-rank parameter matrix and a local second low-rank parameter matrix; Based on the local power incremental data, updating the local first low-rank parameter matrix and the local second low-rank parameter matrix, and sending the update result to the global scheduling server, so that the global scheduling server updates the global first low-rank parameter matrix and the global second low-rank parameter matrix; Update the local power model of the distributed power node based on the product between the updated global first low-rank parameter matrix and the updated global second low-rank parameter matrix from the global scheduling server, and the important parameter matrix.

[0007] According to another aspect of the present invention, there is provided a distributed power model update device based on model incremental training, including: A singular value decomposition module, configured to perform singular value decomposition on the local model parameter matrix in the distributed power node to obtain a left singular matrix, a diagonal matrix, and a right singular matrix when the data volume of the local power incremental data in the distributed power node is less than a preset threshold; An important parameter determination module, configured to extract corresponding column vectors from the left singular matrix and the right singular matrix respectively based on the main singular values in the diagonal matrix to obtain an important parameter matrix; A low-rank parameter determination module, configured to determine an unimportant parameter matrix based on the difference matrix between the local model parameter matrix and the important parameter matrix, and decompose the unimportant parameter matrix to obtain a local first low-rank parameter matrix and a local second low-rank parameter matrix; A global model update module, configured to update the local first low-rank parameter matrix and the local second low-rank parameter matrix in the local model parameter matrix based on the local power incremental data of the distributed power node, and send the update result to the global scheduling server, so that the global scheduling server updates the global first low-rank parameter matrix and the global second low-rank parameter matrix; A local model update module, configured to update the local power model of the distributed power node based on the product between the updated global first low-rank parameter matrix and the updated global second low-rank parameter matrix from the global scheduling server, and the important parameter matrix.

[0008] Adopting the technical solution of the present invention, when the local power increment data of the distributed power node is insufficient, perform singular value decomposition on the local model parameter matrix of the local power model in the distributed power node to obtain a left singular matrix, a diagonal matrix, and a right singular matrix; based on the main singular values in the diagonal matrix, extract the corresponding column vectors from the left singular matrix and the right singular matrix respectively to obtain an important parameter matrix; determine an unimportant parameter matrix based on the difference matrix between the local model parameter matrix and the important parameter matrix, and perform low-rank decomposition on the unimportant parameter matrix to obtain a local first low-rank parameter matrix and a local second low-rank parameter matrix; based on the local power increment data of the distributed power node, update the local first low-rank parameter matrix and the local second low-rank parameter matrix in the local model parameter matrix. In this way, in the case of insufficient increment data, the important parameters in the local model parameters can be retained without participating in the update, and only the unimportant parameters are updated, which can effectively maintain the original task adaptability of the model, avoid drift of the local power model, and reduce the data volume and improve the model update efficiency after performing low-rank decomposition on the unimportant parameter matrix and then updating. Moreover, send the update result to the global scheduling server so that the global scheduling server updates the global first low-rank parameter matrix and the global second low-rank parameter matrix. Subsequently, based on the product between the updated global first low-rank parameter matrix and the updated global second low-rank parameter matrix from the global scheduling server, and the important parameter matrix, update the local power model of the distributed power node. In this way, by aggregating the local power model parameters of each distributed node to obtain global model parameters and using them to update the local power model of the distributed node, the model accuracy of the local power model of each distributed node can be improved, and drift of the local power model can be avoided.

[0009] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present invention, nor is it used to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] The drawings are used to better understand the solution and do not constitute a limitation to the present invention. Among them: Figure 1 is a flowchart of a method for updating a distributed power model based on model incremental training according to an embodiment of the present invention; Figure 2 is a schematic diagram of a model update process according to an embodiment of the present invention; Figure 3 is a structural block diagram of a device for updating a distributed power model based on model incremental training according to an embodiment of the present invention; Figure 4 is a block diagram of an electronic device for implementing the method according to an embodiment of the present invention. Detailed implementation manners

[0011] The following describes exemplary embodiments of the present invention with reference to the accompanying drawings. Various details of the embodiments of the present invention are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope of the present invention. Similarly, descriptions of well-known functions and structures are omitted below for clarity and conciseness.

[0012] Figure 1 is a flowchart of a distributed power model update method based on model incremental training according to an embodiment of the present invention.

[0013] As Figure 1 shown, the distributed power model update method based on model incremental training includes: S110, when the data volume of the local power increment data of the distributed power node is less than a preset threshold, perform singular value decomposition on the local model parameter matrix in the distributed power node to obtain a left singular matrix, a diagonal matrix, and a right singular matrix; S120, based on the main singular values in the diagonal matrix, extract corresponding column vectors from the left singular matrix and the right singular matrix respectively to obtain an important parameter matrix; S130, based on the difference matrix between the local model parameter matrix and the important parameter matrix, determine an unimportant parameter matrix, and perform low-rank decomposition on the unimportant parameter matrix to obtain a local first low-rank parameter matrix and a local second low-rank parameter matrix; S140, based on the local power increment data of the distributed power node, update the local first low-rank parameter matrix and the local second low-rank parameter matrix in the local model parameter matrix, and send the update result to the global scheduling server so that the global scheduling server updates the global first low-rank parameter matrix and the global second low-rank parameter matrix; S150, based on the product of the updated global first low-rank parameter matrix and the updated global second low-rank parameter matrix from the global scheduling server, and the important parameter matrix, update the local power model of the distributed power node.

[0014] In practical applications, the power system may include multiple distributed power nodes and a global scheduling server. As Figure 2 shown, the global scheduling server is Figure 2The scheduling server in it, the distributed power nodes include Node 1 to Node 4. Of course, other nodes can also be included, but they are not shown in the figure. Each node uploads status data, such as parameters like load, temperature, and power consumption, to the scheduling server. At the same time, each node also uploads the locally differentially updated model parameters to the scheduling server. The scheduling server calculates the weights of each node using the status data of each node. And for the nodes with the ARM architecture, the scheduling coefficient is used to update the weights of these nodes. Thus, using the weights of each node, the model parameters provided by each node are aggregated to obtain global model parameters, which are then distributed to each node to update the local model parameters of each node.

[0015] It can be understood that the method provided by the embodiments of the present invention can be applied to any distributed power nodes. Each distributed power node can adopt an edge device equipped with an ARM chip, enabling it to undertake more data preprocessing, feature extraction, and some lightweight inference tasks, and using the low-power consumption characteristics for long-term online caching. Since the important parameter matrix remains unchanged, the ARM side only needs to store this important parameter matrix during the initial loading, and the pressure on communication and energy consumption is relatively small.

[0016] It can be understood that the amount of local power increment data is less than the preset data volume threshold, indicating that the local power increment data of the distributed power node is insufficient.

[0017] It can be understood that the initial local power models of each distributed power node are generally the same, and the local power models obtained after the process of local update, global update, and then local update are also the same on each node.

[0018] Exemplarily, perform singular value decomposition on the local model parameter matrix in the distributed power node to obtain the left singular matrix, diagonal matrix, and right singular matrix, specifically as follows: For the local model parameter matrix , the singular value decomposition result is: ; Among them, is the left singular matrix, is the diagonal matrix, is the right singular matrix. Among them, the left singular matrix and the right singular matrix are orthogonal matrices, and their column vectors are singular vectors.

[0019] In one example, the larger singular value in the diagonal matrix is taken as the main singular value. The column vectors corresponding to the main singular value in the left singular matrix and the right singular matrix are combined to obtain an important parameter matrix. When the local incremental data is insufficient, the important parameter matrix is not updated by training, and only the non-important parameter matrix is updated by training, which can avoid the drift or overfitting of the local power model.

[0020] In one example, statistics are performed according to the weights of the local power model, and the parameters with a large change amplitude during the model training process are selected. According to these parameters, their positions in the diagonal matrix are determined, and the singular value at this position is taken as the main singular value. This can be combined with the larger singular value in the above diagonal matrix as the main singular value. The column vectors corresponding to the main singular value in the left singular matrix and the right singular matrix are combined to obtain an important parameter matrix. When the local incremental data is insufficient, the important parameter matrix is not updated by training, and only the non-important parameter matrix is updated by training, which can avoid the drift or overfitting of the local power model.

[0021] In one example, according to the local power model in the previous training process, the model parameters with a gradient sensitivity greater than the preset threshold are determined, the position of the model parameter in the diagonal matrix is determined, and the singular value at this position is taken as the main singular value. The main singular value here can be combined with the larger singular value in the above diagonal matrix as all the main singular values. In this way, the important parameter matrix can include parameters with a relatively small singular value but actually a relatively high gradient sensitivity. Furthermore, when updating the non-important parameters, these parameters with a relatively small singular value but actually a relatively high gradient sensitivity can be avoided from participating in the model update, thereby further avoiding the drift or overfitting of the model.

[0022] Exemplarily, the difference matrix between the local model parameter matrix and the important parameter matrix is used as the non-important parameter matrix. The low-rank decomposition of the non-important parameter matrix can be: ; where, is the non-important parameter matrix, is the local first low-rank parameter matrix, is the local second low-rank parameter matrix, , , .

[0023] Exemplarily, the local first low-rank parameter matrix and the local second low-rank parameter matrix in the local model parameter matrix are continuously iteratively updated, and the local first low-rank parameter matrix updated in the last two times is subtracted to obtain a local first low-rank parameter difference matrix, and the local second low-rank parameter matrix updated in the last two times is subtracted to obtain a local second low-rank parameter difference matrix. In this way, since the data volumes of the local first low-rank parameter difference matrix and the local second low-rank parameter difference matrix are small, transmitting them to the global scheduling server can improve the transmission speed, and also, due to the small data volume, the calculation speed of global model parameter aggregation can be improved.

[0024] Exemplarily, the global scheduling server globally aggregates the local first low-rank parameter difference matrix and the local second low-rank parameter difference matrix of each distributed power node, and uses the global aggregation result to update the global first low-rank parameter matrix and the global second low-rank parameter matrix stored in the server, and sends the updated global first low-rank parameter matrix and the updated global second low-rank parameter matrix to each distributed power node.

[0025] Exemplarily, the matrix obtained by adding the product between the updated global first low-rank parameter matrix and the updated global second low-rank parameter matrix to the important parameter matrix is used as the local model parameter matrix of the local power model of the distributed power node, specifically as follows: ; Among them, is the local model parameter matrix, is the important parameter matrix, and represent the global first low-rank parameter matrix and the global second low-rank parameter matrix, represents the updated non-important parameter matrix.

[0026] Exemplarily, when the data volume of the local power increment data of the distributed power node is greater than a preset threshold, the local model parameter matrix of the local power model in the distributed power node is low-rank decomposed to obtain a local first low-rank parameter matrix and a local second low-rank parameter matrix. And based on these two matrices, the model training processes of step S140 and step S150 are performed. In this way, when the data volume is sufficient, the local model parameter matrix is directly low-rank decomposed, and the results of the low-rank decomposition are used to update the local model parameters. In this way, not only can the model accuracy and the model training speed be improved, but also, since the local power increment data is sufficient, there will be no model drift and overfitting.

[0027] According to the above embodiments, in the case where the local power increment data of the distributed power node is insufficient, singular value decomposition is performed on the local model parameter matrix of the local power model in the distributed power node to obtain a left singular matrix, a diagonal matrix, and a right singular matrix; based on the main singular values in the diagonal matrix, corresponding column vectors are respectively extracted from the left singular matrix and the right singular matrix to obtain an important parameter matrix; based on the difference matrix between the local model parameter matrix and the important parameter matrix, an unimportant parameter matrix is determined, and low-rank decomposition is performed on the unimportant parameter matrix to obtain a local first low-rank parameter matrix and a local second low-rank parameter matrix; based on the local power increment data of the distributed power node, the local first low-rank parameter matrix and the local second low-rank parameter matrix in the local model parameter matrix are updated. In this way, in the case of insufficient increment data, important parameters in the local model parameters can be retained without participating in the update, and only unimportant parameters are updated, which can effectively maintain the original task adaptability of the model, avoid drift of the local power model, and reduce the data volume and improve the model update efficiency after performing low-rank decomposition on the unimportant parameters. And, the update result is sent to the global scheduling server so that the global scheduling server updates the global first low-rank parameter matrix and the global second low-rank parameter matrix. Subsequently, based on the product between the updated global first low-rank parameter matrix and the updated global second low-rank parameter matrix from the global scheduling server, and the important parameter matrix, the local power model of the distributed power node is updated. In this way, by aggregating the local power model parameters of each distributed node to obtain global model parameters and using them to update the local power model of the distributed node, the model accuracy of the local power model of each distributed node can be improved, and drift of the local power model can be avoided.

[0028] In one embodiment, based on the main singular values in the diagonal matrix, corresponding column vectors are respectively extracted from the left singular matrix and the right singular matrix to obtain an important parameter matrix, including: sorting the diagonal singular values in the diagonal matrix from largest to smallest, and intercepting the first N diagonal singular values as the main singular values, where N is a positive integer greater than 1; determining the important parameter matrix based on the singular vectors in the left singular matrix and the singular vectors in the right singular matrix corresponding to each main singular value.

[0029] For example, determine the column positions of the main singular values, extract the left singular vectors at these column positions in the left singular matrix, and extract the right singular vectors at these column positions in the right singular matrix. The extracted left singular vectors and right singular vectors are respectively formed into matrices, and combined with the matrix formed by the main singular values to perform reverse singular value decomposition to obtain the important parameter matrix, where the number of rows and columns of the important parameter matrix is the same as that of the local model parameter matrix.

[0030] In this way, subtracting the important parameter matrix from the local model parameter matrix subsequently can obtain the unimportant parameter matrix, i.e., the aforementioned difference matrix.

[0031] According to the above embodiment, taking the singular values with larger numerical values as the main singular values, and combining the singular vectors corresponding to these main singular values in the left singular matrix and the right singular matrix to obtain the important parameter matrix. In this way, subtracting the important parameter matrix from the local model parameter matrix can obtain the unimportant parameter matrix. Only updating the local model for the unimportant parameter matrix can avoid model drift or overfitting when the training data of the local power model is insufficient.

[0032] In one embodiment, based on the main singular values in the diagonal matrix, extracting the corresponding column vectors from the left singular matrix and the right singular matrix respectively to obtain the important parameter matrix, including: determining the first position of the first power model parameter corresponding in the diagonal matrix according to the first power model parameter in the local power model whose gradient sensitivity is greater than the preset threshold; determining the main singular value based on the singular value at the first position; determining the important parameter matrix based on the singular vectors of each main singular value in the left singular matrix and the singular vectors in the right singular matrix.

[0033] It can be understood that the first power model parameter can include one or more, so the first position can be one or more, and thus the main singular value can include one or more.

[0034] It can be understood that in the case of insufficient training data, updating the model parameters with too high gradient sensitivity will cause the model to drift, thus affecting the model accuracy. Therefore, in this example, the first position of these model parameters corresponding in the diagonal matrix can be used, and the singular value at this position is used as the main singular value. This main singular value can be combined with the above-mentioned first N diagonal singular values as the final main singular value. In this way, even if the singular value of a certain model parameter is too small but its gradient sensitivity is too high, it is still regarded as an important parameter. In this way, in the case of insufficient training data, avoiding important parameters and only training unimportant parameters can improve the model accuracy while avoiding model drift.

[0035] In one embodiment, the above determining the main singular value based on the singular value at the first position includes: sorting the diagonal singular values in the diagonal matrix from large to small, and determining the first N diagonal singular values, where N is a positive integer greater than 1; obtaining the main singular value based on the first N diagonal singular values and the singular value at the first position.

[0036] Exemplarily, taking the union of the first N diagonal singular values and the singular value at the first position to obtain the main singular value.

[0037] Understandably, for the singular value at the first position, its value can be small and may not be arranged among the top N diagonal singular values. However, since it corresponds to a model parameter with an overly large gradient sensitivity, the singular value at the first position and the top N diagonal singular values need to be regarded as the main singular values. In this way, the important parameter matrix can include not only the parameters with large singular value magnitudes but also the parameters with small singular value magnitudes but actually high gradient sensitivities. Thus, when updating unimportant parameters, it is possible to avoid these parameters with small singular value magnitudes but actually high gradient sensitivities and the parameters with large singular value magnitudes from participating in model updates, further avoiding model drift or overfitting.

[0038] In one implementation, the distributed power node operates in a computing power platform with an acceleration card. Based on the local power increment data of the distributed power node, the local first low-rank parameter matrix and the local second low-rank parameter matrix are updated, including: Dividing the local power increment data into multiple batches of data and performing the following local training batch by batch: training the local power model based on the batch data to obtain a local loss; and updating the local first low-rank parameter matrix and the local second low-rank parameter matrix based on the gradient information of the local loss and the hardware characteristic adjustment efficiency factor corresponding to the computing power platform to obtain the updated local first low-rank parameter matrix and the updated local second low-rank parameter matrix; In the case where the number of local training times does not reach the preset requirement, updating the local power model based on the product of the updated local first low-rank parameter matrix and the updated local second low-rank parameter matrix and the important parameter matrix for the next batch of local training; In the case where the number of local training times reaches the preset requirement, determining the local first low-rank parameter difference matrix based on the difference between the updated local first low-rank parameter matrix and the local first low-rank parameter matrix before update, and determining the local second low-rank parameter difference matrix based on the difference between the updated local second low-rank parameter matrix and the local second low-rank parameter matrix before update.

[0039] Understandably, the computing power platform can use hardware such as Ascend 910B, Kunpeng 920 CPU, or ARM chips for computing, which can accelerate the model training process and improve training accuracy.

[0040] Exemplarily, updating the local first low-rank parameter matrix and the local second low-rank parameter matrix based on the gradient information of the local loss and the hardware characteristic adjustment efficiency factor corresponding to the computing power platform to obtain the updated local first low-rank parameter matrix and the updated local second low-rank parameter matrix is as follows: ; ; Among them, is the local first low-rank parameter matrix, is the gradient information corresponding to the local first low-rank parameter matrix, is the local second low-rank parameter matrix, is the gradient information corresponding to the local second low-rank parameter matrix, represents the hardware characteristic adjustment efficiency factor, represents the learning rate.

[0041] Exemplarily, the hardware characteristic adjustment efficiency factor can be dynamically changed according to conditions such as the temperature, load, and bandwidth of the acceleration card. For example, the value of the characteristic adjustment efficiency factor is greater than or equal to 1. When it is detected that the hardware is in a high-temperature state or an overloaded state, the value of the hardware characteristic adjustment efficiency factor can be converged towards 1 to ensure numerical stability. Otherwise, the value of the hardware characteristic adjustment efficiency factor can be appropriately increased based on 1 to accelerate training.

[0042] Exemplarily, the product of the updated local first low-rank parameter matrix and the updated local second low-rank parameter matrix is added to the important parameter matrix, and the resulting matrix is the model parameter of the local power model updated in this batch. In this way, it can be used for the next batch of local training.

[0043] Exemplarily, the preset requirement can be that the number of local training times reaches the total number of batches of the above batch of data. Or it reaches a preset training time threshold.

[0044] According to the above embodiments, by continuously updating the model, taking the difference between the last two updated local first low-rank parameter matrices as the local first low-rank parameter difference matrix, and taking the difference between the last two updated local second low-rank parameter matrices as the local second low-rank parameter difference matrix, a difference matrix with extremely small data volume can be obtained. Transmitting this to the global scheduling server can minimize the data volume as much as possible and also minimize the computational amount of the model parameters during global aggregation.

[0045] In one embodiment, the above sending the update result to the global scheduling server includes: determining the global synchronization period based on the network bandwidth and network latency between each distributed power node and the global scheduling server; when the time difference between the time of the last sending of the parameter matrix to the global scheduling server and the current time is the same as the global synchronization period, sending the local first low-rank parameter difference matrix and the local second low-rank parameter difference matrix to the global scheduling server.

[0046] Exemplarily, the reciprocal of the network bandwidth plus the network latency is used as the global synchronization period.

[0047] Exemplarily, a lower limit value and an upper limit value of the synchronization period can be adopted to limit the global synchronization period, so as to prevent the model synchronization from being too frequent or overly delayed.

[0048] Exemplarily, when the network bandwidth is large and the latency is small, the sum of the reciprocal of the above network bandwidth and the network latency is small, and the global synchronization period is closer to the lower limit value of the synchronization period. On the contrary, when the network bandwidth is small and the latency is large, that is, when the network is poor, the global synchronization period becomes longer, but does not exceed the upper limit value of the synchronization period. If it exceeds, the upper limit value of the synchronization period is taken as the global synchronization period.

[0049] It can be understood that the global synchronization period determines how often the global model parameter aggregation is performed. If the time from the current time to the last global model parameter aggregation is the global synchronization period, each node transmits its local first low-rank parameter difference matrix and local second low-rank parameter difference matrix to the global scheduling server. In this way, the global scheduling server can perform global aggregation on the local first low-rank parameter difference matrix and local second low-rank parameter difference matrix of each node respectively.

[0050] In one implementation, based on the current network bandwidth and current network latency between each distributed power node and the global scheduling server, determining the global synchronization period includes: based on the network bandwidth hyperparameter and network latency hyperparameter, respectively performing weighted summation on the reciprocal of the current network bandwidth and the current network latency to obtain the predicted synchronization period; taking the maximum value between the lower limit value of the synchronization period and the predicted synchronization period, and taking the minimum value between this maximum value and the upper limit value of the synchronization period to obtain the global synchronization period.

[0051] Exemplarily, the global synchronization period can be: ; wherein, represents the global synchronization period, represents the upper limit value of the synchronization period, represents the lower limit value of the synchronization period, represents the network bandwidth hyperparameter, represents the current network bandwidth, represents the network latency hyperparameter, represents the current network latency.

[0052] According to the above implementation, determining the global synchronization period according to the network bandwidth and network latency can avoid the model synchronization of each distributed power node from being too frequent or overly delayed, and can also avoid the incremental data collected by each node from being too insufficient due to excessive frequency, resulting in insufficient accuracy of model update in one cycle, or avoid overly delaying and causing the incremental data collected by each node to accumulate, resulting in too slow speed of model update in the next cycle.

[0053] In one implementation, the process of the global scheduling server updating the global first low-rank parameter matrix and the global second low-rank parameter matrix includes: the global scheduling server respectively determining the aggregation weights of each distributed power node based on the load, temperature, and power consumption of each distributed power node; the global scheduling server respectively performing weighted aggregation on the local first low-rank parameter difference matrix and the local second low-rank parameter difference matrix provided by each distributed power node based on the aggregation weights of each distributed power node, and using the weighted aggregation results to respectively update the global first low-rank parameter matrix and the global second low-rank parameter matrix of the global scheduling server.

[0054] Exemplarily, the load, temperature, and power consumption of each distributed power node are respectively normalized to obtain the load score, temperature score, and power consumption score of each distributed power node.

[0055] Exemplarily, the load score, temperature score, power consumption score, and error of the distributed power node are weighted and summed, and the reciprocal of the weighted sum result is used as the aggregation weight of the distributed power node.

[0056] Exemplarily, the aggregation weight of the th distributed power node is: where represents the aggregation weight of the , and respectively represent the load score, temperature score, and power consumption score of the th distributed power node, , and respectively represent the load weight, temperature weight, and power consumption weight, represents the error.

[0057] Exemplarily, using the weighted aggregation results to respectively update the global first low-rank parameter matrix and the global second low-rank parameter matrix of the global scheduling server is specifically as follows: ; ; where and represent the global first low-rank parameter matrix and the global second low-rank parameter matrix, represents the total number of distributed power nodes, and represent the The first local first low-rank parameter difference matrix and the local second low-rank parameter difference matrix of a distributed power node.

[0058] Figure 3 It is a structural block diagram of a distributed power model updating device based on model incremental training according to an embodiment of the present invention.

[0059] As Figure 3 shown, a distributed power model updating device based on model incremental training includes: A singular value decomposition module 310, configured to perform singular value decomposition on the local model parameter matrix of the local power model in the distributed power node when the data volume of the local power increment data in the distributed power node is less than a preset threshold, to obtain a left singular matrix, a diagonal matrix, and a right singular matrix; An important parameter determination module 320, configured to extract corresponding column vectors from the left singular matrix and the right singular matrix respectively based on the main singular values in the diagonal matrix to obtain an important parameter matrix; A low-rank parameter determination module 330, configured to determine an unimportant parameter matrix based on the difference matrix between the local model parameter matrix and the important parameter matrix, and decompose the unimportant parameter matrix to obtain a local first low-rank parameter matrix and a local second low-rank parameter matrix; A global model update module 340, configured to update the local first low-rank parameter matrix and the local second low-rank parameter matrix in the local model parameter matrix based on the local power increment data of the distributed power node, and send the update result to a global scheduling server, so that the global scheduling server updates the global first low-rank parameter matrix and the global second low-rank parameter matrix; A local model update module 350, configured to update the local power model of the distributed power node based on the product between the updated global first low-rank parameter matrix and the updated global second low-rank parameter matrix from the global scheduling server, and the important parameter matrix.

[0060] In an implementation manner, the important parameter determination module includes: A first singular value determination unit, configured to sort the diagonal singular values in the diagonal matrix from large to small, and intercept the first N diagonal singular values as the main singular values, where N is a positive integer greater than 1; A first matrix determination unit, configured to determine the important parameter matrix based on the singular vectors of each main singular value in the left singular matrix and the singular vectors in the right singular matrix.

[0061] In an implementation manner, the important parameter determination module matrix includes: A position determination unit, configured to determine a first position corresponding to the first power model parameter in the diagonal matrix according to the first power model parameter in the local power model whose gradient sensitivity is greater than a preset threshold; A second singular value determination unit, configured to determine the main singular value based on the singular value at the first position; A second matrix determination unit, configured to determine the important parameter matrix based on the singular vectors in the left singular matrix and the singular vectors in the right singular matrix for each of the main singular values.

[0062] In one implementation manner, the second singular value determination unit is specifically configured to: Sort the diagonal singular values in the diagonal matrix from largest to smallest, and determine the first N diagonal singular values, where N is a positive integer greater than 1; Obtain the main singular value based on the first N diagonal singular values and the singular value at the first position.

[0063] In one implementation manner, the global model update module includes: [[ID=ID=16]] A local training unit, configured to divide the local power increment data into multiple batches of data, and perform the following local training batch by batch: train the local power model based on the batch data to obtain a local loss; and update the local first low-rank parameter matrix and the local second low-rank parameter matrix based on the gradient information of the local loss and the efficiency factor adjusted according to the hardware characteristics corresponding to the computing power platform, to obtain the updated local first low-rank parameter matrix and the updated local second low-rank parameter matrix; A local model update unit, configured to, when the number of local training times does not reach the preset requirement, update the local power model based on the product between the updated local first low-rank parameter matrix and the updated local second low-rank parameter matrix and the important parameter matrix, for the next batch of local training; A difference calculation unit, configured to, when the number of local training times reaches the preset requirement, determine the local first low-rank parameter difference matrix based on the difference between the updated local first low-rank parameter matrix and the local first low-rank parameter matrix before update, and determine the local second low-rank parameter difference matrix based on the difference between the updated local second low-rank parameter matrix and the local second low-rank parameter matrix before update.

[0064] In one implementation manner, the global model update module includes: A synchronization period determination unit, configured to determine a global synchronization period based on the network bandwidth and network latency between each of the distributed power nodes and the global scheduling server; A differential result sending unit, configured to send the local first low-rank parameter difference matrix and the local second low-rank parameter difference matrix to the global scheduling server when the time difference between the time of the last sending of the parameter matrix to the global scheduling server and the current time is the same as the global synchronization period.

[0065] In one implementation manner, the synchronization period determination unit is specifically configured to: Based on network bandwidth hyperparameters and network latency hyperparameters, respectively perform weighted summation on the reciprocal of the current network bandwidth and the current network latency to obtain an estimated synchronization period; Take the maximum value between the lower limit value of the synchronization period and the estimated synchronization period, and take the minimum value between this maximum value and the upper limit value of the synchronization period to obtain the global synchronization period.

[0066] In one implementation manner, the process of the global scheduling server updating the global first low-rank parameter matrix and the global second low-rank parameter matrix includes: The global scheduling server respectively determines the aggregation weights of each of the distributed power nodes based on the load, temperature, and power consumption of each of the distributed power nodes; The global scheduling server respectively performs weighted aggregation on the local first low-rank parameter difference matrix provided by each of the distributed power nodes and the local second low-rank parameter difference matrix based on the aggregation weights of each of the distributed power nodes, and uses the weighted aggregation results to respectively update the global first low-rank parameter matrix and the global second low-rank parameter matrix of the global scheduling server.

[0067] For the specific functions and examples of each module and sub-module of the system according to the embodiments of the present invention, reference may be made to the relevant descriptions of the corresponding steps in the above method embodiments, which will not be elaborated here.

[0068] In the technical solution of the present invention, the acquisition, storage, and application of user personal information involved all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.

[0069] According to the embodiments of the present invention, the present invention also provides a system and a readable storage medium.

[0070] Figure 4FIG. 0 shows a schematic block diagram of an exemplary electronic device 800 that can be used to implement an embodiment of the present invention. The electronic device is intended to represent various forms of digital computers, such as, for example, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, for example, personal digital assistants, cellular telephones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely exemplary and are not intended to limit the implementation of the present invention described and / or claimed herein.

[0071] As Figure 4 shown, the electronic device 800 includes a computing unit 801 that can perform various appropriate actions and processes in accordance with a computer program stored in a read-only memory (ROM) 802 or a computer program loaded from a storage unit 808 into a random access memory (RAM) 803. In the RAM 803, various programs and data required for the operation of the electronic device 800 can also be stored. The computing unit 801, the ROM 802, and the RAM 803 are connected to each other via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.

[0072] A plurality of components in the electronic device 800 are connected to the I / O interface 805, including: an input unit 806, such as a keyboard, a mouse, etc.; an output unit 807, such as various types of displays, speakers, etc.; a storage unit 808, such as a magnetic disk, an optical disk, etc.; and a communication unit 809, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 809 allows the electronic device 800 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0073] The computing unit 801 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 801 executes the various methods and processes described above, such as the distributed power model update method based on model incremental training. For example, in some embodiments, the distributed power model update method based on model incremental training can be implemented as a computer software program that is tangibly contained in a machine-readable medium, such as the storage unit 808. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 800 via the ROM 802 and / or the communication unit 809. When the computer program is loaded into the RAM 803 and executed by the computing unit 801, one or more steps of the distributed power model update method based on model incremental training described above can be executed. Alternatively, in other embodiments, the computing unit 801 can be configured to execute the distributed power model update method based on model incremental training in any other suitable manner (e.g., by means of firmware).

[0074] Various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor, and can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.

[0075] The program code for implementing the methods of the present invention can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The program codes can be executed entirely on the machine, partially on the machine, as an independent software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0076] In the context of the present invention, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. The machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0077] To provide for interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide for interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic, speech, or tactile input).

[0078] The systems and techniques described herein can be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.

[0079] A computer system may include a client and a server. The client and the server are generally far from each other and usually interact via a communication network. The relationship between the client and the server is generated by computer programs running on the respective computers and having a client-server relationship with each other. The server may be a cloud server, a server of a distributed system, or a server incorporating a blockchain.

[0080] It should be understood that various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in the present invention can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in the present invention can be achieved, and this is not limited herein.

[0081] The above specific embodiments do not constitute a limitation on the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the principles of the present invention shall be included within the protection scope of the present invention.

Claims

1. A distributed power model update method based on model incremental training, characterized in that Including: When the data volume of the local power increment data of the distributed power node is less than a preset threshold, performing singular value decomposition on the local model parameter matrix of the local power model in the distributed power node to obtain a left singular matrix, a diagonal matrix, and a right singular matrix; Based on the main singular values in the diagonal matrix, extracting corresponding column vectors from the left singular matrix and the right singular matrix respectively to obtain an important parameter matrix; Based on the difference matrix between the local model parameter matrix and the important parameter matrix, determining an unimportant parameter matrix, and performing low-rank decomposition on the unimportant parameter matrix to obtain a local first low-rank parameter matrix and a local second low-rank parameter matrix; Based on the local power increment data, updating the local first low-rank parameter matrix and the local second low-rank parameter matrix, and sending the update result to the global scheduling server so that the global scheduling server updates the global first low-rank parameter matrix and the global second low-rank parameter matrix; Based on the product between the updated global first low-rank parameter matrix and the updated global second low-rank parameter matrix from the global scheduling server, and the important parameter matrix, updating the local power model of the distributed power node.

2. The method according to claim 1, characterized in that The step of extracting corresponding column vectors from the left singular matrix and the right singular matrix respectively based on the main singular values in the diagonal matrix to obtain an important parameter matrix includes: Determining a first position of the first power model parameter corresponding in the diagonal matrix according to the first power model parameter in the local power model whose gradient sensitivity is greater than a preset threshold; Based on the singular value at the first position, determining the main singular value; Based on the singular vectors of each of the main singular values in the left singular matrix and the singular vectors in the right singular matrix, determining the important parameter matrix.

3. The method according to claim 2, wherein The step of determining the main singular value based on the singular value at the first position includes: Sorting the diagonal singular values in the diagonal matrix from large to small, and determining the first N diagonal singular values, where N is a positive integer greater than 1; Based on the first N diagonal singular values and the singular value at the first position, obtaining the main singular value.

4. The method according to claim 1, wherein The step of extracting corresponding column vectors from the left singular matrix and the right singular matrix respectively based on the main singular values in the diagonal matrix to obtain an important parameter matrix includes: Sorting the diagonal singular values in the diagonal matrix from large to small, and intercepting the first N diagonal singular values as the main singular values, where N is a positive integer greater than 1; Based on the singular vectors of each of the main singular values in the left singular matrix and the singular vectors in the right singular matrix, determining the important parameter matrix.

5. The method according to claim 1, wherein The distributed power node runs in a computing power platform with an acceleration card. The step of updating the local first low-rank parameter matrix and the local second low-rank parameter matrix based on the local power increment data includes: Divide the local power increment data into multiple batches of data, and perform the following local training batch by batch: Based on the batch data, train the local power model to obtain a local loss; and based on the gradient information of the local loss and the efficiency factor adjusted according to the hardware characteristics corresponding to the computing power platform, update the local first low-rank parameter matrix and the local second low-rank parameter matrix to obtain the updated local first low-rank parameter matrix and the updated local second low-rank parameter matrix; In the case where the number of local training times does not meet the preset requirements, update the local power model based on the product between the updated local first low-rank parameter matrix and the updated local second low-rank parameter matrix, and the important parameter matrix, for the next batch of local training; In the case where the number of local training times meets the preset requirements, determine a local first low-rank parameter difference matrix based on the difference between the updated local first low-rank parameter matrix and the local first low-rank parameter matrix before update, and determine a local second low-rank parameter difference matrix based on the difference between the updated local second low-rank parameter matrix and the local second low-rank parameter matrix before update.

6. The method according to claim 5, wherein The sending the update result to the global scheduling server includes: Determine a global synchronization period based on the network bandwidth and network latency between each distributed power node and the global scheduling server; In the case where the time difference between the time of the last sending of the parameter matrix to the global scheduling server and the current time is the same as the global synchronization period, send the local first low-rank parameter difference matrix and the local second low-rank parameter difference matrix to the global scheduling server.

7. The method according to claim 6, characterized in that, The determining a global synchronization period based on the current network bandwidth and current network latency between each distributed power node and the global scheduling server includes: Based on network bandwidth hyperparameters and network latency hyperparameters, perform weighted summation on the reciprocal of the current network bandwidth and the current network latency respectively to obtain an estimated synchronization period; Take the maximum value between the lower limit value of the synchronization period and the estimated synchronization period, and take the minimum value between this maximum value and the upper limit value of the synchronization period to obtain the global synchronization period; The process of the global scheduling server updating the global first low-rank parameter matrix and the global second low-rank parameter matrix includes: The global scheduling server determines the aggregation weight of each distributed power node based on the load, temperature and power consumption of each distributed power node; The global scheduling server performs weighted aggregation on the local first low-rank parameter difference matrix and the local second low-rank parameter difference matrix provided by each distributed power node based on the aggregation weight of each distributed power node, and uses the weighted aggregation result to update the global first low-rank parameter matrix and the global second low-rank parameter matrix of the global scheduling server respectively.

8. A distributed power model update device based on incremental training of the model, characterized in that, Includes: A singular value decomposition module, configured to perform singular value decomposition on the local model parameter matrix of the local power model in the distributed power node when the data volume of the local power increment data of the distributed power node is less than a preset threshold, to obtain a left singular matrix, a diagonal matrix, and a right singular matrix; An important parameter determination module, configured to extract corresponding column vectors from the left singular matrix and the right singular matrix respectively based on the main singular values in the diagonal matrix, to obtain an important parameter matrix; A low-rank parameter determination module, configured to determine an unimportant parameter matrix based on the difference matrix between the local model parameter matrix and the important parameter matrix, and decompose the unimportant parameter matrix to obtain a local first low-rank parameter matrix and a local second low-rank parameter matrix; A global model update module, configured to update the local first low-rank parameter matrix and the local second low-rank parameter matrix in the local model parameter matrix based on the local power increment data of the distributed power node, and send the update result to a global scheduling server, so that the global scheduling server updates the global first low-rank parameter matrix and the global second low-rank parameter matrix; A local model update module, configured to update the local power model of the distributed power node based on the product between the updated global first low-rank parameter matrix and the updated global second low-rank parameter matrix from the global scheduling server, and the important parameter matrix.

9. An electronic device, characterized in that, Comprising: At least one processor, and a memory communicatively connected to the at least one processor; Wherein, the memory stores instructions executable by the processor, and the processor is configured to obtain the instructions from the memory and execute the instructions, so that the processor can execute the distributed power model update method based on model incremental training according to any one of claims 1-7.

10. A non-transitory computer-readable storage medium storing computer instructions, characterized in that, The computer instructions are used to be provided to a computer to instruct the computer to execute the distributed power model update method based on model incremental training according to any one of claims 1-7.

Citation Information

Patent Citations

  • Construction method and device of non-intrusive load identification model and storage medium

    CN113158134A

  • Federal learning method

    CN115099424A

  • Distributed large language model quantitative deployment method and device, equipment and storage medium

    CN118735001A

  • Large language model training method and system based on elastic federal low-rank adaptation fine tuning

    CN119443311A

  • Fine adjustment method and device for large language model

    CN119990183A