Deployment method and system of large model in power grid field

By employing a parallel decomposition strategy that coordinates computation and storage and automatically allocates hardware resources, the problems of high training time and insufficient scalability of large models have been solved, enabling efficient training and flexible deployment of large models in the power grid field, reducing costs and shortening the R&D cycle.

CN120950089APending Publication Date: 2025-11-14SHANDONG LUNENG SOFTWARE TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510861809.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2025-04-08
Filing Date
2025-06-25
Publication Date
2025-11-14

AI Technical Summary

Technical Problem

Large-scale model training is time-consuming, computationally expensive, and lacks elastic scalability, resulting in slow development of large-scale models in the power grid field.

Method used

By adopting a parallel decomposition strategy that coordinates computing and storage resources, and through the parallel decomposition strategy and automatic hardware resource allocation, distributed training and automatic conversion of large models are achieved, thereby improving training efficiency and reducing costs.

Benefits of technology

It enables efficient training and flexible scaling of large models, reduces training costs, shortens the R&D cycle, and improves the applicability and flexibility of large models in the power grid field.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120950089A_ABST
    Figure CN120950089A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of large model deployment in the power grid field, and provides a large model deployment method and system in the power grid field. The deployment method of the large model in the power grid field comprises the following steps of: based on single parameter memory overhead, single gradient memory overhead, single optimizer state memory overhead, model parameter quantity and data parallelism; according to the relationship between parameters such as the number of working nodes in model parallelism and pipeline parallelism and the memory overhead in the parallel decomposition strategy of calculation and storage collaboration, the parallel decomposition strategy of calculation and storage collaboration with the minimum memory overhead is calculated, and the parallel decomposition strategy of the large model in the power grid field is determined; on the basis of a parallel decomposition strategy of a large model in the power grid field, hardware resource allocation is automatically performed by using hyper-parameter input hardware equipment information and a deep learning framework, then deep learning model structure setting and parallel decomposition scheme configuration are performed by using guidance statements, automatic code translation is completed, and automatic conversion and training of a deep learning model are realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of large-scale power grid deployment technology, and particularly relates to a method and system for deploying large-scale power grid models. Background Technology

[0002] The statements in this section are merely background information related to the present invention and do not necessarily constitute prior art.

[0003] Compared to traditional models, large-scale models, due to their powerful generalization and data analysis capabilities, can not only handle tasks such as target detection, equipment management, and cost control, but also provide decision-making suggestions and constructive solutions for power grid companies using their graphic and text generation capabilities. However, the training and deployment of large-scale models still face the following challenges: (1) Training large models is time-consuming and computationally expensive. How to make full use of existing computing resources to complete the retraining process of transforming general domain large models into power grid domain large models is an urgent problem to be solved.

[0004] (2) The large model has insufficient elastic expansion capability and slow update iteration. How to complete the initial training of the large model based on 2-3 application scenarios, and then carry out functional elastic expansion and agile deployment application in other application scenarios is the key issue that leads to the slow development of large models in the power grid field. Summary of the Invention

[0005] To address the aforementioned technical problems, this invention provides a method and system for deploying large-scale models in the power grid field, which can construct an optimal distributed training strategy for large-scale models, improve training efficiency, and reduce training costs.

[0006] To achieve the above objectives, the present invention adopts the following technical solution: The first aspect of the present invention provides a method for deploying a large-scale model in the field of power grids.

[0007] In one or more embodiments, a method for deploying a large-scale model in the power grid domain is provided, comprising: Based on the relationship between the memory overhead of single parameter, single gradient, single optimizer state, model parameter count, number of working nodes in data parallelism, model parallelism and pipeline parallelism, and the memory overhead in the computation-storage coordinated parallel decomposition strategy, the computation-storage coordinated parallel decomposition strategy with the minimum memory overhead is calculated, and the parallel decomposition strategy for large power grid models is determined. Based on the parallel decomposition strategy of large power grid models, this paper utilizes hyperparameter input hardware device information and deep learning framework to automatically allocate hardware resources. Then, it uses guided statements to set the deep learning model structure and configure the parallel decomposition scheme, completes automatic code translation, and realizes automatic conversion and training of deep learning models.

[0008] As one implementation method, the memory overhead of a single parameter Memory overhead of a single gradient Memory overhead of a single optimizer state Model parameter count and the number of worker nodes in data parallelism These parameters, along with the memory overhead of the parallel decomposition strategy coordinated with storage, The relationship between them is: .

[0009] As one implementation method, the number of model parameters for: ;in, To calculate the precision of floating-point numbers; The total number of model parameters for a deep learning framework.

[0010] As one implementation method, the parallel decomposition strategy for the large power grid model includes inter-group data parallelism, intra-group model parallelism, pipeline parallelism, and optimizer parallelism.

[0011] As one implementation method, the intra-group model parallelism strategy is as follows: ;in, Memory overhead for storage-coordinated parallel decomposition strategies, Due to memory limitations of a single device, Allocate equipment quantities within the group.

[0012] As one implementation method, the inter-group data parallelism strategy is as follows: ;in, Allocate equipment quantities within the group. Allocate equipment quantities between groups. This represents the total number of available devices.

[0013] A second aspect of the present invention provides a deployment system for a large-scale model in the field of power grids.

[0014] In one or more embodiments, a deployment system for a large-scale power grid model includes: The parallel decomposition strategy determination module is used to calculate the parallel decomposition strategy with the minimum memory overhead based on the relationship between parameters such as memory overhead of a single parameter, memory overhead of a single gradient, memory overhead of a single optimizer state, number of model parameters and number of working nodes in data parallelism, model parallelism and pipeline parallelism, and memory overhead in a computation-storage coordinated parallel decomposition strategy, and to determine the parallel decomposition strategy for large models in the power grid field. The automatic model conversion training module is used for parallel decomposition strategies based on large models in the power grid field. It uses hyperparameter input hardware device information and deep learning framework to automatically allocate hardware resources, and then uses guided statements to set the deep learning model structure and configure the parallel decomposition scheme, completes automatic code translation, and realizes automatic conversion and training of deep learning models.

[0015] As one implementation, in the parallel decomposition strategy determination module, the memory overhead of a single parameter is... Memory overhead of a single gradient Memory overhead of a single optimizer state Model parameter count and the number of worker nodes in data parallelism These parameters, along with the memory overhead of the parallel decomposition strategy coordinated with storage, The relationship between them is: .

[0016] A third aspect of the present invention provides a computer-readable storage medium.

[0017] A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps in the deployment method of the large-scale power grid model as described above.

[0018] A fourth aspect of the present invention provides an electronic device.

[0019] An electronic device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the steps in the deployment method of the large-scale power grid model as described above.

[0020] Compared with the prior art, the beneficial effects of the present invention are: (1) This invention proposes a parallel decomposition strategy that coordinates computing and storage resources. Based on the limitations of available computing power and storage capacity, a series of suitable parallel decomposition strategies are selected, and finally the optimal distributed training strategy for large models is constructed to improve training efficiency and reduce training costs.

[0021] (2) This invention proposes an automatic generation technology for training large models in the power grid field, which automatically transforms the traditional artificial intelligence models that have been applied into large models that can be trained in a distributed manner, and completes the automatic configuration of hardware resources, thereby realizing the elastic expansion and agile deployment of large model functions. Attached Figure Description

[0022] The accompanying drawings, which form part of this invention, are used to provide a further understanding of the invention. The illustrative embodiments of the invention and their descriptions are used to explain the invention and do not constitute an improper limitation of the invention.

[0023] Figure 1 This is a schematic diagram of the parallel decomposition strategy for computation and storage collaboration in an embodiment of the present invention. Figure 2 This is a flowchart illustrating the large-scale model deployment method in the power grid field according to an embodiment of the present invention; Figure 3 This is a schematic diagram illustrating the accelerated distributed training of the model according to an embodiment of the present invention; Figure 4 This is a schematic diagram of an electronic device according to an embodiment of the present invention. Detailed Implementation

[0024] The present invention will be further described below with reference to the accompanying drawings and embodiments.

[0025] It should be noted that the following detailed description is illustrative and intended to provide further explanation of the invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.

[0026] It should be noted that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the scope of exemplary embodiments according to the invention. As used herein, the singular form is intended to include the plural form as well, unless the context clearly indicates otherwise. Furthermore, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.

[0027] Figure 1 This is a flowchart illustrating a method for deploying a large-scale power grid model according to an embodiment of the present invention, as shown below. Figure 1 The deployment method of the large-scale power grid model in this embodiment may include: S101, based on the relationship between the memory overhead of a single parameter, a single gradient, a single optimizer state, the number of model parameters and the number of working nodes in data parallelism, model parallelism and pipeline parallelism, and the memory overhead in the computation-storage collaborative parallel decomposition strategy, calculates the computation-storage collaborative parallel decomposition strategy with the minimum memory overhead, and determines the parallel decomposition strategy for large models in the power grid field. S102 is based on a parallel decomposition strategy for large-scale power grid models. It utilizes hyperparameter input hardware device information and a deep learning framework to automatically allocate hardware resources. Then, it uses guided statements to set the deep learning model structure and configure the parallel decomposition scheme, completes automatic code translation, and realizes automatic conversion and training of deep learning models.

[0028] In step S101, during the current large-scale model training process, the memory overhead constraint is far greater than the computational overhead constraint due to the significant increase in model parameters, gradients, optimizer states, and other data. Therefore, a computation-storage coordinated parallel decomposition strategy is proposed. From the perspective of memory constraints, while improving model training efficiency, based on model parameter, gradient, and optimizer state sharding techniques, this embodiment of the large-scale power grid model adopts a storage-coordinated parallel decomposition strategy, such as... Figure 2 As shown.

[0029] Memory overhead of a single parameter Memory overhead of a single gradient Memory overhead of a single optimizer state Model parameter count and the number of worker nodes in data parallelism These parameters, along with the memory overhead of the parallel decomposition strategy coordinated with storage, The relationship between them is: .

[0030] This embodiment is based on a model sharding strategy, and combines the computation time and communication time of the model's forward and backward propagation with the model structure to design an optimal pipeline to improve model training efficiency.

[0031] Among them, the number of model parameters for: ;in, To calculate the precision of floating-point numbers.

[0032] The total number of model parameters for a deep learning framework.

[0033] In this embodiment, the parallel decomposition strategy for the large power grid model includes inter-group data parallelism, intra-group model parallelism, pipeline parallelism, and optimizer parallelism.

[0034] The intra-group model parallelism strategy is as follows: ;in, Memory overhead for storage-coordinated parallel decomposition strategies, Due to memory limitations of a single device, Allocate equipment quantities within the group.

[0035] The inter-group data parallelism strategy is as follows: ;in, Allocate equipment quantities within the group. Allocate equipment quantities between groups. This represents the total number of available devices.

[0036] To address the challenges of flexible expansion and agile iterative updates of large-scale power grid models, this study investigates an automatic generation technology for training large-scale power grid models. This technology can automatically transform AI models trained and run in traditional single-machine, single-card environments into large-scale models in multi-machine, multi-card environments, and can automate hardware resource configuration and training. Currently, it supports deep learning frameworks such as PyTorch and Tensorflow, and is implemented through hyperparameter input and guided statements.

[0037] In step S102, the hyperparameter input is mainly used for framework selection and hardware resource configuration, and supports inputting framework name, computing resource information and network transmission information through hyperparameters.

[0038] Directed statements are simple statements inserted into traditional artificial intelligence model code to programmatically implement parallel decomposition strategies that coordinate computation and storage. Directed statements identify the parallel decomposition strategy used by the model, the specific number of devices for each strategy, and the model structure, thus automatically translating small model code in a single-machine, single-GPU environment into large model code in a multi-machine, multi-GPU environment.

[0039] Finally, by utilizing hyperparameter inputs and code translation results, the automatic configuration and execution of the hardware environment are achieved.

[0040] Figure 3 This is a schematic diagram of the distributed training acceleration of the model in an embodiment of the present invention. This embodiment adopts a parallel decomposition strategy that coordinates computing and storage, which can make full use of computing and storage resources. When using multiple (e.g., 4) GPUs to complete the distributed training of a large model, the memory overhead can be reduced by 50% compared with the previous strategy, and the computing time can achieve a speedup of more than 5.4 times compared with a single GPU. Therefore, the overall cost can be reduced by about 52%.

[0041] This embodiment employs an automated generation technology for training large-scale power grid models, enabling agile iterative updates and flexible functional expansion of large models, thus shortening the development cycle. Assuming that experienced large-scale model development engineers typically go through five stages—dataset construction, domain model design, computing resource adaptation, general large-scale model fine-tuning, and parameter optimization—in developing domain-specific large-scale models, taking a total of three months, the automated generation technology proposed in this project can significantly reduce the time spent on domain model design, computing resource adaptation, and general large-scale model fine-tuning, shortening the development cycle by more than 50%.

[0042] In one or more embodiments, a deployment system for a large-scale power grid model is also provided, comprising: The parallel decomposition strategy determination module is used to calculate the parallel decomposition strategy with the minimum memory overhead based on the relationship between parameters such as memory overhead of a single parameter, memory overhead of a single gradient, memory overhead of a single optimizer state, number of model parameters and number of working nodes in data parallelism, model parallelism and pipeline parallelism, and memory overhead in a computation-storage coordinated parallel decomposition strategy, and to determine the parallel decomposition strategy for large models in the power grid field. The automatic model conversion training module is used for parallel decomposition strategies based on large models in the power grid field. It uses hyperparameter input hardware device information and deep learning framework to automatically allocate hardware resources, and then uses guided statements to set the deep learning model structure and configure the parallel decomposition scheme, completes automatic code translation, and realizes automatic conversion and training of deep learning models.

[0043] Specifically, in the parallel decomposition strategy determination module, the memory overhead of a single parameter is... Memory overhead of a single gradient Memory overhead of a single optimizer state Model parameter count and the number of worker nodes in data parallelism These parameters, along with the memory overhead of the parallel decomposition strategy coordinated with storage, The relationship between them is: .

[0044] It should be noted that each module in the deployment system of the large-scale power grid model corresponds one-to-one with each step in the deployment method of the large-scale power grid model, and their specific implementation processes are the same, so they will not be repeated here.

[0045] Reference Figure 4 A schematic diagram of an electronic device is provided. It should be noted that... Figure 4 The electronic device 400 shown is merely an example and should not be construed as limiting the functionality and scope of use of the embodiments of the present invention.

[0046] like Figure 4 As shown, the electronic device 400 includes a central processing unit (CPU) 401, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 402 or a program loaded from a storage section 408 into a random access memory (RAM) 403. The RAM 403 also stores various programs and data required for system operation. The CPU 401, ROM 402, and RAM 403 are interconnected via a bus 404. An input / output (I / O) interface 405 is also connected to the bus 404.

[0047] The following components are connected to I / O interface 405: an input section 406 including a keyboard, mouse, etc.; an output section 407 including a cathode ray tube (CRT), liquid crystal display (LCD), etc., and speakers, etc.; a storage section 408 including a hard disk, etc.; and a communication section 409 including a network interface card such as a local area network (LAN) card, modem, etc. The communication section 409 performs communication processing via a network such as the Internet. A drive 410 is also connected to I / O interface 405 as needed. A removable medium 411, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed on drive 410 as needed so that computer programs read from it can be installed into storage section 408 as needed.

[0048] When the central processing unit 401 in the electronic device of this embodiment executes the program, it achieves the following: Figure 1 The steps in the deployment method of the large-scale power grid model are shown.

[0049] Specifically, according to embodiments of this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program including functions for executing... Figure 1 The program code for the method shown. In such an embodiment, the computer program can be downloaded and installed from a network via communication section 409, and / or installed from removable medium 411. When the computer program is executed by central processing unit 401, it performs the various functions defined in the apparatus of this application.

[0050] in, Figure 1 The computer program instructions corresponding to the method shown may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to operate in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in the process. Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0051] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. The storage medium can be a magnetic disk, optical disk, read-only memory (ROM), or random access memory (RAM), etc.

[0052] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A method for deploying a large-scale model in the power grid field, characterized in that, include: Based on the relationship between the memory overhead of single parameter, single gradient, single optimizer state, model parameter count, number of working nodes in data parallelism, model parallelism and pipeline parallelism, and the memory overhead in the computation-storage coordinated parallel decomposition strategy, the computation-storage coordinated parallel decomposition strategy with the minimum memory overhead is calculated, and the parallel decomposition strategy for large power grid models is determined. Based on the parallel decomposition strategy of large power grid models, this paper utilizes hyperparameter input hardware device information and deep learning framework to automatically allocate hardware resources. Then, it uses guided statements to set the deep learning model structure and configure the parallel decomposition scheme, completes automatic code translation, and realizes automatic conversion and training of deep learning models.

2. The deployment method for a large-scale power grid model as described in claim 1, characterized in that, Memory overhead of a single parameter Memory overhead of a single gradient Memory overhead of a single optimizer state Model parameter count and the number of worker nodes in data parallelism These parameters, along with the memory overhead of the parallel decomposition strategy coordinated with storage, The relationship between them is: .

3. The deployment method for a large-scale power grid model as described in claim 2, characterized in that, Model parameter count for: ;in, To calculate the precision of floating-point numbers; The total number of model parameters for a deep learning framework.

4. The deployment method for a large-scale power grid model as described in claim 1, characterized in that, The parallel decomposition strategies for large power grid models include inter-group data parallelism, intra-group model parallelism, pipeline parallelism, and optimizer parallelism.

5. The deployment method for a large-scale power grid model as described in claim 4, characterized in that, The intra-group model parallelism strategy is as follows: ;in, Memory overhead for storage-coordinated parallel decomposition strategies, Due to memory limitations of a single device, Allocate equipment quantities within the group.

6. The deployment method for a large-scale power grid model as described in claim 4, characterized in that, The inter-group data parallelism strategy is as follows: ;in, Allocate equipment quantities within the group. Allocate equipment quantities between groups. This represents the total number of available devices.

7. A deployment system for a large-scale model in the power grid field, characterized in that, include: The parallel decomposition strategy determination module is used to calculate the parallel decomposition strategy with the minimum memory overhead based on the relationship between parameters such as memory overhead of a single parameter, memory overhead of a single gradient, memory overhead of a single optimizer state, number of model parameters and number of working nodes in data parallelism, model parallelism and pipeline parallelism, and memory overhead in a computation-storage coordinated parallel decomposition strategy, and to determine the parallel decomposition strategy for large models in the power grid field. The automatic model conversion training module is used for parallel decomposition strategies based on large models in the power grid field. It uses hyperparameter input hardware device information and deep learning framework to automatically allocate hardware resources, and then uses guided statements to set the deep learning model structure and configure the parallel decomposition scheme, completes automatic code translation, and realizes automatic conversion and training of deep learning models.

8. The deployment system for a large-scale power grid model as described in claim 7, characterized in that, In the parallel decomposition strategy determination module, the memory overhead of a single parameter Memory overhead of a single gradient Memory overhead of a single optimizer state Model parameter count and the number of worker nodes in data parallelism These parameters, along with the memory overhead of the parallel decomposition strategy coordinated with storage, The relationship between them is: .

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the steps in the deployment method of the large-scale power grid model as described in any one of claims 1-6.

10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps in the deployment method of the large-scale power grid model as described in any one of claims 1-6.