Deployment method and system of large model in power grid field

Through the parallel decomposition strategy of computing storage collaboration and automated hardware resource allocation, the problem of high time-consuming and cost-effective large-scale model training in the power grid field is solved, and efficient and flexible model deployment and update are achieved.

CN120295659AInactive Publication Date: 2025-07-11SHANDONG LUNENG SOFTWARE TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510433987.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-08
Publication Date
2025-07-11
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The training of large models is time-consuming, high computing power costs, and insufficient elastic expansion capabilities, resulting in slow development of large models in the power grid field.

Method used

The parallel decomposition strategy of computing storage resources is adopted to automatically configure hardware resources through hyperparameter input and guidance statements, automatically convert and train large models in the power grid field, and support distributed training.

Benefits of technology

It improves training efficiency, reduces training costs, and realizes agile deployment of large models and elastic expansion of functions, shortens the R&D cycle.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120295659A_ABST
    Figure CN120295659A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of large model deployment in the power grid field, and provides a large model deployment method and system in the power grid field. The deployment method of the large model in the power grid field comprises the following steps of: based on single parameter memory overhead, single gradient memory overhead, single optimizer state memory overhead, model parameter quantity and data parallelism; according to the relationship between parameters such as the number of working nodes in model parallelism and pipeline parallelism and the memory overhead in the parallel decomposition strategy of calculation and storage collaboration, the parallel decomposition strategy of calculation and storage collaboration with the minimum memory overhead is calculated, and the parallel decomposition strategy of the large model in the power grid field is determined; on the basis of a parallel decomposition strategy of a large model in the power grid field, hardware resource allocation is automatically performed by using hyper-parameter input hardware equipment information and a deep learning framework, then deep learning model structure setting and parallel decomposition scheme configuration are performed by using guidance statements, automatic code translation is completed, and automatic conversion and training of a deep learning model are realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of large model deployment in the power grid field, and particularly relates to a method and system for deploying a large model in the power grid field. Background Art

[0002] The statements in this part merely provide background technical information related to the present invention, and do not necessarily constitute prior art.

[0003] Compared with traditional models, due to its powerful generalization ability and data analysis ability, in addition to completing tasks in scenarios such as object detection, equipment management, and cost control, the large model can also provide decision-making suggestions and constructive solutions for power grid enterprises by using its image and text generation ability. However, there are still the following problems in the current training and deployment of large models:

[0004] (1) The training of large models is time-consuming and the computing power cost is high. How to make full use of the existing computing power resources to complete the retraining process of converting the general domain large model into a large model in the power grid field is an urgent problem to be solved.

[0005] (2) The elastic expansion ability of large models is insufficient and the update and iteration are slow. How to complete the initial training of the large model based on 2-3 application scenarios and then perform the functional elastic expansion and agile deployment application of other application scenarios is the key problem that has led to the slow development of large models in the power grid field. Summary of the Invention

[0006] In order to solve the above technical problems, the present invention provides a method and system for deploying a large model in the power grid field, which can construct an optimal distributed training strategy for the large model, improve the training efficiency, and reduce the training cost.

[0007] In order to achieve the above object, the present invention adopts the following technical solutions:

[0008] The first aspect of the present invention provides a method for deploying a large model in the power grid field.

[0009] In one or more embodiments, a method for deploying a large model in the power grid field is provided, including:

[0010] Based on the relationship between the memory overhead of a single parameter, the memory overhead of a single gradient, the memory overhead of a single optimizer state, the number of model parameters, and the memory overhead in the parallel decomposition strategy of computing and storage coordination among the number of worker nodes in data parallelism, model parallelism, and pipeline parallelism, calculate the parallel decomposition strategy of computing and storage coordination with the smallest memory overhead, and determine the parallel decomposition strategy of the large model in the power grid field;

[0011] Based on the parallel decomposition strategy of large models in the power grid field, by inputting hardware device information and deep learning frameworks with hyperparameters, hardware resource allocation is automatically performed, and then guidance statements are used to set the deep learning model structure and configure the parallel decomposition scheme, completing automatic code translation and realizing the automatic conversion and training of deep learning models.

[0012] As an implementation, the memory overhead P for a single parameter, the memory overhead G for a single gradient, the memory overhead O for a single optimizer state, the number of model parameters ψ, and the number of worker nodes N in data parallelism d The relationship between these parameters and the memory overhead M' of the parallel decomposition strategy for storage collaboration is:

[0013] As an implementation, the number of model parameters ψ is: ψ = Np * Pre / 8; where Pre is the floating-point calculation precision.

[0014] As an implementation, the parallel decomposition strategy of the large model in the power grid field includes inter-group data parallelism, intra-group model parallelism, pipeline parallelism, and optimizer parallelism strategies.

[0015] As an implementation, the intra-group model parallelism strategy is: where M ′ is the memory overhead of the parallel decomposition strategy for storage collaboration, M0 is the memory limit of a single device, and N1 is the intra-group device allocation.

[0016] As an implementation, the inter-group data parallelism strategy is: where N1 is the intra-group device allocation, N2 is the inter-group device allocation, and N is the total number of available devices.

[0017] The second aspect of the present invention provides a deployment system for a large model in the power grid field.

[0018] In one or more embodiments, a deployment system for a large model in the power grid field includes:

[0019] A parallel decomposition strategy determination module, which is used to calculate the parallel decomposition strategy of the large model in the power grid field with the smallest memory overhead based on the relationship between the memory overhead of the parallel decomposition strategy for computing and storage collaboration and parameters such as the memory overhead for a single parameter, the memory overhead for a single gradient, the memory overhead for a single optimizer state, the number of model parameters, and the number of worker nodes in data parallelism, model parallelism, and pipeline parallelism;

[0020] The model automatic conversion training module is used to automatically allocate hardware resources based on the parallel decomposition strategy of the large model in the power grid field, using hyperparameters to input hardware device information and deep learning frameworks, and then using guidance statements to set the deep learning model structure and configure the parallel decomposition scheme, complete automatic code translation, and achieve automatic conversion and training of the deep learning model.

[0021] As an implementation method, in the parallel decomposition strategy determination module, the memory overhead P of a single parameter, the memory overhead G of a single gradient, the memory overhead O of a single optimizer state, the number of model parameters ψ, and the number of worker nodes N in data parallelism d The relationship between these parameters and the memory overhead M' of the parallel decomposition strategy for storage collaboration is:

[0022] The third aspect of the present invention provides a computer-readable storage medium.

[0023] A computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the steps in the deployment method of the large model in the power grid field as described above.

[0024] The fourth aspect of the present invention provides an electronic device.

[0025] An electronic device includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the steps in the deployment method of the large model in the power grid field as described above.

[0026] Compared with the prior art, the beneficial effects of the present invention are:

[0027] (1) The present invention proposes a parallel decomposition strategy for computing and storage resource collaboration. According to the computing power and storage capacity limitations of available computing resources, a series of appropriate parallel decomposition strategies are selected, and finally an optimal large model distributed training strategy is constructed to improve training efficiency and reduce training costs.

[0028] (2) The present invention proposes an automatic generation technology for training large models in the power grid field, automatically converts the currently applied traditional artificial intelligence models into large models that can be distributedly trained, and completes the automatic configuration of hardware resources, thereby realizing the elastic expansion and agile deployment of large model functions. Description of the Drawings

[0029] The specification drawings constituting a part of the present invention are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation to the present invention.

[0030] Figure 1Schematic diagram of the parallel decomposition strategy for computing-storage collaboration according to an embodiment of the present invention;

[0031] Figure 2 Schematic diagram of the process of the large model deployment method in the power grid field according to an embodiment of the present invention;

[0032] Figure 3 Schematic diagram of the acceleration situation of the distributed training of the model according to an embodiment of the present invention;

[0033] Figure 4 Schematic diagram of an electronic device according to an embodiment of the present invention. Detailed implementation manners

[0034] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.

[0035] It should be noted that the following detailed descriptions are all illustrative and are intended to provide further descriptions of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs.

[0036] It should be noted that the terms used herein are only for describing specific implementation manners and are not intended to limit the exemplary embodiments according to the present invention. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprise" and / or "include" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.

[0037] Figure 1 Schematic diagram of the process of a large model deployment method in the power grid field according to an embodiment of the present invention. As Figure 1 shown, the large model deployment method in the power grid field in this embodiment may include:

[0038] S101, based on the relationships between the memory overhead of a single parameter, the memory overhead of a single gradient, the memory overhead of a single optimizer state, the number of model parameters, and the number of worker nodes in data parallelism, model parallelism, and pipeline parallelism and the memory overhead in the parallel decomposition strategy for computing-storage collaboration, calculate the parallel decomposition strategy for computing-storage collaboration with the minimum memory overhead, and determine the parallel decomposition strategy for the large model in the power grid field;

[0039] S102, based on the parallel decomposition strategy of the large model in the power grid field, use the hyperparameters to input the hardware device information and the deep learning framework, automatically perform hardware resource allocation, and then use the guiding statements to set the deep learning model structure and configure the parallel decomposition scheme, complete the automatic code translation, and realize the automatic conversion and training of the deep learning model.

[0040] In step S101, during the current large model training process, due to the significant increase in data such as model parameters, gradients, and optimizer states, the memory overhead limitation is much greater than the computing overhead limitation. Therefore, a parallel decomposition strategy of computing-storage collaboration is proposed. From the perspective of memory limitation, on the premise of improving the model training efficiency, based on the model parameter, gradient, and optimizer state sharding technologies, the parallel decomposition strategy of the large model in the power grid field of this embodiment adopts a parallel decomposition strategy of storage collaboration, as Figure 2 shown.

[0041] The memory overhead P of a single parameter, the memory overhead G of a single gradient, the memory overhead O of a single optimizer state, the number of model parameters ψ, and the number of worker nodes N in data parallelism d The relationship between these parameters and the memory overhead M' of the parallel decomposition strategy of storage collaboration is:

[0042] Based on the model sharding strategy, this embodiment combines the forward and backward propagation calculation time, communication time, and model structure of the model to design an optimal pipeline for improving the model training efficiency.

[0043] Among them, the number of model parameters ψ is: ψ = Np * Pre / 8; where Pre is the floating-point calculation precision.

[0044] The number of model parameters is automatically obtained by the deep learning framework.

[0045] In this embodiment, the parallel decomposition strategy of the large model in the power grid field includes inter-group data parallelism, intra-group model parallelism, pipeline parallelism, and optimizer parallelism strategies.

[0046] Among them, the intra-group model parallelism strategy is: Among them, M ′ is the memory overhead of the parallel decomposition strategy of storage collaboration, M0 is the memory limit of a single device, and N1 is the intra-group device allocation.

[0047] The inter-group data parallelism strategy is: Among them, N1 is the intra-group device allocation, N2 is the inter-group device allocation, and N is the total number of available devices.

[0048] To solve the problems of difficult functional elastic expansion and agile iterative update of large models in the power grid field, an automatic generation technology for large model training in the power grid field has been studied, which can automatically convert an artificial intelligence model trained and run in a traditional single-machine single-card environment in the power grid field into a large model in a multi-machine multi-card environment, and can complete automated hardware resource configuration and training. Currently, it supports deep learning frameworks PyTorch and TensorFlow, and is implemented through hyperparameter input and guiding statements.

[0049] In step S102, the hyperparameter input is mainly used for framework selection and hardware resource configuration, and supports inputting the framework name, computing power resource information, and network transmission information through hyperparameters.

[0050] The guiding statements are to insert some simple guiding statements into the traditional artificial intelligence model code, which are used to implement the parallel decomposition strategy of computing-storage collaboration in the program. The guiding statements identify what kind of parallel decomposition strategy the model adopts, the specific number of devices for each strategy, and the model structure, so that the small model code in the single-machine single-GPU environment can be automatically translated into the large model code in the multi-machine multi-GPU environment.

[0051] Finally, the automatic configuration and execution of the hardware environment are realized by using the hyperparameter input and the code translation result.

[0052] Figure 3 It is a schematic diagram of the accelerated distributed training of the model in the embodiment of the present invention. In this embodiment, the parallel decomposition strategy of computing-storage collaboration is adopted, which can make full use of computing and storage resources. When using multiple (such as 4) GPUs to complete the distributed training of the large model, the memory overhead can be reduced by 50% compared with that before applying this strategy, and the computing time can achieve an acceleration ratio of more than 5.4 times compared with a single GPU. Therefore, the comprehensive cost calculation can be reduced by about 52%.

[0053] This embodiment adopts the automatic generation technology for large model training in the power grid field, which can realize the agile iterative update of the large model and the elastic expansion of functions, and shorten the R & D cycle of the large model. Suppose that an experienced large model R & D engineer needs to go through five stages, namely building a dataset, domain model design, computing power resource adaptation, general large model fine-tuning, and parameter tuning, to carry out the R & D work of the domain large model, which takes a total of 3 months. Using the automatic generation technology proposed in this project, the process of domain model design, computing power resource adaptation, and general large model fine-tuning can be greatly reduced, and the R & D cycle can be shortened by more than 50%.

[0054] In one or more embodiments, a deployment system for a large model in the power grid field is further provided, including:

[0055] A parallel decomposition strategy determination module, which is used to calculate the parallel decomposition strategy of computing-storage collaboration with the smallest memory overhead based on the relationship between these parameters such as the memory overhead of a single parameter, the memory overhead of a single gradient, the memory overhead of a single optimizer state, the number of model parameters, and the number of worker nodes in data parallelism, model parallelism, and pipeline parallelism, and the memory overhead in the parallel decomposition strategy of computing-storage collaboration, and determine the parallel decomposition strategy of the large model in the power grid field;

[0056] The model automatic conversion training module is used to automatically allocate hardware resources based on the parallel decomposition strategy of the large model in the power grid field, using hyperparameters to input hardware device information and deep learning frameworks, and then using guiding statements to set the deep learning model structure and configure the parallel decomposition scheme, complete automatic code translation, and achieve automatic conversion and training of the deep learning model.

[0057] Specifically, in the parallel decomposition strategy determination module, the single-parameter memory overhead P, the single-gradient memory overhead G, the single-optimizer state memory overhead O, the number of model parameters ψ, and the number of worker nodes N in data parallelism d The relationship between these parameters and the memory overhead M' of the parallel decomposition strategy for storage coordination is:

[0058] It should be noted here that each module in the deployment system of the large model in the power grid field corresponds one by one to each step in the deployment method of the large model in the power grid field, and their specific implementation processes are the same, so they will not be elaborated here.

[0059] Refer to Figure 4 , a schematic diagram of an electronic device is given. It should be noted that Figure 4 The illustrated electronic device 400 is only an example and should not impose any limitations on the functions and usage scope of the embodiments of the present invention.

[0060] As Figure 4 shown, the electronic device 400 includes a central processing unit (CPU) 401, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 402 or the program loaded from the storage section 408 into the random access memory (RAM) 403. In the RAM 403, various programs and data required for system operation are also stored. The central processing unit 401, the ROM 402, and the RAM 403 are connected to each other via a bus 404. The input / output (I / O) interface 405 is also connected to the bus 404.

[0061] The following components are connected to the I / O interface 405: an input section 406 including a keyboard, a mouse, etc.; an output section 407 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc. as well as a speaker; a storage section 408 including a hard disk, etc.; and a communication section 409 including a network interface card such as a local area network (LAN) card, a modem, etc. The communication section 409 performs communication processing via a network such as the Internet. A drive 410 is also connected to the I / O interface 405 as needed. A removable medium 411, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 410 as needed so that the computer program read from it can be installed into the storage section 408 as needed.

[0062] When the central processing unit 401 in the electronic device of this embodiment executes the program, it implements the steps in the deployment method of the large model in the power grid field as Figure 1 shown.

[0063] Specifically, according to the embodiments of the present application, the process described above with reference to the flowchart can be implemented as a computer software program. For example, the embodiments of the present application include a computer program product, which includes a computer program carried on a computer-readable medium, and this computer program contains program codes for executing Figure 1 the method shown. In such an embodiment, this computer program can be downloaded and installed from the network through the communication part 409, and / or installed from the removable medium 411. When this computer program is executed by the central processing unit 401, it executes various functions defined in the device of the present application.

[0064] Among them, Figure 1 the computer program instructions corresponding to the method shown can also be stored in a computer-readable memory that can guide a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in this computer-readable memory generate a manufactured product including an instruction device, and this instruction device implements the functions specified in Figure 1 one process or multiple processes and / or Figure 1 one block or multiple blocks.

[0065] Those of ordinary skill in the art can understand that to implement all or part of the processes in the above-mentioned embodiment methods, it can be completed by instructing relevant hardware through a computer program. The said program can be stored in a computer-readable storage medium. When this program is executed, it can include the processes of the embodiments of the above-mentioned various methods. Among them, the said storage medium can be a magnetic disk, an optical disc, a read-only memory (ROM), or a random access memory (RAM), etc.

[0066] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A deployment method for a large model in the power grid field, characterized in that, including: Based on the relationships between the memory overheads in the parallel decomposition strategy for computing-storage collaboration with respect to parameters such as the memory overhead of a single parameter, the memory overhead of a single gradient, the memory overhead of a single optimizer state, the number of model parameters, and the number of worker nodes in data parallelism, model parallelism, and pipeline parallelism, calculate the parallel decomposition strategy for computing-storage collaboration with the minimum memory overhead, and determine the parallel decomposition strategy for the large model in the power grid domain; Based on the parallel decomposition strategy for the large model in the power grid domain, use hyperparameters to input hardware device information and deep learning frameworks, automatically allocate hardware resources, and then use guiding statements to set the deep learning model structure and configure the parallel decomposition scheme, complete automatic code translation, and achieve automatic conversion and training of the deep learning model.

2. The deployment method of the large model in the power grid field according to claim 1, characterized in that, The memory overhead P for a single parameter, the memory overhead G for a single gradient, the memory overhead O for a single optimizer state, the number of model parameters ψ, and the number of worker nodes N in data parallelism d The relationship between these parameters and the memory overhead M′ of the parallel decomposition strategy for storage collaboration is as follows:

3. The deployment method of the large model in the power grid field according to claim 2, characterized in that, The number of model parameters ψ is: ψ = Np * Pre / 8; where Pre is the floating-point precision for calculation.

4. The deployment method of the large model in the power grid field according to claim 1, characterized in that, The parallel decomposition strategy for the large model in the power grid domain includes inter-group data parallelism, intra-group model parallelism, pipeline parallelism, and optimizer parallelism strategies.

5. The deployment method of the large model in the power grid field according to claim 4, wherein, The in-group model parallelism strategy is as follows: where M′ is the memory overhead for storing the parallel decomposition strategy for collaboration, M0 is the memory limit of a single device, and N1 is the in-group device allocation volume.

6. The deployment method of the large model in the power grid field according to claim 4, characterized in that, The inter-group data parallel strategy is as follows: where N1 is the number of devices allocated within a group, N2 is the number of devices allocated between groups, and N is the total number of available devices.

7. A deployment system for a large model in the power grid field, characterized in that, including: A parallel decomposition strategy determination module, which is used to calculate the parallel decomposition strategy for computing-storage collaboration with the minimum memory overhead and determine the parallel decomposition strategy for the large model in the power grid domain based on the relationships between the memory overheads in the parallel decomposition strategy for computing-storage collaboration with respect to parameters such as the memory overhead of a single parameter, the memory overhead of a single gradient, the memory overhead of a single optimizer state, the number of model parameters, and the number of worker nodes in data parallelism, model parallelism, and pipeline parallelism; A model automatic conversion and training module, which is used to automatically allocate hardware resources based on the parallel decomposition strategy for the large model in the power grid domain, use hyperparameters to input hardware device information and deep learning frameworks, and then use guiding statements to set the deep learning model structure and configure the parallel decomposition scheme, complete automatic code translation, and achieve automatic conversion and training of the deep learning model.

8. The deployment system of the large model in the power grid field according to claim 7, characterized in that, In the parallel decomposition strategy determination module, the memory overhead P of a single parameter, the memory overhead G of a single gradient, the memory overhead O of a single optimizer state, the number of model parameters ψ, and the number of worker nodes N in data parallelism d The relationship between these parameters and the memory overhead M' of the parallel decomposition strategy for storage collaboration is as follows:

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by a processor, it implements the steps in the deployment method of the large model in the power grid domain as described in any one of claims 1-6.

10. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps in the deployment method of the large model in the power grid domain as described in any one of claims 1-6.