A Resource Allocation Method and System Based on Tensor Parallelism

By dividing computing devices into multiple computing groups according to their working life, and dividing computing tasks according to data characteristics, optimizing computing resource allocation, the problems of insufficient computing resources and equipment performance decay in the power grid environment are solved, and more efficient computing resource utilization and equipment life extension are achieved.

CN119645665BActive Publication Date: 2025-06-13NINGBO ELECTRIC POWER DESIGN INST
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510174186.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-18
Publication Date
2025-06-13
Estimated Expiration
2045-02-18

AI Technical Summary

Technical Problem

The existing computing resource allocation methods are difficult to meet long-term computing needs in the power grid environment, resulting in accelerated decay of equipment performance, increasing usage costs, and due to the huge amount of data and complex types, the load of computing equipment is unbalanced, reducing the model's inference processing efficiency.

Method used

By dividing it into multiple calculation groups based on the remaining working life of the computing device, hardware information is obtained and the optimal and maximum calculation load of each calculation group is calculated; the data to be calculated is divided into submodules, and divided into tensor modules or pipeline modules according to the data type and hierarchical structure information to determine priority; the submodules are tested in turn, the calculation load amount and communication efficiency of the corresponding calculation group of each submodule are calculated, and the resource allocation is performed according to the load margin and communication efficiency, and the utilization rate of the calculation resources is optimized.

Benefits of technology

By optimizing the management strategy and computing resource allocation of computing devices, we can extend the service life of the equipment, improve the reliability and stability of computing devices, reduce the demand for computing resources, improve computing efficiency, and avoid equipment load overload and inefficient communication.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119645665B_ABST
    Figure CN119645665B_ABST
Patent Text Reader

Abstract

The technology of the present invention relates to the field of large language models, and specifically, to a resource allocation method and system based on tensor parallelism. The problem solved by the present invention is: how to meet computing needs when computing resources are insufficient and extend the service life of equipment. To solve the above problem, the present invention provides a resource allocation method, including: dividing computing groups, calculating the optimal computing load and the maximum computing load of the computing groups; dividing sub-modules into tensor modules or pipeline modules, and dividing priorities; calculating the computing load; calculating the communication efficiency; calculating the load margin; allocating sub-modules to each computing group to obtain a computing resource allocation plan; if the current load is greater than or equal to the maximum computing load, giving priority to calculating some sub-modules, and marking the sub-modules that are not given priority as modules to be allocated; and allocating the modules to be allocated to the computing group for assisting calculation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present invention relates to the field of large language models. Specifically, it relates to a resource allocation method and system based on tensor parallelism. Background Art

[0002] With the rapid development of artificial intelligence technology, the power grid field has also witnessed unprecedented changes. As an important part of artificial intelligence technology, large language models are gradually penetrating into all aspects of the power system, providing new solutions for the intelligence and efficiency of the power grid, and helping power companies predict future power demands and formulate plans, optimize resource allocation, and improve overall service quality in a more intelligent way. However, in actual practice, due to the need to train the model and continuously obtain a large amount of grid operation data such as measurement data and meteorological data for a long time, there is a long-term demand for computing resources. Existing computing resource allocation methods usually give priority to computing efficiency, putting computing devices in a high-load and high-performance state, resulting in an accelerated decline in the performance of computing devices, which does not meet the requirements for long-term use in the power grid environment, greatly increasing the usage cost of large language models. Furthermore, due to the huge amount of system data and complex data types in power grid operation, it is easy to cause uneven device loads of computing devices in each module, resulting in low efficiency in the inference processing process of the model and wasting a large amount of computing resources. Summary of the Invention

[0003] The problem solved by the present invention is: how to meet the computing requirements in the case of insufficient computing resources and extend the service life of the device.

[0004] To solve the above problems, an embodiment of the present invention provides a resource allocation method based on tensor parallelism. The resource allocation method includes: dividing computing devices into multiple computing groups according to the remaining working life of the computing devices, obtaining the hardware information of the computing devices, and calculating the optimal computing load and the maximum computing load of the computing groups according to the hardware information and the computing groups; splitting the data to be calculated into multiple sub-modules according to the data type and the hierarchical structure information, determining the applicable computing method for the sub-modules, dividing the sub-modules into tensor modules or pipeline modules, and dividing priorities; testing the sub-modules sequentially to obtain test data, and calculating the computing load of each sub-module corresponding to each computing group according to the test data; calculating the communication efficiency when data is transmitted between each computing group according to the test data; obtaining the current load of each computing group, and calculating the load margin according to the optimal computing load and the current load; allocating the tensor modules and the pipeline modules to each computing group according to the load margin, the computing load, and the communication efficiency to obtain a computing resource allocation plan; if the current load is less than the maximum computing load, continue to allocate the computing resources according to the computing resource allocation plan; if the current load is greater than or equal to the maximum computing load, re-allocate the computing resources according to the priorities of the sub-modules, perform priority calculation on some sub-modules, and mark the sub-modules that are not prioritized after re-allocation as to-be-allocated modules; if the load margin of the computing group is greater than or equal to the computing load of the to-be-allocated module, allocate the to-be-allocated module to the computing group for assisted calculation.

[0005] Compared with the prior art, the technical effects achieved by adopting this technical solution are as follows: By dividing the computing devices into multiple computing groups according to the working life, the management strategy of the computing devices and the computing resource allocation can be optimized, the reliability and stability of the devices can be improved, and the reduction of computing efficiency caused by too large a performance gap between computing devices can be avoided. Further, it is convenient for subsequent maintenance and update of the computing devices. By dividing the tensor modules and the pipeline modules, the processing efficiency of a large amount of data to be calculated is improved, and the overload of the computing device caused by too large a data volume is avoided, so that different data can select appropriate computing methods according to their own characteristics, reducing the demand for computing resources. The calculation of the computing load can obtain the specific load situation when each sub-module is calculated in each computing group, providing data support for subsequent sub-module allocation. By calculating the communication efficiency, the specific situation of data transmission between each computing device can be obtained, avoiding the reduction of computing efficiency caused by too low communication efficiency in the future. By calculating the load margin, the load situation of each computing group when calculating the sub-modules can be obtained, and further providing data support for subsequent optimization of task allocation. By obtaining the computing resource allocation plan, the consumption of computing resources can be reduced as much as possible while meeting the computing efficiency requirements for the data to be calculated. Allocating the to-be-allocated module to the computing group for assisted calculation can optimize the utilization rate of computing resources and improve the overall computing efficiency.

[0006] In one embodiment of the present invention, computer devices with similar remaining working lives are divided into several computing groups; according to the hardware information, the load capacity that each computing group can withstand when operating at the expected power is calculated to obtain the optimal computing load, and the load capacity that each computing group can withstand when operating at full power is calculated to obtain the maximum computing load.

[0007] Compared with the prior art, the technical effects achieved by adopting this technical solution are as follows: by calculating the optimal computing load, the load capacity of each computing group under ideal conditions can be determined, providing data support for subsequent sub-module allocation, making the load distribution of the computing device more reasonable, and extending the working life of the computing device. By calculating the maximum computing load, the maximum load that each computing group can withstand can be obtained, ensuring that the computing group will not cause damage to the computing device and data loss due to overloading during high-load operation.

[0008] In one embodiment of the present invention, data of the same type and at the same level are merged according to the data type and hierarchical structure information, and then divided into several sub-modules according to the performance of the computing device; according to the data type, the applicable computing method for each sub-module is judged. If a sub-module is more suitable for tensor parallelism, it is marked as a tensor module. If a sub-module is more suitable for pipeline parallelism, it is marked as a pipeline module; according to the hierarchical structure information, the computing subordination structure of each sub-module during the computing process is judged to obtain the priority.

[0009] Compared with the prior art, the technical effects achieved by adopting this technical solution are as follows: by obtaining the priority, the task scheduling and resource allocation during the computing process of the computing device can be made more reasonable, ensuring that key data can be processed preferentially. Further, it provides data support for subsequent normal computing according to the parallel strategy of sub-modules in the case of tight computing resources, optimizing the overall data computing logic.

[0010] In one embodiment of the present invention, all sub-modules are sent to each computing group, and the sub-modules are sequentially tested within each computing group to obtain the load change and computing results of the computing group, thereby obtaining test data; according to the test data, the maximum load value of the computing group during the test time is obtained to obtain the computing load.

[0011] Compared with the prior art, the technical effects achieved by adopting this technical solution are as follows: by obtaining the test data, the actual processing situation of each computing group for the sub-module can be accurately obtained, making the subsequent computing resource allocation plan more reasonable. By obtaining the computing load, it can help allocate each sub-module, further improving the utilization rate of computing resources.

[0012] In an embodiment of the present invention, test data of each sub-module is obtained, and the transmission test of each calculation result is carried out among all calculation groups. The change situation of the transmission speed of the calculation results of each sub-module among each calculation group is obtained to obtain communication data. According to the communication data, the average value of the transmission speed within the test time is calculated to obtain the communication efficiency.

[0013] Compared with the prior art, the technical effects achieved by adopting this technical solution are as follows: By obtaining the communication data, the transmission situation of the calculation results of different sub-modules among different calculation groups can be understood in detail, providing data support for the subsequent calculation of the communication efficiency, making the calculation of the communication efficiency more accurate. By calculating the average value of the transmission speed, the numerical value of the actual transmission rate can be obtained more accurately, improving the accuracy of the communication efficiency.

[0014] In an embodiment of the present invention, according to the load margin and the calculated load amount, the tensor module and the pipeline module are allocated to each calculation group, and the scheme with the minimum total load amount is calculated and recorded as the first allocation scheme. According to the first allocation scheme and the communication efficiency, the calculation efficiency of the first allocation scheme is calculated to obtain the first calculation efficiency. If the first calculation efficiency is greater than or equal to the expected efficiency, the calculation resource allocation plan is planned according to the first allocation scheme. If the first calculation efficiency is less than the expected efficiency, the first allocation scheme is adjusted according to the communication efficiency to obtain the second allocation scheme, and the calculation efficiency of the second allocation scheme is calculated to obtain the second calculation efficiency. When the second calculation efficiency is greater than or equal to the expected efficiency, the calculation resource allocation plan is planned according to the second allocation scheme.

[0015] Compared with the prior art, the technical effects achieved by adopting this technical solution are as follows: By obtaining the first allocation scheme, the allocation logic of the sub-modules can be simplified, and the allocation scheme with the minimum total load amount can be calculated quickly, improving the allocation efficiency of the sub-modules. By calculating the first calculation efficiency and the second calculation efficiency, it can be directly seen whether the allocation scheme meets the minimum requirement of the calculation efficiency, ensuring that the data to be processed can be calculated in time, improving the rationality of the calculation resource allocation, and improving the utilization rate of the calculation resources.

[0016] In an embodiment of the present invention, according to the calculation resource allocation plan, each sub-module is allocated and calculated. If the current load is less than the maximum calculation load, the allocation and calculation of each sub-module are continued according to the calculation resource classification plan. If the current load is greater than or equal to the calculation load, the priorities of all sub-modules in this calculation group are obtained, and according to the priorities and the maximum calculation load, the sub-modules are re-allocated, and some sub-modules are preferentially calculated. All sub-modules that have not been preferentially calculated are obtained to obtain the modules to be allocated.

[0017] Compared with the prior art, the technical effects achieved by adopting this technical solution are as follows: By comparing the current load and the maximum calculated load, it can be ensured that the calculation group will not be overloaded, which may lead to equipment damage or data loss, thus improving the security of the computing device. By preferentially calculating some sub-modules, it is ensured that normal calculation can still be maintained in the case of insufficient computing resources, further guaranteeing the computing efficiency and equipment security. Marking the modules to be allocated enables the sub-modules that have stopped calculating to be re-allocated and calculated in a timely manner, improving the overall computing efficiency.

[0018] In an embodiment of the present invention, the load margins of each calculation group are compared with the calculation load amounts of each module to be allocated; when the load margin is greater than or equal to the calculation load amount, the module to be allocated is allocated to the calculation group for assisted calculation; if there is no load margin less than the calculation load amount, no assisted calculation is performed.

[0019] Compared with the prior art, the technical effects achieved by adopting this technical solution are as follows: By comparing the load margin and the calculation load of the module to be allocated to determine whether to perform assisted calculation on the module to be allocated, the allocation logic of the sub-modules is optimized, the load sizes among the calculation groups are balanced, and the utilization rate of computing resources and the overall computing efficiency are improved.

[0020] In an embodiment of the present invention, a resource allocation system based on tensor parallelism is further provided. The resource allocation method based on tensor parallelism described in the above embodiment is applied to this resource allocation system. The resource allocation system includes: a storage module for storing hardware information and sub-modules; a processing module for splitting and obtaining sub-modules and dividing them into tensor modules and pipeline modules; a calculation module for calculating the calculation load amount, communication efficiency, and load margin; and an allocation module for executing the computing resource allocation plan. This resource allocation system has all the technical features of the above resource allocation method and will not be elaborated here one by one. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] Figure 1 One of the flowcharts of the resource allocation method based on tensor parallelism of the present invention;

[0022] Figure 2 Another flowchart of the resource allocation method based on tensor parallelism of the present invention;

[0023] Figure 3 Another flowchart of the resource allocation method based on tensor parallelism of the present invention;

[0024] Figure 4 Another flowchart of the resource allocation method based on tensor parallelism of the present invention;

[0025] Figure 5System schematic diagram of the resource allocation system based on tensor parallelism of the present invention;

[0026] Explanation of reference numerals:

[0027] 100 - Resource allocation system; 110 - Storage module; 120 - Processing module; 130 - Computing module; 140 - Allocation module. Detailed implementation manners

[0028] To make the above objects, features and advantages of the present invention more obvious and understandable, the following detailed description of the specific embodiments of the present invention will be given with reference to the accompanying drawings.

[0029]

First Embodiment

[0030] Refer to Figure 1 , in a specific embodiment, the present invention provides a resource allocation method based on tensor parallelism. The resource allocation method includes:

[0031] S100. Divide the computing devices into multiple computing groups according to the remaining working life of the computing devices, obtain the hardware information of the computing devices, and calculate the optimal computing load and the maximum computing load of the computing groups according to the hardware information and the computing groups;

[0032] S200. Split the data to be calculated into multiple sub - modules according to the data type and the hierarchical structure information, determine the applicable computing methods for the sub - modules, divide the sub - modules into tensor modules or pipeline modules, and assign priorities;

[0033] S300. Test the sub - modules in sequence to obtain test data, and calculate the computing load of each sub - module corresponding to each computing group according to the test data;

[0034] S400. Calculate the communication efficiency when data is transmitted between each computing group according to the test data;

[0035] S500. Obtain the current load of each computing group, and calculate the load margin according to the optimal computing load and the current load;

[0036] S600. Allocate the tensor modules and pipeline modules to each computing group according to the load margin, the computing load, and the communication efficiency to obtain a computing resource allocation plan;

[0037] S700. If the current load is less than the maximum computing load, continue to allocate computing resources according to the computing resource allocation plan; if the current load is greater than or equal to the maximum computing load, re - allocate computing resources according to the priorities of the sub - modules, perform priority calculation on some sub - modules, and mark the sub - modules that are not preferentially calculated after re - allocation as modules to be allocated;

[0038] S800. If the load margin of the computing group is greater than or equal to the computing load of the module to be allocated, allocate the module to be allocated to this computing group for assisted computing.

[0039] In step S100, usually, all computing devices should be of the same model. If there are computing devices of different models, the computing power needs to be converted according to the hardware of the computing devices. The remaining working life refers to the remaining time that the computing device can work normally. Since with the aging, replacement, and expansion of computing devices, there are differences in the remaining working lives of each computing device, and there are also differences in the computing performance of computing devices with different remaining working lives. Further, there are also differences in the optimal computing load and maximum computing load of each computing device.

[0040] It should be noted that the working life of the computing device depends on the device model and actual usage.

[0041] In step S200, the data to be calculated refers to data related to power grid operation such as large language model training data, measurement data, meteorological data, and power distribution data. Usually, the data types of these data are complex and the data volume is extremely large. If these data are directly calculated, it will generate too much computing load on the computing device. Therefore, the data of the same type or related data in the data to be calculated are integrated and processed and split into multiple sub-modules.

[0042] It should be noted that the data volume of the sub-module is set according to the performance of the computing device. If the performance of the computing device is relatively good, the data volume of the sub-module can be set relatively large. If the performance of the computing device is relatively poor, the data volume of the sub-module can be set relatively small. The specific standard depends on the actual situation, and there may be large differences in the data volumes of each sub-module.

[0043] The tensor module refers to a data group that uses tensor parallelism for computing, usually applicable to high-dimensional data such as images, voices, and texts in power grid operation data. For example, data such as cable laying images and power grid conference voice records. The pipeline module refers to a data group that uses pipeline parallelism for computing, usually applicable to data for training large prediction models. For example, training an intelligent power supply system model, etc.

[0044] In step S300, since there are differences in the performance of computing devices in different computing groups, and there are also differences in the data types and data volumes of different sub-modules, the computing loads of different computing groups for each sub-module are also different. Therefore, select some data in each sub-module as test data and send it to each computing group for testing to avoid subsequent allocation of sub-modules to computing groups with too large computing loads, resulting in excessive consumption of computing resources.

[0045] In step S400, the communication efficiency refers to the proportion of the actual data transmission volume per unit time between computing groups in the expected data transmission volume. The calculation formula is as follows:

[0046] H c =C e / C a ×100%, where H c is the communication efficiency, C e is the actual transmission rate, and C a is the expected transmission rate.

[0047] For example, if the expected transmission rate is 20 gbps and the actual transmission rate is 16.5 gbps, then the communication efficiency H c =C e / C a ×100% = (16.5 gbps / 20 gbps) × 100% = 82.5%.

[0048] Since some sub - modules may need to cooperate in the calculation of the tensor module and the pipeline module, the communication efficiency between each computing group may affect the overall computing efficiency. For example, if computing group A is responsible for calculating tensor module A and computing group B is responsible for calculating tensor module B, and the calculation of tensor module B requires the calculation result of tensor module A as support, if the communication efficiency between computing group A and computing group B for this calculation result is not 100%, it means that the computing efficiency will be affected due to the inability to communicate in real time.

[0049] It should be noted that usually, the communication efficiency during data transmission within a computing group is defaulted to 100%.

[0050] In step S500, due to the long - term usage requirements for computing devices in the power grid environment, under the condition of ensuring to meet the computing requirements, all computing devices should be ensured to work under the optimal load as much as possible to extend the working life of the computing devices. Therefore, the load margin refers to the difference between the current load and the optimal computing load. It should be noted that since the value of the current load usually fluctuates during the actual computing process, the average load value of the computing group within a certain time period is taken as the current load, and the length of the time period depends on specific requirements.

[0051] In step S600, since the load margins of each computing group are different, the computing loads of each sub-module in each computing group are different, and there are also differences in the communication efficiency between each computing group. In order to avoid overloading the computing device due to excessive power grid operation data and affecting the overall system operation, when classifying sub-modules, multiple factors such as load margin, computing load, and communication efficiency should be fully considered. On the premise of ensuring computing efficiency, the total computing load should be reduced to the lowest possible level as much as possible to relieve the pressure of insufficient computing resources and make the computing of the data to be calculated more efficient.

[0052] In step S700, during the actual operation of the computing group, load fluctuations may occur during the sub-module computing process, resulting in the current load not being able to be maintained around the optimal computing load, or even exceeding the maximum computing load. When there is too much data to be calculated, it will also lead to the inability to ensure computing efficiency at the optimal computing load. Therefore, in order to be able to process the data to be calculated in a timely manner, sub-modules are still allocated to the computing group whose current load is greater than the optimal computing load and less than the maximum computing load. However, when the current load is greater than or equal to the maximum computing load, it means that the computing resources of the computing group are no longer able to support the calculation of the current data volume, that is, the computing group cannot bear the current load. In order to avoid equipment damage and data loss caused by overloading, the calculation of one or more sub-modules needs to be suspended, and the sub-modules with higher priorities are calculated first to ensure that the current load is always less than the maximum computing load and protect the safety of equipment and data.

[0053] In step S800, since the data types and data volumes of different sub-modules are different, the computing time required is also different. Some computing groups may complete the calculation of sub-modules faster than other computing groups, and the current load will accordingly decrease, and the load margin will increase. In order to relieve the burden on other computing groups and avoid waste of computing resources, if the load margin is sufficient to support the calculation of the module to be allocated, that is, when the load margin is greater than or equal to the computing load of the module to be allocated, the module to be allocated is allocated to this computing group for assisted calculation.

[0054] By dividing computing devices into multiple computing groups according to their working lifetimes, the management strategies and computing resource allocations of the computing devices can be optimized, the reliability and stability of the devices can be improved, the reduction in computing efficiency caused by excessive performance gaps between computing devices can be avoided, and furthermore, the subsequent maintenance and update of the computing devices can be facilitated. By dividing the tensor module and the pipeline module, the processing efficiency of a large amount of data to be computed is improved, the overload of the computing devices caused by excessive data volume is avoided, different data can select appropriate computing methods according to their own characteristics, the demand for computing resources is reduced, the calculation of the computing load can obtain the specific load conditions of each sub-module when computing in each computing group, providing data support for subsequent sub-module allocation. By calculating the communication efficiency, the specific situation of data transmission between each computing device can be obtained, avoiding the reduction in computing efficiency caused by too low communication efficiency in the future. By calculating the load margin, the load conditions of each computing group when computing the sub-module can be obtained, and then providing data support for subsequent optimization of task allocation. By obtaining the computing resource allocation plan, the data to be computed can minimize the consumption of computing resources as much as possible under the premise of meeting the computing efficiency requirements. Allocating the module to be allocated to this computing group for assisted computing can optimize the utilization rate of computing resources and improve the overall computing efficiency.

[0055]

Second Embodiment

[0056] In a specific embodiment, the computing devices are divided into multiple computing groups according to the remaining working lifetimes of the computing devices, the hardware information of the computing devices is obtained, and according to the hardware information and the computing groups, the optimal computing load and the maximum computing load of the computing groups are calculated, specifically including:

[0057] S110. Divide the computing devices with similar remaining working lifetimes into several computing groups;

[0058] S120. According to the hardware information, calculate the load size that the computing group can bear when operating at the expected power to obtain the optimal computing load, and calculate the load size that the computing group can bear when operating at full power to obtain the maximum computing load.

[0059] In step S110, the computing devices are divided according to the remaining working lifetimes, and the specific division criteria depend on the actual situation. For example, if the working lifetime of the computing device is 10 years, the computing devices can be divided into five computing groups with remaining working lifetimes of (0, 2] years, (2, 4] years, (4, 6] years, (6, 8] years, and (8, 10] years respectively.

[0060] It should be noted that the number of computing groups may be multiple or one. If the remaining working lifetimes of all computing devices are similar, only one computing group needs to be divided.

[0061] In step S120, the setting of the desired power should ensure that the device can operate in an optimal state, taking into account the energy efficiency and lifespan of the computing device. It can be 280W or 300W, and can be adjusted according to specific requirements. The optimal computing load refers to the total load capacity that all computing devices in the computing group can bear when operating at the desired power.

[0062] Full power refers to the maximum rated power of the computing device. The maximum computing load refers to the total load capacity that all computing devices in the computing group can bear when operating at full power. It should be noted that since the load may fluctuate during the data calculation process, in order to avoid exceeding the maximum limit due to the load, a certain margin should be reserved when setting the maximum computing load. The margin can be 10% or 15%, with the principle of ensuring that the maximum value of the fluctuation range does not exceed the maximum computing load.

[0063] By calculating the optimal computing load, the load capacity of each computing group under ideal conditions can be determined, providing data support for the subsequent allocation sub-module, making the load allocation of the computing device more reasonable, and extending the working lifespan of the computing device. By calculating the maximum computing load, the maximum load capacity that each computing group can bear can be obtained, ensuring that the computing group will not cause damage to the computing device and data loss due to overload during high-load operation.

[0064]

Third Embodiment

[0065] See Figure 2 , in a specific embodiment, according to the data type and hierarchical structure information, the data to be calculated is split into multiple sub-modules, the applicable calculation method for the sub-module is judged, the sub-module is divided into a tensor module or a pipeline module, and the priority is divided, specifically including:

[0066] S210. According to the data type and hierarchical structure information, the data of the same type and the same layer are merged, and according to the performance of the computing device, it is divided into several sub-modules;

[0067] S220. According to the data type, judge the applicable calculation method for the sub-module. If the sub-module is more suitable for tensor parallelism, it is marked as a tensor module. If the sub-module is more suitable for pipeline parallelism, it is marked as a pipeline module;

[0068] S230. According to the hierarchical structure information, judge the calculation subordination structure of each sub-module during the calculation process to obtain the priority.

[0069] In step S210, due to the rich data types and complex data structures of grid data, which contain a large number of complex and multi-dimensional tasks that usually require a large amount of computing resources and a heavy computing load, a single computing device cannot complete the calculations independently. Therefore, data of the same type and at the same level are merged, and further, complex tasks are decomposed into multiple sub-modules, with each sub-module solving a specific problem. It should be noted that the size of the sub-module should be considered in light of the load capacity of a single computing device, with the principle that each computing device can independently calculate at least one sub-module.

[0070] In steps S220 to S230, since tensor parallelism involves splitting data into multiple tensor modules and distributing them to different computing devices for parallel computing, and pipeline parallelism distributes different pipeline modules of data to different computing devices to form a pipeline for parallel computing, and tensor parallelism and pipeline parallelism are usually combined in a complex data processing task to achieve finer-grained parallel computing. Therefore, in actual calculations, there may be a dependency relationship between tensor modules and pipeline modules. For example, pipeline module B needs to calculate based on the calculation results of pipeline module A, and tensor module A needs to calculate based on the calculation results of pipeline module B and pipeline module C. Therefore, priorities should be set for each sub-module according to the dependency relationship between the sub-modules to ensure that when there are too many sub-modules and the computing devices cannot calculate all sub-modules simultaneously, the more fundamental sub-modules can be calculated first to ensure that data can be processed normally.

[0071] By obtaining the priorities, the task scheduling and resource allocation during the calculation process of the computing devices can be made more reasonable, ensuring that critical data can be processed first. Further, it provides data support for subsequent normal calculation according to the parallel strategy of sub-modules in the case of tight computing resources, optimizing the overall data calculation logic.

[0072]

Fourth Embodiment

[0073] In a specific embodiment, the sub-modules are tested in sequence to obtain test data. Based on the test data, the computing load of each sub-module corresponding to each computing group is calculated, specifically including:

[0074] S310: Send all sub-modules to each computing group, and test the sub-modules in sequence within each computing group to obtain the load change and calculation results of the computing group, thereby obtaining test data;

[0075] S320: Based on the test data, obtain the maximum load of the computing group during the test time to obtain the computing load.

[0076] In step S310, when conducting the test, it should be ensured that each computing group operates at the corresponding expected power, and the test time for each test is the same. The test data includes the load change situation and calculation results of the computing group, where the calculation results are used to subsequently test the communication efficiency between each computing group.

[0077] In step S320, the test time can be one minute or ten minutes, depending on the specific computing requirements and device performance. The longer the test time, the higher the accuracy of the obtained test data. Since the load fluctuates during the actual computing process, to avoid the load of the computing group exceeding the limit that the device can bear due to load fluctuations, the maximum value in the load change curve within the test time is used as the calculated load amount. The calculated load amount refers to the load size generated when the sub-module is processed in the computing group. It should be noted that since the remaining working life of different computing groups is different, their performances are also different, and the calculated load amount for processing the same sub-module is also different.

[0078] By obtaining the test data, the actual processing situation of each computing group for the sub-module can be accurately obtained, making the subsequent calculation resource allocation plan more reasonable. By obtaining the calculated load amount, it can help allocate each sub-module, further improving the utilization rate of the calculation resources.

[0079]

Fifth Embodiment

[0080] In a specific embodiment, according to the test data, calculate the communication efficiency when data is transmitted between each computing group, specifically including:

[0081] S410. Obtain the test data of each sub-module, conduct a transmission test of each calculation result among all computing groups, obtain the change situation of the transmission speed of the calculation results of each sub-module among each computing group, and obtain the communication data;

[0082] S420. According to the communication data, calculate the average value of the transmission speed within the test time to obtain the communication efficiency.

[0083] In step S410, the communication data includes the changes in all transmission speeds within the test time. Since the data types and data complexities of different sub-modules are different, there are also differences in the transmission speeds of the calculation results for different sub-modules. For example, the transmission speed of text data is much greater than that of image data. Therefore, the transmission tests of each calculation result are carried out among all calculation groups to obtain the changes in the transmission speeds of the calculation results of each sub-module among different calculation groups. For example, for calculation result A, the transmission rate between calculation group A and calculation group B is 20 gbps, and the transmission rate between calculation group B and calculation group C is 16.5 gbps. For calculation result B, the transmission rate between calculation group A and calculation group B is 25 gbps, and the transmission rate between calculation group B and calculation group C is 19.3 gbps.

[0084] In step S420, since the transmission speed fluctuates during the transmission of a certain calculation result, in order to calculate the communication efficiency more accurately, the average value of the transmission speed within the calculation test time is used as the actual transmission rate for calculation.

[0085] By obtaining the communication data, it is possible to understand the transmission situation of the calculation results of different sub-modules among different calculation groups, provide data support for the subsequent calculation of communication efficiency, make the calculation of communication efficiency more accurate, and by calculating the average value of the transmission speed, the value of the actual transmission rate can be obtained more accurately, improving the accuracy of communication efficiency.

[0086]

Sixth Embodiment

[0087] See Figure 3 In a specific embodiment, according to the load margin, calculation load, and communication efficiency, the tensor module and the pipeline module are allocated to each calculation group to obtain a calculation resource allocation plan, which specifically includes:

[0088] S610. According to the load margin and calculation load, allocate the tensor module and the pipeline module to each calculation group, and calculate the plan with the minimum total load, which is recorded as the first allocation plan;

[0089] S620. According to the first allocation plan and communication efficiency, calculate the calculation efficiency of the first allocation plan to obtain the first calculation efficiency;

[0090] S630. If the first calculation efficiency is greater than or equal to the expected efficiency, then plan according to the first allocation plan to obtain a calculation resource allocation plan;

[0091] S640. If the first calculation efficiency is less than the expected efficiency, then adjust the first allocation plan according to the communication efficiency to obtain a second allocation plan, and calculate the calculation efficiency of the second allocation plan to obtain the second calculation efficiency;

[0092] S650. When the second computing efficiency is greater than or equal to the expected efficiency, perform planning according to the second allocation scheme to obtain a computing resource allocation plan.

[0093] In step S610, in order to minimize the consumption of computing resources as much as possible, when allocating sub-modules, the scheme with the smallest total load should be given priority. The total load refers to the sum of the computing loads of all tensor modules and pipeline modules.

[0094] It should be noted that during the allocation process, in order to protect the lifespan of the computing device, it should be ensured as much as possible that the computing load of the sub-modules allocated to each computing group does not exceed the load margin of that computing group. If the load margin is not sufficient to meet the computing load of the sub-module, then it should be ensured that the load of each computing group does not exceed the maximum computing load. If necessary, only some sub-modules are used for computing to relieve the computing load pressure.

[0095] In step S620, in the first allocation scheme, sub-modules with a subordinate relationship may be allocated to different computing groups, and the communication efficiency between computing groups will affect the overall computing efficiency, which may cause the computing efficiency not to reach the expected value. The first computing efficiency refers to the computing efficiency when computing according to the first allocation scheme. For example, for a certain power grid data processing task, tensor module A is allocated to computing group A, tensor module B is allocated to computing group B, pipeline module A is allocated to computing group B, the computation of tensor module B is based on the computation data of tensor module A, and the computation of pipeline module A is based on the computation data of tensor module B. Then the communication efficiency between tensor module A and tensor module B is 85%, and the communication efficiency between tensor module B and pipeline module A is 100%. Then according to the subordinate relationship, the first computing efficiency = 85% × 100% = 85%.

[0096] In steps S630 to S650, the expected efficiency refers to the minimum computational efficiency requirement set to ensure the normal operation of the power grid. For different data to be processed, the expected efficiency may vary, being 50% or 90% for example. If the first computational efficiency is greater than or equal to the expected efficiency, it indicates that the computational efficiency of the first allocation scheme is sufficient to meet the power grid demand and can be directly used as the computational resource allocation plan to perform calculations on all sub-modules. If the first computational efficiency is less than the expected efficiency, it means that the computational efficiency is too low to meet the power grid demand. Under the premise of ensuring that the current load of each computational group does not exceed the maximum computational load, the allocation of sub-modules needs to be adjusted. For example, according to the first allocation scheme, tensor module A and pipeline module A are allocated to computational group A, and tensor module B and pipeline module B are allocated to computational group B. The calculation of tensor module B is based on the calculation data of tensor module A, and the calculation of pipeline module B is based on the calculation data of pipeline module B. The calculated first computational efficiency is 73%, and the expected efficiency is 95%. Then, the first allocation scheme is adjusted by allocating tensor module B to computational group A and pipeline module B to computational group B to obtain the second allocation scheme. Similar to the calculation method of the first computational efficiency, the calculated second computational efficiency is 100%, which is greater than the expected efficiency, and the second allocation scheme is used as the computational resource allocation plan.

[0097] It should be noted that during the process of adjusting the first allocation scheme to obtain the second allocation scheme, the total load should also be minimized as much as possible.

[0098] By obtaining the first allocation scheme, the allocation logic of sub-modules can be simplified, the allocation scheme with the minimum total load can be quickly calculated, and the allocation efficiency of sub-modules can be improved. By calculating the first computational efficiency and the second computational efficiency, it can be intuitively obtained whether the allocation scheme meets the minimum requirement of computational efficiency, ensuring that the data to be processed can be calculated in a timely manner, improving the rationality of computational resource allocation, and improving the utilization rate of computational resources.

[0099]

Seventh Embodiment

[0100] See Figure 4 , in a specific embodiment, if the current load is less than the maximum computational load, the computational resources are continuously allocated according to the computational resource allocation plan; if the current load is greater than or equal to the maximum computational load, the computational resources are reallocated according to the priorities of the sub-modules, and some sub-modules are preferentially calculated, and the sub-modules that are not preferentially calculated after the reallocation are marked as modules to be allocated, which specifically includes:

[0101] S710. Allocate and calculate each sub-module according to the computational resource allocation plan;

[0102] S720. If the current load is less than the maximum calculated load, continue to allocate and calculate each sub-module according to the classification plan of computing resources.

[0103] S730. If the current load is greater than or equal to the calculated load, obtain the priorities of all sub-modules in this calculation group, and re-allocate the sub-modules according to the priorities and the maximum calculated load, and preferentially calculate some sub-modules.

[0104] S740. Obtain all sub-modules that have not been preferentially calculated to get the modules to be allocated.

[0105] In steps S710 to S730, allocate each sub-module according to the computing resource allocation plan, and monitor the current load of each calculation group in real time. Since when multiple sub-modules are calculated within a calculation group, the calculation loads of each sub-module may fluctuate, for some calculation groups where the current load amount is close to the maximum calculated load, there may be a situation where the current load is greater than the maximum calculated load due to load fluctuations, which may damage the computing device or cause data loss. Therefore, compare the values of the current load and the maximum calculated load. If the current load is less than the maximum calculated load, no adjustment is required. If the current load is greater than or equal to the calculated load, the priorities of all sub-modules within this calculation group should be obtained, and the sub-modules with higher priorities are preferentially calculated, while the sub-modules with lower priorities stop calculating, so as to ensure that the current load of this calculation group is always less than the maximum calculated load.

[0106] In step S740, the modules to be allocated refer to the sub-modules that have not been allocated to the calculation group for calculation and are used for subsequent re-allocation. It should be noted that the modules to be allocated still perform calculations according to the originally divided tensor parallelism or pipeline parallelism.

[0107] By comparing the current load and the maximum calculated load, it can be ensured that the calculation group will not be overloaded, leading to equipment damage or data loss, improving the security of the computing device. By preferentially calculating some sub-modules, it is ensured that normal calculation can still be maintained in the case of insufficient computing resources, further guaranteeing the calculation efficiency and equipment safety. Marking the modules to be allocated enables the sub-modules that stop calculating to be re-allocated and calculated in a timely manner, improving the overall calculation efficiency.

[0108]

Eighth Embodiment

[0109] In a specific embodiment, if the load margin of the calculation group is greater than or equal to the calculation load amount of the module to be allocated, allocate the module to be allocated to this calculation group for assisting calculation, specifically including:

[0110] S810. Compare the load margins of each calculation group with the calculation load amounts of each module to be allocated.

[0111] S820. When the load margin is greater than or equal to the calculated load amount, allocate the module to be allocated to the calculation group for assisted calculation;

[0112] S830. If there is no load margin less than the calculated load amount, no assisted calculation is performed.

[0113] In steps S810 to S830, during the actual calculation process, some calculation groups may leave a relatively large load margin due to various reasons such as the low calculation difficulty of sub-modules, which is sufficient to calculate some modules to be allocated. Therefore, when the load margin is greater than or equal to the calculated load amount, the module to be allocated is allocated to the calculation group for assisted calculation. For example, if the module to be allocated A originally belongs to calculation group A, now the module to be allocated A is allocated to calculation group B for calculation. After the calculation of the module to be allocated A is completed, the calculation result is still sent to calculation group A for subsequent processing by calculation group A. If the load margin is not sufficient to support the calculation of the module to be allocated, no change is made and the module to be allocated is not calculated temporarily.

[0114] By comparing the load margin and the calculation load of the module to be allocated, it is determined whether to perform assisted calculation on the module to be allocated, optimizing the allocation logic of sub-modules, balancing the load among different calculation groups, and improving the utilization rate of calculation resources and the overall calculation efficiency.

[0115]

Ninth Embodiment

[0116] In a specific embodiment, the present invention further provides a resource allocation system 100 based on tensor parallelism. The resource allocation method described in the above embodiments is applied to this resource allocation system 100. The resource allocation system 100 includes: a storage module 110 for storing hardware information and sub-modules; a processing module 120 for splitting and obtaining sub-modules and dividing them into tensor modules and pipeline modules; a calculation module 130 for calculating the calculated load amount, communication efficiency, and load margin; and an allocation module 140 for executing the calculation resource allocation plan. This resource allocation system 100 has all the technical features of the above resource allocation method and will not be elaborated here one by one.

[0117] Although the present invention is disclosed as above, the present invention is not limited thereto. Any person skilled in the art can make various changes and modifications without departing from the spirit and scope of the present invention. Therefore, the protection scope of the present invention should be subject to the scope defined by the claims.

Claims

1. A resource allocation method based on tensor parallelism, characterized in that: The resource allocation method comprises: According to the remaining working life of the computing device, the computing device is divided into a plurality of computing groups, hardware information of the computing device is obtained, and the optimal computing load and the maximum computing load of the computing group are calculated according to the hardware information and the computing group; According to the data type and hierarchical structure information, the data to be calculated is divided into multiple sub-modules, the calculation method applicable to the sub-module is determined, the sub-module is divided into a tensor module or a pipeline module, and the priority is divided; Testing the submodules in sequence to obtain test data, and calculating the computing load of each submodule corresponding to each computing group according to the test data; Calculating the communication efficiency of data transmission between the computing groups according to the test data; Acquire the current load of each of the calculation groups, and calculate the load margin according to the optimal calculation load and the current load; Allocating the tensor module and the pipeline module to each of the computing groups according to the load margin, the computing load and the communication efficiency to obtain a computing resource allocation plan; Adjust the computing resource allocation plan according to the current load to obtain a module to be allocated; If the load margin of the computing group is greater than or equal to the computing load of the module to be allocated, the module to be allocated is allocated to the computing group for assisting computing.

2. The resource allocation method based on tensor parallelism according to claim 1, characterized in that: The step of dividing the computing devices into a plurality of computing groups according to the remaining working life of the computing devices, acquiring hardware information of the computing devices, and calculating the optimal computing load and the maximum computing load of the computing groups according to the hardware information and the computing groups specifically includes: Dividing the computing devices with similar remaining working lives into a plurality of computing groups; According to the hardware information, the load size that the computing group can withstand when running at the expected power is calculated to obtain the optimal computing load, and the load size that the computing group can withstand when running at full power is calculated to obtain the maximum computing load.

3. The resource allocation method based on tensor parallelism according to claim 2, characterized in that: The method of dividing the data to be calculated into multiple sub-modules according to the data type and the hierarchical structure information, determining the calculation method applicable to the sub-modules, dividing the sub-modules into tensor modules or pipeline modules, and dividing the priorities specifically includes: According to the data type and the hierarchical structure information, data of the same type and the same layer are merged, and divided into a plurality of the sub-modules according to the performance of the computing device; According to the data type, determine the computing method applicable to the submodule. If the submodule is more applicable to the tensor parallel computing method, mark it as a tensor module; if the submodule is more applicable to the pipeline parallel computing method, mark it as a pipeline module; According to the hierarchical structure information, the calculation subordinate structure of each of the submodules in the calculation process is determined to obtain the priority.

4. The resource allocation method based on tensor parallelism according to claim 3, characterized in that: The testing of the submodules in sequence to obtain test data, and calculating the computing load of each submodule corresponding to each computing group according to the test data, specifically includes: Sending all the submodules to each of the computing groups, testing the submodules in each of the computing groups in turn, acquiring the load changes and calculation results of the computing groups, and obtaining the test data; According to the test data, the maximum load value of the computing group within the test time is obtained to obtain the computing load amount.

5. The resource allocation method based on tensor parallelism according to claim 4, characterized in that: The calculating, according to the test data, the communication efficiency when each of the computing groups performs data transmission between each other specifically includes: Acquire the test data of each of the submodules, perform transmission test on each of the calculation results between all the calculation groups, acquire the change of transmission speed of the calculation results of each of the submodules between the calculation groups, and obtain communication data; The communication efficiency is obtained by calculating an average value of the transmission speed within the test time according to the communication data.

6. The resource allocation method based on tensor parallelism according to claim 5, characterized in that: The allocating the tensor module and the pipeline module to each of the computing groups according to the load margin, the computing load and the communication efficiency to obtain a computing resource allocation plan specifically includes: According to the load margin and the computational load, the tensor modules and the pipeline modules are allocated to the computational groups, and a scheme with the smallest total computational load is recorded as a first allocation scheme; Calculating the computing efficiency of the first allocation scheme according to the first allocation scheme and the communication efficiency to obtain a first computing efficiency; If the first computing efficiency is greater than or equal to the expected efficiency, planning is performed according to the first allocation scheme to obtain the computing resource allocation plan; If the first computing efficiency is less than the expected efficiency, adjusting the first allocation scheme according to the communication efficiency to obtain a second allocation scheme, and calculating the computing efficiency of the second allocation scheme to obtain a second computing efficiency; When the second computing efficiency is greater than or equal to the expected efficiency, planning is performed according to the second allocation scheme to obtain the computing resource allocation plan.

7. The resource allocation method based on tensor parallelism according to claim 6, characterized in that: The adjusting the computing resource allocation plan according to the current load to obtain the module to be allocated specifically includes: Allocate and calculate each of the submodules according to the computing resource allocation plan; If the current load is less than the maximum computing load, continue to allocate and calculate each of the submodules according to the computing resource classification plan; If the current load is greater than or equal to the computing load, the priorities of all the submodules in the computing group are obtained, and the submodules are reallocated according to the priorities and the maximum computing load, and some of the submodules are preferentially computed; All the submodules that have not been calculated preferentially are acquired to obtain the modules to be allocated.

8. The resource allocation method based on tensor parallelism according to claim 7, characterized in that: If the load margin of the computing group is greater than or equal to the computing load of the module to be allocated, the module to be allocated is allocated to the computing group for assisting computing, specifically including: Comparing the load margin of each of the computing groups with the computing load of each of the modules to be allocated; When the load margin is greater than or equal to the computing load, allocating the module to be allocated to the computing group for assisting computing; If the load margin that is less than the calculated load does not exist, no auxiliary calculation is performed.

9. A resource allocation system based on tensor parallelism, characterized in that: The resource allocation method according to any one of claims 1 to 8 is applied to the resource allocation system, and the resource allocation system comprises: A storage module, the storage module is used to store the hardware information and the sub-modules; A processing module, the processing module is used to split and obtain the sub-modules and divide them into the tensor module and the pipeline module; A calculation module, the calculation module is used to calculate the calculation load, the communication efficiency and the load margin; An allocation module, wherein the allocation module is used to execute the computing resource allocation plan.

Citation Information

Patent Citations

  • Data-in-data table-oriented computing resource allocation method and apparatus, and electronic device

    CN114741200A

  • Method for scheduling and allocating resources, and communication device

    WO2018099201A1