A resource allocation method and system based on pipeline parallelism
By dividing the data to be calculated in the power system into multiple data sets, parallel processing and communication efficiency correction, the problem of limited computing resources and high real-time performance in the power system is solved, and efficient computing resource allocation within the time period is achieved.
Patent Information
- Application Number
- CN202510397946.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-01
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-04-01
AI Technical Summary
How to reasonably allocate computing resources in the power system to ensure that computing efficiency is improved under limited resources and meet real-time requirements, especially when the power system has high computational complexity and high real-time requirements, how to process data in pipeline parallel processing to complete calculations within a time period.
The data to be calculated is divided into multiple data sets, the load amount of each data set is calculated based on the ideal calculation time, and the data sets that can be calculated within the ideal time are filtered for initial allocation. The unfinished data sets are processed in parallel through pipelines, the communication efficiency is calculated and corrected, and the allocation plan is adjusted to ensure that all data sets are completed within time.
By reasonably allocating computing tasks, the utilization rate of computing resources is improved, the allocation of computing resources is optimized, the calculation time of data sets that cannot be calculated in time is reduced, and the real-time requirements of the power system are ensured.
Smart Images

Figure CN119917286B_ABST
Abstract
Description
Technical Field
[0001] The technology of the present invention relates to the field of data processing. Specifically, it relates to a resource allocation method and system based on pipeline parallelism. Background Art
[0002] With the rapid development of information technology and the growth of global energy demand, the power system is facing problems of improving efficiency, reducing costs, and enhancing system reliability. To solve these problems, artificial intelligence technology has begun to be widely applied in the power grid field. More and more power systems are using deep learning models to predict future power demands, and then formulate corresponding power consumption strategies to optimize the allocation of power resources. However, due to the particularity of the power system, the accident distributions of different operating states or fault types are usually extremely unbalanced, which may affect the training effect and generalization ability of the deep learning model, and thus consume a large amount of computing resources and occupy a large amount of storage space. Further, due to the high real-time requirement and high computational complexity of the power system, when the computing resources are insufficient, the deep learning model often takes a long time to perform inference and is difficult to meet the real-time response requirements of the power system. Therefore, how to reasonably allocate computing resources according to the computing resource status of the power system and improve the computing efficiency of the data to be calculated as much as possible under limited computing resources is one of the problems that need to be urgently solved by those skilled in the art. Summary of the Invention
[0003] The problem solved by the present invention is: how to reasonably allocate the computing resources of the power grid system, perform pipeline parallel processing on some data, and ensure that the data calculation is completed within the time limit.
[0004] To solve the above problems, an embodiment of the present invention provides a resource allocation method based on pipeline parallelism. The resource allocation method includes: dividing the data to be calculated into multiple data sets to be calculated, and calculating the load of each data set to be calculated according to the ideal calculation time; testing the computing device to obtain the computing rate and maximum load of the computing device, and calculating the actual calculation time according to the computing rate and the data sets to be calculated; marking all the data sets to be calculated whose actual calculation time is less than or equal to the ideal calculation time as the original group, and allocating the data sets to be calculated within the original group to obtain an initial allocation plan; according to the initial allocation plan, obtaining all the unallocated data sets to be calculated to obtain a parallel data set, performing pipeline parallel processing on the parallel data set to obtain a pipeline module; obtaining the initial communication efficiency of the pipeline module according to the parallel mode of the pipeline module, and calculating the initial calculation time of the parallel data set according to the initial communication efficiency; correcting the initial communication efficiency according to the initial calculation time to obtain the theoretical calculation time of the parallel data set; judging whether the initial allocation plan needs to be modified according to the theoretical calculation time; when the initial allocation plan does not need to be modified, obtaining the final allocation plan through a parallel method, and obtaining the final allocation plan according to the theoretical calculation time and the initial allocation plan.
[0005] Compared with the prior art, the technical effects achieved by adopting this technical solution are as follows: By dividing the data to be calculated into multiple data sets to be calculated, the computing tasks are reasonably allocated, avoiding the excessive data processing difficulty caused by the large data volume and complex data types. By calculating the maximum load, it is ensured that the data sets to be calculated can be reasonably allocated to each of the computing devices later, improving the utilization rate of computing resources. Through pipeline parallel processing, the allocation of computing resources is optimized, reducing the calculation time of the data sets to be calculated that cannot be calculated in time. The acquisition and correction of the initial communication efficiency ensure that the theoretical calculation time can effectively parallel the actual calculation time of the data set, improving the rationality of subsequent computing resource allocation. By judging whether the initial allocation plan needs to be modified, the risk that some parallel data sets cannot be calculated within the ideal calculation time is avoided, effectively solving the problem of limited computing resources and high real-time requirements in the power system.
[0006] In an embodiment of the present invention, dividing the data to be calculated into multiple data sets to be calculated and calculating the load of each data set to be calculated according to the ideal calculation time specifically includes: obtaining all the data that needs to be calculated to obtain the data to be calculated; dividing the data to be calculated into multiple data sets to be calculated according to the data type and calculation requirements; setting the ideal calculation time for the data sets to be calculated according to the processing deadline of the data within the data sets to be calculated; calculating the load generated when the data sets to be calculated are processed in the computing device according to the data type and data volume of the data sets to be calculated to obtain the load.
[0007] Compared with the prior art, the technical effects achieved by adopting this technical solution are as follows: By setting the ideal calculation time, subsequent calculation tasks can be better allocated, ensuring that all data can be calculated within the processing deadline, improving the rationality of the system's calculation resource allocation. By calculating the load, data support can be provided for the allocation of subsequent datasets to be calculated, improving the utilization rate of calculation resources.
[0008] In an embodiment of the present invention, all datasets to be calculated with actual calculation time less than or equal to the ideal calculation time are recorded as the original group, and the datasets to be calculated within the original group are allocated to obtain an initial allocation plan, which specifically includes: screening out all calculation devices with actual calculation time less than or equal to the ideal calculation time, and corresponding calculation devices to obtain the original group; allocating the datasets to be calculated within the original group according to the above-mentioned load and the maximum load to obtain alternative plans; calculating the calculation time of each alternative plan, and selecting the alternative plan with the least calculation time as the initial allocation plan.
[0009] Compared with the prior art, the technical effects achieved by adopting this technical solution are as follows: By obtaining the original group, an allocation method that can be directly allocated and calculated within the time limit can be quickly screened out, improving the efficiency of calculation resource allocation. By obtaining the alternative plans and the initial allocation plan, an allocation plan with the minimum calculation time can be quickly screened out, improving the calculation efficiency of the datasets to be calculated and the rationality of calculation resource allocation.
[0010] In an embodiment of the present invention, according to the initial allocation plan, all unallocated calculation datasets are obtained to get a parallel dataset, and pipeline parallel processing is performed on the parallel dataset to obtain pipeline modules, which specifically includes: according to the initial allocation plan, obtaining all unallocated datasets to be calculated to get a parallel dataset; splitting the parallel dataset into multiple pipeline modules according to the data type and hierarchical structure information of the parallel dataset, calculating the load, and dividing priorities; according to the initial allocation plan, calculating the remaining load of each calculation device; allocating the pipeline modules according to the load, priority, and remaining load to obtain a parallel mode.
[0011] Compared with the prior art, the technical effects achieved by adopting this technical solution are as follows: By obtaining the priority, the task scheduling and resource allocation during the calculation process of the calculation device can be made more reasonable, ensuring that key data can be preferentially processed. The calculation of the remaining load provides data support for the normal calculation according to the parallel strategy of the pipeline module when the calculation resources are tight later, optimizing the allocation logic of the calculation resources.
[0012] In an embodiment of the present invention, the initial communication efficiency of the parallel data set is obtained according to the parallel mode of the pipeline module, and the initial calculation time of the parallel data set is calculated according to the initial communication efficiency. Specifically, it includes: calculating the average value of the transmission speed between computing devices according to the parallel mode to obtain the initial communication efficiency; calculating the actual calculation time of each pipeline module according to the calculation rate; calculating the initial calculation time of the parallel data set according to the initial communication efficiency and the actual calculation time.
[0013] Compared with the prior art, the technical effects achieved by adopting this technical solution are as follows: By calculating the initial communication efficiency, the specific situation of data transmission between each computing device can be obtained, avoiding the inconsistency between the predicted initial calculation time of the parallel data set and the actual situation due to too low communication efficiency, improving the accuracy of the initial calculation time calculation. By obtaining the initial calculation time, it can be determined whether the parallel data set can complete the calculation within the time limit, improving the rationality of computing resource allocation.
[0014] In an embodiment of the present invention, the initial communication efficiency is corrected according to the initial calculation time to obtain the theoretical calculation time of the parallel data set. Specifically, it includes: obtaining the change situation of the initial communication efficiency over time according to historical data to obtain the actual communication efficiency; calculating the theoretical calculation time according to the actual communication efficiency and the initial calculation time.
[0015] Compared with the prior art, the technical effects achieved by adopting this technical solution are as follows: By correcting the initial communication efficiency, the impact on computing resource allocation caused by the decrease in communication efficiency due to long-term data transmission can be avoided, ensuring that the calculation time of the parallel data set can be accurately estimated. Obtaining the theoretical calculation time provides data support for subsequent judgment on whether the initial allocation plan needs to be modified, improving the rationality of computing resource allocation.
[0016] In an embodiment of the present invention, when the initial allocation plan does not need to be modified, the final allocation plan is obtained through the parallel mode. When the initial allocation plan needs to be modified, according to the theoretical calculation time, partial data sets to be calculated within the original group are processed in pipeline parallel to obtain the final allocation plan. Specifically, it includes: if the initial allocation plan does not need to be modified, the final allocation plan is obtained according to the initial allocation plan and the parallel mode; if the initial allocation plan needs to be modified, calculate the difference between the theoretical calculation time and the ideal calculation time to obtain the overspending time; estimate the theoretical calculation time of each data set to be calculated within the original group if it is processed in pipeline parallel, and calculate the difference between the theoretical calculation time and the actual calculation time to obtain the surplus time; according to the overspending time and the surplus time, select partial data sets to be calculated for pipeline parallel processing and calculate the corresponding parallel mode; adjust the initial allocation plan according to the parallel mode to obtain the final allocation plan.
[0017] Compared with the prior art, the technical effects achieved by adopting this technical solution are as follows: By calculating the overrun time, it is possible to clarify exactly how much additional time the power grid system needs, which can provide a basis for selecting appropriate data sets to be calculated for pipeline parallel processing. By calculating the surplus time, it is possible to clarify the time benefits that can be obtained if each data set to be calculated undergoes pipeline parallel processing, improving the rationality of subsequent computing resource allocation.
[0018] In an embodiment of the present invention, there is also provided a resource allocation system based on pipeline parallelism. The resource allocation method based on pipeline parallelism described in the above embodiment is applied to this resource allocation system. The resource allocation system includes: a storage module for storing data sets to be calculated and a pipeline module; a processing module for allocating the data sets to be calculated and the pipeline module; a calculation module for calculating the actual calculation time, the initial communication efficiency, and the theoretical calculation time; and an allocation module for executing the final allocation plan. This resource allocation system has all the technical features of the above resource allocation method and will not be elaborated here one by one. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 It is one of the flowcharts of the resource allocation method based on pipeline parallelism of the present invention;
[0020] Figure 2 It is one of the flowcharts of the resource allocation method based on pipeline parallelism of the present invention;
[0021] Figure 3 It is one of the flowcharts of the resource allocation method based on pipeline parallelism of the present invention;
[0022] Figure 4 It is one of the flowcharts of the resource allocation method based on pipeline parallelism of the present invention;
[0023] Figure 5 It is a system schematic diagram of the resource allocation system based on pipeline parallelism of the present invention;
[0024] Description of the Reference Numerals:
[0025] 100 - Resource allocation system; 110 - Storage module; 120 - Processing module; 130 - Calculation module; 140 - Allocation module. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0026] To make the above objects, features, and advantages of the present invention more apparent and understandable, the following detailed description of the specific embodiments of the present invention will be given with reference to the accompanying drawings.
[0027]
First Embodiment
[0028] See Figure 1, in a specific embodiment, the present invention provides a resource allocation method based on pipeline parallelism. The resource allocation method includes:
[0029] S100: Divide the data to be calculated into multiple data sets to be calculated, and calculate the load of each data set to be calculated according to the ideal calculation time;
[0030] S200: Test the computing device to obtain the computing rate and maximum load of the computing device, and calculate the actual computing time according to the computing rate and the data sets to be calculated;
[0031] S300: Denote all the data sets to be calculated whose actual computing time is less than or equal to the ideal computing time as the original group, and allocate the data sets to be calculated within the original group to obtain an initial allocation plan;
[0032] S400: According to the initial allocation plan, obtain all the unallocated data sets to be calculated to get a parallel data set, and perform pipeline parallel processing on the parallel data set to obtain a pipeline module;
[0033] S500: Obtain the initial communication efficiency of the pipeline module according to the parallel mode of the pipeline module, and calculate the initial computing time of the parallel data set according to the initial communication efficiency;
[0034] S600: Correct the initial communication efficiency according to the initial computing time to obtain the theoretical computing time of the parallel data set;
[0035] S700: Obtain the final allocation plan according to the theoretical computing time and the initial allocation plan.
[0036] In step S100, in the power grid system, since the computing resources are limited and the real-time requirement is high, the types of data to be calculated are complex and the data volume is large, including various power system-related data such as measurement data, power grid operation data, and distribution network data. If the data to be calculated is directly calculated, it will consume a large amount of computing resources, the computing efficiency is low, and it will cause a large amount of computing load on the computing device. Therefore, in order to improve the computing efficiency, the data to be calculated is divided into multiple data sets, that is, data sets to be calculated, for reasonable allocation of computing tasks.
[0037] It should be noted that the load of a single data set to be calculated should be less than the maximum load of a single designed computing device.
[0038] The ideal computing time refers to the remaining time when the dataset to be computed can be processed in a timely manner according to the requirements of the power grid system. For example, in the power grid system, if the dataset to be computed obtained at 8:00 am on November 4, 2024 includes meteorological data for meteorological prediction on November 5, 2024, the ideal computing time for this dataset to be computed can be set to 10h or 12h. The specific setting of the ideal computing time can be adjusted according to actual requirements.
[0039] The load refers to the computing device load generated when the dataset to be computed is processed within the ideal computing time. It should be noted that due to the aging, replacement, and expansion of computing devices, there are differences in the computing performance of each computing device. Therefore, for the same dataset to be computed, the loads generated when processed on different computing devices may also be different and need to be calculated separately.
[0040] In step S200, usually, all computing devices should be of the same model. If there are computing devices of different models, the computing power needs to be converted according to the hardware of the computing devices. By testing the computing devices, calculate the average computing speed of the computing devices within the test time to obtain the computing rate. By calculating the average value of the load size when the computing devices are running at full power within the test time, obtain the maximum load. It should be noted that the test duration for calculating the computing rate and the maximum load should be the same.
[0041] In step S300, since pipelining parallelism requires multiple computing devices to perform collaborative computing and has high requirements for communication efficiency, if all datasets to be computed are processed using pipelining parallelism, it will greatly increase the communication burden between computing devices, possibly resulting in the inability to continue communication and affecting the processing efficiency of the datasets to be computed, wasting a large amount of computing resources. Therefore, for datasets to be computed that can be completed within the ideal computing time even without using pipelining parallelism, they are directly allocated and computed to obtain the initial allocation plan.
[0042] In step S400, since only some of the datasets to be computed are allocated in the initial allocation plan, the remaining unallocated computing datasets cannot be processed in a timely manner because the actual computing time is greater than the ideal computing time. Therefore, pipelining parallelism is used to shorten the computing time. The parallel dataset refers to the dataset to be computed that needs to be processed using pipelining parallelism. The pipeline module refers to each module obtained after splitting the parallel dataset through pipelining parallel processing. It should be noted that the ideal computing time of the parallel dataset is the same as that of its corresponding dataset to be computed.
[0043] In step S500, in pipeline parallelism, pipeline modules may be assigned to different computing devices for calculation. Therefore, the communication efficiency between computing devices affects the calculation efficiency of the parallel dataset. Due to the performance differences of computing devices and the data type differences within the parallel dataset, the communication efficiency between different computing devices is different, and the communication efficiency between different pipeline modules on the same computing devices is also different, which need to be calculated separately. The initial communication efficiency is the communication efficiency when the data of the pipeline module is transmitted between computing devices under normal circumstances. For example, for pipeline module A, if the calculation result is obtained in computing device A and transmitted to computing device B, the initial communication efficiency may be 90%. If the calculation result is obtained in computing device B and transmitted to computing device C, the initial communication efficiency may be 75%. For pipeline module B, if the calculation result is obtained in computing device A and transmitted to computing device B, the communication efficiency may be 63%.
[0044] The initial calculation time refers to the time required for the parallel dataset to be calculated through pipeline parallelism considering the impact of the initial communication efficiency.
[0045] In step S600, in the actual use of computing devices, over time, computing devices may be affected by factors such as temperature, making it difficult to maintain the initial communication efficiency for a long time. Therefore, for parallel datasets with a long initial calculation time, the communication efficiency may continuously decrease over time during the calculation process, resulting in the initial calculation time being unable to reflect the true calculation time of the parallel dataset. To ensure an accurate judgment of whether the parallel dataset can be calculated within the ideal calculation time, it is necessary to correct the initial communication efficiency based on the initial calculation time and recalculate the time required for the parallel dataset calculation to obtain the theoretical calculation time.
[0046] In step S700, if the theoretical calculation time of at least one parallel dataset is greater than the ideal calculation time, the initial assignment calculation needs to be modified. If the theoretical calculation time of all parallel datasets is less than or equal to the ideal calculation time, there is no need to modify the initial assignment plan.
[0047] If the initial assignment plan does not need to be modified, it means that all datasets to be calculated can be completed within the ideal calculation time. Therefore, combining the initial assignment plan and the parallel method, clarify the processing and assignment methods of each dataset to be calculated to obtain the final assignment plan.
[0048] If the theoretical calculation time of at least one parallel data set is greater than the ideal calculation time, it means that there is a parallel data set that cannot complete the calculation within the time limit. Therefore, it is necessary to perform pipeline parallel processing on some of the data sets to be calculated in the initial allocation plan to reduce the calculation time, so that the parallel data sets that cannot complete the calculation in time can start processing in advance, ensuring that the calculation can be completed within the time limit. For example, according to the initial allocation plan, the data set A to be calculated is calculated on device A and is expected to be completed at 3 pm. The parallel data set A is calculated after the data set A to be calculated. The theoretical calculation time is 3 hours, the ideal calculation time is 2 hours, and it is expected to be completed at 6 pm. It needs to be completed before 5 pm. It is judged that the calculation cannot be completed in time. Therefore, pipeline parallel processing is performed on the data set A to be calculated, reducing the calculation time and completing the calculation at 2 pm in advance. Further, the parallel data set A can start calculating at 2 pm in advance and is expected to be completed at 5 pm, so that the calculation can be completed in time.
[0049] By dividing the data to be calculated into multiple data sets to be calculated and reasonably allocating the calculation tasks, the data processing difficulty caused by excessive data volume and complex data types is reduced. By calculating the maximum load, it is ensured that the data sets to be calculated can be reasonably allocated to each calculation device in the subsequent work, improving the utilization rate of computing resources. Through pipeline parallel processing, the allocation of computing resources is optimized, and the calculation time of the data sets to be calculated that cannot be calculated in time is reduced. The acquisition and correction of the initial communication efficiency ensure that the theoretical calculation time can effectively parallel the actual calculation time of the data set, improving the rationality of the subsequent computing resource allocation. By judging whether the initial allocation plan needs to be modified, the risk that some parallel data sets cannot complete the calculation within the ideal calculation time is avoided, effectively solving the problem of limited computing resources and high real-time requirements in the power system.
[0050]
Second Embodiment
[0051] See Figure 2 , in a specific embodiment, the data to be calculated is divided into multiple data sets to be calculated, and the load of each data set to be calculated is calculated according to the ideal calculation time, which specifically includes:
[0052] S110. Obtain all the data that needs to be calculated to obtain the data to be calculated;
[0053] S120. Divide the data to be calculated into multiple data sets to be calculated according to the data type and calculation requirements;
[0054] S130. Set the ideal calculation time for the data set to be calculated according to the processing deadline of the data in the data set to be calculated;
[0055] S140. Calculate the load generated when the data set to be calculated is processed within the computing device according to the data type and data volume of the data set to be calculated, and obtain the load amount.
[0056] In steps S110 to S120, due to the complexity and real-time nature of the power grid system, the data types, data acquisition methods, and data acquisition times of different system structures are different. Therefore, the data to be calculated needs to be acquired in real time and periodically sorted according to actual requirements. For example, for a certain power grid system, the distribution data is updated every five minutes, and the meteorological data is updated every fifteen minutes. However, due to their different calculation requirements, the distribution data is sorted every half hour to obtain the data to be calculated, and the meteorological data is sorted every 1h to obtain the data to be calculated.
[0057] The calculation requirement refers to the necessary calculation conditions that need to be met when processing the data to be calculated. According to the calculation, the data to be calculated is divided into multiple data sets to be calculated. For example, in a certain power grid system, if it is necessary to predict the weather conditions in area A on September 25th based on the meteorological data in areas A and B on September 24th, then all the meteorological data in areas A and B on September 24th need to be divided into the same data set to be calculated for subsequent calculations.
[0058] In step S130, since in the actual calculation process, it is impossible to ensure that all data sets to be calculated are completed within the expected time, when setting the ideal calculation time, some margins should be set according to the processing deadline to avoid being unable to meet the processing deadline due to accidents. For example, in the power grid system, a data set to be calculated obtained at 8:00 am on November 4, 2024, includes meteorological data for meteorological prediction on November 5, 2024. This data set to be calculated has 16h available for calculation, but in order to avoid an increase in calculation time due to unexpected situations, the ideal calculation is set to 12h.
[0059] In step S140, the load amount includes parameters such as CPU usage rate, memory occupancy rate, and disk I / O rate. Since there are differences in the data types of the data sets to be calculated, and different computing devices also have different capabilities in processing different data types, therefore, when different computing devices calculate the same data set to be calculated, the load amounts generated are different. For example, for a certain data set to be calculated A, which occupies more CPU for processing during the calculation process, computing device A has better memory performance and poorer CPU performance, while computing device B has poorer memory performance and better CPU performance. Then the load amount of data set to be calculated A when calculated on computing device A is greater than that on computing device B.
[0060] By setting the ideal computing time, subsequent allocation of computing tasks can be better carried out, ensuring that all data can complete the calculation within the processing deadline, improving the rationality of the system's computing resource allocation. By calculating the load, data support can be provided for the subsequent allocation of the dataset to be calculated, improving the utilization rate of computing resources.
[0061]
Third Embodiment
[0062] In a specific embodiment, all datasets to be calculated with actual computing time less than or equal to the ideal computing time are recorded as the original group, and the datasets to be calculated within the original group are allocated to obtain an initial allocation plan, which specifically includes:
[0063] S310. Screen out all datasets to be calculated with actual computing time less than or equal to the ideal computing time, and obtain the corresponding computing devices to get the original group;
[0064] S320. Allocate the datasets to be calculated within the original group according to the above-mentioned load and the maximum load to obtain alternative plans;
[0065] S330. Calculate the computing time of each alternative plan, and select the alternative plan with the least computing time as the initial allocation plan.
[0066] In step S310, since the time required for calculating the datasets to be calculated is different in different computing devices, it is necessary to screen out the datasets to be calculated with actual computing time less than or equal to the ideal computing time and the corresponding computing devices, which are recorded as the original group. For example, for the dataset to be calculated A, the ideal computing time is 30 minutes. When calculating on computing device A, the actual computing time is 1 hour, and when calculating on computing device B, the actual computing time is 20 minutes. Among them, when calculating on computing device B, the actual computing time of the dataset to be calculated A is less than or equal to the ideal computing time. Therefore, the computing device A to be calculated and the computing device B are recorded as the original group.
[0067] In step S320, obtain all the datasets to be calculated included in the original group. For the same dataset to be calculated, preferentially allocate the corresponding dataset to be calculated according to the original group with the least actual computing time. During the allocation process, for the same computing device, the total load of the datasets to be calculated allocated to it shall not be greater than the maximum load of the computing device. Further, if for a certain dataset to be calculated, the computing device corresponding to the original group with the least actual computing time does not have enough remaining load for the calculation of this dataset to be calculated, then allocate it according to the original group with the least computing time among the remaining original groups of this dataset to be calculated, and so on. If this dataset to be calculated cannot be allocated to any corresponding computing device according to the original group, then this dataset to be calculated is not allocated.
[0068] For example, for the dataset A to be calculated, there are two corresponding original groups. The original group A is calculated on the computing device A, and the actual calculation time is 30 minutes. The original group B is calculated on the computing device B, and the actual calculation time is 45 minutes. Then, the dataset A to be calculated is preferentially allocated to the computing device A. If the remaining load of the computing device A is less than the load of the dataset A to be calculated, the dataset A to be calculated is allocated to the computing device B. If the remaining load of the computing device B is also less than the load of the dataset A to be calculated, the dataset A to be calculated is not allocated.
[0069] In step S330, since there are multiple data groups to be calculated, there may be multiple candidate solutions due to differences in the allocation order. Calculate the calculation time of each candidate solution, and select the candidate solution with the shortest calculation time as the initial allocation plan. The calculation time refers to the time required to complete all the datasets to be calculated that are allocated.
[0070] For example, if the candidate solution A allocates the dataset A to be calculated to the computing device A and the dataset B to be calculated to the computing device B, and the calculation time is 3 hours, and the candidate solution B allocates the dataset A to be calculated to the computing device B and the dataset B to be calculated to the computing device A, and the calculation time is 2.5 hours, then select the candidate solution B as the initial allocation plan.
[0071] By obtaining the original groups, an allocation method that can be directly allocated and complete the calculation within the time limit can be quickly screened out, improving the efficiency of computing resource allocation. By obtaining the candidate solutions and the initial allocation plan, an allocation plan with the minimum calculation time can be quickly screened out, enhancing the calculation efficiency of the data to be calculated and the rationality of computing resource allocation.
[0072]
Fourth Embodiment
[0073] See Figure 3 , in a specific embodiment, according to the initial allocation plan, obtain all the unallocated computing datasets to obtain a parallel dataset, and perform pipeline parallel processing on the parallel dataset to obtain pipeline modules, specifically including:
[0074] S410. According to the initial allocation plan, obtain all the unallocated datasets to be calculated to obtain a parallel dataset;
[0075] S420. According to the data type and hierarchical structure information of the parallel dataset, split the parallel dataset into multiple pipeline modules, calculate the load, and divide the priorities;
[0076] S430. According to the initial allocation plan, calculate the remaining load of each computing device;
[0077] S440. Allocate the pipeline modules according to the load amount, priority, and remaining load amount to obtain the parallel mode.
[0078] In step S410, in the power grid system, due to the performance limitations of computing devices, there may be some datasets to be calculated that cannot be processed within the time limit through direct calculation. Therefore, the initial allocation plan may not allocate all the datasets to be calculated. The remaining unallocated datasets to be calculated need to be processed in parallel through the pipeline in order to complete the calculation within the ideal computing time. Therefore, these datasets to be calculated are marked as parallel datasets.
[0079] In step S420, usually, the data types of the parallel datasets are complex and the data volume is large. If these data are directly calculated, it will cause too much load on the computing devices. Therefore, the data of the same type or related in the parallel datasets are integrated and processed, and split into multiple pipeline modules.
[0080] It should be noted that the data volume of the pipeline module is set according to the performance of the computing device. If the performance of the computing device is relatively good, the data volume of the pipeline module can be set relatively large. If the performance of the computing device is relatively poor, the data volume of the pipeline module can be set relatively small. The specific standard depends on the actual situation, and there may be significant differences in the data volume of each pipeline module.
[0081] Since pipeline parallelism distributes different pipeline modules of data to different computing devices to form a pipeline for parallel computing, in actual calculations, there may be a subordinate relationship between pipeline modules. For example, pipeline module B needs to calculate based on the calculation result of pipeline module A. Therefore, priorities should be set for each sub-module according to the subordinate relationship between each sub-module to ensure that when there are too many sub-modules and the computing device cannot calculate all sub-modules simultaneously, the more basic sub-modules can be calculated first to ensure that the data can be processed normally.
[0082] In step S430, the remaining load amount refers to the remaining load amount of the computing device that can be used for processing the datasets to be calculated. According to the initial allocation plan, the load amount of the datasets to be calculated allocated to each computing device is known, and the difference between the total load amount of the datasets to be calculated and the maximum load amount of the corresponding computing device is calculated to obtain the remaining load amount.
[0083] It should be noted that the remaining load amount does not refer to a single performance parameter of the computing device, but includes parameters such as CPU usage rate, memory occupancy rate, and disk I / O rate.
[0084] In step S440, the parallel mode refers to the specific allocation plan for all pipeline modules, which needs to ensure that according to the remaining load of the computing device and the load of the pipeline modules, the total load of the computing device does not exceed its maximum load. Further, in order to ensure the computing efficiency of pipeline modules with a subordinate relationship, allocation should be prioritized for pipeline modules with a higher priority according to the priority, and allocation for pipeline modules with a lower priority should be postponed. For example, if a parallel data set is divided into two pipeline modules, namely pipeline module A and pipeline module B, and the calculation of pipeline module B requires the calculation result data support of pipeline module A, pipeline module A should be allocated first to ensure that pipeline module B can be calculated in a timely manner according to the parallel mode.
[0085] By obtaining the priority, the task scheduling and resource allocation during the calculation process of the computing device can be made more reasonable, ensuring that key data can be processed first. The calculation of the remaining load provides data support for the normal calculation according to the parallel strategy of the pipeline modules when the subsequent computing resources are tense, and optimizes the allocation logic of the computing resources.
[0086]
Fifth Embodiment
[0087] In a specific embodiment, the initial communication efficiency of the parallel data set is obtained according to the parallel mode of the pipeline module, and the initial calculation time of the parallel data set is calculated according to the initial communication efficiency, specifically including:
[0088] S510. Calculate the average value of the transmission speeds between computing devices according to the parallel mode to obtain the initial communication efficiency;
[0089] S520. Calculate the actual calculation time of each pipeline module according to the calculation rate;
[0090] S530. Calculate the initial calculation time of the parallel data set according to the initial communication efficiency and the actual calculation time.
[0091] In step S510, the initial communication efficiency refers to the proportion of the actual data transmission volume per unit time between computing groups in the expected data transmission volume, and the calculation formula is:
[0092] H c =C e / C a ×100%, H c is the initial communication efficiency, C e is the actual transmission rate, C a is the expected transmission rate.
[0093] For example, if the expected transmission rate is 30 gbps and the actual transmission rate is 27 gbps, then the initial communication efficiency Hc =C e / C a ×100% = (27 gbps / 30 gbps) × 100% = 90%.
[0094] In steps S520 to S530, the calculation method of the actual calculation time of the pipeline module is the same as that of the dataset to be calculated. Since the time for the pipeline module to complete the calculation may be affected by the communication efficiency, in order to ensure the accuracy of the initial calculation time of the parallel dataset, when calculating the initial calculation time of the parallel dataset through the actual calculation time of the pipeline module, the impact of communication efficiency on the calculation efficiency needs to be considered. For example, a certain parallel dataset is divided into three pipeline modules: pipeline module A, pipeline module B, and pipeline module C. Among them, pipeline module A can directly perform calculations, and the actual calculation time is 20 mins. The calculation of pipeline module B requires the calculation result of pipeline module A as data support, the initial communication efficiency is 85%, and the actual calculation time is 30 mins. The calculation of pipeline module C requires the calculation result of pipeline module B as data support, the initial communication efficiency is 90%, and the actual calculation time is 40 mins. Then, the initial calculation time of this parallel dataset can be calculated as = 20 mins + 30 mins ÷ 85% + 40 mins ÷ 90% ≈ 99.74 mins.
[0095] By calculating the initial communication efficiency, the specific situation of data transmission between each computing device can be obtained, avoiding the prediction of the initial calculation time of the parallel dataset not matching the actual situation due to too low communication efficiency, improving the accuracy of the initial calculation time calculation. By obtaining the initial calculation time, it can be determined whether the parallel dataset can complete the calculation within the time limit, improving the rationality of the calculation resource allocation.
[0096]
Sixth Embodiment
[0097] In a specific embodiment, the initial communication efficiency is corrected according to the initial calculation time to obtain the theoretical calculation time of the parallel dataset, which specifically includes:
[0098] S610. Obtain the change situation of the initial communication efficiency over time according to historical data to obtain the actual communication efficiency;
[0099] S620. Calculate the theoretical calculation time according to the actual communication efficiency and the initial calculation time.
[0100] In steps S610 to S620, as the transmission time increases, the computing device may experience a problem of reduced communication efficiency. Therefore, it is necessary to correct the initial communication efficiency according to the change of communication efficiency, recalculate the required time, and obtain the theoretical calculation time. For example, when the technical result of pipeline module A needs to be transmitted from computing device A to computing device B, the initial calculation time is 2h. The initial communication efficiency is maintained for the first 15 minutes, and the communication efficiency is 95%. The communication efficiency drops to 80% of the initial communication efficiency from 15 minutes to 40 minutes, and the communication efficiency has been maintained at 75% of the initial communication efficiency after 40 minutes. Then, the theoretical calculation time can be calculated as follows: theoretical calculation time = 15 minutes + 25 minutes ÷ 80% + 80 minutes ÷ 75% ≈ 152.92 minutes.
[0101] By correcting the initial communication efficiency, it is possible to avoid the impact of the decrease in communication efficiency caused by long-term data transmission on the calculation resource allocation, ensure the accurate estimation of the calculation time of the parallel data set, obtain the theoretical calculation time, provide data support for subsequent judgment on whether to modify the initial allocation plan, and improve the rationality of the calculation resource allocation.
[0102]
Seventh Embodiment
[0103] See Figure 4 , in a specific embodiment, the final allocation plan is obtained according to the theoretical calculation time and the initial allocation plan, which specifically includes:
[0104] S710. If the initial allocation plan does not need to be modified, the final allocation plan is obtained according to the initial allocation plan and the parallel mode;
[0105] S720. If the initial allocation plan needs to be modified, calculate the difference between the theoretical calculation time and the ideal calculation time to obtain the overspending time;
[0106] S730. Estimate the theoretical calculation time of each data set to be calculated in the original group if pipeline parallel processing is performed, and calculate the difference between the theoretical calculation time and the actual calculation time to obtain the surplus time;
[0107] S740. According to the overspending time and the surplus time, select some data sets to be calculated for pipeline parallel processing, and calculate the corresponding parallel mode;
[0108] S750. Adjust the initial allocation plan according to the parallel mode to obtain the final allocation plan.
[0109] In step S710, if the initial allocation plan does not need to be modified, it means that all datasets to be calculated can be calculated within the ideal calculation time. Therefore, the datasets to be calculated can be directly allocated and calculated according to the initial allocation plan and the parallel mode.
[0110] In steps S720 to S730, the overrun time refers to how much extra time the computing device needs to exceed the ideal calculation time to complete the calculation of all datasets to be calculated. The surplus time refers to how much time can be advanced to complete the calculation of the dataset to be calculated relative to the actual calculation time if the dataset to be calculated is calculated using the pipeline parallel mode. If the initial allocation plan needs to be modified, it means that there are datasets to be calculated that cannot be calculated within the ideal calculation time.
[0111] In steps S740 to S750, due to the performance limitations of the computing device, if there are datasets to be calculated that are calculated using the pipeline parallel mode, it may lead to excessive communication pressure between computing devices, resulting in too low communication efficiency and affecting the overall computing efficiency. At the same time, too many pipeline modules will increase the difficulty of the system in allocating computing tasks. Therefore, the sum of the surplus times of the selected datasets to be calculated should be as close as possible to the overrun time. Further, in order to avoid fluctuations in the calculation time caused by other factors, a certain margin should be set for the surplus time. For example, when a power grid system allocates a batch of datasets to be calculated, the overrun time is 5h. If the datasets to be calculated module A, datasets to be calculated module B, and datasets to be calculated module C are processed in pipeline parallel, the sum of the surplus times is 5h. If the datasets to be calculated module B, datasets to be calculated module C, and datasets to be calculated module D are processed in pipeline parallel, the sum of the surplus times is 6h. Then, the datasets to be calculated module B, datasets to be calculated module C, and datasets to be calculated module D should be selected for pipeline parallel processing. The specific reserved amount can be adjusted according to actual requirements and actual situations.
[0112] By calculating the overrun time, it is possible to clarify how much additional time the power grid system specifically needs, which can provide a basis for selecting appropriate datasets to be calculated for pipeline parallel processing. By calculating the surplus time, it is possible to clarify the time benefits that can be obtained if each dataset to be calculated is processed in pipeline parallel, improving the rationality of subsequent computing resource allocation.
[0113]
Eighth Embodiment
[0114] See Figure 5, in a specific embodiment, the present invention further provides a resource allocation system 100 based on pipeline parallelism. The resource allocation method described in the above embodiment is applied to this resource allocation system 100. The resource allocation system 100 includes: a storage module 110, which is used to store the data set to be calculated and the pipeline module; a processing module 120, which is used to allocate the data set to be calculated and the pipeline module; a calculation module 130, which is used to calculate the actual calculation time, the initial communication efficiency, and the theoretical calculation time; an allocation module 140, which is used to execute the final allocation plan. This resource allocation system 100 has all the technical features of the above resource allocation method, and will not be elaborated here one by one.
[0115] Although the present invention is disclosed as above, the present invention is not limited thereto. Any person skilled in the art can make various changes and modifications without departing from the spirit and scope of the present invention. Therefore, the protection scope of the present invention should be subject to the scope defined by the claims.
Claims
1. A resource allocation method based on pipeline parallelism, characterized in that, The resource allocation method includes: Dividing the data to be calculated into multiple data sets to be calculated, and calculating the load of each data set to be calculated according to the ideal calculation time; Testing the computing device to obtain the computing rate and maximum load of the computing device, and calculating the actual calculation time according to the computing rate and the data set to be calculated; Denoting all the data sets to be calculated whose actual calculation time is less than or equal to the ideal calculation time as the original group, and allocating the data sets to be calculated within the original group to obtain an initial allocation plan; According to the initial allocation plan, obtaining all the unallocated data sets to be calculated to obtain a parallel data set, and performing pipeline parallel processing on the parallel data set to obtain a pipeline module; Obtaining the initial communication efficiency of the pipeline module according to the parallel mode of the pipeline module, and calculating the initial calculation time of the parallel data set according to the initial communication efficiency; Correcting the initial communication efficiency according to the initial calculation time to obtain the theoretical calculation time of the parallel data set; Obtaining a final allocation plan according to the theoretical calculation time and the initial allocation plan.
2. The resource allocation method based on pipeline parallelism according to claim 1, wherein, The step of dividing the data to be calculated into multiple data sets to be calculated and calculating the load of each data set to be calculated according to the ideal calculation time specifically includes: Obtaining all the data that needs to be calculated to obtain the data to be calculated; Dividing the data to be calculated into multiple data sets to be calculated according to the data type and calculation requirements; Setting an ideal calculation time for the data set to be calculated according to the processing deadline of the data in the data set to be calculated; Calculating the load generated when the data set to be calculated is processed in the computing device according to the data type and data volume of the data set to be calculated to obtain the load.
3. The resource allocation method based on pipeline parallelism according to claim 2, wherein The step of denoting all the data sets to be calculated whose actual calculation time is less than or equal to the ideal calculation time as the original group, and allocating the data sets to be calculated within the original group to obtain an initial allocation plan specifically includes: Screening out all the data sets to be calculated whose actual calculation time is less than or equal to the ideal calculation time, and obtaining the corresponding computing device to obtain the original group; Allocating the data sets to be calculated within the original group according to the load and the maximum load to obtain a candidate plan; Calculating the calculation time of each candidate plan, and selecting the candidate plan with the shortest calculation time as the initial allocation plan.
4. The resource allocation method based on pipeline parallelism according to claim 3, wherein The step of obtaining all the unallocated data sets to be calculated according to the initial allocation plan to obtain a parallel data set, and performing pipeline parallel processing on the parallel data set to obtain a pipeline module specifically includes: Obtaining all the unallocated data sets to be calculated according to the initial allocation plan to obtain a parallel data set; Splitting the parallel data set into multiple pipeline modules according to the data type and hierarchical structure information of the parallel data set, calculating the load, and dividing the priorities; Calculating the remaining load of each computing device according to the initial allocation plan; Allocate the pipeline modules according to the load, the priority, and the remaining load to obtain the parallel mode.
5. The resource allocation method based on pipeline parallelism according to claim 4, wherein Obtain the initial communication efficiency of the pipeline module according to the parallel mode of the pipeline module, and calculate the initial calculation time of the parallel data set according to the initial communication efficiency. Specifically, it includes: Calculate the average value of the transmission speeds between the computing devices according to the parallel mode to obtain the initial communication efficiency; Calculate the actual calculation time of each pipeline module according to the calculation rate; Calculate the initial calculation time of the parallel data set according to the initial communication efficiency and the actual calculation time.
6. The resource allocation method based on pipeline parallelism according to claim 5, wherein Modify the initial communication efficiency according to the initial calculation time to obtain the theoretical calculation time of the parallel data set. Specifically, it includes: Obtain the change of the initial communication efficiency over time according to historical data to obtain the actual communication efficiency; Calculate the theoretical calculation time according to the actual communication efficiency and the initial calculation time.
7. The resource allocation method based on pipeline parallelism according to claim 6, wherein Obtain the final allocation plan according to the theoretical calculation time and the initial allocation plan. Specifically, it includes: If the initial allocation plan does not need to be modified, obtain the final allocation plan according to the initial allocation plan and the parallel mode; If the initial allocation plan needs to be modified, calculate the difference between the theoretical calculation time and the ideal calculation time to obtain the overrun time; Estimate the theoretical calculation time of each data set to be calculated in the original group if pipeline parallel processing is performed, and calculate the difference between the theoretical calculation time and the actual calculation time to obtain the surplus time; Select some of the data sets to be calculated for pipeline parallel processing according to the overrun time and the surplus time, and calculate the corresponding parallel mode; Adjust the initial allocation plan according to the parallel mode to obtain the final allocation plan.
8. A resource allocation system based on pipeline parallelism, characterized in that, The resource allocation method according to any one of claims 1 to 7 is applied to the resource allocation system, and the resource allocation system includes: A storage module for storing the data sets to be calculated and the pipeline modules; A processing module for allocating the data sets to be calculated and the pipeline modules; A calculation module for calculating the actual calculation time, the initial communication efficiency, and the theoretical calculation time; An allocation module for executing the final allocation plan.
Citation Information
Patent Citations
Calibration of resource allocation during parallel processing
CN102043673A
Model real-time monitoring and distribution method based on pipeline parallel training
CN118277209A