Computing resource allocation method and system based on data dimension conversion
By using data dimension conversion technology in distributed training, computing tasks are allocated to hardware devices, the problem of uneven allocation of computing tasks is solved, computing performance and efficiency are improved, and more efficient resource utilization and computing stability are achieved.
Patent Information
- Application Number
- CN202510638569.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-19
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2045-05-19
AI Technical Summary
In distributed training, the problem of how to reasonably allocate computing tasks to different hardware devices through tensor parallelism and improve the utilization rate of hardware devices has become a key issue that needs to be solved urgently.
By obtaining the minimum alignment granularity of all computing devices in the work library, the current computed data of different dimensions is converted into one-dimensional tensors, and the memory space in the computing device is dynamically divided according to the memory usage and minimum alignment granularity of each tensor data block in the one-dimensional tensor to obtain the predetermined storage layout of the one-dimensional tensor. The fixed calculation data is planned according to the predetermined storage layout, the pre-allocation information of the newly added calculation data is determined, and the pre-allocation information is adjusted according to the theoretical load and the average access hit rate of the tensor data block.
Effectively organize and manage various computing resources, make resource scheduling and allocation more efficient and intelligent, reduce access delays, optimize device usage, avoid resource waste, ensure the stability and continuity of computing equipment, improve the efficiency of memory usage of computing equipment, avoid the waste of memory resources, and ensure the efficient progress of computing tasks.
Smart Images

Figure CN120179415A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and in particular, to a computing resource allocation method and system based on data dimension conversion. Background Art
[0002] With the rapid development of artificial intelligence technology, the parameter scale of deep learning models has been increasing day by day, and the number of model parameters ranges from millions to trillions. Single-device training can no longer meet the requirements, and distributed training has become an inevitable choice. Tensor parallelism and pipeline parallelism are the core strategies of distributed training. Tensor parallelism performs parallel computing in a single operation. By splitting and distributing large-scale tensors to multiple computing units, computing acceleration can be achieved, such as matrix-matrix multiplication. Therefore, from another perspective, tensor parallelism can be regarded as intra-layer parallelism.
[0003] Currently, significant progress has been made in distributed training through tensor parallelism technology, but there are still some challenges. First, when partitioning tensors for a large amount of data, the fixed alignment granularity easily leads to a mismatch between tensor data blocks and the memory pool, resulting in memory fragmentation and resource waste. Second, the task allocation strategy overly relies on the static logical structure of tensor data while ignoring the hardware load and cache efficiency, which may cause uneven loads on execution units. In addition, the fixed data layout may not be able to dynamically adapt to changes in hardware states, leading to access conflicts and bandwidth waste.
[0004] Therefore, in distributed training, the problem of how to reasonably allocate computing tasks to different hardware devices through tensor parallelism and improve the utilization rate of hardware devices has become a key problem to be solved urgently. Summary of the Invention
[0005] The problem solved by the present invention is: how to reasonably allocate computing tasks to different hardware devices through data dimension conversion, thereby improving the overall computing performance and efficiency.
[0006] To solve the above problems, an embodiment of the present invention provides a computing resource allocation method based on data dimension conversion. The computing resource allocation method includes: obtaining the minimum alignment granularity of all computing devices in the working library, and converting the current computing data in different dimensions into a one-dimensional tensor according to the minimum alignment granularity; dynamically partitioning the memory space in the computing device according to the memory occupancy of each tensor data block in the one-dimensional tensor and the minimum alignment granularity to obtain a predetermined storage layout of the one-dimensional tensor; planning the fixed computing data according to the predetermined storage layout, recording the computing device for computing the fixed computing data as the locked device, and calculating the available memory capacity of the locked device; when the working library receives new computing data, determining the pre-allocation information of the new computing data according to the computing situation of the fixed computing data and the available memory capacity; determining the theoretical load of each computing device according to the pre-allocation information, and adjusting the pre-allocation information according to the theoretical load and the average access hit rate of the tensor data block.
[0007] Compared with the prior art, the technical effects achieved by adopting this technical solution are as follows: The working library can effectively organize and manage various computing resources, making resource scheduling and allocation more efficient and intelligent. By dividing the current computing data according to the minimum alignment granularity, it can ensure the correct alignment of tensor data blocks with the memory, thereby reducing access latency. The memory occupancy reflects the actual space occupied by each tensor data block in the memory, which helps to accurately allocate memory resources, optimize device usage, and avoid resource waste. Through the dynamic partitioning of the memory, it can more flexibly respond to different computing requirements. The early determination of fixed data and locked devices can effectively avoid performance jitter caused by frequent data migration, ensuring the stability and continuity of computing. The determination of the available memory capacity helps to reasonably allocate computing resources and prevent memory overload. The reasonable allocation of new computing data can avoid waste of memory resources and ensure the efficient execution of computing tasks. Understanding the computing situation helps to timely adjust the memory resource allocation and avoid local bottlenecks, thus achieving a smooth transition in dynamically changing computing requirements. Calculating the theoretical load helps to determine the maximum carrying capacity of the computing device, so that tasks can be reasonably allocated to avoid overloading a single device. The average access hit rate reflects the latency of memory access and helps to adjust the pre-allocation information.
[0008] In an embodiment of the present invention, obtaining the minimum alignment granularity of all computing devices in the working library and converting the current computing data in different dimensions into a one-dimensional tensor according to the minimum alignment granularity specifically includes: obtaining the memory alignment granularity of each computing device and screening to obtain the minimum alignment granularity corresponding to the working library; performing tensor parallel partitioning on the current computing data in different dimensions according to the minimum alignment granularity to obtain multiple tensor data blocks; splicing the multiple tensor data blocks to convert the current computing data into a one-dimensional tensor.
[0009] Compared with the prior art, the technical effects achieved by adopting this technical solution are as follows: By performing tensor parallel partitioning on data in different dimensions with the minimum alignment granularity, it can ensure that the computing device can process multiple tensor data blocks in parallel, thereby reducing the computing time. By converting multi-dimensional data into one-dimensional tensors, the storage layout of the current computing data can be simplified, thereby reducing the memory management complexity and improving the computing efficiency of tensor parallelism.
[0010] In an embodiment of the present invention, the memory space in the computing device is dynamically partitioned according to the memory occupation of each tensor data block in the one-dimensional tensor and the minimum alignment granularity to obtain a predetermined storage layout for the one-dimensional tensor, which specifically includes: screening the tensor data blocks according to the memory occupation to obtain full-load data blocks and missing data blocks; allocating the full-load data blocks to each computing device according to the memory occupation; calculating the missing memory of each missing data block, merging the missing data blocks according to the missing memory to obtain merged data blocks; allocating the merged data blocks to each computing device according to the memory occupation, and obtaining the predetermined storage layout according to the distribution of the merged data blocks and the full-load data blocks.
[0011] Compared with the prior art, the technical effects achieved by adopting this technical solution are as follows: By identifying the full-load data blocks and missing data blocks, the tensor data blocks that need to be merged are clarified, which helps to optimize the memory allocation strategy and improve the utilization efficiency of the memory of the computing device. The missing memory reflects the degree of memory waste and can effectively guide the merging strategy. By merging the missing data blocks into merged data blocks, the memory fragmentation can be effectively reduced, avoiding excessive memory waste, thereby improving the overall utilization rate of memory resources.
[0012] In an embodiment of the present invention, the fixed computing data is planned according to the predetermined storage layout, the computing device for calculating the fixed computing data is denoted as the locked device, and the available memory capacity of the locked device is calculated, which specifically includes: denoting the memory occupation of the fixed computing data as the target occupation, and calculating the memory margin of each computing device according to the predetermined storage layout; obtaining the processing plan of the fixed computing data. When the fixed computing data is planned to be processed by a single locked device, calculating the available memory capacity according to the target occupation and the memory margin; when the fixed computing data is planned to be processed by multiple locked devices, calculating the average occupation according to the target occupation and the target number of the locked devices; calculating the theoretical occupation corresponding to the reduction of the locked devices according to the target number and the target occupation; calculating the available memory capacity of the locked device according to the average occupation, the theoretical occupation and the memory margin.
[0013] Compared with the prior art, the technical effects achieved by adopting this technical solution are as follows: The target occupancy reflects the memory requirements for fixed calculation data, providing a benchmark for subsequent resource allocation of memory resources. The calculation of the memory margin reflects the available memory resources of the computing device, facilitating the dynamic allocation of memory resources, ensuring that the computing device does not malfunction or experience performance degradation due to insufficient memory. The clarity of the processing plan helps to ensure the rationality of memory resource allocation, avoiding or reducing unnecessary occupancy of the computing device, thereby reducing the number of merges of the calculation results of tensor data blocks and improving the calculation efficiency.
[0014] In an embodiment of the present invention, the available memory capacity of the locking device is calculated based on the average occupancy, the theoretical occupancy, and the memory margin, specifically including: when the memory margin of each locking device is greater than or equal to the average occupancy and less than each theoretical occupancy, the available memory capacity is determined based on the average occupancy and the memory margin; when there is a memory margin greater than or equal to the theoretical occupancy, it is determined whether the theoretical occupancy is reasonable based on the theoretical occupancy and the number of locking devices corresponding to the theoretical occupancy; if so, the available memory capacity is calculated based on the theoretical occupancy and the memory margin; if not, the reasonable occupancy is determined based on the theoretical occupancy and the memory margin, and the available memory capacity is calculated based on the reasonable occupancy and the memory margin.
[0015] Compared with the prior art, the technical effects achieved by adopting this technical solution are as follows: When multiple devices cooperate to process tasks, the average occupancy is used as a unified allocation benchmark to ensure that each device undertakes the same load, avoiding resource waste or overload caused by differences in device performance. Through the matching analysis of the theoretical occupancy and the memory margin, the utilization rate of memory resources can be optimized, thereby dynamically adjusting the number of devices participating in the calculation task. Through the sub-scenario decision-making of reasonable and unreasonable theoretical occupancies, precise allocation of memory resources is achieved, avoiding resource waste or overload caused by a "one-size-fits-all" strategy.
[0016] In an embodiment of the present invention, when the working library receives new calculation data, the pre-allocation information of the new calculation data is determined based on the calculation situation of the fixed calculation data and the available memory capacity, specifically including: regarding the computing devices other than the locking devices as conventional devices. When the fixed calculation data has not been calculated, the new calculation data is allocated according to the available memory capacity corresponding to the conventional devices to obtain the pre-allocation information; after the fixed calculation data starts to be calculated, the allocable margin of the locking device is determined according to the completion degree of the fixed calculation data; the new calculation data is allocated according to the available memory capacity corresponding to the conventional devices and the allocable margin to obtain the pre-allocation information.
[0017] Compared with the prior art, the technical effects achieved by adopting this technical solution are as follows: Classifying the locked devices and conventional devices can effectively prevent key computing tasks from being interfered, reduce memory fragmentation caused by task competition, thereby improving the utilization rate of memory resources. By tracking the execution progress of fixed tasks in real time, it provides a quantitative basis for the allocation strategy of newly added computing data, realizes the time-sharing multiplexing of memory resources, and dynamically allocates newly added computing data according to the available margin, which can effectively prevent the resources of locked devices from being idle under low load and improve the utilization efficiency of locked devices.
[0018] In an embodiment of the present invention, when the fixed computing data is not being calculated, the newly added computing data is allocated according to the available memory capacity corresponding to the conventional device to obtain pre-allocation information, which specifically includes: obtaining the newly added memory capacity corresponding to the newly added computing data, and when the newly added memory capacity is less than or equal to the available memory capacity, allocating the newly added computing data to the conventional device; when the newly added memory capacity is greater than the available memory capacity, screening the newly added computing data according to the available memory capacity to obtain excess computing data; dividing the excess computing data into the corresponding conventional devices according to the data processing progress of the conventional device.
[0019] Compared with the prior art, the technical effects achieved by adopting this technical solution are as follows: The newly added memory capacity quantifies the memory occupancy of the newly added computing data and is the basic basis for the allocation of available memory resources. The comparison between the newly added memory capacity and the available memory capacity of the conventional device is the key condition for determining the trigger of the excess data screening mechanism. Tracking the data processing progress of the conventional device can effectively prevent the conventional device from being overloaded or idle, and achieve load balancing by dynamically adjusting the division of the newly added computing data, thereby maximizing the utilization of the memory resources of the conventional device and improving the computing efficiency.
[0020] In an embodiment of the present invention, the theoretical load of each computing device is determined according to the pre-allocation information, and the pre-allocation information is adjusted according to the theoretical load and the average access hit rate of the tensor data block, which specifically includes: correcting the theoretical load according to the average access hit rate and communication efficiency to obtain the actual load of the memory of each computing device; obtaining the normal operating load of the computing device and determining whether the actual load is less than or equal to the normal operating load; if so, allocating the newly added computing data according to the pre-allocation information; if not, marking the computing device with an actual load greater than the normal operating load as an overloaded device, and marking the computing device with an actual load less than the normal operating load as an idle device; obtaining the overloaded memory and overloaded data blocks corresponding to the newly added computing data in the overloaded device; migrating the overloaded data blocks to the idle device according to the device distance and overloaded memory between the overloaded device and the idle device.
[0021] Compared with the prior art, the technical effects achieved by adopting this technical solution are as follows: The theoretical load is corrected through the average access hit rate and communication efficiency to obtain the actual load that is closer to the actual operating state of the computing device. By comparing the actual load with the normal operating load, it is determined whether the actual load is within the normal operating load of the computing device, thereby effectively avoiding the decline in computing efficiency or equipment failure caused by improper data resource allocation. The identification of overloaded devices and the determination of overloaded data blocks provide clear goals and directions for the adjustment and migration of data resources, thereby effectively reducing the probability of the computing device malfunctioning and improving the stability and reliability of the computing device. Selecting idle devices based on the device distance can reduce data transmission costs and improve communication efficiency.
[0022] In an embodiment of the present invention, there is also provided a computing resource allocation system based on data dimension conversion. The computing resource allocation method described in the above embodiment is applied to this allocation system. The allocation system includes: a storage module for storing the minimum alignment granularity and current computing data of all computing devices; a processing module for converting the current computing data in different dimensions into a one-dimensional tensor; a computing module for calculating the available memory capacity of the locked device; and an allocation module for processing the pre-allocation information of newly added computing data. This allocation system has all the technical features of the above computing resource allocation method and will not be elaborated here one by one. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] Figure 1 is one of the flowcharts of the computing resource allocation method; Figure 2 is the second flowchart of the computing resource allocation method; Figure 3 is the third flowchart of the computing resource allocation method; Figure 4 is the fourth flowchart of the computing resource allocation method; Figure 5 is the fifth flowchart of the computing resource allocation method; Figure 6 is the system schematic diagram of the computing resource allocation system; Description of the Reference Numerals: 100 - allocation system; 110 - storage module; 120 - processing module; 130 - computing module; 140 - allocation module. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0024] To make the above objects, features, and advantages of the present invention more obvious and understandable, the following detailed description of the specific embodiments of the present invention will be given with reference to the accompanying drawings.
[0025]
First Embodiment
[0026] In steps S100 and S200, the working library refers to a storage and scheduling system that centrally manages computing tasks, data resources, and device information. The working library usually contains information such as computing devices, computing data, and their corresponding resource configurations required for computing tasks. The computing device refers to a hardware device or hardware unit used to execute computing tasks, including CPUs, GPUs, and TPUs, etc. The current computing data refers to the core data set that is being used or processed when executing a computing task, such as model parameters or activation values during training. A one-dimensional tensor is a continuous one-dimensional data structure obtained by expanding multi-dimensional data according to the minimum alignment granularity. The memory occupancy refers to the actual space occupied by the tensor data block in the memory. The predetermined storage layout refers to the memory layout method specified for the tensor data block based on the memory occupancy and the minimum alignment granularity before the start of the computing task.
[0027] It should be noted that when dividing the one-dimensional tensor according to the minimum alignment granularity, since the actual memory occupancy of the tensor data block and the minimum alignment granularity may not be an integer multiple relationship, therefore, the memory space occupied by the divided tensor data block may not be exactly the same as the minimum alignment granularity. Therefore, the predetermined storage layout needs to be determined jointly according to the memory occupancy of the tensor data block and the minimum alignment granularity.
[0028] For example, when the minimum alignment granularity is 64 bytes and the actual size of the tensor data block is 50 bytes, after aligning to 64 bytes, it occupies 64 bytes, wasting 14 bytes. When the actual size of the tensor data block is 70 bytes, it needs to be aligned to 128 bytes (64×2), wasting 58 bytes.
[0029] In step S300, fixed calculation data refers to the steady-state data set that needs to occupy the computing device for a long time during the calculation process, including data such as the weights of the deep learning model, the training set, and the validation set. Locking the device refers to the dedicated computing device designated to process the fixed calculation data. The available memory capacity is the remaining allocable memory space in the normal operating memory of the locked device after deducting the memory occupancy of the fixed calculation data.
[0030] In steps S400 and S500, new calculation data refers to the newly added data that needs to participate in the calculation during the calculation task, usually including new data samples or information that enter after the start of the calculation task, as well as new intermediate data generated during the calculation process. The calculation situation refers to the working state of the device, the calculation load, and the task progress during the calculation process. The pre-allocation information refers to the memory allocation plan pre-planned for the new calculation data according to the calculation requirements and the available memory capacity. The theoretical load refers to the estimated load capacity of the computing device based on factors such as the complexity of the calculation task and the performance of the computing device. The average access hit rate refers to the proportion of successfully accessing the required tensor data block within a reasonable time.
[0031] It should be noted that the average access hit rate is usually closely related to the access frequency and the access distance. When the physical or logical distance between the computing unit and the required tensor data block in the storage hierarchy is closer, the access time is shorter and the access hit rate is higher. When the required tensor data block is highly competitively accessed by multiple computing units, the access waiting time of some computing units will be extended, resulting in a decrease in the access hit rate.
[0032] The working library can effectively organize and manage various computing resources, making resource scheduling and allocation more efficient and intelligent. By dividing the current computing data with the minimum alignment granularity, it can ensure the correct alignment of tensor data blocks with memory, thereby reducing access latency. The memory occupancy reflects the actual space occupied by each tensor data block in memory, which helps to accurately allocate memory resources, optimize device usage, and avoid resource waste. Through the dynamic partitioning of memory, it can more flexibly handle different computing requirements. The prior determination of fixed data and locked devices can effectively avoid performance jitter caused by frequent data migration, ensuring the stability and continuity of computing. The determination of available memory capacity helps to reasonably allocate computing resources, prevent memory overload, and the reasonable allocation of newly added computing data can avoid waste of memory resources, ensuring the efficient progress of computing tasks. Understanding the computing situation helps to timely adjust the memory resource allocation, avoid local bottlenecks, and thus achieve a smooth transition in dynamically changing computing requirements. Calculating the theoretical load helps to determine the maximum carrying capacity of computing devices, so that tasks can be reasonably allocated to avoid overloading a single device. The average access hit rate reflects the latency of memory access, which helps to adjust the pre-allocation information.
[0033]
Second Embodiment
[0034] In steps S110 to S130, the memory alignment granularity refers to the minimum unit for data alignment in memory, and the minimum alignment granularity refers to the minimum unit of memory allocation required by the hardware device when storing tensor data blocks, that is, the minimum value of the memory alignment granularity.
[0035] By performing tensor parallel partitioning on data of different dimensions with the minimum alignment granularity, it can ensure that the computing device can process multiple tensor data blocks in parallel, thereby reducing the computing time. By converting multi-dimensional data into a one-dimensional tensor, the storage layout of the current computing data can be simplified, thereby reducing the memory management complexity and improving the computing efficiency of tensor parallelism.
[0036]
Third Embodiment
[0037] In step S210, a full-load data block refers to a tensor data block whose memory occupancy is a multiple of the minimum alignment granularity, and a missing data block refers to a tensor data block whose memory occupancy is not a multiple of the minimum alignment granularity.
[0038] For example, when the minimum alignment granularity is 64 bytes, if the memory occupancy of a tensor data block is 64 bytes or 128 bytes, then this tensor data block is a full-load data block; if the memory occupancy of a tensor data block is 54 bytes or 74 bytes, then this tensor data block is a missing data block.
[0039] In steps S220 to S240, the missing memory refers to the memory vacancy caused by the alignment requirement of the missing data block, that is, the memory amount in the computing device that is not allocated to any data block. A merged data block is a new data block formed by splicing multiple missing data blocks, and the total memory occupancy of the merged data block is close to or equal to an integer multiple of the alignment granularity. The distribution refers to the allocation and layout of the full-load data blocks and the merged data blocks in each computing device.
[0040] For example, when the minimum alignment granularity is 64 bytes, if there are four missing data blocks, and the memory occupancies of the four missing data blocks are 28 bytes, 54 bytes, 74 bytes, and 100 bytes respectively, then when allocating memory according to the minimum alignment granularity, the missing memories corresponding to the three missing data blocks are 36 bytes, 10 bytes, 54 bytes, and 28 bytes. Therefore, according to the missing memory, the missing data blocks of 28 bytes and 100 bytes are merged, and at the same time, the missing data blocks of 54 bytes and 74 bytes are merged, to obtain two merged memory blocks with a total memory occupancy of 128 bytes each. At this time, the missing memory is 0 bytes.
[0041] It should be noted that although the space saved each time after merging missing data blocks is at the byte level, when the data volume is large and there are many computing devices, the memory space saved by merging missing data blocks can reach gigabytes or even higher levels. The saved memory space is used to allocate tensor data, thereby effectively reducing the number of computing devices. In addition, when allocating tensor data blocks to each computing device, the total memory occupied by the tensor data blocks in each computing device needs to be less than or equal to the normal operating memory of each computing device, so as to ensure the normal calculation of the computing device.
[0042] By identifying full data blocks and missing data blocks, the tensor data blocks that need to be merged are determined, which helps to optimize the memory allocation strategy and improve the utilization efficiency of the memory of the computing device. The missing memory reflects the degree of memory waste and can effectively guide the merging strategy. By merging missing data blocks into merged data blocks, memory fragmentation can be effectively reduced, excessive memory waste can be avoided, and thus the overall utilization rate of memory resources can be improved.
[0043]
Fourth Embodiment
[0044] In steps S310 and S320, the target occupancy refers to the size of the space required for the fixed computing data to occupy in the memory, the memory margin is the currently remaining available memory space in the computing device calculated according to the predetermined storage layout, the processing plan refers to the specific execution plan formulated for the fixed computing data, and the processing plan usually includes information such as the computing device used, the processing steps, and the time nodes.
[0045] In steps S330 to S350, the target quantity refers to the number of locked devices used to process fixed calculation data in the processing plan. The average occupancy refers to the average value of the target occupancy that each locked device needs to bear in the scenario of collaborative calculation by multiple locked devices. The theoretical occupancy is the memory occupancy of the fixed calculation data in each locked device after the reduction of the locked devices.
[0046] It should be noted that when the fixed calculation data plan is completed by a single locked device, the available memory capacity is the difference between the sum of the memory margins of each computing device and the target occupancy. When the fixed calculation data plan is completed by multiple locked devices, the average occupancy is the ratio of the target occupancy to the target quantity. The reduction number of locked devices is determined according to the target occupancy of the fixed calculation data and the memory margin of each computing device.
[0047] For example, when the number of computing devices is 5, the normal operating memory is 128GB, among which there is 1 locked device, and the memory margins of the computing devices are 25GB, 18GB, 24GB, 31GB, and the memory margin of the locked device is 108GB. If the target occupancy of the fixed calculation data is 90GB, then the fixed calculation data is allocated to the memory of the locked device. At this time, the available memory capacity of the computing device is 116GB. When the number of computing devices is 9, among which there are 3 locked devices, if the sum of the memory margins of the computing devices other than the locked devices is 125GB, and the memory margins of the locked devices are all 80GB, when the target occupancy of the fixed calculation data is 90GB, then one locked device can be reduced, and the number of locked devices becomes 2, and the theoretical occupancy is 40GB.
[0048] The target occupancy reflects the memory requirements of the fixed calculation data, provides a benchmark for the subsequent resource allocation of memory resources. The calculation of the memory margin reflects the available memory resources of the computing device, helps the dynamic allocation of memory resources, ensures that the computing device will not malfunction or have performance degradation due to insufficient memory. The clarity of the processing plan helps to ensure the rationality of memory resource allocation, avoid or reduce unnecessary occupancy of computing devices, thereby reducing the number of merges of the calculation results of tensor data blocks and improving the calculation efficiency.
[0049]
Fifth Embodiment
[0050] In steps S351 to S354, the reasonable occupation amount is the memory amount evenly allocated to each locked device for the fixed calculation data after calculating the reasonable number of locked devices according to the theoretical occupation amount and the memory surplus.
[0051] It should be noted that the reduced locked devices are recorded as adjusted devices. When the memory surplus of each locked device is greater than or equal to the average occupation amount and less than each theoretical occupation amount, it means that the theoretical occupation amount of the fixed calculation data in the adjusted devices has exceeded the memory surplus of the locked devices, that is, the fixed calculation data cannot be completely allocated to the adjusted devices. At this time, reducing the number of locked devices will further increase the single-device load. Therefore, the original number of locked devices needs to be maintained; when there is memory surplus greater than or equal to the theoretical occupation amount, it means that the theoretical occupation amount of the fixed calculation data in the adjusted devices does not exceed the memory surplus of the locked devices, that is, the fixed calculation data can be completely allocated to the adjusted devices. However, the number of adjusted devices is different, and the corresponding theoretical occupation amounts are also different. Therefore, it is necessary to judge whether the theoretical occupation amount is reasonable according to the theoretical occupation amount and the number of adjusted devices, and select the corresponding calculation method for the available memory capacity.
[0052] For example, when the target number of locked devices is 5, the average occupation amount is 10 GB, the number of adjusted devices can be 4, 3, 2, and 1, and the corresponding theoretical occupation amounts are 12.5 GB, 16.7 GB, 25 GB, and 50 GB respectively. If the theoretical occupation amount is 12.5 GB and the number of adjusted devices is 4, and the memory surplus of the adjusted devices is 18 GB each, at this time the theoretical occupation amount is less than the memory surplus. The 4 locked devices carrying a 50 GB task will cause resource redundancy. Therefore, this theoretical occupation amount and the number of adjusted devices corresponding to the theoretical occupation amount are determined to be unreasonable. At this time, determine the reasonable occupation amount according to the theoretical occupation amount and the memory surplus. When the memory surplus is 18 GB each, the closest theoretical occupation amount is 16.7 GB. Therefore, the reasonable occupation amount is 16.7 GB, and the minimum number of adjusted devices can be 3, that is, the maximum number of locked devices can be reduced by 2. When the theoretical occupation amount is unreasonable, the available memory capacity is the difference between the total memory surplus and the total reasonable occupation amount, that is, 4 GB.
[0053] When multiple devices collaborate to process tasks, the average occupancy is used as a unified allocation benchmark to ensure that each device undertakes the same load, avoiding resource waste or overload caused by differences in device performance. Through the matching analysis of the theoretical occupancy and the remaining memory, the utilization rate of memory resources can be optimized, thereby dynamically adjusting the number of devices participating in the computing task. Through the scenario-based decision-making of reasonable and unreasonable theoretical occupancies, precise allocation of memory resources can be achieved, avoiding resource waste or overload caused by a "one-size-fits-all" strategy.
[0054]
Sixth Embodiment
[0055] In steps S410 to S430, the regular device refers to other computing devices except the locked device, which are used to process new calculation data or non-fixed tasks. The completion degree refers to the calculation progress of the fixed calculation data on the locked device, usually expressed as a percentage. The allocable surplus refers to the available memory capacity released by the locked device during the process of processing the fixed calculation data, and the allocable surplus is determined by the completion degree of the fixed calculation data.
[0056] It should be noted that the memory of the locked device always gives priority to meeting the needs of the fixed calculation data. When the fixed calculation data has not started to be calculated, the locked device is in a pre-occupied state. If the new calculation data occupies the locked device before the fixed calculation data starts to be calculated, it may cause the fixed calculation task to be delayed or failed due to insufficient memory resources, affecting the calculation efficiency. Therefore, the new calculation data can only be allocated to regular devices. After the calculation of the fixed calculation data starts, as the computing task progresses, part of the memory is gradually released, but the computing unit of the locked device may still be completely occupied by the fixed task. At this time, even if the memory is idle, there is no remaining computing power to process new tasks. Therefore, when allocating new calculation data to the locked device, the theoretical calculation process of the fixed calculation data needs to be considered.
[0057] For example, according to the theoretical calculation process of fixed calculation data, when the completion degree of the fixed calculation data is determined to be 60% - 80%, the calculation unit of the locked device will be completely occupied by the fixed task again. Therefore, the safety threshold of the allocable margin in the locked device is 60% - 80% of the memory occupancy of the fixed calculation data. Therefore, when the allocable margin is 50%, the newly added calculation data can be allocated to the locked device, while when the allocable margin is 70%, the newly added calculation data cannot be allocated to the locked device.
[0058] The classification of the locked device and the conventional device can effectively prevent key calculation tasks from being interfered, reduce memory fragmentation caused by task competition, thereby improving the utilization rate of memory resources. By tracking the execution progress of the fixed task in real time according to the completion degree, it provides a quantitative basis for the allocation strategy of the newly added calculation data, realizes the time-sharing multiplexing of memory resources, and dynamically allocates the newly added calculation data according to the allocable margin, which can effectively prevent the resources of the locked device from being idle under low load and improve the utilization efficiency of the locked device.
[0059]
Seventh Embodiment
[0060] In steps S411 to S413, the newly added memory capacity refers to the memory amount required by the newly added calculation data, the excess calculation data refers to the data part where the memory requirement of the newly added calculation data exceeds the available memory capacity of the conventional device, and the data processing progress refers to the calculation progress when calculating the current calculation data on the conventional device, usually expressed as a percentage.
[0061] It should be noted that when the memory requirement of newly added computing data does not exceed the available memory capacity of conventional devices, the principle of "minimizing the number of devices" is preferentially followed for allocation, that is, the newly added computing data is concentrated and allocated to one or a few conventional devices as much as possible to avoid resource fragmentation caused by dispersion to multiple conventional devices. When the newly added memory capacity exceeds the available memory capacity of conventional devices, the computing tasks in the newly added computing data are sorted by priority, and tasks are reserved in order from high to low priority until the total occupied memory is close to the available memory of conventional devices. The remaining unallocated computing tasks enter the overage computing data queue, and the method for allocating overage computing data according to the data processing progress of conventional devices refers to step S420.
[0062] The newly added memory capacity quantifies the memory occupancy of newly added computing data and is the basic basis for the allocation of available memory resources. The comparison between the newly added memory capacity and the available memory capacity of conventional devices is the key condition for determining the triggering of the overage data screening mechanism. Tracking the data processing progress of conventional devices can effectively avoid overloading or idling of conventional devices, and achieve load balancing by dynamically adjusting the division of newly added computing data, thereby maximizing the utilization of the memory resources of conventional devices and improving computing efficiency.
[0063]
Eighth Embodiment
[0064] In steps S510 to S540, the communication efficiency refers to the ratio of the actual data transmission volume per unit time between computing devices to the theoretical data transmission volume in the ideal state. The communication efficiency is usually affected by factors such as network bandwidth and device interface performance. The actual load is the real-time load of the computing device obtained by correcting the theoretical load through the communication efficiency and the average access hit rate. The normal operating load is the maximum load threshold that the computing device can withstand for a long time under normal and stable operating conditions. An overloaded device is a computing device whose actual load exceeds the normal operating load, and an idle device is a computing device whose actual load is lower than the normal operating load. The formula for calculating the actual load is as follows: Actual load = Theoretical load × (1 + Average access hit rate + Communication efficiency).
[0065] Among them, the value ranges of both the average access hit rate and the communication efficiency are 0% to 1%.
[0066] It should be noted that when the average access hit rate and the communication efficiency are low, the theoretical load of the computing device will be lower than the actual load. Therefore, when calculating the actual load, the average access hit rate and the communication efficiency need to be considered. After data allocation is performed according to the pre-allocation information of the newly added calculation data, although the memory occupancy in each computing device is less than or equal to the normal operating memory, due to low communication efficiency or the tensor data blocks of the newly added calculation data need to be frequently accessed in the computing device, it may result in a situation where although the memory occupancy does not exceed the limit, the actual calculation load may exceed the normal operating load. Therefore, it is necessary to confirm whether the pre-allocation information of the newly added calculation data is reasonable according to the actual load and the normal operating load.
[0067] In steps S550 and S560, an overloaded data block refers to the tensor data block in the newly added calculation data that causes the actual load of the computing device to exceed the normal operating load. The overloaded memory refers to the newly added memory capacity occupied by the overloaded data block. The device distance is a logical indicator used to measure the communication efficiency or resource sharing degree between computing devices in tensor parallel computing. The device distance is usually proportional to the communication efficiency and the resource sharing degree.
[0068] It should be noted that it is preferred to select an idle device that is close to the overloaded device and has sufficient load to migrate the overloaded data block. A short device distance means lower network latency and higher bandwidth utilization, which can significantly reduce the data migration time and network congestion during long-distance migration. Sufficient load can avoid secondary overload caused by insufficient computing load of the target idle device.
[0069] The theoretical load is corrected by the average access hit rate and communication efficiency to obtain the actual load that is closer to the actual operating state of the computing device. By comparing the actual load with the normal operating load, it is determined whether the actual load is within the normal operating load of the computing device, thereby effectively avoiding the reduction of computing efficiency or equipment failure caused by improper data resource allocation. The identification of overloaded devices and the determination of overloaded data blocks provide clear goals and directions for the adjustment and migration of data resources, thereby effectively reducing the probability of equipment failure in the computing device and improving the stability and reliability of the computing device. Selecting idle devices according to the device distance can reduce data transmission costs and improve communication efficiency.
[0070]
Ninth Embodiment
[0071] Although the present invention is disclosed as above, the present invention is not limited thereto. Any person skilled in the art can make various changes and modifications without departing from the spirit and scope of the present invention. Therefore, the protection scope of the present invention should be subject to the scope defined by the claims.
Claims
1. A computing resource allocation method based on data dimension conversion, characterized in that: The computing resource allocation method comprises: Obtaining the minimum alignment granularity of all computing devices in the working library, and converting current computing data of different dimensions into one-dimensional tensors according to the minimum alignment granularity; Dynamically partitioning the memory space in the computing device according to the memory occupancy of each tensor data block in the one-dimensional tensor and the minimum alignment granularity to obtain a predetermined storage layout of the one-dimensional tensor; Planning the fixed computing data according to the predetermined storage layout, recording the computing device for computing the fixed computing data as a locked device, and calculating the available memory capacity of the locked device; When the working library receives the newly added computing data, the pre-allocation information of the newly added computing data is determined according to the computing status of the fixed computing data and the available memory capacity; The theoretical load of each of the computing devices is determined according to the pre-allocation information, and the pre-allocation information is adjusted according to the theoretical load and the average access hit rate of the tensor data block.
2. The computing resource allocation method according to claim 1, characterized in that: The step of obtaining the minimum alignment granularity of all computing devices in the working library and converting current computing data of different dimensions into one-dimensional tensors according to the minimum alignment granularity specifically includes: Obtaining the memory alignment granularity of each of the computing devices, and screening to obtain the minimum alignment granularity corresponding to the working library; Dividing the current calculation data on different dimensions in tensor parallel according to the minimum alignment granularity to obtain a plurality of tensor data blocks; The plurality of tensor data blocks are concatenated to convert the current calculation data into the one-dimensional tensor.
3. The computing resource allocation method according to claim 2, characterized in that: The dynamically dividing the memory space in the computing device according to the memory occupancy of each tensor data block in the one-dimensional tensor and the minimum alignment granularity to obtain a predetermined storage layout of the one-dimensional tensor specifically includes: The tensor data blocks are screened according to the memory occupancy to obtain fully loaded data blocks and missing data blocks; Allocating the fully loaded data blocks to each of the computing devices according to the memory occupancy; Calculating the missing memory of each missing data block, and merging the missing data blocks according to the missing memory to obtain a merged data block; The merged data blocks are allocated to each of the computing devices according to the memory occupancy, and the predetermined storage layout is obtained according to the distribution of the merged data blocks and the fully loaded data blocks.
4. The computing resource allocation method according to claim 3, characterized in that: The step of planning the fixed computing data according to the predetermined storage layout, recording the computing device for computing the fixed computing data as a locked device, and calculating the available memory capacity of the locked device specifically includes: Recording the memory occupancy of the fixed computing data as a target occupancy, and calculating the memory margin of each computing device according to the predetermined storage layout; Obtaining a processing plan for the fixed computing data, and when the fixed computing data is planned to be processed by a single locking device, calculating the available memory capacity according to the target occupancy and the memory margin; When the fixed calculation data is planned to be processed by a plurality of the locking devices, an average occupancy is calculated according to the target occupancy and the target number of the locking devices; Calculating the theoretical occupancy corresponding to the reduction of the locked device according to the target number and the target occupancy; The available memory capacity of the locking device is calculated according to the average occupancy, the theoretical occupancy and the memory margin.
5. The computing resource allocation method according to claim 4, characterized in that: The calculating the available memory capacity of the locking device according to the average occupancy, the theoretical occupancy and the memory margin specifically includes: When the memory margin of each of the locked devices is greater than or equal to the average occupancy and less than each of the theoretical occupancy, determining the available memory capacity according to the average occupancy and the memory margin; When the memory margin is greater than or equal to the theoretical occupancy, judging whether the theoretical occupancy is reasonable according to the theoretical occupancy and the number of locked devices corresponding to the theoretical occupancy; If yes, then calculating the available memory capacity according to the theoretical occupancy and the memory margin; If not, a reasonable occupancy is determined according to the theoretical occupancy and the memory margin, and the available memory capacity is calculated according to the reasonable occupancy and the memory margin.
6. The computing resource allocation method according to claim 4, characterized in that: When the working library receives the newly added computing data, determining the pre-allocation information of the newly added computing data according to the computing status of the fixed computing data and the available memory capacity specifically includes: Recording the computing devices other than the locked device as regular devices; When the fixed calculation data has not been calculated, the newly added calculation data is allocated according to the available memory capacity corresponding to the conventional device to obtain the pre-allocation information; When the fixed calculation data starts to be calculated, determining the allocatable margin of the locking device according to the completion degree of the fixed calculation data; The newly added computing data is allocated according to the available memory capacity and the allocatable margin corresponding to the conventional device to obtain the pre-allocation information.
7. The computing resource allocation method according to claim 6, characterized in that: When the fixed calculation data is not calculated, allocating the newly added calculation data according to the available memory capacity corresponding to the conventional device to obtain the pre-allocation information specifically includes: Acquire the newly added memory capacity corresponding to the newly added computing data, and when the newly added memory capacity is less than or equal to the available memory capacity, allocate the newly added computing data to the conventional device; When the newly added memory capacity is greater than the available memory capacity, the newly added calculation data is screened according to the available memory capacity to obtain excess calculation data; The excess calculation data is divided into corresponding conventional devices according to the data processing progress of the conventional devices.
8. The computing resource allocation method according to claim 7, characterized in that: Determining the theoretical load of each computing device according to the pre-allocation information, and adjusting the pre-allocation information according to the theoretical load and the average access hit rate of the tensor data block specifically includes: The theoretical load is corrected according to the average access hit rate and the communication efficiency to obtain the actual load of the memory of each computing device; Obtaining the normal operating load of the computing device, and determining whether the actual load is less than or equal to the normal operating load; If yes, allocate the newly added computing data according to the pre-allocation information; If not, the computing device whose actual load is greater than the normal operating load is recorded as an overloaded device, and the computing device whose actual load is less than the normal operating load is recorded as an idle device; Obtaining the overload memory and overload data block corresponding to the newly added calculation data in the overload device; The overloaded data block is migrated to the idle device according to the device distance between the overloaded device and the idle device and the overloaded memory.
9. A computing resource allocation system based on data dimension conversion, characterized in that: The computing resource allocation method according to any one of claims 1 to 8 is applied to the allocation system, the allocation system comprising: A storage module, the storage module is used to store the minimum alignment granularity and the current computing data of all the computing devices; A processing module, the processing module is used to convert the current calculation data of different dimensions into the one-dimensional tensor; A calculation module, the calculation module is used to calculate the available memory capacity of the locking device; An allocation module, wherein the allocation module is used to process the pre-allocation information of the newly added computing data.
Citation Information
Patent Citations
Data processing method and device, computer equipment and storage medium
CN111401511A
Deep neural network accelerator based on dynamic reconfigurable pulsation tensor operation engine
CN114781632A
Method, device and medium for converting layout of tensor data
CN117170588A
Cloud data processing system based on artificial intelligence algorithm
CN119473645A
Method and device for realizing task splitting based on multi-core processor and related product
CN119847730A