Sparse data load balancing method, device, equipment and storage medium
Patent Information
- Application Number
- CN202611072328.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-17
- Publication Date
- 2026-09-25
AI Technical Summary
[0004]为了降低整体的通信量,优化方案仅对稀疏数据进行传输,但由于各个任务稀疏度不同,存在负载不均衡的问题
[0020]本申请提供一种稀疏数据负载均衡方法、装置、设备、存储介质,该方法获取稀疏数据任务序列;根据稀疏数据任务序列中各任务所包含的有效非零数据量及硬件通信代价参数,确定各任务的通信负载权重;根据处理器数量、稀疏数据任务序列中任务数量、各任务的通信负载权重生成分割位置矩阵;根据分割位置矩阵确定分割位置;按分割位置将稀疏数据任务序列中的任务负载均衡至各处理器。本申请的方法可在严禁打乱物理内存顺序的前提下,精准寻找到任务序列的最佳切割点,进行负载均衡,从而消除并行计算中的“木桶效应”,最大化系统的整体吞吐率。
Smart Images

Figure CN122816896A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular to a sparse data load balancing method, apparatus, device, and storage medium. Background Technology
[0002] 3D-FFT (Three-dimensional Fast Fourier Transform) is the core operation for realizing bidirectional time-frequency domain transformation and plays an irreplaceable role in the field of quantum chemistry.
[0003] For example, in the Pseudopotential Plane Wave (PPW) calculation scenario in new material simulation, due to the physical "energy cutoff" mechanism, the effective data in the reciprocal space is sparsely distributed as an "ellipsoid" rather than filling the entire cuboid computational grid.
[0004] To reduce overall communication volume, the optimization scheme only transmits sparse data, but due to the different sparsity of each task, there is a problem of unbalanced load. Summary of the Invention
[0005] To address one of the aforementioned technical deficiencies, this application provides a sparse data load balancing method, apparatus, device, and storage medium.
[0006] A first aspect of this application provides a sparse data load balancing method, the method comprising: Sequence of tasks for acquiring sparse data; The communication load weight of each task is determined based on the amount of effective non-zero data contained in each task in the sparse data task sequence and the hardware communication cost parameters. Based on the number of processors Number of tasks in a sparse data task sequence The communication load weights of each task are used to generate a segmentation location matrix; where the generated segmentation location matrix is... OK Column matrix; Determine the segmentation position based on the segmentation position matrix; The task load in the sparse data task sequence is balanced across the processors according to the partition position.
[0007] Optionally, based on the number of processors Number of tasks in a sparse data task sequence The communication load weights for each task generate a segmentation location matrix, including: Determine the cumulative weight of each cumulative position in the sparse data task sequence; where the cumulative position in the sparse data task sequence is... Cumulative weights , To accumulate positions, , For the summation index, , For the sparse data task sequence, the first Communication load weights for each task; Based on the cumulative weights of each cumulative position in the sparse data task sequence, determine the first position in the sparse data task sequence. Tasks assigned to The segmentation location of each processor; where, Number of processors , This represents the sequence number of the task within the sparse data task sequence. ; Generate a segmentation position matrix from all the segmentation positions.
[0008] Optionally, based on the cumulative weight of each cumulative position in the sparse data task sequence, the first... Tasks are assigned to The segmentation locations of each processor include: like Then determine the first part of the sparse data task sequence. Tasks are assigned to The segmentation position of each processor ,and ;in, To make the first part of the sparse data task sequence Tasks are assigned to The globally optimal bottleneck value for each processor; like Then determine the first part of the sparse data task sequence. Tasks assigned to The segmentation position of each processor ,and ;in, This refers to the sequence number of the segmentation location within the sparse data task sequence. To find the function with the maximum value, To make the function reach its minimum value, the corresponding value, To make the first part of the sparse data task sequence Tasks assigned to The globally optimal bottleneck value for each processor for The minimum value, for The maximum value that can be obtained.
[0009] Optionally, Or, when hour ,when hour
[0010] Optionally, Or, if it exists ,but If it does not exist ,but ; in, for The effective value, and hour, ,or, and hour, for Integer values in, and and .
[0011] Optionally, determining the segmentation position based on the segmentation position matrix includes: Initialize the current line identifier Current column identifier ; In the segmentation position matrix OK The value of the column is determined by the current position; renew ; like Then update At the current position, repeating the process will segment the position matrix. OK The steps for determining the column value as the split position and subsequent steps, until... ; like Then all current positions will be determined as the dividing positions.
[0012] A second aspect of this application provides a sparse data load balancing device, the device comprising: The acquisition module is used to acquire sparse data task sequences; The first determining module is used to determine the communication load weight of each task based on the effective non-zero data volume contained in each task in the sparse data task sequence and the hardware communication cost parameters. The generation module is used to determine the number of processors. Number of tasks in a sparse data task sequence The communication load weights of each task are used to generate a segmentation location matrix; where the generated segmentation location matrix is... OK Column matrix; The second determining module is used to determine the segmentation position based on the segmentation position matrix; The load balancing module is used to distribute the workload of tasks in a sparse data task sequence to each processor according to the partition position.
[0013] Optionally, a generation module is used to determine the cumulative weight of each cumulative position in the sparse data task sequence; wherein, the cumulative position in the sparse data task sequence is... Cumulative weights , To accumulate positions, , For the summation index, , For the sparse data task sequence, the first Communication load weights for each task; Based on the cumulative weights of each cumulative position in the sparse data task sequence, determine the first position in the sparse data task sequence. Tasks assigned to The segmentation location of each processor; where, Number of processors , This represents the sequence number of the task within the sparse data task sequence. ; Generate a segmentation position matrix from all the segmentation positions.
[0014] Optionally, based on the cumulative weight of each cumulative position in the sparse data task sequence, the first... Tasks assigned to The segmentation locations of each processor include: like Then determine the first part of the sparse data task sequence. Tasks assigned to The segmentation position of each processor ,and ;in, To make the first part of the sparse data task sequence Tasks assigned to The globally optimal bottleneck value for each processor; like Then determine the first part of the sparse data task sequence. Tasks assigned to The segmentation position of each processor ,and ;in, This refers to the sequence number of the segmentation location within the sparse data task sequence. To find the function with the maximum value, To make the function reach its minimum value, the corresponding value, To make the first part of the sparse data task sequence Tasks assigned to The globally optimal bottleneck value for each processor for The minimum value, for The maximum value that can be obtained.
[0015] Optionally, Or, when hour ,when hour
[0016] Optionally, Or, if it exists ,but If it does not exist ,but ; in, for The effective value, and hour, ,or, and hour, for Integer values in, and and .
[0017] Optionally, determining the segmentation position based on the segmentation position matrix includes: Initialize the current line identifier Current column identifier ; In the segmentation position matrix OK The value of the column is determined by the current position; renew ; like Then update At the current position, repeating the process will segment the position matrix. OK The steps for determining the column value as the split position and subsequent steps, until... ; like Then all current positions will be determined as the dividing positions.
[0018] A third aspect of this application provides an electronic device, comprising: Memory; Processor; and Computer programs; The computer program is stored in the memory and configured to be executed by the processor to implement the method described in the first aspect above.
[0019] In a fourth aspect, this application provides a computer-readable storage medium having a computer program stored thereon; the computer program is executed by a processor to implement the method described in the first aspect above.
[0020] This application provides a sparse data load balancing method, apparatus, device, and storage medium. The method obtains a sparse data task sequence; determines the communication load weight of each task based on the effective non-zero data volume and hardware communication cost parameters of each task in the sparse data task sequence; generates a segmentation position matrix based on the number of processors, the number of tasks in the sparse data task sequence, and the communication load weight of each task; determines the segmentation position based on the segmentation position matrix; and balances the task load in the sparse data task sequence to each processor according to the segmentation position. This method can accurately find the optimal segmentation point of the task sequence and perform load balancing under the premise of strictly prohibiting disruption of the physical memory order, thereby eliminating the "weakest link" effect in parallel computing and maximizing the overall throughput of the system. Attached Figure Description
[0021] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings: Figure 1 A flowchart illustrating a sparse data load balancing method provided in an embodiment of this application; Figure 2 A schematic diagram of the structure of a sparse data load balancing device provided in an embodiment of this application; Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0022] To make the technical solutions and advantages of the embodiments of this application clearer, the exemplary embodiments of this application will be described in further detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not an exhaustive list of all embodiments. It should be noted that, unless otherwise specified, the embodiments and features in the embodiments of this application can be combined with each other.
[0023] In developing this application, the inventors discovered that 3D-FFT (Three-dimensional Fast Fourier Transform) is the core operation for achieving bidirectional time-frequency domain conversion and plays an irreplaceable role in the field of quantum chemistry. To reduce overall communication load, the optimized scheme only transmits sparse data; however, due to the varying sparsity of different tasks, there is a problem of unbalanced load.
[0024] To address the aforementioned issues, this application provides a sparse data load balancing method, apparatus, device, and storage medium. The method obtains a sparse data task sequence; determines the communication load weight of each task based on the effective non-zero data volume and hardware communication cost parameters of each task in the sparse data task sequence; generates a segmentation position matrix based on the number of processors, the number of tasks in the sparse data task sequence, and the communication load weight of each task; determines the segmentation position based on the segmentation position matrix; and balances the task load in the sparse data task sequence to each processor according to the segmentation position. This method can accurately find the optimal segmentation point of the task sequence for load balancing while strictly prohibiting disruption of the physical memory order, thereby eliminating the "weakest link" effect in parallel computing and maximizing the overall system throughput.
[0025] See Figure 1 This embodiment provides a sparse data load balancing method, the implementation process of which is as follows: 101, Obtain the sparse data task sequence.
[0026] The sparse data task sequence obtained in step 101 is the sparse data task sequence to be processed.
[0027] If the total length of the sparse data task sequence is This indicates that the number of tasks in the sparse data task sequence is... .
[0028] 102. Based on the effective non-zero data volume and hardware communication cost parameters contained in each task in the sparse data task sequence, determine the communication load weight of each task.
[0029] Due to physical limitations, 3D-FFT data is truncated at geometric boundaries, resulting in significant differences in the effective computational load across different task sequences. In step 102, the specific attribute information of each task in the sparse data task sequence can be dynamically read by parsing the compressed index table of the sparse data, thereby obtaining the effective non-zero data volume of each task. Based on the effective non-zero data volume and hardware communication cost parameters, the communication load weight of each task is determined. For example, to obtain the effective non-zero data volume of the sparse data task sequence... The communication load weight of a task is determined by the amount of valid non-zero data contained in the task, combined with the hardware communication costs such as DDR SDRAM (Double Data Rate Synchronous Dynamic Random Access Memory, DDR memory) reads and on-chip data distribution. .
[0030] in, This represents the sequence number of the task within the sparse data task sequence. .
[0031] In step 102, the first The communication cost of each task is expressed by the extracted weights. Based on this, the total pipeline delay of DDR SDRAM read and on-chip distribution is comprehensively considered.
[0032] 103, based on the number of processors Number of tasks in a sparse data task sequence The communication load weights of each task are used to generate a segmentation location matrix.
[0033] The generated segmentation position matrix is as follows: OK Column matrix.
[0034] The implementation process of step 103 is as follows: 103-1, Determine the cumulative weight of each cumulative position in the sparse data task sequence.
[0035] Among them, the cumulative position in the sparse data task sequence is Cumulative weights , To accumulate positions, , For the summation index, , For the sparse data task sequence, the first Communication load weights for each task.
[0036] Given the sparse data task sequence obtained in step 101 as {task 1, task 2, task 3, task 4, task 5}, the communication load weight of task 1 determined in step 102 is... Communication load weights for Task 2 Communication load weights for Task 3 Communication load weights for Task 4 Communication load weight for Task 5 For example, step 103-1 determines the cumulative weight of the cumulative position 1. The cumulative weight at position 2 The cumulative weight at position 3 The cumulative weight at position 4 The cumulative weight at position 5 As shown in Table 1.
[0037] Table 1
[0038] By using the cumulative weights at each cumulative position, the continuous task range that any processor can undertake can be determined. Total load directly through This avoids redundant accumulation calculations and enables rapid evaluation of load in any interval within O(1) time, reducing the computational complexity from O(n) to constant time O(1), thus supporting rapid planning of large-scale data.
[0039] 103-2, Based on the cumulative weight of each cumulative position in the sparse data task sequence, determine the first position in the sparse data task sequence. Tasks assigned to The segmentation position of each processor.
[0040] in, Number of processors , This represents the sequence number of the task within the sparse data task sequence.
[0041] To eliminate search redundancy in conventional global optimization algorithms and reduce computational overhead, pruning is performed based on hardware physical limitations. For example, when traversing the number of tasks... When setting the lower limit of the search range to (ensure before) (Number of processors not idle), with a maximum of [number] processors. (Ensure that follow-up) Each processor reserves at least one task, thus obtaining... . The limitation eliminates a large number of invalid allocation states, significantly improving the running speed of the sparse data load balancing method provided in this embodiment.
[0042] Step 103-2 implements Min-Max dynamic programming for multi-core processors while maintaining the physical address continuation of the task sequence. By mathematically recursively balancing the "preceding cumulative bottleneck" and the "current node's immediate load," the globally optimal set of split points is found, ensuring that the time taken for the maximum load in the system reaches the theoretical minimum.
[0043] The implementation process of step 103-2 is as follows: A. If Then determine the first part of the sparse data task sequence. Tasks assigned to The segmentation position of each processor ,and .
[0044] in, To make the first part of the sparse data task sequence Tasks assigned to The globally optimal bottleneck value for each processor, which is about to be... Each task is assigned to the previous The "minimum maximum bottleneck value" that can be achieved with one processor.
[0045] The processor count is 1, meaning all the previous Each task is handled independently by this single processor, without any preceding processor. Therefore, the equivalent preceding cut point at this time... The corresponding bottleneck value is the front. The sum of the weights of all tasks, i.e. .
[0046] B. If Then determine the first part of the sparse data task sequence. Tasks assigned to The segmentation position of each processor ,and .
[0047] in, This refers to the sequence number of the segmentation location within the sparse data task sequence. To find the function with the maximum value, To make the function reach its minimum value, the corresponding value, To make the first part of the sparse data task sequence Tasks assigned to The globally optimal bottleneck value for each processor for The minimum value, for The maximum value that can be obtained.
[0048] The Min-Max-based state transition calculation can be performed at all valid partition locations. In the middle, find a point such that "the front" "Preceding bottleneck value of the current processor" and "Current bottleneck value of the current processor" The larger of the two factors, "instantaneous load of each processor" and "overall bottleneck," is minimized. To meet the physical constraint that compute nodes cannot be idle, the partitioning location... (i.e., the preceding sequence) The range of the total number of tasks handled by each processor is strictly limited: (Right now lower bound To ensure the preceding At least one processor can be allocated Each task (at least one per processor). (Right now The upper realm ), to ensure that the current number Each processor must reserve at least one task (i.e., the first task). (One task).
[0049] Based on this, the global optimal bottleneck value from the preceding sequence is used... Monotonic and unchanging, current load follows Due to the monotonically decreasing characteristic (i.e., based on the monotonically increasing load, when a new task is added to the right, the load on the tail processor will inevitably increase sharply, the optimal split position will remain unchanged or move to the right, and the optimal balanced split position will exhibit a monotonically non-decreasing characteristic), during the state transition process, the optimization pointer does not need to be reset to zero every time, but inherits the optimal split position from the previous state. The optimal solution is found by directly extending unidirectionally to the right. For the same number of processors... ,according to Calculate in ascending order. When At that time, if there is no previous valid state for that number of processors, take... ;when At that time, the optimization pointer inherits the optimal segmentation position from the previous task-scale state. Starting from that position, search for the optimal solution to the right, that is... This way it will no longer be... The previous position was processed (i.e., invalid evaluation was directly proposed, and the same position was inherited). The optimization pointer of the previous state (i.e., the optimal split position) (Continuing to search for optimization to the right), monotonicity optimization is achieved. If conventional dynamic programming is used, it is necessary to enumerate the historical splitting points in each state, with a complexity of O(mn²). The sparse data load balancing method provided in this embodiment can optimize this to O(mn) by utilizing monotonicity optimization. This completely avoids the conventional full traversal and ensures the real-time performance of scheduling.
[0050] The sparse data load balancing method provided in this embodiment utilizes a monotonic pointer mechanism. When searching for the optimal cut point, the optimization pointer only extends unidirectionally to the right. Once overtaking occurs (the preceding bottleneck is greater than the current load), the process is immediately interrupted without backtracking. The optimization pointer reduces the inner complexity of state transitions to an extremely low level, significantly improving planning speed and enabling the solution to maintain high real-time scheduling performance even when facing a massive number of tasks.
[0051] Additionally, if This indicates that the bottleneck at the beginning is much smaller than the load at the end, so the current segmentation position is too far to the left and needs to be shifted further to the right. If If this indicates that the preceding bottleneck has overtaken the tail load, it is determined that the optimal equilibrium intersection point has been passed, and the search is immediately terminated. It can also be restricted to satisfying of .
[0052] In other words, Or, when hour ,when hour, That is, when At that time, because there is no situation where the number of tasks is the same under the same number of processors, The valid state is considered, and the lower limit of the segmentation position is thus determined. ;when At that time, the same processor Quantity, previous task scale Corresponding optimal segmentation position As the lower bound of the candidate segmentation positions in the current state, the lower bound of the segmentation positions is therefore determined as follows. .
[0053] in, Or, if it exists ,but If it does not exist ,but .
[0054] in, for The effective value, and hour, ,or, and hour, for Integer values in, and and .
[0055] Through the above and By imposing constraints, the upper and lower bounds of the task allocation search range can be dynamically set during the state transition traversal process, taking into account the characteristic that the processor cannot be idle. This method significantly reduces the invalid state space and substantially lowers the computational cost of the planning process.
[0056] Example 1: Given the sparse data task sequence obtained in step 101 as {task 1, task 2, task 3, task 4, task 5}, the number of processors... Number of tasks in a sparse data task sequence The cumulative weight of each cumulative position ( As shown in Table 1, , Taking this as an example, the implementation process of step 103-2 will be explained.
[0057] I. For a single processor, that is, the number of processes handled by the processor.
[0058] In this situation .
[0059] Depend on hour ,and Therefore, we can conclude that: hour, , .
[0060] hour, , .
[0061] hour, , .
[0062] II. For two processors, i.e., the number of processors processing...
[0063] In this situation . .
[0064] Depend on hour ,and Therefore, we can conclude that: 1.
[0065] hour, , .
[0066] but .
[0067] .
[0068] 2.
[0069] hour, , or .
[0070] but .
[0071] .
[0072] 3.
[0073] hour, , or or .
[0074] but .
[0075] .
[0076] Right now hour, , .
[0077] hour, , .
[0078] hour, , .
[0079] III. For three processors, i.e., the number of processors processing...
[0080] In this situation . .
[0081] Depend on hour ,and Therefore, we can conclude that: 1.
[0082] hour, , .
[0083] but .
[0084] .
[0085] 2.
[0086] hour, , or .
[0087] but .
[0088] .
[0089] 3.
[0090] hour, , or or .
[0091] but .
[0092] .
[0093] Right now hour, , .
[0094] hour, , .
[0095] hour, , .
[0096] Example 2: Given the sparse data task sequence obtained in step 101 as {task 1, task 2, task 3, task 4, task 5}, the number of processors... Number of tasks in a sparse data task sequence The cumulative weight of each cumulative position ( As shown in Table 1, when hour, ;when hour, ; Taking this as an example, the implementation process of step 103-2 will be explained.
[0097] I. For a single processor, that is, the number of processes handled by the processor.
[0098] The handling process for this situation is the same as in Example 1. The processing procedure is the same, and will not be repeated here. Please refer to Example 1. The processing steps are straightforward.
[0099] II. For two processors, i.e., the number of processors processing...
[0100] In this situation .
[0101] Depend on hour ,and Therefore, we can conclude that: 1.
[0102] ,therefore , , .
[0103] but .
[0104] .
[0105] 2.
[0106] ,therefore , , or .
[0107] but .
[0108] .
[0109] 3.
[0110] ,therefore , , or .
[0111] but .
[0112] .
[0113] Right now hour, , .
[0114] hour, , .
[0115] hour, , .
[0116] III. For three processors, i.e., the number of processors processing...
[0117] In this situation .
[0118] Depend on hour ,and Therefore, we can conclude that: 1.
[0119] ,therefore , , .
[0120] but .
[0121] .
[0122] 2.
[0123] ,therefore , , .
[0124] but .
[0125] .
[0126] 3.
[0127] ,therefore , , .
[0128] but .
[0129] .
[0130] Right now hour, , .
[0131] hour, , .
[0132] hour, , .
[0133] This method follows Increasing, in the same Based on the previous state, the search starts from the previous state and searches to the right, thereby reducing the computational cost of candidate segmentation positions.
[0134] Example 3: Given the sparse data task sequence obtained in step 101 as {task 1, task 2, task 3, task 4, task 5}, the number of processors... Number of tasks in a sparse data task sequence The cumulative weight of each cumulative position ( As shown in Table 1, If it exists ,but If it does not exist ,but , and hour, ,or, and hour, for Integer values in, and and Taking this as an example, the implementation process of step 103-2 will be explained.
[0135] I. For a single processor, that is, the number of processes handled by the processor.
[0136] The handling process for this situation is the same as in Example 1. The processing procedure is the same, and will not be repeated here. Please refer to Example 1. The processing steps are straightforward.
[0137] II. For two processors, i.e., the number of processors processing...
[0138] In this situation . .
[0139] Depend on hour ,and Therefore, we can conclude that: 1.
[0140] hour, In other words, ,and That is, satisfying and Then determine ,Right now , .
[0141] but .
[0142] .
[0143] 2.
[0144] hour, In other words, ,and If it exists , Only for Integer values in, i.e. for The integer value in the value is 1 or 2. The existence of this value is then determined using the following (1) and (2). : (1) When hour, , . ,Right now Not satisfied The conditions, therefore, .
[0145] (2) When hour, , . ,Right now ,satisfy The conditions, therefore, .
[0146] That is, it exists .
[0147] , or .
[0148] but .
[0149] .
[0150] 3.
[0151] hour, In other words, 3, and If it exists , Only for Integer values in, i.e. for The integer value in the value is 1, 2, or 3. The existence of the value is then determined using the following methods (1), (2), and (3). : (1) When hour, , . ,Right now Not satisfied The conditions, therefore, .
[0152] (2) When hour, , . ,Right now Not satisfied The conditions, therefore, .
[0153] (3) When hour, , . ,Right now ,satisfy The conditions, therefore, .
[0154] That is, it exists .
[0155] , or or .
[0156] but .
[0157] .
[0158] Right now hour, , .
[0159] hour, , .
[0160] hour, , .
[0161] III. For three processors, i.e., the number of processors processing...
[0162] In this situation . .
[0163] Depend on hour ,and Therefore, we can conclude that: 1.
[0164] hour, In other words, ,and That is, satisfying and Then determine ,Right now , .
[0165] but .
[0166] .
[0167] 2.
[0168] hour, In other words, ,and If it exists , Only for Integer values in, i.e. for The integer value is 2 or 3. Then, we determine whether it exists by using the following (1), (2), and (3). : (1) When hour, , . ,Right now Not satisfied The conditions, therefore, .
[0169] (2) When hour, , . ,Right now Not satisfied The conditions. Therefore, .
[0170] (3) At this point, it is determined that it does not exist. ,but , or .
[0171] but .
[0172] .
[0173] 3.
[0174] hour, In other words, ,and If it exists , Only for Integer values in, i.e. for The integer value in the value is 2, 3, or 4. The existence of this value is then determined using the following methods (1), (2), and (3). : (1) When hour, , . ,Right now Not satisfied The conditions, therefore, .
[0175] (2) When hour, , . ,Right now Not satisfied The conditions, therefore, .
[0176] (3) When hour, , . ,Right now ,satisfy The conditions, therefore, .
[0177] That is, it exists .
[0178] , or or .
[0179] but .
[0180] .
[0181] Right now hour, , .
[0182] hour, , .
[0183] hour, , .
[0184] Example 4: Given the sparse data task sequence obtained in step 101 as {task 1, task 2, task 3, task 4, task 5}, the number of processors... Number of tasks in a sparse data task sequence The cumulative weight of each cumulative position ( As shown in Table 1, when hour, ;when hour, If it exists ,but If it does not exist ,but , and hour, ,or, and hour, for Integer values in, and and Taking this as an example, the implementation process of step 103-2 will be explained.
[0185] I. For a single processor, that is, the number of processes handled by the processor.
[0186] The handling process for this situation is the same as in Example 1. The processing procedure is the same, and will not be repeated here. Please refer to Example 1. The processing steps are straightforward.
[0187] II. For two processors, i.e., the number of processors processing...
[0188] In this situation .
[0189] Depend on hour ,and Therefore, we can conclude that: 1.
[0190] ,therefore .
[0191] In other words, ,and That is, satisfying and Then determine ,Right now , .
[0192] but .
[0193] .
[0194] 2.
[0195] ,therefore .
[0196] In other words, ,and If it exists , Only for Integer values in, i.e. for The integer value in the value is 1 or 2. The existence of this value is then determined using the following (1) and (2). : (1) When hour, , . ,Right now Not satisfied The conditions, therefore, .
[0197] (2) When hour, , . ,Right now ,satisfy The conditions, therefore, .
[0198] That is, it exists .
[0199] , or .
[0200] but .
[0201] .
[0202] 3.
[0203] ,therefore .
[0204] In other words, 3, and If it exists , Only for Integer values in, i.e. for The integer value is 2 or 3. Then, the existence of the integer value is determined by the following (1) and (2). : (1) When hour, , . ,Right now Not satisfied The conditions, therefore, .
[0205] (2) When hour, , . ,Right now ,satisfy The conditions, therefore, .
[0206] That is, it exists .
[0207] , or .
[0208] but .
[0209] .
[0210] Right now hour, , .
[0211] hour, , .
[0212] hour, , .
[0213] III. For three processors, i.e., the number of processors processing...
[0214] In this situation .
[0215] Depend on hour ,and Therefore, we can conclude that: 1.
[0216] ,therefore .
[0217] In other words, ,and That is, satisfying and Then determine ,Right now , .
[0218] but .
[0219] .
[0220] 2.
[0221] ,therefore .
[0222] In other words, 3, and If it exists , Only for Integer values in, i.e. for The integer value is 2 or 3. Then, the existence of the integer value is determined by the following (1) and (2). : (1) When hour, , . ,Right now Not satisfied The conditions, therefore, .
[0223] (2) When hour, , . ,Right now Not satisfied The conditions, therefore, .
[0224] At this time, there are no conditions that are met. Therefore , or .
[0225] but
[0226]
[0227] 3.
[0228] ,therefore .
[0229] In other words, ,and If it exists , Only for Integer values in, i.e. for The integer value is 3 or 4. At this point, we determine whether it exists using the following (1) and (2). : (1) When hour, , . ,Right now Not satisfied The conditions, therefore, .
[0230] (2) When hour, , . ,Right now ,satisfy The conditions, therefore, .
[0231] That is, it exists
[0232] , or .
[0233] but .
[0234] .
[0235] Right now hour, , .
[0236] hour, , .
[0237] hour, , .
[0238] Step 103-2 abandons the "greedy" approach that only considers the current optimal value. Instead, it uses a state transition equation to simultaneously weigh the accumulated bottleneck of the preceding processor against the immediate load of the current processor in each splitting decision. This ensures that a global load-balanced solution is obtained based on the Min-Max model without disrupting the original order of the data.
[0239] 103-3, generate a segmentation position matrix for all segmentation positions.
[0240] In step 103-3, during implementation, storage space can be allocated on the host side to store all... Construct a state transition matrix, and put all of them into a state transition matrix. The resulting segmentation location matrix is stored in the storage space allocated on the host side.
[0241] For example, initialize one state transition matrix and one segmentation position matrix, where both the state transition matrix and the segmentation position matrix are... OK The column matrix, the initialized state transition matrix and the segmentation position matrix are empty.
[0242] Then all the results obtained in step 103-2 Store all the states obtained in step 103-2 into the initialized state transition matrix. The data is stored in the initialized segmentation position matrix (the positions where no data is stored are still empty). Then, the stored state transition matrix (as shown in Table 2) and the stored segmentation position matrix (as shown in Table 3) are both stored in the storage space allocated on the Host side.
[0243] Table 2
[0244] Table 3
[0245] State transition matrix Used to record the first [item] in the sparse data task sequence. Tasks are continuously assigned to The globally optimal bottleneck value achievable with one processor. ; Segmentation position matrix Used to record the optimal segmentation position corresponding to obtaining the globally optimal bottleneck value. This is so that the segmentation position can be backtracked according to the segmentation position matrix later.
[0246] 104. Determine the segmentation position based on the segmentation position matrix.
[0247] The implementation process of step 104 is as follows: 104-1, Initialize the current line identifier Current column identifier .
[0248] 104-2, divide the position matrix into... OK The value of the column is determined as the current position.
[0249] 104-3, Update .
[0250] 104-4, if Then update At the current position, repeating the process will segment the position matrix. OK The step of determining the column value as the split position (i.e., step 104-2) and subsequent steps, until... .
[0251] 104-5, if Then all current positions will be determined as the dividing positions.
[0252] In terms of the number of processors Number of tasks in a sparse data task sequence The segmentation position matrix is shown in Table 3. The following process is then executed: 1. Initialize the current row identifier through step 104-1. Current column identifier .
[0253] 2. Through step 104-2, determine the value 4 in the 5th row and 3rd column of the segmentation position matrix shown in Table 3 as the current position.
[0254] At this point, there is only one current position of 4, which is the end cut point. This means that the last cut position is determined to be after the 4th task, which means that the 3rd processor independently undertakes task 5.
[0255] 3. Update via step 104-3. .
[0256] 4. Then it is necessary to determine how the remaining four tasks should be allocated to the first two processors. This is done through step 104-4. Current position (i.e.) Repeat step 104-2 to determine the current position as the value 3 in the 4th row and 2nd column of the segmentation position matrix shown in Table 3.
[0257] At this point, there are two current positions, 4 and 3. Position 3 indicates that the first segmentation position in the preceding sequence is after the third task, meaning the first processor handles the first three tasks (task 1, task 2, task 3), and the second processor handles task 4. 5. Update via step 104-3. .
[0258] 6. Then, in step 104-5, all current positions are determined as segmentation positions.
[0259] That is, the cutting positions are 4 and 3.
[0260] Step 104 enables the backtracking of the segmentation positions in the segmentation position matrix, i.e., from... Initially, a reverse tracing method was used, that is, based on the index of the current task position. and processor number Find its optimal preorder position for segmentation. Loop backtracking until all are obtained. The optimal segmentation position.
[0261] 105. Distribute the task load across processors in the sparse data task sequence according to the partition position.
[0262] Taking a sparse data task sequence of {task 1, task 2, task 3, task 4, task 5}, with processors 1, 2, and 3, and split positions 4 and 3 as an example, the sequence number in the sparse data task sequence is 4 (i.e., If the task is task 4, then the tasks after task 4 in the sparse data task sequence constitute a task segment (i.e., {task 5}), and the remaining tasks are {task 1, task 2, task 3, task 4}.
[0263] The sequence number in the sparse data task sequence is 3 (i.e. If the task is task 3, then the tasks after task 3 in the remaining tasks constitute a task segment (i.e., {task 4}), and the remaining tasks constitute another task segment {task 1, task 2, task 3}.
[0264] In this way, the sparse data task sequence is continuously divided into three task segments, and these three segments are assigned to three processors respectively, so that the communication load borne by each processor is globally load-balanced. For example, task 1, task 2, and task 3 are load-balanced to processor 1, task 4 to processor 2, and task 5 to processor 3. The host generates a DMA transfer instruction table for each processor according to the segmentation position and drives DMA (Direct Memory Access) to perform continuous data transfer according to the corresponding transfer instruction table.
[0265] For example, the host side restores the globally optimal allocation scheme as: [10,20,30]|
[40] |
[50] . The actual communication load of each node is 60, 40, and 50 respectively, and the maximum bottleneck is limited to 60, eliminating the "barrel effect". The host side generates a DMA transfer instruction table based on this split point, driving the underlying hardware to execute efficiently and concurrently.
[0266] The sparse data load balancing method provided in this embodiment is a dynamic programming-based method for sparse data task allocation and load balancing, which can be used for parallel scheduling of sparse data on multi-core processors. This sparse data load balancing method can accurately find the optimal cut-off point of the task sequence without disrupting the physical memory order, thereby eliminating the "weakest link" effect in parallel computing and maximizing the overall system throughput.
[0267] The sparse data load balancing method provided in this embodiment breaks through the limitations of traditional scheduling and can achieve optimal load balancing. Specifically, the sparse data load balancing method in this embodiment models the complex on-chip and off-chip bus transmissions as cumulative weights. Under the constraint of maintaining data linearity and continuity, it ensures that the state transition process can obtain a globally optimal load allocation strategy, overcoming the problems of idle computing power and bus congestion caused by non-uniform data distribution.
[0268] The sparse data load balancing method provided in this embodiment eliminates search redundancy and achieves rapid planning. Specifically, addressing the high computational complexity of traditional dynamic programming algorithms under large-scale tasks, this method delves into the intersection of the "monotonically increasing cumulative bottleneck" and the "monotonically decreasing newly added load," introducing a monotonic pointer mechanism in the state transition. This effectively eliminates the traversal of invalid states, enabling rapid calculation of the task allocation table. This mechanism eliminates the need to reset the optimization pointer each time, effectively eliminating redundant state traversal in inner optimization and significantly reducing planning latency.
[0269] The sparse data load balancing method provided in this embodiment offers a fast, multi-core load balancing underlying scheduling method for scenarios involving sparse data 3D-FFT computation, such as quantum chemistry, significantly improving the overall energy efficiency of the hardware platform.
[0270] The sparse data load balancing method provided in this embodiment is more adapted to the underlying hardware architecture, adapting to the discrete characteristics of non-uniform data while ensuring memory continuity constraints. It can obtain the globally optimal load distribution strategy under hardware constraints that strictly prohibit disrupting the order of memory data; it can eliminate the "weakest link" effect caused by data space non-uniformity in multi-core collaborative computing, minimizing maximum communication latency; and it can eliminate redundant searches caused by solving for the global optimum on the resource-constrained host, completing boundary partitioning of large-scale tasks with extremely low time complexity.
[0271] The sparse data load balancing method provided in this embodiment uses Min-Max dynamic programming to accurately solve for the optimal split point of the task sequence while ensuring the physical continuity of data storage, thereby completely eliminating the "barrel effect" in parallel computing.
[0272] The sparse data load balancing method provided in this embodiment combines dynamic programming algorithm with hardware characteristics to achieve fast and optimal task allocation and load balancing when performing 3D-FFT calculations on large-scale sparse data.
[0273] Compared with the traditional 3D-FFT scheduling method, the advantages of the sparse data load balancing method provided in this embodiment are shown in Table 4.
[0274] Table 4
[0275] The sparse data load balancing method provided in this embodiment combines dynamic programming algorithm with the physical constraint of continuous memory access, providing a robust and efficient parallel scheduling mechanism with high practical value in non-uniform computing scenarios such as large-scale 3D-FFT.
[0276] This embodiment provides a sparse data load balancing method. It obtains a sparse data task sequence; determines the communication load weight of each task based on the effective non-zero data volume and hardware communication cost parameters; generates a segmentation position matrix based on the number of processors, the number of tasks in the sparse data task sequence, and the communication load weight of each task; determines the segmentation position based on the segmentation position matrix; and balances the task load in the sparse data task sequence across processors according to the segmentation position. This method can accurately find the optimal segmentation point of the task sequence and perform load balancing without disrupting the physical memory order, thereby eliminating the "weakest link" effect in parallel computing and maximizing the overall system throughput.
[0277] Based on the same inventive concept as the above method, this embodiment provides a device, see [link to device]. Figure 2 The device includes: The acquisition module 201 is used to acquire sparse data task sequences.
[0278] The first determining module 202 is used to determine the communication load weight of each task based on the effective non-zero data volume contained in each task in the sparse data task sequence and the hardware communication cost parameters.
[0279] Generation module 203 is used to generate based on the number of processors Number of tasks in a sparse data task sequence The communication load weights of each task are used to generate a segmentation location matrix. The generated segmentation location matrix is as follows: OK Column matrix.
[0280] The second determining module 204 is used to determine the segmentation position based on the segmentation position matrix.
[0281] The load balancing module 205 is used to load balance the tasks in the sparse data task sequence to each processor according to the partition position.
[0282] The generation module 203 is used to determine the cumulative weight of each cumulative position in the sparse data task sequence. The cumulative position in the sparse data task sequence is... Cumulative weights , To accumulate positions, , For the summation index, , For the sparse data task sequence, the first Communication load weights for each task.
[0283] Based on the cumulative weights of each cumulative position in the sparse data task sequence, determine the first position in the sparse data task sequence. Tasks assigned to The segmentation location of each processor. Number of processors , This represents the sequence number of the task within the sparse data task sequence. .
[0284] Generate a segmentation position matrix from all the segmentation positions.
[0285] Specifically, based on the cumulative weight of each cumulative position in the sparse data task sequence, the first position in the sparse data task sequence is determined. Tasks assigned to The segmentation locations of each processor include: like Then determine the first part of the sparse data task sequence. Tasks assigned to The segmentation position of each processor ,and .in, To make the first part of the sparse data task sequence Tasks assigned to The globally optimal bottleneck value for each processor.
[0286] like Then determine the first part of the sparse data task sequence. Tasks assigned to The segmentation position of each processor ,and .in, This refers to the sequence number of the segmentation location within the sparse data task sequence. To find the function with the maximum value, To make the function reach its minimum value, the corresponding value, To make the first part of the sparse data task sequence Tasks assigned to The globally optimal bottleneck value for each processor for The minimum value, for The maximum value that can be obtained.
[0287] in, Or, when hour ,when hour, .
[0288] in, Or, if it exists ,but If it does not exist ,but .
[0289] in, for The effective value, and hour, ,or, and hour, for Integer values in, and and .
[0290] Determining the segmentation position based on the segmentation position matrix includes: Initialize the current line identifier Current column identifier .
[0291] In the segmentation position matrix OK The value of the column is determined as the current position.
[0292] renew .
[0293] like Then update At the current position, repeating the process will segment the position matrix. OK The steps for determining the column value as the split position and subsequent steps, until... .
[0294] like Then all current positions will be determined as the dividing positions.
[0295] The device provided in this embodiment can accurately find the optimal cutting point of the task sequence and perform load balancing under the premise that the physical memory order is strictly prohibited from being disrupted, thereby eliminating the "barrel effect" in parallel computing and maximizing the overall throughput of the system.
[0296] Based on the same inventive concept as the above method, this embodiment provides an electronic device, which is as follows: Figure 3 As shown, it includes: a memory 301, a processor 302, and a computer program.
[0297] The computer program is stored in memory 301 and configured to be executed by processor 302 to implement the above method.
[0298] Specifically, Sequence of tasks for obtaining sparse data.
[0299] The communication load weight of each task is determined based on the amount of effective non-zero data contained in each task in the sparse data task sequence and the hardware communication cost parameters.
[0300] Based on the number of processors Number of tasks in a sparse data task sequence The communication load weights of each task are used to generate a segmentation location matrix. The generated segmentation location matrix is as follows: OK Column matrix.
[0301] The segmentation position is determined based on the segmentation position matrix.
[0302] The task load in the sparse data task sequence is balanced across the processors according to the partition position.
[0303] Among them, based on the number of processors Number of tasks in a sparse data task sequence The communication load weights for each task generate a segmentation location matrix, including: Determine the cumulative weight of each cumulative position in the sparse data task sequence. Wherein, the cumulative position in the sparse data task sequence is... Cumulative weights , To accumulate positions, , For the summation index, , For the sparse data task sequence, the first Communication load weights for each task.
[0304] Based on the cumulative weights of each cumulative position in the sparse data task sequence, determine the first position in the sparse data task sequence. Tasks assigned to The segmentation location of each processor. Number of processors , This represents the sequence number of the task within the sparse data task sequence. .
[0305] Generate a segmentation position matrix from all the segmentation positions.
[0306] Specifically, based on the cumulative weight of each cumulative position in the sparse data task sequence, the first position in the sparse data task sequence is determined. Tasks assigned to The segmentation locations of each processor include: like Then determine the first part of the sparse data task sequence. Tasks assigned to The segmentation position of each processor ,and .in, To make the first part of the sparse data task sequence Tasks assigned to The globally optimal bottleneck value for each processor.
[0307] like Then determine the first part of the sparse data task sequence. Tasks assigned to The segmentation position of each processor ,and .in, This refers to the sequence number of the segmentation location within the sparse data task sequence. To find the function with the maximum value, To make the function reach its minimum value, the corresponding value, To make the first part of the sparse data task sequence Tasks assigned to The globally optimal bottleneck value for each processor for The minimum value, for The maximum value that can be obtained.
[0308] in, Or, when hour, ;when hour, .
[0309] in, Or, if it exists ,but If it does not exist ,but .
[0310] in, for The effective value, and hour, ,or, and hour, for Integer values in, and and .
[0311] Determining the segmentation position based on the segmentation position matrix includes: Initialize the current line identifier Current column identifier .
[0312] In the segmentation position matrix OK The value of the column is determined as the current position.
[0313] renew .
[0314] like Then update At the current position, repeating the process will segment the position matrix. OK The steps for determining the column value as the split position and subsequent steps, until... .
[0315] like Then all current positions will be determined as the dividing positions.
[0316] The electronic device provided in this embodiment has a computer program executed by a processor to accurately find the optimal cutting point of the task sequence and perform load balancing under the premise of strictly prohibiting the disruption of the physical memory order, thereby eliminating the "barrel effect" in parallel computing and maximizing the overall throughput of the system.
[0317] Based on the same inventive concept as the above method, this embodiment provides a computer-readable storage medium on which a computer program is stored. The computer program is executed by a processor to implement the above method.
[0318] Specifically, Sequence of tasks for obtaining sparse data.
[0319] The communication load weight of each task is determined based on the amount of effective non-zero data contained in each task in the sparse data task sequence and the hardware communication cost parameters.
[0320] Based on the number of processors Number of tasks in a sparse data task sequence The communication load weights of each task are used to generate a segmentation location matrix. The generated segmentation location matrix is as follows: OK Column matrix.
[0321] The segmentation position is determined based on the segmentation position matrix.
[0322] The task load in the sparse data task sequence is balanced across the processors according to the partition position.
[0323] Among them, based on the number of processors Number of tasks in a sparse data task sequence The communication load weights for each task generate a segmentation location matrix, including: Determine the cumulative weight of each cumulative position in the sparse data task sequence. Wherein, the cumulative position in the sparse data task sequence is... Cumulative weights , To accumulate positions, , For the summation index, , For the sparse data task sequence, the first Communication load weights for each task.
[0324] Based on the cumulative weights of each cumulative position in the sparse data task sequence, determine the first position in the sparse data task sequence. Tasks assigned to The segmentation location of each processor. Number of processors , This represents the sequence number of the task within the sparse data task sequence. .
[0325] Generate a segmentation position matrix from all the segmentation positions.
[0326] Specifically, based on the cumulative weight of each cumulative position in the sparse data task sequence, the first position in the sparse data task sequence is determined. Tasks assigned to The segmentation locations of each processor include: like Then determine the first part of the sparse data task sequence. Tasks assigned to The segmentation position of each processor ,and .in, To make the first part of the sparse data task sequence Tasks assigned to The globally optimal bottleneck value for each processor.
[0327] like Then determine the first part of the sparse data task sequence. Tasks assigned to The segmentation position of each processor ,and .in, This refers to the sequence number of the segmentation location within the sparse data task sequence. To find the function with the maximum value, To make the function reach its minimum value, the corresponding value, To make the first part of the sparse data task sequence Tasks assigned to The globally optimal bottleneck value for each processor for The minimum value, for The maximum value that can be obtained.
[0328] in, Or, when hour, ;when hour, .
[0329] in, Or, if it exists ,but If it does not exist ,but .
[0330] in, for The effective value, and hour, ,or, and hour, for Integer values in, and and .
[0331] Determining the segmentation position based on the segmentation position matrix includes: Initialize the current line identifier Current column identifier .
[0332] In the segmentation position matrix OK The value of the column is determined as the current position.
[0333] renew .
[0334] like Then update At the current position, repeating the process will segment the position matrix. OK The steps for determining the column value as the split position and subsequent steps, until... .
[0335] like Then all current positions will be determined as the dividing positions.
[0336] The computer-readable storage medium provided in this embodiment allows the computer program thereon to be executed by a processor to accurately find the optimal cut-off point of the task sequence and perform load balancing, under the premise of strictly prohibiting disruption of the physical memory order, thereby eliminating the "barrel effect" in parallel computing and maximizing the overall throughput of the system.
[0337] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code. The solutions in the embodiments of this application can be implemented in various computer languages, such as the object-oriented programming language Java and the interpreted scripting language JavaScript.
[0338] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0339] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0340] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0341] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of this application, "multiple" means at least two, such as two, three, etc., unless otherwise explicitly specified.
[0342] Although preferred embodiments of this application have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of this application.
[0343] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application. Therefore, if such modifications and variations fall within the scope of the claims of this application and their equivalents, this application also intends to include such modifications and variations.
Claims
1. A sparse data load balancing method, characterized in that, The method includes: Sequence of tasks for acquiring sparse data; Based on the effective non-zero data volume and hardware communication cost parameters contained in each task in the sparse data task sequence, the communication load weight of each task is determined. Based on the number of processors The number of tasks in the sparse data task sequence The communication load weights of each task are used to generate a segmentation location matrix; wherein, the generated segmentation location matrix is... OK Column matrix; The segmentation position is determined based on the segmentation position matrix; The task load in the sparse data task sequence is balanced among the processors according to the segmentation position.
2. The method according to claim 1, characterized in that, According to the number of processors The number of tasks in the sparse data task sequence The communication load weights for each task generate a segmentation location matrix, including: Determine the cumulative weight of each cumulative position in the sparse data task sequence; wherein, the cumulative position in the sparse data task sequence is... Cumulative weights , To accumulate positions, , For the summation index, , For the sparse data task sequence, the first... Communication load weights for each task; Based on the cumulative weights of each cumulative position in the sparse data task sequence, determine the first position in the sparse data task sequence. Tasks assigned to The segmentation location of each processor; where, Number of processors , Let be the sequence number of the task within the sparse data task sequence. ; Generate a segmentation position matrix from all the segmentation positions.
3. The method according to claim 2, characterized in that, The step involves determining the cumulative weight of each cumulative position in the sparse data task sequence, based on the cumulative weight of each position. Tasks assigned to The segmentation locations of each processor include: like Then determine the first part of the sparse data task sequence. Tasks assigned to The segmentation position of each processor ,and ;in, To make the first part of the sparse data task sequence Tasks assigned to The globally optimal bottleneck value for each processor; like Then determine the first part of the sparse data task sequence. Tasks assigned to The segmentation position of each processor ,and ;in, The segmentation position is the sequence number corresponding to the location in the sparse data task sequence. To find the function with the maximum value, To make the function reach its minimum value, the corresponding value, To make the first part of the sparse data task sequence Tasks assigned to The globally optimal bottleneck value for each processor for The minimum value, for The maximum value that can be obtained.
4. The method according to claim 3, characterized in that, Or, when hour ,when hour, .
5. The method according to claim 3, characterized in that, Or, if it exists ,but If it does not exist ,but ; in, for The effective value, and hour, ,or, and hour, for Integer values in, and and .
6. The method according to claim 1, characterized in that, Determining the segmentation position based on the segmentation position matrix includes: Initialize the current line identifier Current column identifier ; In the segmentation position matrix OK The value of the column is determined by the current position; renew ; like Then update At the current position, repeat the process of segmenting the position matrix. OK The steps for determining the column value as the split position and subsequent steps, until... ; like Then all current positions will be determined as the dividing positions.
7. A sparse data load balancing device, characterized in that, The device includes: The acquisition module is used to acquire sparse data task sequences; The first determining module is used to determine the communication load weight of each task based on the effective non-zero data volume contained in each task in the sparse data task sequence and the hardware communication cost parameters. The generation module is used to determine the number of processors. The number of tasks in the sparse data task sequence The communication load weights of each task are used to generate a segmentation location matrix; wherein, the generated segmentation location matrix is... OK Column matrix; The second determining module is used to determine the segmentation position based on the segmentation position matrix; A load balancing module is used to distribute the task load of the sparse data task sequence to each processor according to the segmentation position.
8. The apparatus according to claim 7, characterized in that, The generation module is used to determine the cumulative weight of each cumulative position in the sparse data task sequence; wherein, the cumulative position in the sparse data task sequence is... Cumulative weights , To accumulate positions, , For the summation index, , For the sparse data task sequence, the first... Communication load weights for each task; Based on the cumulative weights of each cumulative position in the sparse data task sequence, determine the first position in the sparse data task sequence. Tasks assigned to The segmentation location of each processor; where, Number of processors , Let be the sequence number of the task within the sparse data task sequence. ; Generate a segmentation position matrix from all the segmentation positions.
9. The apparatus according to claim 8, characterized in that, The step involves determining the cumulative weight of each cumulative position in the sparse data task sequence, based on the cumulative weight of each position. Tasks assigned to The segmentation locations of each processor include: like Then determine the first part of the sparse data task sequence. Tasks assigned to The segmentation position of each processor ,and ;in, To make the first part of the sparse data task sequence Tasks assigned to The globally optimal bottleneck value for each processor; like Then determine the first part of the sparse data task sequence. Tasks assigned to The segmentation position of each processor ,and ;in, The segmentation position is the sequence number corresponding to the location in the sparse data task sequence. To find the function with the maximum value, To make the function reach its minimum value, the corresponding value, To make the first part of the sparse data task sequence Tasks assigned to The globally optimal bottleneck value for each processor for The minimum value, for The maximum value that can be obtained.
10. The apparatus according to claim 9, characterized in that, Or, when hour ,when hour .
11. The apparatus according to claim 9, characterized in that, Or, if it exists ,but If it does not exist ,but ; in, for The effective value, and hour, ,or, and hour, for Integer values in, and and .
12. The apparatus according to claim 7, characterized in that, Determining the segmentation position based on the segmentation position matrix includes: Initialize the current line identifier Current column identifier ; In the segmentation position matrix OK The value of the column is determined by the current position; renew ; like Then update At the current position, repeat the process of segmenting the position matrix. OK The steps for determining the column value as the split position and subsequent steps, until... ; like Then all current positions will be determined as the dividing positions.
13. An electronic device, characterized in that, include: Memory; processor; as well as Computer programs; The computer program is stored in the memory and configured to be executed by the processor to implement the method as described in any one of claims 1-6.
14. A computer-readable storage medium, characterized in that, It stores a computer program thereon; the computer program is executed by a processor to implement the method as described in any one of claims 1-6.