Data processing method and related device
By converting sparse tensors into sequences and distributing them to multiple computing devices based on their sparsity characteristics, the problem of unbalanced load in distributed computing with sparse tensors is solved, achieving more efficient resource utilization.
Patent Information
- Application Number
- CN202410509420.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-04-25
- Publication Date
- 2025-10-28
AI Technical Summary
Sparse tensors are difficult to split in distributed computing, have large differences in computation time, and are difficult to merge. Existing solutions lead to unbalanced computing load and low resource utilization.
The sparse tensor is transformed into a sparse tensor sequence, the computational load is determined based on the sparsity characteristics, and the load is divided into multiple computing devices by optimizing the grouping algorithm, thereby achieving load-balanced distributed computing.
The load balancing of computing devices has been optimized, reducing computation time and improving the utilization of computing resources.
Smart Images

Figure CN120849074A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence, and more particularly to a data processing method and related apparatus. Background Technology
[0002] Sparse tensors are an efficient data structure for handling large-scale sparse data, widely used in recommender systems, machine learning, and data analysis. The number of non-zero elements in a sparse tensor is much smaller than the total number of elements in the tensor, making it suitable for representing high-dimensional arrays where most elements are zero. Meanwhile, with the rapid development of deep learning, the amount of training data is increasing, and model sizes are growing larger. Traditional single-machine training methods can no longer meet performance and time requirements; therefore, distributed training has gradually become one of the most popular research directions in this field.
[0003] Due to the irregularity and dynamic shape of sparse tensors, as well as the imperfect support for sparse tensors in deep learning frameworks and libraries, sparse tensors face challenges in distributed computing, including difficulty in splitting them before computation, large differences in computation time, and difficulty in merging them after computation. To address these challenges and enable sparse tensors to better integrate with deep learning frameworks and achieve distributed parallel computing capabilities, this paper proposes a sparse distributed parallel scheme based on the JAX deep learning framework. This scheme primarily addresses the difficulty in splitting sparse tensors in distributed computing by adding a batch dimension to the coordinate format (COO) storage method of the sparse tensors, and solves the problem of data misalignment during data merging by padding with zero values.
[0004] However, existing solutions introduce an explicit storage parameter for the number of specified elements (nse) when addressing the difficulty of splitting sparse tensors, and require nse alignment to construct sparse tensors adapted to distributed computing. In addition, existing solutions divide sparse tensors based on the batch dimension and the sparse tensor index order, which may lead to uneven computing load on devices in scenarios with high computational demands. Summary of the Invention
[0005] This application provides a data processing method and related apparatus for grouping multiple sparse tensors according to their computational load, and for allocating the multiple sparse tensors to multiple computing devices according to the grouping results, so as to realize distributed computing of sparse tensors and improve the utilization of computing resources.
[0006] In view of this, in a first aspect, this application provides a data processing method, comprising: first, converting N input sparse tensors into N sparse tensor sequences; determining the computational load of each sparse tensor sequence based on the sparsity characteristics of each sparse tensor, wherein the computational load indicates the amount of computation required when each sparse tensor sequence is executed by a computing device; subsequently, dividing the N sparse tensor sequences into M groups based on the number of computing devices M and the computational load of each sparse tensor sequence, wherein each group includes at least one sparse tensor sequence; and sending the M groups of sparse tensor sequences to the M computing devices respectively to run an artificial intelligence model, thereby realizing distributed computation of sparse tensors.
[0007] In this embodiment of the application, the input sparse tensor can be transformed into a sparse tensor sequence so that the subsequent partitioning of the sparse tensor can be performed using the sparse tensor as the data granularity. This avoids splitting the data stored in the sparse tensor and also avoids adding the batch dimension when partitioning the sparse tensor, thereby avoiding the limitations brought by high-dimensional sparse tensors with batch dimension.
[0008] In one possible implementation, the difference in computational load between the aforementioned M groups is the smallest among all schemes for dividing the N sparse tensor sequences into the M groups.
[0009] In this embodiment, multiple sparse tensor sequences can be divided according to the computational load of each sparse tensor, so that the sum of the optimization indices of the sparse tensors obtained by each computing device is as close as possible, thereby balancing the computational load of each computing device, reducing computation time, and improving the utilization rate of computing resources of each computing device.
[0010] In one possible implementation, the aforementioned determination of the computational load of each sparse tensor sequence based on the sparsity characteristics of each sparse tensor may include: determining the computational load of each sparse tensor sequence based on the number of non-zero elements in each sparse tensor.
[0011] In one possible implementation, the aforementioned determination of the computational load of each sparse tensor sequence based on the sparsity characteristics of each sparse tensor may include: determining the computational load of each sparse tensor sequence based on the sparsity uniformity of each sparse tensor, where the sparsity uniformity represents the degree of uniformity of the distribution of non-zero elements in each sparse tensor.
[0012] In one possible implementation, the method may further include: after allocating M sets of sparse tensor sequences to M computing devices, recording the sparse tensor sequence allocated to each computing device;
[0013] After M computing devices complete the calculations on N sparse tensor sequences and obtain the calculation results, the calculation result corresponding to each sparse tensor sequence is obtained from the calculation results based on the records.
[0014] In one possible implementation, the method may further include: identifying operators in the artificial intelligence model that require parallel computation based on user-input computation operator identifiers; and deploying the operators in M computing devices.
[0015] In this embodiment, the part of the sparse tensor calculation process that needs to be parallelized can be determined according to the calculation function, thereby determining the operation nodes that need to be parallelized. The sparse tensor can be distributed and distributed only when parallelization is required, so as to reduce the communication cost brought about by distributed computing.
[0016] In one possible implementation, the method may further include: obtaining the shapes of the computation results of M computing devices through the allgather communication operator to obtain a shape sequence; and aggregating the computation results corresponding to each computing device according to the shape sequence through the allgatherv communication operator.
[0017] In this embodiment, the shape of the computation result of each computing device can be obtained through the allgather communication operator to estimate the amount of data of each computing device. Thus, the computation result of each computing device can be aggregated using the allgatherv communication operator to achieve the aggregation of sparse tensors of different shapes.
[0018] Secondly, this application provides a data processing apparatus, comprising:
[0019] The processing module is used to transform N input sparse tensors into a sequence of N sparse tensors;
[0020] The determination module is used to determine the computational load of each sparse tensor sequence based on the sparsity characteristics of each sparse tensor. The computational load indicates the amount of computation required when each sparse tensor sequence is executed by a computing device.
[0021] The partitioning module is used to divide N sparse tensor sequences into M groups based on the number of computing devices M and the computational load of each sparse tensor sequence, with each group including at least one sparse tensor sequence.
[0022] The sending module is used to send M sets of sparse tensor sequences to M computing devices to run artificial intelligence models.
[0023] In one possible implementation, the difference in computational load between the aforementioned M groups is the smallest among all schemes for dividing the N sparse tensor sequences into the M groups.
[0024] In one possible implementation, the aforementioned determining module is specifically used to: determine the computational load of each sparse tensor sequence based on the number of non-zero elements in each sparse tensor.
[0025] In one possible implementation, the aforementioned determining module is specifically used to: determine the computational load of each sparse tensor sequence based on the sparsity uniformity of each sparse tensor, where sparsity uniformity represents the degree of uniformity of the distribution of non-zero elements in each sparse tensor.
[0026] In one possible implementation, the device may further include:
[0027] The aforementioned processing module is also used to record the sparse tensor sequence allocated to each computing device after allocating the M sets of sparse tensor sequences to the M computing devices.
[0028] The acquisition module is used to retrieve the calculation result corresponding to each sparse tensor sequence from the calculation results after the calculations are completed on N sparse tensor sequences by M computing devices.
[0029] In one possible implementation, the device may further include:
[0030] The aforementioned processing module is also used to identify operators in the artificial intelligence model that require parallel computation based on the computation operator identifiers input by the user;
[0031] The deployment module is used to deploy operators across M computing devices.
[0032] In one possible implementation, the device may further include:
[0033] The aforementioned acquisition module is used to obtain the shapes of the computation results from M computing devices through the allgather communication operator, and obtain a shape sequence;
[0034] The aggregation module is used to aggregate the computation results corresponding to each computing device based on the shape sequence using the allgatherv communication operator.
[0035] Thirdly, this application provides a processor that can be connected to a memory for executing instructions stored in the memory, so that the processor performs the method as described in the first aspect or any possible implementation thereof.
[0036] Fourthly, embodiments of this application provide a computer-readable storage medium. This computer-readable storage medium stores computer instructions; when these computer instructions are executed on a computer, they cause the computer to perform the method as described in any of the possible implementations of the first aspect.
[0037] Fifthly, embodiments of this application provide a computer program product. This computer program product includes a computer program or instructions that, when executed on a computer, cause the computer to perform the method as described in any of the possible implementations of the first aspect.
[0038] The technical effects of the second to fifth aspects or any of their possible implementations can be found in the first aspect or the related possible implementations of the first aspect, and will not be repeated here. Attached Figure Description
[0039] Figure 1 A schematic diagram of a system architecture is provided for this application;
[0040] Figure 2 A flowchart illustrating the distributed computation process for sparse tensors, along with schematics of the relevant core modules.
[0041] Figure 3 A schematic diagram of the implementation process of the solution provided in this application;
[0042] Figure 4 A flowchart illustrating a data processing method provided in this application;
[0043] Figure 5 A diagram illustrating the overall computation time of each computing device after optimization for computational load;
[0044] Figure 6 A schematic diagram of copying mappings for computing nodes;
[0045] Figure 7 A flowchart illustrating the communication aggregation of computation results from various computing devices;
[0046] Figure 8 This is a schematic diagram of the structure of a data processing device provided in this application. Detailed Implementation
[0047] The technical solutions of the embodiments of this application will now be described with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application.
[0048] The method provided in this application can be applied to artificial intelligence (AI) scenarios. AI is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to obtain optimal results. In other words, artificial intelligence is a branch of computer science that attempts to understand the essence of intelligence and produce a new kind of intelligent machine that can react in a way similar to human intelligence. Artificial intelligence studies the design principles and implementation methods of various intelligent machines, enabling them to have the functions of perception, reasoning, and decision-making. Research in the field of artificial intelligence includes robotics, natural language processing, computer vision, decision-making and reasoning, human-computer interaction, recommendation and search, and fundamental AI theories.
[0049] First, the overall workflow of an artificial intelligence system is described. The following sections elaborate on the aforementioned AI framework from two dimensions: the "Intelligent Information Chain" (horizontal axis) and the "IT Value Chain" (vertical axis). The "Intelligent Information Chain" reflects a series of processes from data acquisition to processing. For example, it could be the general process of intelligent information perception, intelligent information representation and formation, intelligent reasoning, intelligent decision-making, and intelligent execution and output. In this process, data undergoes a condensation process of "data—information—knowledge—wisdom." The "IT Value Chain" reflects the value that artificial intelligence brings to the information technology industry, from the underlying infrastructure of human intelligence and information (provided and processed by technology) to the industrial ecosystem of the system.
[0050] (1) Infrastructure
[0051] Infrastructure provides computing power to support artificial intelligence systems, enabling communication with the external world and providing support through a basic platform. This communication occurs through sensors; computing power is provided by intelligent chips (hardware acceleration chips such as CPUs, NPUs, GPUs, ASICs, and FPGAs); and the basic platform includes distributed computing frameworks and related platform guarantees and support, which may include cloud storage and computing, interconnected networks, etc. For example, sensors communicate with the outside world to acquire data, and this data is provided to intelligent chips in the distributed computing system provided by the basic platform for computation.
[0052] (2) Data
[0053] The data at the next layer of infrastructure is used to represent the data sources in the field of artificial intelligence. The data involves graphics, images, voice, text, and IoT data from traditional devices, including business data from existing systems and sensor data such as force, displacement, liquid level, temperature, and humidity.
[0054] (3) Data processing
[0055] Data processing typically includes methods such as data training, machine learning, deep learning, search, reasoning, and decision-making.
[0056] Among them, machine learning and deep learning can perform intelligent information modeling, extraction, preprocessing, and training on data, including symbolization and formalization.
[0057] Reasoning refers to the process in which, in a computer or intelligent system, the machine thinks and solves problems by simulating human intelligent reasoning, based on reasoning control strategies and using formalized information. Typical functions include search and matching.
[0058] Decision-making refers to the process of making decisions based on intelligent information after reasoning, and it typically provides functions such as classification, sorting, and prediction.
[0059] (4) General ability
[0060] After the data processing mentioned above, the results of the data processing can be used to form some general capabilities, such as algorithms or a general system, for example, translation, text analysis, computer vision processing, speech recognition, image recognition, etc.
[0061] (5) Smart Products and Industry Applications
[0062] Intelligent products and industry applications refer to products and applications of artificial intelligence systems in various fields. They are the encapsulation of overall artificial intelligence solutions, productizing intelligent information decision-making and realizing practical applications. Their application areas mainly include: intelligent terminals, intelligent transportation, intelligent healthcare, autonomous driving, smart cities, etc.
[0063] This application involves the application of tensors. In order to better understand the solutions of this application, the relevant terms and concepts of tensors that may be involved in this application will be introduced below.
[0064] (1) Sparse tensor
[0065] A sparse tensor is a tensor in which most of the data elements are 0, non-zero elements make up only a very small part, and it is stored in a sparse format.
[0066] (2) Tensor
[0067] A tensor, also known as a dense tensor, is a multidimensional array that stores all elements according to consecutive indices.
[0068] (3) Sparse matrix
[0069] Two-dimensional sparse tensors are called sparse matrices, which are generally stored in formats such as COO, CSR, CSC, LIL, DIA, BSR, DOK, and SPArray.
[0070] (4) COO format
[0071] The COO format (coordinate format) refers to the way sparse tensors are stored in terms of data, indices, and shape (where shape does not necessarily need to be explicitly passed by the user and can be inferred by the framework). data stores the values of all non-zero elements; indices stores the indices (also known as coordinate positions) of all non-zero elements; and shape represents the shape of the corresponding dense matrix.
[0072] (5) CSR format
[0073] CRS (compressed sparse row format) refers to the way sparse tensors are stored in terms of indptr, indices, data, and shape (where shape does not necessarily need to be explicitly passed by the user and can be inferred by the framework). The size of indptr is the corresponding dense matrix shape[0]+1, which represents the starting position of each row of non-zero elements in indices and data; indices represents the column coordinates corresponding to each non-zero element; data represents the specific value of each non-zero element; and shape represents the shape of the corresponding dense matrix.
[0074] (6)Non-zero dollar quantity
[0075] The number of non-zero elements (nnz) refers to the number of non-zero elements stored in a sparse tensor.
[0076] (7) Display the number of stored elements
[0077] The number of specified elements (nse) is explicitly stored, including the number of non-zero elements and the explicitly stored non-zero elements. Therefore, for the same sparse tensor, nse is greater than nnz.
[0078] The method provided in this application can be applied to distributed computing scenarios of sparse tensors, especially for distributed computing scenarios of sparse tensors that are computationally complex and / or have a large computational load.
[0079] The following describes some possible system architectures provided in the embodiments of this application.
[0080] See Figure 1 This is a schematic diagram of the structure of a distributed computing system 100. (See diagram below.) Figure 1 As shown, the distributed computing system 100 may include a client 110 and a server cluster 120. The client 110 and the server cluster 120 can interact with each other, such as through Hypertext Transfer Protocol (HTTP). The server cluster 120 includes at least one server, which can be a cloud server, a central server, an edge server, or a local server in a local data center.
[0081] Specifically, the processors used for data processing and distribution of sparse tensors can be distributed in any one or more servers in the aforementioned distributed computing system 100. The processors can be central processing units (CPUs), data processing units (DPUs), or other general-purpose processors, etc., without any specific limitation.
[0082] After the processor transforms multiple input sparse tensors into multiple sparse tensor sequences, these sequences can be distributed across multiple computing devices (also known as accelerator cards). These computing devices, used for distributed computation of the multiple sparse tensor sequences, can be located on any one or more servers within the aforementioned distributed computing system 100. The computing devices can be graphics processing units (GPUs), neural network processing units (NPUs), Ascend processors, etc., and are not specifically limited here.
[0083] It is worth noting that, Figure 1 The system architecture shown is merely an example and is not intended to limit its specific implementation to this example. For example, in other possible system architectures, the distributed computing system 100 may not include the client 110, and the sparse tensor may be obtained directly from the network by the servers in the server cluster 120; or, the distributed computing system 100 may be deployed on other types of devices, such as desktop computers or other terminal devices, and is not limited to the servers mentioned above. This embodiment does not limit this.
[0084] In this embodiment, client 110 can obtain sparse tensors and provide them to server cluster 120. Server cluster 120 can then perform distributed parallel computation of the sparse tensors based on the sparse tensors sent by client 110. In deep learning frameworks, tensor computation generally involves five main processes: user code parsing and computation graph establishment, compiler front-end optimization and compiler back-end optimization, and actual computation. The flowchart of tensor computation based on a deep learning framework, the distributed computation flowchart of sparse tensors in this embodiment, and related core modules are as follows: Figure 2 As shown.
[0085] like Figure 2 As shown, the source program decoding module 210 can parse the source program to obtain the sparse tensor sequence and computation function in the program code; then the computation graph generation module 220 can generate a computation graph according to the computation function, so that the computation nodes can be marked according to the computation graph. The computation nodes are the operators that need to be computed in parallel among the input computation operators.
[0086] This embodiment adds a load optimization module 231 and a node replication and mapping module 232 to the compiler front-end optimization module 230. The load optimization module 231 can calculate the optimization index of the sparse tensor based on the acquired sparse tensor sequence, and can group the sparse tensor sequence using an optimization grouping algorithm based on the optimization index to obtain the grouping result, so that the data distribution module 241 can distribute the sparse tensor sequence according to the grouping result; the node replication and mapping module 232 can mark the operation nodes that need to be computed in parallel in the computation graph, and can replicate the operation nodes according to the number of computing devices. The operation nodes can also be pre-deployed in multiple computing devices; specifically, as shown... Figure 3 As shown, marking computation nodes can be done in the custom submodule `definepattern`, while copying and mapping computation nodes can be done in the process submodule.
[0087] The compiler backend optimization module 240 adds a data distribution module 241 and a communication and reordering module 242. After obtaining the grouping result of the sparse tensor sequence, the data distribution module 241 can divide the multiple sparse tensors in the sparse tensor sequence into multiple computing devices according to the grouping result. The communication and reordering module 242 can aggregate the computing results of multiple computing devices and broadcast the aggregated result to multiple computing devices through the allgatherv communication operator so that each computing device has the full computing result. In addition, it can also reorder the aggregated computing results according to the sparse tensor sequence correspondence table output by the load optimization module 231 so that the order of the obtained aggregated computing results is consistent with the user input.
[0088] For distributed parallelism of sparse tensors, an existing solution has proposed a distributed parallelism scheme based on the JAX deep learning framework. However, this scheme is not applicable to distributed computing scenarios with increased NSE, and requires NSE alignment during aggregation after distributed computing, which makes the scheme quite restrictive. In addition, the scheme suffers from unbalanced load in high computing scenarios, resulting in low resource utilization of computing devices.
[0089] To address the current problem, this application proposes a data processing method that can group multiple sparse tensors based on the computational load of each sparse tensor. Subsequently, based on the grouping results, the multiple sparse tensors can be distributed to multiple computing devices so that the sum of the optimization indices of the sparse tensors obtained in each computing device is as close as possible, thereby optimizing the computational load of each computing device, achieving load balancing of computing devices, reducing computation time, and improving the utilization rate of computing resources.
[0090] The method flow provided in this application will be described below in conjunction with the aforementioned system architecture.
[0091] See Figure 4 The following is a flowchart illustrating a data processing method provided in this application.
[0092] Step 401: Transform the N input sparse tensors into a sequence of N sparse tensors;
[0093] First, due to the irregular data distribution of sparse tensors, it is difficult to ensure that the amount of data obtained by each computing device is balanced when splitting sparse tensors for distributed parallel computing. Furthermore, sparse tensors are usually indexed to record the position of non-zero elements and related information, and it is difficult to guarantee the consistency of the index when splitting sparse tensors across multiple computing devices.
[0094] Therefore, current distributed parallel computation of sparse tensors typically requires adding a batch dimension to the sparse tensor as a partitioning dimension for distributed computation. However, the most widespread application of sparse tensors is two-dimensional sparse tensors, i.e., sparse matrices. Most deep learning frameworks and libraries typically support two-dimensional sparse structures, and their application in distributed computation of high-dimensional sparse tensors is not yet mature or universal enough.
[0095] Based on this, in this embodiment, N input sparse tensors can be transformed into N sparse tensor sequences. This allows subsequent partitioning of the sparse tensors to be performed using the sparse tensors as the data granularity, avoiding the splitting of the data stored in the sparse tensors and also avoiding adding batch dimension to the sparse tensors. This transforms the processing from high-dimensional sparse tensors to low-dimensional sparse tensors. The sparse tensor sequence can be a tuple, list, or dictionary, or it can be a built-in type array or vector; the specific type is not limited here.
[0096] Step 402: Determine the computational load of each sparse tensor sequence based on the sparsity characteristics of each sparse tensor;
[0097] Factors affecting the computational load of sparse tensors include the number of nonzero elements (nnz) involved in the computation and the uniformity of sparsity. Higher sparsity results in smaller nnz and a lower computational load. The computational load is negatively correlated with the memory access time of specified indices and the degree of uniformity of sparsity; higher uniformity of sparsity results in a lower computational load.
[0098] Optionally, the computational load of each sparse tensor sequence can be determined based on the number of non-zero elements in each sparse tensor. The computational load can be expressed as δ0 = nnz.
[0099] Optionally, the computational load for each sparse tensor sequence can be determined based on the sparsity uniformity of each sparse tensor. The computational load can be expressed as follows: It represents the ratio of the sum of the squares of the number of non-zero elements in each row of a sparse tensor to the total number of non-zero elements.
[0100] Furthermore, the expression for calculating the computational load for each sparse tensor sequence can also be a user-defined calculation formula based on the characteristics of sparse tensors, and no specific limitations are imposed here.
[0101] Step 403: Divide the N sparse tensor sequences into M groups according to the number of computing devices M and the computational load of each sparse tensor sequence;
[0102] This involves dividing N sparse tensor sequences into M groups using an optimized grouping algorithm, based on the number of computing devices M and the computational load of each calculated sparse tensor sequence. The optimized grouping algorithm can be either bucket sort or K-means sort; the specific algorithm is not limited here. Furthermore, the difference in computational load between the resulting M groups of sparse tensor sequences should be minimized among all methods for grouping the N sparse tensor sequences.
[0103] Specifically, the greedy bucketing algorithm in the bucket sort algorithm can be used to group the N sparse tensor sequences. The N sparse tensor sequences are sorted from smallest to largest according to their computational load. Then, the sparse tensor sequence with the smallest computational load can be grouped with the sparse tensor sequence with the largest computational load, the sparse tensor sequence with the second smallest computational load can be grouped with the sparse tensor sequence with the second largest computational load, and so on, to complete the grouping of the N sparse tensor sequences and make the computational load of each computing device as close as possible.
[0104] In this embodiment, by optimizing the grouping algorithm to group multiple sparse tensors, the sum of the optimization indices of the sparse tensors obtained by each computing device can be made as close as possible, thereby optimizing the computational load of each computing device, reducing the waiting time of the computing devices, and thus improving the utilization of computing resources. For example, load optimization is performed using four computing devices as an example. Figure 5 As can be seen, after optimizing the computing load of each computing device, since the computing load of the computing devices is as close as possible, the overall computing time of each computing device is shortened. Therefore, the waiting time of the computing devices is also shortened, thereby improving the utilization rate of computing resources.
[0105] Step 404: Send the M sets of sparse tensor sequences to M computing devices to run the artificial intelligence model.
[0106] Before sending M sets of sparse tensor sequences to M computing devices to run the artificial intelligence model, it is also possible to obtain the operators in the artificial intelligence model that need to be computed in parallel.
[0107] Optionally, operators that need to be computed in parallel in the artificial intelligence model can be identified based on the operator identifier input by the user; then, the operator can be deployed in M computing devices according to the number of computing devices.
[0108] Optionally, a computation function can be obtained to instruct the sparse tensor to be computed; subsequently, a computation node can be determined based on the computation function, which is an operator that needs to be computed in parallel during the computation of the sparse tensor; after determining the computation node, the computation node can be copied according to the number of computing devices performing parallel computation, resulting in multiple computation nodes.
[0109] For example, taking four computing devices as an example, the schematic diagram of the parallel computing node replication mapping is as follows: Figure 6 As shown, based on the identifiers of the computation operators, computation nodes 1 and 2 that require parallel computation are identified. Subsequently, computation nodes 1 and 2 can be replicated according to the number of computing devices. Furthermore, computation nodes 1 and 2 can be pre-deployed on other computing devices to facilitate distributed computation of sparse tensors later. In this example, both the preceding and succeeding nodes are non-parallel computation nodes.
[0110] After obtaining M sets of sparse tensor sequences, the M sets of sparse tensor sequences can be sent to M computing devices to achieve distributed computing.
[0111] Optionally, after dividing the M groups of sparse tensor sequences, each computing device can perform distributed computation based on the obtained sparse tensor sequences to obtain the computation results of the M computing devices. Subsequently, the computation results of each computing device can be aggregated. The computation results of each computing device can be sorted according to the sparse tensor sequence correspondence table, which includes the correspondence between the sparse tensor sequences in the M groups and the first N sparse tensor sequences of the group. This table can represent the order change of the N sparse tensor sequences after grouping. After reordering the computation results of each computing device according to the sparse tensor sequence correspondence table, the order of the aggregated sparse tensor sequences can be consistent with the order in the initial sparse tensor sequences, which facilitates the comparison of the changes between the initial sparse tensors and the computed sparse tensors.
[0112] Since the number of non-zero elements in a sparse tensor may change during computation, resulting in a dynamic change in the shape of the sparse tensor computation result, in this embodiment, the allgatherv communication operator can be used to aggregate the computation results of each computing device, which can realize the aggregation of sparse tensors with misaligned data shapes. However, before aggregation, it is necessary to clarify the data estimate of each computing device.
[0113] Therefore, after data distribution (i.e., partitioning multiple sparse tensors), the `allgather` communication operator can be used to obtain the shape of the first computation result from each computing device, resulting in a shape sequence. This shape is used to estimate the data volume of each computing device, and the shape sequence represents the estimated data volume of each computing device. Subsequently, the `allgatherv` communication operator can be used to aggregate the computation results from each computing device, and the aggregated computation results can be broadcast to each computing device so that each computing device has the full computation results. The flowchart illustrating the communication aggregation process after data distribution is shown below. Figure 7 As shown.
[0114] In this embodiment, multiple sparse tensor sequences can be grouped according to the computational load of the sparse tensor, and multiple sparse tensor sequences can be divided into multiple computing devices according to the number of computing devices and the computational load of each sparse tensor sequence. The computational load can be obtained according to a custom function or a user-inputted calculation method, so that the computational load of each computing device is as close as possible, thereby optimizing the computational load of the computing devices and improving the utilization of computing resources.
[0115] The method flow provided in this application has been described above. The apparatus provided in this application will now be described based on the aforementioned method flow.
[0116] See Figure 8 This application provides a schematic diagram of the structure of a processor 800, including:
[0117] Processing module 801 is used to transform N input sparse tensors into a sequence of N sparse tensors;
[0118] The determination module 802 is used to determine the computational load of each sparse tensor sequence based on the sparse characteristics of each sparse tensor. The computational load indicates the amount of computation required when each sparse tensor sequence is executed by a computing device.
[0119] The partitioning module 803 is used to divide N sparse tensor sequences into M groups according to the number of computing devices M and the computing load of each sparse tensor sequence, with each group including at least one sparse tensor sequence.
[0120] The sending module 804 is used to send M sets of sparse tensor sequences to M computing devices to run artificial intelligence models.
[0121] In one possible implementation, the difference in computational load between the aforementioned M groups is the smallest among all schemes for dividing the N sparse tensor sequences into the M groups.
[0122] In one possible implementation, the aforementioned determining module 802 is specifically used to: determine the computational load of each sparse tensor sequence based on the number of non-zero elements in each sparse tensor.
[0123] In one possible implementation, the aforementioned determining module 802 is specifically used to: determine the computational load of each sparse tensor sequence based on the sparsity uniformity of each sparse tensor, where sparsity uniformity represents the degree of uniformity of the distribution of non-zero elements in each sparse tensor.
[0124] In one possible implementation, the device may further include:
[0125] The aforementioned processing module 801 is also used to record the sparse tensor sequence allocated to each computing device after allocating the M sets of sparse tensor sequences to the M computing devices.
[0126] The acquisition module 805 is used to obtain the calculation result corresponding to each sparse tensor sequence from the calculation results after the calculations are completed on N sparse tensor sequences by M computing devices.
[0127] In one possible implementation, the device may further include:
[0128] The aforementioned processing module 801 is also used to identify operators in the artificial intelligence model that need to be computed in parallel based on the computation operator identifiers input by the user;
[0129] Deployment module 806 is used to deploy operators in M computing devices.
[0130] In one possible implementation, the device may further include:
[0131] The aforementioned acquisition module 805 is used to acquire the shapes of the calculation results of M computing devices through the allgather communication operator, and obtain a shape sequence;
[0132] The aggregation module 807 is used to aggregate the computation results corresponding to each computing device based on the shape sequence using the allgatherv communication operator.
[0133] This application also provides a computer-readable storage medium storing a program that, when run on a computer, causes the computer to perform the aforementioned actions. Figure 4 The steps in the method described in the illustrated embodiment.
[0134] This application also provides a computer program product that, when run on a computer, causes the computer to perform the aforementioned actions. Figure 4 The method steps described in the illustrated embodiment.
[0135] It should also be noted that the device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. In addition, in the device embodiment drawings provided in this application, the connection relationship between modules indicates that they have a communication connection, which can be implemented as one or more communication buses or signal lines.
[0136] Through the above description of the embodiments, those skilled in the art can clearly understand that, for the sake of convenience and brevity, the specific working process of the systems, devices and units described above can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0137] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be an indirect coupling or communication connection between apparatuses or units through some interfaces, and may be electrical, mechanical, or other forms.
[0138] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0139] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0140] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0141] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a particular order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments described herein can be implemented in a sequence other than that illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0142] Finally, it should be noted that the above are only specific embodiments of this application, but the protection scope of this application is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the protection scope of this application.
Claims
1. A data processing method, characterized in that, include: Transform N input sparse tensors into a sequence of N sparse tensors; The computational load of each sparse tensor sequence is determined based on the sparsity characteristics of each sparse tensor, wherein the computational load indicates the amount of computation required when each sparse tensor sequence is executed by a computing device; The N sparse tensor sequences are divided into M groups based on the number of computing devices M and the computational load of each sparse tensor sequence, with each group including at least one sparse tensor sequence. The M sets of sparse tensor sequences are sent to M computing devices to run the artificial intelligence model.
2. The method according to claim 1, characterized in that, The difference in computational load among the M groups is the smallest among all schemes for dividing the N sparse tensor sequences into the M groups.
3. The method according to any one of claims 1 or 2, characterized in that, The computational load for determining each sparse tensor sequence based on the sparse characteristics of each sparse tensor includes: The computational load of each sparse tensor sequence is determined based on the number of non-zero elements in each sparse tensor.
4. The method according to any one of claims 1 or 2, characterized in that, The computational load for determining each sparse tensor sequence based on the sparse characteristics of each sparse tensor includes: The computational load of each sparse tensor sequence is determined based on the sparsity uniformity of each sparse tensor, where the sparsity uniformity represents the degree of uniformity of the distribution of non-zero elements in each sparse tensor.
5. The method according to any one of claims 1 to 4, characterized in that, The method further includes: After allocating the M sets of sparse tensor sequences to the M computing devices, record the sparse tensor sequence allocated to each computing device; After the M computing devices complete the calculations on the N sparse tensor sequences and obtain the calculation results, the calculation result corresponding to each sparse tensor sequence is obtained from the calculation results according to the record.
6. The method according to any one of claims 1 to 5, characterized in that, The method further includes: Identify the operators that need to be computed in parallel in the artificial intelligence model based on the computation operator identifiers input by the user; The operator is deployed in the M computing devices.
7. The method according to any one of claims 1 to 6, characterized in that, The method further includes: The shape of the computation results of the M computing devices is obtained through the allgather communication operator, resulting in a shape sequence. Based on the shape sequence, the calculation results corresponding to each computing device are aggregated using the allgatherv communication operator.
8. A processor, characterized in that, The processor is connected to a memory, and the processor is configured to execute instructions stored in the memory to cause the processor to perform the method as described in any one of claims 1 to 7.
9. A computer-readable storage medium comprising instructions that, when executed on a computer, cause the computer to perform the method as claimed in any one of claims 1 to 7.
10. A computer program product comprising instructions that, when run on a computer, cause the computer to perform the method as described in any one of claims 1 to 7.
Citation Information
Cited By
Sparse tensor processing method and device, electronic equipment and storage medium
CN122242571A
A sparse tensor processing method and device, electronic equipment and storage medium
CN122242571B