Data processing method and device, equipment and storage medium
By segmenting in the target tensor according to the dimension with the largest storage position span, the problem of low data reading and computing efficiency caused by unreasonable tensor segmentation in the prior art is solved, and more efficient data processing is achieved.
Patent Information
- Application Number
- CN202510566117.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-30
- Publication Date
- 2025-08-15
AI Technical Summary
In the prior art, the segmentation of tensors is not reasonable enough, resulting in high-performance devices such as graphics processors need to go to multiple storage areas when reading sub-tensors, and the data reading efficiency is low, which leads to bottlenecks in improving computing efficiency.
By judging the storage location span of each dimension in the target tensor, and dividing the target tensor into multiple sub-tensters according to the first target dimension with the largest storage location span, ensuring that the element distribution area of the sub-tenster is relatively concentrated, thereby improving data reading efficiency.
This improves data reading efficiency, thereby improving computing efficiency, and solving the problem of bottlenecks in improving computing efficiency.
Smart Images

Figure CN120495061A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and in particular to data processing methods, devices, equipment, and storage media. Background Art
[0002] In the field of deep learning and large model computing, as the model scale expands, the dimension of the input tensor grows exponentially, which leads to a significant increase in the computational complexity of the underlying operators (such as matrix multiplication, convolution, etc.).
[0003] Currently, some technologies divide tensors into multiple submodules and then use high-performance devices such as graphics processing units (GPUs) and tensor processing units (TPUs) to process tensor operations of these multiple submodules in parallel. This improves computational efficiency and, in turn, enhances model training and inference performance. However, in these technologies, the division of tensors is not reasonable, and there are bottlenecks in improving computational efficiency. Summary of the Invention
[0004] The present application provides a data processing method, a data processing device, an electronic device, a computer-readable storage medium, and a computer program product to at least solve the bottleneck problem in improving computing efficiency in related technologies.
[0005] This application provides a data processing method, including:
[0006] Obtaining a target tensor on which a target mathematical operation is to be performed and a storage layout of the target tensor, wherein the target tensor has multiple dimensions, each element in the target tensor has an index value in each dimension, and the storage layout represents a storage location distribution of the elements in the target tensor in a storage medium;
[0007] When the target tensor satisfies a preset segmentation condition, determining a storage location span of each dimension based on the storage layout and the index value, where the storage location span refers to the minimum number of elements in the storage medium that must be spaced apart when the index value of each dimension changes;
[0008] The target tensor is divided into a plurality of sub-tensors according to at least a first target dimension having the largest storage location span, and the target mathematical operation is performed using the sub-tensors.
[0009] The present application also provides a data processing device, comprising:
[0010] a tensor acquisition module, configured to acquire a target tensor to be subjected to a target mathematical operation and a storage layout of the target tensor, wherein the target tensor has multiple dimensions, each element in the target tensor has an index value in each dimension, and the storage layout represents a storage location distribution of the elements in the target tensor in a storage medium;
[0011] a span search module, configured to determine, if the target tensor satisfies a preset segmentation condition, a storage location span of each dimension based on the storage layout and the index value, wherein the storage location span refers to the minimum number of elements in the storage medium that must be spaced apart when the index value of each dimension changes;
[0012] A splitting module is used to split the target tensor into multiple sub-tensors according to at least a first target dimension with the largest storage location span, and use the sub-tensors to perform the target mathematical operation.
[0013] The present application also provides an electronic device, comprising: a memory for storing a computer program; and a processor for implementing the steps of any of the above-mentioned data processing methods when executing the computer program.
[0014] The present application also provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the steps of any of the above-mentioned data processing methods are implemented.
[0015] In the technical solutions of some embodiments of the present application, when a target tensor to be subjected to a target mathematical operation is obtained, the target tensor is divided into multiple sub-tensors by determining the storage location spans of each dimension in the target tensor and at least according to the first target dimension with the largest storage location span. In this way, the element distribution areas of each sub-tensor can be relatively concentrated, so that when a device such as a graphics processor that performs the target data operation obtains the value of each element in the sub-tensor, it does not need to read the allocated sub-tensors from multiple storage areas. Therefore, the data reading efficiency can be improved, and the operation efficiency can be improved, thereby solving the bottleneck problem of improving the operation efficiency in the related art. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] In order to more clearly illustrate the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0017] Figure 1 is a sample image of the first color channel;
[0018] Figure 2is an example image of the second color channel;
[0019] Figure 3 is an example image of the third color channel;
[0020] Figure 4 A schematic diagram of the three-dimensional tensor space of an image;
[0021] Figure 5 A flowchart of a data processing method provided in some embodiments of the present application;
[0022] Figure 6 A schematic diagram of a module of a data processing device provided in some embodiments of the present application;
[0023] Figure 7 A schematic diagram of a module of an electronic device provided for some embodiments of the present application. DETAILED DESCRIPTION
[0024] The following will be combined with the accompanying drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0025] It should be noted that, in the description of this application, the terms "comprises," "includes," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements but also other elements not explicitly listed, or also includes elements inherent to such process, method, article, or device. The terms "first," "second," etc., in this application are used to distinguish similar objects and are not used to describe a specific order or precedence.
[0026] In order to enable those skilled in the art to better understand the present application, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.
[0027] A tensor is a data structure consisting of multiple dimensions. In the field of artificial intelligence, tensors are used to describe various types of data. For example, image data can be represented as a three-dimensional tensor (color channels, height, width), and video data can be represented as a four-dimensional tensor (frames, height, width, and color channels). For ease of understanding, the following explanation uses image data as an example.
[0028] Assume that image A has three color channels and its height and width are 5 (i.e., the pixels of image A are 5×5). Figures 1 to 3, are sample images of image A in the first color channel, the second color channel, and the third color channel, respectively. Figures 1 to 3 As can be seen, Image A has 25 pixels in each color channel. By fusing the color values of the corresponding pixels in the three color channels, we can obtain the color value of the corresponding pixel in Image A. For example, by fusing the color value at the 1×1 position in the first color channel, the color value at the 1×1 position in the second color channel, and the color value at the 1×1 position in the third color channel, we can obtain the color value at the 1×1 position in Image A.
[0029] It can be seen that the image data of image A can be represented by three dimensions: color channel, height, and width. That is, the image data of image A can be represented as a three-dimensional tensor (color channel, height, width). In the three-dimensional tensor representing image A, multiple elements can be included. These elements are used to represent the color value of each pixel in each color channel of image A. For example, element P1 represents the color value of image A at the 0×0 position of the first color channel, element P2 represents the color value of image A at the 0×1 position of the first color channel, and element Pn represents the color value of image A at the 4×4 position of the third color channel.
[0030] Further, see Figure 4 , is an example of the three-dimensional tensor space of image A. Figure 4 In the three-dimensional tensor space shown, each dimension has its own corresponding index value range. The so-called index value range is the value range of each dimension. For example, the index range of the color channel is 0 to 2, and the index value range of the height is 0 to 4. Each element has an index value in each dimension, and the index value is used to limit the position of the element in the three-dimensional tensor space. For example, the index value of the above element P1 is (0,0,0), and the index value of the above element P2 is (0,0,1). Simply put, in the three-dimensional tensor space, the value of the element is used to represent the color value of image A at the corresponding position. For example, assuming that the value of element P1 is 20 and the index value is (0,0,0), it means that the color value of image A at the 0×0 position of the first color channel is 20. It can be seen that in the three-dimensional tensor of image A, the color values and index values represented by each element can be exemplarily shown in Table 1.
[0031] Table 1 Color values and index values
[0032]
[0033] In the fields of deep learning and large-scale model computing, when performing mathematical operations based on multidimensional tensors, to improve computational efficiency, a multidimensional tensor is often split into multiple sub-tensors. These sub-tensors are then computed in parallel using high-performance devices such as graphics processors and tensor processors. Finally, the computational results of these sub-tensors are aggregated to obtain the computational result of the multidimensional tensor. For example, Table 1 above shows all the elements included in the three-dimensional tensor of image A. When segmenting the three-dimensional tensor of image A, elements P1 to P15 can be segmented into sub-tensor A1, elements P16 to P30 into sub-tensor A2, and so on, resulting in five sub-tensors: A1, A2, A3, A4, and A5. Sub-tensor A1 can be assigned to graphics processor U1, sub-tensor A2 can be assigned to graphics processor U2, and so on. Each graphics processor can perform mathematical operations on its assigned sub-tensors and obtain the computational results. Finally, the computational results of all graphics processors are aggregated to obtain the computational result of the three-dimensional tensor of image A.
[0034] Furthermore, when segmenting a multi-dimensional tensor, the segmentation is usually performed according to one or more dimensions. The following still uses the three-dimensional tensor of image A as an example for explanation.
[0035] Refer to Table 1. For example, assuming that the three-dimensional tensor of image A is segmented according to the color channel dimension, then all elements of the height and width corresponding to each color channel can be segmented into the same sub-tensor. For example, the elements of the first color channel can be segmented into sub-tensor B1, the elements of the second color channel can be segmented into sub-tensor B2, and the elements of the third color channel can be segmented into sub-tensor B3. Among them, the elements of the first color channel refer to the elements with the color channel index value of 0, such as (0,0,0) and (0,4,1). The elements of the second color channel refer to the elements with the color channel index value of 1, such as the elements with the index values of (1,0,0) and (1,4,1). The elements of the third color channel refer to the elements with the color channel index value of 2, such as the elements with the index values of (2,0,0) and (2,4,1).
[0036] For example, suppose the three-dimensional tensor of image A is segmented according to the height dimension. The height can be divided into multiple height ranges, such as 0-2 and 3-4. Elements of all widths of all color channels corresponding to each height range can be segmented into the same sub-tensor. For example, elements with heights of 0-2 can be segmented into sub-tensor C1, and elements with heights of 3-4 can be segmented into sub-tensor C2. Elements with heights of 0-2 are those with height indices of 0-2, such as (0,2,0), (1,2,1), and (2,2,4). Similarly, elements with heights of 3-4 are those with height indices of 3-4, such as (0,3,0), (1,3,1), and (2,4,4).
[0037] Currently, some technologies lack a robust approach to segmenting multidimensional tensors. When high-performance devices like graphics processors and tensor processors retrieve allocated sub-tensors from storage media (such as memory) for mathematical operations, they must access data from multiple storage areas. This inefficient data retrieval leads to bottlenecks in improving computational efficiency. The following uses the three-dimensional tensor of image A as an example to illustrate this.
[0038] For example, assume that in a storage medium, the storage order of each element in the three-dimensional tensor of image A is as shown in Table 2. For the sake of simplicity, Table 2 only records the index value of each element, not the element value.
[0039] Table 2 Storage order in storage medium
[0040]
[0041]
[0042] In Table 2, each table represents a storage unit in the storage medium. It should be noted that in actual storage, the storage units may not be continuous. For example, there may be a storage unit for storing other data between P2 (0, 0, 1) and P3 (0, 0, 2).
[0043] Based on Table 2, assume that the three-dimensional tensor of image A is segmented according to the height dimension. For example, elements with heights of 0 to 2 are segmented into sub-tensor C1, and elements with heights of 3 to 4 are segmented into sub-tensor C2. At the same time, sub-tensor C1 is assigned to graphics processor U1 for mathematical operations, and sub-tensor C2 is assigned to graphics processor U2 for mathematical operations. Therefore, when graphics processor U1 reads sub-tensor C1 from the storage medium, it needs to read the values of each element in rows 1 to 3, rows 6 to 8, and rows 11 to 13 of Table 2. Similarly, when graphics processor U2 reads sub-tensor C2 from the storage medium, it needs to read the values of each element in rows 2 to 4, rows 9 to 10, and rows 14 to 15 of Table 2. This example shows that graphics processors U1 and U2 need to access multiple storage areas to read the values of each element in the assigned sub-tensors. This region-by-region reading significantly reduces data reading efficiency, creating a bottleneck in improving data calculation efficiency.
[0044] In view of this, the present application provides a data processing method that can solve the bottleneck problem of improving data computing efficiency in some technologies. The data processing method can be applied to data processing devices such as central processing units, graphics processing units, and tensor processing units, or can also be applied to electronic devices that integrate data processing devices such as central processing units, graphics processing units, and tensor processing units. Among them, electronic devices can include but are not limited to tablet computers, laptop computers, desktop computers, servers, etc. Figure 5 , which is a flow chart of the data processing method provided in some embodiments of the present application. Figure 5 In the data processing method, the data processing method includes the following steps:
[0045] Step S501, obtain a target tensor to be subjected to a target mathematical operation and a storage layout of the target tensor, wherein the target tensor has multiple dimensions, and each element in the target tensor has an index value in each dimension. The storage layout represents the storage location distribution of the elements in the target tensor in the storage medium.
[0046] Among them, regarding the tensor, the dimension of the tensor, and the index value of each element in the tensor, please refer to the above description, which will not be repeated here. In addition, regarding the storage layout of the target tensor, please refer to Table 2 above, which will not be repeated here.
[0047] In this embodiment, the data processing device can obtain information such as the storage address, tensor shape, storage type, and element data type of the target tensor while obtaining the target tensor.
[0048] The tensor shape refers to the size or size of the tensor in each dimension, that is, the maximum index value in each dimension. For example, the tensor shape of the three-dimensional tensor of image A above is (3, 5, 5).
[0049] The storage type refers to how the elements of the target tensor are organized on the storage medium. These organization methods can include, but are not limited to, row-major and column-major. Row-major means that the elements of the target tensor are first arranged by the index of the lowest dimension, then by the index of the next-lowest dimension, and so on. Column-major means that the elements of the target tensor are first arranged by the index of the highest dimension, then by the index of the next-highest dimension, and so on. The lowest dimension refers to the dimension at the bottom of the target tensor's dimensional hierarchy. The highest dimension refers to the dimension at the top of the target tensor's dimensional hierarchy. For example, consider Table 2. Assume that, among the three dimensions of color channel, width, and height, color channel is the bottom dimension, and height is the top dimension. As shown in Table 2, the elements of the three-dimensional tensor of image A are first arranged by the index of the highest dimension, so the elements of the three-dimensional tensor of image A are stored on the storage medium in a column-major organization.
[0050] The element data type specifies the amount of storage space that the element's value will occupy. Data types include, but are not limited to, int and float.
[0051] Based on the storage address, tensor shape, storage type, element data type and other information of the target tensor, the index value of each element in each dimension of the target tensor can be determined.
[0052] Step S502: When the target tensor meets the preset segmentation conditions, the storage location span of each dimension is determined based on the storage layout and index value. The storage location span refers to the minimum number of elements required to be separated when the index value of each dimension changes in the storage medium.
[0053] In this embodiment, the preset splitting condition includes at least one of the following conditions: 1) the number of elements in the target tensor is greater than a data volume threshold; 2) the storage space required by the elements in the target tensor is greater than a space threshold. When the target tensor meets any of the above conditions, it means that the target tensor meets the preset splitting condition.
[0054] The number of elements in the target tensor is equal to the product of the maximum index values of each dimension. For example, in the three-dimensional tensor of image A above, the number of elements = 3*5*5 = 75.
[0055] The storage space required for each element in the target tensor is equal to the product of the number of elements and the element data type. For example, if the element data type in the 3D tensor of image A is int (i.e., each element occupies 4 bytes of storage space), the storage space required for each element in the 3D tensor of image A = 75 * 4 = 300 bytes.
[0056] Specifically, given that data processing devices often have hardware limitations when processing large amounts of data, such as memory and index limits, the aforementioned data volume and space thresholds can be determined based on the hardware limitations of the data processing devices. This helps avoid issues such as memory overflow and int32_t index overflow during operations on the target tensor.
[0057] The following examples illustrate how to determine the storage span for each dimension. Refer to Table 2. For example, in the height dimension, the height index values of P1-P5 are all 0, the height index values of P6-P10 are all 1, the height index values of P11-P15 are all 2, and so on. This means that when the index value of the height dimension changes, there must be a minimum interval of 5 elements. Therefore, the storage span of the height dimension is 5. For another example, in the width dimension, the width index value of P1 is 0, the width index value of P2 is 1, the width index value of P3 is 2, and so on. This means that when the index value of the width dimension changes, there must be a minimum interval of 0 elements. Therefore, the storage span of the width dimension is 0. For another example, in the color channel dimension, the color channel index values of P1-P25 are all 0, the color channel index values of P26-P50 are all 1, and so on. This means that when the color channel index value changes, there must be a minimum interval of 25 elements. Therefore, the storage span of the color channel dimension is 25.
[0058] Step S503 : Split the target tensor into multiple sub-tensors according to at least the first target dimension with the largest storage location span, and use the sub-tensors to perform the target mathematical operation.
[0059] For example, in Table 2, since the storage location span of the color channel dimension of the color channel is the largest, the target tensor can be first divided into multiple sub-tensors according to the color channel dimension. For example, the elements of the first color channel are divided into sub-tensor B1, the elements of the second color channel are divided into sub-tensor B2, and the elements of the third color channel are divided into sub-tensor B3. In this way, the elements in sub-tensor B1 are distributed in rows 1 to 5 of Table 2, the elements in sub-tensor B2 are distributed in rows 6 to 10 of Table 2, and the elements in sub-tensor B3 are distributed in rows 11 to 15 of Table 2. The element distribution area in each sub-tensor is relatively concentrated, which can greatly improve the reading efficiency when reading the element values in each sub-tensor, thereby improving the computing efficiency.
[0060] In summary, in the technical solutions of some embodiments of the present application, when a target tensor to be subjected to a target mathematical operation is obtained, the target tensor is divided into multiple sub-tensors by judging the storage location spans of each dimension in the target tensor and at least according to the first target dimension with the largest storage location span. In this way, it is possible to ensure that the element distribution areas of each sub-tensor are relatively concentrated, so that when a graphics processor or other device that performs target data operations obtains the value of each element in the sub-tensor, it is not necessary to read the allocated sub-tensors from multiple storage areas. Therefore, the data reading efficiency can be improved, and then the operation efficiency can be improved, solving the bottleneck problem of improving the operation efficiency in the related art.
[0061] Continue referring to Table 2. After the target tensor is split into multiple sub-tensors according to the first target dimension (i.e., the color channel dimension), the number of elements in at least some of the sub-tensors may exceed the data volume threshold, or the storage space required by the elements in at least some of the sub-tensors may exceed the space threshold. In this case, to avoid exceeding the hardware limitations of the data processing device during the calculation process, the sub-tensors can be split again.
[0062] Specifically, in some embodiments, splitting the target tensor into multiple sub-tensors according to at least a first target dimension having the largest storage location span may include:
[0063] According to the first target dimension, split the target tensor into multiple first sub-tensors;
[0064] For any first sub-tensor, determine whether the first sub-tensor meets the preset segmentation condition. If so, further segment the first sub-tensor into multiple second sub-tensors.
[0065] For any second sub-tensor, if the second sub-tensor does not meet the preset segmentation condition, the segmentation of the second sub-tensor is stopped; if the second sub-tensor meets the preset segmentation condition, the segmentation of the second sub-tensor is continued.
[0066] In this embodiment, the target tensor can be divided into two first sub-tensors according to the first target dimension, and then whether the first sub-tensor needs to be divided into multiple second sub-tensors is determined based on the number of elements in each first sub-tensor and the storage space required by the elements.
[0067] Similarly, for the second sub-tensor obtained by segmentation, it is also possible to determine whether the second sub-tensor needs to be further segmented based on the number of elements in each second sub-tensor and the storage space required by the elements. This process is repeated until all the sub-tensors obtained by segmentation do not meet the preset segmentation conditions.
[0068] The sub-tensor used to perform the target mathematical operation is a sub-tensor obtained by segmentation that does not meet the preset segmentation conditions. For example, according to the first target dimension, after the target tensor is segmented into the first sub-tensor M1 and the first sub-tensor M2, if the first sub-tensor M2 meets the preset segmentation conditions, the first sub-tensor M2 is further segmented into the second sub-tensor M21 and the second sub-tensor M22. In the second sub-tensor M21 and the second sub-tensor M22, if the second sub-tensor M21 meets the preset segmentation conditions, the second sub-tensor M21 is further segmented into the third sub-tensor M211 and the third sub-tensor M212. If both the third sub-tensor M211 and the third sub-tensor M212 do not meet the preset segmentation conditions, the segmentation ends, and mathematical operations can be performed based on the first sub-tensor M1, the second sub-tensor M22, the third sub-tensor M211, and the third sub-tensor M212. In this way, it can be ensured that the operations of each sub-tensor will not exceed the hardware limitations of the data processing device, avoiding the risk of errors and memory leaks caused by excessive tensor data volume.
[0069] In some embodiments, when further segmenting any sub-tensor obtained by segmentation, the method of the present application may further include:
[0070] In the storage layout, obtain the sub-storage layout corresponding to the sub-tensor, where the sub-storage layout represents the storage location distribution of the elements in the sub-tensor in the storage medium;
[0071] According to the sub-storage layout and the index value of each element in the sub-tensor, find the second target dimension with the largest storage position span in the sub-tensor;
[0072] The sub-tensors are further split according to the second target dimension.
[0073] Specifically, the second target dimension may be the same as the first target dimension, or may be different from the first target dimension. Let's take Table 2 as an example. Assuming that according to the dimension of the color channel, the elements of the first color channel are divided into sub-tensor B1, the elements of the second color channel are divided into sub-tensor B2, and the elements of the third color channel are divided into sub-tensor B3, if the sub-tensor B1 meets the preset segmentation conditions, then the second target dimension with the largest storage position span can be found based on the storage location distribution of P1 to P25 and the index value of each element in P1 to P25. It can be seen from Table 2 that the second target dimension is height. According to the height dimension, each row of elements in sub-tensor B1 can be divided into a sub-tensor. In this way, in the next-level sub-tensor obtained based on the sub-tensor segmentation, the storage area of the elements is still relatively concentrated, which can improve the data reading efficiency.
[0074] In some embodiments, when segmenting the target tensor according to the first target dimension, the target tensor can be segmented according to the minimum segmentation unit in the first target dimension. The so-called minimum segmentation unit refers to the set of elements with the same index value in the first target dimension. This can reduce the number of subsequent segmentations.
[0075] In some embodiments, using subtensors to perform target mathematical operations may include:
[0076] Based on the sub-tensor, the target mathematical operation is divided into multiple sub-mathematical operations;
[0077] Creating at least two task flows, each task flow having a corresponding computing resource for performing a mathematical operation;
[0078] At least some of the sub-mathematical operations are distributed to different task flows, and at least some of the sub-mathematical operations are executed in parallel based on computing resources corresponding to the different task flows.
[0079] Specifically, a task stream can be created based on the Compute Unified Device Architecture (CUDA). Multiple sub-mathematical operations can be assigned to different task streams, allowing the sub-mathematical operations to be executed in parallel, thereby improving computational efficiency.
[0080] In some embodiments, when sub-mathematical operations are allocated to different task flows, a round-robin allocation mechanism may be adopted, thereby ensuring load balancing among the task flows.
[0081] In some embodiments, at least some sub-mathematical operations may have dependencies. Dependencies refer to the fact that some sub-mathematical operations rely on the results of other sub-mathematical operations. To ensure the proper execution of each sub-mathematical operation, an event chain provided by the unified device architecture can be used to ensure that dependent sub-mathematical operations are executed sequentially.
[0082] Before explaining the event chain, let's first explain the events. In the unified device architecture, each task flow can have its own corresponding event, and the event has an event status. The event status is used to mark whether the sub-mathematical operation in the corresponding task flow has been completed. For example, you can create event S. When the sub-mathematical operation O starts to be executed in task flow L, event S can be used to mark the sub-mathematical operation O. At this time, event S is associated with the sub-mathematical operation O. By calling functions such as cudaEventQuery, you can query the status of event S. The so-called querying the status of event S is actually querying whether the operation associated with event S is completed. If the operation associated with event S is completed, it returns completed; if the operation associated with event S is not completed, it returns incomplete. In this way, by querying the event status, you can determine whether the sub-mathematical operation in each task flow has been completed.
[0083] An event chain consists of a series of events with dependencies, which occur in a specific order, and the occurrence of the latter event depends on the completion of the previous event. For sub-mathematical operations with dependencies, the dependency between the sub-mathematical operations can be identified in an event chain manner so that the sub-mathematical operations with dependencies can be executed in sequence. Specifically, when a first sub-mathematical operation depends on a second sub-mathematical operation, the event status of the task flow where the second sub-mathematical operation is located is monitored; if the monitored event status indicates that the second sub-mathematical operation has been completed, the first sub-mathematical operation is executed based on the computing resources of the task flow where the first sub-mathematical operation is located; if the monitored event status indicates that the second sub-mathematical operation has not been completed, the execution of the first sub-mathematical operation is suspended.
[0084] For example, suppose that in an event chain, event S1 precedes event S2. Event S1 marks the execution of sub-mathematical operation O1, while event S2 marks the execution of sub-mathematical operation O2. After assigning sub-mathematical operation O1 to task flow 1 and sub-mathematical operation O2 to task flow 2, the status of event S1 can be monitored. If the status of event S1 indicates that sub-mathematical operation O1 has completed, sub-mathematical operation O2 can be executed based on the computing resources of task flow 2. If the status of event S1 indicates that sub-mathematical operation O1 has not completed, the execution of sub-mathematical operation O2 is suspended.
[0085] Compared with explicit stream synchronization, the dependencies between sub-computation tasks are identified through events and event chains, and only the predecessor stream needs to be waited for, which greatly reduces the synchronization overhead and avoids the performance bottleneck of global synchronization.
[0086] In some embodiments, the method of the present application further comprises:
[0087] Query the event status of the events corresponding to each task flow, and determine whether the sub-mathematical operations in each task flow have been completed based on the event status;
[0088] When all sub-mathematical operations in all task flows are completed, obtain the sub-operation results of each task flow;
[0089] The result of the target mathematical operation is determined based on the sub-operation results of each task flow.
[0090] Specifically, the sub-operation results of each task flow can be stored in a designated storage area. When the sub-mathematical operations in all task flows are completed, the sub-operation results of each task flow can be obtained from the designated storage area, and the operation result of the target mathematical operation can be obtained based on the sub-operation results (e.g., summary).
[0091] In the above embodiment, after all sub-mathematical operations in all task flows are completed, the sub-operation results are obtained, thereby ensuring the integrity of the operation results.
[0092] In some embodiments, the data processing device of the present application may provide a developed universal interface. When splitting a target tensor, these universal interfaces may be called to complete the process. For example, a tensor preprocessing interface may be provided. Before splitting the target tensor, the tensor preprocessing interface may be called to preprocess the target tensor, for example, to unify the data types of the elements in the target tensor to the same type.
[0093] Corresponding to the data processing method, the present application also provides a data processing device. Figure 6 , which is a module schematic diagram of a data processing device provided in some embodiments of the present application. Figure 6 In the data processing device, the data processing device includes:
[0094] A tensor acquisition module 601 is configured to acquire a target tensor to be subjected to a target mathematical operation and a storage layout of the target tensor, wherein the target tensor has multiple dimensions, each element in the target tensor has an index value in each dimension, and the storage layout represents the storage location distribution of the elements in the target tensor in a storage medium;
[0095] Span search module 602 is used to determine the storage location span of each dimension based on the storage layout and index value when the target tensor meets the preset segmentation conditions. The storage location span refers to the minimum number of elements required to be separated when the index value of each dimension changes in the storage medium;
[0096] The splitting module 603 is configured to split the target tensor into multiple sub-tensors according to at least a first target dimension having the largest storage location span, and perform a target mathematical operation using the sub-tensors.
[0097] In some embodiments, the preset segmentation condition includes at least one of the following conditions:
[0098] The number of elements in the target tensor is greater than the data amount threshold;
[0099] The storage space required for the elements in the target tensor is greater than the space threshold.
[0100] In some embodiments, the segmentation module 603 is specifically configured to:
[0101] According to the first target dimension, split the target tensor into multiple first sub-tensors;
[0102] For any first sub-tensor, determine whether the first sub-tensor meets the preset segmentation condition. If so, further segment the first sub-tensor into multiple second sub-tensors.
[0103] For any second sub-tensor, if the second sub-tensor does not meet the preset segmentation condition, then stop segmenting the second sub-tensor; if the second sub-tensor meets the preset segmentation condition, then continue segmenting the second sub-tensor;
[0104] The sub-tensor used for performing the target mathematical operation is a sub-tensor obtained by segmentation that does not meet the preset segmentation conditions.
[0105] In some embodiments, when further segmenting any sub-tensor obtained by segmentation, the segmentation module 603 is further configured to:
[0106] In the storage layout, obtain the sub-storage layout corresponding to the sub-tensor, where the sub-storage layout represents the storage location distribution of the elements in the sub-tensor in the storage medium;
[0107] According to the sub-storage layout and the index value of each element in the sub-tensor, find the second target dimension with the largest storage position span in the sub-tensor;
[0108] The sub-tensors are further split according to the second target dimension.
[0109] In some embodiments, the segmentation module 603 is specifically configured to:
[0110] Based on the sub-tensor, the target mathematical operation is divided into multiple sub-mathematical operations;
[0111] Creating at least two task flows, each task flow having a corresponding computing resource for performing a mathematical operation;
[0112] At least some of the sub-mathematical operations are distributed to different task flows, and at least some of the sub-mathematical operations are executed in parallel based on computing resources corresponding to the different task flows.
[0113] In some embodiments, each task flow has its own corresponding event, and the event has an event status, which is used to mark whether the sub-mathematical operation in the corresponding task flow has been completed. The segmentation module 603 is further used to:
[0114] Query the event status of the events corresponding to each task flow, and determine whether the sub-mathematical operations in each task flow have been completed based on the event status;
[0115] When all sub-mathematical operations in all task flows are completed, obtain the sub-operation results of each task flow;
[0116] The result of the target mathematical operation is determined based on the sub-operation results of each task flow.
[0117] In some embodiments, each task flow has its own corresponding event, and the event has an event status. The event status is used to mark whether the sub-mathematical operation in the corresponding task flow has been completed, and there are dependencies between at least some of the sub-mathematical operations. The segmentation module 603 is specifically used to:
[0118] When the first sub-mathematical operation depends on the second sub-mathematical operation, monitoring the event status of the task flow where the second sub-mathematical operation is located;
[0119] If the monitored event status indicates that the second sub-mathematical operation has been completed, then executing the first sub-mathematical operation based on the computing resources of the task flow where the first sub-mathematical operation is located;
[0120] If the monitored event status indicates that the second sub-mathematical operation has not been completed, the execution of the first sub-mathematical operation is suspended.
[0121] For the description of the features in the embodiment corresponding to the data processing device, reference can be made to the relevant description of the embodiment corresponding to the sample data processing method, which will not be repeated here.
[0122] See also Figure 7 An embodiment of the present application further provides an electronic device, comprising a memory 10 and a processor 20, wherein the memory 10 stores a computer program, and the processor 20 is configured to run the computer program to execute the steps in any one of the above-mentioned data processing method embodiments.
[0123] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored, wherein the computer program is configured to execute the steps of any of the above-mentioned data processing method embodiments when run.
[0124] In an exemplary embodiment, the computer-readable storage medium may include, but is not limited to, various media that can store computer programs, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk, or an optical disk.
[0125] An embodiment of the present application further provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the steps in any one of the above data processing method embodiments are implemented.
[0126] An embodiment of the present application further provides another computer program product, including a non-volatile computer-readable storage medium, wherein the non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in any of the above-mentioned data processing method embodiments are implemented.
[0127] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0128] The above is a detailed introduction to a data processing method, device, equipment and storage medium provided by the present application. Specific examples are used herein to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method and core ideas of the present application. It should be pointed out that for ordinary technicians in this technical field, without departing from the principles of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the scope of protection of the claims of the present application.
Claims
1. A data processing method, characterized in that: The method comprises: Obtaining a target tensor on which a target mathematical operation is to be performed and a storage layout of the target tensor, wherein the target tensor has multiple dimensions, each element in the target tensor has an index value in each dimension, and the storage layout represents a storage location distribution of the elements in the target tensor in a storage medium; When the target tensor satisfies a preset segmentation condition, determining a storage location span of each dimension based on the storage layout and the index value, where the storage location span refers to the minimum number of elements in the storage medium that must be spaced apart when the index value of each dimension changes; The target tensor is divided into a plurality of sub-tensors according to at least a first target dimension having the largest storage location span, and the target mathematical operation is performed using the sub-tensors.
2. The method according to claim 1, characterized in that The preset segmentation condition includes at least one of the following conditions: The number of elements in the target tensor is greater than the data amount threshold; The storage space required by the elements in the target tensor is greater than the space threshold.
3. The method according to claim 1 or 2, characterized in that The step of dividing the target tensor into a plurality of sub-tensors according to at least a first target dimension having a largest storage location span comprises: Splitting the target tensor into a plurality of first sub-tensors according to the first target dimension; For any of the first sub-tensors, determining whether the first sub-tensor satisfies the preset segmentation condition; if so, further segmenting the first sub-tensor into a plurality of second sub-tensors; For any second sub-tensor, if the second sub-tensor does not meet the preset segmentation condition, stop segmenting the second sub-tensor; if the second sub-tensor meets the preset segmentation condition, continue segmenting the second sub-tensor; The sub-tensor used to perform the target mathematical operation is a sub-tensor obtained by segmentation that does not meet the preset segmentation condition.
4. The method according to claim 3, characterized in that When further segmenting any sub-tensor obtained by segmentation, the method further includes: In the storage layout, obtaining a sub-storage layout corresponding to a sub-tensor, the sub-storage layout representing a storage location distribution of elements in the sub-tensor in the storage medium; Searching, in the sub-tensor, for a second target dimension having the largest storage position span according to the sub-storage layout and the index value of each element in the sub-tensor; The sub-tensor is further divided according to the second target dimension.
5. The method according to claim 1, wherein The using the sub-tensor to perform the target mathematical operation includes: Based on the sub-tensor, split the target mathematical operation into a plurality of sub-mathematical operations; Creating at least two task flows, each task flow having a corresponding computing resource for performing a mathematical operation; At least part of the sub-mathematical operations are distributed to different task flows, and at least part of the sub-mathematical operations are executed in parallel based on computing resources corresponding to the different task flows.
6. The method according to claim 5, characterized in that Each of the task flows has a corresponding event, and the event has an event status, and the event status is used to mark whether the sub-mathematical operation in the corresponding task flow is completed. The method further includes: querying the event status of the event corresponding to each of the task flows, and judging whether the sub-mathematical operations in each of the task flows have been completed according to the event status; When all sub-mathematical operations in all task flows are completed, obtaining sub-operation results of each task flow; The operation result of the target mathematical operation is determined according to the sub-operation results of each of the task flows.
7. The method according to claim 5, characterized in that Each of the task flows has a corresponding event, each of which has an event status, and the event status is used to mark whether the sub-mathematical operations in the corresponding task flow have been completed, and at least some of the sub-mathematical operations have dependencies. The method further includes: When the first sub-mathematical operation depends on the second sub-mathematical operation, monitoring the event status of the task flow where the second sub-mathematical operation is located; If the monitored event status indicates that the second sub-mathematical operation has been completed, executing the first sub-mathematical operation based on the computing resources of the task flow where the first sub-mathematical operation is located; If the monitored event status indicates that the second sub-mathematical operation has not been completed, the execution of the first sub-mathematical operation is suspended.
8. A data processing device, characterized in that: The device comprises: a tensor acquisition module, configured to acquire a target tensor to be subjected to a target mathematical operation and a storage layout of the target tensor, wherein the target tensor has multiple dimensions, each element in the target tensor has an index value in each dimension, and the storage layout represents a storage location distribution of the elements in the target tensor in a storage medium; a span search module, configured to determine, if the target tensor satisfies a preset segmentation condition, a storage location span of each dimension based on the storage layout and the index value, wherein the storage location span refers to the minimum number of elements in the storage medium that must be spaced apart when the index value of each dimension changes; A splitting module is used to split the target tensor into multiple sub-tensors according to at least a first target dimension with the largest storage location span, and use the sub-tensors to perform the target mathematical operation.
9. An electronic device, characterized in that: include: Memory for storing computer programs; A processor, configured to implement the steps of the data processing method according to any one of claims 1 to 7 when executing the computer program.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, wherein the computer program, when executed by a processor, implements the steps of the data processing method according to any one of claims 1 to 7.
Citation Information
Cited By
Convolution calculation method, electronic equipment, storage medium and program product
CN121233886A