Parallel Computing Method, System and Computer Device for Tensor Reduction

By dynamically segmenting tensors on non-criterion dimensions and performing parallel specifications, the problem of low computational efficiency of tensor criterion methods in the prior art in multi-core environment is solved, and efficient and dynamically adaptable tensor criterion operations are realized.

CN119718687BActive Publication Date: 2025-06-17青岛国实科技集团有限公司
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510227901.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2025-06-17
Estimated Expiration
2045-02-28

AI Technical Summary

Technical Problem

The existing tensor regulation methods have problems such as low computing efficiency, load imbalance, and data transmission bottlenecks in multi-core parallel computing environments, and cannot fully utilize the parallel capabilities of multi-core processors.

Method used

By dynamic tensor segmentation on a non-criterion dimension, multiple sub-tensors are generated and dynamically allocated to multiple computing processing units based on the number of sub-tensors. Parallel regulation is performed using recursive algorithms, and the final result of the regulation is written back to the main memory.

Benefits of technology

Improves computing efficiency, simplifies the synchronization mechanism, ensures data consistency when the result is written back, avoids potential write conflict problems, improves the spatial and temporal locality of the data, and reduces the number of interactions between the kernel group and memory.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119718687B_ABST
    Figure CN119718687B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of information processing technology. More specifically, the present invention relates to a parallel computing method, system, and computer device for tensor reduction, including: a tensor splitting step: dynamically splitting a high-dimensional tensor based on non-reduction dimensions to obtain multiple sub-tensors; a sub-tensor allocation step: allocating the multiple sub-tensors to multiple computing processing units based on the number of sub-tensors; a sub-tensor parallel reduction step: the multiple computing processing units performing parallel reduction on the multiple sub-tensors based on a recursive algorithm to obtain a reduced result tensor; and a reduced result tensor write-back step: writing the reduced result tensor back to the main memory.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of information processing technology. More specifically, the present invention relates to a parallel computing method, system, and computer device for tensor reduction. Background Art

[0002] With the rapid development of artificial intelligence and high-performance computing, the demand for computer computing power has increased sharply. As the core data structure in deep learning, the efficiency of tensor reduction operations (such as summation, maximum value calculation, etc.) directly affects the training and inference speed of the model. However, existing tensor reduction methods have many problems in multi-core parallel computing environments, such as low computing efficiency, load imbalance, data transmission bottlenecks, etc. For example, existing methods usually use simple multi-threaded parallel computing, which cannot fully utilize the parallel capabilities of multi-core processors, and lack dynamic adaptability in tensor slicing and distribution, resulting in some computing units being idle while others are overloaded. In addition, there are complex data transmission and synchronization operations in tensor reallocation and dynamic adjustment, further reducing the computing efficiency.

[0003] Therefore, providing a parallel computing method, system, and computer device for tensor reduction with high computing efficiency and strong dynamic adaptability is an urgent problem to be solved currently. Summary of the Invention

[0004] The present invention provides a parallel computing method, system, and computer device for tensor reduction to at least solve the problems of low computing efficiency and lack of dynamic adaptability of existing tensor reduction methods.

[0005] To achieve the above object, the present invention provides a parallel computing method for tensor reduction, including:

[0006] Tensor splitting step: Dynamically split a high-dimensional tensor based on non-reduction dimensions to obtain multiple sub-tensors;

[0007] Sub-tensor allocation step: Allocate multiple sub-tensors to multiple computing processing units based on the number of the sub-tensors;

[0008] Sub-tensor parallel reduction step: Multiple computing processing units perform parallel reduction on multiple sub-tensors based on a recursive algorithm to obtain a reduced result tensor;

[0009] Reduced result tensor write-back step: Write the reduced result tensor back to the main memory.

[0010] Further, the tensor splitting step includes:

[0011] Traverse all non-reduction dimensions from the last dimension of the high-dimensional tensor forward to the first dimension that satisfies the storage space requirement of the sub-tensor formed from this dimension to the last dimension is not less than the local memory capacity as the splitting dimension;

[0012] Obtain a segmentation step size based on the segmentation dimension and the local storage capacity;

[0013] Segment the high-dimensional tensor based on the segmentation dimension and the segmentation step size to obtain a plurality of sub-tensors.

[0014] Further, it is characterized in that the sub-tensor allocation step includes:

[0015] If the number of the sub-tensors is greater than the total number of slave cores, allocate the plurality of sub-tensors to a plurality of slave cores based on cyclic allocation;

[0016] If the number of the sub-tensors is less than the total number of slave cores, perform secondary segmentation on the sub-tensors based on the reduction dimension, divide the slave cores into working groups based on the number of the sub-tensors after secondary segmentation, and allocate the plurality of sub-tensors after secondary segmentation to a plurality of working groups based on cyclic allocation.

[0017] Further, the sub-tensor parallel reduction step includes:

[0018] A plurality of the computing processing units perform a recursive decomposition operation on the allocated sub-tensors in parallel until the dimension of the sub-tensors is not greater than a set threshold; perform a reduction operation on the sub-tensors that reach the set threshold to obtain a reduced result tensor, and record the reduced result tensor at a target storage location.

[0019] Further, the sub-tensor parallel reduction step further includes:

[0020] Reduction index calculation step: During the reduction process, calculate the reduction index of each element in the reduction dimension in real time based on the source data offset and the data stride, and record the reduced result tensor and the corresponding reduction index at the target storage location.

[0021] Further, the reduced result tensor write-back step includes:

[0022] If a single sub-tensor is processed by a single computing processing unit, the first slave core of each computing processing unit writes the reduced result tensor back to the main memory;

[0023] If a single sub-tensor is processed by a plurality of computing processing units, sequentially traverse the plurality of computing processing units participating in the processing of the sub-tensor, the first slave core of each computing processing unit writes the reduced result tensor back to the main memory, and perform a memory synchronization operation after each write-back.

[0024] The present invention provides a parallel computing system for tensor reduction, which is applied to the above-mentioned parallel computing method for tensor reduction, and includes:

[0025] Tensor splitting module: Dynamically split a high-dimensional tensor based on non-reduced dimensions to obtain multiple sub-tensors;

[0026] Sub-tensor allocation module: Allocate multiple sub-tensors to multiple computing processing units based on the number of the sub-tensors;

[0027] Sub-tensor parallel reduction module: Multiple computing processing units perform parallel reduction on multiple sub-tensors based on a recursive algorithm to obtain a reduced result tensor;

[0028] Reduced result tensor write-back module: Write the reduced result tensor back to the main memory.

[0029] Furthermore, the tensor splitting module includes:

[0030] Traverse all non-reduced dimensions forward from the last dimension of the high-dimensional tensor until the first dimension that satisfies the storage space requirement of the sub-tensor formed from this dimension to the last dimension is not less than the local memory capacity is used as the splitting dimension;

[0031] Obtain a splitting step based on the splitting dimension and the local memory capacity;

[0032] Split the high-dimensional tensor based on the splitting dimension and the splitting step to obtain multiple tensors.

[0033] The present invention provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the above-mentioned parallel computing method for tensor reduction is implemented.

[0034] The present invention provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, the above-mentioned parallel computing method is implemented.

[0035] Compared with the prior art, the advantages and positive effects of the present invention are as follows:

[0036] Based on the structure of the tensor and the characteristics of the reduction operation, the present invention realizes efficient parallel processing and improves the computing efficiency by dynamically splitting tensors in non-reduced dimensions. At the same time, the synchronization mechanism is simplified, ensuring data consistency when writing back the results and effectively avoiding potential write conflict problems. Meanwhile, the spatial and temporal locality of data is improved, thereby reducing the number of interactions between the slave core group and the memory and improving the memory access efficiency.

[0037] The present invention adopts a load balancing strategy of dynamically determining the parallel granularity according to the number of sub-tensors, making full use of the data exchange ability between slave cores in a multi-core parallel computing environment and improving the parallel computing efficiency.

[0038] The present invention performs recursive parallel reduction on the segmented sub-tensors, allowing data segmentation and reduction to be processed simultaneously in one traversal, thereby improving the computing efficiency. At the same time, the reduction calculation can also return the index of each element in the reduction dimension to achieve data tracking and traceability, which helps to better understand and optimize the algorithm, and also helps to locate potential errors and performance bottlenecks. Description of the Drawings

[0039] Figure 1 It is a schematic flowchart of the parallel computing method for tensor reduction in an embodiment of the present invention;

[0040] Figure 2 It is a schematic diagram of the recursive equivalence relationship for reducing high-dimensional tensors in an embodiment of the present invention;

[0041] Figure 3 It is a schematic diagram of parallel reduction of high-dimensional tensors in an embodiment of the present invention;

[0042] Figure 4 It is a schematic diagram of the structure of the parallel computing system for tensor reduction in an embodiment of the present invention;

[0043] Figure 5 It is a schematic diagram of the computer device provided in an embodiment of the present invention.

[0044] In the above figures:

[0045] 40. Bus; 41. Processor; 42. Memory; 43. Communication interface; 100. Tensor segmentation module; 200. Sub-tensor allocation module; 300. Sub-tensor parallel reduction module; 400. Reduction result tensor write-back module. Detailed Embodiments

[0046] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0047] In the description of the present application, it should be understood that the orientation or positional relationships indicated by the terms "center", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc. are based on the orientation or positional relationships shown in the drawings. These are only for the convenience of describing the present application and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus should not be construed as a limitation to the present application. The terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the present application, unless otherwise stated, the meaning of "plural" is two or more.

[0048] In the description of the present application, it should be noted that unless otherwise clearly defined and limited, the terms "installed", "connected", and "coupled" should be understood in a broad sense. For example, it may be a fixed connection, a detachable connection, or an integral connection; it may be a mechanical connection or an electrical connection; it may be directly connected or indirectly connected through an intermediate medium, and it may be the communication inside two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present application can be understood according to specific circumstances.

[0049] As Figures 1-5 shown, the present invention provides a parallel computing method, system and computer device for tensor reduction, so as to at least solve the problems of low computing efficiency and lack of dynamic adaptability in the existing tensor reduction methods. The following will Figures 1-5 describe in detail the parallel computing method, system and computer device for tensor reduction provided by the present invention.

[0050] Embodiment 1:

[0051] Figure 1 is a schematic flowchart of the parallel computing method for tensor reduction in an embodiment of the present invention. As Figure 1 shown, the present invention provides a parallel computing method for tensor reduction, including a tensor splitting step S100, a sub-tensor allocation step S200, a sub-tensor parallel reduction step S300, and a reduced result tensor write-back step S400.

[0052] The tensor splitting step S100: Dynamically split a high-dimensional tensor based on non-reduced dimensions to obtain a plurality of sub-tensors. Specifically, the non-reduced splitting dimensions can be found based on the reduced high-dimensional tensor.

[0053] Since tensor splitting occurs in non-reduced dimensions, each split sub-tensor is complete in the reduced dimension. This means that the results generated after reduction of the sub-tensors processed by each workgroup are unique to that sub-tensor and do not overlap with the results of other sub-tensors. This not only achieves efficient parallel processing and improves computational efficiency, but also simplifies the synchronization mechanism, ensures data consistency when writing back the results, effectively avoids potential write conflict problems, and at the same time improves the spatial and temporal locality of data, thereby reducing the number of interactions between the slave core group and memory and improving the memory access efficiency.

[0054] Preferably, the tensor splitting step S100 includes: determining the splitting dimension S101, calculating the splitting stride S102, and generating sub-tensors S103.

[0055] Determining the splitting dimension S101: Traverse all dimensions from the last dimension of the high-dimensional tensor forward to the first dimension that satisfies the condition that the storage space requirement of the sub-tensor formed from this dimension to the last dimension is not less than the local memory capacity as the splitting dimension;

[0056] Calculating the splitting stride S102: Obtain the splitting stride based on the splitting dimension and the local memory capacity; the splitting dimension and the splitting stride together determine the amount of data processed by each parallel task. To ensure that the space occupied by each split sub-tensor is less than or equal to the local memory, the data space occupied by the tensor formed by accumulating the splitting strides to the last dimension needs to be less than or equal to the local memory;

[0057] Generating sub-tensors S103: Split the high-dimensional tensor based on the splitting dimension and the splitting stride to obtain multiple sub-tensors. This splitting strategy is applicable to input tensors of different shapes and sizes. Through dynamic calculation and adjustment, it ensures that the size of the sub-tensors is always within the local memory capacity of the slave core, avoids unnecessary over-splitting, and at the same time maintains the computational efficiency.

[0058] In some embodiments, taking the original tensor T with dimensions [3, N, C, H, W] as an example for tensor splitting, assuming that the dimension numbers of T are 0, 1, 2, 3, 4 from left to right in sequence, the 0th dimension is used as the reduced dimension for reduction, and the reduced tensor T1 obtained after reduction is [1, N, C, H, W].

[0059] Determine the splitting dimension based on the reduced tensor T1. Search for the splitting dimension k in the non-reduced dimensions (i.e., the 1st to 4th dimensions). Traverse all dimensions from the 4th dimension forward to the first dimension that satisfies the condition that the cumulative storage space occupied by the sub-tensor [k,..., H, W] formed from k to the 4th dimension is greater than or equal to the local memory capacity as the splitting dimension, that is:

[0060]

[0061] Among them, Indicates the storage space accumulated by the sub-tensor formed from dimension k to the fourth dimension; M represents the local memory capacity.

[0062] Assume that the second dimension is the dimension that meets the condition, then the second dimension is the splitting dimension.

[0063] Calculate the splitting step S on the splitting dimension. To ensure that the space occupied by each sub-tensor after splitting is less than or equal to the local memory capacity, the data space occupied by the tensor formed by accumulating the splitting strides to the last dimension needs to be less than or equal to the local memory capacity, that is:

[0064]

[0065] Then the calculation formula for the splitting step S is as follows:

[0066]

[0067] Assume that the second dimension is the splitting dimension and the local memory capacity is 1024. Then the sub-tensor after splitting based on the second dimension is [1, 1, S, H, W], and the splitting step is:

[0068]

[0069] Based on the splitting dimension and the splitting step, the original tensor T before reduction is sliced into n sub-tensors , and the dimension of each sub-tensor is [3, 1, S, H, W].

[0070] Sub-tensor allocation step S200: Allocate multiple sub-tensors to multiple computing processing units based on the number of sub-tensors; the allocation strategy is dynamically adjusted according to the number of sub-tensors and the number of slave cores to ensure load balancing and efficient use of computing resources.

[0071] Preferably, the sub-tensor allocation step S200 includes:

[0072] If the number of sub-tensors is greater than the total number of slave cores, then allocate multiple sub-tensors to multiple slave cores based on cyclic allocation;

[0073] If the number of sub-tensors is less than the total number of slave cores, then perform secondary splitting of the sub-tensors based on the reduction dimension, divide the slave cores into working groups based on the number of sub-tensors after secondary splitting, and allocate multiple sub-tensors after secondary splitting to multiple working groups based on cyclic allocation.

[0074] In some embodiments, if the number of sub-tensors is greater than the total number of slave cores, then adopt a cyclic allocation strategy and allocate the sub-tensors to each slave core in turn. Each slave core processes one sub-tensor to ensure that all sub-tensors are processed.

[0075] Suppose the above-mentioned original tensor T[3, N, C, H, W] is divided into 100 sub-tensors after segmentation. Suppose there are 24 slave cores in the system. Then, the 100 sub-tensors are sequentially assigned to the 24 slave cores in a cyclic distribution manner. Slave core 0 processes sub-tensor 0, slave core 1 processes sub-tensor 1, and so on. After slave core 23 finishes processing sub-tensor 23, slave core 0 continues to process sub-tensor 24 until all sub-tensors are assigned.

[0076] If the total number of slave cores is more than the number of sub-tensors, the sub-tensors are further divided based on the reduction dimension, and a group of slave cores, that is, a workgroup, is used to process a sub-tensor to achieve finer-grained parallel processing.

[0077] After the slave cores within the workgroup finish parallel processing of the sub-tensor, they also need to communicate to further merge the reduction results. To efficiently exchange data between slave cores, there are 3 ways to select a workgroup, which are a row of slave cores in the slave core array, all slave cores in the slave core array, and all slave cores on a chip. If the workgroup is a row of slave cores in the slave core array, it can perform high-speed data exchange through RMA row broadcast. If the workgroup is all slave cores in the slave core array or all slave cores on a chip, it can perform high-speed data exchange through RMA intra-group broadcast.

[0078] When a row of slave cores is used as a workgroup, a chip can be divided into 48 workgroups, that is, 48 sub-blocks can be processed in parallel. When an entire slave core array is used as a workgroup, a chip can be divided into 6 workgroups, that is, 6 sub-blocks can be processed in parallel. From this, a load balancing strategy can be obtained. When the number of sub-tensors is greater than or equal to 48, a row of slave cores is used as a workgroup; when the number of sub-tensors is greater than or equal to 6 and less than 48, an entire slave core array is used as a workgroup. When the number of sub-tensors is less than 6, all slave cores on a chip are used as a workgroup.

[0079] After completing the workgroup division, multiple sub-tensors after secondary segmentation are still assigned to multiple workgroups based on cyclic distribution.

[0080] Parallel reduction step S300 of sub-tensors: Multiple computing processing units perform parallel reduction on multiple sub-tensors based on a recursive algorithm to obtain a reduced result tensor;

[0081] Preferably, the parallel reduction step S300 of sub-tensors includes:

[0082] Multiple computing processing units perform recursive decomposition operations on the assigned sub-tensors in parallel until the dimension of the sub-tensor is not greater than a set threshold; perform a reduction operation on the sub-tensor that reaches the set threshold to obtain a reduced result tensor, and record the reduced result tensor at the target storage location.

[0083] Figure 2Schematic diagram for reducing the recurrence equivalence relationship of high-dimensional tensors. As shown in Figure 2 FIG. [0000187], in some embodiments, the computing and processing unit performs parallel reduction on the segmented sub-tensors, and the reduction problem can be solved by a recursive algorithm. Specifically, the reduction problem of an n-dimensional tensor can be recursively transformed into the solution of sub-problems of n - 1 dimensional tensors. If the sub-tensor is a single element in all dimensions, it is directly returned; otherwise, it is recursively processed until the dimension of the sub-tensor is less than or equal to a set threshold, usually 3 dimensions, and then these sub-tensors with dimensions less than 3 are combined through a reduction operation.

[0084] In some embodiments, still taking the original tensor T with dimensions [3, N, C, H, W] as an example, the dimensions of the sub-tensor split_src_shape after splitting along the non-reduction dimension are [3, 1, S, H, W], where the 0th dimension is the reduction dimension. The reduction problem of the 5-dimensional tensor [3, 1, S, H, W] can be recursively transformed into the solution of sub-problems of reducing 3 4-dimensional tensors [1, S, H, W].

[0085] Figure 3 Schematic diagram for parallel reduction of high-dimensional tensors. As shown in Figure 3 FIG. [0000195], in some embodiments, taking the slave core 0 as an example, during the reduction calculation, the slave core 0 sequentially traverses the 3 4-dimensional tensors obtained by transforming split_src_shape, accumulates the reduction operations into the local memory, and the dimensions of the local memory are [1, 1, S, H, W]. That is to say, in the reduction dimension, these elements will be accumulated into one result, and finally the obtained result is written back to the corresponding target storage location. This design allows data segmentation and reduction to be processed simultaneously in one traversal, improving the computing efficiency.

[0086] Preferably, the sub-tensor parallel reduction step S300 further includes:

[0087] Reduction index calculation step: During the reduction process, based on the source data offset and data stride, calculate the reduction index of each element in the reduction dimension in real time, and record the reduction result tensor and the corresponding reduction index in the target storage location together.

[0088] Introducing the calculation of the reduction index not only returns the reduction result, but also returns the index of each element in the reduction dimension, providing richer information for subsequent data analysis and processing to achieve data tracking and traceability, which helps to better understand and optimize the algorithm, and also helps to locate potential errors and performance bottlenecks.

[0089] In some embodiments, the reduction index can be calculated from the source data offset offset and data stride src_stride. The specific algorithm steps are as follows:

[0090] Starting from the first dimension, loop until the reduction dimension (excluding): Calculate the data offset divided by the data stride of the current dimension offset / src_stride[i] to obtain the index in the current dimension i; Multiply this index by the stride of the current dimension src_stride[i] to obtain the offset contributed by this dimension; Subtract this offset from the data offset offset; Finally, divide the remaining offset by the stride of the reduction dimension to obtain the index of the reduction dimension.

[0091] In some embodiments, assume there is a 3D tensor (2, 3, 4), the reduction dimension is 1 (the middle dimension), the source data offset is 14, and the data stride src_stride is [12, 4, 1]. Calculate the index in the reduction dimension 1 for the offset 14 (corresponding to the element [1, 0, 2]).

[0092] Calculate the index in the current dimension dimension 0 as 14 / 12, multiply this index by the stride of the current dimension to obtain the offset contributed by this dimension (14 / 12)*12, subtract this offset from the data offset offset to obtain the remaining offset as 14 - (14 / 12)*12 = 2, and finally divide the remaining data offset offset by the stride of the reduction dimension to obtain the index of the reduction dimension as 2 / 4 = 0, returning 0, which correctly indicates the index in the reduction dimension (dimension 1).

[0093] Reduction result tensor write-back step S400: Write the reduction result tensor back to the main memory.

[0094] Preferably, the reduction result tensor write-back step S400 includes:

[0095] If a single sub-tensor is processed by a single computing processing unit, the first slave core of each computing processing unit writes the reduction result tensor back to the main memory; This avoids conflicts that may be caused by multiple slave cores writing simultaneously, simplifies the write-back logic, and improves the write-back efficiency.

[0096] If a single sub-tensor is processed by multiple computing processing units, sequentially traverse the multiple computing processing units participating in the processing of this sub-tensor. The first slave core of each computing processing unit writes the reduction result tensor back to the main memory and performs a memory synchronization operation after each write-back. This ensures the orderliness of data write-back. At the same time, performing a memory synchronization operation after each write-back guarantees data consistency, enabling subsequent processing to accurately obtain the reduction result and avoiding errors caused by data asynchronization.

[0097] Dynamically adjust the write-back strategy according to the processing method of sub-tensors, flexibly cope with tensor reduction tasks of different scales and complexities, be applicable to various computing scenarios, and improve the adaptability and robustness of the system. At the same time, make full use of the parallel computing power of the multi-core architecture, and maximize the advantages of parallel processing through reasonable task division and write-back coordination.

[0098] Embodiment 2:

[0099] Figure 4 It is a schematic structural diagram of a parallel computing system for tensor reduction in an embodiment of the present invention. As Figure 4 shown, the present invention provides a parallel computing method for tensor reduction, including a tensor splitting module 100, a sub-tensor allocation module 200, a sub-tensor parallel reduction module 300, and a reduced result tensor write-back module 400.

[0100] Tensor splitting module 100: Dynamically split a high-dimensional tensor based on non-reduced dimensions to obtain multiple sub-tensors. Specifically, non-reduced splitting dimensions can be found based on the reduced high-dimensional tensor.

[0101] Sub-tensor allocation module 200: Allocate multiple sub-tensors to multiple computing processing units based on the number of sub-tensors; the allocation strategy is dynamically adjusted according to the number of sub-tensors and the number of slave cores to ensure load balancing and efficient use of computing resources.

[0102] Sub-tensor parallel reduction module 300: Multiple computing processing units perform parallel reduction on multiple sub-tensors based on a recursive algorithm to obtain a reduced result tensor.

[0103] Reduced result tensor write-back module 400: Write the reduced result tensor back to the main memory.

[0104] The tensor splitting module 100 includes:

[0105] Traverse all non-reduced dimensions from the last dimension of the high-dimensional tensor forward to the first dimension that satisfies that the storage space requirement of the sub-tensor formed from this dimension to the last dimension is not less than the local memory capacity as the splitting dimension;

[0106] Obtain the splitting step size based on the splitting dimension and the local memory capacity;

[0107] Split the high-dimensional tensor based on the splitting dimension and the splitting step size to obtain multiple tensors.

[0108] Embodiment 3:

[0109] Combined with Figure 5 shown, this embodiment discloses a specific implementation manner of a computer device. The computer device may include a processor 41 and a memory 42 storing computer program instructions.

[0110] Specifically, the above-mentioned processor 41 may include a central processing unit (CPU), or an application specific integrated circuit (ASIC), or may be configured as one or more integrated circuits for implementing the embodiments of the present application.

[0111] Among them, the memory 42 may include a mass memory for data or instructions. By way of example and not limitation, the memory 42 may include a hard disk drive (HDD), a floppy disk drive, a solid state drive (SSD), a flash memory, an optical disc, a magneto-optical disc, a magnetic tape, or a universal serial bus (USB) drive, or a combination of two or more of these. In appropriate cases, the memory 42 may include removable or non-removable (or fixed) media. In appropriate cases, the memory 42 may be internal or external to the data processing device. In a particular embodiment, the memory 42 is a non-volatile memory. In a particular embodiment, the memory 42 includes a read-only memory (ROM) and a random access memory (RAM). In appropriate cases, the ROM may be a mask-programmed ROM, a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), an electrically alterable ROM (EAROM), or a flash memory, or a combination of two or more of these. In appropriate cases, the RAM may be a static random access memory (SRAM) or a dynamic random access memory (DRAM), where the DRAM may be a fast page mode dynamic random access memory (FPMDRAM), an extended date out dynamic random access memory (EDODRAM), a synchronous dynamic random access memory (SDRAM), etc.

[0112] The memory 42 can be used to store or cache various data files required for processing and / or communication, as well as possible computer program instructions executed by the processor 41.

[0113] The processor 41 reads and executes the computer program instructions stored in the memory 42 to implement the parallel computing method for tensor reduction in the above embodiments.

[0114] In some of the embodiments, the computer device may further include a communication interface 43 and a bus 40. Among them, as Figure 5 shown, the processor 41, the memory 42, the communication interface 43 and the bus 40 are connected and communicate with each other.

[0115] The communication interface 43 is used to implement communication between the various modules, devices, units and / or devices in the embodiments of the present application.

[0116] The communication interface 43 can also implement data communication with other components such as external devices, image / data acquisition devices, databases, external storage, and image / data processing workstations, etc.

[0117] The bus 40 includes hardware, software, or both, and couples components of a computer device to each other. The bus 40 includes, but is not limited to, at least one of the following: a data bus, an address bus, a control bus, an expansion bus, and a local bus. By way of example and not limitation, the bus 40 may include an Accelerated Graphics Port (AGP) or other graphics bus, an Extended Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), a Hyper Transport (HT) interconnect, an Industry Standard Architecture (ISA) bus, an InfiniBand interconnect, a Low Pin Count (LPC) bus, a memory bus, a Micro Channel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCI-X) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local Bus (VLB) bus, or other suitable bus or a combination of two or more of these. In suitable cases, the bus 40 may include one or more buses. Although the embodiments of the present application describe and illustrate specific buses, the present application contemplates any suitable bus or interconnect.

[0118] In addition, in combination with the parallel computing method for tensor reduction in the above embodiments, an embodiment of the present application can provide a computer-readable storage medium to implement. Computer program instructions are stored on the computer-readable storage medium; when the computer program instructions are executed by a processor, any one of the parallel computing methods for tensor reduction in the above embodiments is implemented.

[0119] The above are only the preferred embodiments of the present invention, and do not limit the present invention in other forms. Any person skilled in the art may use the technical content disclosed above to make changes or modifications into equivalent embodiments with equivalent changes and apply them to other fields. However, as long as it does not depart from the technical solution content of the present invention, any simple modification, equivalent change and modification made to the above embodiments according to the technical essence of the present invention still belong to the protection scope of the technical solution of the present invention.

Claims

1. A parallel computing method for tensor reduction, characterized in that: include: Tensor segmentation step: dynamically segment the high-dimensional tensor based on the unreduced dimension to obtain multiple sub-tensors; A sub-tensor allocation step: allocating a plurality of the sub-tensors to a plurality of computing processing units based on the number of the sub-tensors; Sub-tensor parallel reduction step: multiple computing processing units perform parallel reduction on multiple sub-tensors based on a recursive algorithm to obtain a reduction result tensor; The reduction result tensor writing back step: writing the reduction result tensor back to the main memory; Wherein, the tensor segmentation step comprises: Traverse all non-reduced dimensions from the last dimension of the high-dimensional tensor forward to the first dimension that satisfies the storage space requirement of the sub-tensor formed from the dimension to the last dimension is not less than the local memory capacity as the splitting dimension; Acquire a segmentation step length based on the segmentation dimension and the local memory capacity; Splitting the high-dimensional tensor based on the segmentation dimension and the segmentation step to obtain a plurality of sub-tensors; The sub-tensor allocation step comprises: If the number of the sub-tensors is greater than the total number of slave cores, allocating the plurality of sub-tensors to the plurality of slave cores based on a cyclic allocation; If the number of the sub-tensors is less than the total number of slave cores, the sub-tensors are split twice based on the reduced dimension, the slave cores are divided into working groups based on the number of sub-tensors after the second split, and multiple sub-tensors after the second split are allocated to multiple working groups based on cyclic allocation; each working group processes one sub-tensor after the second split.

2. The parallel computing method for tensor reduction according to claim 1, characterized in that: The sub-tensor parallel reduction step comprises: The plurality of computing processing units perform recursive decomposition operations on the allocated sub-tensors in parallel until the dimension of the sub-tensors is no greater than a set threshold; a reduction operation is performed on the sub-tensors that reach the set threshold to obtain a reduced result tensor, and the reduced result tensor is recorded in a target storage location.

3. The parallel computing method for tensor reduction according to claim 2, characterized in that: The sub-tensor parallel reduction step further includes: Reduction index calculation step: During the reduction process, the reduction index of each element in the reduction dimension is calculated in real time based on the source data offset and the data stride, and the reduction result tensor and the corresponding reduction index are recorded in the target storage location.

4. The parallel computing method of tensor reduction according to claim 1, characterized in that: The step of writing back the reduced result tensor includes: If a single sub-tensor is processed by a single computing processing unit, the first slave core of each computing processing unit writes the reduced result tensor back to the main memory; If a single sub-tensor is processed by multiple computing processing units, the multiple computing processing units involved in the processing of the sub-tensor are traversed in turn, and the first slave core of each computing processing unit writes the reduced result tensor back to the main memory, and performs a memory synchronization operation after each write back.

5. A tensor-reduced parallel computing system, characterized in that: A parallel computing method for tensor reduction applied to any one of claims 1 to 4, comprising: Tensor segmentation module: dynamically segment high-dimensional tensors based on non-reduced dimensions to obtain multiple sub-tensors; A sub-tensor allocation module: allocating a plurality of the sub-tensors to a plurality of computing processing units based on the number of the sub-tensors; Sub-tensor parallel reduction module: multiple computing processing units perform parallel reduction on multiple sub-tensors based on a recursive algorithm to obtain a reduction result tensor; A reduction result tensor write-back module: writes the reduction result tensor back to the main memory; Wherein, the tensor segmentation module includes: Traverse all non-reduced dimensions from the last dimension of the high-dimensional tensor forward to the first dimension that satisfies the storage space requirement of the sub-tensor formed from the dimension to the last dimension is not less than the local memory capacity as the splitting dimension; Acquire a segmentation step length based on the segmentation dimension and the local memory capacity; Splitting the high-dimensional tensor based on the segmentation dimension and the segmentation step to obtain a plurality of sub-tensors; The sub-tensor allocation module includes: If the number of the sub-tensors is greater than the total number of slave cores, allocating the plurality of sub-tensors to the plurality of slave cores based on a cyclic allocation; If the number of sub-tensors is less than the total number of slave cores, the sub-tensors are divided twice based on the reduced dimension, the slave cores are divided into work groups based on the number of sub-tensors after the second division, and multiple sub-tensors after the second division are allocated to multiple work groups based on cyclic allocation; each work group processes one sub-tensor after the second division.

6. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the parallel computing method for tensor reduction as claimed in any one of claims 1 to 4 is implemented.

7. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the parallel computing method for tensor reduction as claimed in any one of claims 1 to 4 is implemented.

Citation Information

Patent Citations

  • Operator segmentation method and device and operator compiling system

    CN118277711A