Inter-core data exchange method and apparatus, electronic device, and storage medium

CN122817142APending Publication Date: 2026-09-25SHANGHAI LIXIANG AUTOMOBILE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510358958.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-24
Publication Date
2026-09-25

AI Technical Summary

Benefits of technology

[0015]本申请实施例提供的核间数据交换方法、装置、电子设备及存储介质,通过根据源地址中数据张量的第一切分参数和目标地址中数据张量的第二切分参数,确定目标地址中与源地址中的各张量块分别对应的目标核心和目标张量块地址,在第一切分轴大于第二切分轴的情况下,遍历源地址中各切分数据下的各张量块,并针对遍历到的当前张量块,基于带宽容量从源地址中读取当前张量块的数据,并将读取到的数据写入到与当前张量块对应的目标核心中的目标张量块地址中,对写入各目标张量块地址中的张量块进行转置操作,并对各目标核心中的切分数据分别进行转置操作,由于进行核间数据交换(即将源地址中的张量块写入目标核心并使得写入目标核心的数据符合数据排布要求)时,不需要将数据张量写出到外部存储,可以从源地址遍历读取张量块内的数据并写入目标核心,避免了通过写出到外部存储引入的巨大延迟,而且在第一切分轴大于第二切分轴时,通过直接传输张量块内符合带宽容量的数据,并在传输完成后对张量块进行转置操作,之后再对目标核心下的切分数据进行整体的转置操作,使得目标地址中的张量块在所有转置操作前变得连续,从而可以直接读取张量块内符合带宽容量的多行数据进行传输,而不需要单独传输张量块内每一行的数据,从而提高了带宽利用率,而通过两次转置操作使得目标核心下的数据符合目标地址的数据排布要求。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122817142A_ABST
    Figure CN122817142A_ABST
Patent Text Reader

Abstract

Embodiments of the present application disclose a method and device for inter-core data exchange, electronic equipment and storage medium. The method comprises: determining target cores and target tensor block addresses corresponding to each tensor block in a source address according to a first splitting parameter of a data tensor in the source address and a second splitting parameter of a data tensor in a target address; in the case that a first splitting axis is greater than a second splitting axis, traversing each tensor block under each split data in the source address, and for a current tensor block traversed, reading data of the current tensor block from the source address based on bandwidth capacity, and writing the read data into a target tensor block address in a target core corresponding to the current tensor block; performing a transpose operation on each tensor block written into each target tensor block address, and performing a transpose operation on each split data in each target core. Embodiments of the present application avoid the huge delay introduced by writing out to external storage, and improve the bandwidth utilization.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of processor technology, and in particular to an inter-core data exchange method, apparatus, electronic device, and storage medium. Background Technology

[0002] With the development of large-scale and autonomous driving models, the demands on chip computing power are increasing. NPU (Neural-network Processing Unit) chips and GPGPU (General-purpose computing on graphics processing units) chips account for a large share of the market. Especially as autonomous driving models become more mature, their sensitivity to latency is increasing. Therefore, both NPU and GPGPU architectures experience high latency in accessing external / global memory. Thus, GPGPU / NPU chips should minimize interaction with external memory units during computation to avoid introducing excessive latency.

[0003] However, due to the limited memory access capacity within GPGPU / NPUs and the inconsistent data parallelism among operators, inconsistencies in the way data tensors are partitioned are common. This leads to differences in memory layout between the input and output tensors of operators with topological relationships in neural networks. The traditional solution is to write the data tensors to external storage, rearrange the data, and then load them into internal storage via the AXI (Advanced eXtensible Interface) bus, which introduces significant latency. Summary of the Invention

[0004] This application provides an inter-core data exchange method, apparatus, electronic device, and storage medium, which helps to avoid the huge latency introduced by writing to external storage and can improve bandwidth utilization.

[0005] To address the aforementioned problems, in a first aspect, embodiments of this application provide an inter-core data exchange method, comprising:

[0006] Based on the first partitioning parameter of the data tensor in the source address and the second partitioning parameter of the data tensor in the target address, the target core and target tensor block addresses corresponding to each tensor block in the source address are determined respectively. The data tensor includes multiple tensor blocks, and the target tensor block address is the address of the tensor block mapped to the target address.

[0007] When the first segmentation axis in the first segmentation parameter is greater than the second segmentation axis in the second segmentation parameter, the tensor blocks under each segmented data in the source address are traversed, and for the current tensor block that is traversed, the data of the current tensor block is read from the source address based on the bandwidth capacity, and the read data is written to the target tensor block address in the target core corresponding to the current tensor block. The segmented data in the source address is the data obtained by segmenting the data tensor according to the first segmentation parameter.

[0008] The tensor blocks written to the addresses of each of the target tensor blocks are transposed, and the segmented data in each of the target cores are also transposed.

[0009] Secondly, embodiments of this application provide an inter-core data exchange apparatus, comprising:

[0010] The data mapping module is used to determine the target core and target tensor block addresses corresponding to each tensor block in the source address based on the first partitioning parameter of the data tensor in the source address and the second partitioning parameter of the data tensor in the target address. The data tensor includes multiple tensor blocks, and the target tensor block address is the address of the tensor block mapped to the target address.

[0011] The data exchange module is used to traverse each tensor block under each segmented data in the source address when the first segmentation axis in the first segmentation parameter is greater than the second segmentation axis in the second segmentation parameter, and for the current tensor block that has been traversed, read the data of the current tensor block from the source address based on the bandwidth capacity, and write the read data into the target tensor block address in the target core corresponding to the current tensor block. Each segmented data in the source address is data obtained by segmenting the data tensor according to the first segmentation parameter.

[0012] The transpose module is used to transpose the tensor blocks written to the addresses of each of the target tensor blocks, and to transpose the segmented data in each of the target cores.

[0013] Thirdly, embodiments of this application also provide an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the inter-core data exchange method described in embodiments of this application.

[0014] Fourthly, embodiments of this application provide a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the inter-core data exchange method disclosed in embodiments of this application.

[0015] The inter-core data exchange method, apparatus, electronic device, and storage medium provided in this application determine the target core and target tensor block addresses corresponding to each tensor block in the target address and the source address based on the first partitioning parameter of the data tensor in the source address and the second partitioning parameter of the data tensor in the target address. When the first partitioning axis is greater than the second partitioning axis, the tensor blocks under each partitioned data in the source address are traversed. For the current tensor block being traversed, data of the current tensor block is read from the source address based on bandwidth capacity, and the read data is written to the target tensor block address in the target core corresponding to the current tensor block. The tensor blocks written to each target tensor block address are transposed, and the partitioned data in each target core is also transposed. Because inter-core data exchange (i.e., writing tensor blocks from the source address to the target tensor block address) is performed, the data exchange is achieved. When writing data to the target core (ensuring the data layout meets the requirements), it is not necessary to write the data tensor to external storage. Data can be read from the source address and written to the target core, avoiding the huge latency introduced by writing to external storage. Moreover, when the first split axis is greater than the second split axis, by directly transmitting data within the tensor block that meets the bandwidth capacity, and performing a transpose operation on the tensor block after transmission, and then performing a whole transpose operation on the split data under the target core, the tensor block in the target address becomes continuous before all transpose operations. This allows for direct reading and transmission of multiple rows of data within the tensor block that meet the bandwidth capacity, without needing to transmit each row of data within the tensor block separately, thereby improving bandwidth utilization. Furthermore, the two transpose operations ensure that the data under the target core meets the data layout requirements of the target address. Attached Figure Description

[0016] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0017] Figure 1 This is a flowchart of an inter-core data exchange method provided in an embodiment of this application;

[0018] Figure 2 This is an example diagram illustrating the mapping relationship between the source and destination addresses of a tensor block when the first segmentation axis is greater than the second segmentation axis in an embodiment of this application.

[0019] Figure 3 This is an example diagram illustrating data exchange in an embodiment of this application when the first dividing axis is greater than the second dividing axis;

[0020] Figure 4This is a flowchart of an inter-core data exchange method provided in an embodiment of this application;

[0021] Figure 5 This is a schematic diagram of the structure of an inter-core data exchange device provided in an embodiment of this application;

[0022] Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0023] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0024] This application addresses the issue that when the first partition axis of the data tensor at the source address is greater than the second partition axis of the data tensor at the destination address, data sent from a source core to a destination core is discontinuous on the destination side, leading to significant waste of hardware bandwidth. To ensure full utilization of hardware bandwidth, this application inserts two transpose operations after transmission, making the previously discontinuous data continuous on the destination side. This application resolves the problems of significant latency introduced by writing to external storage and insufficient bandwidth utilization in specific scenarios where the first partition axis of the data tensor at the source address is greater than the second partition axis of the data tensor at the destination address.

[0025] Figure 1 This is a flowchart illustrating an inter-core data exchange method provided in an embodiment of this application. This inter-core data exchange method is applicable to scenarios involving the exchange of data between cores within a processor. Figure 1 As shown, the method includes steps 110 to 130.

[0026] Step 110: Based on the first partitioning parameter of the data tensor in the source address and the second partitioning parameter of the data tensor in the target address, determine the target core and target tensor block addresses corresponding to each tensor block in the source address, respectively. The data tensor includes multiple tensor blocks.

[0027] The target tensor block address is the address that the tensor block maps to in the target address.

[0028] In an exemplary embodiment, the first segmentation parameter may include a first segmentation axis and a first segmentation quantity. The first segmentation axis is the segmentation axis of the data tensor in the source address, and the first segmentation quantity is the number of segments of the data tensor in the source address. The second segmentation parameter may include a second segmentation axis and a second segmentation quantity. The second segmentation axis is the segmentation axis of the data tensor in the target address, and the second segmentation quantity is the number of segments of the data tensor in the target address. The data tensor is segmented based on the first segmentation parameter to obtain the first number of segmented data. Each segmented data is stored in a source core, meaning the source address includes the addresses in the first number of source cores. The second segmentation parameter characterizes the data arrangement requirements in the target address. Each segmented data obtained by segmenting the data tensor based on the second segmentation parameter needs to be stored in a target core, and the target address includes the addresses in the second number of target cores.

[0029] In an exemplary embodiment, the tensor block is a common tensor block included after the data tensor is segmented according to the first segmentation parameter and the second segmentation parameter, respectively. For example, when the first segmentation axis in the first segmentation parameter and the second segmentation axis in the second segmentation parameter are the same, and the first segmentation number in the first segmentation parameter is greater than or equal to the second segmentation number in the second segmentation parameter, the tensor block is a segmented data obtained after segmenting the data tensor according to the first segmentation parameter, and this segmented data is also included in the segmented data obtained after segmenting the data tensor according to the second segmentation parameter; when the first segmentation axis in the first segmentation parameter and the second segmentation axis in the second segmentation parameter are the same, and the first segmentation number in the first segmentation parameter is less than or equal to the second segmentation number in the second segmentation parameter, the tensor block is a segmented data obtained after segmenting the data tensor according to the second segmentation parameter, and this segmented data is also included in the segmented data obtained after segmenting the data tensor according to the first segmentation parameter. For example, when the first segmentation axis in the first segmentation parameter and the second segmentation axis in the second segmentation parameter are inconsistent, the tensor block is the data obtained by segmenting the data tensor according to the first segmentation parameter (i.e., the data located on a source core, and the source address corresponds to multiple source cores), and then segmenting it again according to the second segmentation parameter. For example, the shape of the data tensor on a single source core (i.e., a segmented data obtained by segmenting the data tensor according to the first segmentation parameter) is 1x1x30x256, and the segmentation axis on the destination address is 3, and it is segmented into 32 parts. Then the shape of the tensor block is 1x1x30x256 / 32, that is, the size of the block is 30x8.

[0030] In one exemplary embodiment, the data tensor at the source address can be either already prepared data or a data stream. A data stream can be interpreted as data transfer starting before the data in the tensor block is fully ready, requiring synchronization via a synchronization signal.

[0031] In one exemplary embodiment, a source address corresponds to multiple source cores, and a target address corresponds to multiple target cores. The source cores and target cores are cores located in the same processor. The processor can be, for example, a CPU, GPU, NPU, etc.

[0032] In one exemplary embodiment, a mapping relationship between each source core and each target core is determined based on a first segmentation parameter and a second segmentation parameter. Then, based on this mapping relationship, tensor blocks in the source cores can be mapped to the target tensor block address of the corresponding target core. The target tensor block address corresponding to the tensor block refers to the target address where the tensor block needs to be written.

[0033] Step 120: When the first segmentation axis in the first segmentation parameter is greater than the second segmentation axis in the second segmentation parameter, traverse each tensor block under each segmentation data in the source address, and for the current tensor block that has been traversed, read the data of the current tensor block from the source address based on the bandwidth capacity, and write the read data into the target tensor block address in the target core corresponding to the current tensor block.

[0034] The data segments in the source address are obtained by segmenting the data tensor according to the first segmentation parameter.

[0035] In one exemplary embodiment, each segment of data in the source address is stored in a different source core.

[0036] In an exemplary embodiment, when the first slicing axis in the first slicing parameter is greater than the second slicing axis in the second slicing parameter, a tensor block is the data obtained by slicing the data tensor according to the first slicing parameter (i.e., data located on one source core, with the source address corresponding to multiple source cores), and then slicing it again according to the second slicing parameter. That is, a tensor block is the data obtained by slicing the data tensor into a first number of slicing data according to the first slicing axis, and then slicing each slicing data again according to the second slicing axis and slicing it into a second number of slicing data, where one piece of data is a tensor block.

[0037] Figure 2 This is an example diagram illustrating the mapping relationship between the source and destination addresses of a tensor block when the first segmentation axis is greater than the second segmentation axis in an embodiment of this application. Figure 2As shown, the shape (original size) of the data tensor is 1x32x60x256, the first partition axis is 3, and the first partition number is 8. That is, the shape of the data tensor on a single source core (i.e., a partitioned data) is 1x32x60x32. The second partition axis is 1, and the second partition number is 32. Based on this second partition parameter, the data tensor on a single source core is partitioned again, resulting in a tensor block with a shape of 1x32 / 32x60x32. That is, the size of the tensor block is 60x32, meaning the tensor block contains 1920 elements (the specific elements are not shown in the diagram). Figure 2 The middle arrow represents the mapping relationship between the source address and the destination address of the tensor block.

[0038] In one exemplary embodiment, the first split axis is larger than the second split axis, meaning the first and second split axes are not aligned. In this case, the data mapping relationship between the source address and the target address is a full mapping, meaning the split data in each source core is mapped to all target cores. For example... Figure 2 As shown, in the source address, the rows of a tensor block are stored contiguously; while in the destination address, the rows of a tensor block are not stored contiguously because the rows of a tensor block are separated by rows of data from other tensor blocks. Therefore, the data within a tensor block in the source address is contiguous, while the data within a tensor block in the destination address is discontinuous.

[0039] In an exemplary embodiment, since the data inside the tensor block in the target address is discontinuous, each line of data in the tensor block needs to be transmitted separately when transmitting data, resulting in low bandwidth utilization. In order to improve bandwidth utilization, the data inside the tensor block in the target address needs to be made continuous. At this time, the data in the tensor block in the source address can be transmitted first, and then the transmitted data can be transposed twice to make the tensor block in the target address continuous before the transpose operation. Figure 3 This is an example diagram illustrating data exchange in an embodiment of this application when the first dividing axis is larger than the second dividing axis, as shown below. Figure 3 As shown, the tensor blocks in the segmented data are transmitted, and the tensor blocks transmitted to the target address are transposed. Then, the data on each target core is transposed again, so that the final arrangement of the data tensors in the target address meets the requirements.

[0040] In an exemplary embodiment, when the first partition axis is greater than the second partition axis, each piece of partitioned data in the source address can be transmitted to the target address separately. When transmitting one piece of partitioned data, each tensor block within that partitioned data is transmitted separately. When transmitting a tensor block under that partitioned data, the data of that tensor block is read based on the bandwidth capacity, and the read data is written to the target tensor block address. When reading and transmitting data of a tensor block, it can be read and transmitted according to the bandwidth capacity. That is, if the bandwidth capacity is greater than or equal to the data size of the tensor block, the entire data in the tensor block can be read directly, and the entire data of the tensor block can be written to the target tensor block address; if the bandwidth capacity is less than the data size of the tensor block, the bandwidth capacity data can be read from the tensor block, and the read data can be written to the target tensor block address. Then, data is read from the tensor block and written to the target tensor block address each time based on the bandwidth capacity, until all the data in the tensor block is written to the target tensor block address. Here, the bandwidth capacity refers to the bandwidth capacity for data transmission between the cores within the processor.

[0041] In some embodiments of this application, the step of traversing each tensor block under each segmented data in the source address, for the current tensor block being traversed, reading data of the single data length from the source address based on the bandwidth capacity, and writing the read data into the target tensor block address in the target core corresponding to the current tensor block, may include:

[0042] Based on the address compensation value between two adjacent segments of data in the source address, traverse each of the segmented data;

[0043] For the currently segmented data that has been traversed, each tensor block in the currently segmented data is traversed according to the address compensation value between two adjacent tensor blocks in the currently segmented data;

[0044] For the current tensor block traversed in the current segmented data, the data of the current tensor block is read from the source address based on the bandwidth capacity, and the read data is written to the target tensor block address in the target core corresponding to the current tensor block.

[0045] In an exemplary embodiment, transmitting the segmented data from the source address to the target address can be accomplished using a three-layer loop. The first layer, the outermost loop, iterates through each segmented data based on the address compensation value (source stride3) between adjacent segments in the source address, transmits the read segmented data, and writes the data to the target address based on the address compensation value (destination stride3) of each segmented data in the target address. The second layer, based on the address compensation value (source stride2) between adjacent tensor blocks within the segmented data, iterates through each tensor block within a segmented data, transmits the read tensor block, and writes the tensor block to the target address based on the address compensation value (destination stride2) of each adjacent tensor block within the segmented data in the target address. The third layer, the innermost loop, is used to read the tensor blocks within a segmented data based on the address compensation value (source stride2) between adjacent tensor blocks within the segmented data. Stride1 iterates through the tensor blocks, reading data in a loop. It reads data of the appropriate length from the tensor block according to bandwidth capacity, and writes the read data back to the target tensor block address corresponding to that tensor block based on the address compensation value between rows within the tensor block at the target address (destination Stride1). Stride1, Stride2, and Stride3 refer to the source address. Stride1 describes the address compensation relationship between rows within a tensor block, Stride2 describes the address compensation relationship between tensor blocks within the same core, and Stride3 describes the address compensation relationship between tensor blocks across cores.

[0046] The specific process of the above three-level loop can be represented as follows:

[0047] Step A: Based on the address compensation value (source stride3) between each adjacent segmented data in the source address, traverse and read one segmented data as the current segmented data;

[0048] Step B: Based on the address compensation value (source Stride2) between each adjacent tensor block under the segmented data, traverse and read a tensor block under the current segmented data as the current tensor block;

[0049] Step C: Based on the address compensation value between rows within the tensor block (Source Stride1), traverse and read the data under the current tensor block. At this time, data of the corresponding length can be read from the tensor block according to the bandwidth capacity and used as the current transmission data.

[0050] Step D: Based on the address compensation value (destinationStride1) between rows within the tensor block in the target address, write the current transmission data into the target tensor block address corresponding to the tensor block in the target address;

[0051] Step E, repeat steps C and D until the data transfer under the current tensor block is complete;

[0052] Step F involves repeatedly executing steps B, C, D, and E to take the next tensor block as the current tensor block. When writing to the target address, the tensor block to be transmitted to the target address is written based on the address compensation value (destination Stride2) of each adjacent tensor block under the segmented data in the target address, until the data transmission of all tensor blocks under the current segmented data is completed.

[0053] Step G involves repeatedly executing steps A, B, C, D, E, and F to use the next piece of split data as the current split data. When writing to the target address, the data to be transmitted to the target address is written based on the address compensation value (destination stride3) of each split data in the target address, until the data transmission of all split data in the source address is completed.

[0054] By using a three-layer loop to traverse and read data, there is no need to write the data tensor to external storage for rearrangement, which reduces the latency of the chip in model inference and improves the chip's data throughput and utilization.

[0055] Step 130: Transpose the tensor blocks written to the addresses of each target tensor block, and transpose the segmented data in each target core.

[0056] In an exemplary embodiment, after all data reading and writing are completed, it can be observed that the data tensor layout in the target address does not meet the requirements. It is necessary to perform a transpose operation on each tensor block in each target tensor block address separately, and then perform an overall transpose operation on the segmented data in each target core to obtain the target data tensor layout.

[0057] by Figure 2Taking a tensor block as an example, assuming the bandwidth is 100 and the size of a tensor block is 60x32, which includes 1920 elements, without transposition, each row in the tensor block needs to be transmitted separately, transmitting 32 elements each time, resulting in a waste of bandwidth resources. However, through the technical solution of this application embodiment, the data of the entire tensor block is transmitted directly, transmitting the data that the bandwidth can hold each time, that is, 100 elements each time. One tensor block can be transmitted 20 times, which greatly improves the bandwidth utilization. After the transmission is completed, the tensor blocks are transposed, and then the segmented data under each target core is transposed as a whole, so that the data tensor under the target address meets the arrangement requirements.

[0058] The inter-core data exchange method provided in this application determines the target core and target tensor block addresses corresponding to each tensor block in the target address and the source address based on the first partitioning parameter of the data tensor in the source address and the second partitioning parameter of the data tensor in the target address. When the first partitioning axis is greater than the second partitioning axis, it traverses each tensor block under each partitioned data in the source address. For the current tensor block being traversed, it reads the data of the current tensor block from the source address based on bandwidth capacity and writes the read data into the target tensor block address in the target core corresponding to the current tensor block. It then performs a transpose operation on the tensor blocks written to each target tensor block address and performs a transpose operation on the partitioned data in each target core. Since inter-core data exchange is performed... There's no need to write the data tensor to external storage. Data within the tensor block can be read from the source address and written to the target core, avoiding the significant latency introduced by writing to external storage. Furthermore, when the first split axis is greater than the second split axis, by directly transmitting data within the tensor block that meets the bandwidth capacity, and then transposing the tensor block after transmission, followed by a complete transpose operation on the split data under the target core, the tensor block at the target address becomes contiguous before all transpose operations. This allows for direct reading and transmission of multiple rows of data within the tensor block that meet the bandwidth capacity, without needing to transmit each row of data within the tensor block separately, thus improving bandwidth utilization. The transpose operation ensures that the data under the target core conforms to the data arrangement requirements of the target address.

[0059] Based on the above technical solution, determining the target core and target tensor block addresses corresponding to each tensor block in the source address according to the first partitioning parameter of the data tensor in the source address and the second partitioning parameter of the data tensor in the target address may include:

[0060] Based on the first segmentation parameter and the second segmentation parameter, the mapping relationship between each source core and each target core is determined. Each source core is a processor core corresponding to the source address. Each source core stores the segmentation data. Each target core is a processor core corresponding to the target address.

[0061] The shape and size of the tensor block are determined based on the first segmentation parameter and the second segmentation parameter;

[0062] Based on the shape and size of the tensor blocks and the mapping relationship, the target core and target tensor block addresses corresponding to each tensor block in the source address are determined.

[0063] In an exemplary embodiment, the data tensor is segmented according to a first segmentation parameter of the data tensor in the source address, and each segmented data is stored in a source core. When exchanging data, a second segmentation parameter is used to transform the arrangement of the data tensor in the source address into the data tensor arrangement required by the target address, and write it into each target core corresponding to the target address. In this way, a mapping relationship between each source core and each target core can be established based on the first and second segmentation parameters, that is, to determine which target core each source core is mapped to.

[0064] In an exemplary embodiment, the data mapping relationship between the source address and the target address is determined to be a partial mapping or a full mapping based on whether the first split axis and the second split axis are consistent. When the first split axis and the second split axis are consistent, the data mapping relationship between the source address and the target address is a partial mapping; when the first split axis and the second split axis are inconsistent, the data mapping relationship between the source address and the target address is a full mapping.

[0065] For example, when the data mapping relationship between the source address and the target address is a full mapping, the data tensor is segmented according to the first segmentation axis in the first segmentation parameter of the data tensor in the source address, dividing the data tensor into a first number of segmented data. Then, each segmented data is further segmented based on the second segmentation axis and the second number of segmentations in the second segmentation parameter, that is, each segmented data is divided into a second number of tensor blocks based on the second segmentation axis, thus obtaining the shape and size of the tensor blocks. Figure 2 Taking the data tensor shown as an example, such as Figure 2As shown, the shape of the data tensor is 1x32x30x256, the first partition axis is 3, and the first partition number is 8, meaning that the shape of the data tensor on a single source core (i.e., a partitioned data) is 1x32x30x32. The second partition axis is 1, and the second partition number is 32. Based on this second partition parameter, the data tensor on a single source core is partitioned again, resulting in a tensor block with a shape of 1x32 / 32x60x32, meaning the size of the tensor block is 60x32, and the tensor block contains 1920 elements.

[0066] For example, when the first and second partition axes coincide, the data mapping relationship between the source and destination addresses is a partial mapping. In this case, a portion of the data segmented by the largest partition number among the first and second partition numbers can be considered as a tensor block. For instance, if the shape of the data tensor is 1x16x30x256, the first partition axis is 1, the first partition number is 16, the second partition axis is 1, and the second partition number is 8, then a portion of the data segmented according to the first partition parameters can be considered as a tensor block.

[0067] In an exemplary embodiment, after determining the mapping relationship between the source core and the target core, as well as the shape and size of the tensor blocks, each tensor block in the source core can be mapped to the target tensor block address of the target core, thereby obtaining the target core and target tensor block addresses corresponding to each tensor block in the target address and the source address, respectively, and establishing an overall data mapping relationship between the source address and the target address.

[0068] By determining the mapping relationship between each source core and each target core based on the first and second segmentation parameters, and determining the shape and size of the tensor block, and based on the shape and size of the tensor block and the mapping relationship, the target core and target tensor block addresses corresponding to each tensor block in the target address and the source address are determined respectively, so that the target tensor block address to be written to each tensor block can be accurately determined.

[0069] Based on the above technical solution, determining the mapping relationship between each source core and each target core according to the first segmentation parameter and the second segmentation parameter may include: determining the data mapping relationship between the source address and the target address according to the first segmentation axis in the first segmentation parameter and the second segmentation axis in the second segmentation parameter, wherein the data mapping relationship includes partial mapping or full mapping; and determining the mapping relationship between each source core and each target core according to the data mapping relationship, the first segmentation number in the first segmentation parameter and the second segmentation number in the second segmentation parameter.

[0070] In one exemplary embodiment, the consistency of the first segmentation axis in the first segmentation parameter and the second segmentation axis in the second segmentation parameter can be compared to determine whether the data mapping relationship between the source address and the target address is a partial mapping or a full mapping. When the data mapping relationship is a partial mapping, one or more target cores corresponding to each source core can be determined based on the first segmentation number and the second segmentation number, thus obtaining the mapping relationship between each source core and each target core. When the data mapping relationship is a full mapping, the mapping relationship between each source core and each target core can be determined to be that each source core is mapped to each target core respectively.

[0071] Based on the above technical solution, determining the data mapping relationship between the source address and the target address according to the first segmentation axis in the first segmentation parameter and the second segmentation axis in the second segmentation parameter may include:

[0072] If the first slicing axis and the second slicing axis are consistent, the data mapping relationship between the source address and the target address is determined to be a partial mapping; or

[0073] If the first segmentation axis and the second segmentation axis are inconsistent, the data mapping relationship between the source address and the target address is determined to be a full mapping.

[0074] Based on the above technical solution, determining the mapping relationship between each source core and each target core according to the data mapping relationship, the first segmentation number in the first segmentation parameter, and the second segmentation number in the second segmentation parameter may include:

[0075] In the case where the data mapping relationship is a partial mapping, the mapping relationship between each source core and each target core is determined based on the first segmentation number and the second segmentation number; or

[0076] When the data mapping relationship is a full mapping, the mapping relationship between each source core and each target core is determined as follows: each source core is mapped to each target core.

[0077] In an exemplary embodiment, when the data mapping relationship is a partial mapping, if the number of the first segmentation is greater than the number of the second segmentation, the number of the first segmentation can be divided by the number of the second segmentation to obtain the same target core corresponding to multiple source cores; if the number of the first segmentation is less than the number of the second segmentation, the number of the second segmentation can be divided by the number of the first segmentation to obtain multiple target cores corresponding to one source core. That is, the mapping relationship between the source core and the target core is 0<->[0,D-1],1<->[D,2D-1]..., where D represents the number of the first segmentation. In other words, the mapping relationship between the source core and the target core is that the 0th source core is mapped to the 0th target core to the (D-1)th target core, the 1st source core is mapped to the (D)th target core to the (2D-1)th target core, and so on, to obtain the target core mapped to each source core.

[0078] In another exemplary embodiment, when the data mapping relationship is a full mapping, the mapping relationship between each source core and each target core is that each source core is mapped to all target cores respectively. This mapping relationship can be represented as 0<->[0,N-1],1<->[0,N-1],..., where N represents the second number of divisions. That is, the mapping relationship between the source core and the target core is that the 0th source core is mapped to all target cores, the 1st source core is mapped to all target cores, and so on, so that each source core is mapped to all target cores respectively.

[0079] In this embodiment of the application, when the first split axis is greater than the second split axis, the first split axis and the second split axis are not consistent. The data mapping relationship between the source address and the target address is a full mapping. The mapping relationship between each source core and each target core is that each source core is mapped to all target cores respectively.

[0080] Based on the above technical solution, before traversing each tensor block under each segmented data in the source address, the method further includes: determining the single data length as the data length of the tensor block, wherein the single data length is the length of a single data read and a single data write.

[0081] In an exemplary embodiment, the tensor block within the data segmentation at the source address is contiguous, and the tensor block at the target address, after transposition, is also contiguous. In this case, the length of the entire tensor block is used as the length of a single data read and a single data write (data_len_per_iteration), meaning the length of the entire tensor block is used as the length of a single data read. Figure 2 Taking the data tensor shown as an example, the size of the tensor block in the transposed data segment is 60x32, so the length of a single data transmission is 60x32 elements. During data transmission, the entire tensor block can be read to improve bandwidth utilization.

[0082] Figure 4 This is a flowchart of an inter-core data exchange method provided in an embodiment of this application. This embodiment uses... Figure 2 The data tensors shown and Figure 3 The data exchange process shown is illustrated using an example. Figure 4 As shown, the method includes steps 401 to 409.

[0083] Step 401: Determine the mapping relationship between each source core and each target core based on the first segmentation parameter and the second segmentation parameter.

[0084] Each source core is a processor core corresponding to the source address, and each source core stores the segmented data. Each target core is a processor core corresponding to the target address.

[0085] like Figure 2 As shown, the first split axis is 3, the first split number is 8, the second split axis is 1, and the second split number is 32. The first split axis and the second split axis are different, which can determine that the data mapping relationship between the source address and the target address is a full mapping. Thus, the mapping relationship between each source core and each target core is that each source core is mapped to all target cores respectively.

[0086] Step 402: Determine the shape and size of the tensor block based on the first segmentation parameter and the second segmentation parameter.

[0087] like Figure 2 As shown, the shape of the data tensor is 1x32x30x256, the first partition axis is 3, and the first partition number is 8, meaning the shape of the data tensor on a single source core (i.e., a partitioned data) is 1x32x60x32. The second partition axis is 1, and the second partition number is 32. Based on this second partition parameter, the data tensor on a single source core is partitioned again, resulting in a tensor block with a shape of 1x32 / 32x60x32, meaning the size of the tensor block is 60x32, and the tensor block contains 1920 elements.

[0088] Step 403: Based on the shape and size of the tensor block and the mapping relationship, determine the target core and target tensor block addresses corresponding to each tensor block in the source address.

[0089] like Figure 2 and Figure 3As shown, the source cores are Core 0, Core 1, ..., Core 7 (not all are shown in the figure), and the target cores are Core 0, Core 1, Core 2, ..., Core 31 (not all are shown in the figure). The segmented data in each source core includes 32 tensor blocks. Each tensor block is mapped to all target cores. That is, each tensor block in Core 0 is mapped to the position of the 0th tensor block of all target cores (as shown by the arrows in the figure; the mapping between the first two source cores and target cores is used as an example, and the mapping between other source cores and target cores is omitted). Each tensor block in Core 1 is mapped to the position of the 1st tensor block of all target cores (as shown by the arrows in the figure; the mapping between the first two source cores and target cores is used as an example, and the mapping between other source cores and target cores is omitted). And so on, finally, each tensor block in Core 7 is mapped to the position of the 7th tensor block of all target cores (not shown in the figure). Figure 2 and Figure 3 In the example, ta_b (a = 0, 1, ..., 7; b = 0, 1, 2, ..., 31) represents the b-th tensor block in the source core a.

[0090] Step 404: If the first segmentation axis in the first segmentation parameter is greater than the second segmentation axis in the second segmentation parameter, traverse each of the segmented data according to the address compensation value between two adjacent segments of the source address.

[0091] When each segment of data in the source address resides in a different source core, and the data addresses of different source cores are individually addressed, the address compensation value between two adjacent transposed segments of data can be determined to be 0. When the data addresses of different source cores are uniformly addressed, the address compensation value between two adjacent transposed segments of data can be determined to be the address interval between two adjacent source cores. Figure 2 and Figure 3 As can be seen, the tensor blocks in a target core come from different source cores, so the address compensation value (destination stride3) of the inter-core tensor block in the target address is 60x32.

[0092] When the first segmentation axis is greater than the second segmentation axis, each segmented data is traversed and read according to the address compensation value between two adjacent segments in the source address. Based on the address compensation value of the inter-core tensor block in the target address, the traversed and read segmented data is written to the target address. For the currently traversed segmented data, each tensor block is traversed and read according to step 405. The traversal of reading and writing segmented data under different source cores is the first loop.

[0093] Step 405: For the currently segmented data that has been traversed, traverse each tensor block in the currently segmented data according to the address compensation value between two adjacent tensor blocks in the currently segmented data.

[0094] For the currently segmented data that has been traversed, such as Figure 3 As shown, the size of a tensor block is 60x32. From a memory contiguity perspective, the positions of two adjacent tensor blocks are 60x32 elements apart. This means the address offset (source Stride2) between two adjacent tensor blocks is 60x32 elements. At the destination address, the address offset (destination Stride2) between two tensor blocks is equal to 0 because Stride2 describes the offset relationship (offset) between tensor blocks. Figure 2 If tensor blocks in the same segment of data in the source address are distributed to different destination cores, then this compensation relationship is not expressed by Stride2, so it can be determined that destination Stride2 equals 0.

[0095] Based on the address compensation value between two adjacent tensor blocks in the current segmented data, each tensor block in the current segmented data is traversed and read. Then, based on the address compensation value between two tensor blocks at the target address, the traversed tensor blocks are written to the target address. For each traversed tensor block, the data within the tensor block is traversed and read according to step 406. This process is repeated until all tensor blocks in the current segmented data have been read and written. Figure 3 The entire segmented data [1x32x60x32] has been read and written. The traversal of reading and writing each tensor block under the segmented data is a second loop. After the second loop ends, all data in a source core has been read and written.

[0096] Step 406: For the current tensor block traversed in the current segmented data, read the data of the current tensor block from the source address based on the bandwidth capacity, and write the read data into the target tensor block address in the target core corresponding to the current tensor block.

[0097] like Figure 3As shown, after the data transpose operation, the single data length (data_len_per_iteration) is determined to be 60x32 elements. Since the data in the tensor block (source block) is continuous in the split data, Stride1 equals 0 (i.e., Source Stride1 equals 0). Based on data_len_per_iteration and SourceStride1, the tensor block (60x32) can be read and traversed. To determine the data write location, based on the single data length and the address compensation value between adjacent rows in the target core (destination Stride1, which equals 0 because the data within the tensor block is continuous after the transpose operation), one write traversal of the tensor block can be completed. The traversal of reading and writing within the tensor block is the innermost loop.

[0098] Step 407: Perform transpose operations on the tensor blocks written into the addresses of each target tensor block, and perform transpose operations on the segmented data in each target core.

[0099] like Figure 3 As shown, after completing the reading and writing of all data, it can be observed that the data layout does not meet the requirements. It is necessary to perform a transpose operation on each tensor block separately, and then perform an overall transpose operation on the segmented data in each target core to obtain the target's data tensor layout.

[0100] This application utilizes a three-layer loop to abstract the address compensation value (stride) of the three-layer loop corresponding to the source / destination, thereby completing the transformation of data tensor layout within a limited space. This scheme solves the problem of the universality of tensor layout traversal, avoids the huge latency introduced by writing to external storage, greatly reduces the latency of chip model inference, and improves the chip's data throughput and utilization. This application proposes a unified solution that completes the reading and writing operations of tensors divided by arbitrary axes / tensors of arbitrary shapes by establishing a mathematical model, avoiding the case-by-case analysis and processing (i.e., it does not require separate modeling of the transformation of data tensor layout from a fixed number of source cores to a fixed number of target cores), and realizes the transformation of data tensor layout from an arbitrary number of source cores to an arbitrary number of target cores. This application embodiment addresses the scenario where the first split axis is larger than the second split axis. It can directly transmit the entire tensor block. After the transmission is completed, each tensor block is transposed. Then, the split data under each target core is transposed separately, so that the data tensors in the target address meet the arrangement requirements, thereby improving bandwidth utilization.

[0101] Figure 5 This is a schematic diagram of the structure of an inter-core data exchange device provided in an embodiment of this application, as shown below. Figure 5 As shown, the device includes:

[0102] The data mapping module 510 is used to determine the target core and target tensor block addresses corresponding to each tensor block in the source address based on the first segmentation parameter of the data tensor in the source address and the second segmentation parameter of the data tensor in the target address. The data tensor includes multiple tensor blocks, and the target tensor block address is the address of the tensor block mapped to the target address.

[0103] The data exchange module 520 is used to traverse each tensor block under each segmented data in the source address when the first segmentation axis in the first segmentation parameter is greater than the second segmentation axis in the second segmentation parameter, and for the current tensor block that has been traversed, read the data of the current tensor block from the source address based on the bandwidth capacity, and write the read data into the target tensor block address in the target core corresponding to the current tensor block. The segmented data in the source address is the data obtained by segmenting the data tensor according to the first segmentation parameter.

[0104] The transpose module 530 is used to transpose the tensor blocks written to the addresses of each of the target tensor blocks, and to transpose the segmented data in each of the target cores.

[0105] Optionally, the data exchange module includes:

[0106] The first traversal unit is used to traverse each of the segmented data according to the address compensation value between two adjacent segments of the source address;

[0107] The second traversal unit is used to traverse each tensor block in the current segmented data according to the address compensation value between two adjacent tensor blocks in the current segmented data.

[0108] The third traversal unit is used to read the data of the current tensor block from the source address based on the bandwidth capacity for the current tensor block traversed in the current segmented data, and write the read data into the target tensor block address in the target core corresponding to the current tensor block.

[0109] Optionally, the data mapping module includes:

[0110] The core mapping determination unit is used to determine the mapping relationship between each source core and each target core according to the first segmentation parameter and the second segmentation parameter, wherein each source core is a processor core corresponding to the source address, each source core stores the segmentation data, and each target core is a processor core corresponding to the target address.

[0111] Tensor block parameter determination unit, used to determine the shape and size of the tensor block according to the first segmentation parameter and the second segmentation parameter;

[0112] The data mapping unit is used to determine the target core and target tensor block addresses corresponding to each tensor block in the source address, based on the shape and size of the tensor block and the mapping relationship.

[0113] Optionally, the core mapping determination unit includes:

[0114] The data mapping relationship determination subunit is used to determine the data mapping relationship between the source address and the target address based on the first segmentation axis in the first segmentation parameter and the second segmentation axis in the second segmentation parameter. The data mapping relationship includes partial mapping or full mapping.

[0115] The core mapping determination subunit is used to determine the mapping relationship between each source core and each target core based on the data mapping relationship, the first number of segments in the first segmentation parameter, and the second number of segments in the second segmentation parameter.

[0116] Optionally, the data mapping relationship determining subunit is specifically used for:

[0117] If the first slicing axis and the second slicing axis are consistent, the data mapping relationship between the source address and the target address is determined to be a partial mapping; or

[0118] If the first segmentation axis and the second segmentation axis are inconsistent, the data mapping relationship between the source address and the target address is determined to be a full mapping.

[0119] Optionally, the core mapping determining subunit is specifically used for:

[0120] In the case where the data mapping relationship is a partial mapping, the mapping relationship between each source core and each target core is determined based on the first segmentation number and the second segmentation number; or

[0121] When the data mapping relationship is a full mapping, the mapping relationship between each source core and each target core is determined as follows: each source core is mapped to each target core.

[0122] Optionally, the device further includes:

[0123] The data length determination module is used to determine the single data length as the data length of the tensor block, wherein the single data length is the length of a single data read and a single data write.

[0124] The inter-core data exchange apparatus provided in this application embodiment is used to implement the steps of the inter-core data exchange method described in this application embodiment. The specific implementation of each module of the apparatus is described in the corresponding steps, and will not be repeated here.

[0125] The inter-core data exchange apparatus provided in this application determines the target core and target tensor block addresses corresponding to each tensor block in the target address and the source address, respectively, based on the first partitioning parameter of the data tensor in the source address and the second partitioning parameter of the data tensor in the target address. When the first partitioning axis is greater than the second partitioning axis, it traverses each tensor block under each partitioned data in the source address. For the current tensor block being traversed, it reads the data of the current tensor block from the source address based on the bandwidth capacity and writes the read data into the target tensor block address in the target core corresponding to the current tensor block. It then performs a transpose operation on the tensor blocks written to each target tensor block address and performs a transpose operation on the partitioned data in each target core. Since inter-core data exchange is performed... There's no need to write the data tensor to external storage. Data within the tensor block can be read from the source address and written to the target core, avoiding the significant latency introduced by writing to external storage. Furthermore, when the first split axis is greater than the second split axis, by directly transmitting data within the tensor block that meets the bandwidth capacity, and then transposing the tensor block after transmission, followed by a complete transpose operation on the split data under the target core, the tensor block at the target address becomes contiguous before all transpose operations. This allows for direct reading and transmission of multiple rows of data within the tensor block that meet the bandwidth capacity, without needing to transmit each row of data within the tensor block separately, thus improving bandwidth utilization. The transpose operation ensures that the data under the target core conforms to the data arrangement requirements of the target address.

[0126] Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application, such as... Figure 6 As shown, the electronic device 600 may include one or more processors 610 and one or more memories 620 connected to the processors 610. The electronic device 600 may also include an input interface 630 and an output interface 640 for communicating with another device or system. Program code executed by the processor 610 may be stored in the memory 620.

[0127] The processor 610 in the electronic device 600 calls the program code stored in the memory 620 to execute the inter-core data exchange method in the above embodiment.

[0128] This application also provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the steps of the inter-core data exchange method as described in this application.

[0129] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the inter-core data exchange method as described in this application.

[0130] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the apparatus embodiments, since they are substantially similar to the method embodiments, the description is relatively simple; relevant parts can be referred to the descriptions of the method embodiments.

[0131] The foregoing has provided a detailed description of an inter-core data exchange method, apparatus, electronic device, and storage medium provided in the embodiments of this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.

[0132] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

Claims

1. A method for inter-core data exchange, characterized in that, include: Based on the first partitioning parameter of the data tensor in the source address and the second partitioning parameter of the data tensor in the target address, the target core and target tensor block addresses corresponding to each tensor block in the source address are determined respectively. The data tensor includes multiple tensor blocks, and the target tensor block address is the address of the tensor block mapped to the target address. When the first segmentation axis in the first segmentation parameter is greater than the second segmentation axis in the second segmentation parameter, the tensor blocks under each segmented data in the source address are traversed, and for the current tensor block that is traversed, the data of the current tensor block is read from the source address based on the bandwidth capacity, and the read data is written to the target tensor block address in the target core corresponding to the current tensor block. The segmented data in the source address is the data obtained by segmenting the data tensor according to the first segmentation parameter. The tensor blocks written to the addresses of each of the target tensor blocks are transposed, and the segmented data in each of the target cores are also transposed.

2. The method according to claim 1, characterized in that, The process of traversing each tensor block under each segment of data in the source address, and for the current tensor block being traversed, reading the data of the current tensor block from the source address based on the bandwidth capacity, and writing the read data into the target tensor block address in the target core corresponding to the current tensor block, includes: Based on the address compensation value between two adjacent segments of data in the source address, traverse each of the segmented data; For the currently segmented data that has been traversed, each tensor block in the currently segmented data is traversed according to the address compensation value between two adjacent tensor blocks in the currently segmented data; For the current tensor block traversed in the current segmented data, the data of the current tensor block is read from the source address based on the bandwidth capacity, and the read data is written to the target tensor block address in the target core corresponding to the current tensor block.

3. The method according to claim 1 or 2, characterized in that, The step of determining the target core and target tensor block addresses corresponding to each tensor block in the source address based on the first partitioning parameter of the data tensor in the source address and the second partitioning parameter of the data tensor in the target address includes: Based on the first segmentation parameter and the second segmentation parameter, the mapping relationship between each source core and each target core is determined. Each source core is a processor core corresponding to the source address. Each source core stores the segmentation data. Each target core is a processor core corresponding to the target address. The shape and size of the tensor block are determined based on the first segmentation parameter and the second segmentation parameter; Based on the shape and size of the tensor blocks and the mapping relationship, the target core and target tensor block addresses corresponding to each tensor block in the source address are determined.

4. The method according to claim 3, characterized in that, The step of determining the mapping relationship between each source core and each target core based on the first segmentation parameter and the second segmentation parameter includes: Based on the first segmentation axis in the first segmentation parameter and the second segmentation axis in the second segmentation parameter, a data mapping relationship between the source address and the target address is determined, wherein the data mapping relationship includes partial mapping or full mapping; Based on the data mapping relationship, the first number of segments in the first segmentation parameter, and the second number of segments in the second segmentation parameter, the mapping relationship between each source core and each target core is determined.

5. The method according to claim 4, characterized in that, Determining the data mapping relationship between the source address and the target address based on the first segmentation axis in the first segmentation parameters and the second segmentation axis in the second segmentation parameters includes: If the first slicing axis and the second slicing axis are consistent, the data mapping relationship between the source address and the target address is determined to be a partial mapping; or If the first segmentation axis and the second segmentation axis are inconsistent, the data mapping relationship between the source address and the target address is determined to be a full mapping.

6. The method according to claim 4 or 5, characterized in that, The step of determining the mapping relationship between each source core and each target core based on the data mapping relationship, the first segmentation number in the first segmentation parameter, and the second segmentation number in the second segmentation parameter includes: In the case where the data mapping relationship is a partial mapping, the mapping relationship between each source core and each target core is determined based on the first segmentation number and the second segmentation number; or When the data mapping relationship is a full mapping, the mapping relationship between each source core and each target core is determined as follows: each source core is mapped to each target core.

7. The method according to any one of claims 1-6, characterized in that, Before traversing each tensor block under each segment of data in the source address, the method further includes: The single data length is determined to be the data length of the tensor block, where the single data length is the length of a single data read and a single data write.

8. An inter-core data exchange device, characterized in that, include: The data mapping module is used to determine the target core and target tensor block addresses corresponding to each tensor block in the source address based on the first partitioning parameter of the data tensor in the source address and the second partitioning parameter of the data tensor in the target address. The data tensor includes multiple tensor blocks, and the target tensor block address is the address of the tensor block mapped to the target address. The data exchange module is used to traverse each tensor block under each segmented data in the source address when the first segmentation axis in the first segmentation parameter is greater than the second segmentation axis in the second segmentation parameter, and for the current tensor block that has been traversed, read the data of the current tensor block from the source address based on the bandwidth capacity, and write the read data into the target tensor block address in the target core corresponding to the current tensor block. Each segmented data in the source address is data obtained by segmenting the data tensor according to the first segmentation parameter. The transpose module is used to transpose the tensor blocks written to the addresses of each of the target tensor blocks, and to transpose the segmented data in each of the target cores.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the inter-core data exchange method according to any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by the processor, the program implements the steps of the inter-core data exchange method according to any one of claims 1 to 7.