Data processing method and device, storage medium and program product

By continuously storing the same column of data of three-dimensional tensors in the buffer, the problem of waste of bus bandwidth in the three-dimensional tensor transpose operation is solved, and data transmission efficiency and access speed are improved.

CN120122889APending Publication Date: 2025-06-10ARM TECH CHINA CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510280566.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-10
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

In the transpose operation of three-dimensional tensors, electronic devices read data in a row by row, resulting in wasting of bus bandwidth and unable to write multiple data at the same time, reducing data transmission efficiency.

Method used

By obtaining part of the data of the tensor to be processed and storing it in a preset buffer, the data arrangement in the buffer is the same as that in the tensor to be processed, ensuring that the data of the same column is continuously stored in the buffer, and the electronic device can read multiple data at the same time and write to the target tensor.

Benefits of technology

This improves the bandwidth of tensor transpose operations, reduces bus bandwidth waste, and improves data transmission efficiency and access speed.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120122889A_ABST
    Figure CN120122889A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data processing, and discloses a data processing method and device, a storage medium and a program product. The data processing method comprises the following steps: the electronic equipment can prefetch at least part of data of a tensor to be processed into a preset buffer area line by line; wherein the arrangement mode of the pre-fetched data in the to-be-processed tensor is the same as the arrangement mode of the pre-fetched data in the buffer area, and the data in the same column of the to-be-processed tensor are continuously stored in the buffer area. Next, a plurality of data of the same column of the tensor to be processed can be read simultaneously in the buffer area, and the total data volume of the plurality of data can be matched with the bus bandwidth. And finally, writing the plurality of read data into the same row of the target tensor. Therefore, the electronic equipment does not need to frequently switch among different memory areas to access the data, so that the access speed or the data transmission efficiency cannot be reduced. In addition, multiple pieces of data are transposed to the target tensor at the same time, and the situation of bus bandwidth waste can be reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of data processing, and particularly relates to a data processing method, device, storage medium, and program product. Background Art

[0002] Currently, the transpose operation of three-dimensional tensors is widely used in machine learning or neural network models so that an electronic device can process data such as images and audio. Among them, a three-dimensional tensor has three dimensions and can be used to represent data with multiple attributes such as space and time. For example, as Figure 1 shown, the three-dimensional tensor 102 can be used to represent the color image 101 before processing. Among them, the three dimensions of the three-dimensional tensor 102 can respectively represent the height, width, and color channel of the color image 101 (corresponding to the three planes f0, f1, and f2). When the electronic device transposes the three-dimensional tensor 102 to obtain the three-dimensional tensor 103, for example, transposes the data in the i-th (i is a positive integer) row and j-th (j is a positive integer) column in each plane of the three-dimensional tensor 102 to the j-th row and i-th column data, it can represent that the electronic device rotates the color image 101 clockwise by 90 degrees and then horizontally flips it, so as to obtain the processed color image 104.

[0003] Among them, when transposing a three-dimensional tensor, if the electronic device sequentially reads each data in the three-dimensional tensor before transposition row by row, the electronic device can only write one data each time in the target three-dimensional tensor after transposition, resulting in a waste of bus bandwidth. For example, in the above Figure 1 , when the electronic device transposes the three-dimensional tensor 102, it can only sequentially read and write one data each time, and the data volume of these data is usually 8 bits (bit), 16 bit, or 32 bit. When these data are transmitted on a wider bus (such as 512 bit), there will be a situation of waste of bus bandwidth. Summary of the Invention

[0004] In order to solve the above problems, this application provides a data processing method, device, storage medium, and program product.

[0005] In a first aspect, the present application provides a data processing method applied to an electronic device. The method includes: obtaining a to-be-processed tensor corresponding to to-be-processed data, where the to-be-processed tensor may include a row direction dimension and a column direction dimension. Among them, in the row direction dimension, each piece of data arranged in sequence in each row of data in the to-be-processed tensor is continuously stored in the memory space of the electronic device; reading at least part of the data of the to-be-processed tensor along the row direction dimension and storing the at least part of the data in a preset buffer. Among them, the data arrangement manner of the at least part of the data in the buffer is the same as the data arrangement manner of the at least part of the data in the to-be-processed tensor, and among the at least part of the data, each piece of data arranged in sequence in each column of data is continuously stored in the memory space of the buffer; reading a plurality of data in the buffer along the column direction dimension; and using the plurality of data as the data in the row direction dimension of the target tensor.

[0006] It should be understood that the data arrangement manner of the at least part of the data in the buffer being the same as the data arrangement manner of the at least part of the data in the to-be-processed tensor means that: the data in the same row in the to-be-processed tensor is also the data in the same row in the buffer; and the data in the same column in the to-be-processed tensor is also the data in the same column in the buffer.

[0007] In some embodiments, each piece of data is stored in the to-be-processed tensor in a "row-continuous storage" manner and is stored in a "column-continuous storage" manner after being prefetched into the buffer.

[0008] In this way, through the above method, when the electronic device reads the data of the to-be-processed tensor, a row-by-row reading method is adopted. Since the data in the same row of the to-be-processed tensor can be stored in a continuous memory space, the electronic device does not need to frequently switch between different memory areas to access data, and thus will not reduce the access speed or data transmission efficiency. At the same time, since the memory space occupied by the data in the same column of the to-be-processed tensor in the buffer is also continuous, when the electronic device reads a plurality of data in the buffer at the same time, it also does not need to frequently switch between different memory areas to access data, and thus will not reduce the access speed or data transmission efficiency. In addition, the process of the electronic device prefetching the remaining data of the to-be-processed tensor into the buffer and the process of writing the existing data in the buffer into the target tensor can be carried out simultaneously, which is beneficial to the pipelined parallelization of production and consumption and can improve the transpose operation bandwidth of the tensor.

[0009] In a possible implementation of the above first aspect, the absolute value of the difference between the total data volume of the plurality of data read from the buffer along the column direction dimension and the bus bandwidth of the electronic device is less than a preset difference threshold.

[0010] In some embodiments, the electronic device may read multiple data of the same column of the tensor to be processed in the buffer based on the bus bandwidth, and the total data volume of the multiple data may match the bus bandwidth, thereby avoiding bandwidth waste.

[0011] In a possible implementation of the foregoing first aspect, reading at least part of the data of the tensor to be processed along the row direction dimension and storing at least part of the data in a preset buffer includes: along the column direction dimension, dividing the tensor to be processed into k data blocks row by row, where k is a positive integer; wherein, in each of the k data blocks, the absolute value of the difference between the total data volume of the data in each column and the bus bandwidth of the electronic device is less than a preset difference threshold; setting k buffers in the electronic device based on the k data blocks, and each of the k buffers corresponds to one of the k data blocks one by one; reading at least part of the data in each of the k data blocks along the row direction dimension and storing at least part of the data in each data block in the corresponding k buffers.

[0012] For example, if each plane in the tensor to be processed has M rows (M is a positive integer) of data, the electronic device may divide each plane of the tensor to be processed into k data blocks row by row along the column direction dimension, such as data block E1, data block E2, data block E3... data block Ek. And each data block may include M / k rows of data. Among them, the absolute value of the difference between the total data volume of the data in each column of each of the k data blocks of each plane and the bus bandwidth may be less than the preset difference threshold.

[0013] In this way, when storing the data in each data block in the corresponding buffer, the electronic device may directly write one column of data in the buffer into the same row of the target tensor.

[0014] In a possible implementation of the foregoing first aspect, the s-th data block among the k data blocks has q rows of data, and the s-th buffer corresponding to the s-th data block includes q sub-buffers for correspondingly storing q rows of data, 1≤s≤k, and both s and q are positive integers; reading at least part of the data in each of the k data blocks along the row direction dimension and storing at least part of the data in each data block in the preset k buffers includes: reading at least part of the data in the q rows of data along the row direction dimension and correspondingly storing each data in the at least part of the data in the q sub-buffers, wherein when two adjacent data in the same row of the s-th data block are stored in the corresponding sub-buffers, the memory interval between the two adjacent data in the same row of data is x bits, x>0; when two adjacent data in the same column of the s-th data block are stored in the corresponding sub-buffers, the two data in the same column are located in adjacent memory spaces in the s-th buffer.

[0015] In some embodiments, when the electronic device stores the data of each row (e.g., the i-th row) in each read data block, a fixed memory space can be interposed between adjacent data. Then, when storing the data of the next row (e.g., the (i + 1)-th row), the data of this row (e.g., the (i + 1)-th row) can be stored adjacent to the data of the same column in the previous row, so that the data are stored in the buffer area in a "column continuous storage" manner.

[0016] For example, when the electronic device reads each row of the tensor to be processed, the multiple data of the first row read can be (a1, b1, c1), and the multiple data of the second row can be (a2, b2, c2). At this time, when storing the multiple data (a1, b1, c1) of the first row, in the buffer area, the memory address of data a1 and the memory address of data b1 can be separated by a fixed memory space (e.g., x bits), and the memory address of data b1 and the memory address of data c1 can be separated by x bits. Then, when storing the multiple data (a2, b2, c2) of the second row, data a2 and data a1 can be adjacent, data b2 and data b1 can be adjacent, and data c2 and data c1 can also be adjacent. Similarly, when storing the data (a3, b3, c3) of the third row, data a3 and data a2 can be adjacent, data b3 and data b2 can be adjacent, and data c3 and data c3 can also be adjacent. In this way, at least part of the data in the tensor to be processed can be stored in the buffer area in a "column continuous storage" manner.

[0017] In a possible implementation of the first aspect above, reading at least part of the data of q rows along the row direction dimension and storing each data in at least part of the data in q sub-buffers respectively includes: among the q sub-buffers, determining that there are t sub-buffers whose remaining memory space is greater than or equal to a preset memory space threshold, 1 ≤ t ≤ q and t is a positive integer, where the t sub-buffers correspond one-to-one to t rows of data among the q rows of data; along the row direction dimension, reading adjacent n data in each row of the t rows of data in a parallel reading manner or a row-by-row reading manner, where the total data volume between the n data in each row of data is the preset memory space threshold, and n is a positive integer; storing the adjacent n data in each row of the t rows of data in the t sub-buffers through a polling arbitration method.

[0018] In some embodiments, when prefetching each piece of data in each data block into a corresponding buffer, it may include a read operation of reading data from the data block and a write operation of writing the data into a corresponding sub-buffer. Among them, when performing the read operation, the electronic device may adopt a row-by-row reading method (for example, reading the second row after reading the first row), or may adopt a parallel reading method (for example, reading three rows simultaneously). And for each row of data, the electronic device may read all the data of one row at a time, or may read a part of the data of one row at a time (for example, reading multiple data with a data volume of 128K bits at a time). Then, when performing the write operation, the electronic device may write the read data into the corresponding sub-buffer row by row. In this way, multiple rows of data in each data block can be stored in the corresponding sub-buffer.

[0019] In a possible implementation of the foregoing first aspect, storing adjacent n pieces of data in each row of t rows of data in t sub-buffers through a polling arbitration method includes: based on the order of the rows corresponding to the t rows of data, sequentially storing adjacent n pieces of data in each row of the t rows of data in the corresponding t sub-buffers.

[0020] For example, when the order of the rows corresponding to the t rows of data is the i-th row, the (i + 1)-th row, and the (i + 2)-th row respectively, n pieces of data in the i-th row, n pieces of data in the (i + 1)-th row, and n pieces of data in the (i + 2)-th row can be sequentially stored in the corresponding sub-buffers. Or, when the order of the rows corresponding to the t rows of data is the i-th row, the (i + 1)-th row, and the (i + 2)-th row respectively, n pieces of data in the (i + 2)-th row, n pieces of data in the (i + 1)-th row, and n pieces of data in the i-th row can also be sequentially stored in the corresponding sub-buffers. The present application does not limit the specific implementation manner of the polling arbitration.

[0021] In a possible implementation of the foregoing first aspect, the method further includes: after storing adjacent n pieces of data in each row of t rows of data in the corresponding t sub-buffers, recording the (n + 1)-th memory space in the t sub-buffers for storing the (n + 1)-th piece of data; determining that the remaining memory space in the t sub-buffers is greater than or equal to a preset memory space threshold; reading adjacent n pieces of data from the remaining data of the t rows of data along the row direction dimension, and storing the adjacent n pieces of data from the remaining data of the t rows of data starting from the (n + 1)-th memory space in the t sub-buffers.

[0022] For example, for the first sub-buffer corresponding to the first row of data (a1, b1, c1...) in a data block, after the electronic device writes the first two data a1 and b1 in the first row of data into the first sub-buffer, the electronic device can record a position after data b1 in the first sub-buffer (that is, record the position corresponding to data c1). When writing data c1, data d1, etc. into the first sub-buffer next time, it can start writing from the recorded position. In this way, all the data in the i-th row of the s-th data block can be written into the i-th sub-buffer.

[0023] In a second aspect, the present application provides an electronic device, including: a memory and a processor, the memory is coupled to the processor; the memory is used to store computer program code / instructions; when the computer program code / instructions are executed by the processor, the electronic device is caused to execute the data processing method mentioned in the present application.

[0024] In a third aspect, the present application provides a readable storage medium, on which instructions are stored, and when the instructions are executed on an electronic device, the electronic device is caused to execute the data processing method mentioned in the present application.

[0025] In a fourth aspect, the present application provides a computer program product, including: computer instructions, and when the computer instructions run on an electronic device, the electronic device is caused to implement the data processing method mentioned in the present application.

[0026] For the beneficial effects of the second to fourth aspects above, reference can be made to the relevant descriptions in the first aspect and various possible implementations of the first aspect, and details are not described herein again. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] Figure 1 According to some embodiments of the present application, a schematic diagram of an application scenario of three-dimensional tensor transpose is shown;

[0028] Figure 2 According to some embodiments of the present application, a schematic diagram of the structure of three-dimensional tensor transpose is shown;

[0029] Figure 3A According to some embodiments of the present application, a schematic diagram of the structure for dividing a tensor to be processed into data blocks is shown;

[0030] Figure 3B According to some embodiments of the present application, a schematic diagram of the structure for prefetching data of a tensor to be processed into a buffer is shown;

[0031] Figure 4 According to some embodiments of the present application, a schematic diagram of the flow of a data processing method is shown;

[0032] Figure 5According to some embodiments of the present application, a schematic diagram of the hardware structure of an electronic device is shown. Detailed implementation manners

[0033] The illustrative embodiments of the present application include, but are not limited to, a data processing method, device, storage medium, and program product.

[0034] As mentioned above, the transpose operation of tensors is widely used in machine learning or neural network models so that an electronic device can process data such as images and audio. It should be understood that a tensor is a data representation method and can be used to represent various types of data. Among them, the tensors mentioned in the present application can be data representation methods corresponding to any type of data such as image data, audio data, and communication data. In addition, the dimension of the tensor can also be any dimension. For example, the tensor can be a one-dimensional tensor, a two-dimensional tensor, a three-dimensional tensor, a four-dimensional tensor, or any other tensor.

[0035] For example, in the above Figure 1 , the three-dimensional tensor 102 can be used to represent the color image 101 before processing, and the three-dimensional tensor 103 can be used to represent the color image 104 after processing. Another example is that stereo sound includes at least two independent sound channels. Therefore, in the process of audio processing, stereo data can also be represented by a three-dimensional tensor, where the three dimensions of the three-dimensional tensor can respectively represent data such as the time, frequency, and audio channel of the audio. Another example is that in the process of communication data processing, the two dimensions of a two-dimensional tensor can be used to represent the signal frequency and signal strength respectively. Or, in the process of communication data processing, the four dimensions of a four-dimensional tensor can be used to represent time, signal frequency, signal strength, and noise respectively. It can be seen that the present application does not limit the data type corresponding to the tensor.

[0036] It should be understood that a p (p≥2 and p is a positive integer)-dimensional tensor can be composed of multiple p-1-dimensional tensors. For example, a three-dimensional tensor can be composed of multiple two-dimensional tensors. For example, the three-dimensional tensor (a, b, c) can be composed of c two-dimensional tensors (a, b). Another example is that a four-dimensional tensor can be composed of multiple three-dimensional tensors. For example, the four-dimensional tensor (a, b, c, d) can be composed of d three-dimensional tensors (a, b, c). Among them, a, b, c, and d are all positive integers.

[0037] Therefore, when transposing a p-dimensional tensor, it can be gradually split into multiple two-dimensional tensors, and then each of the split two-dimensional tensors is transposed to complete the transposition of the p-dimensional tensor. For example, when transposing a three-dimensional tensor (a, b, c), the c two-dimensional tensors (a, b) split from the three-dimensional tensor can be transposed. Another example is that when transposing a four-dimensional tensor (a, b, c, d), the four-dimensional tensor can first be split into d three-dimensional tensors (a, b, c), and then each three-dimensional tensor (a, b, c) is split into c two-dimensional tensors (a, b), resulting in c*d two-dimensional tensors (a, b). In this way, after transposing the c*d two-dimensional tensors (a, b), the transposition of the four-dimensional tensor (a, b, c, d) can be completed.

[0038] The following takes the transposition of a three-dimensional tensor as an example to describe the implementation manner of this application.

[0039] Refer to Figure 2 As shown, the three-dimensional tensor 201 can be represented by 3 two-dimensional planes f0, f1, f2. When performing tensor transposition on the three-dimensional tensor 201, the data A(i, j) in the i-th row and j-th column of each two-dimensional plane of the three-dimensional tensor 201 can be used as the data A′(j, i) in the j-th row and i-th column of the corresponding two-dimensional plane in the three-dimensional tensor 202.

[0040] It should be understood that when the three-dimensional tensor 201 (or the three-dimensional tensor 202) is stored in the memory of an electronic device such as a computer, the data in the W direction (or W′ direction) dimension of each two-dimensional plane can be stored continuously, that is, the data in the same row can be stored in a continuous memory space. However, the data in the H direction (or H′ direction) dimension of each two-dimensional plane and the data on different two-dimensional planes are not necessarily stored continuously, that is, the data between the same column and the data between different two-dimensional planes can be stored in a discontinuous memory space.

[0041] Among them, when performing a tensor transpose on the three-dimensional tensor 201, data can be read row by row along the W dimension (which can also be referred to as the row dimension). The electronic device can write the corresponding data along the H' dimension (which can also be referred to as the column dimension) of the three-dimensional tensor 202. However, since the data in the H' direction of each two-dimensional plane is not necessarily stored continuously, when the electronic device reads multiple data simultaneously row by row along the W dimension of the three-dimensional tensor 201, the electronic device cannot write the corresponding multiple data simultaneously along the H' dimension of the three-dimensional tensor 202, nor can it determine the memory addresses corresponding to the multiple data along the H' dimension of the three-dimensional tensor 202. That is to say, when the electronic device performs a transpose on the three-dimensional tensor 201, it can only sequentially read and write one data at a time. However, the amount of these data is usually 8bit, 16bit, or 32bit. When these data are transmitted on a wider bus (such as 512bit), there will be a situation of bus bandwidth waste.

[0042] Therefore, in some examples, the electronic device can also read multiple data simultaneously along the H dimension in the three-dimensional tensor 201, and write multiple data simultaneously along the W' dimension in the three-dimensional tensor 202. Among them, the total amount of the multiple data matches the bus bandwidth, thereby reducing the situation of bandwidth waste. For example, referring to Figure 2 , along the H dimension, the electronic device can simultaneously read the first data in the 0th row, the first data in the 1st row... the first data in the ith row, and the total amount of the multiple data read (such as 512bit) is the same as the bus bandwidth (such as 512bit). Next, the electronic device can write the multiple data in the same column into the same row along the W' dimension of the three-dimensional tensor 202.

[0043] However, since the data in the H dimension of each two-dimensional plane is not necessarily stored continuously, when the electronic device reads multiple data simultaneously along the H dimension in the three-dimensional tensor 201, the electronic device may need to frequently switch between different memory areas to access the data, thereby reducing the access speed or data transmission efficiency.

[0044] To solve the above problems, the present application provides a data processing method. In the present application, the electronic device can prefetch at least part of the data of the tensor to be processed row by row into a preset buffer. Among them, when the data is prefetched into the buffer, the data arrangement manner in the tensor to be processed is the same as the data arrangement manner in the buffer, and the memory spaces occupied by the data in the same column of the tensor to be processed in the buffer are continuous. Next, the electronic device can also simultaneously read multiple data in the same column of the tensor to be processed in the buffer, and the total amount of the multiple data can match the bus bandwidth. Finally, the electronic device can write the multiple data read into the same row of the target tensor.

[0045] It should be understood that the data arrangement manner of the above-mentioned data in the tensor to be processed is the same as that in the buffer, which means that the data in the same row in the tensor to be processed is also in the same row in the buffer; and, the data in the same column in the tensor to be processed is also in the same column in the buffer.

[0046] In this way, through the above method, when the electronic device reads the data of the tensor to be processed, it adopts the method of reading row by row. Since the data in the same row of the tensor to be processed can be stored in a continuous memory space, the electronic device does not need to frequently switch between different memory areas to access the data, and thus will not reduce the access speed or data transmission efficiency. At the same time, since the memory space occupied by the data in the same column of the tensor to be processed in the buffer is also continuous, when the electronic device reads multiple data in the buffer at the same time, it also does not need to frequently switch between different memory areas to access the data, and thus will not reduce the access speed or data transmission efficiency. In addition, the process of the electronic device prefetching the remaining data of the tensor to be processed into the buffer and the process of writing the existing data in the buffer into the target tensor can be carried out simultaneously, which is beneficial to the pipelining parallelization of production and consumption and can improve the transpose operation bandwidth of the tensor.

[0047] It can be understood that the electronic device can read multiple data along the column direction dimension of the tensor to be processed in the buffer. Among them, the absolute value of the difference between the total data volume of the multiple read data and the bus bandwidth can be less than a preset difference threshold. For example, the total data volume of the multiple read data and the bus bandwidth can be the same. In this way, when transposing the multiple read data into the data in the row direction dimension of the target tensor, the situation of bandwidth waste can be avoided.

[0048] Based on this, in some embodiments, the electronic device can divide the tensor to be processed into k (k is a positive integer) data blocks row by row along the column direction dimension. And, in each of the k data blocks, the absolute value of the difference between the total data volume of the data in each column and the bus bandwidth can be less than a preset difference threshold. In this way, a whole column of data in a data block can match the bus bandwidth of the electronic device, and the electronic device can directly read a whole column of data in a data block when transposing the columns of data in each data block onto the target tensor.

[0049] Next, the electronic device can set up k buffers in the memory based on the k divided data blocks, so that each buffer in the k buffers corresponds one-to-one to each data block in the k data blocks. Among them, the buffer corresponding to each data block can also be called a collect data queue (CLDQ). In this way, the electronic device can read at least part of the data in each of the k data blocks along the row direction dimension, and store at least part of the data in each data block in the corresponding k buffers in a "column storage manner".

[0050] For example, as Figure 3A shown, if each plane 3011 in the tensor 301 to be processed has M (M is a positive integer) rows of data, the electronic device can divide each plane 3011 of the tensor 301 to be processed into k data blocks row by row along the column direction dimension. For example, data block E1, data block E2, data block E3... data block Ek. And each data block can include M / k rows of data. For example, L1, L2... L64, and each data block has 64 rows. Among them, the absolute value of the difference between the total data volume of each column of data in each data block (such as the data volume sum between the 64 data in the first column data E1' of data block E1) and the bus bandwidth can be less than a preset difference threshold. In this way, after storing at least part of the column data in each data block in the corresponding buffer, the electronic device can directly write one column of data in the buffer into the same row of the target tensor.

[0051] As described above, the electronic device can read multiple data of the corresponding data block along the column direction dimension in the buffer, that is, the data in each data block is stored in the buffer in a "column continuous storage" manner. The process of realizing "column continuous storage" of the data in each data block in the buffer is described below.

[0052] In some embodiments, among the k data blocks divided from each plane of the tensor to be processed, when storing multiple adjacent data in the i-th row (i is a positive integer) of the s-th data block (1 ≤ s ≤ k, and s is a positive integer) into the s-th buffer corresponding to the s-th data block, a fixed memory space (e.g., x bits, x > 0) can be reserved between every two adjacent data in the s-th buffer. When storing multiple adjacent data in the (i + 1)-th row of the s-th data block into the s-th buffer corresponding to the s-th data block, the multiple data in the (i + 1)-th row and the multiple data in the i-th row that are in the same column are stored adjacent to each other in the s-th buffer. For example, if a certain first data among the multiple data in the i-th row and a certain second data among the multiple data in the (i + 1)-th row are in the same column of the s-th data block, then in the s-th buffer, the second data can be stored in the memory space adjacent to the first data.

[0053] Similarly, when storing multiple adjacent data in the (i + 2)-th row of the s-th data block into the s-th buffer corresponding to the s-th data block, the multiple data in the (i + 2)-th row and the multiple data in the i-th row that are in the same column are stored adjacent to each other in the s-th buffer. For example, if a certain first data among the n data in the i-th row, a certain second data among the n data in the (i + 1)-th row, and a certain third data among the n data in the (i + 1)-th row are in the same column of the s-th data block, then in the s-th buffer, the third data can be stored in the memory adjacent to the second data. That is to say, when storing the i-th row of data, the memory space between two adjacent data can be used to store the column data adjacent to each data.

[0054] For example, as Figure 3B shown, when the electronic device reads each row in the s-th data block of the tensor 301 to be processed, the data read from the first row can be (a1, b1, c1), and the data read from the second row can be (a2, b2, c2). At this time, when storing the first row of data (a1, b1, c1), in the buffer 302, the memory addresses of data a1 and data b1 can be separated by a fixed memory space (e.g., x bits), and the memory addresses of data b1 and data c1 can be separated by x bits. Then, when storing the second row of data (a2, b2, c2), data a2 and data a1 can be adjacent, data b2 and data b1 can be adjacent, and data c2 and data c1 can also be adjacent. Similarly, when storing the data (a3, b3, c3) in the third row, it can also be made that data a3 and data a2 are adjacent, data b3 and data b2 are adjacent, and data c3 and data c3 are also adjacent.

[0055] In this way, through the above method, at least part of the data in each data block of the tensor to be processed can be stored in the corresponding buffer in the form of "column - continuous storage".

[0056] It can be understood that the memory interval between adjacent data in the same row (that is, the x bits in Figure 3B ) is used to store the data in the same column. In this embodiment, the memory interval between two adjacent data can be set arbitrarily. For example, the interval memory can be greater than or equal to the bus bandwidth of the electronic device, or can be greater than or equal to the total data volume in one column of the corresponding data block. Specifically, regardless of the capacity of the interval memory, as long as the sum of the data volumes of multiple adjacent data (corresponding to the column data in the same data block) in the buffer matches the bus bandwidth of the electronic device, multiple data can be directly transposed to the same row of the target tensor.

[0057] For example, in the above Figure 3B , if the sum of the data volumes of data a1 and data a2 is the same as the bus bandwidth, the electronic device can directly transpose data a1 and data a2 to the same row of the target tensor (for example, transpose them to the first row). Then, next time, the electronic device can continue to transpose data a3 and data a4 adjacent to data a2 in the buffer to the first row of the target tensor. In this way, when setting the memory interval between adjacent data in the same row, only the bus bandwidth of the electronic device needs to be considered, without considering the total data volume of the column data of different tensors to be processed. That is to say, for any tensor to be processed, in the same electronic device, the same interval memory can be set in the buffer based on the bus bandwidth of the electronic device.

[0058] Based on the above method, when the electronic device reads the data in the same column in the buffer, there is no need to frequently switch between different memory areas to access the data, and thus the access speed or data transmission efficiency will not be reduced. At this time, since the data in each plane of the target tensor is also "row - continuous storage", when the electronic device writes multiple data read from the buffer into the same row of the target tensor, there is still no need to frequently switch between different memory areas.

[0059] It should be understood that since the electronic device stores the data of each data block in the corresponding buffer along the row direction dimension. Therefore, each row of data in each data block can correspond one-to-one with each sub-buffer (e.g., ring buffer) in the buffer corresponding to the data block. For example, if the s-th data block among k data blocks has q rows of data, the electronic device can divide the s-th buffer corresponding to the s-th data block into q sub-buffers, so that each row of data among the q rows of data can correspond one-to-one with each sub-buffer among the q sub-buffers, in order to store the q rows of data in the corresponding q sub-buffers respectively. Wherein, 1 ≤ s ≤ k, and both s and q are positive integers.

[0060] Among them, as described above, if the data in each data block is to be stored column-continuously in the corresponding buffer, then adjacent data of each row of data can have a memory interval of x bits in the buffer. Therefore, in this embodiment, when dividing the s-th data block into q sub-buffers, each sub-buffer can be a discontinuous memory space, and in each sub-buffer, there can be a memory interval of x bit between the memory spaces for storing adjacent data in the corresponding row of data. In this way, when storing two adjacent data in the same row of data in the s-th data block in the corresponding sub-buffer, there can be a memory interval of x bits (x > 0) between the two adjacent data in the same row of data. And, when storing two adjacent data in the same column of data in the s-th data block in the corresponding sub-buffer, the two adjacent data in the same column of data can be located in adjacent memory spaces in the s-th buffer. That is to say, by the above method of dividing each sub-buffer, when storing each row of data in the corresponding sub-buffer, each row of data can be stored in a "column storage mode".

[0061] For example, as described above Figure 3B As shown, when dividing the sub-buffers of buffer 302, if the data volume of each data is x' bit, the electronic device can use the regions at intervals of x bits of memory space such as the x'-th bit, the (x'+x)-th bit, the (x'+x+x)-th bit, etc. as the sub-buffers corresponding to the first row of data. Then, use the regions at intervals of x bits of memory space such as the 2x'-th bit, the (2x'+x)-th bit, the (2x'+x+x)-th bit, etc. as the sub-buffers corresponding to the second row of data.

[0062] It can be understood that when prefetching each piece of data in each data block into the corresponding buffer, it may include a read operation of reading data from the data block and a write operation of writing the data into the corresponding sub-buffer. Among them, when performing the read operation, the electronic device can adopt a row-by-row reading method (for example, reading the second row after reading the first row), or a parallel reading method (for example, reading three rows simultaneously). And for each row of data, the electronic device can read all the data of one row at a time, or read a part of the data of one row at a time (for example, reading multiple data with a data volume of 128K bits at a time). Then, when performing the write operation, the electronic device can write the read data into the corresponding sub-buffer in a row-by-row manner. The following specifically introduces the process of reading the data in each data block and storing it in the corresponding sub-buffer.

[0063] In some embodiments, in the q sub-buffers of the s-th buffer, if it is determined that there are t (1≤t≤q and t is a positive integer) sub-buffers in which the remaining memory space of each sub-buffer is greater than or equal to the preset memory space threshold, and the t sub-buffers correspond one-to-one to t rows of data among the q rows of data. Then the electronic device can read the adjacent n pieces of data in each row of the t rows of data in a parallel reading manner or a row-by-row reading manner along the row direction dimension, where the total data volume between the n (n is a positive integer) pieces of data in each row of data can be the preset memory space threshold.

[0064] For example, the electronic device can initiate a prefetch request to each of the t rows of data, and one prefetch request is used to read n pieces of data in one row. In this way, based on the t prefetch requests, the electronic device can read the adjacent n pieces of data in each row of the t rows of data along the row direction dimension.

[0065] Next, when obtaining the adjacent n pieces of data in each row of the t rows of data, the electronic device can store the adjacent n pieces of data in each row of the t rows of data in the corresponding t sub-buffers through a polling arbitration method. For example, the arbitration rule of the polling arbitration method can be to sequentially store the adjacent n pieces of data in each row of the t rows of data in the corresponding t sub-buffers based on the order of the rows corresponding to the t rows of data. For example, when the orders of the rows corresponding to the t rows of data are the i-th row, the (i + 1)-th row, and the (i + 2)-th row respectively, the n pieces of data in the i-th row, the n pieces of data in the (i + 1)-th row, and the n pieces of data in the (i + 2)-th row can be sequentially stored in the corresponding sub-buffers. Or, when the orders of the rows corresponding to the t rows of data are the i-th row, the (i + 1)-th row, and the (i + 2)-th row respectively, the n pieces of data in the (i + 2)-th row, the n pieces of data in the (i + 1)-th row, and the n pieces of data in the i-th row can also be sequentially stored in the corresponding sub-buffers. The present application does not limit the specific implementation manner of the polling arbitration.

[0066] It can be understood that after the electronic device stores adjacent n data in each row of the t rows of data in the corresponding t sub-buffers respectively, it can also record the (n + 1)-th memory space corresponding to the (n + 1)-th data in each sub-buffer of the t sub-buffers. Next, when the electronic device determines again that the remaining memory space in each sub-buffer of the t sub-buffers is greater than or equal to the preset memory space threshold, the electronic device can also read adjacent n data in the remaining data in each row of the t rows of data along the row direction dimension again. At this time, the electronic device can store adjacent n data in the remaining data in each row of the t rows of data starting from the (n + 1)-th memory space in the corresponding sub-buffer of the t sub-buffers.

[0067] For example, in the above Figure 3B , for the first sub-buffer corresponding to the first row of data (a1, b1, c1...), when the electronic device writes the first two data a1 and b1 in the first row of data into the first sub-buffer, the electronic device can record a position after data b1 in the first sub-buffer (that is, record the position corresponding to data c1). The next time when writing data c1, data d1, etc. into the first sub-buffer, it can start writing from the recorded position. In this way, all the data in the i-th row of the s-th data block can be written into the i-th sub-buffer.

[0068] In this way, through the above method, when the electronic device reads each row of data of the tensor to be processed and writes it into the corresponding sub-buffer, and writes the data in the buffer corresponding to each data block to the target tensor, there is no need to frequently switch between different memory areas to access data, and thus the access speed or data transmission efficiency will not be reduced. In addition, the process of prefetching some of the remaining data in each row of the tensor to be processed into the sub-buffer and the process of writing the existing data in the sub-buffer to the target tensor can be carried out simultaneously, which is beneficial to the pipelining parallelization of production and consumption and can improve the transpose operation bandwidth of the tensor.

[0069] It can be understood that the above data processing method of the present application can be applied to any electronic device. Among them, the electronic device includes but is not limited to a mobile station (MS), a mobile terminal (MT), etc. For example, the electronic device can be a mobile phone, a smart TV, a wearable device, a tablet computer (Pad), a desktop computer, a laptop computer, a virtual reality (VR) device, an augmented reality (AR) device, a terminal in industrial control, a terminal in self-driving, a terminal in remote medical surgery, a terminal in a smart grid, a terminal in transportation safety, a terminal in a smart city, a terminal in a smart home, etc. The specific form of the electronic device in the embodiments of the present application is not limited.

[0070] Based on Figure 4 the following flow schematic diagram, a brief introduction to the data processing method mentioned in the embodiments of the present application is given. Among them, this data processing method can be applied to an electronic device, such as any electronic device such as a computer mentioned above. As Figure 4 shown, specifically, the method is as follows:

[0071] S401: Obtain a to-be-processed tensor corresponding to the to-be-processed data.

[0072] It can be understood that the to-be-processed tensor submitted in the present application can be a tensor corresponding to any data such as image data, audio data, industrial data, etc. Moreover, the to-be-processed tensor mentioned in the present application can be any multi-dimensional tensor, such as a two-dimensional tensor, a three-dimensional tensor, a four-dimensional tensor, etc.

[0073] Among them, the tensor to be processed includes at least a row direction dimension and a column direction dimension. For example, when the tensor to be processed is a two-dimensional tensor, the tensor to be processed is a plane with only a row direction dimension and a column direction dimension. Another example is that when the tensor to be processed is a three-dimensional tensor (a, b, c), the three-dimensional tensor can be regarded as a set composed of multiple two-dimensional tensors. For example, it can be regarded as a set composed of c two-dimensional tensors (a, b). At this time, the tensor to be processed can have a row direction dimension, a column direction dimension, and a depth direction dimension. The row direction dimension and the column direction dimension can form a plane, and the number of data on the same depth direction dimension can represent the number of planes. Another example is that when the tensor to be processed is a four-dimensional tensor (a, b, c, d), the four-dimensional tensor can be regarded as a set composed of multiple three-dimensional tensors. For example, it can be regarded as a set composed of d three-dimensional tensors (a, b, c). At this time, the tensor to be processed can have four dimensions: a row direction dimension, a column direction dimension, a spatial direction dimension 1, and a spatial direction dimension 2. The row direction dimension and the column direction dimension can form a plane. Therefore, the data processing method mentioned in this application can be applied to process any multi-dimensional tensor, and this application does not make any limitations.

[0074] It can be understood that in this embodiment, in the row direction dimension, each data arranged in sequence in each row of data in the tensor to be processed is continuously stored in the memory space of the electronic device. That is to say, in this embodiment, the dimension direction in which data is continuously stored can be used as the row direction dimension.

[0075] S402: Read at least part of the data of the tensor to be processed along the row direction dimension.

[0076] In some embodiments, since in the tensor to be processed, each data arranged in sequence in each row of data is continuously stored in the memory space of the electronic device. Therefore, when the electronic device reads at least part of the data of the tensor to be processed along the row direction dimension, there is no need to frequently switch between different memory areas to access data, and thus the access speed or data transmission efficiency will not be reduced.

[0077] In some embodiments, before reading the data of the tensor to be processed, the electronic device can first divide each plane of the tensor to be processed into k data blocks row by row along the column direction dimension, and then read at least part of the data in each of the k data blocks along the row direction dimension. Among them, each row of data in each data block can correspond one by one to each sub-buffer in the buffer corresponding to the data block. In this way, the electronic device can read multiple rows of data in the buffer along the row direction dimension and store them in the corresponding sub-buffers.

[0078] Specifically, the total capacity of each sub-buffer can be set arbitrarily. For example, when the total capacity of each sub-buffer is less than the total data volume of the corresponding row of data, the electronic device can directly prefetch each data in the row of data into the corresponding sub-buffer. Or, when the total capacity of a certain sub-buffer is less than the total data volume of the corresponding row of data, the electronic device can only store a partial amount of data in the row of data into the sub-buffer at a time. However, after the electronic device writes at least a part of the existing data in the sub-buffer into the target tensor, there will be remaining memory space in the sub-buffer. Once the remaining memory space is greater than the preset capacity threshold, the electronic device can initiate a prefetch request for the row of data again and store the remaining partial data in the sub-buffer. In this way, the process of the electronic device prefetching row data into the sub-buffer and the process of writing the existing data in the sub-buffer into the target tensor can be carried out simultaneously, which is beneficial to the pipelined parallelization of production and consumption and can improve the transpose operation bandwidth of the tensor.

[0079] For example, as described above Figure 3A As shown, the capacity of the sub-buffer corresponding to each row of data can be set to 128 * 8 * 3 bits (i.e., 3K bits). Among them, when the total data volume of each row of data is less than or equal to 3K bits, the electronic device can directly store all the data of each row in the corresponding sub-buffer. Or, when the total data volume of each row of data is greater than 3K bits, for each row of data (for example, taking the first row of data as an example), the electronic device can store at most 3K bits of data in the corresponding sub-buffer at a time. However, when the electronic device writes at least a part (such as 2K bits) of the existing data in the sub-buffer into the target tensor, there will be remaining memory space (such as 2K bits) in the sub-buffer. Once the remaining memory space is greater than the preset memory space threshold (such as greater than 1K bits), the electronic device can initiate a prefetch request for the first row of data again and store at least a part of the remaining data in the sub-buffer.

[0080] It can be understood that when the remaining memory space of a certain sub-buffer is greater than or equal to the preset memory space threshold, a prefetch request can be initiated. Among them, the sum of the data volumes of a group of data prefetched by a prefetch request can be the same as the preset memory space threshold. For example, the sum of the prefetched data volumes can be 1K bits. Or, the sum of the data volumes of a group of data prefetched by a prefetch request can also be the same as the remaining memory space in the corresponding sub-buffer, which is not limited in this application.

[0081] S403: Store at least part of the data in a preset buffer, where, among the at least part of the data, the data arranged in sequence in each column of data are continuously stored in the memory space of the buffer.

[0082] It can be understood that the data arrangement manner of at least part of the pre-fetched data in the buffer is the same as the data arrangement manner of the at least part of the data in the tensor to be processed. That is to say, the data in the same row in the tensor to be processed is also the data in the same row in the buffer; and the data in the same column in the tensor to be processed is also the data in the same column in the buffer.

[0083] In some embodiments, as described in S402 above, in each sub-buffer of the s-th buffer, if the electronic device determines that there is at least one sub-buffer with the remaining memory space of each sub-buffer greater than the preset memory space threshold, at least one pre-fetch request can be issued, and each pre-fetch request is used to read a set of data corresponding to a sub-buffer. At this time, the electronic device can sequentially select each set of data from the read sets of data corresponding to each sub-buffer by the round-robin method and store them in the corresponding sub-buffer.

[0084] For example, referring to the above Figure 3A As shown, if the number of rows of each data block is 64 rows, each data block can correspond to 64 sub-buffers (which can also be called ring buffers). At this time, if the remaining memory space of each sub-buffer in the 64 sub-buffers corresponding to data block E1 is greater than the preset memory space threshold (for example, greater than 1K bits), the electronic device can initiate 64 pre-fetch requests to read at least part of the data in each row of the 64 rows of data in data block E1. Among them, the burst length of one pre-fetch request can be 2 and the data bit width can be 512 bits. That is to say, one pre-fetch request can pre-fetch 1K bits of data in one row. In addition, the memory access request splitting module in the electronic device can receive at most 64 pre-fetch requests.

[0085] Then, for the 64 sets of data obtained from the 64 pre-fetch requests, the electronic device can arbitrate and select one pre-fetch request for splitting and sending through the round-robin method, and a total of 128 memory access requests can be issued. That is to say, the electronic device can sequentially select each set of data by the round-robin method and store them in the corresponding sub-buffer. For example, the first time, select the first set of data and store it in the corresponding first sub-buffer, the second time, select the second set of data and store it in the corresponding second sub-buffer, until the 64th time, select the 64th set of data and store it in the corresponding 64th sub-buffer. In this way, the 64 sets of data obtained from the pre-fetch requests can be stored in the corresponding sub-buffers.

[0086] Exemplarily, in the implementation where the number of rows of each of the above-mentioned data blocks is 64 rows, a storage block (bank) such as a static random-access memory (SRAM) and other hardware can be used for implementation. For example, it can be implemented by the hardware of 8 SRAM banks.

[0087] S404: Read multiple data in the buffer along the column direction dimension.

[0088] In some embodiments, the absolute value of the difference between the total data volume of the multiple read data and the bus bandwidth of the electronic device can be less than a preset difference threshold. Among them, the preset difference threshold can be set arbitrarily. In this way, the situation of bandwidth waste can be reduced.

[0089] It should be understood that since the memory spaces occupied by the data in the same column of the tensor to be processed in the buffer are also continuous, when the electronic device reads multiple data in the buffer at the same time, it does not need to frequently switch between different memory areas to access the data, and thus will not reduce the access speed or data transmission efficiency.

[0090] In addition, in some embodiments, the electronic device can also control the reading and writing of the buffer by maintaining a table (which can be called the T linestatus table). Specifically, the table can record the positions of the read and write pointers (such as the read pointer line_rptr and the write pointer line_wptr) of each sub-buffer in the sub-buffer to be used to track the current readable and writable data positions.

[0091] For example, in the above Figure 3B For the first sub-buffer corresponding to the first row of data (a1, b1, c1...), when the electronic device writes the first two data a1 and b1 in the first row of data into the first sub-buffer, the electronic device can point the write pointer to a position after data b1 in the first sub-buffer (that is, the position corresponding to data c1). The next time data c1, data d1, etc. are written into the first sub-buffer, writing can start from the position pointed to by the write pointer. And if the first sub-buffer can only write three data (such as a1, b1, c1), when the electronic device writes three data into the first sub-buffer, it can point the write pointer to the memory address corresponding to the first data in the first sub-buffer (that is, the position corresponding to data a1), so that when there is remaining memory space in the first sub-buffer next time, writing can start from the memory address corresponding to the first data.

[0092] In addition, in the above Figure 3BIn the case of the first sub-buffer corresponding to the first row of data (a1, b1, c1...), when the electronic device reads the first two data a1 and b1 in the first row of data in the first sub-buffer and writes them into the target tensor, the electronic device can point the read pointer to a position after b1 in the first sub-buffer (that is, to the position corresponding to data c1). When reading data c1 and data d1 in the first sub-buffer and writing them into the target tensor next time, it can start reading from the position pointed to by the read pointer. Moreover, when all the data in the first sub-buffer are read at once and the last data read is located at the last memory address of the first sub-buffer, the read pointer can be pointed to the memory address corresponding to the first data (that is, the position corresponding to data a1) again, so that when storing data in the first sub-buffer next time, it can start reading data from the memory address corresponding to the first data.

[0093] In addition, since the head and tail of the circular buffer may cross the storage boundary, resulting in some bytes being invalid, valid bytes can be marked by flag fields (such as the swi_rbe field) in the table. Additionally, the last data in each row of data can be marked by a flag field (such as the line_last field) in the table; and the last row of data in the data block can be marked by a flag field (such as the last_line field) in the table. Moreover, in the table, a flag field (such as the line_rvld field) can be used to record whether each row of data contains valid data, so as to distinguish which data can be read and processed.

[0094] S405: Use multiple data as data in the row direction dimension of the target tensor.

[0095] It can be understood that when transposing a tensor, the data A(i, j) in the i-th row and j-th column of each two-dimensional plane in the tensor to be processed needs to be transposed to the data A′(j, i) in the j-th row and i-th column of the corresponding two-dimensional plane in the target tensor. Among them, due to the arrangement of the data in the buffer being the same as the arrangement of the data in the tensor to be processed, the electronic device can use the multiple data read along the column direction dimension in S404 above as the data in the row direction dimension on the target tensor. In this way, the operation of tensor transposition can be completed.

[0096] In this way, through the above method, read requests are issued at the granularity of rows, and the data to be transposed is prefetched into the internal buffer, and then the data is read from the buffer for transposition operations. By prefetching to compensate for the delay caused by reading double data rate (DDR) data and the nature of this operation itself, the full pipelining of production and consumption is achieved, and the bandwidth of the transposition operation is increased.

[0097] In some embodiments, the present application provides a readable storage medium, on which instructions are stored. When the instructions are executed on an electronic device, the electronic device implements the data processing method mentioned in the present application.

[0098] In some other embodiments, the present application further provides a computer program product, which includes computer instructions. When the computer instructions run on an electronic device, the electronic device implements the data processing method mentioned in the present application.

[0099] In some other embodiments, the present application further provides an electronic device, which includes a memory and a processor, and the memory is coupled to the processor. The memory is used to store computer program code / instructions, and when the computer program code / instructions are executed by the processor, the electronic device can implement the data processing method mentioned in the present application.

[0100] As Figure 5 shown, a schematic diagram of the hardware structure of the electronic device 1200 according to an embodiment of the present application is exemplarily illustrated. As Figure 5 shown, the electronic device 1200 may include one or more processors 1202, a system control logic 1201 connected to at least one of the processors 1202, a system memory 1205 connected to the system control logic 1201, a memory 1203 connected to the system control logic 1201, and a network interface 1207 connected to the system control logic 1201.

[0101] It can be understood that the structure schematically shown in the embodiments of the present application does not constitute a limitation on the only implementable manner of the electronic device 1200. In some other embodiments of the present application, the electronic device 1200 may include more or fewer components than shown in the figure, or combine certain components, or split certain components, or have different component arrangements. The components shown in the figure can be implemented in hardware, software, or a combination of software and hardware.

[0102] The processor 1202 may include one or more single-core or multi-core processors. In some embodiments, the processor 1202 may include any combination of a general-purpose processor and a dedicated processor (for example, an application processor, a baseband processor, etc.). It can be understood that in the embodiments of the present application, the processor 1202 may be configured to execute the executable instructions 1204 stored in the memory 1203 to implement the data processing method of the embodiments of the present application. When at least one of the processors 1202 executes the instructions, the electronic device 1200 implements the data processing method of the embodiments of the present application.

[0103] The system control logic 1201 may include any suitable interface controllers to provide any suitable interfaces to at least one of the processors 1202 and / or any suitable devices or components communicating with the system control logic 1201. The system control logic 1201 may include one or more memory controllers to provide an interface connected to the system memory 1205. The system memory 1205 may be used to load and store data and / or instructions. In some embodiments, the system memory 1205 of the electronic device 1200 may include any suitable volatile memory, such as a suitable dynamic random access memory.

[0104] The memory 1203 may include one or more tangible, non-transitory computer-readable media for storing data and / or instructions. In some embodiments, the memory 1203 may include any suitable volatile memory and / or any suitable non-volatile storage device. For example, the memory 1203 may include: a random access memory (RAM) and / or a cache storage unit, and may further include a read-only memory (ROM).

[0105] The memory 1203 may include a part of the storage resources installed on the device of the electronic device 1200, or it may be accessible by the device but not necessarily part of the device. For example, the memory 1203 may be accessed via the network interface 1207 through a network.

[0106] In particular, the system memory 1205 and the memory 1203 may respectively include: a temporary copy and a permanent copy of the instructions 1204. The instructions 1204 may include: data processing methods that cause the electronic device 1200 to implement the embodiments of the present application when executed by at least one of the processors 1202. In some embodiments, the instructions 1204, hardware, firmware, and / or its software components may additionally / alternatively be placed in the system control logic 1201, the network interface 1207, and / or the processor 1202.

[0107] The network interface 1207 may include a transceiver for providing a radio interface for the electronic device 1200, and then communicating with any other suitable devices (such as a front-end module, an antenna, etc.) through one or more networks. In some embodiments, the network interface 1207 may be integrated into other components of the electronic device 1200. For example, the network interface 1207 may be integrated into at least one of the processor 1202, the system memory 1205, the memory 1203, and a firmware device (not shown) having instructions.

[0108] The network interface 1207 may further include any suitable hardware and / or firmware to provide a multiple-input multiple-output radio interface. For example, the network interface 1207 may be a network adapter, a wireless network adapter, a telephone modem, and / or a wireless modem.

[0109] The electronic device 1200 may further include: an input / output (I / O) device 1206. The I / O device 1206 may include a user interface that enables a user to interact with the electronic device 1200; the design of the peripheral component interface enables peripheral components to also interact with the electronic device 1200. In some embodiments, the electronic device 1200 further includes sensors for determining at least one of environmental conditions and location information related to the electronic device 1200.

[0110] In some embodiments, the user interface may include, but is not limited to, a display (e.g., a liquid crystal display, a touch screen display, etc.), a speaker, a microphone, one or more cameras (e.g., a still image camera and / or a video camera), a flashlight (e.g., a light-emitting diode flash), and a keyboard.

[0111] In some embodiments, the peripheral component interface may include, but is not limited to, a non-volatile memory port, an audio jack, and a power interface.

[0112] In some embodiments, the sensors may include, but are not limited to, a gyroscope sensor, an accelerometer, a proximity sensor, an ambient light sensor, and a positioning unit. The positioning unit may also be part of the network interface 1207 or interact with the network interface 1207 to communicate with components of a positioning network (e.g., global positioning system (GPS) satellites).

[0113] The embodiments disclosed in the present application may be implemented in hardware, software, firmware, or a combination of these implementation methods. The embodiments of the present application may be implemented as a computer program or program code executed on a programmable system, the programmable system including at least one processor, a storage system (including volatile and non-volatile memories and / or storage elements), at least one input device, and at least one output device.

[0114] The program code may be applied to input instructions to perform the various functions described in the present application and generate output information. The output information may be applied to one or more output devices in a known manner. For the purposes of the present application, a processing system includes any system having a processor such as, for example, a digital signal processor, a microcontroller, an application-specific integrated circuit, or a microprocessor.

[0115] The program code can be implemented in a high-level procedural language or an object-oriented programming language to communicate with the processing system. When necessary, the program code can also be implemented in assembly language or machine language. In fact, the mechanisms described in this application are not limited to the scope of any specific programming language. In any case, the language can be a compiled language or an interpreted language.

[0116] In some cases, the disclosed embodiments can be implemented in hardware, firmware, software, or any combination thereof. The disclosed embodiments can also be implemented as instructions carried or stored on one or more transient or non-transitory machine-readable (e.g., computer-readable) storage media, which can be read and executed by one or more processors. For example, the instructions can be distributed via a network or via other computer-readable media. Thus, machine-readable media can include any mechanism for storing or transmitting information in a form readable by a machine (e.g., a computer), including but not limited to, floppy disks, optical disks, optical discs, magneto-optical discs, read-only memory (ROM), random access memory (RAM), magnetic or optical cards, or tangible machine-readable memories for transmitting information (e.g., carrier waves, infrared signals, digital signals, etc.) using electrical, optical, acoustic, or other forms of propagated signals via the Internet. Thus, machine-readable media include any type of machine-readable media suitable for storing or transmitting electronic instructions or information in a form readable by a machine (e.g., a computer).

[0117] In the drawings, some structural or method features may be shown in a particular arrangement and / or order. However, it should be understood that such a particular arrangement and / or ordering may not be required. Rather, in some embodiments, these features may be arranged in a different manner and / or order than shown in the illustrative drawings. Additionally, the inclusion of a structural or method feature in a particular figure does not imply that such a feature is required in all embodiments, and in some embodiments, these features may not be included or may be combined with other features.

[0118] It should be noted that each unit / module mentioned in the device embodiments of this application is a logical unit / module. Physically, a logical unit / module can be a physical unit / module, a part of a physical unit / module, or can be implemented as a combination of multiple physical units / module. The physical implementation manner of these logical units / module themselves is not the most important. The combination of the functions implemented by these logical units / module is the key to solving the technical problems proposed in this application. In addition, in order to highlight the innovative part of this application, the above device embodiments of this application do not introduce units / modules that are not closely related to solving the technical problems proposed in this application. This does not mean that there are no other units / modules in the above device embodiments.

[0119] It should be noted that in the examples and the description of this application, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising one" does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.

[0120] Although this application has been illustrated and described by reference to certain preferred embodiments thereof, those of ordinary skill in the art should understand that various changes in form and detail may be made therein without departing from the scope of this application.

Claims

1. A data processing method, applied to an electronic device, characterized in that: include: A tensor to be processed corresponding to the data to be processed is obtained, wherein the tensor to be processed includes a row dimension and a column dimension, wherein in the row dimension, each data sequentially arranged in each row of data in the tensor to be processed is stored continuously in the memory space of the electronic device; Reading at least part of the data of the tensor to be processed along the row dimension, and storing the at least part of the data in a preset buffer; The data arrangement of at least part of the data in the buffer is the data arrangement of at least part of the data in the tensor to be processed; Furthermore, in the at least part of the data, the data sequentially arranged in each column of data are stored continuously in the memory space of the buffer; Reading a plurality of data in the buffer along the column direction dimension; The multiple data are used as data in the row direction dimension of the target tensor.

2. The method according to claim 1, characterized in that An absolute value of a difference between a total data volume of the plurality of data and a bus bandwidth of the electronic device is smaller than a preset difference threshold.

3. The method according to claim 1 or 2, characterized in that: The step of reading at least part of the data of the tensor to be processed along the row dimension and storing the at least part of the data in a preset buffer includes: Divide the tensor to be processed into k data blocks by row along the column dimension, where k is a positive integer; Wherein, in each of the k data blocks, an absolute value of a difference between a total amount of data between each data in each column of data and a bus bandwidth of the electronic device is less than a preset difference threshold; Based on the k data blocks, k buffers are set in the electronic device, and each buffer in the k buffers corresponds to each data block in the k data blocks one by one; At least part of the data in each of the k data blocks is read along the row dimension, and at least part of the data in each of the k data blocks is stored in the corresponding k buffers.

4. The method according to claim 3, characterized in that The sth data block among the k data blocks has q rows of data, and the sth buffer corresponding to the sth data block includes q sub-buffers for correspondingly storing the q rows of data, 1≤s≤k, and s and q are both positive integers; The step of reading at least part of the data in each of the k data blocks along the row direction dimension, and storing at least part of the data in each of the k data blocks in the preset k buffers, comprises: Read at least part of the q rows of data along the row direction dimension, and store each data in the at least part of the data in the q sub-buffers accordingly, When two adjacent data in the same row of data in the s-th data block are stored in the corresponding sub-buffer, the memory interval between the two adjacent data in the same row of data is x bits, where x>0; When two adjacent data in the same column of data in the sth data block are stored in the corresponding sub-buffer, the two data in the same column of data are located in adjacent memory spaces in the sth buffer.

5. The method according to claim 4, characterized in that The step of reading at least part of the q rows of data along the row direction dimension, and storing each data of the at least part of the data in the q sub-buffers accordingly, comprises: Among the q sub-buffers, it is determined that there are t sub-buffers whose remaining memory space is greater than or equal to a preset memory space threshold, 1≤t≤q and t is a positive integer, wherein the t sub-buffers correspond one-to-one to t rows of data among the q rows of data; Along the row direction dimension, read n adjacent data in each row of data in the t rows of data in a parallel reading manner or a row-by-row reading manner, wherein the total amount of data between the n data in each row of data is the preset memory space threshold, and n is a positive integer; The adjacent n data in each row of data in the t rows of data are stored in the t sub-buffers by a polling arbitration method.

6. The method according to claim 5, characterized in that The storing n adjacent data in each row of data in the t rows of data in the t sub-buffers by a polling arbitration method includes: The adjacent n data in each row of the t rows of data are sequentially stored in the corresponding t sub-buffers based on the order of rows corresponding to the t rows of data.

7. The method according to claim 6, characterized in that Also includes: After storing n adjacent data in each row of data in the t rows of data in the corresponding t sub-buffers, recording the n+1th memory space in the t sub-buffers corresponding to the n+1th data in each row of data in the t rows of data; Determine that the remaining memory space of the t sub-buffers is greater than or equal to a preset memory space threshold; Along the row direction dimension, n adjacent data in the remaining data of the t rows of data are read, and the n adjacent data in the remaining data of the t rows of data are stored starting from the (n+1)th memory space in the t sub-buffers.

8. An electronic device, characterized in that: include: A memory and a processor, wherein the memory is coupled to the processor; the memory is used to store computer program code / instructions; when the computer program code / instructions are executed by the processor, the electronic device executes the data processing method according to any one of claims 1 to 7.

9. A readable storage medium, characterized in that: The readable storage medium stores instructions, and when the instructions are executed on an electronic device, the electronic device executes the data processing method according to any one of claims 1 to 7.

10. A computer program product, characterized in that include: Computer instructions, when the computer instructions are executed on an electronic device, enable the electronic device to implement the data processing method according to any one of claims 1 to 7.