Data processing method, readable storage medium, program product and electronic equipment
By staggering the N data sets in the data group during the data transposition process, the problem of memory bank conflict is solved and the efficiency of data transposition is improved.
Patent Information
- Application Number
- CN202510323348.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-18
- Publication Date
- 2025-06-20
AI Technical Summary
When transposing data in neural network models, electronic devices are prone to memory bank conflicts, resulting in delays in reading intermediate data, which in turn makes the transposition process take a long time.
By staggering at least one memory unit in the N data sets in the data group and storing them one by one in N subspaces one by one, N access requests are avoided to access the same memory bank at the same time, thereby solving the memory bank conflict problem.
It effectively avoids memory bank conflicts, improves the access speed of electronic devices when transposing data, and reduces the time consumption of the transposition process.
Smart Images

Figure CN120179178A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of data processing, and in particular, to a data processing method, a readable storage medium, a program product, and an electronic device. Background Art
[0002] Currently, in neural network models (for example, in various generative models based on the Transformer model), the transpose operation of multi-dimensional data is very common. For example, in some neural network layers (such as convolutional layers) of a neural network model for processing images, it is necessary to transpose image data in the form of height×width into image data in the form of width×height, and then further process the transposed image data.
[0003] When performing data transposition, an electronic device usually stores source data (the source data is stored in the form of height×width) into multiple memory banks of a shared memory. It can be understood that each memory bank can independently serve an access request within a single clock cycle. The multiple memory banks of the shared memory are, for example, divided by columns, that is, there are multiple columns in a row in the shared memory, and one memory bank can store one column of data, for example. During the process of the electronic device storing the source data into the shared memory, the electronic device can read N data in one row of the source data within a single clock cycle, and then the electronic device can access N consecutive memory banks in one row of the shared memory within a single clock cycle and store the N data in one row of the source data into the N memory banks. It can be understood that no memory bank conflict will occur in the above process.
[0004] Then, when the electronic device reads the intermediate data in the shared memory into the transpose cache in the transpose format to obtain transposed data (the transposed data is stored in the form of width×height), due to the change in the data form of the transposed data, the electronic device needs to read N consecutive column data from the shared memory by columns, and then store the N data by rows into the transpose cache. However, one column of data in the shared memory is stored in one memory bank, resulting in that during the process of the electronic device reading N consecutive column data, N access requests will simultaneously access the same memory bank, thereby causing a memory bank conflict, resulting in a delay in the electronic device reading the intermediate data, and thus making the transpose process take a longer time. Summary of the Invention
[0005] Embodiments of this application provide a data processing method, a readable storage medium, a program product, and an electronic device.
[0006] In a first aspect, an embodiment of the present application provides a data processing method, which is applied to an electronic device. The electronic device includes first data and a first storage space. The first data includes at least one data group, each data group includes N data sets, and a data set includes multiple sub-data. N is greater than 1. The first storage space includes at least N sub-spaces, each sub-space is composed of multiple consecutive storage units, and the i-th storage units of each sub-space in the N sub-spaces jointly form the i-th storage bank. Among them, one storage bank supports one access request for reading and writing data, and the storage space of one storage unit can store at most one sub-data. In response to a request to store the first data in the first storage space, the electronic device stores the N data sets in a data group one by one corresponding to the N sub-spaces, where the starting storage positions between any two of the N data sets in a data group are at least separated by one storage unit.
[0007] In some embodiments of the present application, when storing the first data in the first storage space, the electronic device can stagger at least one storage unit between any two of the N data sets in a data group of the first data, and store them one by one corresponding to the N sub-spaces in the first storage space. In this way, when the electronic device accesses the N sub-spaces of the first storage space through N access requests at the same time, the N access requests can access different storage banks, thereby avoiding storage bank conflicts and ensuring the access speed of the electronic device to the first data.
[0008] In a possible implementation of the above first aspect, the bit width of the storage space of one storage unit of the above storage bank is B, and the electronic device can read M data from the first storage space within one clock cycle, where M / B = N, so that the electronic device can access the data in N storage units through N access requests at the same time, where the N storage units are the storage units at the same position in the N data sets of a data group.
[0009] In some embodiments of the present application, since M / B = N, the electronic device can read N sub-data within one clock cycle. Therefore, a data group in the first data can be set to N data sets. In this way, the electronic device can read one sub-data from each of the N data sets in a data group within one clock cycle, thus avoiding storage bank conflicts. And there is no requirement for the storage positions of the N data sets in the second data group (if any) in the first data in the first storage space compared with the storage positions of the N data sets in the first data group in the first space. For example, the initial position of the first data set in the second data group in the first storage space and the initial position of the first data set in the first data group in the first storage space can be in different storage units of the same storage bank.
[0010] In a possible implementation of the first aspect above, the first data is matrix data. Corresponding to a data set being a row of data in the first data, the method further includes: in a first clock cycle, the electronic device reads, from N data sets of a first data group in a first storage space, sub-data in one storage unit respectively as N sub-data in a row of the transposed first data. Alternatively, corresponding to a data set being a column of data in the first data, in the first clock cycle, the electronic device reads, from N data sets of a first data group in a first storage space, sub-data in one storage unit respectively as N sub-data in a column of the transposed first data.
[0011] In some embodiments of the present application, the first electronic data may be matrix data. A row of the first data may be a data set, or a column of the first data may be a data set. When the electronic device transposes the first data, the first data may be stored in the first storage space, so that when the electronic device reads data from the first storage space, there will be no memory bank conflict when accessing data sets in different sub-spaces of the first storage space, thereby improving the speed of the electronic device transposing the first data. For example, a row of data in the first data is stored in a sub-space in the first storage space as a data set. When the electronic device transposes the first data, it is necessary to read the data in the same column of each row in the first data and use the column data as a row of data after the first transposition. And the data of each row in a data group in the first data is stored in different sub-spaces, and there is at least one storage unit interval between each row. That is to say, the data at the same position in each row of the data in a data group in the first data is located in different memory banks. Thus, the electronic device can read one sub-data from N row data of a data group in the first data in one clock cycle and use the N sub-data as N sub-data in a row after the first data is transposed.
[0012] In a possible implementation of the first aspect above, the first clock cycle is the k-th clock cycle when the electronic device reads data from the first storage space. In the first clock cycle, the electronic device reads, from N data sets of a first data group in the first storage space, a sub- including: the electronic device reads one sub-data from the k-th storage unit after the initial storage position corresponding to each data set from each sub-space respectively.
[0013] It can be understood that in the embodiments of the present application, when the electronic device transposes the first data, it reads the sub-data from each subspace in the first storage space in the order of the sub-data in the data set, so as to ensure the continuity of the data after the first data is transposed. That is to say, in the k-th clock cycle, the electronic device reads a sub-data from the k-th storage unit after the initial storage position of a data set.
[0014] In a possible implementation of the above first aspect, the addresses corresponding to multiple sub-data in each data set are misaligned. The electronic device regards the data before the initial addresses of multiple sub-data in each data set as invalid data. And the electronic device moves multiple sub-data in each data set to the lower bits and moves the invalid data to the higher bits. Then, the electronic device clears the invalid data that fills a storage unit in each subspace.
[0015] In some embodiments of the present application, the first data may also include some invalid data. When the electronic device stores the first data in the first storage space, it can also remove the invalid data in the first data. For example, the electronic device can move the invalid data in a data set to the higher bits, move multiple sub-data to the lower bits, and then remove the invalid data in a data set in the first data in units of a storage unit. In this way, it can avoid the invalid data occupying the first storage space and, during the transposition of the first data, can avoid all the transposed data being valid data.
[0016] In a possible implementation of the above first aspect, the storage units at the initial positions of each data set in each subspace include invalid data, and the starting storage positions between any two of the N data sets in a data group are at least two storage units apart.
[0017] In some embodiments of the present application, since the electronic device removes the invalid data in a dataset of the first data in units of a storage unit, that is to say, if the invalid data is less than one storage unit, then this storage unit includes both invalid data and a part of a sub-data. Therefore, the other part of this sub-data will be stored in the next storage unit, that is to say, a sub-data may occupy two storage units. Thus, in the first data, the N datasets of a data group are at least two storage units apart from each other in the first storage space. In this way, when the electronic device transposes the first data, a memory bank conflict will not occur. For example, the sub-data at the initial position of the first dataset is in the first storage unit in the first subspace. Since there is still a part of invalid data in the first storage unit, the end position of this sub-data is in the second storage unit. If the sub-data at the initial position of the second dataset is in the second storage unit in the second subspace, then when the electronic device accesses the first sub-data in the first dataset and the first sub-data in the second dataset simultaneously, it will access the second storage units in different subspaces at the same time. And the second storage units in different subspaces belong to one memory bank, so a memory bank conflict will occur. However, in the embodiments of the present application, in the case where the first data has invalid data, the N datasets of a data group can be stored in their respective corresponding subspaces in the first storage space with two storage units apart from each other. In this way, even if a sub-data occupies two storage units, the electronic device will not have a memory bank conflict when reading the sub-data in different datasets simultaneously.
[0018] In a possible implementation of the foregoing first aspect, the foregoing first data is matrix data, and the corresponding dataset is a row of data in the first data. The method further includes: in the first clock cycle, the electronic device reads one sub-data from each of the N datasets of the first data group in the first storage space, and uses them as the N sub-datas in a row of the transposed first data. Alternatively, corresponding to the dataset being a column of data in the first data, in the first clock cycle, the electronic device reads one sub-data from each of the N datasets of the first data group in the first storage space, and uses them as the N sub-datas in a column of the transposed first data.
[0019] In a possible implementation of the above first aspect, the above first clock cycle is the k-th clock cycle when the electronic device reads data from the first storage space. In the first clock cycle, the electronic device reads a sub in a storage unit from each of the N data sets in the first data set in the first storage space, including: the electronic device reads the high-order part of a sub-data from the k-th storage unit after the initial storage position corresponding to each data set in each sub-space respectively, and reads the low-order part of a sub-data from the (k + 1)-th storage unit after the initial storage position of each data set respectively, and splices the high-order part of a sub-data and the low-order part of a sub-data into a sub-data.
[0020] In a second aspect, the present application provides an electronic device, including: a memory for storing instructions; at least one processor for executing the instructions to enable the device to implement the method provided in the above first aspect and any possible implementation of the above first aspect. The beneficial effects that can be achieved by the second aspect can refer to the beneficial effects of the method provided in any embodiment of the first aspect, which will not be elaborated here.
[0021] In a third aspect, the present application provides a computer-readable storage medium, in which instructions are stored, and when the instructions are executed by a device, the computer is enabled to implement the method provided in the above first aspect and any possible implementation of the above first aspect. The beneficial effects that can be achieved by the third aspect can refer to the beneficial effects of the method provided in any embodiment of the first aspect, which will not be elaborated here.
[0022] In a fourth aspect, the present application provides a computer program product, which, when running on a device, enables the device to implement the method provided in the above first aspect and any possible implementation of the above first aspect. The beneficial effects that can be achieved by the fourth aspect can refer to the beneficial effects of the method provided in any embodiment of the first aspect, which will not be elaborated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] Figure 1 Shows a format of two-dimensional data transposition;
[0024] Figure 2 Shows a schematic diagram of the storage structure of a memory;
[0025] Figure 3 Shows a schematic diagram of an electronic device optimizing an image through a neural network;
[0026] Figure 4 Shows a schematic diagram of the transposition process of two-dimensional data;
[0027] Figure 5 Shows an implementation flowchart of a data processing method according to an embodiment of the present application;
[0028] Figure 6A According to some embodiments of the present application, a schematic diagram of a first data is shown;
[0029] Figure 6B According to some embodiments of the present application, a schematic diagram of a first storage space is shown;
[0030] Figure 6C According to some embodiments of the present application, a schematic diagram of an electronic device clearing invalid data is shown;
[0031] Figure 7 According to some embodiments of the present application, a schematic diagram of a first storage space storing a first data is shown;
[0032] Figure 8 According to some embodiments of the present application, another schematic diagram of a first storage space storing a first data is shown;
[0033] Figure 9 According to some embodiments of the present application, a schematic diagram of the structure of an electronic device 100 is shown. Detailed implementation manners
[0034] The illustrative embodiments of the present application include but are not limited to data processing methods, readable storage media, program products, and electronic devices.
[0035] In order to make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be described in detail below in conjunction with the accompanying drawings of the specification and specific implementation manners.
[0036] For ease of understanding, some terms and related technologies involved in the present application are explained below.
[0037] 1. Data transposition:
[0038] Data transposition refers to interchanging the rows and columns of data stored in the form of a table or matrix, thereby converting the data from one format to another.
[0039] For example, Figure 1 A format of two-dimensional data transposition is shown.
[0040] As Figure 1 shown, after the electronic device transposes the source data source flat, the target data destination flat can be obtained.
[0041] For example, the width (W) of the data in the source data source flat is src_width, and the height (H) is dst_width. In the source data source flat, the addresses of the data in the same row are consecutive, but the addresses of the data between two rows are not necessarily consecutive, that is, the addresses between the data in the same column are not consecutive.
[0042] Similarly, the width of the transposed destination data destination flat is dst_width, and the height is src_width. It can be understood that in the destination data destination flat, the addresses of the data in the same row are consecutive, but the addresses of the data between two rows are not necessarily consecutive, that is, the addresses between the data in the same column are not consecutive.
[0043] Data transposition means writing the data A at the i-th row and j-th column in the source data source fla to the data B at the j-th row and i-th column in the destination data destination flat.
[0044] 2. Memory bank:
[0045] A memory bank is a logical division in memory used to manage and optimize memory access. In the presence of multiple memory banks, different memory banks can perform independent read and write operations simultaneously, thus improving the parallel access ability of memory. This design helps to increase memory bandwidth and access speed.
[0046] For example, Figure 2 shows a schematic diagram of the memory storage structure.
[0047] As Figure 2 shown, the memory M can be divided into 4 memory banks (banks) by column, such as bank0, bank1, bank2, and bank3. That is to say, one row in the memory M includes 4 memory banks. Among them, each memory bank can store 128-bit data, so one row in the memory can store 4 × 128-bit = 512-bit data.
[0048] It can be understood that the memory M can support 4 access requests to access 4 memory banks simultaneously, thus increasing the throughput of the memory M. However, if more than two access requests need to access different data in the same memory bank, a memory bank conflict will occur, resulting in the access requests needing to queue up, increasing the latency of reading data. For example, if there are two access requests to read the data in the first column and the second column of bank0 respectively, these two access requests need to queue up to access bank0, resulting in a decrease in the reading speed of the memory M.
[0049] Next, the process of an electronic device processing an image based on a neural network will be introduced.
[0050] For example, Figure 3 A schematic diagram showing an electronic device optimizing an image through a neural network is shown.
[0051] As Figure 3 shown, the image optimization interface 10 of the electronic device 100 includes an image P0 and an optimization control 11. In some embodiments of the present application, the electronic device 100 can, for example, configure an image optimization model. After the electronic device 100 detects that the user clicks the optimization control 11, it can call the image model to optimize the image P0, thereby generating an optimized image P1.
[0052] In the above process of optimizing the image, the image optimization model can collect the image information in the image P0. This image information can, for example, include two-dimensional data of the image features of each pixel point of the image P0. During the process of the electronic device 100 processing the image information through the image optimization model, the two-dimensional array in the image information may be transposed. If a memory bank conflict occurs during the transposition process, it will cause the image optimization model to take a longer time to optimize the image P0.
[0053] Next, the transposition process of the two-dimensional data will be introduced.
[0054] For example, Figure 4 A schematic diagram showing a transposition process of two-dimensional data is shown.
[0055] As Figure 4 shown, during the process of transposing the source data of the two-dimensional data, the electronic device can first store the source data in the memory M. The memory M can, for example, be divided into 4 banks by column. Exemplarily, the source data can be a 4×4 data matrix, that is, the source data includes 4 rows of data and 4 columns of data. In some embodiments of the present application, the bit width of a data in the source data is the same as the bit width of a bank. Therefore, after the source data is stored in the memory M, each bank in the memory M stores a single column of data in the source data. The electronic device can also process the data in the memory M, such as clearing the invalid information read, etc.
[0056] Then, the electronic device can read the data in the memory M in a transposed format into the transposed cache to obtain the target data. It can be understood that since the rows and columns of the source data and the target data are interchanged, it is necessary to read a column of data from the memory M as a row of data in the target data. However, the data in each column in the memory M is stored in the same bank. Therefore, in the case where the electronic device simultaneously accesses multiple data on a column in the memory M in a parallel manner, a bank conflict will occur, resulting in the inability of the electronic device to access the memory M in parallel to read data, reducing the transpose speed of the source data, and thus affecting the optimization efficiency of the image optimization model.
[0057] As mentioned above, during the process of storing data in the memory, if multiple data that need to be read in parallel are stored in the same bank of the memory, when the electronic device reads data from the memory in parallel, it is easy to have the problem of bank conflict, thus affecting the efficiency of data reading.
[0058] To solve the problem of low data reading efficiency of the electronic device, the present application proposes a data processing method. The electronic device includes first data and a first storage space. The first data includes at least one data group, each data group includes N data sets, a data set includes multiple sub-data, and N is greater than 1. The first storage space includes at least N sub-spaces, each sub-space is composed of multiple consecutive storage units, and the i-th storage units in each of the N sub-spaces together form the i-th bank, where a bank supports one access request for reading and writing data, and the storage space of one storage unit can store at most one sub-data. The method includes:
[0059] In response to a request to store the first data in the first storage space, the electronic device stores the N data sets in a data group one-to-one in the N sub-spaces, where the starting storage positions between any two of the N data sets in a data group are at least separated by one storage unit.
[0060] Through the above solution, the data at the same positions between any two of the N data sets in each data group in the first data are at least separated by one bank. Therefore, within one clock cycle, in a data group, the electronic device can simultaneously read data from N banks. That is to say, there will be no bank conflict during the process of the electronic device reading data from the first storage space, thereby improving the transpose efficiency of the electronic device.
[0061] In some embodiments of the present application, the first data may be, for example, matrix data, and a data set of the first data may be, for example, a row of data or a column of data in the first data. The number of data sets in a data group is related to the bit width of the data that the electronic device can read from the first storage space within one clock cycle and the bit width of each memory bank in the first storage space. For example, if the electronic device can read 512 bits of data from the first storage space within one clock cycle, and a memory bank in the first storage space can store 128 bits of data (i.e., B is 128), then the number of data sets in the first data is 512 / 128 = 4 (i.e., N = 4). It can be understood that in other embodiments, the bit width of the data that the electronic device reads from the first storage space within one clock cycle and the bit width of a memory bank in the first storage space may also be other values. For example, the bit width of the read data may be 128 bits, 256 bits, 1024 bits, etc., and the bit width of each memory bank may be 16 bits, 32 bits, 64 bits, etc. The present application does not limit the bit width of the data that the electronic device reads from the first storage space within one clock cycle and the bit width of a memory bank in the first storage space.
[0062] Exemplarily, in some embodiments, the electronic device may remove the invalid data in the first data from the first storage space. For example, within each clock cycle, after the electronic device reads N×B bits of data from the first data each time, it may remove the invalid data in units of B bits. In some embodiments of the present application, invalid data only appears when the addresses of the valid data in the data sets of the first data are misaligned. When there is invalid data in the first data, the starting storage positions of the N data sets in a data group are at least two storage units apart from each other.
[0063] It can be understood that each data type (such as int, float, double, etc.) in the electronic device usually has a fixed size in the memory (such as 4 bytes, 8 bytes, etc., and one byte is 8 bits of data). If this type of data is stored at an address that is not an integer multiple of its size, an address misalignment occurs. If the address of the valid data in the data set of the first data is misaligned, then the data before the address of the valid data in this data set can be determined as invalid data. Or if the address of the valid data in the data set of the first data is not at the 0th bit of the data set, then the data before the address of the valid data in this data set can be determined as invalid data.
[0064] In the process of removing invalid data, for each data set, the electronic device can move the valid data to the lower positions and the invalid data to the higher positions. After removing the invalid data of a data set, if the invalid data in the higher positions of the data set reaches or exceeds the bit width of a memory bank, that is, reaches or exceeds B bits of data, then the B-bit data can be cleared. That is to say, after clearing the invalid data, there will be no situation where the invalid data in a data set reaches or exceeds B bits. It can be understood that if the invalid data in a data set does not reach B bits, then in the first storage space, in the memory bank storing the data set (for example, the memory bank storing the higher B bits of the data set), there is a situation where there are both invalid data and valid data. Therefore, in this data set, the B-bit valid data may be stored across two memory banks. If the data at the same position in two data sets is separated by only one memory bank, there will still be a situation where two access requests access the same memory bank simultaneously. Therefore, in the embodiments of the present application, the data at the same position in two data sets of a data group needs to be separated by two memory banks.
[0065] It can be understood that in some embodiments of the present application, since the invalid data refers to the valid data generated due to misaligned addresses, the invalid data generally appears in the higher positions of the data set. And, in order to meet the storage requirements of the first data as a matrix, if there is invalid data in the first data set of the first data, then in order to align the bit width of the second data set of the first data, invalid data will also be added to the higher positions of the second data set. That is to say, after removing the invalid data, the bit widths of each data set are the same.
[0066] Next, the data processing method in the embodiments of the present application will be introduced.
[0067] For example, Figure 5 According to the embodiments of the present application, a flowchart of an implementation of a data processing method is shown.
[0068] Exemplarily, the execution subject of each of the following processes can be an electronic device. It should be noted that the specific form of the electronic device is not limited in the present application. The electronic device can be a mobile phone, a laptop computer, a tablet, a desktop computer, a large-screen device, a wearable device (such as a watch, smart glasses, a helmet), an augmented reality (AR) / virtual reality (VR) device, a personal digital assistant (PDA), and other devices.
[0069] As Figure 5 shown, the process includes:
[0070] S501, in response to a request to store the first data in the first storage space.
[0071] Exemplarily, in some embodiments of the present application, the first data may be, for example, a two-dimensional matrix. Taking the transposition of the first data by an electronic device as an example, during the process of transposing the first data, the electronic device may first store the first data in the first storage space. During this process, the electronic device may receive a request to store the first data in the first storage space, and the electronic device may respond to this request.
[0072] In some embodiments of the present application, the first data includes at least one data group, and each data group includes N data sets, where N is greater than 1.
[0073] For example, Figure 6A According to some embodiments of the present application, a schematic diagram of the first data is shown.
[0074] As Figure 6A shown, the first data is, for example, a 4×8 matrix, and each sub-data in the first data may be 128 bits, that is to say, the number of bits of a row of data in the first data is 128×4 bits. A data set in the first data may be a row of data. For example, the first data set in the first data is line0, the second data set is line1, the third data set is line2, the fourth data set is line3, the fifth data set is line4, the sixth data set is line5, the seventh data set is line6, and the eighth data set is line7. Among them, in the first data, every 4 rows of data form a data group (i.e., N is 4), that is to say, the first data set line0, the second data set line1, the third data set line2, and the fourth data set line3 form a data group, that is, data group 1, and the fifth data set line4, the sixth data set line5, the seventh data set line6, and the eighth data set line7 form another data group, that is, data group 2.
[0075] In some embodiments of the present application, the first storage space includes at least N sub-spaces, each sub-space is composed of a plurality of consecutive storage units, and the i-th storage units of each sub-space in the N sub-spaces jointly form the i-th memory bank.
[0076] For example, Figure 6B According to some embodiments of the present application, a schematic diagram of the first storage space is shown.
[0077] As Figure 6BAs shown, the first storage space includes eight sub-spaces, for example, the first sub-space to the eighth sub-space. Each sub-space is composed of multiple consecutive storage units, and the sub-units at the same position in each sub-space together form a bank in the first storage space. That is to say, the banks and sub-spaces in the first storage space are interlaced vertically and horizontally. In some embodiments of the present application, the bit width of data stored in each bank can be, for example, 128 bits (for example, B is 128). For example, in some embodiments of the present application, the first storage space can be divided into 16 banks, then one row of the first storage space can store 128×16 bits of data. The electronic device can read at most 512 bits of data from the first storage space in each clock cycle. That is to say, the electronic device can simultaneously access 4 banks in one clock cycle and read 128 bits of data from each of the 4 banks respectively. In some other embodiments, the 128 bits of data read by the electronic device may be stored in different banks. Therefore, when the electronic device reads a 128-bit data, it can simultaneously access two banks. When the electronic device reads 4 128-bit data in one clock cycle, it can simultaneously access 8 banks.
[0078] S502, store the N data sets in a data group into N sub-spaces one by one, where the starting storage positions between any two of the N data sets in a data group are at least separated by one storage unit.
[0079] In some embodiments of the present application, a data group in the first data may include N data sets. In some embodiments of the present application, N can be 4. That is to say, in the first data, every 4 data sets form a data group, and one data set can be, for example, one row of data in the first data.
[0080] Store the 4 data sets in a data group into 4 sub-spaces in the first storage space one by one, where the starting storage positions between any two of the 4 data sets in a data group are at least separated by one storage unit.
[0081] It can be understood that since the rows of data in data group 1 of the first data are separated by at least one bank in the first storage space, when the electronic device reads data from the first storage space through multiple access requests in one clock cycle, the data in the data sets required to be accessed by each access request are located in different banks, thus avoiding the problem of bank conflict, and improving the speed at which the electronic device obtains the first data. In this way, when the electronic device transposes the first data, the transposition efficiency can also be improved.
[0082] Exemplarily, in some cases, there may also be invalid data in the first data. Then, during the process of the electronic device storing the first data in the first storage space, the invalid data in the first data can also be removed.
[0083] Exemplarily, if the valid data in the first data set line0 in the first data has an unaligned address, the data before the first bit of the valid data in the first data set line0 is all invalid data, and the electronic device can clear the invalid data in the first data set line0.
[0084] For example, Figure 6C According to some embodiments of the present application, a schematic diagram of an electronic device clearing invalid data is shown.
[0085] As Figure 6C shown, when the electronic device obtains the first data set line0 in the first data through the bus network, the invalid data in the first data set line0 can be determined every 128 bits. It can be understood that since the invalid data in a data set is the data before the first valid data when the address of the valid data is unaligned. Therefore, referring to Figure 6C the first data set line0, there is no invalid data in the two 128-bit data at the lower position, and all the 128-bit data at the higher position is invalid data. A part of the second 128-bit data at the higher position is invalid data, and the other part is valid data. When the electronic device clears the invalid data in the first data set line0, the invalid data is cleared in units of the bit width of a memory bank. That is to say, the invalid data will be completely cleared only when it reaches 128 bits. Referring to Figure 6B , after the invalid data in the first data set line0 is cleared, there are only three sub-data left. Among them, there is no invalid data in the two sub-data at the lower position, and there is a part of invalid data in the one sub-data at the higher position.
[0086] Exemplarily, the invalid data in subsequent data sets is cleared in a similar manner, and then the electronic device can store the first data after clearing the invalid data in the first storage space.
[0087] For example, Figure 7 According to some embodiments of the present application, a schematic diagram of the first storage space storing the first data is shown.
[0088] As Figure 7As shown, the first storage space includes 16 memory banks (16 banks, only 10 banks are shown in the figure, bank0 to bank9). Among them, the first data set line0 in the first data occupies three memory banks, that is, from the first memory bank bank0 to the third memory bank bank2. That is to say, the first bit of data in the first data set line0 is located at the initial position of the first memory bank bank0, and the last bit of data in the first data set line0 is located at the end position of the third memory bank bank2. In some embodiments of the present application, the initial positions of different data sets in the same data group stored in the first storage space may be spaced two memory banks apart from each other. Therefore, in data group 1, the data at the same position between different data sets is at least spaced two memory banks apart. Therefore, the first bit of data in the second data set line1 is located at the initial position of the third memory bank bank2, and the second data set line1 also occupies 3 memory banks, so the last bit of data in the second data set line1 is located at the last bit of the fifth memory bank bank4. Similarly, the first bit of data in the third data set line2 is located at the initial position of the fifth memory bank bank4, and the last bit of data in the third data set line2 is located at the end position of the seventh memory bank bank6. The first bit of data in the fourth data set line3 is located at the initial position of the seventh memory bank bank6, and the last bit of data in the fourth data set line3 is located at the end position of the ninth memory bank bank8. In this way, the data at the same position between different data sets is at least spaced two memory banks apart.
[0089] It can be understood that since there may be cases where the addresses in the data set are not aligned, 128 bits of valid data in a data set may span two memory banks. For example, continue to refer to Figure 7, the 128-bit data of the first valid data in the first data set line0 is located in the first memory bank0 and the second memory bank1. Therefore, when the electronic device obtains the first 128-bit valid data of the first data set line0, it needs to access the first memory bank0 and the second memory bank1 simultaneously. Similarly, the 128-bit data of the valid data in the second data set line1 is located in the third memory bank2 and the fourth memory bank3. The electronic device needs to access the third memory bank2 and the fourth memory bank3 simultaneously when obtaining the first 128-bit valid data of the second data set line1. Similarly, when the electronic device obtains the first 128-bit valid data of the third data set line2, it needs to access the fifth memory bank4 and the sixth memory bank5 simultaneously. The electronic device needs to access the seventh memory bank6 and the eighth memory bank7 simultaneously when obtaining the first 128-bit valid data of the fourth data set line3. It can be understood that in the above process, when the electronic device obtains 128-bit data from 4 data sets respectively in one clock cycle, it accesses different memory banks respectively. Therefore, there will be no memory bank conflict when the electronic device obtains 4 128-bit data, that is, 512-bit data, from the first storage space.
[0090] In some other embodiments, the first storage space may also be divided into memory banks in other forms.
[0091] For example, Figure 8 According to some embodiments of the present application, another schematic diagram of storing the first data in the first storage space is shown.
[0092] As Figure 8 shown, the first storage space is divided into 16 memory banks in row form (only 10 are shown in the figure, bank0 to bank9), that is, the first storage space includes a column of 16 rows of memory banks. Then, in the process of storing the first data into the first storage space, one data set of the first data can be a column of data. After storing the first data into the first storage space, the data at the same position between different data sets in the same data group is also spaced two memory banks apart. When the electronic device reads 128-bit data from each data set in the first storage space simultaneously, there will also be no problem of memory bank conflict.
[0093] The electronic device involved in each of the above embodiments will be introduced below.
[0094] For example, Figure 9 According to some embodiments of the present application, a schematic structural diagram of an electronic device 100 is shown.
[0095] The electronic device 100 can be used to implement the view processing method provided in the foregoing embodiments.
[0096] As Figure 9 shown, the electronic device 100 includes one or more processors 101, a system memory 102, a non-volatile memory (NVM) 103, a communication interface 104, an input / output device 105, and a system control logic unit 106 for coupling the processor 101, the system memory 102, the non-volatile memory 103, the communication interface 104, and the input / output (I / O) device 105. Among them:
[0097] The processor 101 may include one or more processing units. For example, it may include a central processing unit (CPU), a graphics processing unit (GPU), a digital signal processor (DSP), a micro-programmed control unit (MCU), an artificial intelligence (AI) processor, or a field programmable gate array (FPGA), a neural-network processing unit (NPU), etc. The processing module or processing circuit may include one or more single-core or multi-core processors. In some embodiments, the CPU may be used to optimize the neural network model to be run. For example, in some embodiments of the present application, the neural network model may transpose the first data, and the NPU may be used to run the neural network model to be run.
[0098] The system memory 102 is a volatile memory, such as a random-access memory (RAM), a double data rate synchronous dynamic random access memory (DDR SDRAM), etc. The system memory is used to temporarily store data and / or instructions. For example, in some embodiments, the system memory 102 may be used to store the data provided by the foregoing different services, such as the first data, etc., and may also be used to store the instructions of the data processing method provided in the foregoing embodiments.
[0099] The non-volatile memory 103 may include one or more tangible, non-transitory computer-readable media for storing data and / or instructions. In some embodiments, the non-volatile memory 103 may include any suitable non-volatile memory such as flash memory and / or any suitable non-volatile storage device, such as a hard disk drive (HDD), a compact disc (CD), a digital versatile disc (DVD), a solid-state drive (SSD), etc. In some embodiments, the non-volatile memory 103 may also be a removable storage medium, such as a secure digital (SD) memory card, etc. In other embodiments, the non-volatile memory 103 may be used to store instructions of the data processing methods provided in the foregoing embodiments, etc.
[0100] Specifically, the system memory 102 and the non-volatile memory 103 may respectively include: a temporary copy and a permanent copy of the instructions 107. The instructions 107 may include: when executed by at least one of the processors 101, enabling the electronic device 100 to implement the data processing methods provided in the embodiments of the present application.
[0101] The communication interface 104 may include a transceiver for providing a wired or wireless communication interface for the electronic device 100, and thus communicating with any other suitable device through one or more networks. In some embodiments, the communication interface 104 may be integrated into other components of the electronic device 100. For example, the communication interface 104 may be integrated into the processor 101. In some embodiments, the electronic device 100 may communicate with other devices through the communication interface 104. For example, the electronic device 100 may obtain corresponding data from other devices through the communication interface 104.
[0102] The input / output (I / O) device 105 may include input devices such as a keyboard, a mouse, etc., and output devices such as a display, etc. The user may interact with the electronic device 100 through the input / output (I / O) device 105.
[0103] The system control logic unit 106 may include any suitable interface controller to provide any suitable interface to other modules of the electronic device 100. For example, in some embodiments, the system control logic unit 106 may include one or more memory controllers to provide an interface connected to the system memory 102 and the non-volatile memory 103.
[0104] In some embodiments, at least one of the processors 101 may be logically encapsulated with one or more controllers for the system control logic unit 106 to form a system in package (SiP). In other embodiments, at least one of the processors 101 may also be integrated with the logic of one or more controllers for the system control logic unit 106 on the same chip to form a system-on-chip (SoC).
[0105] It can be understood that Figure 9 The structure of the illustrated electronic device 100 is merely an example. In other embodiments, the electronic device 100 may include more or fewer components than shown in the figure, or combine certain components, or split certain components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.
[0106] It can be understood that the electronic device 100 may be, including but not limited to, a mobile phone, a car infotainment system, a terminal in self-driving, a wireless terminal in transportation safety, a terminal in a smart city, and so on.
[0107] The embodiments of the present application also provide a program product. When executed on an electronic device, the program product can enable the electronic device to implement the methods provided in the foregoing embodiments.
[0108] The embodiments of the present application also provide a readable storage medium. One or more programs are stored in the readable storage medium. When the one or more programs are executed by the electronic device, the electronic device is enabled to implement the methods provided in the foregoing embodiments.
[0109] The embodiments of the mechanisms disclosed in the present application can be implemented in hardware, software, firmware, or a combination of these implementation methods. The embodiments of the present application can be implemented as a computer program or program code executed on a programmable system. The programmable system includes at least one processor, a storage system (including volatile and non-volatile memories and / or storage elements), at least one input device, and at least one output device.
[0110] The program code can be applied to the input instructions to perform the various functions described in the present application and generate output information. The output information can be applied to one or more output devices in a known manner. For the purposes of the present application, the processing system includes any system having a processor such as, for example, a digital signal processor, a microcontroller, an application-specific integrated circuit, or a microprocessor.
[0111] The program code can be implemented in a high-level procedural language or an object-oriented programming language to communicate with the processing system. When necessary, the program code can also be implemented in assembly language or machine language. In fact, the mechanisms described in this application are not limited to the scope of any specific programming language. In any case, the language can be a compiled language or an interpreted language.
[0112] In some cases, the disclosed embodiments can be implemented in hardware, firmware, software, or any combination thereof. The disclosed embodiments can also be implemented as instructions carried or stored on one or more transitory or non-transitory machine-readable (e.g., computer-readable) storage media, which can be read and executed by one or more processors. For example, the instructions can be distributed via a network or via other computer-readable media. Thus, machine-readable media can include any mechanism for storing or transmitting information in a form readable by a machine (e.g., a computer), including but not limited to, floppy disks, optical disks, optical discs, compact disc-read only memory (CD-ROMs), magneto-optical discs, read only memory (ROM), random-access memory (RAM), erasable programmable read only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), magnetic or optical cards, flash memory, or tangible machine-readable memories for transmitting information (e.g., carrier waves, infrared signals, digital signals, etc.) in electrical, optical, acoustic, or other forms via the Internet. Thus, machine-readable media include any type of machine-readable media suitable for storing or transmitting electronic instructions or information in a form readable by a machine (e.g., a computer).
[0113] In the drawings, some structural or method features may be shown in a particular arrangement and / or order. However, it should be understood that such a particular arrangement and / or ordering may not be required. Rather, in some embodiments, these features may be arranged in a different manner and / or order than shown in the illustrative drawings. Additionally, the inclusion of a structural or method feature in a particular figure does not imply that such a feature is required in all embodiments, and in some embodiments, these features may not be included or may be combined with other features.
[0114] It should be noted that each unit / module mentioned in the device embodiments of the present application is a logical unit / module. Physically, a logical unit / module can be a physical unit / module, a part of a physical unit / module, or can be implemented as a combination of multiple physical units / module. The physical implementation manner of these logical units / modules themselves is not the most important. The combination of the functions implemented by these logical units / modules is the key to solving the technical problems proposed in the present application. In addition, in order to highlight the innovative part of the present application, the above-mentioned device embodiments of the present application do not introduce units / modules that are not closely related to solving the technical problems proposed in the present application. This does not mean that there are no other units / modules in the above-mentioned device embodiments.
[0115] It should be noted that in the examples and the specification of this patent, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "comprising", "including" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising one" does not exclude the presence of additional identical elements in the process, method, article or device comprising the said element.
[0116] Although the present application has been illustrated and described by referring to some preferred embodiments of the present application, those of ordinary skill in the art should understand that various changes can be made in form and detail without departing from the scope of the present application.
Claims
1. A data processing method, applied to an electronic device, characterized in that: The electronic device includes first data and a first storage space; The first data includes at least one data group, each data group includes N data sets, one data set includes multiple sub-data, and N is greater than 1; The first storage space includes at least N subspaces, each subspace is composed of a plurality of continuous storage units, the i-th storage unit of each subspace of the N subspaces together forms the i-th storage body, wherein one storage body supports one access request for reading and writing data, and the storage space of one storage unit can store at most one sub-data; In response to a request to store the first data into the first storage space; The N data sets in one data group are stored in the N subspaces in a one-to-one correspondence, wherein the starting storage positions of the N data sets in one data group are separated by at least one storage unit.
2. The method according to claim 1, characterized in that The bit width of the storage space of a storage unit of the storage body is B, and the electronic device can read M data from the first storage space within one clock cycle, wherein M / B=N, so that the electronic device can simultaneously access the data in the N storage units respectively through N access requests, wherein the N storage units are the storage units at the same position in the N data sets of the one data group.
3. The method according to claim 2, characterized in that The first data is matrix data; The data set corresponds to a row of data in the first data; The method further comprises: In a first clock cycle, sub-data in one storage unit is read from each of the N data sets of the first data group in the first storage space as N sub-data in a row of data after the first data is transposed; Or, the data set corresponds to a column of data in the first data; In a first clock cycle, sub-data in a storage unit are respectively read from the N data sets of the first data group in the first storage space as N sub-data in a column of data after the first data is transposed.
4. The method according to claim 3, characterized in that: The first clock cycle is the kth clock cycle in which the electronic device reads data from the first storage space; The step of reading, in the first clock cycle, a sub-unit in each storage unit from the N data sets of the first data group in the first storage space comprises: One sub-data is read from the k-th storage unit corresponding to the initial storage position of each data set in each sub-space.
5. The method according to claim 2, characterized in that: Corresponding to the misalignment of the addresses of the plurality of sub-data in each data set, treating the data before the initial addresses of the plurality of sub-data in each data set as invalid data; And, the plurality of sub-data in each data set are moved to a low position, and the invalid data are moved to a high position; The invalid data that fills up one storage unit in each subspace is cleared.
6. The method according to claim 5, characterized in that The storage units at the initial positions of the respective data sets in the respective subspaces include invalid data, and the starting storage positions of the N data sets in the one data group are spaced at least two storage units apart from each other.
7. The method according to claim 6, characterized in that The first data is matrix data; The data set corresponds to a row of data in the first data; The method further comprises: In a first clock cycle, one sub-data is read from each of the N data sets of the first data group in the first storage space as the N sub-data in a row of data after the first data is transposed; Or, the data set corresponds to a column of data in the first data; In a first clock cycle, one sub-data is read from each of the N data sets of the first data group in the first storage space as the N sub-data in a column of data after the first data is transposed.
8. The method according to claim 7, characterized in that The first clock cycle is the kth clock cycle in which the electronic device reads data from the first storage space; The step of reading, in the first clock cycle, a sub-unit in each storage unit from the N data sets of the first data group in the first storage space comprises: The high-order part of a sub-data is read from the k-th storage unit corresponding to the initial storage position of each data set in each subspace, and the low-order part of the sub-data is read from the k+1-th storage unit corresponding to the initial storage position of each data set, and the high-order part of the sub-data and the low-order part of the sub-data are spliced into the sub-data.
9. An electronic device, characterized in that: The invention comprises: a memory for storing instructions; At least one processor is used to execute the instructions so that the electronic device implements the method according to any one of claims 1 to 8.
10. A computer-readable storage medium, characterized in that: The readable storage medium stores instructions, and when the instructions are executed on a computer, the computer is caused to execute the method according to any one of claims 1 to 8.
11. A computer program product, characterized in that When the computer program product is executed on a device, the device is caused to perform the method according to any one of claims 1 to 8.