Data processing method and device
By calculating in the processor and continuously storing input data blocks in shared memory using the target logical address, the problem of low coordinate mapping and loading efficiency during bilinear upsampling is solved, and more efficient data processing is achieved.
Patent Information
- Application Number
- CN202410418712.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-04-08
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2044-04-08
AI Technical Summary
During the bilinear upsampling process, the coordinate mapping operations output to the input are complex and the misaligned input data addresses are slow to load, resulting in low kernel operation efficiency and affecting the performance of computer vision models.
In a processor that supports tensor layout, the logical address of the target input data in the input tensor is determined by calculating the positioning data logical coordinates of each data block in the output tensor, and continuously stores the input data blocks in shared memory, reducing address conversion and loading time.
By reducing the load time and address conversion time of input data, the efficiency of the bilinear upsampling process is improved, and is suitable for various memory architectures that support tensor layout.
Smart Images

Figure CN118331542B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technology, and in particular to a data processing method and device. Background Art
[0002] Upsampling is widely used in the process of tensor operation processing. It refers to sampling and interpolating the elements in the input tensor in two or more dimensions to obtain an output tensor enlarged in two or more dimensions, such as bilinear upsampling. In this process, there are two key steps. The first is the coordinate mapping step: according to the logical coordinates of the data to be generated on the output tensor and the sampling method, the number of source data on the input tensor to be selected and the mapping coordinates of each source data point are different. The second is the sampling calculation step. According to the different sampling methods, the processing and calculation process for the selected source data set is also different. In the process of bilinear upsampling, one output data needs to be generated by four adjacent input data. In the related art, the calculation process of the coordinate mapping operation from output to input is relatively complicated, and loading according to the address of the obtained unaligned input data is usually slow, which will make the generated assembly code kernel run inefficiently, and further cause the performance of some computer vision models that use a large number of bilinear upsampling operators to be far below expectations. How to improve the efficiency and speed of the entire bilinear upsampling process is a technical problem that needs to be solved urgently. Summary of the invention
[0003] In view of this, the present disclosure proposes a data processing method and device.
[0004] According to one aspect of the present disclosure, a data processing method is provided, which is applied to a processor of a memory architecture supporting tensor layout, and the method includes:
[0005] For each first data block in the output tensor, according to the logical coordinates of the positioning data in the first data block, calculate a target logical address of the target input data corresponding to the positioning data in the input tensor, where the positioning data is one of the plurality of output data included in the first data block;
[0006] For each input data loaded, the input data acquired according to the target logical address is stored in the shared memory in the order indicated by the corresponding logical coordinates, so as to store a second data block including the plurality of input data in a continuous space of the shared memory, wherein the size of the second data block and the size of the first data block are both the target size;
[0007] For the calculation of each output data, four input data corresponding to the output data are obtained from the second data block according to the corresponding offset information and the target size, and the four input data are used as interpolation reference points combined with pre-calculated and stored interpolation weights to perform bilinear interpolation calculation to obtain the corresponding output data.
[0008] In a possible implementation, the method further includes:
[0009] Output and store a first data block formed according to the respective output data in the order of corresponding logical coordinates; or
[0010] For the calculation of each output data, after the output data is obtained, it is output and stored in order according to the corresponding logical coordinates.
[0011] In a possible implementation, multiple threads in the execution unit in the processor are used to load the input data, the number of output data in the first data block is the same as the number of threads in the execution unit, and each input data in the second data block comes from the input tensor;
[0012] The input data acquired according to the target logical address is stored in the shared memory in the order indicated by the corresponding logical coordinates, including:
[0013] Controlling a first thread among the multiple threads to obtain input data from a target physical address corresponding to the target logical address, and storing the input data to a shared memory in an order indicated by corresponding logical coordinates; or
[0014] Controlling each thread among the multiple threads different from the first thread, determining a logical address of input data to be obtained according to a thread offset of the thread and the target logical address, and then obtaining the input data from a physical address determined according to the logical address, and storing the input data to a shared memory in an order indicated by corresponding logical coordinates;
[0015] The offset information corresponding to the input data to be loaded by each thread is the thread offset of the thread.
[0016] In a possible implementation, multiple threads in the execution unit in the processor are used to load input data, the number of output data in the first data block is the same as the number of threads in the execution unit, and among the multiple input data in the second data block, only input data that needs to be used as interpolation reference points comes from the input tensor;
[0017] The input data acquired according to the target logical address is stored in the shared memory in the order indicated by the corresponding logical coordinates, including:
[0018] Controlling each of the threads to obtain the input data of the interpolation reference point from the target physical address corresponding to the target logical address when it is determined according to the thread offset of the thread that the input data of the interpolation reference point needs to be obtained, or determining the logical address of the input data to be obtained according to the thread offset of the thread and the target logical address, and then obtaining the input data from the physical address determined according to the logical address, and storing the input data to the shared memory in the order indicated by the corresponding logical coordinates;
[0019] Control each of the threads to obtain input data from a target physical address corresponding to the target logical address when it is determined according to the thread offset of the thread that there is no need to obtain input data of an interpolation reference point, and store the input data in a shared memory according to a storage location corresponding to the thread offset;
[0020] The offset information corresponding to the input data to be loaded by each thread is the thread offset of the thread.
[0021] In a possible implementation, multiple threads in the execution unit in the processor are used to load input data, the number of output data in the first data block is the same as the number of threads in the execution unit, and among the multiple input data in the second data block, only input data that needs to be used as interpolation reference points comes from the input tensor;
[0022] The input data acquired according to the target logical address is stored in the shared memory in the order indicated by the corresponding logical coordinates, including:
[0023] Controlling each of the threads to obtain the input data of the interpolation reference point from the physical address corresponding to the target logical address when it is determined according to the thread offset of the thread that the input data of the interpolation reference point needs to be obtained, or to determine the logical address of the input data to be obtained according to the thread offset of the thread and the target logical address, and then obtain the input data from the physical address determined according to the logical address, and store the input data to the shared memory in the order indicated by the corresponding logical coordinates;
[0024] Control each of the threads to determine a preset value obtained from a preset physical address as input data, or directly determine the preset value as input data, when it is determined according to the thread offset of the thread that there is no need to obtain input data of an interpolation reference point, and store the input data in a shared memory according to a storage location corresponding to the thread offset;
[0025] The offset information corresponding to the input data to be loaded by each thread is the thread offset of the thread.
[0026] In a possible implementation, multiple threads in an execution unit in the processor are used to calculate output data.
[0027] Wherein, according to the corresponding offset information and the target size, four input data corresponding to the output data are obtained from the second data block, including:
[0028] Control each of the threads to obtain input data at an upper left coordinate position among four input data corresponding to the output data from the second data block according to the thread offset of the thread;
[0029] Control each of the threads to obtain the remaining three input data of the four input data corresponding to the output data from the second data block according to the thread offset of the thread and the target size;
[0030] The offset information corresponding to the output data to be calculated by each thread is the thread offset of the thread.
[0031] In a possible implementation, when the input tensor includes multiple layers, the output tensor where the first data block is located is the 0th layer input tensor, and the number of first data blocks in the input tensors of each layer is the target number, and the method further includes:
[0032] In the case where it is determined that the input tensor includes a plurality of layers, a target physical address corresponding to the target logical address is also determined;
[0033] For the loading of each input data, when it is determined that the input data to be obtained does not belong to the input tensor of the 0th layer, the input data obtained from the physical address determined according to the target physical address, the first offset corresponding to the number of layers of the input tensor where the input data to be obtained is located, and the offset information corresponding to the input data, shall be stored in the shared memory in the order indicated by the corresponding logical coordinates; or, the input data obtained from the physical address determined according to the target logical address, the first offset corresponding to the number of layers of the input tensor where the input data to be obtained is located, and the offset information corresponding to the input data shall be stored in the shared memory in the order indicated by the corresponding logical coordinates.
[0034] In a possible implementation, the positioning data is output data located at the upper left position on the boundary among multiple output data of the first data block, and / or
[0035] If there is a corresponding relationship between the difference in storage positions of the input data in the shared memory and the difference in indexes between the threads, the thread offset is the index of the thread.
[0036] In a possible implementation, the method further includes:
[0037] The determined interpolation weights are stored in a preset register or in the shared memory.
[0038] According to another aspect of the present disclosure, there is provided a data processing device, which is applied to a processor of a memory architecture supporting tensor layout, the device comprising a control unit and at least one execution unit, each of the execution units comprising a plurality of threads;
[0039] The control unit calculates, for each first data block in the output tensor, a target logical address of target input data corresponding to the positioning data in the input tensor according to the logical coordinates of the positioning data in the first data block, wherein the positioning data is one of the plurality of output data included in the first data block;
[0040] Each of the threads, for input data that needs to be loaded by the thread, stores the input data that needs to be acquired by the thread according to the target logical address in the shared memory in the order indicated by the corresponding logical coordinates, so as to store a second data block including multiple input data in a continuous space of the shared memory, wherein the size of the second data block and the size of the first data block are both the target size;
[0041] Each of the threads, for the output data to be calculated by the thread, obtains four input data corresponding to the output data from the second data block according to the corresponding offset information and the target size, and uses the four input data as interpolation reference points in combination with pre-calculated and stored interpolation weights to perform bilinear interpolation calculation to obtain the corresponding output data.
[0042] In a possible implementation, the method further includes:
[0043] The control unit outputs and stores a first data block formed according to the output data in the order of corresponding logical coordinates; or
[0044] After obtaining the output data to be calculated, each thread outputs and stores it in order according to the corresponding logical coordinates.
[0045] In a possible implementation, the number of output data in the first data block is the same as the number of threads in the execution unit, and each input data in the second data block comes from the input tensor;
[0046] Each of the threads, for input data that the thread needs to load, stores the input data that the thread needs to obtain according to the target logical address into the shared memory in the order indicated by the corresponding logical coordinates, including:
[0047] A first thread among the multiple threads obtains input data from a target physical address corresponding to the target logical address, and stores the input data to a shared memory in an order indicated by corresponding logical coordinates;
[0048] Each thread among the multiple threads different from the first thread determines a logical address of the input data to be obtained according to the thread offset of the thread and the target logical address, and then obtains the input data from a physical address determined according to the logical address, and stores the input data to the shared memory in an order indicated by the corresponding logical coordinates;
[0049] The offset information corresponding to the input data to be loaded by each thread is the thread offset of the thread.
[0050] In a possible implementation, the number of output data in the first data block is the same as the number of threads in the execution unit, and among the multiple input data in the second data block, only the input data that needs to be used as an interpolation reference point comes from the input tensor;
[0051] Each of the threads, for input data that the thread needs to load, stores the input data that the thread needs to obtain according to the target logical address into the shared memory in the order indicated by the corresponding logical coordinates, including:
[0052] Each of the threads, when determining that input data of an interpolation reference point needs to be obtained according to the thread offset of the thread, obtains the input data from the target physical address corresponding to the target logical address, or determines the logical address of the input data to be obtained according to the thread offset of the thread and the target logical address, and then obtains the input data from the physical address determined according to the logical address, and stores the input data to the shared memory in the order indicated by the corresponding logical coordinates;
[0053] When it is determined according to the thread offset of the thread that there is no need to obtain input data of the interpolation reference point, the input data is obtained from the target physical address corresponding to the target logical address, and the input data is stored in the shared memory according to the storage location corresponding to the thread offset;
[0054] The offset information corresponding to the input data to be loaded by each thread is the thread offset of the thread.
[0055] In a possible implementation, the number of output data in the first data block is the same as the number of threads in the execution unit, and among the multiple input data in the second data block, only the input data that needs to be used as an interpolation reference point comes from the input tensor;
[0056] Each of the threads, for input data that the thread needs to load, stores the input data that the thread needs to obtain according to the target logical address into the shared memory in the order indicated by the corresponding logical coordinates, including:
[0057] Each of the threads, when determining that input data of an interpolation reference point needs to be obtained according to the thread offset of the thread, obtains the input data from a physical address corresponding to the target logical address, or determines a logical address of the input data to be obtained according to the thread offset of the thread and the target logical address, and then obtains the input data from a physical address determined according to the logical address, and stores the input data to the shared memory in an order indicated by the corresponding logical coordinates;
[0058] In a case where it is determined according to the thread offset of the thread that there is no need to obtain input data of the interpolation reference point, a preset value obtained from a preset physical address is determined as the input data, or the preset value is directly determined as the input data, and the input data is stored in the shared memory according to the storage position corresponding to the thread offset;
[0059] The offset information corresponding to the input data to be loaded by each thread is the thread offset of the thread.
[0060] In a possible implementation, each of the threads, for output data to be calculated by the thread, obtains four input data corresponding to the output data from the second data block according to corresponding offset information and the target size, including:
[0061] For output data to be calculated by the thread, the input data at the upper left coordinate position among four input data corresponding to the output data is acquired from the second data block according to the thread offset of the thread;
[0062] For output data to be calculated by the thread, obtaining the remaining three input data of the four input data corresponding to the output data from the second data block according to the thread offset of the thread and the target size;
[0063] The offset information corresponding to the output data to be calculated by each thread is the thread offset of the thread.
[0064] In a possible implementation, when the input tensor includes multiple layers, the output tensor where the first data block is located is the 0th layer input tensor, and the number of first data blocks in the input tensors of each layer is the target number.
[0065] The control unit, when determining that the input tensor includes a plurality of layers, further determines a target physical address corresponding to the target logical address;
[0066] When it is determined that the input data to be obtained does not belong to the input tensor of the 0th layer, each of the threads will store the input data obtained from the physical address determined according to the target physical address, the first offset corresponding to the number of layers of the input tensor where the input data to be obtained is located, and the offset information corresponding to the input data in the shared memory in the order indicated by the corresponding logical coordinates; or, store the input data obtained from the physical address determined according to the target logical address, the first offset corresponding to the number of layers of the input tensor where the input data to be obtained is located, and the offset information corresponding to the input data in the shared memory in the order indicated by the corresponding logical coordinates.
[0067] In a possible implementation, the positioning data is output data located at the upper left position on the boundary among multiple output data of the first data block, and / or
[0068] If there is a corresponding relationship between the difference in storage positions of the input data in the shared memory and the difference in indexes between the threads, the thread offset is the index of the thread.
[0069] In a possible implementation manner, the control unit is further configured to store the determined interpolation weights in a preset register or in the shared memory.
[0070] According to another aspect of the present disclosure, an electronic device is provided, comprising: a processor; and a memory for storing processor-executable instructions; wherein the processor is configured to implement the above method when executing the instructions stored in the memory.
[0071] According to another aspect of the present disclosure, a non-volatile computer-readable storage medium is provided, on which computer program instructions are stored, wherein the computer program instructions implement the above method when executed by a processor.
[0072] According to another aspect of the present disclosure, a computer program product is provided, including a computer-readable code, or a non-volatile computer-readable storage medium carrying the computer-readable code. When the computer-readable code runs in a processor of an electronic device, the processor in the electronic device executes the above method.
[0073] A data processing method and device provided by the embodiment of the present disclosure, for each first data block in the output tensor, according to the logical coordinates of the positioning data in the first data block, calculates the target logical address of the target input data corresponding to the positioning data in the input tensor, and the positioning data is one of the multiple output data included in the first data block. For the loading of each input data, the input data obtained according to the target logical address is stored in the shared memory in the order indicated by the corresponding logical coordinates, so as to store the second data block including multiple input data in the continuous space of the shared memory, and the size of the second data block and the size of the first data block are both the target size. For the calculation of each output data, according to the corresponding offset information and the target size, four input data corresponding to the output data are obtained from the second data block, and the four input data are used as interpolation reference points combined with the pre-calculated and stored interpolation weights to perform bilinear interpolation calculation to obtain the corresponding output data. The loading time of the input data is reduced, the time occupied by the address conversion is reduced, and it is suitable for application scenarios of various memory architectures that support tensor layout.
[0074] Further features and aspects of the present disclosure will become apparent from the following detailed description of exemplary embodiments with reference to the attached drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0075] The accompanying drawings, which are incorporated in and constitute a part of the specification, illustrate exemplary embodiments, features, and aspects of the disclosure and, together with the description, serve to explain the principles of the disclosure.
[0076] Figure 1 A schematic diagram of the upsampling process in the related art is shown.
[0077] Figure 2 A flow chart of a data processing method according to an embodiment of the present disclosure is shown.
[0078] Figure 3 A block diagram of a data processing device according to an embodiment of the present disclosure is shown.
[0079] Figure 4 A schematic diagram showing the correspondence between output data and interpolation reference points in a one-dimensional tensor scenario in a data processing method according to an embodiment of the present disclosure.
[0080] Figure 5 A schematic diagram showing a second data block loaded by loading mode 1 in a data processing method according to an embodiment of the present disclosure.
[0081] Figure 6 A schematic diagram showing a second data block loaded by loading mode 2 in a data processing method according to an embodiment of the present disclosure.
[0082] Figure 7A schematic diagram showing a second data block loaded by loading mode three in a data processing method according to an embodiment of the present disclosure is shown.
[0083] Figure 8 A schematic diagram showing the correspondence between output data and interpolation reference points in a multidimensional tensor scenario in a data processing method according to an embodiment of the present disclosure.
[0084] Fig. 9 It is a block diagram of a device 1900 for data processing according to an exemplary embodiment. DETAILED DESCRIPTION
[0085] Various exemplary embodiments, features and aspects of the present disclosure will be described in detail below with reference to the accompanying drawings. The same reference numerals in the accompanying drawings represent elements with the same or similar functions. Although various aspects of the embodiments are shown in the accompanying drawings, the drawings are not necessarily drawn to scale unless otherwise specified.
[0086] The word “exemplary” is used exclusively herein to mean “serving as an example, example, or illustration.” Any embodiment described herein as “exemplary” is not necessarily to be construed as preferred or advantageous over other embodiments.
[0087] In addition, in order to better illustrate the present disclosure, numerous specific details are given in the following specific embodiments. It should be understood by those skilled in the art that the present disclosure can also be implemented without certain specific details. In some examples, methods, means, components and circuits well known to those skilled in the art are not described in detail in order to highlight the main purpose of the present disclosure.
[0088] Operator is a computing unit in an artificial intelligence neural network model, which is used to perform corresponding mathematical calculations on tensors transmitted in the network model. There are various types of operators for different computing requirements. Among them, upsampling operators can sample and interpolate (such as nearest neighbor interpolation, bilinear interpolation, and bicubic interpolation) the input tensor based on the set sampling mode (or sampling method) to obtain an enlarged output vector.
[0089] For the upsampling operator based on bilinear interpolation, since one output data needs to be generated by interpolation calculation based on four input data, the entire upsampling process is as follows: Figure 1As shown, specifically: obtain the logical coordinates (Ox, Oy) of the output data at a certain point in the output tensor, and then calculate the logical coordinates of the corresponding four interpolation reference points based on the logical coordinates. Assuming that the four interpolation reference points are s00, s01, s10 and s11, the logical coordinates of the four interpolation reference points s00, s01, s10 and s11 can be expressed as (x, y), (x+1, y), (x, y+1) and (x+1, y+1). Then calculate the memory addresses (ie, physical addresses) of the corresponding four input data based on the logical coordinates of the interpolation reference points, and load the input data in the memory address as the value of the interpolation reference point based on the memory address. After obtaining the values of the four interpolation reference points, it is also necessary to calculate the interpolation coefficient (ie, interpolation weight) based on the logical coordinates of the output data and the magnification, and then calculate the value of the corresponding output data based on the value of the interpolation reference point and the interpolation coefficient. Then perform sampling and interpolation calculations on the next output data until all output calculations are completed. The detailed calculation process is as follows:
[0090] x_ratio=[(Ox+0.5) / scale_fator_x-0.5]-floor((Ox+0.5) / scale_factor_x-0.5)
[0091] y_ratio=[(Oy+0.5) / scale_fator_y-0.5]-floor((Oy+0.5) / scale_factor_y-0.5)
[0092] result=(1-y_ratio)*(1-x_ratio)*s00+y_ratio*(1-x_ratio)*s10+x_ratio*(1-y_ratio)*s01+x_ratio*y_ratio*s11
[0093] Among them, scale_factor_x and scale_factor_y represent the magnification of the current upsampling operation in the x and y directions. x_ratio and y_ratio represent the interpolation coefficients when calculating the interpolation of the output data at this point. result represents the interpolation calculation result (also the output data of this point) of the point with logical coordinates (Ox, Oy) to be stored in the output tensor. s00, s01, s10 and s11 represent the values of the four interpolation reference points respectively.
[0094] For GPGPU (General-Purpose Graphics Processing Unit) chip software stack, it refers to the software architecture used to perform general computing tasks on the graphics processing unit (GPU). The common hardware implementation in the related art is a memory architecture with tensor layout. Under this memory architecture, the most efficient way to access data is often "aligned layout". In other words, if the memory layout is N-byte aligned, it can be regarded as divided into multiple N-byte blocks. Then, for a continuous N or integer multiple of N bytes, and the starting point is also the starting point of the block, the hardware can access the memory faster than the memory access of other sizes and starting coordinates. This article hereinafter refers to "block access". Correspondingly, the hardware in the related art generally accesses data more directly in the form of buffer, that is, linearly accesses a string of memory addresses. In contrast, the efficiency of data access in block form is generally much better than the efficiency of access in buffer form.
[0095] In the related art, the Bilinear upsampling operator is implemented as follows: Figure 1 The direct calculation process performs sampling and interpolation calculations. Due to different algorithms, the address mapping of a continuous block of data points on the output to the input data is often not a whole block. Even under different logical coordinates, the same set of algorithms will have different mapping rules or parameters, resulting in the address mapping result being not continuous. Therefore, the "loading interpolation reference points according to the address" process implemented by the bilinear upsampling operator in the related technology cannot be directly "loaded by block", but can only select point-to-point loading to calculate the input data in the input tensor required. Here, a large number of "loading in buffer form" are often used, directly converting the address according to the logical coordinates of the output point to obtain the corresponding physical address, and then executing the required point-to-point loading process based on the physical address.
[0096] In the bilinear interpolation upsampling process, one piece of output data requires four pieces of input data as interpolation reference points, and the bilinear interpolation algorithm needs to consider multiple ways of reordering at the boundary due to the align-corner rule, resulting in excessive performance loss (warp divergence) of the execution unit in the calculation process, affecting the execution efficiency. Furthermore, the calculation process of the coordinate mapping operation from output to input is relatively complicated, and loading according to the address of the obtained misaligned input data is usually slow. This leads to the bilinear upsampling operator under this idea to contain a large amount of "loading in buffer form" and a large amount of "logical address to physical address" conversion calculations. These will make the generated assembly code kernel run inefficiently, and further cause the performance of some computer vision models that use this operator to be far lower than expected.
[0097] For this situation, the practice in the related technology is mostly to use shared memory to complete the rearrangement of input data, so as to ensure that the memory access to input data and output data is "block-by-block" as much as possible. However, due to the complexity and breadth of the algorithm itself, and limited by the support form of shared memory on the specific hardware platform, there is no set of effective and universal optimization methods to implement this process. It can only be optimized for specific scenarios, including the input and output of operators, algorithm modes, and the characteristics of the hardware platform. Specifically, in the process of linear interpolation upsampling, it is generally "loaded in buffer form" to shared memory, but because there are many overlaps in the interpolation reference points required for adjacent output data, this method has a large number of repeated loading and repeated address calculations. Since "loading in buffer form" is slow in itself, and the buffer address calculation is slow in "memory architecture with tensor layout", this approach also has low performance. Another approach in the related art is to calculate the entire address calculation result in advance by the host and put it in a preloaded tensor, and convert the device calculation process to read the tensor. However, the problem with this approach is that the pre-calculation process takes a long time and the host compilation is slow. At the same time, this method requires time-consuming re-arrangement of data to import it into the device, and occupies additional video memory of the device, which is not a feasible solution in business scenarios.
[0098] Therefore, in the process of linear interpolation upsampling, how to reduce the number of "loading in buffer form" and the number of "buffer address calculation" as much as possible is a technical problem that needs to be solved urgently.
[0099] In order to solve the above technical problems, the embodiment of the present disclosure provides a data processing method and device. For each first data block in the output tensor, according to the logical coordinates of the positioning data in the first data block, the target logical address of the target input data corresponding to the positioning data in the input tensor is calculated, and the positioning data is one of the multiple output data included in the first data block. For the loading of each input data, the input data obtained according to the target logical address is stored in the shared memory in the order indicated by the corresponding logical coordinates, so as to store the second data block including multiple input data in the continuous space of the shared memory, and the size of the second data block and the size of the first data block are both the target size. For the calculation of each output data, according to the corresponding offset information and the target size, four input data corresponding to the output data are obtained from the second data block, and the four input data are used as interpolation reference points combined with the pre-calculated and stored interpolation weights to perform bilinear interpolation calculation to obtain the corresponding output data. The loading time of the input data is reduced, the time occupied by the address conversion is reduced, and it is suitable for application scenarios of various memory architectures that support tensor layout.
[0100] The present disclosure provides a data processing method, which includes being applied to a processor of a memory architecture supporting tensor layout, such as Figure 2 As shown, the method includes steps S101 to S103. The method can be Figure 3 The data processing device shown executes, as Figure 3 As shown, the data processing device includes a control unit 11 and at least one execution unit (warp) 12, each of which includes multiple threads. The number of execution units 12 in the device can be one or more. Figure 2 Only one execution unit is schematically drawn. The device can be applied to a processor of a memory architecture that supports tensor layout. In this embodiment, the entire data processing process can be implemented based on a single instruction multiple threads (SIMT) method to complete bilinear upsampling. In this way, the data processing method is applicable to all SIMT architectures with tensor memory layout and efficient block memory access capabilities. It is only necessary to adjust the amount of loaded data (that is, the target size described below) and the multiplexing process of adjusting the interpolation weight and memory address calculation according to the size of the memory block in the specific hardware scenario and the mapping relationship between the memory data block and the logical data block.
[0101] In step S101, for each first data block in the output tensor, the target logical address of the target input data corresponding to the positioning data in the input tensor is calculated according to the logical coordinates of the positioning data in the first data block, and the positioning data is one of the multiple output data included in the first data block. In some embodiments, in order to simplify the address calculation process of subsequent input data loading, the positioning data is the output data at the upper left position on the boundary of the input data required for bilinear interpolation calculation of multiple output data of the first data block.
[0102] In step S102, for each input data loading, the input data obtained according to the target logical address is stored in the shared memory in the order indicated by the corresponding logical coordinates, so as to store a second data block including multiple input data in the continuous space of the shared memory, and the size of the second data block and the size of the first data block are both the target size. In this way, the four times of "buffer loading" of the four interpolation reference point input data of one output data in the related art is reduced to one time, which greatly reduces the time consumption.
[0103] In step S103, for the calculation of each output data, four input data corresponding to the output data are obtained from the second data block according to the corresponding offset information and the target size, and the four input data are used as interpolation reference points in combination with the pre-calculated and stored interpolation weights to perform bilinear interpolation calculation to obtain the corresponding output data. In this way, each input data is stored in an independent location of the shared memory, ensuring that the shared memory is aligned to avoid possible bank conflicts.
[0104] In this embodiment, step S101 can be performed by Figure 3 The control unit 11 shown is executed, and the loading of each input data in step S102 and the calculation of each output data in step S103 can be executed by the corresponding thread in the execution unit 12.
[0105] In some embodiments, the interpolation weights may be determined in advance by the control unit 11 before step S103 is executed, and stored in a preset register or in the shared memory.
[0106] In some embodiments, the method may further include: outputting and storing the first data block formed according to the corresponding logical coordinate sequence of each output data, and this step may be performed by the control unit 11. Alternatively, in the calculation of each output data in step S103, after the output data is obtained, the output data is output and stored according to the corresponding logical coordinate sequence. In this way, each output data in the first data block can be stored in blocks, thereby improving the overall efficiency and speed.
[0107] In order to more intuitively and clearly illustrate the working principle and process of the data processing method and device, the following is combined with Figure 4-Figure 8 The example of “using a device including an execution unit 12 provided with 32 threads to perform a bilinear upsampling process with a magnification of 2 times” is given to schematically illustrate the method and device provided in the embodiment of the present disclosure.
[0108] like Figure 4 As shown, in the process of bilinear upsampling the input tensor to obtain a 2-fold larger output tensor according to the method of the embodiment of the present disclosure, when the number of output data in the first data block is the same as the number of threads in the execution unit 12, each thread can load a specified input data and calculate an output data, which can improve the overall efficiency and speed, and each input data is stored in an independent location of the shared memory to ensure that the shared memory is aligned with the number of threads to avoid possible memory conflicts. At this time, the target size can be set according to the number of threads, such as the number of data in the target size can be set to be the same as the number of threads, and the length and width of the target size can be set according to the needs of the interpolation reference point to ensure that the interpolation reference points (i.e., input data) for calculating each output data in the first data block are loaded into the second data block. Then the output tensor can be divided into multiple first data blocks according to the target size, such as multiple 4*8 first data blocks. For any output data in each first data block, its four input data in the input tensor as interpolation reference points can be calculated, and then the second data blocks corresponding to each first data block can be obtained from the input tensor according to the calculation results and stored in the shared memory in a 4*8 arrangement. For example, the four input data as interpolation reference points of output data 1 in the first data block are input data 1, 2, 9, 10 in the second data block; the four input data as interpolation reference points of output data 2 in the first data block are input data 1, 2, 9, 10 in the second data block; the four input data as interpolation reference points of output data 3 in the first data block are input data 2, 3, 10, 11 in the second data block, and so on. In fact, the input data of each output data as an interpolation reference point in the second data block can be determined through mapping transformation based on the magnification and the logical coordinates of the output data. Then a reasonable calculation is performed, such as Figure 4 As shown, we can determine that when the multiple output data of the first data block are magnified by 2 times and arranged in 4*8 format, the interpolation reference points corresponding to all the output data in the first data block have the pattern shown in Table 1 below, that is, when the multiple output data of the first data block are magnified by 2 times and arranged in 4*8 format, in the second data block of the same size, the reference interpolation points required for calculating each output data are only the input data within the upper left 3*5 range in the second data block.
[0109] Table 1 Examples of interpolation reference points
[0110] All upper left interpolation points 1、2、3、4、9、10、11、12 All upper right interpolation points 2、3、4、5、10、11、12、13 All lower left interpolation points 9、10、11、12、17、18、19、20 All lower right interpolation points 10、11、12、13、18、19、20、21
[0111] It can be seen that, in the case of specifying the magnification and the number of threads, obtaining a second data block that is equal in size to the first data block can ensure that all interpolation reference points for calculating each output data in the first data block are loaded. In the process of using multiple threads in the execution unit in the processor to load input data, if the number of output data in the first data block is the same as the number of threads in the execution unit, there are the following different implementation methods for loading each input data in step S102.
[0112] Loading method 1:
[0113] If it is set that each input data in the second data block can all come from the input tensor. For example, Figure 4 The input data 1-32 are all data from the input tensor. Then, the loading steps for each input data in step 102 may include:
[0114] Controlling a first thread among the multiple threads to obtain input data from a target physical address corresponding to the target logical address, and storing the input data to a shared memory in an order indicated by corresponding logical coordinates; or
[0115] Controlling each thread among the multiple threads different from the first thread, determining a logical address of input data to be obtained according to a thread offset of the thread and the target logical address, and then obtaining the input data from a physical address determined according to the logical address, and storing the input data to a shared memory in an order indicated by corresponding logical coordinates;
[0116] The offset information corresponding to the input data to be loaded by each thread is the thread offset of the thread.
[0117] For example, for Figure 4 , Figure 5In the example shown, input data is loaded according to loading method 1, and thread 1, as the first thread, loads input data 1. The process is to obtain input data 1 from the target physical address corresponding to the target logical address (such as (0,0)), and store the input data 1 to the shared memory in the order indicated by the corresponding logical coordinates. Thread 2 determines the logical address (1,0) of the input data to be obtained based on the thread offset "such as 2" and the target logical address (0,0), and then obtains input data 2 from the physical address determined according to the logical address (1,0), and stores the input data 2 to the shared memory in the order indicated by the corresponding logical coordinates; and so on, threads 3-32 also complete the loading of the corresponding input data. Finally, we get Figure 5 The second data block is shown.
[0118] Loading method 2:
[0119] If only the input data needed as the interpolation reference point among the multiple input data in the second data block comes from the input tensor. Figure 4 In the example shown, it is only necessary to ensure that the input data in the upper left 3*5 area of the second data block is actually from the input tensor. Since the input data in the remaining areas will not be used in step S103, they can be loaded according to the settings, for example, all of them load a certain specified input data in the input tensor, and in order to improve efficiency and speed, the specified input data can all be target input data. At this time, the loading steps for each input data in step 102 may include:
[0120] Controlling each of the threads to obtain the input data of the interpolation reference point from the target physical address corresponding to the target logical address when it is determined according to the thread offset of the thread that the input data of the interpolation reference point needs to be obtained, or determining the logical address of the input data to be obtained according to the thread offset of the thread and the target logical address, and then obtaining the input data from the physical address determined according to the logical address, and storing the input data to the shared memory in the order indicated by the corresponding logical coordinates;
[0121] Control each of the threads to obtain input data from a target physical address corresponding to the target logical address when it is determined according to the thread offset of the thread that there is no need to obtain input data of an interpolation reference point, and store the input data in a shared memory according to a storage location corresponding to the thread offset;
[0122] The offset information corresponding to the input data to be loaded by each thread is the thread offset of the thread.
[0123] For example, for Figure 4 , Figure 6In the example shown, the input data is loaded according to loading method 2, and thread 1, as the first thread, loads input data 1. The process is to obtain input data 1 from the target physical address corresponding to the target logical address (such as (0,0)), and store the input data 1 to the shared memory in the order indicated by the corresponding logical coordinates. Thread 2 determines the logical address (1,0) of the input data to be obtained based on the thread offset "such as 2" of thread 2 and the target logical address (0,0), and then obtains input data 2 from the physical address determined according to the logical address (1,0), and stores the input data 2 to the shared memory in the order indicated by the corresponding logical coordinates; and so on, threads 3-5, 9-13, 17-18 also complete the loading of the corresponding input data. The remaining threads can obtain input data 1 from the target physical address corresponding to the target logical address (such as (0,0)) like thread 1, and store the input data 1 to the shared memory in the order indicated by the logical coordinates corresponding to the thread itself. Finally, we get Figure 6 The second data block is shown.
[0124] Loading method three:
[0125] The same is true when only the input data that is needed as the interpolation reference point among the multiple input data in the second data block comes from the input tensor. Figure 4 In the example shown, it is only necessary to ensure that the input data in the upper left 3*5 area of the second data block is actually from the input tensor. Since the input data in the remaining areas will not be used in step S103, they can be loaded according to the settings, for example, a preset value is loaded as input data and stored in the shared memory. The preset value can be directly stored in the shared memory or a specified register for each thread to obtain and store, or a specific value of the preset value such as "0" can be pre-set, and each thread can directly store the preset value. At this time, step 102 for each input data loading step may include:
[0126] Controlling each of the threads to obtain the input data of the interpolation reference point from the physical address corresponding to the target logical address when it is determined according to the thread offset of the thread that the input data of the interpolation reference point needs to be obtained, or to determine the logical address of the input data to be obtained according to the thread offset of the thread and the target logical address, and then obtain the input data from the physical address determined according to the logical address, and store the input data to the shared memory in the order indicated by the corresponding logical coordinates;
[0127] Control each of the threads to determine a preset value obtained from a preset physical address as input data, or directly determine the preset value as input data, when it is determined according to the thread offset of the thread that there is no need to obtain input data of an interpolation reference point, and store the input data in a shared memory according to a storage location corresponding to the thread offset;
[0128] The offset information corresponding to the input data to be loaded by each thread is the thread offset of the thread.
[0129] For example, for Figure 4 , Figure 7 In the example shown, input data is loaded according to loading method three, and thread 1, as the first thread, loads input data 1. The process is to obtain input data 1 from the target physical address corresponding to the target logical address (such as (0,0)), and store the input data 1 to the shared memory in the order indicated by the corresponding logical coordinates. Thread 2 determines the logical address (1,0) of the input data to be obtained based on the thread offset "such as 2" of thread 2 and the target logical address (0,0), and then obtains input data 2 from the physical address determined according to the logical address (1,0), and stores the input data 2 to the shared memory in the order indicated by the corresponding logical coordinates; and so on, threads 3-5, 9-13, 17-18 also complete the loading of the corresponding input data. As for the remaining threads, the preset value can be used as input data like thread 1, and the preset value can be stored in the shared memory in the order indicated by the logical coordinates corresponding to the thread itself. Finally, we get Figure 7 The second data block is shown.
[0130] In some embodiments, if there is a corresponding relationship between the difference in storage position of the input data in the shared memory and the difference in index between the threads, the thread offset is the index of the thread. Wherein, the difference in storage position of the input data in the shared memory may be the same as the difference in index between the threads, and the index of the thread may be the sequential number of the thread. In the implementation of the above-mentioned loading methods 1, 2, and 3, the address calculation may be performed directly based on the index of the thread.
[0131] Continue with Figure 4 Taking the "32 threads performing bilinear upsampling with a magnification of 2 times" as an example, in order to improve efficiency and speed, each thread can perform a calculation of a specified output data, wherein the input data loading and output data calculation performed by each thread can be specified as a data with the same logical coordinates, for example, thread 1 performs Figure 4 In the case where the number of output data in the first data block is the same as the number of threads in the execution unit 12, step S103 may include:
[0132] Control each of the threads to obtain input data at an upper left coordinate position among four input data corresponding to the output data from the second data block according to the thread offset of the thread;
[0133] Acquire the remaining three input data of the four input data corresponding to the output data from the second data block according to the thread offset of the thread and the target size;
[0134] The offset information corresponding to the output data to be calculated by each thread is the thread offset of the thread.
[0135] For example, if Figure 4 As shown, thread 1, thread 1 calculates output data 1, firstly calculates the address of the input data of the upper left coordinate position to be obtained in the shared memory according to its own thread index and the base address of the second data block in the shared memory, and then obtains the input data 1 of the upper left coordinate position; then, based on its own thread index, the offset between other difference reference points and the interpolation reference point of the upper left coordinate position calculated according to the target size, and the base address of the second data block in the shared memory, calculates the addresses of the other three input data to be obtained in the shared memory, and then obtains the remaining three input data, namely input data 2, 9, and 10, and then calculates to obtain output data 1. Similarly, other threads also synchronously obtain the input data of the difference reference point and the subsequent corresponding output data calculation.
[0136] In this way, the above steps S101 to S103 are continuously repeated for each first data block until all output data in the output tensor have been calculated, and then the calculation can be stopped.
[0137] The above describes the process of bilinear upsampling of a layer of input tensor to obtain a layer of output tensor. When the input tensor includes multiple layers, the output tensor where the first data block is located is the 0th layer input tensor, and the number of first data blocks in each layer of input tensor is the target number. The number of layers of a tensor can refer to the dimension of the tensor, such as Figure 8 As shown in the figure, for the bilinear interpolation upsampling process on a plane composed of multiple dimensions, all points on the coordinates of an axis perpendicular to the plane (called the equivalent magnification position) have the same interpolation weight, so the tensors of other layers except the 0th layer can reuse the logical coordinate calculation process of the 0th layer. For hardware memory with tensor layout, the target physical address plus the offset can also be directly reused to speed up the calculation process of "converting logical coordinates to physical addresses" of different "blocks". Then:
[0138] In step S101, when it is determined that the input tensor includes multiple layers, a target physical address corresponding to the target logical address is also determined.
[0139] In step S102, if the physical addresses of input tensors of different layers can be calculated by offset, then for the loading of each input data, when it is determined that the input data to be obtained belongs to the 0th layer input tensor, the input data obtained according to the target logical address is stored in the shared memory in the order indicated by the corresponding logical coordinates. When it is determined that the input data to be obtained does not belong to the 0th layer input tensor, the input data obtained from the physical address determined according to the target physical address, the first offset corresponding to the number of layers of the input tensor where the input data to be obtained is located, and the offset information corresponding to the input data is stored in the shared memory in the order indicated by the corresponding logical coordinates. That is, based on the target physical address Addr_Mem_O1 (Ox0, Oy0), the physical address of the equivalent magnification position of each layer is obtained in the manner of Addr_Mem_O1 (Ox0, Oy0) + number of layers * total number of first data blocks of a single layer * amount of data corresponding to the target size, thereby realizing the loading of each non-0th layer input data.
[0140] If the physical addresses of input tensors of different layers can be calculated by offset, then for the loading of each input data, when it is determined that the input data to be obtained belongs to the input tensor of the 0th layer, the input data obtained according to the target logical address will be stored in the shared memory in the order indicated by the corresponding logical coordinates. When it is determined that the input data to be obtained does not belong to the input tensor of the 0th layer, the input data obtained from the physical address determined according to the target logical address, the first offset corresponding to the number of layers of the input tensor where the input data to be obtained is located, and the offset information corresponding to the input data will be stored in the shared memory in the order indicated by the corresponding logical coordinates.
[0141] In this way, the time-consuming calculation of "coordinate mapping to physical address" in the bilinear interpolation upsampling process for each output data can be reused in all "equivalent magnification positions" (that is, corresponding positions on different tensor layers). This maximizes the data reuse rate and reduces the overall workload of physical address calculation.
[0142] For the data processing method provided in the embodiment of the present disclosure, the target sizes of the first data block and the second data block can be adjusted according to different application scenarios such as the actual hardware architecture and interpolation weights, and the address mapping relationship calculation method can be adaptively set, thereby realizing the adaptive setting of bilinear upsampling in different application scenarios.
[0143] In some embodiments, the functions or modules included in the device provided by the embodiments of the present disclosure can be used to execute the method described in the above method embodiments. The specific implementation can refer to the description of the above method embodiments, and for the sake of brevity, it will not be repeated here.
[0144] The embodiment of the present disclosure also provides a computer-readable storage medium on which computer program instructions are stored, and the computer program instructions implement the above method when executed by a processor. The computer-readable storage medium can be a volatile or non-volatile computer-readable storage medium.
[0145] An embodiment of the present disclosure further proposes an electronic device, comprising: a processor; and a memory for storing instructions executable by the processor; wherein the processor is configured to implement the above method when executing the instructions stored in the memory.
[0146] The embodiments of the present disclosure also provide a computer program product, including a computer-readable code, or a non-volatile computer-readable storage medium carrying the computer-readable code. When the computer-readable code runs in a processor of an electronic device, the processor in the electronic device executes the above method.
[0147] Fig. 9 1 is a block diagram of a device 1900 for data processing according to an exemplary embodiment. For example, the device 1900 may be provided as a server or an electronic device. Fig. 9 , the apparatus 1900 includes a processing component 1922, which further includes one or more processors, and a memory resource represented by a memory 1932 for storing instructions, such as an application, that can be executed by the processing component 1922. The application stored in the memory 1932 may include one or more modules, each corresponding to a set of instructions. In addition, the processing component 1922 is configured to execute instructions to perform the above method.
[0148] The device 1900 may also include a power supply component 1926 configured to perform power management of the device 1900, a wired or wireless network interface 1950 configured to connect the device 1900 to a network, and an input / output (I / O) interface 1958. The device 1900 may operate based on an operating system stored in the memory 1932, such as Windows Server™, MacOS X™, Unix™, Linux™, FreeBSD™, or the like.
[0149] In an exemplary embodiment, a non-volatile computer-readable storage medium is also provided, such as a memory 1932 including computer program instructions, which can be executed by the processing component 1922 of the device 1900 to perform the above method.
[0150] The present disclosure may be a system, a method and / or a computer program product. The computer program product may include a computer-readable storage medium carrying computer-readable program instructions for causing a processor to implement various aspects of the present disclosure.
[0151] A computer-readable storage medium may be a tangible device that can hold and store instructions used by an instruction execution device. A computer-readable storage medium may be, for example, but not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples of computer-readable storage media (a non-exhaustive list) include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disk read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanical encoding device, such as a punch card or a raised structure in a groove on which instructions are stored, and any suitable combination of the foregoing. As used herein, a computer-readable storage medium is not to be interpreted as a transient signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., a light pulse through a fiber optic cable), or an electrical signal transmitted through a wire.
[0152] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to each computing / processing device, or downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network can include copper transmission cables, optical fiber transmissions, wireless transmissions, routers, firewalls, switches, gateway computers, and / or edge servers. The network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in the computer-readable storage medium in each computing / processing device.
[0153] The computer program instructions for performing the operation of the present disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-related instructions, microcode, firmware instructions, state setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages, such as Smalltalk, C++, etc., and conventional procedural programming languages, such as "C" language or similar programming languages. Computer-readable program instructions may be executed completely on a user's computer, partially on a user's computer, as an independent software package, partially on a user's computer, partially on a remote computer, or completely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., using an Internet service provider to connect via the Internet). In some embodiments, an electronic circuit, such as a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA), may be customized by utilizing the state information of the computer-readable program instructions, and the electronic circuit may execute the computer-readable program instructions, thereby realizing various aspects of the present disclosure.
[0154] Various aspects of the present disclosure are described herein with reference to the flowcharts and / or block diagrams of the methods, devices (systems) and computer program products according to the embodiments of the present disclosure. It should be understood that each box in the flowchart and / or block diagram and the combination of each box in the flowchart and / or block diagram can be implemented by computer-readable program instructions.
[0155] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a machine, so that when these instructions are executed by the processor of the computer or other programmable data processing device, a device that implements the functions / actions specified in one or more boxes in the flowchart and / or block diagram is generated. These computer-readable program instructions can also be stored in a computer-readable storage medium, and these instructions cause the computer, programmable data processing device, and / or other equipment to work in a specific manner, so that the computer-readable medium storing the instructions includes a manufactured product, which includes instructions for implementing various aspects of the functions / actions specified in one or more boxes in the flowchart and / or block diagram.
[0156] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device so that a series of operating steps are performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to implement the functions / actions specified in one or more boxes in the flowchart and / or block diagram.
[0157] The flow chart and block diagram in the accompanying drawings show the possible architecture, function and operation of the system, method and computer program product according to multiple embodiments of the present disclosure. In this regard, each square box in the flow chart or block diagram can represent a part of a module, program segment or instruction, and a part of the module, program segment or instruction includes one or more executable instructions for realizing the specified logical function. In some alternative implementations, the function marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two continuous square boxes can actually be executed substantially in parallel, and they can sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each square box in the block diagram and / or flow chart, and the combination of the square boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs the specified function or action, or can be implemented with a combination of special hardware and computer instructions.
[0158] The embodiments of the present disclosure have been described above, and the above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and changes will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The selection of terms used herein is intended to best explain the principles of the embodiments, practical applications, or technical improvements in the market, or to enable other persons of ordinary skill in the art to understand the embodiments disclosed herein.
Claims
1. A data processing method, characterized in that: Applied in a processor having a memory architecture supporting tensor layout, the method comprises: For each first data block in the output tensor, according to the logical coordinates of the positioning data in the first data block, calculate a target logical address of the target input data corresponding to the positioning data in the input tensor, where the positioning data is one of the plurality of output data included in the first data block; For each input data loaded, the input data acquired according to the target logical address is stored in the shared memory in the order indicated by the corresponding logical coordinates, so as to store a second data block including the plurality of input data in a continuous space of the shared memory, wherein the size of the second data block and the size of the first data block are both the target size; For calculation of each output data, four input data corresponding to the output data are obtained from the second data block according to the corresponding offset information and the target size, and the four input data are used as interpolation reference points in combination with pre-calculated and stored interpolation weights to perform bilinear interpolation calculation to obtain the corresponding output data; The input data are loaded respectively using a plurality of threads in the execution unit in the processor, and the number of output data in the first data block is the same as the number of threads in the execution unit.
2. The method according to claim 1, characterized in that The method further comprises: Output and store a first data block formed according to the respective output data in the order of corresponding logical coordinates; or For the calculation of each output data, after the output data is obtained, it is output and stored in order according to the corresponding logical coordinates.
3. The method according to claim 1, characterized in that Each input data in the second data block comes from the input tensor; The input data acquired according to the target logical address is stored in the shared memory in the order indicated by the corresponding logical coordinates, including: Controlling a first thread among the multiple threads to obtain input data from a target physical address corresponding to the target logical address, and storing the input data to a shared memory in an order indicated by corresponding logical coordinates; or Controlling each thread among the multiple threads different from the first thread, determining a logical address of input data to be obtained according to a thread offset of the thread and the target logical address, and then obtaining the input data from a physical address determined according to the logical address, and storing the input data to a shared memory in an order indicated by corresponding logical coordinates; The offset information corresponding to the input data to be loaded by each thread is the thread offset of the thread.
4. The method according to claim 1, characterized in that: Among the plurality of input data in the second data block, only the input data required to be used as interpolation reference points comes from the input tensor; The input data acquired according to the target logical address is stored in the shared memory in the order indicated by the corresponding logical coordinates, including: Controlling each of the threads to obtain the input data of the interpolation reference point from the target physical address corresponding to the target logical address when it is determined according to the thread offset of the thread that the input data of the interpolation reference point needs to be obtained, or determining the logical address of the input data to be obtained according to the thread offset of the thread and the target logical address, and then obtaining the input data from the physical address determined according to the logical address, and storing the input data to the shared memory in the order indicated by the corresponding logical coordinates; Control each of the threads to obtain input data from a target physical address corresponding to the target logical address when it is determined according to the thread offset of the thread that there is no need to obtain input data of an interpolation reference point, and store the input data in a shared memory according to a storage location corresponding to the thread offset; The offset information corresponding to the input data to be loaded by each thread is the thread offset of the thread.
5. The method according to claim 1, characterized in that Among the plurality of input data in the second data block, only the input data required to be used as interpolation reference points comes from the input tensor; The input data acquired according to the target logical address is stored in the shared memory in the order indicated by the corresponding logical coordinates, including: Controlling each of the threads to obtain the input data of the interpolation reference point from the physical address corresponding to the target logical address when it is determined according to the thread offset of the thread that the input data of the interpolation reference point needs to be obtained, or to determine the logical address of the input data to be obtained according to the thread offset of the thread and the target logical address, and then obtain the input data from the physical address determined according to the logical address, and store the input data to the shared memory in the order indicated by the corresponding logical coordinates; Control each of the threads to determine a preset value obtained from a preset physical address as input data, or directly determine the preset value as input data, when it is determined according to the thread offset of the thread that there is no need to obtain input data of an interpolation reference point, and store the input data in a shared memory according to a storage location corresponding to the thread offset; The offset information corresponding to the input data to be loaded by each thread is the thread offset of the thread.
6. The method according to any one of claims 3 to 5, characterized in that: Utilizing multiple threads in the execution unit within the processor to calculate output data, Wherein, according to the corresponding offset information and the target size, four input data corresponding to the output data are obtained from the second data block, including: Control each of the threads to obtain the input data at the upper left coordinate position of the four input data corresponding to the output data from the second data block according to the thread offset of the thread; Control each of the threads to obtain the remaining three input data of the four input data corresponding to the output data from the second data block according to the thread offset of the thread and the target size; The offset information corresponding to the output data to be calculated by each thread is the thread offset of the thread.
7. The method according to claim 1, characterized in that In the case where the input tensor includes multiple layers, the output tensor where the first data block is located is the 0th layer input tensor, and the number of first data blocks in the input tensors of each layer is the target number, and the method further includes: In the case where it is determined that the input tensor includes a plurality of layers, a target physical address corresponding to the target logical address is also determined; For the loading of each input data, when it is determined that the input data to be obtained does not belong to the input tensor of the 0th layer, the input data obtained from the physical address determined according to the target physical address, the first offset corresponding to the number of layers of the input tensor where the input data to be obtained is located, and the offset information corresponding to the input data, shall be stored in the shared memory in the order indicated by the corresponding logical coordinates; or, the input data obtained from the physical address determined according to the target logical address, the first offset corresponding to the number of layers of the input tensor where the input data to be obtained is located, and the offset information corresponding to the input data shall be stored in the shared memory in the order indicated by the corresponding logical coordinates.
8. The method according to claim 6, characterized in that The positioning data is output data located at the upper left position on the boundary among the multiple output data of the first data block, and / or If there is a corresponding relationship between the difference in storage positions of the input data in the shared memory and the difference in indexes between the threads, the thread offset is the index of the thread.
9. The method according to claim 1, characterized in that: The method further comprises: The determined interpolation weights are stored in a preset register or in the shared memory.
10. A data processing device, characterized in that: Applied in a processor of a memory architecture supporting tensor layout, the device comprises a control unit and at least one execution unit, each of the execution units comprises a plurality of threads; The control unit calculates, for each first data block in the output tensor, a target logical address of target input data corresponding to the positioning data in the input tensor according to the logical coordinates of the positioning data in the first data block, wherein the positioning data is one of the plurality of output data included in the first data block; Each of the threads, for input data that needs to be loaded by the thread, stores the input data that needs to be acquired by the thread according to the target logical address in the shared memory in the order indicated by the corresponding logical coordinates, so as to store a second data block including multiple input data in a continuous space of the shared memory, wherein the size of the second data block and the size of the first data block are both the target size; Each of the threads, for output data to be calculated by the thread, obtains four input data corresponding to the output data from the second data block according to the corresponding offset information and the target size, and uses the four input data as interpolation reference points in combination with pre-calculated and stored interpolation weights to perform bilinear interpolation calculation to obtain corresponding output data; The number of output data in the first data block is the same as the number of threads in the execution unit.
11. An electronic device, characterized in that: include: processor; a memory for storing processor-executable instructions; Wherein, the processor is configured to implement the method described in any one of claims 1 to 9 when executing the instructions stored in the memory.
12. A non-volatile computer-readable storage medium having computer program instructions stored thereon, characterized in that: When the computer program instructions are executed by a processor, the method according to any one of claims 1 to 9 is implemented.
13. A computer program product comprising computer readable code, or a non-volatile computer readable storage medium carrying computer readable code, characterized in that: When the computer readable code runs in a processor of an electronic device, the processor in the electronic device executes the method according to any one of claims 1 to 9.
Citation Information
Patent Citations
Image processing method and related product
CN116385714A
Method and device for performing acceleration operation on feature data, medium and equipment
CN117575888A