Image data processing method, device, computer equipment and storage medium
By obtaining the position information of the target data block in the convolutional neural network and loading the corresponding pixel data, the characteristic target matrix is constructed, and the problems of low image data processing efficiency and high memory usage in the prior art are solved, thereby achieving more efficient image data processing.
Patent Information
- Application Number
- CN202310717124.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-06-16
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2043-06-16
AI Technical Summary
The image data processing methods in existing convolutional neural networks need to store large matrices, resulting in high memory footprint and low computing efficiency.
By obtaining the position information of the target data block, determining its pixel band coordinates in the preset data storage mode, and loading the corresponding pixel data based on these coordinates to build a feature target matrix. This method does not need to store a complete feature matrix, reducing memory footprint.
This improves the efficiency of image data processing, reduces memory usage, and simplifies the convolutional calculation process.
Smart Images

Figure CN116797444B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to an image data processing method, apparatus, computer equipment, storage medium and computer program product. Background Art
[0002] With the development of computer technology, the most important thing in convolutional neural networks is convolution calculation. Improving the calculation speed of convolution can greatly improve the training and reasoning time of neural network.
[0003] GEMM (general matrix multiplication) + im2col (from image to matrix) is widely used as a common neural network acceleration method. GEMM + im2col first uses im2col to expand the feature map (input feature map) and filter (convolution kernel) into matrices A and B respectively through dedicated hardware or software, and then performs A x B to obtain the result matrix C. Finally, col2img (the inverse process of im2col) is used to convert C into an output feature map. im2col needs to store the expanded large matrix A, resulting in more memory usage, which affects the operation efficiency to a certain extent. Summary of the invention
[0004] Based on this, it is necessary to provide an image data processing method, apparatus, computer equipment, computer-readable storage medium and computer program product that can improve the efficiency of image data processing in order to address the above technical problems.
[0005] In a first aspect, the present application provides an image data processing method. The method comprises:
[0006] Get the location information of the target data block;
[0007] Determining pixel band coordinates of a preset data storage mode corresponding to a first pixel band in the target data block according to the position information of the target data block;
[0008] determining first pixel band address information of the first pixel band based on the pixel band coordinates of the first pixel band;
[0009] Based on the first pixel band address information, determining second pixel band address information of other pixel bands in the target data block except the first pixel band;
[0010] According to the first pixel band address information and the second pixel band address information, pixel data corresponding to the target data block is loaded, and a characteristic target matrix of the target data block is determined according to the pixel data. In one embodiment, the pixel band coordinates of the preset data storage mode corresponding to the first pixel band in the target data block are determined according to the position information of the target data block, including:
[0011] Dividing the target data block to obtain multiple target sub-data blocks;
[0012] Selecting any one of the target sub-data blocks as a current target sub-data block, and determining a first pixel band from the current target sub-data block;
[0013] According to the position information of the target sub-data block, the pixel band position of the first pixel band is determined, and the pixel band position is mapped to a preset data storage mode to obtain the pixel band coordinates of the preset data storage mode corresponding to the first pixel band in the target data block.
[0014] In one embodiment, the preset data storage mode includes a channel priority storage mode or a channel folding storage mode; mapping the pixel band position to the preset data storage mode to obtain the pixel band coordinates of the preset data storage mode corresponding to the first pixel band in the target data block includes:
[0015] If the preset data storage mode is a channel priority storage mode, mapping the pixel band position to the channel priority storage mode to obtain the channel dimension coordinates, width dimension coordinates, height dimension coordinates and quantity dimension coordinates of the first pixel band in the channel priority storage mode;
[0016] If the preset data storage mode is a channel folding storage mode, the pixel band position is mapped to the channel folding storage mode to obtain the channel dimension coordinates, width dimension coordinates, height dimension coordinates, quantity dimension coordinates and grouping index dimension coordinates of the first pixel band in the channel folding storage mode.
[0017] In one embodiment, determining first pixel band address information of the first pixel band based on the pixel band coordinates of the first pixel band includes:
[0018] When the memory coordinates of the first pixel band include channel dimension coordinates, width dimension coordinates, height dimension coordinates, and quantity dimension coordinates in the channel priority storage mode, based on the width dimension coordinates and the width dimension step, the height dimension coordinates and the height dimension step, the quantity dimension coordinates and the quantity dimension step, and the channel dimension coordinates in the channel priority storage mode, determining the channel priority address information of the first pixel band;
[0019] When the memory coordinates of the first pixel band include channel dimension coordinates, width dimension coordinates, height dimension coordinates, quantity dimension coordinates and group index dimension coordinates in the channel folding storage mode, the channel folding address information of the first pixel band is determined based on the width dimension coordinates and the width dimension step, the height dimension coordinates and the height dimension step, the quantity dimension coordinates and the quantity dimension step, the group index dimension coordinates and the group index dimension step and the channel dimension coordinates in the channel folding storage mode.
[0020] In one embodiment, the first pixel band address information determines the second pixel band address information of other pixel bands in the target data block except the first pixel band based on the first pixel band address information, including:
[0021] When the first pixel band address information is channel priority address information, determining second pixel band address information of other pixel bands in the target data block except the first pixel band based on the channel priority address information and the width dimension coordinate change amount and width dimension step of the width dimension, the height dimension coordinate change amount and height dimension step of the height dimension, and the number dimension coordinate change amount and number dimension step of the pixel band coordinates of the first pixel band in the channel priority storage mode;
[0022] When the first pixel band address information is channel folding address information, based on the channel folding address information, and the width dimension coordinate change and width dimension step of the width dimension, the height dimension coordinate change and height dimension step of the height dimension, the quantity dimension coordinate change and quantity dimension step of the quantity dimension, the grouping index dimension coordinate change and grouping index dimension step of the grouping index dimension, and the channel dimension coordinate change of the pixel band coordinates of the first pixel band in the channel folding storage mode, determine the second pixel address information corresponding to other pixel bands in the target data block except the first pixel band.
[0023] In one embodiment, before determining the characteristic target matrix of the target data block according to the pixel data, the method further includes:
[0024] After loading the pixel data corresponding to the first pixel band address information, the second pixel band address information and the target data block, the pixel data of the pixel bands belonging to the same target sub-data block are assembled.
[0025] In one embodiment, loading the pixel data corresponding to the target data block according to the first pixel band address information and the second pixel band address information includes:
[0026] Selecting a height coordinate value of a height dimension coordinate of the first pixel band address information as a current height coordinate value, and selecting a width coordinate value of a width dimension coordinate of the first pixel band address information as a current width coordinate value;
[0027] Comparing the current height coordinate value with a first height filling threshold and a second height filling threshold, the first height filling threshold being less than the second height filling threshold;
[0028] When the current height coordinate value is less than the first height filling threshold or the current height coordinate value is not less than the second height filling threshold, obtaining filling data, using the filling data as pixel data, and updating the height coordinate value of the height dimension coordinate of the second pixel band address information to the current height coordinate value, and updating the width coordinate value of the width dimension coordinate of the second pixel band address information to the current width coordinate value, returning to the step of comparing the current height coordinate value with the first height filling threshold and the second height filling threshold and continuing to execute;
[0029] When the current height coordinate value is less than the first height filling threshold or the current height coordinate value is less than the second height filling threshold, comparing the current width coordinate value with the first width filling threshold and the second width filling threshold, the first width filling threshold being less than the second width filling threshold;
[0030] When the current width coordinate value is less than the first width fill threshold or the current width coordinate value is greater than or equal to the second width fill threshold, the filling data is obtained, the filling data is used as pixel data, and the height coordinate value of the height dimension coordinate of the second pixel band address information is updated to the current height coordinate value, and the width coordinate value of the width dimension coordinate of the second pixel band address information is updated to the current width coordinate value, and the step of comparing the current height coordinate value with the first height fill threshold and the second height fill threshold is returned and continued.
[0031] In one embodiment, loading the pixel data corresponding to the target data block according to the first pixel band address information and the second pixel band address information includes:
[0032] Compare the grouping index coordinate value of the grouping index dimension coordinate in the first pixel band address information and the second pixel band address information with the grouping threshold to determine whether the grouping index exceeds the number of divided groups;
[0033] If the group index exceeds the number of divided groups, the padding data is directly returned and the padding data is used as the pixel data corresponding to the target data block;
[0034] If the group index does not exceed the number of divided groups, determine whether there is pixel data in the memory based on the height coordinate value of the height dimension coordinate and the width coordinate value of the width dimension coordinate in the first pixel band address information and the second pixel band address information, and if there is no pixel data, obtain filling data and use the filling data as pixel data.
[0035] In a second aspect, the present application also provides an image data processing device. The device comprises:
[0036] A data acquisition module is used to obtain the location information of the target data block;
[0037] A coordinate determination module, configured to determine, according to the position information of the target data block, the pixel band coordinates of the preset data storage mode corresponding to the first pixel band in the target data block;
[0038] A first address determining module, configured to determine first pixel band address information of the first pixel band based on the pixel band coordinates of the first pixel band;
[0039] A second address determination module, configured to determine second pixel band address information of other pixel bands in the target data block except the first pixel band based on the first pixel band address information;
[0040] The matrix construction module is used to load the pixel data corresponding to the target data block according to the first pixel band address information and the second pixel band address information, and determine the characteristic target matrix of the target data block according to the pixel data.
[0041] In a third aspect, the present application further provides a computer device, wherein the computer device comprises a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of the above-mentioned image data processing method when executing the computer program.
[0042] In a fourth aspect, the present application further provides a computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the steps of the above-mentioned image data processing method are implemented.
[0043] In a fifth aspect, the present application further provides a computer program product, which includes a computer program, and when the computer program is executed by a processor, the steps of the above-mentioned image data processing method are implemented.
[0044] The above-mentioned image data processing method, device, computer equipment, storage medium and computer program product obtain the position information of the target data block; determine the pixel band coordinates of the preset data storage mode corresponding to the first pixel band in the target data block according to the position information of the target data block; determine the first pixel band address information of the first pixel band based on the pixel band coordinates of the first pixel band; determine the second pixel band address information of other pixel bands in the target data block except the first pixel band based on the first pixel band address information; load the pixel data corresponding to the target data block according to the first pixel band address information and the second pixel band address information, and determine the characteristic target matrix of the target data block according to the pixel data. According to the position information of the target data block, the pixel band coordinates corresponding to the target image data in the preset data storage mode can be calculated, and the pixel band address information of each pixel band can be further determined according to the pixel band coordinates, so as to load the pixel data according to the pixel band address information, so as to obtain the characteristic target matrix. In the convolution calculation process of the neural network model, there is no need to store the characteristic matrix in the memory, and there is no need to consider the convolution parameters too much. It is only necessary to load the pixel data according to the determined pixel address information to construct the characteristic target matrix, thereby improving the processing efficiency of the image data. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] Figure 1 A schematic diagram of the structure of an image data processing system of an image data processing method in one embodiment;
[0046] Figure 2 A schematic diagram of the structure of an image data processing system of an image data processing method in one embodiment;
[0047] Figure 3 is a schematic flow chart of an image data processing method in one embodiment;
[0048] Figure 4 A schematic diagram of a characteristic matrix structure in one embodiment;
[0049] Figure 5 A schematic diagram of a characteristic matrix structure in another embodiment;
[0050] Figure 6 is a schematic diagram of an input feature map in an embodiment;
[0051] Figure 7 A schematic diagram of a cache structure of an image data processing method in one embodiment;
[0052] Figure 8 is a schematic diagram of a cache structure of an image data processing method in another embodiment;
[0053] Fig. 9 is a schematic diagram of the structure of a target sub-data block in an embodiment;
[0054] Fig.10 A schematic diagram of a filling situation of an image data processing method in one embodiment;
[0055] Fig.11 is a schematic flow chart of an image data processing method in one embodiment;
[0056] Fig.12 is a schematic diagram of a filling situation of an image data processing method in another embodiment;
[0057] Fig.13 is a flowchart of an image data processing method in another embodiment;
[0058] Fig.14 is a flowchart of an image data processing method in another embodiment;
[0059] Fig.15 A schematic diagram of pixel band loading in one embodiment;
[0060] Fig.16 is a flowchart of an image data processing method in another embodiment;
[0061] Fig.17 is a flowchart of an image data processing method in another embodiment;
[0062] Fig.18 is a flowchart of an image data processing method in another embodiment;
[0063] Fig.19 A schematic diagram of pixel band loading in another embodiment;
[0064] Fig. 20 is a structural block diagram of an image data processing device in one embodiment;
[0065] Fig.21 FIG. 4 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION
[0066] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0067] The image data processing method provided in the embodiment of the present application can be applied to Figure 1 In the image data processing system shown. Figure 1The image data processing system shown may include modules such as an execution unit, a feature matrix construction unit, a cache, and a channel-priority video memory layout. The execution unit may include a matrix calculation unit, a texture sampling module, and a high-speed shared cache. The matrix calculation unit may be used to multiply the feature matrix and the weight matrix of im2col (from image to matrix). The feature matrix construction unit may include an address generator, a pixel band loader, and a pixel band assembler. Among them, during the neural network acceleration process, the matrix calculation unit can divide the feature matrix of the input feature map into multiple target data blocks, and send the location information of the target data block, such as FM (X, Y), to the feature matrix construction unit. After receiving the location information of the target data block, the feature matrix construction unit calculates the pixel band coordinates corresponding to the pixel band through the address generator, and calculates the corresponding pixel band address information. After receiving the pixel band address information, the pixel band loader loads the pixel data corresponding to the pixel band address information into the cache. If the cache hits, the pixel data will be returned to the pixel band assembler. Otherwise, the pixel data will be loaded from the channel-priority video memory layout, cached in the current cache and returned to the pixel band assembler for assembly. When assembling, the pixel band assembler will assemble according to the size of the target data block requested by the matrix calculation unit, and write it to the high-speed shared cache for the execution unit to calculate after the assembly is completed, so as to construct the feature target matrix.
[0068] In one embodiment, the image data processing method provided in the embodiment of the present application can also be applied to Figure 2 In the image data processing system shown. Figure 2The image data processing system shown may include modules such as an execution unit, a feature matrix construction unit, a cache, and a channel folded video memory layout, wherein the execution unit may include a matrix calculation unit, a texture sampling module, and a high-speed shared cache, etc. The matrix calculation unit may be used to perform multiplication operations on the feature matrix and weight matrix of im2col (from image to matrix). Among them, the feature matrix construction unit includes a grouping index calculator, an address generator, a pixel band loader and a pixel band assembler. In the acceleration process of the neural network model, the matrix calculation unit can divide the feature matrix of the input feature map into multiple target data blocks, and send the position information of the target data block, such as FM (X, Y), to the feature matrix construction unit. After receiving the position information of the target data block, the feature matrix construction unit can calculate the grouping index of the current target data block through the grouping index calculator, and generate the corresponding pixel band address information by combining the grouping index calculated by the grouping index calculator through the address generator. After receiving the pixel band address information, the pixel band loader loads the pixel data corresponding to the pixel band address information into the cache. If the cache hits, the pixel data will be returned to the pixel band assembler. Otherwise, the pixel data will be loaded from the channel folded video memory layout, cached in the current cache and returned to the pixel band assembler for assembly. When assembling, the pixel band assembler will assemble according to the size of the target data block requested by the matrix calculation unit, and write it to the high-speed shared cache for the execution unit to calculate after the assembly is completed, so as to construct the feature target matrix.
[0069] in, Figure 1 , Figure 2 The image data processing system in the system can be independently deployed in a terminal or in a server. It is also possible to deploy each module in the image data processing system in a computer device including both a terminal and a server according to actual needs, wherein the terminal communicates with the server through a network. The data storage system can store data that the server needs to process. The data storage system can be integrated on a server or placed on a cloud or other network server. Among them, the terminal can be, but is not limited to, various personal computers, laptops, smart phones, tablet computers, Internet of Things devices, and portable wearable devices. The Internet of Things devices can be smart speakers, smart TVs, smart air conditioners, smart car-mounted devices, etc. Portable wearable devices can be smart watches, smart bracelets, head-mounted devices, etc. The server can be implemented with an independent server or a server cluster consisting of multiple servers.
[0070] In one embodiment, Figure 3 As shown, a method for processing image data is provided, and the method is applied to a computer device as an example for explanation. The computer device is deployed with Figure 1 or Figure 2The image data processing system in the embodiment of the present invention, wherein the computer device may be a server or a terminal, comprises the following steps:
[0071] Step 302: Obtain location information of the target data block.
[0072] The target data block may be determined after dividing the feature matrix, the feature matrix may be obtained after the computer device expands the input feature map according to im2col, the computer device may divide the feature matrix to obtain a plurality of target data blocks, wherein the size of the target data block may be arbitrarily specified, for example, the size of the target data block may be 32x16, 32x32, 64x32, 64x64, etc. The position information of the target data block may refer to the position of the target data block in the feature matrix, and the position information may specifically be the coordinates of the target data block.
[0073] In one embodiment, Figure 4 As shown in the figure, when the input feature map is arranged in NHWC (channel priority) mode, the feature matrix is obtained after the input feature map is expanded according to im2col, where the columns of the feature matrix are kh*kw*C in , the rows of the feature matrix are N*H out *W out , where N is the batch size, (H out , W out ) is the width and height of the output feature map obtained by convolution of the input feature map and the convolution kernel, (kh, kw) is the width and height of the convolution kernel, C in To input the feature map and the number of channels corresponding to the convolution kernel, the computer device can divide the feature matrix to obtain the target data block. Figure 4 In the figure, the size of the target data block is a*b, where a is the width of the target data block and b is the height of the target data block.
[0074] In one embodiment, Figure 5 As shown in the figure, when the input feature map is arranged in N(C / x)HWx (channel folding), the feature matrix is obtained by expanding the input feature map according to im2col, where the columns of the feature matrix are (kh*kw*x)x(C in / x), the behavior of the feature matrix is N*Hout*Wout, where N is the batch size, (H out ,W out ) is the width and height of the output feature map obtained by convolution of the input feature map and the convolution kernel, (kh, kw) is the width and height of the convolution kernel, C in is the number of channels corresponding to the input feature map and the convolution kernel, relative to Figure 4The characteristic matrix in kh*kw*C in Direction: All channels C corresponding to a pixel in After loading, all channels of the next pixel are loaded. Specifically, the computer device can calculate the C of the pixel according to the channel folding number x. in The number of groups can be divided by C in / x is calculated, and for each group, x channel data of a pixel is loaded, and then x channel data of the next pixel is loaded, until the kh*kw*x data corresponding to the current group is loaded.
[0075] Step 304 : determining pixel band coordinates of a preset data storage mode corresponding to a first pixel band in the target data block according to the position information of the target data block.
[0076] Among them, the target data block can be composed of multiple pixels, each pixel can constitute a pixel band of the target data block, and the pixels in the same row in the target data block can correspond to each pixel band of the target data block, such as taking the pixel band composed of the pixels in the first row of the target data block as the first pixel band. Specifically, if the size of the target data block is a*b, it means that the target data block can be composed of b a*1 pixel bands, so the first pixel band can be the pixel band corresponding to the a*1 row in the target data block. The preset data storage mode can be the storage mode of the input feature map, that is, the memory layout of the input feature map in the memory. The preset data storage mode can be the NHWC (channel priority) storage mode, or it can be the N(C / x)HWx (channel folding) storage mode. The pixel band coordinates can refer to the coordinates of the first pixel band corresponding to the memory layout. Different data storage modes, correspondingly, the pixel band coordinates corresponding to the first pixel band will also be different.
[0077] In one embodiment, Figure 6 As shown in the figure, in the acceleration process of the neural network model, the corresponding batch size N = 1 and the number of channels C in =64, feature map height H = 5, feature map width W = 4, Figure 7 In order to use the NHWC method, after storing the input feature map, the memory layout of the input feature map in the memory is obtained. Figure 8 The memory layout of the input feature map obtained after the input feature map is stored in the N(C / x)HWx mode, wherein the value of the folded channel number X can be 32 when the N(C / x)HWx mode is adopted. Figure 7As shown, when storing the input feature map according to NHWC, first store the pixels at the corresponding positions along the channel C direction, such as c0:0, c1:20, ..., C63:1260; then move to the next pixel such as c0:1 along the width W direction, and repeat the first step; then, after the elements in the W direction are stored, move to the next pixel along the height H direction, such as after the W=0 pixel is stored, it will move to C0:4, that is, W=1; finally, after the HWC storage is completed, move along the batch N direction until the entire NxHxWxC pixels are stored. The entire NHWC memory layout can be regarded as a one-dimensional buffer. From Figure 8 It can be seen that when storing the input feature map according to N(C / x)HWx, C is first grouped according to the channel folding number x, such as Figure 8 C0~C31 is group0, C32~C63 is group1; for each group, storage is performed according to HW32, that is, along the W direction, 32 channels of each pixel are stored in order (C0~C31 or C32~C63 in this embodiment); then, after the W direction elements are stored, move to the next pixel along the high H direction, such as W=0 pixel storage, it will move to C0:4, that is, W=1; then, after HW32 storage is completed, move to the next group until the (C / 32)HW32 pixels of the current batch are stored; finally, along the N direction, store the pixel data of the next batch until all batches are stored. The entire N(C / x)HWx memory layout can be regarded as a one-dimensional buffer.
[0078] Step 306 : Determine first pixel band address information of the first pixel band based on the pixel band coordinates of the first pixel band.
[0079] The first pixel band address information refers to the address information of the first pixel band corresponding to the one-dimensional buffer. The one-dimensional buffer obtained by different data storage modes has different addresses corresponding to the one-dimensional buffer. After obtaining the first pixel band coordinates, the computer device can calculate the first pixel band address information according to the first pixel band coordinates.
[0080] Step 308: Determine second pixel band address information of other pixel bands in the target data block except the first pixel band based on the first pixel band address information.
[0081] Among them, the target data block is composed of multiple pixel points, and each pixel point can constitute a pixel band of the target data block, and the pixel points in the same row in the target data block can correspond to each pixel band of the target data block. Specifically, if the size of the target data block is a*b, it means that the target data block can be composed of b pixel bands of a*1 rows. Therefore, the first pixel band can be the pixel band of a*1 row in the target data block, and other pixel bands can refer to the pixel bands of other rows in the target data block except a*1 row. Other pixel bands all correspond to corresponding pixel band address information. Other pixel bands can be one or more, and the specific number of other pixel bands can be determined according to the actual division situation. In this embodiment, the address information of other pixel bands except the first pixel band is collectively referred to as the second pixel band address information.
[0082] Step 310: Load pixel data corresponding to the target data block according to the first pixel band address information and the second pixel band address information, and determine a characteristic target matrix of the target data block according to the pixel data.
[0083] Among them, pixel data can refer to data in a one-dimensional buffer, and the characteristic target matrix refers to a matrix constructed based on the pixel data. The computer device can load the pixel data corresponding to the target data block based on the first pixel band address information and the second pixel band address information, and determine the characteristic target matrix of the target data block based on the pixel data.
[0084] In the above-mentioned image data processing method, the position information of the target data block is obtained; according to the position information of the target data block, the pixel band coordinates of the preset data storage mode corresponding to the first pixel band in the target data block are determined; based on the pixel band coordinates of the first pixel band, the first pixel band address information of the first pixel band is determined; based on the first pixel band address information, the second pixel band address information of other pixel bands in the target data block except the first pixel band is determined; according to the first pixel band address information and the second pixel band address information, the pixel data corresponding to the target data block is loaded, and according to the pixel data, the characteristic target matrix of the target data block is determined. Among them, according to the position information of the target data block, the pixel band coordinates corresponding to the target image data in the preset data storage mode can be calculated, and the pixel band address information of each pixel band can be further determined according to the pixel band coordinates, so as to load the pixel data according to the pixel band address information, so as to obtain the characteristic target matrix. In the convolution calculation process of the neural network model, there is no need to store the characteristic matrix in the memory, and there is no need to consider the convolution parameters too much. It is only necessary to load the pixel data according to the determined pixel address information to construct the characteristic target matrix, thereby improving the processing efficiency of the image data.
[0085] In one embodiment, according to the position information of the target data block, the pixel band coordinates of the preset data storage mode corresponding to the first pixel band in the target data block are determined, including: dividing the target data block to obtain multiple target sub-data blocks; selecting any one of the target sub-data blocks as the current target sub-data block, and determining the first pixel band from the current target sub-data block; determining the pixel band position of the first pixel band according to the position information of the target sub-data block, and mapping the pixel band position to the preset data storage mode to obtain the pixel band coordinates of the preset data storage mode corresponding to the first pixel band in the target data block.
[0086] The computer device may further divide the target data block according to the actual processing situation of the matrix calculation unit to obtain multiple target sub-data blocks. The size of the target sub-data block is related to the number of pre-divided target sub-data blocks. When determining the size of the target sub-data block, the calculation may be carried out according to the following formula:
[0087] st w =a / sn w
[0088] st h =b / sn h
[0089] Among them, sn w is the number of target sub-data blocks divided along the width direction of the target data block, sn h is the number of target sub-data blocks divided along the height direction of the target data block, st w is the width of each target sub-data block, st h is the size corresponding to the height of each target sub-data block. Specifically, Fig. 9 As shown, the target sub-data block 0, the target sub-data block 1, the target sub-data block 2 and the target sub-data block 3 are involved. In this embodiment, the size of the target data block axb is set to 64x64, and the target sub-data block sn w xsn h =2x2, st w x st h =32x32 as an example for explanation.
[0090] Specifically, the computer device can select any one from the target sub-data blocks as the current target sub-data block, and determine the first pixel band from the current target sub-data block. According to the position information of the target data block, the position information of the target sub-data block can be determined accordingly, and the pixel band position of the first pixel band can be further determined. For example, if the target sub-data block 0 is selected as the current target sub-data block, the coordinate information of the target sub-data block 0 can be (X, Y), (X, Y+1), ..., (X, Y+n-1). The computer device can map (X, Y) to the NHWC memory layout or to the N(C / x)HWx memory layout, so as to obtain the pixel band coordinates of the preset data storage mode corresponding to the first pixel band in the target data block.
[0091] In one embodiment, in order to improve the loading channel C in Data efficiency, before storing the input feature map in the memory according to the preset data storage mode, the computer device can store the target sub-data block according to the width st w C in After alignment, (k h xk w )xC in It can just divide the width of the target sub-data block.
[0092] In this embodiment, the computer device further divides the target data block to obtain the target sub-data block, so that the target sub-data block can be divided into the target sub-data block according to its width. w C in Alignment to improve loading channel C in Data efficiency.
[0093] In one embodiment, the preset data storage mode includes a channel priority storage mode or a channel folding storage mode; mapping the pixel band position to the preset data storage mode to obtain the pixel band coordinates of the preset data storage mode corresponding to the first pixel band in the target data block, including: if the preset data storage mode is the channel priority storage mode, mapping the pixel band coordinates to the channel priority storage mode to obtain the channel dimension coordinates, width dimension coordinates, height dimension coordinates and quantity dimension coordinates of the first pixel band in the channel priority storage mode; if the preset data storage mode is the channel folding storage mode, mapping the pixel band coordinates to the channel folding storage mode to obtain the channel dimension coordinates, width dimension coordinates, height dimension coordinates, quantity dimension coordinates and grouping index dimension coordinates of the first pixel band in the channel folding storage mode.
[0094] The pixel band coordinates of the first pixel band in the target sub-data block need to be mapped to the preset data storage mode according to the pixel band position of the pixel band, such as (X, Y), (X, Y+1), ..., (X, Y+n-1). When the preset data storage mode is the NHWC memory layout, due to the C in According to the width w Therefore, the computer device only needs to map the first pixel band to the corresponding channel dimension coordinates, width dimension coordinates, height dimension coordinates and quantity dimension coordinates in the NHWC memory layout according to the pixel band position, that is, the channel dimension coordinates, width dimension coordinates, height dimension coordinates and quantity dimension coordinates are used as pixel band coordinates. When the preset data storage mode is N(C / x)HWx memory layout, since C is also in According to the width tw Alignment is performed, so the computer device only needs to map the first pixel band to the corresponding channel dimension coordinates, width dimension coordinates, height dimension coordinates, quantity dimension coordinates and grouping index dimension coordinates in the N(C / x)HWx memory layout according to the pixel band position, that is, the channel dimension coordinates, width dimension coordinates, height dimension coordinates, quantity dimension coordinates and grouping index dimension coordinates are used as pixel band coordinates.
[0095] In this embodiment, when the computer device determines the pixel band coordinates corresponding to the first pixel band, due to the in According to the width w After alignment, the pixel band coordinates can be obtained by simply mapping the pixel band position of the first pixel band to the corresponding preset data storage mode, thereby improving the efficiency of image data processing.
[0096] In one of the embodiments, first pixel band address information of the first pixel band is determined based on the pixel band coordinates of the first pixel band, including: when the memory coordinates of the first pixel band include the channel dimension coordinates, width dimension coordinates, height dimension coordinates and quantity dimension coordinates in the channel priority storage mode, the channel priority address information of the first pixel band is determined based on the width dimension coordinates and the width dimension step, the height dimension coordinates and the height dimension step, the quantity dimension coordinates and the quantity dimension step, and the channel dimension coordinates in the channel priority storage mode; when the memory coordinates of the first pixel band include the channel dimension coordinates, width dimension coordinates, height dimension coordinates, quantity dimension coordinates and the group index dimension coordinates in the channel folding storage mode, the channel folding address information of the first pixel band is determined based on the width dimension coordinates and the width dimension step, the height dimension coordinates and the height dimension step, the quantity dimension coordinates and the quantity dimension step, the group index dimension coordinates and the group index dimension step, and the channel dimension coordinates in the channel folding storage mode.
[0097] When the memory coordinates of the first pixel band include the channel dimension coordinates, width dimension coordinates, height dimension coordinates, and quantity dimension coordinates in the channel priority storage mode, there are corresponding dimension strides. Specifically, for the channel priority storage mode, the corresponding dimension stride is calculated as follows: Batch1 , stride H1 , stride W1 )=(H in *W in *C in , W in *C in , C in ), these dimension strides are fixed constants for each convolutional layer, where stride Batch1 Refers to the number dimension step in the channel priority storage mode, stride H1 refers to the height dimension step in the channel priority storage mode, stride W1 Refers to the height dimension step in channel priority storage mode, H in is the height corresponding to the input feature map, W in is the width corresponding to the input feature map.
[0098] When the memory coordinates of the first pixel band include the channel dimension coordinates, width dimension coordinates, height dimension coordinates, quantity dimension coordinates, and group index dimension coordinates in the channel folding storage mode, there are corresponding dimension strides. Specifically, for the channel folding storage mode, the corresponding dimension stride is calculated as follows: Batch2 , stride C , stride H2 , stride W2 )=((C in / x)*H in *W in *x,H in *W in *x,W in *x, x), which are fixed constants for each convolutional layer, stride Batch2 refers to the number dimension step in the channel folding storage mode, stride H2 refers to the height dimension step in the channel folding storage mode, stride W2 refers to the height dimension step in the channel folding storage mode, stride C Refers to the grouping index dimension step in channel folding mode.
[0099] In one embodiment, for the channel priority storage mode, the computer device may use the following formula to calculate the first pixel band address information:
[0100] strip 1D addr1 =batch idx1 *stride Batch1 +h idx1 *stride H1 +w idx1 *stride W1 +c_start idx1
[0101] Among them, strip 1D addr1 Indicates the first pixel with address information, batch idx1 is the number dimension coordinate, stride Batch1 is the step size of the quantity dimension, h idx1 is the height dimension coordinate, stride H1 is the height dimension step, w idx1 is the width dimension coordinate, stride W1 is the width dimension step, c_start idx1 It is the channel dimension coordinate, that is, the starting address of the channel.
[0102] In one embodiment, for the channel folding storage mode, the computer device may use the following formula to calculate the first pixel band address information:
[0103] strip 1D addr2 =batch idx2 *stride Batch2 +group idx *stride c +h idx2 *stride H2 +w idx2 *stride W2 +c_start idx2
[0104] Among them, strip 1D addr2 Indicates the first pixel with address information, batch idx2 is the number dimension coordinate, stride Batch2 is the step size of the quantity dimension, group idx The stride is the grouping index dimension coordinate. C h is the step length of the grouping index dimension, idx2 is the height dimension coordinate, stride H2 is the height dimension step, w idx2 is the width dimension coordinate, stride W2 is the width dimension step, c_start idx2It is the channel dimension coordinate, that is, the starting address of the channel.
[0105] In this embodiment, the computer device determines the corresponding dimensional steps for different storage modes, and can accurately calculate the first pixel band address information based on the dimensional coordinates and the dimensional steps.
[0106] In one of the embodiments, the first pixel band address information determines the second pixel band address information of other pixel bands in the target data block except the first pixel band based on the first pixel band address information. The first pixel band address information includes: when the first pixel band address information is the channel priority address information, based on the channel priority address information, and the width dimension coordinate change amount and width dimension step of the width dimension, the height dimension coordinate change amount and height dimension step of the height dimension, and the quantity dimension coordinate change amount and quantity dimension step of the pixel band coordinate of the first pixel band in the channel priority storage mode, determining the second pixel band address information of other pixel bands in the target data block except the first pixel band second pixel band address information of other pixel bands other than the first pixel band in the target data block; when the first pixel band address information is channel folding address information, based on the channel folding address information, and the width dimension coordinate change and the width dimension step of the width dimension, the height dimension coordinate change and the height dimension step of the height dimension, the quantity dimension coordinate change and the quantity dimension step of the quantity dimension, the grouping index dimension coordinate change and the grouping index dimension step of the grouping index dimension, and the channel dimension coordinate change of the pixel band coordinates of the first pixel band in the channel folding storage mode, determine the second pixel address information corresponding to other pixel bands other than the first pixel band in the target data block.
[0107] Among them, after determining the first pixel band address information, the computer device does not need to repeatedly determine the pixel band coordinates of other pixel bands in the same current target sub-data block, that is, there is no need to recalculate the pixel band address based on the pixel band coordinates and dimension step, etc., and only needs to use the first pixel band address of the current target sub-data block to deduce the second pixel band address of other pixel bands in the current target sub-data block.
[0108] Specifically, for the channel folding address information, the computer device may use the following formula to expand when determining the second pixel band address:
[0109]
[0110] Wherein, curstrip 1D Addr indicates the address information of the second pixel strip, 1 st strip 1D addr2 Indicates channel folding address information, Δgroup idx Indicates the change in the coordinates of the grouping index dimension, Δbatch idx Indicates the change in the coordinates of the quantity dimension, Δh idxIndicates the change in height dimension coordinates, Δw idx Indicates the change in width dimension coordinates, In this embodiment, for the same target sub-data block, the coordinates of the grouping index dimension in its coordinates remain unchanged. In addition, in this embodiment, the batch size Batch Size N = 1, so batch idx2 When the computer device stores the input feature map, the number of channels C in The number of channels in the group is aligned, and the width of the target sub-data block is st w In this example, it is equal to x, so Therefore, for the second pixel band address information of the same target sub-data block, only the address information of other pixel bands (h idx , w idx ) to update, the specific calculation formula is as follows:
[0111] h idx =Q2*stride h +Q4
[0112] w idx =Rem2*stride w +Rem4 / 32
[0113] Q 4 =(X-group idx *kh*kw*x) / (kw*x)
[0114] Rem 4 =(X-groupid x *kh*kw*x)%(kw*x)
[0115] Q 2 =(Y%(H out *Wou t )) / Wou t
[0116] Rem 2 =(Y%(H out *W out ))%W out
[0117] Among them, Q 2 Represents the quotient 2, Rem 2 Represents remainder 2, Q 4 Represents quotient 4, Rem 4Representing the remainder 4, the computer device can perform a series of division operations based on the Y-dimensional coordinates and X-dimensional coordinates of the target sub-data block, combined with the width and height corresponding to the convolution output feature map, i.e., Hout and Wout, the width and height corresponding to the convolution kernel, i.e., kh and kw, and the channel folding number x, to obtain the quotient and the remainder. Furthermore, after calculating the height dimension coordinates and the width dimension coordinates, the computer device can calculate the change in the height dimension coordinates and the change in the width dimension coordinates according to the following formula:
[0118] Δh idx =cur_strip Q2 -pre_strip Q2
[0119] Δw idx =cur_strip Rem2 -pre_strip Rem2
[0120] Among them, cur_strip Q 2 is the Q 2 corresponding to the current pixel strip, pre_strip Q 2 is Q2 corresponding to the previous pixel strip, cur_strip Rem2 Rem2 corresponding to the current pixel band, pre_strip Rem2 is the Rem2 corresponding to the previous pixel band, that is, (Δh idx , Δw idx ) is the difference between (Q2, Rem2) corresponding to the current strip and the previous strip.
[0121] Among them, for the channel priority storage mode, the computer device determines the second pixel band address information of other pixel bands in the target data block except the first pixel band based on the width dimension coordinate change and width dimension step of the width dimension, the height dimension coordinate change and height dimension step of the height dimension, and the quantity dimension coordinate change and quantity dimension step of the quantity dimension. When determining the second pixel band address information, it can also adopt the calculation method of calculating the second pixel address information under the channel folding storage mode, which will not be repeated here.
[0122] In one embodiment, after the pixel band address information of other pixel bands of the same target sub-data block is calculated, the computer device records the pixel band address information corresponding to the last pixel band when the current target sub-data block is loaded, Q 2 、Rem 2 Etc., for use by other target sub-data blocks when loading pixel bands and calculating pixel band address information.
[0123] In this embodiment, the computer device can efficiently and quickly derive the second pixel band address information in each target sub-data block based on the address of the first pixel band in the target sub-data block without additionally calculating the pixel band coordinates of other pixel bands.
[0124] In one of the embodiments, before determining the characteristic target matrix of the target data block based on the pixel data, it also includes: after loading the pixel data corresponding to the first pixel band address information, the second pixel band address information and the target data block, assembling the pixel data of the pixel bands belonging to the same target sub-data block.
[0125] After receiving the pixel band address information, the computer device will load the pixel band address information into the cache. If the cache hits, the pixel data will be returned to the pixel band assembler, and the pixel band assembler in the computer device will assemble the data belonging to the same target sub-data block.
[0126] In this embodiment, the computer device assembles data belonging to the same target sub-data block so that the final data is the same size as the requested characteristic target matrix, thereby improving the accuracy of the target characteristic matrix construction.
[0127] In one of the embodiments, pixel data corresponding to the target data block is loaded according to the first pixel band address information and the second pixel band address information, including: selecting the height coordinate value of the height dimension coordinate of the first pixel band address information as the current height coordinate value, and selecting the width coordinate value of the width dimension coordinate of the first pixel band address information as the current width coordinate value; comparing the current height coordinate value with the first height fill threshold and the second height fill threshold, the first height fill threshold being less than the second height fill threshold; when the current height coordinate value is less than the first height fill threshold or the current height coordinate value is not less than the second height fill threshold, obtaining filling data, using the filling data as pixel data, and updating the height coordinate value of the height dimension coordinate of the second pixel band address information to the current height coordinate value, and updating the width coordinate value of the width dimension coordinate of the second pixel band address information to the current width coordinate value, and returning The step of comparing the current height coordinate value with the first height filling threshold and the second height filling threshold is continued; when the current height coordinate value is less than the first height filling threshold or the current height coordinate value is less than the second height filling threshold, the current width coordinate value is compared with the first width filling threshold and the second width filling threshold, and the first width filling threshold is less than the second width filling threshold; when the current width coordinate value is less than the first width filling threshold or the current width coordinate value is greater than or equal to the second width filling threshold, the filling data is obtained, the filling data is used as pixel data, and the height coordinate value of the height dimension coordinate of the second pixel band address information is updated to the current height coordinate value, and the width coordinate value of the width dimension coordinate of the second pixel band address information is updated to the current width coordinate value, and the step of comparing the current height coordinate value with the first height filling threshold and the second height filling threshold is returned to continue.
[0128] In order to further improve the efficiency of loading each pixel band, the computer device may not store the padding data in the memory if there is padding data in the convolution layer parameters, so as to achieve the effect of reducing the storage space. In the case where the preset data storage mode is the channel priority storage mode, the first height padding threshold and the second height padding threshold can be used to characterize the padding situation in the height direction, and the first width padding threshold and the second width padding threshold can be used to characterize the padding situation in the width direction. When the computer device obtains the pixel data of the pixel band each time, it can determine whether the required pixel data exists in the memory according to the height coordinate value and the width coordinate value, and after the pixel data corresponding to each pixel band is obtained, it will further obtain the pixel data corresponding to the next pixel band. Specifically, for the same target sub-data block, the pixel data corresponding to each pixel band can be obtained in sequence, and after the pixel data of the same target sub-data block is obtained, the pixel data corresponding to the pixel bands of other target sub-data blocks are obtained.
[0129] In one embodiment, Fig.10 As shown in the figure, it is a schematic diagram of the filling situation in the corresponding NHWC memory layout. In the one-dimensional buffer of the memory, the filled data is not stored, which reduces the storage space. The computer device stores the corresponding (h idx , w idx ), namely, the height dimension coordinate and the width dimension coordinate, specifically, the height coordinate value corresponding to the height dimension coordinate is compared with the first height fill threshold and the second height fill threshold, and the width coordinate value is compared with the first width fill threshold and the second width fill threshold to load pixel data.
[0130] In one embodiment, if Fig.11 As shown, it is a flow chart for loading pixel data. When the computer device loads the pixel band data corresponding to each pixel band, it can select the height coordinate value of the height dimension coordinate of the first pixel band address information as the current height coordinate value, and select the width coordinate value of the width dimension coordinate of the first pixel band address information as the current width coordinate value, and compare the current height coordinate value with the first height filling threshold and the second height filling threshold. When the previous height coordinate value is less than the first height filling threshold or the current height coordinate value is not less than the second height filling threshold, it means that the height direction coordinate meets the condition, then it can be directly judged that the required data does not exist in the one-dimensional buffer. At this time, The filled number can be directly returned, such as directly returning all-0 data for the pixel band. If the condition in the height direction is not met, the computer device further compares the current width coordinate value with the first width filling threshold and the second width filling threshold. When the previous width coordinate value is less than the first width filling threshold or the current width coordinate value is greater than or equal to the second width filling threshold, it means that the width direction coordinate meets the condition. It can be directly determined that the required data does not exist in the one-dimensional buffer. At this time, the filled number can be directly returned, such as directly returning all-0 data for the pixel band. Otherwise, it means that the required data exists in the one-dimensional buffer, and the computer device loads the pixel data normally.
[0131] In this embodiment, by not storing the padding data, the storage space is reduced. When the computer device loads the pixel data, it only needs to make a judgment based on the height coordinate value of the height dimension coordinate and the width coordinate value of the width dimension coordinate to efficiently complete the pixel data loading.
[0132] In one of the embodiments, pixel data corresponding to the target data block is loaded according to the first pixel band address information and the second pixel band address information, including: comparing the grouping index coordinate value of the grouping index dimension coordinate in the first pixel band address information and the second pixel band address information with the grouping threshold to determine whether the grouping index exceeds the number of divided groups; if the grouping index exceeds the number of divided groups, directly returning the filling data, and using the filling data as the pixel data corresponding to the target data block; if the grouping index does not exceed the number of divided groups, based on the height coordinate value of the height dimension coordinate and the width coordinate value of the width dimension coordinate in the first pixel band address information and the second pixel band address information, determining whether there is pixel data in the memory, and if there is no pixel data, obtaining the filling data, and using the filling data as the pixel data.
[0133] In order to further improve the efficiency of loading each pixel band, the computer device may not store the padding data in the memory if there is padding data in the convolution layer parameters, so as to achieve the effect of reducing the storage space. In the case where the preset data storage mode is the channel folding storage mode, the computer device can determine whether the required pixel data exists in the memory according to the grouping index coordinate value of the grouping index dimension coordinate, the height coordinate value of the height dimension coordinate, and the width coordinate value of the width dimension coordinate each time the pixel data of the pixel band is obtained, and after the pixel data corresponding to each pixel band is obtained, the pixel data corresponding to the next pixel band will be further obtained. Specifically, for the same target sub-data block, the pixel data corresponding to each pixel band can be obtained in sequence, and after the pixel data of the same target sub-data block is obtained, the pixel data corresponding to the pixel bands of other target sub-data blocks can be obtained.
[0134] In one embodiment, Fig.12 As shown, it is a schematic diagram of the filling situation in the corresponding N(C / x)HWx memory layout. In the one-dimensional buffer of the memory, the filled data is not stored, which reduces the storage space. The computer device determines whether there is the required pixel data in the memory through the pixel band coordinates corresponding to each pixel band involved in the pixel address information, that is, the group index dimension coordinates, the height dimension coordinates, and the width dimension coordinates. If not, the filled data is directly returned as the pixel data.
[0135] In one embodiment, if Fig.13As shown, it is a flow chart for loading pixel data. When loading the pixel band data corresponding to each pixel band, the computer device can first compare the grouping index coordinate value with the grouping threshold to determine whether the grouping index exceeds the number of divided groups. If the grouping index exceeds the number of divided groups, the required pixel data does not exist in the current target sub-data block as a whole, and the pixel band data of all 0s can be directly returned. If the grouping index does not exceed the number of divided groups, the computer device can directly determine whether the required pixel data exists in the N(C / x)HWx memory layout based on the height coordinate value of the height dimension coordinate and the width coordinate value of the width dimension coordinate. If it does not exist, the filling data can be directly returned, such as directly returning all 0 data. Otherwise, it means that the required data exists in the one-dimensional buffer, and the computer device loads the pixel data normally.
[0136] In this embodiment, by not storing the padding data, the storage space is reduced. When the computer device loads the pixel data, it only needs to make a judgment based on the grouping index coordinate value of the grouping index dimension coordinate, the height coordinate value of the height dimension coordinate, and the width coordinate value of the width dimension coordinate, so as to efficiently complete the pixel data loading.
[0137] In one embodiment, Fig.13 As shown, it is a flowchart when the preset data storage mode is the channel priority storage mode:
[0138] Step 1: Given a target data block, the computer device can divide the target data block into target sub-data block 0, target sub-data block 1, target sub-data block 2 and target sub-data block 3. Target sub-data block 0, target sub-data block 1, target sub-data block 2 and target sub-data block 3 can be loaded in parallel. In this embodiment, target sub-data block 0 and target sub-data block 1 are loaded in parallel. AG (address generator) can calculate the pixel band coordinates corresponding to the pixel band. For the channel priority storage mode, the pixel band coordinates are the 4-dimensional left side, and calculate the corresponding pixel band address;
[0139] Step 2: According to step 1, SL (pixel band loader) loads the pixel data corresponding to the pixel band, such as Fig.15 As shown, it is a schematic diagram of a computer device loading n m*1 pixels from the original NHWC memory layout. In this embodiment, for the target sub-data block 0, the computer device loads the pixel band from the original memory layout, thereby implicitly forming a characteristic target matrix. For the pixel bands of the same target sub-data block, they are different in the memory layout, and the pixel band addresses of the pixel bands are required.
[0140] Step 3: SA assembles the data returned by SL;
[0141] Step 4: Repeat steps 1 to 3 to load the remaining target sub-data blocks.
[0142] When loading the pixel bands of the target sub-data block 0 and the target sub-data block 1, it is only necessary to calculate the first pixel band address information, Q2 and Rem2 calculated based on the first pixel band, without recalculating the subtile 0 The 4D coordinates of the remaining pixel bands in . Fig.16 FIG. 1 is a schematic diagram of a process for updating pixel band address information of other pixel bands in target sub-data block 0 and target sub-data block 1. Fig.16 In (Q 2 , Rem 2 ) is the intermediate result generated when AG calculates the address information of the first pixel band. The specific calculation formula is as follows: out Is the width corresponding to the output feature map:
[0143] Q 2 =(Y%(H out *W out )) / W out
[0144] Rem 2 =(Y%(H out *W out ))%W out
[0145] When loading the target sub-data block 2 and the target sub-data block 3, it is necessary to update Q3, i.e., quotient 3; Rem3, remainder 3; Q4, quotient 4; and Rem4, remainder 4, such as Fig.17 As shown, (Q3, Rem3) and (Q4, Rem4) are the intermediate results of AG when calculating the pixel band address:
[0146] Q 3 =X / (kw*C in )
[0147] Rem 3 =X%(kw*C in )
[0148] Q 4 =(X%(kw*C in )) / C in
[0149] Rem 4 =(X%(kw*C in ))%C in
[0150] In one embodiment, Fig.18As shown, it is a flowchart when the preset data storage mode is the channel folding storage mode:
[0151] Step 1: Given a target data block ∈[kh*kw*C in , N*H out *W out ], the computer device can divide the target data block into target sub-data block 0, target sub-data block 1, target sub-data block 2 and target sub-data block 3, and the GIC (group index calculator) can calculate the group index group corresponding to the target sub-data block 0 and the target sub-data block 1 under N(C / x)HWx in parallel idx ;
[0152] Step 2: Combine the group from step 1 idx AG calculates the 5-dimensional index coordinates of N(C / x)HWx corresponding to the first pixel in the target sub-data block 0, and calculates the corresponding pixel band address;
[0153] Step 3: According to step 2, SL Load the pixel strip, such as Fig.19 As shown, it is a schematic diagram showing that the target sub-data block 0mxn loads n mx1 pixel bands from the original N(C / x)HWx memory layout, thereby implicitly forming the target sub-data block, that is, the characteristic target matrix. For the pixel bands of the same sub-target data block, their positions in the N(C / x)HWx memory layout are different, and the pixel band addresses of the pixel bands are required.
[0154] Step 4: SA assembles the data returned by SL;
[0155] Step 5: Repeat steps 2 to 4 to load the remaining target sub-data blocks 2 and 3.
[0156] It should be understood that, although the various steps in the flowcharts involved in the above-mentioned embodiments are displayed in sequence according to the indication of the arrows, these steps are not necessarily executed in sequence according to the order indicated by the arrows. Unless there is a clear explanation in this article, the execution of these steps does not have a strict order restriction, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-mentioned embodiments can include multiple steps or multiple stages, and these steps or stages are not necessarily executed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a part of the steps or stages in other steps.
[0157] Based on the same inventive concept, the embodiment of the present application also provides an image data processing device for implementing the image data processing method involved above. The implementation solution provided by the device to solve the problem is similar to the implementation solution recorded in the above method, so the specific limitations in one or more image data processing method device embodiments provided below can refer to the above limitations on the image data processing method, and will not be repeated here.
[0158] In one embodiment, Fig. 20 As shown, an image data processing method and device 2000 is provided, comprising: a data acquisition module, a coordinate determination module, a first address determination module, a second address determination module and a matrix construction module, wherein:
[0159] The data acquisition module 2002 is used to acquire the location information of the target data block.
[0160] The coordinate determination module 2004 is used to determine the pixel band coordinates of the preset data storage mode corresponding to the first pixel band in the target data block according to the position information of the target data block.
[0161] The first address determining module 2006 is configured to determine first pixel band address information of the first pixel band based on the pixel band coordinates of the first pixel band.
[0162] The second address determining module 2008 is configured to determine second pixel band address information of other pixel bands except the first pixel band in the target data block based on the first pixel band address information.
[0163] The matrix construction module 2010 is used to load the pixel data corresponding to the target data block according to the first pixel band address information and the second pixel band address information, and determine the characteristic target matrix of the target data block according to the pixel data.
[0164] In one of the embodiments, the coordinate determination module is also used to divide the target data block to obtain multiple target sub-data blocks; select any one of the target sub-data blocks as the current target sub-data block, and determine the first pixel band from the current target sub-data block; determine the pixel band position of the first pixel band based on the position information of the target sub-data block, and map the pixel band position to the preset data storage mode to obtain the pixel band coordinates of the preset data storage mode corresponding to the first pixel band in the target data block.
[0165] In one of the embodiments, the coordinate determination module is also used to map the pixel band position to the channel priority storage mode if the preset data storage mode is the channel priority storage mode, and obtain the channel dimension coordinates, width dimension coordinates, height dimension coordinates and quantity dimension coordinates of the first pixel band in the channel priority storage mode; if the preset data storage mode is the channel folding storage mode, map the pixel band position to the channel folding storage mode, and obtain the channel dimension coordinates, width dimension coordinates, height dimension coordinates, quantity dimension coordinates and grouping index dimension coordinates of the first pixel band in the channel folding storage mode.
[0166] In one of the embodiments, the first address determination module is further used to determine the channel priority address information of the first pixel band based on the width dimension coordinate and the width dimension step, the height dimension coordinate and the height dimension step, the quantity dimension coordinate and the quantity dimension step, and the channel dimension coordinate in the channel priority storage mode when the memory coordinates of the first pixel band include the channel dimension coordinates, width dimension coordinates, height dimension coordinates, and the quantity dimension coordinates in the channel priority storage mode; and to determine the channel folding address information of the first pixel band based on the width dimension coordinate and the width dimension step, the height dimension coordinate and the height dimension step, the quantity dimension coordinate and the quantity dimension step, the grouping index dimension coordinate and the grouping index dimension step, and the channel dimension coordinate in the channel folding storage mode when the memory coordinates of the first pixel band include the channel dimension coordinates, width dimension coordinates, height dimension coordinates, the quantity dimension coordinates, and the grouping index dimension coordinates in the channel folding storage mode.
[0167] In one of the embodiments, the second address determination module is further used to determine the second pixel band address information of other pixel bands in the target data block except the first pixel band when the first pixel band address information is channel priority address information based on the channel priority address information, and the width dimension coordinate change and width dimension step of the width dimension, the height dimension coordinate change and height dimension step of the height dimension, and the quantity dimension coordinate change and quantity dimension step of the quantity dimension of the pixel band coordinates of the first pixel band in the channel priority storage mode; when the first pixel band address information is channel folding address information, determine the second pixel address information corresponding to other pixel bands in the target data block except the first pixel band based on the channel folding address information, and the width dimension coordinate change and width dimension step of the width dimension, the height dimension coordinate change and height dimension step of the height dimension, the quantity dimension coordinate change and quantity dimension step of the quantity dimension, the grouping index dimension coordinate change and grouping index dimension step of the grouping index dimension, and the channel dimension coordinate change of the pixel band coordinates of the first pixel band in the channel folding storage mode.
[0168] In one of the embodiments, the apparatus further comprises an assembly module;
[0169] The assembling module is used to assemble the pixel data of the pixel bands belonging to the same target sub-data block after loading the pixel data corresponding to the first pixel band address information, the second pixel band address information and the target data block.
[0170] In one of the embodiments, the matrix construction module is further used to select the height coordinate value of the height dimension coordinate of the first pixel band address information as the current height coordinate value, and select the width coordinate value of the width dimension coordinate of the first pixel band address information as the current width coordinate value; compare the current height coordinate value with the first height filling threshold and the second height filling threshold, and the first height filling threshold is less than the second height filling threshold; when the current height coordinate value is less than the first height filling threshold or the current height coordinate value is not less than the second height filling threshold, obtain filling data, use the filling data as pixel data, and update the height coordinate value of the height dimension coordinate of the second pixel band address information to the current height coordinate value, and update the width coordinate value of the width dimension coordinate of the second pixel band address information to the current width coordinate value, return to the step of comparing the current height coordinate value with the first height filling threshold and the second height filling threshold and continue to execute; when the current height coordinate value is less than the first height filling threshold or the current height coordinate value is less than the second height filling threshold, compare the current width coordinate value with the first width filling threshold and the second width filling threshold, and the first width filling threshold is less than the second width filling threshold;
[0171] When the current width coordinate value is less than the first width fill threshold or the current width coordinate value is greater than or equal to the second width fill threshold, the filling data is obtained, the filling data is used as pixel data, and the height coordinate value of the height dimension coordinate of the second pixel band address information is updated to the current height coordinate value, and the width coordinate value of the width dimension coordinate of the second pixel band address information is updated to the current width coordinate value, and the step of comparing the current height coordinate value with the first height fill threshold and the second height fill threshold is returned and continued.
[0172] In one of the embodiments, the matrix construction module is also used to compare the grouping index coordinate value of the grouping index dimension coordinate in the first pixel band address information and the second pixel band address information with the grouping threshold to determine whether the grouping index exceeds the number of divided groups; if the grouping index exceeds the number of divided groups, the filling data is directly returned and the filling data is used as the pixel data corresponding to the target data block; if the grouping index does not exceed the number of divided groups, based on the height coordinate value of the height dimension coordinate and the width coordinate value of the width dimension coordinate in the first pixel band address information and the second pixel band address information, it is determined whether there is pixel data in the memory, and when there is no pixel data, the filling data is obtained and the filling data is used as the pixel data.
[0173] Each module in the above-mentioned image data processing device can be implemented in whole or in part by software, hardware or a combination thereof. Each module can be embedded in or independent of a processor in a computer device in the form of hardware, or can be stored in a memory in a computer device in the form of software, so that the processor can call and execute the operations corresponding to each module.
[0174] In one embodiment, a computer device is provided. The computer device may be a server or a terminal. The internal structure diagram thereof may be as follows: Fig.21 As shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, referred to as I / O) and a communication interface. The processor, the memory and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store image data. The input / output interface of the computer device is used to exchange information between the processor and an external device. The communication interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, an image data processing method is implemented.
[0175] Those skilled in the art will understand that Fig.21 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.
[0176] In one embodiment, a computer device is provided, including a memory and a processor, wherein a computer program is stored in the memory, and the processor implements the steps of the above-mentioned image data processing method when executing the computer program.
[0177] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps of the above-mentioned image data processing method are implemented.
[0178] In one embodiment, a computer program product is provided, comprising a computer program, which implements the steps of the above-mentioned image data processing method when executed by a processor.
[0179] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant laws, regulations and standards of relevant countries and regions.
[0180] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to the memory, database or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. As an illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The database involved in each embodiment provided in this application may include at least one of a relational database and a non-relational database. Non-relational databases may include distributed databases based on blockchains, etc., but are not limited to this. The processor involved in each embodiment provided in this application may be a general-purpose processor, a central processing unit, a graphics processor, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, etc., but are not limited to this.
[0181] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0182] The above-described embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the present application. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the attached claims.
Claims
1. A method for processing image data, characterized in that: The method comprises: Get the location information of the target data block; Determining pixel band coordinates of a preset data storage mode corresponding to a first pixel band in the target data block according to the position information of the target data block; determining first pixel band address information of the first pixel band based on the pixel band coordinates of the first pixel band; Based on the first pixel band address information, determining second pixel band address information of other pixel bands in the target data block except the first pixel band; According to the first pixel band address information and the second pixel band address information, pixel data corresponding to the target data block is loaded, and according to the pixel data, a characteristic target matrix of the target data block is determined.
2. The method according to claim 1, characterized in that: The step of determining the pixel band coordinates of the preset data storage mode corresponding to the first pixel band in the target data block according to the position information of the target data block comprises: Dividing the target data block to obtain multiple target sub-data blocks; Selecting any one of the target sub-data blocks as a current target sub-data block, and determining a first pixel band from the current target sub-data block; According to the position information of the target sub-data block, the pixel band position of the first pixel band is determined, and the pixel band position is mapped to a preset data storage mode to obtain the pixel band coordinates of the preset data storage mode corresponding to the first pixel band in the target data block.
3. The method according to claim 2, characterized in that The preset data storage mode includes a channel priority storage mode or a channel folding storage mode; the mapping of the pixel band position to the preset data storage mode to obtain the pixel band coordinates of the preset data storage mode corresponding to the first pixel band in the target data block includes: If the preset data storage mode is a channel priority storage mode, mapping the pixel band position to the channel priority storage mode to obtain the channel dimension coordinates, width dimension coordinates, height dimension coordinates and quantity dimension coordinates of the first pixel band in the channel priority storage mode; If the preset data storage mode is a channel folding storage mode, the pixel band position is mapped to the channel folding storage mode to obtain the channel dimension coordinates, width dimension coordinates, height dimension coordinates, quantity dimension coordinates and grouping index dimension coordinates of the first pixel band in the channel folding storage mode.
4. The method according to claim 1, characterized in that: The determining first pixel band address information of the first pixel band based on the pixel band coordinates of the first pixel band includes: When the memory coordinates of the first pixel band include channel dimension coordinates, width dimension coordinates, height dimension coordinates, and quantity dimension coordinates in the channel priority storage mode, based on the width dimension coordinates and the width dimension step, the height dimension coordinates and the height dimension step, the quantity dimension coordinates and the quantity dimension step, and the channel dimension coordinates in the channel priority storage mode, determining the channel priority address information of the first pixel band; When the memory coordinates of the first pixel band include channel dimension coordinates, width dimension coordinates, height dimension coordinates, quantity dimension coordinates and group index dimension coordinates in the channel folding storage mode, the channel folding address information of the first pixel band is determined based on the width dimension coordinates and the width dimension step, the height dimension coordinates and the height dimension step, the quantity dimension coordinates and the quantity dimension step, the group index dimension coordinates and the group index dimension step and the channel dimension coordinates in the channel folding storage mode.
5. The method according to claim 1, characterized in that The first pixel band address information, based on the first pixel band address information, determines the second pixel band address information of other pixel bands in the target data block except the first pixel band, and the first pixel band address information includes: When the first pixel band address information is channel priority address information, determining second pixel band address information of other pixel bands in the target data block except the first pixel band based on the channel priority address information and the width dimension coordinate change amount and width dimension step of the width dimension, the height dimension coordinate change amount and height dimension step of the height dimension, and the number dimension coordinate change amount and number dimension step of the pixel band coordinates of the first pixel band in the channel priority storage mode; When the first pixel band address information is channel folding address information, based on the channel folding address information, and the width dimension coordinate change and width dimension step of the width dimension, the height dimension coordinate change and height dimension step of the height dimension, the quantity dimension coordinate change and quantity dimension step of the quantity dimension, the grouping index dimension coordinate change and grouping index dimension step of the grouping index dimension, and the channel dimension coordinate change of the pixel band coordinates of the first pixel band in the channel folding storage mode, determine the second pixel address information corresponding to other pixel bands in the target data block except the first pixel band.
6. The method according to claim 1, characterized in that Before determining the characteristic target matrix of the target data block according to the pixel data, the method further includes: After loading the pixel data corresponding to the first pixel band address information, the second pixel band address information and the target data block, the pixel data of the pixel bands belonging to the same target sub-data block are assembled.
7. The method according to claim 1, characterized in that The loading of pixel data corresponding to the target data block according to the first pixel band address information and the second pixel band address information comprises: Selecting a height coordinate value of a height dimension coordinate of the first pixel band address information as a current height coordinate value, and selecting a width coordinate value of a width dimension coordinate of the first pixel band address information as a current width coordinate value; Comparing the current height coordinate value with a first height filling threshold and a second height filling threshold, the first height filling threshold being less than the second height filling threshold; When the current height coordinate value is less than the first height filling threshold or the current height coordinate value is not less than the second height filling threshold, obtaining filling data, using the filling data as pixel data, and updating the height coordinate value of the height dimension coordinate of the second pixel band address information to the current height coordinate value, and updating the width coordinate value of the width dimension coordinate of the second pixel band address information to the current width coordinate value, returning to the step of comparing the current height coordinate value with the first height filling threshold and the second height filling threshold and continuing to execute; When the current height coordinate value is less than the first height filling threshold or the current height coordinate value is less than the second height filling threshold, comparing the current width coordinate value with the first width filling threshold and the second width filling threshold, the first width filling threshold being less than the second width filling threshold; When the current width coordinate value is less than the first width fill threshold or the current width coordinate value is greater than or equal to the second width fill threshold, the filling data is obtained, the filling data is used as pixel data, and the height coordinate value of the height dimension coordinate of the second pixel band address information is updated to the current height coordinate value, and the width coordinate value of the width dimension coordinate of the second pixel band address information is updated to the current width coordinate value, and the step of comparing the current height coordinate value with the first height fill threshold and the second height fill threshold is returned and continued.
8. The method according to claim 1, characterized in that The loading of pixel data corresponding to the target data block according to the first pixel band address information and the second pixel band address information comprises: Compare the grouping index coordinate value of the grouping index dimension coordinate in the first pixel band address information and the second pixel band address information with the grouping threshold to determine whether the grouping index exceeds the number of divided groups; If the group index exceeds the number of divided groups, the padding data is directly returned and the padding data is used as the pixel data corresponding to the target data block; If the group index does not exceed the number of divided groups, determine whether there is pixel data in the memory based on the height coordinate value of the height dimension coordinate and the width coordinate value of the width dimension coordinate in the first pixel band address information and the second pixel band address information, and if there is no pixel data, obtain filling data and use the filling data as pixel data.
9. An image data processing device, characterized in that: The device comprises: A data acquisition module is used to obtain the location information of the target data block; A coordinate determination module, configured to determine, according to the position information of the target data block, the pixel band coordinates of the preset data storage mode corresponding to the first pixel band in the target data block; A first address determining module, configured to determine first pixel band address information of the first pixel band based on the pixel band coordinates of the first pixel band; A second address determination module, configured to determine second pixel band address information of other pixel bands in the target data block except the first pixel band based on the first pixel band address information; The matrix construction module is used to load the pixel data corresponding to the target data block according to the first pixel band address information and the second pixel band address information, and determine the characteristic target matrix of the target data block according to the pixel data.
10. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 8 are implemented.
11. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 8 are implemented.
12. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 8 are implemented.
Citation Information
Patent Citations
Feature map processing method and device, electronic equipment and computer readable medium
CN113888390A
Flexible access instructions for efficient access to ML data
CN114648104A