Image processing method and device, electronic equipment and storage medium
By generating and using a scaling convolution kernel group to process bilinear interpolation, the problem of long-term bilinear interpolation calculation in neural network processing units is solved, and more efficient image processing is achieved.
Patent Information
- Application Number
- CN202510357779.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-06-13
AI Technical Summary
In the prior art, the calculation of bilinear interpolation takes a long time, resulting in inefficient processing efficiency of neural network processing units.
The convolution kernel size and channel number of the scaled convolution kernel group are determined based on the size of the original image and the target image, and the bilinear interpolation algorithm is used to generate the scaled convolution kernel group that satisfies these sizes and channel number, and the filling image is used for convolution operations to generate the target image.
Converting vector operations of bilinear interpolation algorithm into convolutional operations makes full use of the parallel capabilities of neural network processors, significantly reducing computing time and improving computing efficiency.
Smart Images

Figure CN120147115A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technology. Specifically, the present disclosure relates to an image processing method, apparatus, electronic device, and storage medium. Background Art
[0002] A neural network chip is a hardware device designed specifically for accelerating neural network computations. Since the calculation of bilinear interpolation requires the use of surrounding neighborhood data for interpolation operations, it is characterized by a relatively high computational density, data access locality, and low parallelism. It is usually implemented using a vector processing unit (a type of vector processing unit). Due to the relatively low computing power of the vector unit, the processing time is relatively long. Summary of the Invention
[0003] Embodiments of the present disclosure provide an image processing method, apparatus, electronic device, and storage medium, which can solve the problem of relatively long calculation time for bilinear interpolation in the prior art. The technical solutions provided by the present disclosure are as follows: According to one aspect of the embodiments of the present disclosure, there is provided an image processing method applied to a neural network processing unit. The method includes: Based on the original image size of the original image and the target image size of the target image, determining the convolution kernel size and number of channels of the scaling convolution kernel group to be generated; the target image is an image obtained by scaling the original image based on the bilinear interpolation algorithm; Padding the original image based on a preset padding number to obtain a padded image; Generating a scaling convolution kernel group that satisfies the convolution kernel size and the number of channels based on the bilinear interpolation algorithm; Obtaining the target image based on the scaling convolution kernel group and the padded image.
[0004] Optionally, the generating a scaling convolution kernel group that satisfies the convolution kernel size and the number of channels based on the bilinear interpolation algorithm includes: For each target pixel in the target image, determining a preset number of first pixels in the padded image that are associated with the target pixel; Based on the bilinear interpolation algorithm, determining the correlation coefficient corresponding to each of the preset number of first pixels, and constructing a convolution kernel for generating the target pixel based on the correlation coefficients corresponding to the respective first pixels among the preset number of first pixels; Generating the scaling convolution kernel group based on the convolution kernels corresponding to the respective target pixels.
[0005] Optionally, the determining the correlation coefficient corresponding to each of the preset number of first pixels based on the bilinear interpolation algorithm includes: Determine the reference coordinates of the target pixel based on the coordinates and scaling ratio of the target pixel in the target image; Input the coordinates corresponding to each first pixel in the filled image and the reference coordinates of the target pixel into the calculation formula of the bilinear interpolation algorithm to obtain the correlation coefficients corresponding to each first pixel.
[0006] Optionally, the determining the convolution kernel size and number of channels of the scaling convolution kernel group to be generated based on the original image size of the original image and the target image size of the target image includes: Take the square of the target image height or the square of the target image width as the number of channels of the scaling convolution kernel group; Based on any one of the original image height and the original image width of the original image and the preset padding quantity, determine the convolution kernel size of the scaling convolution kernel group.
[0007] Optionally, the obtaining the target image based on the scaling convolution kernel group and the filled image includes: Take the original image height or the original image width in the original image size of the original image as the convolution stride; Perform a convolution operation on the filled image based on the scaling convolution kernel group according to the convolution stride to obtain the feature map; Perform a size transformation on the feature map, and use the transformed feature map as the target image.
[0008] Optionally, the filling the original image based on the preset padding quantity to obtain a filled image includes: Determine at least two expansion directions in the original image; For each expansion direction, determine the boundary corresponding to the original image in this expansion direction, and expand the corresponding boundary by the preset padding quantity of rows or columns along this expansion direction; Generate the filled image based on the image of the original image after expansion in each expansion direction.
[0009] Optionally, the generating the filled image based on the image of the original image after expansion in each expansion direction includes: Take the image of the original image after expansion in each expansion direction as the expanded image; Based on a preset filling strategy, fill the unfilled pixels in the expanded image, and use the filled expanded image as the filled image.
[0010] According to another aspect of the embodiments of the present disclosure, there is provided an image processing apparatus, and the apparatus includes: A determination module, configured to determine the kernel size and number of channels of a scaling convolution kernel group to be generated based on the original image size of the original image and the target image size of the target image; the target image is an image obtained by scaling the original image based on the bilinear interpolation algorithm. A padding module, configured to pad the original image based on a preset padding number to obtain a padded image. A convolution kernel generation module, configured to generate a scaling convolution kernel group that meets the kernel size and the number of channels based on the bilinear interpolation algorithm. An image scaling module, configured to obtain the target image based on the scaling convolution kernel group and the padded image.
[0011] According to another aspect of the embodiments of the present disclosure, an electronic device is provided. The electronic device includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps of any one of the above image processing methods are implemented.
[0012] According to still another aspect of the embodiments of the present disclosure, a computer-readable storage medium is provided. A computer program is stored on the computer-readable storage medium. When the computer program is executed by a processor, the steps of any one of the above image processing methods are implemented.
[0013] According to one aspect of the embodiments of the present disclosure, a computer program product is provided, which includes a computer program. When the computer program is executed by a processor, the steps of any one of the above image processing methods are implemented.
[0014] The beneficial effects brought by the technical solutions provided by the embodiments of the present disclosure are as follows: Based on the original image size of the original image and the target image size of the target image, determine the kernel size and number of channels of the scaling convolution kernel group to be generated, and generate a scaling convolution kernel group that meets the kernel size and the number of channels based on the bilinear interpolation algorithm. Obtain the target image based on the scaling convolution kernel group and the padded image, convert the vector operation of the bilinear interpolation algorithm into a convolution operation using the scaling convolution kernel group by a neural network processor, so as to make full use of the efficient parallel ability of the neural network processor for convolution operations, reduce the calculation time of the bilinear interpolation algorithm, and greatly improve the calculation efficiency of the bilinear interpolation algorithm. Description of the Drawings
[0015] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure, the following will briefly introduce the drawings required for the description in the embodiments of the present disclosure.
[0016] Figure 1 It is a schematic flowchart of an image processing method provided by an embodiment of the present disclosure; Figure 2A schematic diagram of image scaling provided by an embodiment of the present disclosure; Figure 3 A schematic diagram of the bilinear interpolation principle; Figure 4 A schematic diagram of the mapping relationship between target pixels and candidate pixels provided by an embodiment of the present disclosure; Figure 5 A schematic diagram of a scaling convolution kernel group provided by an embodiment of the present disclosure; Figure 6 A schematic flowchart of image scaling based on convolution operation provided by an embodiment of the present disclosure; Figure 7 Another schematic flowchart of image scaling based on convolution operation provided by an embodiment of the present disclosure; Figure 8 A schematic flowchart of RGB image scaling based on convolution operation provided by an embodiment of the present disclosure; Figure 9 A schematic diagram of the structure of an image processing apparatus provided by an embodiment of the present disclosure; Figure 10 A schematic diagram of the structure of an electronic device provided by an embodiment of the present disclosure. Detailed implementation manners
[0017] The embodiments of the present disclosure will be described below with reference to the accompanying drawings in the present disclosure. It should be understood that the embodiments described below with reference to the accompanying drawings are exemplary descriptions for explaining the technical solutions of the embodiments of the present disclosure, and do not constitute limitations on the technical solutions of the embodiments of the present disclosure.
[0018] Those skilled in the art of the present technology can understand that unless specifically stated otherwise, the singular forms "a", "an", "the", and "said" used herein may also include the plural forms. It should be further understood that the terms "including" and "comprising" used in the embodiments of the present disclosure mean that the corresponding features can be implemented as the presented features, information, data, steps, operations, elements, and / or components, but do not exclude the implementation of other features, information, data, steps, operations, elements, components, and / or their combinations supported by the art of the present technology. It should be understood that when we say an element is "connected" or "coupled" to another element, the one element can be directly connected or coupled to the other element, or it can mean that the one element and the other element establish a connection relationship through an intermediate element. In addition, the "connection" or "coupling" used herein can include a wireless connection or wireless coupling. The term "and / or" used herein indicates at least one of the items defined by the term, for example, "A and / or B" or "A, B" indicates implementation as "A", or implementation as "B", or implementation as "A and B".
[0019] To make the objectives, technical solutions, and advantages of the present disclosure clearer, the embodiments of the present disclosure will be further described in detail below with reference to the accompanying drawings.
[0020] The technical solutions of the embodiments of the present disclosure and the technical effects produced by the technical solutions of the present disclosure will be described below by describing several exemplary embodiments. It should be noted that the following embodiments can refer to, draw on, or combine with each other. For the same terms, similar features, and similar implementation steps in different embodiments, they will not be described repeatedly.
[0021] Figure 1 A flowchart of an image processing method provided by an embodiment of the present disclosure is shown as Figure 1 shown, and the method includes: Step S110, determining the kernel size and number of channels of the scaling convolution kernel group to be generated based on the original image size of the original image and the target image size of the target image; the target image is an image obtained by scaling the original image based on the bilinear interpolation algorithm.
[0022] Specifically, the execution subject of the image processing method provided by the embodiment of the present disclosure is a neural network processor (Neural Network Processing Unit, NPU).
[0023] The original image can be an image to be processed, and the target image is an image obtained by scaling the original image based on the bilinear interpolation algorithm. The original image size can include the original image height and the original image width, and the target image size can include the target image height and the target image width. The image height can be the number of pixels included in each column of the image, and the image width can be the number of pixels included in each row of the image.
[0024] The scaling ratio can be the ratio between the original image size and the target image size. For example, if the original image size of the original image is 240×240 and it needs to be interpolated to the target image size of 640×640 using the bilinear interpolation algorithm, the scaling ratio is 3:8.
[0025] It should be noted that the original image height and the original image width are the same, and the target image height and the target image width are the same.
[0026] The scaling convolution kernel group can be a convolution kernel group used to perform scaling processing on the original image. That is to say, in the embodiments of the present disclosure, the vector operation of the bilinear interpolation algorithm is converted into a convolution operation using the scaling convolution kernel group by the neural network processor, so as to make full use of the efficient parallel ability of the neural network processor for convolution operations, reduce the calculation time of the bilinear interpolation algorithm, and greatly improve the calculation efficiency of the bilinear interpolation algorithm.
[0027] Optionally, based on the original image size of the original image and the target image size of the target image, determine the convolution kernel size and the number of channels of the scaling convolution kernel group to be generated, including: Take the square of the target image height or the square of the target image width as the number of channels of the scaling convolution kernel group; Based on either the original image height or the original image width and the preset padding amount, determine the convolution kernel size of the scaling convolution kernel group.
[0028] Specifically, after determining the original image size and the target image size, the square of the target image height or the square of the target image width can be taken as the number of channels of the scaling convolution kernel group, where the number of channels of the scaling convolution kernel group is the number of convolution kernels included in the scaling convolution kernel group.
[0029] Based on either the original image height or the original image width and the preset padding amount, determine the convolution kernel size of the scaling convolution kernel group.
[0030] Optionally, the sum of the original image height and the preset padding amount can be taken as the height of the convolution kernel; the sum of the original image width and the preset padding amount can be taken as the width of the convolution kernel.
[0031] Assume the scaling ratio is 3:8. Figure 2 For a schematic diagram of image scaling provided by an embodiment of the present disclosure, the embodiments of the present disclosure and subsequent embodiments are described by taking an image region of 3×3 as the original image as an example. Then, the original image height and the original image width of the original image (i.e., Figure 2 the red image in Figure 2 are 3, and the target image height and the target image width of the target image (i.e.,
[0032] Step S120, pad the original image based on the preset padding amount to obtain a padded image.
[0033] Specifically, after determining the original image, the original image can be padded based on the preset padding amount, and the padded original image is used as the padded image.
[0034] Optionally, padding the original image based on the preset padding amount to obtain a padded image includes: Determine at least two expansion directions in the original image; For each expansion direction, determine the boundary of the original image corresponding to this expansion direction, and expand the corresponding boundary by the preset padding amount of rows or columns along this expansion direction; Generate a padded image based on the image obtained by expanding the original image in each expansion direction.
[0035] Specifically, at least two expansion directions of the original image are determined. One expansion direction can be selected from the left side and the right side, and one expansion direction can be selected from the upper side and the lower side. That is to say, the at least two expansion directions can include any combination of any one of the left side and the right side and any one of the upper side and the lower side.
[0036] There is a boundary corresponding to each expansion direction. The boundary corresponding to the right side is the rightmost column, the boundary corresponding to the left side is the leftmost column, the boundary corresponding to the upper side is the uppermost row, and the boundary corresponding to the lower side is the lowermost row.
[0037] After determining at least two expansion directions, for each expansion direction, the boundary corresponding to the expansion direction can be determined, and the boundary is expanded by a preset number of rows or columns along the expansion direction. When the boundary is a column, a preset number of columns are expanded; when the boundary is a row, a preset number of rows are expanded.
[0038] For example, taking the original image of 3×3 as an example, at least two expansion directions include the right side and the lower side, the corresponding boundaries include the rightmost column and the lowermost row, and the preset filling number is 1. Then, the rightmost column of the original image can be expanded by one column along the right side, and the lowermost row of the original image can be added by one row along the lower side, thereby obtaining a 4×4 image.
[0039] After expanding the original image in each expansion direction, a filled image is generated based on the image of the original image after expansion in each expansion direction.
[0040] Optionally, generating a filled image based on the image of the original image after expansion in each expansion direction includes: Taking the image of the original image after expansion in each expansion direction as the expanded image; Based on a preset filling strategy, filling the unfilled pixels in the expanded image, and taking the filled expanded image as the filled image.
[0041] Specifically, taking the image of the original image after expansion in each expansion direction as the expanded image, based on a preset filling strategy, filling the pixel values of the unfilled pixels (i.e., the pixels in the expanded part) in the expanded image, and taking the filled expanded image as the filled image.
[0042] Among them, when performing image scaling in the Asymmetric mode, the preset filling strategy can be set to mirror padding, that is, copying the pixel values in the boundary corresponding to each expansion direction into the expanded row or column.
[0043] By filling the original image, a filled image is obtained. By expanding the boundaries of the input image, the edge information of the original image can be effectively retained.
[0044] Step S130: Based on the bilinear interpolation algorithm, generate a scaled convolution kernel group that meets the convolution kernel size and number of channels.
[0045] Specifically, the bilinear interpolation algorithm is an algorithm commonly used for spatial data interpolation in image processing and signal processing. The bilinear interpolation algorithm is widely used in image processing scenarios that require smooth pixel transitions. For example, when adjusting the image resolution (such as enlarging or reducing), bilinear interpolation is used to generate smooth new pixel values to avoid jagged or blocky artifacts; when generating an image pyramid (such as a Gaussian pyramid), bilinear interpolation is used for resolution conversion between different levels to support feature matching or object detection; in a semantic segmentation model, the decoder part uses bilinear interpolation to upsample the low-resolution feature map to restore detailed information.
[0046] The image processing method provided by the embodiments of the present disclosure implements the bilinear interpolation algorithm based on convolution operations. That is to say, the image processing method provided by the embodiments of the present disclosure can be applied to the application scenarios applicable to the bilinear interpolation algorithm.
[0047] The principle of the bilinear interpolation algorithm is to estimate the value of an unknown point using the four neighboring pixels of known points (forming a rectangle in two-dimensional space). The interpolation process is divided into two steps: first, linear interpolation is performed in the horizontal direction, and then another linear interpolation is performed in the vertical direction, so it is called bilinear interpolation.
[0048] Figure 3 For a schematic diagram of the bilinear interpolation principle, as Figure 3 shown, , , , are four known points. First, perform two single linear interpolations in the x direction to obtain and two temporary points, and then perform a single linear interpolation in the y direction to obtain .
[0049] Therefore, the coordinate calculation formula (hereinafter referred to as Formula 1) of can be obtained as:
[0050] After determining the convolution kernel size and number of channels of the scaled convolution kernel group, a scaled convolution kernel group that meets the convolution kernel size and number of channels can be generated based on the bilinear interpolation algorithm.
[0051] Optionally, based on the bilinear interpolation algorithm, a scaled convolution kernel group that meets the convolution kernel size and the number of channels is generated, including: For each target pixel in the target image, determine a preset number of first pixels in the padded image that are associated with the target pixel; Based on the bilinear interpolation algorithm, determine the correlation coefficient corresponding to each of the preset number of first pixels, and based on the correlation coefficients corresponding to each of the preset number of first pixels, construct a convolution kernel for generating the target pixel; Generate a scaled convolution kernel group based on the convolution kernels corresponding to each target pixel respectively.
[0052] Specifically, taking the pixels in the target image as target pixels and the pixels in the padded image as candidate pixels, for each target pixel in the target image, a preset number of first pixels associated with the second pixel can be determined from multiple candidate pixels in the padded image.
[0053] The principle of the bilinear interpolation algorithm is to calculate the coordinates of an unknown point based on the four nearest pixels. The preset number of first pixels associated with the second pixel can be the preset number of candidate pixels in the padded image used to calculate the target pixel in the bilinear interpolation algorithm. Among them, the preset number can be 4.
[0054] Figure 4 For a schematic diagram of the mapping relationship between a target pixel and a candidate pixel provided by an embodiment of the present disclosure, as Figure 4 shown, based on the example provided in Figure 3 , using the least common multiple 24 of 3 and 8, the mapping relationship between the output green cells (i.e., target pixels) and the input red cells (i.e., candidate pixels) is shown in a 24×24 grid. Figure 4 The dashed arrows in
[0055] In Figure 4 show the relationship between the red cells in the padded image and the green cells in the target image in the 24×24 grid. 、 、 and , where Denote the candidate pixel at the \(i\)-th row and \(j\)-th column in the filled image. It can be understood that the preset number of first pixels associated with the target pixel include the nearest preset number of candidate pixels centered on the target pixel and surrounding the target pixel.
[0056] For each target pixel in the target image, after determining the preset number of first pixels associated with the target pixel, based on the bilinear interpolation algorithm, determine the correlation coefficient corresponding to each first pixel among the preset number of first pixels, where the correlation coefficient corresponding to the first pixel can be used to reflect the relationship between the pixel value of the first pixel and the pixel value of the target pixel in the bilinear interpolation algorithm.
[0057] After obtaining the correlation coefficients corresponding to each of the preset number of first pixels, based on the position of each first pixel in the filled image and the correlation coefficient corresponding to the first pixel, determine multiple parameter values in the convolution kernel used to generate the target pixel.
[0058] Construct the convolution kernel corresponding to each target pixel based on the bilinear interpolation algorithm, such that the target image generated by the convolution operation based on the scaled convolution kernel group is consistent with the result obtained by the original bilinear interpolation algorithm.
[0059] Based on the determination method of the convolution kernel size, it can be known that the convolution kernel size of the convolution kernel is consistent with the filled image size of the filled image. Assuming the original image size is \(M\times M\) and the preset filling quantity is \(L\), both the convolution kernel size and the filled image size are \((M + L)\times(M + L)\).
[0060] Therefore, each candidate pixel in the filled image has a corresponding position in the convolution kernel. For each target pixel, the correlation coefficients of the preset number of first pixels associated with the target pixel can be used as the parameter values at the corresponding positions of the first pixels in the convolution kernel. The parameter values at other positions in the convolution kernel can be set to 0, that is, when calculating the pixel value of the second pixel, the pixel values of other candidate pixels except the first pixels associated with the second pixel are not considered.
[0061] Optionally, based on the bilinear interpolation algorithm, determining the correlation coefficient corresponding to each first pixel among the preset number of first pixels includes: Based on the coordinates and scaling ratio of the target pixel in the target image, determine the reference coordinates of the target pixel; Input the coordinates corresponding to each first pixel in the filled image and the reference coordinates of the target pixel into the calculation formula of the bilinear interpolation algorithm to obtain the correlation coefficients corresponding to each first pixel.
[0062] Optionally, for each target pixel in the target image, based on the coordinates and the scaling ratio of the target pixel in the target image, the reference coordinates of the target pixel are obtained. Wherein, the coordinates of the target pixel in the target image may include the row and column where the target pixel is located in the target image, and the product of the coordinates of the target pixel in the target image and the scaling ratio may be used as the reference coordinates of the target pixel.
[0063] For example, in combination with Figure 4 the example in as an example, represents the pixel at the i-th row and j-th column in the target image, the coordinates in the target image are (1, 2), and the scaling ratio is 3:8, then the reference coordinates of are (
[0064] The coordinates of a preset number of first pixels in the filled image and the reference coordinates of the target pixel are input into the calculation formula of the bilinear interpolation algorithm (i.e., formula 1), and the respective correlation coefficients corresponding to each first pixel are obtained based on the calculation results.
[0065] In combination with Figure 4 the example in as an example, represents the pixel at the i-th row and j-th column in the target image, and the 4 first pixels associated with include , , and , substituting the coordinates of , , and , as well as the reference coordinates of into formula 1, that is, substituting , , and into the corresponding positions in formula 1 based on the relative position relationship , , , , and using the reference coordinates of as in formula 1.
[0066] The pixel value of and the pixel values , , and corresponding to , , The relationship between, that is , where , , , correspond to the correlation coefficients of these four first pixels respectively. Based on the correlation coefficients corresponding to these four first pixels, multiple parameter values in the convolution kernel corresponding to the target pixel can be determined.
[0067] Figure 5 is a schematic diagram of a scaled convolution kernel group provided by an embodiment of the present disclosure. As Figure 5 shown, Figure 5 the red cells at the four corners of the blue square in [figure number] correspond to the four candidate pixels in the upper left corner of the filled image. The 4 first pixels associated with the green cell (i.e., the second pixel) in the blue square include , , and in the filled image; the red cells at the four corners of the yellow square correspond to the four candidate pixels in the upper right corner of the filled image. The 4 first pixels associated with the green cell in the yellow square include , , and in the filled image; the red cells at the four corners of the green square correspond to the four candidate elements in the lower left corner of the filled image. The 4 first pixels associated with the green cell in the green square include , , and in the filled image; the red cells at the four corners of the orange square correspond to the four candidate elements in the lower right corner of the filled image. The 4 first pixels associated with the green cell in the orange square include , , and in the filled image.
[0068] For the target pixel in the blue square, the parameter values of the four grids in the upper left corner of the corresponding convolution kernel respectively correspond to the correlation coefficients corresponding to the four first pixels in the upper left corner of the filled image. For the target pixel in the yellow square, the parameter values of the four grids in the upper right corner of the corresponding convolution kernel respectively correspond to the correlation coefficients corresponding to the four first pixels in the upper right corner of the filled image. For the target pixel in the green square, the parameter values of the four grids in the lower left corner of the corresponding convolution kernel respectively correspond to the correlation coefficients corresponding to the four first pixels in the lower left corner of the filled image. For the target pixel , the parameter values of the four grids in the lower right corner area of its corresponding convolution kernel respectively correspond to the correlation coefficients corresponding to the four first pixels in the lower right corner of the filled image.
[0069] After obtaining the convolution kernels corresponding to each target pixel, multiple convolution kernels can be combined in a preset order, and the combined convolution kernel group is used as the scaled convolution kernel group.
[0070] Among them, the number of target pixels (i.e., the square of the height of the target image or the square of the width of the target image) is equal to the number of convolution kernels in the scaled convolution kernel group, that is, the number of channels of the scaled convolution kernel group.
[0071] Step S140, obtaining a target image based on the scaled convolution kernel group and the filled image.
[0072] Specifically, after obtaining the scaled convolution kernel group, a convolution operation can be performed on the filled image based on the scaled convolution kernel group, and a target image can be obtained based on the result of the convolution operation.
[0073] Optionally, obtaining a target image based on the scaled convolution kernel group and the filled image includes: Taking the original image height or the original image width in the original image size of the original image as the convolution stride; Performing a convolution operation on the filled image based on the scaled convolution kernel group according to the convolution stride to obtain the feature map; Performing a size transformation on the feature map, and using the transformed feature map as the target image.
[0074] Specifically, taking the original image height or the original image width in the original image size of the original image as the convolution stride (i.e., stride), sliding the convolution kernels in the scaled convolution kernel group on the filled image according to the convolution stride to obtain a feature map, and the size of the feature map is 1×1×(N×N), where N represents the height or width of the target image. Performing a size transformation (i.e., reshape) on the feature map to convert it into an image size of N×N, and using the image after the size transformation as the target image.
[0075] It should be noted that the generation of the scaled convolution kernel group in the embodiments of the present disclosure can be generated offline or online. If the computing power of the neural network processor is not high enough and the application has high requirements for latency, the scaled convolution kernel group can be generated offline and directly used when performing image scaling, which can better meet the application requirements. If the computing power of the neural network processor is relatively high and the application has low requirements for latency, the scaled convolution kernel group can be generated online, and generating the scaled convolution kernel group online can save the storage space of the neural network processor. It can be set according to the actual hardware environment and application requirements of the neural network processor, and the embodiments of the present disclosure do not make any limitations in this regard.
[0076] In the embodiments of the present disclosure, based on the original image size of the original image and the target image size of the target image, the convolution kernel size and the number of channels of the scaling convolution kernel group to be generated are determined, and based on the bilinear interpolation algorithm, a scaling convolution kernel group that meets the convolution kernel size and the number of channels is generated. Based on the scaling convolution kernel group and the padded image, the target image is obtained, and the vector operation of the bilinear interpolation algorithm is converted into a convolution operation using the scaling convolution kernel group by a neural network processor, so that the efficient parallelism of the neural network processor for convolution operations can be fully utilized, the calculation time of the bilinear interpolation algorithm is reduced, and the calculation efficiency of the bilinear interpolation algorithm is greatly improved.
[0077] Figure 6 It is a schematic flowchart of image scaling based on convolution operation provided by the embodiments of the present disclosure. As Figure 6 shown, a 3×3 region in the original image is enlarged. For the rightmost and bottommost boundary parts of the original image, one row and one column are expanded using mirror-pad. For the 4×4 red region, a depth-wise convolution kernel with a kernel size of 4×4 and a channel of 64 is configured for convolution operation, and the output result is a feature map of 1×1×64. After reshaping the output result of 1×1×64, it becomes 8×8, thus realizing enlarging the original image to the target image according to the scaling ratio of 3:8.
[0078] Figure 7 It is a schematic flowchart of image scaling based on convolution operation provided by the embodiments of the present disclosure. As Figure 7 shown, for the original image, convolution is performed using a 4×4 kernel with a stride of 3, that is, the 3×3 input in the original image is scaled to the desired 8×8. The whole process only contains depth-wise convolution and reshape, without additional overhead and calculation.
[0079] Figure 8 It is a schematic flowchart of RGB image scaling based on convolution operation provided by the embodiments of the present disclosure. As Figure 8 shown, Figure 8 the red, green, and blue in Figure 8 respectively correspond to the three RGB channels,
[0080] After testing, it can support bilinear resize with any ratio (on the premise that the maximum convolution kernel supported by the neural network processor, that is, the size of the convolution kernels in the generated scaled convolution kernel group cannot exceed the maximum convolution kernel size supported by the neural network processor), and can achieve exactly the same result as the original bilinear interpolation. Compared with the vector-based implementation, using depth-wise convolution can improve the time performance by 20 - 30 times.
[0081] Figure 9 The structural schematic diagram of an image processing device provided by an embodiment of the present disclosure is as follows Figure 9 As shown, the device of this embodiment may include: A determination module 210, configured to determine the convolution kernel size and the number of channels of the to-be-generated scaled convolution kernel group based on the original image size of the original image and the target image size of the target image; the target image is an image obtained by scaling the original image based on the bilinear interpolation algorithm; A padding module 220, configured to pad the original image based on a preset padding number to obtain a padded image; A convolution kernel generation module 230, configured to generate a scaled convolution kernel group that meets the convolution kernel size and the number of channels based on the bilinear interpolation algorithm; An image scaling module 240, configured to obtain the target image based on the scaled convolution kernel group and the padded image.
[0082] As an optional embodiment, when the convolution kernel generation module generates a scaled convolution kernel group that meets the convolution kernel size and the number of channels based on the bilinear interpolation algorithm, it is specifically configured to: For each target pixel in the target image, determine a preset number of first pixels in the padded image associated with the target pixel; Based on the bilinear interpolation algorithm, determine the correlation coefficient corresponding to each first pixel among the preset number of first pixels, and construct a convolution kernel for generating the target pixel based on the correlation coefficients corresponding to each of the preset number of first pixels; Generate the scaled convolution kernel group based on the convolution kernels corresponding to each target pixel respectively.
[0083] As an optional embodiment, when the convolution kernel generation module determines the correlation coefficient corresponding to each first pixel among the preset number of first pixels based on the bilinear interpolation algorithm, it is specifically configured to: Based on the coordinates and the scaling ratio of the target pixel in the target image, determine the reference coordinates of the target pixel; Input the coordinates corresponding to each first pixel in the filled image and the reference coordinates of the target pixel into the calculation formula of the bilinear interpolation algorithm to obtain the correlation coefficient corresponding to each first pixel.
[0084] As an alternative embodiment, when determining the convolution kernel size and the number of channels of the scaling convolution kernel group to be generated based on the original image size of the original image and the target image size of the target image, the determination module is specifically configured to: Use the square of the target image height or the square of the target image width as the number of channels of the scaling convolution kernel group; Based on either the original image height or the original image width of the original image and the preset padding quantity, determine the convolution kernel size of the scaling convolution kernel group.
[0085] As an alternative embodiment, when obtaining the target image based on the scaling convolution kernel group and the filled image, the image scaling module is specifically configured to: Use the original image height or the original image width in the original image size of the original image as the convolution stride; Perform a convolution operation on the filled image based on the scaling convolution kernel group according to the convolution stride to obtain the feature map; Perform a size transformation on the feature map, and use the transformed feature map as the target image.
[0086] As an alternative embodiment, when padding the original image based on the preset padding quantity to obtain a filled image, the padding module is specifically configured to: Determine at least two expansion directions in the original image; For each expansion direction, determine the boundary corresponding to the original image in this expansion direction, and expand the corresponding boundary by the preset padding quantity of rows or columns along this expansion direction; Generate the filled image based on the image of the original image after expansion in each expansion direction.
[0087] As an alternative embodiment, when generating the filled image based on the image of the original image after expansion in each expansion direction, the padding module is specifically configured to: Use the image of the original image after expansion in each expansion direction as the expanded image; Based on a preset padding strategy, fill the unfilled pixels in the expanded image, and use the filled expanded image as the filled image.
[0088] The device according to an embodiment of the present disclosure can execute the method provided by the embodiment of the present disclosure. Their implementation principles are similar and they have corresponding technical effects. The actions performed by each module in the device according to each embodiment of the present disclosure correspond to the steps in the method according to each embodiment of the present disclosure. For the detailed function descriptions of each module of the device, reference can specifically be made to the descriptions in the corresponding methods shown above, and details are not repeated here.
[0089] In an embodiment of the present disclosure, the term "module" or "unit" refers to a computer program with a predetermined function or a part of a computer program, which works together with other related parts to achieve a predetermined goal, and can be fully or partially implemented by using software, hardware (such as a processing circuit or a memory), or a combination thereof. Similarly, one processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of an overall module or unit that includes the function of the module or unit.
[0090] An electronic device is provided in an embodiment of the present disclosure, including a memory, a processor, and a computer program stored on the memory. The processor executes the above computer program to implement the steps of the method provided by any optional embodiment of the present disclosure. Compared with the prior art, it can be achieved that: based on the original image size of the original image and the target image size of the target image, the convolution kernel size and the number of channels of the scaling convolution kernel group to be generated are determined, and based on the bilinear interpolation algorithm, a scaling convolution kernel group that meets the convolution kernel size and the number of channels is generated. Based on the scaling convolution kernel group and the padded image, the target image is obtained, and the vector operation of the bilinear interpolation algorithm is converted into a convolution operation using the scaling convolution kernel group by a neural network processor, so that the efficient parallel computing ability of the neural network processor for convolution operations can be fully utilized, the computing time of the bilinear interpolation algorithm is reduced, and the computing efficiency of the bilinear interpolation algorithm is greatly improved.
[0091] In an optional embodiment, an electronic device is provided, as Figure 10 shown, Figure 10 the electronic device 4000 shown includes: a processor 4001 and a memory 4003. Among them, the processor 4001 and the memory 4003 are connected, such as connected through a bus 4002. Optionally, the electronic device 4000 may further include a transceiver 4004, and the transceiver 4004 can be used for data interaction between the electronic device and other electronic devices, such as data sending and / or data receiving, etc. It should be noted that in practical applications, the transceiver 4004 is not limited to one, and the structure of the electronic device 4000 does not constitute a limitation to the embodiments of the present disclosure.
[0092] The processor 4001 may be a CPU (Central Processing Unit), a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It can implement or execute various exemplary logic blocks, modules, and circuits described in connection with the present disclosure. The processor 4001 may also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, etc.
[0093] The bus 4002 may include a path for transmitting information between the above components. The bus 4002 may be a PCI (Peripheral Component Interconnect) bus, an EISA (Extended Industry Standard Architecture) bus, or the like. The bus 4002 may be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, Figure 10 only a thick line is shown herein, but it does not mean that there is only one bus or one type of bus.
[0094] The memory 4003 may be a ROM (Read Only Memory) or other type of static storage device that can store static information and instructions, a RAM (Random Access Memory), or other type of dynamic storage device that can store information and instructions. It may also be an EEPROM (Electrically Erasable Programmable Read Only Memory), a CD-ROM (Compact Disc Read Only Memory), or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media, other magnetic storage devices, or any other medium that can be used to carry or store computer programs and can be read by a computer, which is not limited herein.
[0095] The memory 4003 is used to store the computer program for implementing the embodiments of the present disclosure, and is controlled by the processor 4001 to execute. The processor 4001 is used to execute the computer program stored in the memory 4003 to implement the steps shown in the foregoing method embodiments.
[0096] Among them, the electronic device includes but is not limited to: mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Tablet Computers), PMPs (Portable Multimedia Players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), wearable devices, etc., and fixed terminals such as digital TVs, desktop computers, etc.
[0097] The embodiments of the present disclosure provide a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps and corresponding contents shown in the foregoing method embodiments can be implemented.
[0098] The embodiments of the present disclosure also provide a computer program product, including a computer program. When the computer program is executed by a processor, the steps and corresponding contents shown in the foregoing method embodiments can be implemented.
[0099] It should be understood that although the flowcharts in the embodiments of the present disclosure indicate various operation steps by arrows, the execution order of these steps is not limited to the order indicated by the arrows. Unless otherwise clearly stated in this document, in some implementation scenarios of the embodiments of the present disclosure, the implementation steps in each flowchart can be executed in other orders according to requirements. In addition, some or all of the steps in each flowchart may include multiple sub-steps or multiple stages based on the actual implementation scenario. Some or all of these sub-steps or stages can be executed at the same time, and each sub-step or stage among these sub-steps or stages can also be executed at different times respectively. In the scenario where the execution times are different, the execution order of these sub-steps or stages can be flexibly configured according to requirements, and the embodiments of the present disclosure do not limit this.
[0100] The above are only optional implementation manners of some implementation scenarios of the present disclosure. It should be noted that for those of ordinary skill in the art in this technical field, without departing from the technical concept of the solution of the present disclosure, using other similar implementation means based on the technical idea of the present disclosure also belongs to the protection scope of the embodiments of the present disclosure.
Claims
1. An image processing method, characterized in that: Applied to a neural network processing unit, the method comprises: Determine the convolution kernel size and the number of channels of the scaled convolution kernel group to be generated based on the original image size of the original image and the target image size of the target image; the target image is an image obtained by scaling the original image based on a bilinear interpolation algorithm; Filling the original image based on a preset filling amount to obtain a filled image; Based on a bilinear interpolation algorithm, a scaled convolution kernel group satisfying the convolution kernel size and the number of channels is generated; Based on the scaled convolution kernel group and the padded image, the target image is obtained.
2. The method according to claim 1, characterized in that: The step of generating a scaled convolution kernel group satisfying the convolution kernel size and the number of channels based on a bilinear interpolation algorithm includes: For each target pixel in the target image, determining a preset number of first pixels in the filling image that are associated with the target pixel; Based on a bilinear interpolation algorithm, determining a correlation coefficient corresponding to each first pixel in the preset number of first pixels, and constructing a convolution kernel for generating the target pixel based on the correlation coefficient corresponding to each first pixel in the preset number of first pixels; The scaled convolution kernel group is generated based on the convolution kernels corresponding to the target pixels.
3. The method according to claim 2, characterized in that The determining, based on the bilinear interpolation algorithm, the correlation coefficient corresponding to each first pixel in the preset number of first pixels comprises: Determining a reference coordinate of the target pixel based on the coordinates of the target pixel in the target image and the scaling ratio; The coordinates corresponding to each first pixel in the filling image and the reference coordinates of the target pixel are input into the calculation formula of the bilinear interpolation algorithm to obtain the correlation coefficients corresponding to each first pixel.
4. The method according to claim 1, characterized in that: The step of determining the convolution kernel size and the number of channels of the scaled convolution kernel group to be generated based on the original image size of the original image and the target image size of the target image includes: Taking the square of the height of the target image or the square of the width of the target image as the number of channels of the scaled convolution kernel group; Based on any one of the original image height and the original image width and the preset padding amount, determine the convolution kernel size of the scaled convolution kernel group.
5. The method according to claim 1, characterized in that: The step of obtaining the target image based on the scaled convolution kernel group and the padded image comprises: Using an original image height or an original image width in an original image size of the original image as a convolution step; According to the convolution step size, a convolution operation is performed on the padded image based on the scaled convolution kernel group to obtain the feature map; The feature map is resized and the resized feature map is used as the target image.
6. The method according to any one of claims 1 to 5, characterized in that The filling the original image based on a preset filling amount to obtain a filled image includes: Determining at least two expansion directions in the original image; For each expansion direction, determining a boundary of the original image corresponding to the expansion direction, and expanding the corresponding boundary by a preset filling number of rows or columns along the expansion direction; The filling image is generated based on the image after the original image is expanded in each expansion direction.
7. The method according to claim 6, characterized in that: The step of generating the filled image based on the image after the original image is expanded in each expansion direction comprises: The image after the original image is expanded in each expansion direction is used as the expanded image; Based on a preset filling strategy, unfilled pixels in the extended image are filled, and the filled extended image is used as a filled image.
8. An image processing device, characterized in that: include: A determination module, used for determining a convolution kernel size and a number of channels of a scaled convolution kernel group to be generated based on an original image size of an original image and a target image size of a target image; The target image is an image obtained by scaling the original image based on a bilinear interpolation algorithm; A filling module, used for filling the original image based on a preset filling amount to obtain a filled image; A convolution kernel generation module, used to generate a scaled convolution kernel group satisfying the convolution kernel size and the number of channels based on a bilinear interpolation algorithm; An image scaling module is used to obtain the target image based on the scaled convolution kernel group and the padding image.
9. An electronic device comprising a memory, a processor and a computer program stored in the memory, characterized in that: The processor executes the computer program to implement the steps of the method according to any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
11. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.