Acceleration system for carrying out geometric transformation on digital image
By designing an acceleration system in which the processor and accelerator work together in the computing device, the problem of low image geometric transformation operation efficiency in the prior art is solved, and efficient geometric transformation of multiple images is achieved, which significantly shortens the computing time.
Patent Information
- Application Number
- PCT/CN2024/134287
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-28
- Filing Date
- 2024-11-25
- Publication Date
- 2025-06-05
AI Technical Summary
The prior art has low computational efficiency when performing geometric transformation of images, especially when processing multiple images, pixel-by-pixel computing leads to low underlying logic efficiency of computing devices.
An acceleration system is designed, using the processor and the accelerator to work together, and by determining the target transformation matrix and the original coordinate position matrix, geometric transformation of multiple digital images simultaneously is achieved, thereby improving the computing efficiency.
Compared with the geometric transformation of a single image by pixel, the proposed acceleration system significantly improves the computing efficiency, especially in the geometric transformation scenarios of large batches of images, which shortens the computing time.
Smart Images

Figure CN2024134287_05062025_PF_FP_ABST
Abstract
Description
An acceleration system for geometric transformation of digital images
[0001] This application claims priority to the Chinese patent application filed with the China Patent Office on November 28, 2023, with application number 202311615134.X and application name “An acceleration system for geometric transformation of digital images”, the entire contents of which are incorporated by reference into this application. Technical Field
[0002] The present application relates to the field of image processing, and in particular to an acceleration system for performing geometric transformation on digital images. Background Art
[0003] Geometric transformation (such as affine transformation and transmission transformation) is a common algorithm in image processing and is widely used in the field of computer vision, for example, in face recognition, image scaling, image rotation, image registration and correction, etc.
[0004] Currently, when computing devices are used to perform geometric transformations on images, pixel-by-pixel operations are performed in the underlying logic of the computing devices. This is also known as vector operations (addition, subtraction, multiplication, division, etc.) in the industry, and has low computational efficiency. Summary of the Invention
[0005] The present application provides an acceleration system for performing geometric transformation on digital images. The acceleration system performs geometric transformation on multiple digital images simultaneously, thereby improving computational efficiency compared to performing geometric transformation pixel by pixel on a single image.
[0006] In a first aspect, the present application provides an acceleration system for performing geometric transformation on a digital image, wherein the digital image includes a first original image and a second original image, and the acceleration system includes a processor and an accelerator.
[0007] The processor is used to:
[0008] Determine a target transformation matrix, where the target transformation matrix is used for geometric transformation between the first original image and the first target image, and for geometric transformation between the second original image and the second target image, wherein the first target image and the second target image have the same size;
[0009] Determining each coordinate point on the first target image, wherein each coordinate point on the second target image is the same as each coordinate point on the first target image;
[0010] Accelerators are used to:
[0011] Determine an original coordinate position matrix based on the target transformation matrix and each coordinate point on the first target image, wherein the original coordinate position matrix includes the coordinate positions of each coordinate point on the first target image on the first original image and the coordinate positions of each coordinate point on the second target image on the second original image;
[0012] According to the original coordinate position matrix, the pixel values of each coordinate point on the first original image and the pixel values of each coordinate point on the second original image, the pixel values of each coordinate point on the first target image and the pixel values of each coordinate point on the second target image are determined to obtain the first target image and the second target image.
[0013] The present application scheme is described using the example of a digital image comprising a first original image and a second original image. A target transformation matrix is calculated, and in the underlying logic of the computing device, the target transformation matrix is used to simultaneously perform geometric transformation operations on the first and second original images. This improves computational efficiency compared to performing geometric transformation operations on the original images one by one. It should be noted that digital images can include a larger number of images. In practical applications, many images can be processed in batches, each batch comprising multiple images. The target transformation matrix corresponding to a batch of images is first calculated, and then geometric transformation operations are performed on the multiple images in the batch simultaneously using the target transformation matrix. This improves computational efficiency compared to performing geometric transformation operations on the images one by one. Furthermore, performing geometric transformations on a batch of images simultaneously involves matrix or vector operations, for which accelerators are more suitable. Using an accelerator to perform matrix or vector operations allows the processor and accelerator to each utilize their respective capabilities and strengths. This shortens computation time and further improves computational efficiency for geometric transformation scenarios involving large batches of images.
[0014] Based on the first aspect, in a possible implementation, the accelerator is used to:
[0015] Determine a pixel matrix of neighboring coordinate points based on the original coordinate position matrix, the pixel values of each coordinate point on the first original image, and the pixel values of each coordinate point on the second original image, wherein the pixel matrix of neighboring coordinate points includes a matrix consisting of pixel values of neighboring coordinate points of each coordinate point on the first target image at the coordinate position on the first original image, and pixel values of neighboring coordinate points of each coordinate point on the second target image at the coordinate position on the second original image;
[0016] Determine an interpolation weight matrix, where the interpolation weight matrix includes weights of each pixel value in a pixel matrix of adjacent coordinate points;
[0017] The pixel values of each coordinate point on the first target image and the pixel values of each coordinate point on the second target image are determined according to the pixel matrix of the adjacent coordinate points and the interpolation weight matrix to obtain the first target image and the second target image.
[0018] Based on the first aspect, in a possible implementation, the accelerator includes a matrix multiplication acceleration unit and a vector operation acceleration unit, wherein:
[0019] The matrix multiplication acceleration unit is used to determine the original coordinate position matrix and the interpolation weight matrix;
[0020] The vector operation acceleration unit is used to determine the pixel matrix of adjacent coordinate points, and, based on the pixel matrix of adjacent coordinate points and the interpolation weight matrix, determine the pixel values of each coordinate point on the first target image and the pixel values of each coordinate point on the second target image to obtain the first target image and the second target image.
[0021] The accelerator includes a matrix multiplication accelerator unit (MMU) and a vector operation accelerator unit (VMU). The MMU is a parallel operation unit for matrix cross multiplications, while the VMU is a parallel operation unit for vector operations. Operations involving matrix cross multiplications are performed by the MMU, while operations involving vectors are performed by the VMU. By effectively utilizing these two units, each unit can maximize its capabilities and improve operational efficiency.
[0022] Based on the first aspect, in a possible implementation, the processor is configured to:
[0023] Determine a first transformation matrix, where the first transformation matrix is used for geometric transformation between the first original image and the first target image, and the dimension of the first transformation matrix is k*3;
[0024] Determine a second transformation matrix, where the second transformation matrix is used for geometric transformation between the second original image and the second target image, and the dimension of the second transformation matrix is k*3;
[0025] The first transformation matrix and the second transformation matrix are superimposed in the row dimension to obtain a target transformation matrix, where the dimension of the target transformation matrix is 2k*3.
[0026] It can be understood that the description here is based on the first transformation matrix and the second transformation matrix as an example. When a batch includes multiple images, the multiple images correspond to multiple transformation matrices. The target transformation matrix is determined based on the multiple transformation matrices. The target transformation matrix is used to perform geometric transformations on multiple images in a batch at the same time.
[0027] Based on the first aspect, in a possible implementation, the dimension of the matrix formed by the coordinate points on the reference image is 3*m, and the dimension of the original coordinate position matrix is 2k*m.
[0028] Based on the first aspect, in a possible implementation, the geometric transformation includes one of an affine transformation and a transmission transformation.
[0029] In a second aspect, the present application provides an acceleration method for performing geometric transformation on a digital image, wherein the digital image includes a first original image and a second original image, and the method includes:
[0030] The accelerator obtains a target transformation matrix, where the target transformation matrix is used for geometric transformation between the first original image and the first target image, and geometric transformation between the second original image and the second target image, wherein the first target image and the second target image have the same size;
[0031] Acquire each coordinate point on the first target image, wherein each coordinate point on the second target image is the same as each coordinate point on the first target image;
[0032] Determine an original coordinate position matrix based on the target transformation matrix and each coordinate point on the first target image, wherein the original coordinate position matrix includes the coordinate positions of each coordinate point on the first target image on the first original image and the coordinate positions of each coordinate point on the second target image on the second original image;
[0033] According to the original coordinate position matrix, the pixel values of each coordinate point on the first original image and the pixel values of each coordinate point on the second original image, the pixel values of each coordinate point on the first target image and the pixel values of each coordinate point on the second target image are determined to obtain the first target image and the second target image.
[0034] The present application scheme is introduced by taking the target transformation matrix as an example for the geometric transformation between the first original image and the first target image, and the geometric transformation between the second original image and the second target image. In practical applications, the target transformation matrix can be used for the geometric transformation of multiple images. Using the target transformation matrix to perform geometric transformation operations on multiple images at the same time improves computational efficiency compared to performing geometric transformation operations on the original images one by one.
[0035] Based on the second aspect, in a possible implementation, determining the pixel values of each coordinate point on the first target image and the pixel values of each coordinate point on the second target image based on the original coordinate position matrix, the pixel values of each coordinate point on the first original image, and the pixel values of each coordinate point on the second original image to obtain the first target image and the second target image includes:
[0036] Determine a pixel matrix of neighboring coordinate points based on the original coordinate position matrix, the pixel values of each coordinate point on the first original image, and the pixel values of each coordinate point on the second original image, wherein the pixel matrix of neighboring coordinate points includes a matrix consisting of pixel values of neighboring coordinate points of each coordinate point on the first target image at the coordinate position on the first original image, and pixel values of neighboring coordinate points of each coordinate point on the first target image at the coordinate position on the second original image;
[0037] Determine an interpolation weight matrix, where the interpolation weight matrix includes weights of each pixel value in a pixel matrix of adjacent coordinate points;
[0038] The pixel values of each coordinate point on the first target image and the pixel values of each coordinate point on the second target image are determined according to the pixel matrix of the adjacent coordinate points and the interpolation weight matrix to obtain the first target image and the second target image.
[0039] Based on the second aspect, in a possible implementation, the geometric transformation includes one of an affine transformation and a transmission transformation.
[0040] Based on the second aspect, before the accelerator obtains the target transformation matrix, the method further includes:
[0041] The processor determines a target transformation matrix;
[0042] The processor sends the target transformation matrix to the accelerator.
[0043] Based on the second aspect, in a possible implementation, the processor determining the target transformation matrix includes:
[0044] The processor determines a first transformation matrix, where the first transformation matrix is used for geometric transformation between the first original image and the first target image, and the dimension of the first transformation matrix is k*3;
[0045] Determine a second transformation matrix, where the second transformation matrix is used for geometric transformation between the second original image and the second target image, and the dimension of the second transformation matrix is k*3;
[0046] The first transformation matrix and the second transformation matrix are superimposed in the row dimension to obtain a target transformation matrix, where the dimension of the target transformation matrix is 2k*3.
[0047] Based on the second aspect, in a possible implementation, the dimension of the matrix formed by the coordinate points on the reference image is 3*m, and the dimension of the original coordinate position matrix is 2k*m.
[0048] In a third aspect, the present application provides an accelerator, comprising:
[0049] an acquisition unit, configured to acquire a target transformation matrix, wherein the target transformation matrix is used for geometric transformation between the first original image and the first target image, and for geometric transformation between the second original image and the second target image, wherein the first target image and the second target image have the same size;
[0050] The acquisition unit is further configured to acquire each coordinate point on the first target image, wherein each coordinate point on the second target image is the same as each coordinate point on the first target image;
[0051] The matrix multiplication acceleration unit is used to determine an original coordinate position matrix based on the target transformation matrix and each coordinate point on the first target image, wherein the original coordinate position matrix includes the coordinate positions of each coordinate point on the first target image on the first original image and the coordinate positions of each coordinate point on the second target image on the second original image;
[0052] The matrix multiplication acceleration unit and the vector operation acceleration unit are used in collaboration to determine the pixel values of each coordinate point on the first target image and the pixel values of each coordinate point on the second target image based on the original coordinate position matrix, the pixel values of each coordinate point on the first original image and the pixel values of each coordinate point on the second original image, so as to obtain the first target image and the second target image.
[0053] Based on the third aspect, in a possible implementation,
[0054] The vector operation acceleration unit is used to determine a pixel matrix of neighboring coordinate points based on the original coordinate position matrix, the pixel values of each coordinate point on the first original image, and the pixel values of each coordinate point on the second original image, wherein the pixel matrix of neighboring coordinate points includes a matrix composed of pixel values of coordinate points neighboring each coordinate point on the first target image at its respective coordinate position on the first original image, and pixel values of coordinate points neighboring each coordinate point on the second target image at its respective coordinate position on the second original image;
[0055] The matrix multiplication acceleration unit is used to determine an interpolation weight matrix, where the interpolation weight matrix includes weights of each pixel value in a pixel matrix of adjacent coordinate points;
[0056] The vector operation acceleration unit is used to determine the pixel values of each coordinate point on the first target image and the pixel values of each coordinate point on the second target image according to the pixel matrix of the adjacent coordinate points and the interpolation weight matrix, so as to obtain the first target image and the second target image.
[0057] Based on the third aspect, in a possible implementation, the geometric transformation includes one of an affine transformation and a transmission transformation.
[0058] The functional modules of the third aspect are used to implement the method described on the accelerator side in the second aspect, as well as the method described on the accelerator side in any possible implementation manner of the second aspect.
[0059] In a fourth aspect, the present application provides a host, comprising:
[0060] a processor, configured to determine a target transformation matrix, wherein the target transformation matrix is used for a geometric transformation between the first original image and a first target image, and a geometric transformation between the second original image and a second target image, wherein the first target image and the second target image have the same size;
[0061] The processor is further configured to send the target transformation matrix to the accelerator, so that the accelerator performs geometric transformation on the first original image and the second original image according to the target transformation matrix.
[0062] Based on the fourth aspect, in a possible implementation, the processor is configured to:
[0063] Determine a first transformation matrix, where the first transformation matrix is used for geometric transformation between the first original image and the first target image, and the dimension of the first transformation matrix is k*3;
[0064] Determine a second transformation matrix, where the second transformation matrix is used for geometric transformation between the second original image and the second target image, and the dimension of the second transformation matrix is k*3;
[0065] The first transformation matrix and the second transformation matrix are superimposed in the row dimension to obtain a target transformation matrix, where the dimension of the target transformation matrix is 2k*3.
[0066] Based on the fourth aspect, in a possible implementation, the dimension of the matrix formed by the coordinate points on the reference image is 3*m, and the dimension of the original coordinate position matrix is 2k*m.
[0067] The functional modules of the fourth aspect are used to implement the method described on the processor side in the second aspect, as well as the method described on the processor side in any possible implementation manner of the second aspect.
[0068] In a fifth aspect, the present application provides a host comprising a processor and a memory, the memory being used to store instructions, and the processor being used to execute the instructions stored in the memory to implement the method described on the processor side in the above-mentioned second aspect, as well as the method described on the processor side in any possible implementation of the above-mentioned second aspect.
[0069] In a sixth aspect, the present application provides an accelerator comprising a processor and a memory, the memory being used to store instructions, and the processor being used to execute the instructions stored in the memory to implement the method described on the accelerator side in the above-mentioned second aspect, as well as the method described on the accelerator side in any possible implementation of the above-mentioned second aspect.
[0070] In the seventh aspect, the present application provides a computer storage medium comprising program instructions. When the program instructions are executed by a device with computing functions, the device with computing functions executes the method described by the accelerator and / or processor in the above-mentioned second aspect, as well as the method described by the accelerator and / or processor in any possible implementation of the above-mentioned second aspect.
[0071] In an eighth aspect, the present application provides a computer program product comprising program instructions. When the program instructions are executed by a device with computing functionality, the device with computing functionality executes the method described by the accelerator and / or processor in the second aspect above, as well as the method described by the accelerator and / or processor in any possible implementation of the second aspect above.
[0072] In a ninth aspect, the present application provides a chip, which can be used to implement the method described on the accelerator side in the second aspect or any possible implementation of the second aspect. BRIEF DESCRIPTION OF THE DRAWINGS
[0073] FIG1A is a schematic diagram of a reference image given in this application;
[0074] FIG1B is a schematic diagram of an image to be processed given in this application;
[0075] FIG2 is a schematic flow chart of an acceleration method for performing geometric transformation on a digital image provided by the present application;
[0076] FIG3 is an example diagram provided by this application;
[0077] FIG4 is a flow chart of a method for calculating pixel values at non-integer coordinate positions using a bilinear interpolation method provided by the present application;
[0078] FIG5 is a schematic diagram of the structure of an acceleration system for performing geometric transformation on a digital image provided by the present application;
[0079] FIG6 is a schematic structural diagram of an accelerator provided in this application;
[0080] FIG7 is a schematic structural diagram of a device with computing function provided by the present application. DETAILED DESCRIPTION
[0081] The following is an introduction to the technical terms involved in the method embodiments of this application.
[0082] Vector operation acceleration unit refers to a parallel computing unit that specializes in vector operations.
[0083] The matrix multiplication acceleration unit refers to a parallel computing unit that can only be used to perform matrix cross multiplication operations.
[0084] The core computing power resource of the current mainstream artificial intelligence (AI) accelerator is the matrix multiplication acceleration unit, which accounts for most or even the vast majority of the AI accelerator's resources.
[0085] Many applications can be implemented through matrix cross multiplication operations, or some operations in the applications can be implemented through matrix cross multiplication, but none of them are implemented through the matrix multiplication acceleration unit. Instead, they are implemented through vector operation acceleration units using vectors, or through other forms of operations. This makes it impossible to reasonably utilize the core computing power resources in the accelerator.
[0086] This application provides a method and system for geometrically transforming multiple images based on a reference image. Before introducing the method and system provided by this application, the application scenarios involved in this application scheme are first introduced.
[0087] The present application scheme is applicable to application scenarios in which a given reference image is used to perform geometric transformations on multiple images based on the reference image, wherein the geometric transformation may be, for example, an affine transformation or a transmission transformation, and an affine transformation includes translation, rotation, scaling, shearing, and reflection. For example, in an application scenario of target recognition, such as face recognition or license plate recognition, the captured image must first be geometrically transformed into an image of a uniform format with the same size and shape as the reference image, and then the uniform format image must be further processed to achieve target recognition.
[0088] For example, referring to Figures 1A and 1B, Figure 1A is a schematic diagram of a reference image given in this application, and Image 1, Image 2, and Image 3 in Figure 1B are images to be processed, wherein Image 1, Image 2, and Image 3 are images of different sizes and with different poses of the objects in the images. It is required to perform geometric transformations on Image 1, Image 2, and Image 3 based on the given reference image to obtain a target image of the same size, shape, and placement as the reference image. The target image can also be further processed to obtain an image that meets the application requirements.
[0089] It should be noted that the different sizes of the two images refer to the different number of rows, columns, or both rows and columns that make up the two images. The different poses of the objects in the two images refer to the different poses and positions of the objects in the two images. For example, the pose of the person in image 1 is different from the pose of the person in image 2, and the poses of the person in image 1 and image 2 are different from the poses of the person in image 3, respectively.
[0090] In the application scenarios shown in Figures 1A and 1B, image 1, image 2, and image 3 can be affine transformed according to the reference image to obtain target images with the same size, shape, and image placement as the reference image, such as target image 1, target image 2, and target image 3 in Figure 1B.
[0091] Optionally, the faces in image 1, the faces in image 2, and the faces in image 3 may also be located on one image, that is, one image includes multiple faces, and any one or more of the sizes, shapes, and postures of the multiple faces may be different. The multiple faces on this one image are geometrically transformed separately to obtain a target image with the same size, shape, and posture as the reference image.
[0092] The following introduces an acceleration method for geometric transformation of digital images provided by the method of the present application. See Figure 2, which is a flow chart of an acceleration method for geometric transformation of digital images provided by the present application. The method is applied to an acceleration system for geometric transformation of digital images, and the acceleration system includes a processor and an accelerator. The method includes but is not limited to the description below.
[0093] It should be noted that the method of the present application can be used to perform geometric transformations on multiple images at the same time. However, for the convenience of describing the scheme, the following mainly introduces the example of performing geometric transformations on the first original image and the second original image at the same time. The method of performing geometric transformations on multiple images at the same time is similar to this.
[0094] S101 : Determine a target transformation matrix, where the target transformation matrix is used for geometric transformation between a first original image and a first target image, and for geometric transformation between a second original image and a second target image.
[0095] This step may be performed by a processor in the acceleration system.
[0096] First, determine the first transformation matrix and the second transformation matrix. The first transformation matrix is used for the geometric transformation between the first original image and the first target image, and the second transformation matrix is used for the geometric transformation between the second original image and the second target image. Then, superimpose the first and second transformation matrices in the row dimension to obtain the target transformation matrix. If the dimensions of the first and second transformation matrices are both k*3, the dimensions of the target transformation matrix are 2k*3. For example, if the geometric transformation is an affine transformation, k is 2, and if the geometric transformation is a transmission transformation, k is 3.
[0097] The first target image and the second target image are the same size. In one implementation, the first target image and the second target image are both images of the same size as a given reference image. In another implementation, the first target image and the second target image are the same size, but different from the reference image size. The same size of the first target image and the second target image can be specifically set according to the actual application requirements. For example, if the height of the reference image is H rows, i.e., [0…H], and the width is W columns, i.e., [0…W], in actual application, we only need to obtain the [0…H / 2] rows and [0…W / 2] columns of the target image, that is, the output image only needs the [0…H / 2] rows and [0…W / 2] columns of this area, then the size of the first target image and the second target image can be set to [0…H / 2] rows and [0…W / 2] columns. Here, [0…H / 2] rows and [0…W / 2] columns are just examples. In actual application, they can be any other values and do not constitute a limitation of this application. Taking affine transformation as an example, if the first transformation matrix is M1 and the second transformation matrix is M2, then the target transformation matrix is M, where
[0098] In the first transformation matrix, m 00 、m 01 、m 10 and m 11 Used to control the rotation and scaling of the first original image, m 02 and m 12 Used to control the translation of the first original image. Similarly, in the second transformation matrix, m 20 、m 21 、m 30 and m 31 Used to control the rotation and scaling of the second original image, m 22 and m 32 Used to control the translation of the second original image.
[0099] The first transformation matrix may be determined by calculating the matrix using a least squares method based on three or more corresponding point pairs on the reference image and the first original image. Similarly, the second transformation matrix may be determined by calculating the matrix using a least squares method based on three or more corresponding point pairs on the reference image and the second original image. The first transformation matrix and the second transformation matrix may also be calculated using other methods, and this application does not limit the calculation method of the first transformation matrix and the second transformation matrix.
[0100] In this embodiment, the first original image and the second original image are merely examples. In actual applications, geometric transformations are typically performed on a large number of images. When geometric transformations are performed on a large number of images, the large number of images can be divided into multiple batches, each batch including multiple images. The transformation matrix corresponding to each image in each batch is calculated separately, and then the target transformation matrix corresponding to each batch is determined. This method can be referred to as batch processing. If the number of original images in a batch is n and the geometric transformation is an affine transformation, the dimension of the target transformation matrix is 2n*3.
[0101] S102: Determine the coordinates of each point on the first target image.
[0102] This step may be performed by a processor in the acceleration system. Alternatively, this step may also be performed by a vector operation acceleration unit.
[0103] It can be understood that the first target image and the second target image to be obtained are images of the same size, and the positions of the coordinate points included in the first target image are the same as the positions of the coordinate points on the second target image.
[0104] Determining the coordinates of each point on the first target image effectively determines the coordinates of each point on the second target image. For each original image in a batch, the corresponding target image is the same size as the first target image. Therefore, the coordinates of each point on the target image are the same as those on the first target image.
[0105] The coordinate points on the first target image can be generated by grid generation. For example, S can be used to represent the matrix composed of the coordinate points on the first target image, where S includes src x Coordinates, src y coordinates and homogeneous coordinates 1,
[0106] The first target image includes m pixels, and the src on the same column in the S matrix x Coordinates and src y The coordinates represent the coordinates of a pixel point on the first target image, for example, srcx1 and src y1 Indicates the coordinates of a pixel on the first target image, src xm and src ym Represents the coordinates of another pixel point on the first target image. The dimension of the S matrix is 3*m, where m is the number of pixels included in the first target image or each target image, that is, the product of the width and height of the first target image.
[0107] S103 : Determine an original coordinate position matrix according to the target transformation matrix and each coordinate point on the first target image.
[0108] This step is performed by the matrix multiplication acceleration unit.
[0109] The original coordinate position matrix includes the coordinate positions of each coordinate point on the first target image on the first original image and the coordinate positions of each coordinate point on the second target image on the second original image.
[0110] The target transformation matrix is cross-multiplied with the matrix of each coordinate point on the first target image / second target image to obtain the original coordinate position matrix. Taking affine transformation as an example, if a batch includes n original images, the target transformation matrix corresponding to this batch is M 2n×3 , assuming that the first target image or each target image includes m pixels, the matrix formed by the coordinate points on the first target image is S 3×m , then the original coordinate position matrix is: D 2n×m =M 2n×3 ×S 3×m (5)
[0111] Among them, D 2n×m is the original coordinate position matrix, the dimension of the original coordinate position matrix is 2n*m, and the original coordinate position matrix includes the corresponding coordinate positions of each coordinate point on the first target image on each of the n original images in this batch.
[0112] S104. Determine the pixel values of each coordinate point on the first target image and the pixel values of each coordinate point on the second target image according to the original coordinate position matrix, the pixel values of each coordinate point on the first original image, and the pixel values of each coordinate point on the second original image, to obtain the first target image and the second target image.
[0113] After determining the coordinate positions of each coordinate point on the first target image on the first original image, the pixel values of the corresponding coordinate positions on the first original image are assigned to the corresponding coordinate points on the first target image, that is, the pixel values of each coordinate point on the first target image are determined, and the first target image is obtained. Similarly, after determining the coordinate positions of each coordinate point on the second target image on the second original image, the pixel values of the corresponding coordinate positions on the second original image are assigned to the corresponding coordinate points on the second target image, that is, the pixel values of each coordinate point on the second target image are determined, and the second target image is obtained. It should be noted that the original coordinate position matrix includes the coordinate positions of each coordinate point on the first target image on the first original image and the coordinate positions of each coordinate point on the second target image on the second original image. The pixel values of each coordinate point on the first target image and the pixel values of each coordinate point on the second target image are also represented by a matrix, and the matrix is obtained by calculation using a formula.
[0114] For each coordinate point on the first target image, if its coordinate position on the first original image is an integer, the pixel value at that integer coordinate position on the first original image can be directly assigned to the corresponding coordinate point on the first target image. For example, if the coordinate point (1,1) on the first target image is at (8,8) on the first original image, the pixel value A at (8,8) on the first original image can be directly assigned to the coordinate point (1,1) on the first target image as the pixel value of the (1,1) coordinate point.
[0115] When the coordinate positions of each point in the first target image on the first original image are non-integer, it is necessary to calculate the pixel value at that non-integer coordinate position on the first original image. This is typically done using linear interpolation. For ease of understanding, the following describes how to calculate pixel values at non-integer coordinate positions, assuming the original image is a two-dimensional image and bilinear interpolation is used as an example. The implementation steps can be found in the schematic flow chart of the method shown in Figure 4.
[0116] S1041: Determine adjacent coordinate points of non-integer coordinate positions and a pixel matrix of the adjacent coordinate points.
[0117] This step can be performed by a vector operation acceleration unit.
[0118] In bilinear interpolation, the neighboring coordinate points of a non-integer coordinate position include the four coordinate points or pixels closest to the non-integer coordinate position. For example, referring to the example diagram shown in Figure 3, the black dot represents the non-integer coordinate position, and the four coordinate points closest to the non-integer coordinate position are shown as the light-colored dots in Figure 3. x and y are the coordinate values of the horizontal and vertical coordinates of the non-integer coordinate point rounded down, respectively. ox and oy are the remainders of the horizontal and vertical coordinates of the non-integer coordinate point rounded down. The four neighboring coordinate points of the non-integer coordinate point are (x, y), (x+1, y), (x, y+1), and (x+1, y+1). The pixel values of the four neighboring coordinate points are p(x, y), p(x+1, y), p(x, y+1), and p(x+1, y+1), respectively, where p(x, y) represents the pixel value of the coordinate point (x, y) in the original image.
[0119] The pixel matrix of adjacent coordinate points refers to the matrix composed of the pixel values of adjacent coordinate points. If Q is used to represent the pixel matrix of adjacent coordinate points, then Q = [p(x,y),p(x+1,y),p(x,y+1),p(x+1,y+1)] (6)
[0120] Here we take the neighboring coordinate points of a non-integer coordinate position as an example. In practical applications, when a batch includes n original images and each original image includes m pixels, the pixel matrix of the neighboring coordinate points corresponding to the batch is Q nmC×4 , that is, the dimension of Q is nmC*4, where C is the number of channels of the original image.
[0121] S1042: Determine an interpolation weight matrix.
[0122] The step of determining the interpolation weight matrix can be performed by a matrix multiplication acceleration unit.
[0123] The interpolation weight matrix refers to a matrix composed of the weights of the pixel values of each adjacent coordinate point. In bilinear interpolation, the interpolation weight matrix can be obtained by cross-producting the bilinear interpolation factor matrix K with the bilinear interpolation distance matrix N, that is, N×K, where K is the matrix expression coefficient of the bilinear interpolation algorithm and N is the matrix expression variable of the bilinear interpolation algorithm.
[0124] Here we take a distance matrix of non-integer coordinate positions as an example. In practical applications, when a batch includes n original images and each original image includes m pixels, the distance matrix corresponding to the batch is N nm×4 , that is, the dimension of N is nm*4.
[0125] S1043 : Determine the pixel value of each coordinate point on the first target image and the pixel value of each coordinate point on the second target image according to the pixel matrix of the neighboring coordinate points and the interpolation weight matrix.
[0126] This step can be performed by a vector operation acceleration unit.
[0127] Multiply the pixel matrix of the neighboring coordinate points by the interpolation weight matrix to obtain the pixel values of the non-integer coordinate positions in the first original image and the pixel values of the non-integer coordinate positions in the second original image, that is, obtain the pixel values of the corresponding coordinate points on the first target image and the pixel values of the corresponding coordinate points on the second target image, that is, P n×m×C =reduce_sum(Q nmC×4 ·(N nm×4 ×K 4×4 ),axis=1) (9)
[0128] Here, reduce_sum is a function for calculating the sum of tensor elements, and axis=1 means performing a sum operation on the first dimension to obtain the first target image and the second target image.
[0129] Here, the first original image and the second original image are used as examples to introduce the first target image and the second target image. If a geometric transformation is performed on a batch of original images, the target images corresponding to each original image in the batch are obtained. As can be seen, the present application provides an accelerated method for geometric transformation of digital images. In this method, multiple images can be divided into multiple batches, each batch including several images, and the target transformation matrix corresponding to the batch of images is calculated. The target transformation matrix is used to perform an affine transformation on the batch of images together, and the target image corresponding to the batch of images is obtained at the same time. Compared with frame-by-frame calculations at the bottom layer of the computing device, the batch processing method increases the speed of geometric transformation of multiple images, shortens the time required for geometric transformation of multiple images, and improves computing efficiency. Multiple images are divided into multiple batches. When performing geometric transformations on the images in each batch, the matrix multiplication acceleration unit in the accelerator is used to perform operations involving matrix cross multiplications, so that the core computing power resources in the accelerator are reasonably and efficiently utilized. In the geometric transformation of images, the matrix multiplication acceleration unit and the vector operation acceleration unit are reasonably used to perform operations, which can improve computing efficiency.
[0130] This application provides an acceleration system for performing geometric transformations on digital images, as shown in Figure 5. Figure 5 is a schematic diagram of the structure of an acceleration system for performing geometric transformations on digital images provided by this application. The acceleration system includes a host and an accelerator. The accelerator can be plugged into the host in the form of a smart network card, or the accelerator can be integrated and deployed on the host. The host includes a processor, and the accelerator includes a matrix multiplication acceleration unit and a vector operation acceleration unit.
[0131] The processor is configured to group the multiple images into multiple batches, determine the target transformation matrix corresponding to each batch, and determine the coordinates of each point on the reference image. Specifically, the processor can be configured to execute steps S101 and S102 in the above method embodiment.
[0132] The matrix multiplication acceleration unit is used to determine the original coordinate position matrix according to the target transformation matrix and each coordinate point on the reference image. Specifically, the matrix multiplication acceleration unit can be used to execute step S103 in the above method embodiment.
[0133] The vector operation acceleration unit is used to determine the pixel matrix of adjacent coordinate points based on the plurality of image data. Specifically, the vector operation acceleration unit can be used to execute step S1041 in the above-described method embodiment. The matrix multiplication acceleration unit is also used to determine the interpolation weight matrix, which can be used to execute step S1042 in the above-described method embodiment. The vector operation acceleration unit is also used to determine the pixel values of each coordinate point on the first target image and the pixel values of each coordinate point on the second target image based on the pixel matrix of the adjacent coordinate points and the interpolation weight matrix, thereby obtaining target images for each batch, which can be used to execute step S1043 in the above-described method embodiment.
[0134] In this application, multiple images are divided into multiple batches. When geometric transformations are performed on the images in each batch, operations involving matrix cross multiplications are performed using the matrix multiplication acceleration unit in the accelerator, allowing the core computing power resources in the accelerator to be used rationally and efficiently. In the geometric transformation of images, the rational use of the matrix multiplication acceleration unit and the vector operation acceleration unit can improve computing efficiency.
[0135] The above describes the system structure and method embodiments. The following introduces the virtual device corresponding to the above system and method.
[0136] Referring to FIG. 6 , FIG. 6 is a schematic structural diagram of an accelerator 600 provided in this application. The accelerator 600 includes:
[0137] an acquiring unit 610, configured to acquire a target transformation matrix, wherein the target transformation matrix is used for geometric transformation between the first original image and the first target image, and for geometric transformation between the second original image and the second target image, wherein the first target image and the second target image have the same size;
[0138] The acquisition unit 610 is further configured to acquire coordinate points on the first target image / the second target image;
[0139] The matrix multiplication acceleration unit 620 is configured to determine an original coordinate position matrix based on the target transformation matrix and each coordinate point on the first target image / the second target image, wherein the original coordinate position matrix includes the coordinate positions of each coordinate point on the first target image on the first original image and the coordinate positions of each coordinate point on the second target image on the second original image.
[0140] The matrix multiplication acceleration unit 620 and the vector operation acceleration unit 630 are used in collaboration to determine the pixel values of each coordinate point on the first target image and the pixel values of each coordinate point on the second target image based on the original coordinate position matrix, the pixel values of each coordinate point on the first original image and the pixel values of each coordinate point on the second original image, so as to obtain the first target image and the second target image.
[0141] In a possible implementation,
[0142] The vector operation acceleration unit 630 is configured to determine a pixel matrix of neighboring coordinate points based on the original coordinate position matrix, the pixel values of each coordinate point on the first original image, and the pixel values of each coordinate point on the second original image, wherein the pixel matrix of neighboring coordinate points includes a matrix consisting of pixel values of each coordinate point on the reference image at its respective coordinate position on the first original image and pixel values of each coordinate point on the reference image at its respective coordinate position on the second original image.
[0143] The matrix multiplication acceleration unit 620 is used to determine an interpolation weight matrix, where the interpolation weight matrix includes the weights of each pixel value in the pixel matrix of adjacent coordinate points;
[0144] The vector operation acceleration unit 630 is used to determine the pixel values of each coordinate point on the first target image and the pixel values of each coordinate point on the second target image based on the pixel matrix of the adjacent coordinate points and the interpolation weight matrix to obtain the first target image and the second target image.
[0145] In a possible implementation, the geometric transformation includes one of an affine transformation and a transmission transformation.
[0146] Each functional module in the accelerator 600 is used to implement the method described on the accelerator side in the above method embodiment. For details, please refer to the description of the above method embodiment. For the sake of brevity, it will not be repeated here.
[0147] Each functional module in the accelerator 600 can be implemented by software or hardware. The division of each functional module is only an example. In actual applications, network devices can be divided into more or fewer functional modules, and this application does not limit this.
[0148] The present application also provides a device with computing function, as shown in Figure 7, which is a structural diagram of a device 700 with computing function provided by the present application. The device 700 with computing function can be an accelerator in the above method embodiment, used to implement the method described on the accelerator side in the above method embodiment; the device 700 with computing function can also be a physical machine, which includes a processor, used to implement the method described on the processor side in the above method embodiment; the device 700 with computing function can also be a physical machine, which includes a processor and an accelerator, and the physical machine is used to implement all the steps in the above method embodiment.
[0149] The computing device 700 includes a bus 702, a processor 704, a memory 706, and a communication interface 708. The processor 704, the memory 706, and the communication interface 708 communicate with each other via the bus 702. It should be understood that this application does not limit the number of processors and memories in the computing device 700.
[0150] Bus 702 may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, among others. Buses may be classified as address buses, data buses, control buses, and the like. For ease of illustration, FIG7 shows a single bus line, but this does not imply a single bus or type of bus. Bus 702 may include a path for transmitting information between various components of computing device 700 (e.g., memory 706, processor 704, and communication interface 708).
[0151] The processor 704 may include any one or more processors such as a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP).
[0152] The memory 706 may include volatile memory, such as random access memory (RAM). The processor 704 may also include non-volatile memory, such as read-only memory (ROM), flash memory, hard disk drive (HDD), or solid state drive (SSD).
[0153] Memory 706 stores executable code. When device 700 is configured as an accelerator in the method embodiment, processor 704 executes the executable code to implement the functions of the aforementioned acquisition unit 610, matrix multiplication acceleration unit 620, and vector operation acceleration unit 630, thereby implementing the accelerator-side method described in the accelerated method for performing geometric transformations on multiple images based on a reference image. In other words, memory 706 stores instructions for executing the accelerator-side method described in the accelerated method for performing geometric transformations on digital images.
[0154] When device 700 is configured as a physical machine, which does not include an accelerator, processor 704 executes the executable code to implement the processor functionality of the method embodiment, thereby implementing the processor-side method described in the accelerated method for performing geometric transformations on multiple images based on a reference image. Specifically, memory 706 stores instructions for executing the processor-side method described in the accelerated method for performing geometric transformations on digital images.
[0155] Optionally, when device 700 is configured as a physical machine, the physical machine includes an accelerator, and processor 704 executes the executable code to implement the functions of the processor and accelerator in the method embodiment, thereby performing all steps of the method described in the accelerated method for performing geometric transformations on multiple images based on a reference image. In other words, memory 706 stores all instructions for executing the method described in the accelerated method for performing geometric transformations on digital images.
[0156] The communication interface 708 uses a transceiver module such as, but not limited to, a network interface card or a transceiver to implement communication between the device with computing function 700 and other devices or a communication network.
[0157] The present application also provides a computer program product containing instructions. This computer program product can be software or a program product containing instructions that can be run on a device with computing capabilities or stored in any available medium. When this computer program product is run on at least one device with computing capabilities, it causes the at least one device with computing capabilities to perform the method described in the accelerated method for performing geometric transformations on digital images.
[0158] The present application also provides a computer-readable storage medium. The computer-readable storage medium can be any available medium that can be stored by a device with computing capabilities, or a data storage device such as a data center that contains one or more available media. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive). The computer-readable storage medium includes instructions that instruct the device with computing capabilities to execute the method described in the method for accelerating geometric transformation of digital images.
[0159] The present application also provides a chip that can be used to implement some or all of the steps in the above method embodiments. For example, the chip can include a matrix multiplication acceleration unit and a vector operation acceleration unit, each of which can be used to implement some of the steps in the above method embodiments.
[0160] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the protection scope of the technical solutions of the various embodiments of the present invention.
Claims
1. An acceleration system for geometric transformation of digital images, characterized in that: The digital image includes at least a first original image and a second original image, and the acceleration system includes a processor and an accelerator. The processor is used to: Determine a target transformation matrix, where the target transformation matrix is used for geometric transformation between the first original image and a first target image, and geometric transformation between the second original image and a second target image, wherein the first target image and the second target image have the same size; Determine each coordinate point on the first target image, each coordinate point on the second target image is the same as each coordinate point on the first target image; The accelerator is used to: Determine an original coordinate position matrix according to the target transformation matrix and each coordinate point on the first target image, wherein the original coordinate position matrix includes the coordinate positions of each coordinate point on the first target image on the first original image and the coordinate positions of each coordinate point on the second target image on the second original image; According to the original coordinate position matrix, the pixel values of each coordinate point on the first original image and the pixel values of each coordinate point on the second original image, the pixel values of each coordinate point on the first target image and the pixel values of each coordinate point on the second target image are determined to obtain the first target image and the second target image.
2. The acceleration system according to claim 1, characterized in that: The accelerator is used to: Determine a pixel matrix of neighboring coordinate points according to the original coordinate position matrix, the pixel values of each coordinate point on the first original image, and the pixel values of each coordinate point on the second original image, wherein the pixel matrix of neighboring coordinate points includes a matrix composed of pixel values of neighboring coordinate points of each coordinate point on the first target image at the coordinate position on the first original image, and pixel values of neighboring coordinate points of each coordinate point on the second target image at the coordinate position on the second original image; Determine an interpolation weight matrix, wherein the interpolation weight matrix includes weights of each pixel value in the pixel matrix of the neighboring coordinate points; According to the pixel matrix of the neighboring coordinate points and the interpolation weight matrix, the pixel values of each coordinate point on the first target image and the pixel values of each coordinate point on the second target image are determined to obtain the first target image and the second target image.
3. The acceleration system according to claim 2, characterized in that: The accelerator includes a matrix multiplication acceleration unit and a vector operation acceleration unit, wherein: The matrix multiplication acceleration unit is used to determine the original coordinate position matrix and determine the interpolation weight matrix; The vector operation acceleration unit is used to determine the pixel matrix of the neighboring coordinate points, and to determine the pixel values of each coordinate point on the first target image and the pixel values of each coordinate point on the second target image based on the pixel matrix of the neighboring coordinate points and the interpolation weight matrix, so as to obtain the first target image and the second target image.
4. The acceleration system according to any one of claims 1 to 3, characterized in that: The processor is used to: Determine a first transformation matrix, where the first transformation matrix is used for geometric transformation between the first original image and the first target image, and the dimension of the first transformation matrix is k*3; Determine a second transformation matrix, where the second transformation matrix is used for geometric transformation between the second original image and the second target image, and the dimension of the second transformation matrix is k*3; The first transformation matrix and the second transformation matrix are superimposed in the row dimension to obtain the target transformation matrix, where the dimension of the target transformation matrix is 2k*3.
5. The acceleration system according to claim 4, characterized in that: The dimension of the matrix formed by the coordinate points on the reference image is 3*m, and the dimension of the original coordinate position matrix is 2k*m.
6. The acceleration system according to any one of claims 1 to 5, characterized in that: The geometric transformation includes one of an affine transformation and a transmission transformation.
7. A method for accelerating geometric transformation of digital images, characterized in that: The digital image comprises at least a first original image and a second original image, and the method comprises: Acquire a target transformation matrix, where the target transformation matrix is used for geometric transformation between the first original image and a first target image, and geometric transformation between the second original image and a second target image, wherein the first target image and the second target image have the same size; Acquire each coordinate point on the first target image, wherein each coordinate point on the second target image is the same as each coordinate point on the first target image; Determine an original coordinate position matrix according to the target transformation matrix and each coordinate point on the first target image, wherein the original coordinate position matrix includes the coordinate positions of each coordinate point on the first target image on the first original image and the coordinate positions of each coordinate point on the second target image on the second original image; According to the original coordinate position matrix, the pixel values of each coordinate point on the first original image and the pixel values of each coordinate point on the second original image, the pixel values of each coordinate point on the first target image and the pixel values of each coordinate point on the second target image are determined to obtain the first target image and the second target image.
8. The method according to claim 7, characterized in that The step of determining the pixel values of each coordinate point on the first target image and the pixel values of each coordinate point on the second target image according to the original coordinate position matrix, the pixel values of each coordinate point on the first original image, and the pixel values of each coordinate point on the second original image, and obtaining the first target image and the second target image comprises: Determine a pixel matrix of neighboring coordinate points according to the original coordinate position matrix, the pixel values of each coordinate point on the first original image, and the pixel values of each coordinate point on the second original image, wherein the pixel matrix of neighboring coordinate points includes a matrix composed of pixel values of neighboring coordinate points of each coordinate point on the first target image at the coordinate position on the first original image, and pixel values of neighboring coordinate points of each coordinate point on the second target image at the coordinate position on the second original image; Determine an interpolation weight matrix, wherein the interpolation weight matrix includes weights of each pixel value in the pixel matrix of the neighboring coordinate points; According to the pixel matrix of the neighboring coordinate points and the interpolation weight matrix, the pixel values of each coordinate point on the first target image and the pixel values of each coordinate point on the second target image are determined to obtain the first target image and the second target image.
9. The method according to claim 7 or 8, characterized in that: The geometric transformation includes one of an affine transformation and a transmission transformation.
10. An acceleration device for geometric transformation of digital images, characterized in that: The digital image comprises at least a first original image and a second original image, including: an acquisition unit, configured to acquire a target transformation matrix, wherein the target transformation matrix is used for a geometric transformation between the first original image and a first target image, and a geometric transformation between the second original image and a second target image, wherein the first target image and the second target image have the same size; The acquisition unit is further used to acquire each coordinate point on the first target image, wherein each coordinate point on the second target image is the same as each coordinate point on the first target image; The matrix multiplication acceleration unit is used to determine an original coordinate position matrix according to the target transformation matrix and each coordinate point on the first target image, wherein the original coordinate position matrix includes the coordinate positions of each coordinate point on the first target image on the first original image and the coordinate positions of each coordinate point on the second target image on the second original image; The matrix multiplication acceleration unit and the vector operation acceleration unit are used in collaboration to determine the pixel values of each coordinate point on the first target image and the pixel values of each coordinate point on the second target image based on the original coordinate position matrix, the pixel values of each coordinate point on the first original image and the pixel values of each coordinate point on the second original image, so as to obtain the first target image and the second target image.
11. The device according to claim 10, characterized in that The vector operation acceleration unit is used to determine a pixel matrix of neighboring coordinate points according to the original coordinate position matrix, the pixel values of each coordinate point on the first original image and the pixel values of each coordinate point on the second original image, wherein the pixel matrix of neighboring coordinate points includes a matrix composed of pixel values of neighboring coordinate points of each coordinate point on the first target image at the coordinate position on the first original image and pixel values of neighboring coordinate points of each coordinate point on the second target image at the coordinate position on the second original image; The matrix multiplication acceleration unit is used to determine an interpolation weight matrix, where the interpolation weight matrix includes the weights of each pixel value in a pixel matrix of adjacent coordinate points; The vector operation acceleration unit is used to determine the pixel values of each coordinate point on the first target image and the pixel values of each coordinate point on the second target image according to the pixel matrix of the neighboring coordinate points and the interpolation weight matrix, so as to obtain the first target image and the second target image.
12. The device according to claim 10 or 11, characterized in that The geometric transformation includes one of an affine transformation and a transmission transformation.
13. A chip, characterized in that: The chip is used to implement the method according to any one of claims 7 to 9.
Citation Information
Patent Citations
Nearly copied image detection method based on multi-target matching
CN104766084A
A method and apparatus for determining a geometric transformation relationship between images
CN109242892A
Video processing method and related equipment thereof
CN115546043A
Image difference comparison method and device
CN116012617A
Template orientation estimation device, method, and program
US20210217196A1