Efficient image translation registration method and device
By using CUDA acceleration and optimization technical means in image translation registration, the problems of insufficient processing speed, low accuracy and poor robustness in the prior art are solved, and efficient, accurate and robust image registration effects are achieved.
Patent Information
- Application Number
- CN202510250118.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-04
- Publication Date
- 2025-06-27
AI Technical Summary
The prior art has problems such as insufficient processing speed, low accuracy and poor robustness in image translation registration, especially in real-time processing of high-resolution images and applications in complex contexts.
Using the image translation registration method based on CUDA acceleration, a masked image is generated through preprocessing, an affine transformation matrix based on translation parameters is constructed, a gradient map is used to calculate the difference and normalize iteratively, the loss function is optimized, and the translation parameters are iteratively updated through step size adjustment.
It significantly improves the processing speed and accuracy of image registration, enhances the robustness of the method, and can more effectively process high-resolution images and complex backgrounds, meeting the needs of real-time applications.
Smart Images

Figure CN120219448A_ABST
Abstract
Description
Technical Field
[0001] It relates to the field of processing technology, especially an image translation registration method based on CUDA acceleration. Background Art
[0002] Image registration technology is an important technology in the fields of computer vision and image processing, and is widely used in many fields such as medical image analysis, remote sensing image processing, computer vision, industrial inspection, etc. The basic task of image registration is to align images taken from different sources or at different times for further analysis, such as fusion, difference analysis, 3D reconstruction, etc.
[0003] Research Status of the Existing Technology
[0004] Image registration methods can be roughly divided into feature-based methods, region-based methods, and pixel-based methods. Feature-based image registration methods usually match by extracting key points or features in the image. Common feature extraction methods include SIFT (Scale-Invariant Feature Transform), SURF (Speeded-Up Robust Features), ORB (Oriented FAST and Rotated BRIEF), etc. These methods usually have high robustness and accuracy and are suitable for some images with complex structures. However, feature matching methods are sensitive to noise and occlusion and have relatively slow processing speeds.
[0005] Region-based image registration methods perform registration by calculating the similarity between image grayscale or color values. Common algorithms include the mutual information method and the mean squared error method (MSE). These methods are suitable for images with fewer feature points but have high requirements for geometric transformation of the image.
[0006] With the development of computer hardware, especially GPU acceleration technology in the field of image processing, image registration based on CUDA (Compute Unified Device Architecture) has become a common solution. GPU acceleration can significantly improve the processing speed of image registration algorithms, especially for real-time processing of high-resolution images. For example, the CUDA programming platform launched by NVIDIA greatly improves the computing speed through parallel computing, enabling image registration to be completed in a shorter time, thus meeting the needs of real-time applications.
[0007] However, although existing GPU-accelerated image registration methods can achieve significant improvements in speed, there are still certain limitations. First, traditional CPU-based methods usually cannot meet the real-time processing requirements of high-resolution images. Especially in scenarios such as industrial inspection that have high requirements for speed and accuracy, there are still problems of low efficiency and long response time. Second, although existing GPU-accelerated methods have improved in processing speed, there is still room for improvement in terms of accuracy and robustness. Especially for images with noise and complex backgrounds, the matching accuracy of existing methods is still insufficient to meet the high requirements in practical applications.
[0008] Technical problems in the prior art
[0009] The following are several main problems in the prior art regarding image translation registration:
[0010] Insufficient processing speed: Although GPU-accelerated methods have improved the calculation speed to a certain extent, in the real-time processing scenario of high-resolution images, there is still a problem that the calculation speed is not fast enough, especially in applications that require large-scale data processing.
[0011] Low accuracy: Existing image registration methods have insufficient registration accuracy when processing images with noise, complex backgrounds, or reflective light sources, etc. Registration errors are likely to occur and cannot meet the requirements of high-precision applications.
[0012] Poor robustness: For complex backgrounds or low-contrast regions, existing registration methods are often sensitive to noise and image changes, resulting in unsatisfactory image matching effects. Summary of the invention
[0013] To solve the technical defects in the prior art, namely the deficiencies in processing speed, accuracy, and robustness in the existing image translation registration, the technical solution provided by the present invention is as follows:
[0014] An efficient image translation registration method, including:
[0015] A step of preprocessing the image to be registered;
[0016] A step of generating a mask image using an annotation file based on the contour of the image workpiece to define the registration area;
[0017] A step of constructing an affine transformation matrix based on translation parameters and performing an affine transformation on the gradient maps of the template and the target image using this matrix;
[0018] A step of calculating the difference between the template gradient map after affine transformation and the target gradient map, restricting the calculation area to the pixels within the mask range, and normalizing the error;
[0019] Steps of using CUDA parallel computing to accelerate the calculation process, optimizing the loss function and iteratively updating the translation parameters by step size adjustment until the loss value reaches a predetermined threshold or the maximum number of iterations is reached;
[0020] Steps of outputting the optimal translation parameters to achieve translational registration of the image.
[0021] Furthermore, a preferred embodiment is provided, wherein the preprocessing includes converting the image into a grayscale image, performing threshold processing to remove noise and enhance edges, calculating the horizontal and vertical gradients of the image using the Sobel operator, and performing threshold processing and morphological dilation operations on the gradient map to strengthen edge information.
[0022] Furthermore, a preferred embodiment is provided, wherein the preprocessing includes using Gaussian blur to smooth the gradient map and downsampling the gradient map to improve the processing speed.
[0023] Furthermore, a preferred embodiment is provided, using the labelme tool to draw the workpiece contour and generate a JSON file, thereby defining the region to be registered and generating the mask image.
[0024] Furthermore, a preferred embodiment is provided, using the inverse Newton method to calculate the Hessian matrix and adjust the step size coefficient.
[0025] Furthermore, a preferred embodiment is provided, wherein the maximum number of iterations is 40, the maximum loss value is 1e8, the step size coefficient is 0.1, and the step size decay rate is 0.95.
[0026] Based on the same inventive concept, the present invention also provides an efficient image translational registration device, including:
[0027] A module for preprocessing the image to be registered;
[0028] A module for generating a mask image using the annotation file based on the workpiece contour of the image to define the registration region;
[0029] A module for constructing an affine transformation matrix based on the translation parameters and performing an affine transformation on the gradient maps of the template and the target image using this matrix;
[0030] A module for calculating the difference between the template gradient map after affine transformation and the target gradient map, restricting the calculation region to the pixels within the mask range, and normalizing the error;
[0031] A module for using CUDA parallel computing to accelerate the calculation process, optimizing the loss function and iteratively updating the translation parameters by step size adjustment until the loss value reaches a predetermined threshold or the maximum number of iterations is reached;
[0032] A module for outputting the optimal translation parameters to achieve translational registration of the image.
[0033] Based on the same inventive concept, the present invention also provides a computer storage medium for storing a computing program, and when the computer program is read by a computer, the computer executes the method described above.
[0034] Based on the same inventive concept, the present invention also provides a computer, including a processor and a storage medium, and when the processor reads the computer program stored in the storage medium, the computer executes the method described above.
[0035] Based on the same inventive concept, the present invention also provides a computer program product. As a computer program, when the computer program is executed, the method described above is implemented.
[0036] Compared with the prior art, the beneficial effects of the technical solution provided by the present invention are as follows:
[0037] In the present invention, the proposed image translation registration method based on CUDA acceleration effectively solves the deficiencies in the prior art through optimized steps and technical means, bringing the following effects.
[0038] First, the use of CUDA parallel computing accelerates the image registration process and significantly improves the processing speed. Traditional CPU computing methods are slow in processing high-resolution images, especially in application scenarios such as industrial inspection and real-time image processing, often unable to meet the requirements. However, the present invention performs parallel computing on the GPU and utilizes CUDA stream optimization, enabling image registration to be completed in a shorter time, greatly improving the real-time performance and processing efficiency. This effect is particularly prominent in practical applications, especially in the processing of high-resolution images, showing obvious advantages compared with traditional CPU-based computing methods.
[0039] Second, the adoption of an optimized loss function and anti-noise processing technology significantly improves the registration accuracy. Existing image registration methods, such as region registration based on mutual information method and mean square error method, often have insufficient accuracy when there is noise and complex background in the image. However, the present invention calculates the loss using the gradient map and performs multiple preprocessing and anti-noise processing on the image, effectively reducing the interference of noise on the registration result and ensuring high-precision registration in complex backgrounds. This optimization enables the registration process to handle more complex actual image data and improves the accuracy, especially in applications with high accuracy requirements such as industrial inspection and medical image analysis, being superior to traditional methods.
[0040] In addition, restricting the calculation range using the mask area enhances the robustness of the method. In traditional image registration methods, noise and background complexity have a greater impact on the registration effect. Especially in images with reflective light sources or low-contrast regions, the matching accuracy will decrease significantly. However, in the present invention, by generating and applying a mask image, the registration area is restricted within the workpiece contour, avoiding the influence of unnecessary areas on the registration result. Through the mask, the calculation of the loss function only considers the key information area, effectively reducing noise interference and improving the robustness of registration. In the face of actual scenarios with complex backgrounds and a large amount of noise, this method makes the registration effect more stable and reliable.
[0041] These effects together make the image translation registration method of the present invention have obvious advantages over the prior art in terms of speed, accuracy, and robustness. It is particularly suitable for real-time processing and industrial applications of high-resolution images, making up for the deficiencies of traditional methods.
[0042] It can be used in scenarios that require high-precision images and real-time performance, such as medical image analysis, remote sensing image processing, and industrial inspection. Description of the Drawings
[0043] Figure 1 Schematic diagram of the workpiece to be measured;
[0044] Figure 2 Schematic diagram of the registration effect of the workpiece to be measured. Detailed Embodiments
[0045] To make the advantages and beneficial effects of the technical solution provided by the present invention more clearly demonstrated, the technical solution provided by the present invention will be further described in detail below with reference to the accompanying drawings. Specifically:
[0046] Embodiment 1. This embodiment provides an efficient image translation registration method, including:
[0047] The step of preprocessing the image to be registered;
[0048] The step of generating a mask image using the annotation file based on the workpiece contour of the image to limit the registration area;
[0049] The step of constructing an affine transformation matrix based on the translation parameters and using this matrix to perform an affine transformation on the gradient maps of the template and the target image;
[0050] The step of calculating the difference between the template gradient map after affine transformation and the target gradient map, restricting the calculation area to the pixels within the mask range, and normalizing the error;
[0051] The step of using CUDA parallel computing to accelerate the calculation process, optimizing the loss function, and iteratively updating the translation parameters by adjusting the step size until the loss value reaches a predetermined threshold or the maximum number of iterations is reached;
[0052] Steps to output the optimal translation parameters to achieve image translation registration.
[0053] Specifically:
[0054] In this embodiment, a CUDA-accelerated image translation registration method and its system are provided, which combines multiple optimization techniques for image processing and solves the deficiencies of the prior art in terms of real-time performance, accuracy, and robustness. The following are the specific implementation steps of this solution, showing the processing process of each step and the relationship between each step.
[0055] Step 1: Image preprocessing
[0056] Image preprocessing is the first step in this embodiment. The main purpose is to provide higher-quality image data for the subsequent registration process, remove noise, enhance edge information, and optimize the processing speed.
[0057] Grayscale processing: Convert the template image and the target image to be registered into grayscale images respectively, simplify the calculation amount, and prepare for subsequent gradient calculation.
[0058] Threshold processing: Perform threshold processing on the grayscale image, set a reasonable threshold (for example, set pixel values greater than 72 to 72) to remove noise and enhance the edge information of the image. This step helps to eliminate the interference of high-reflection areas, especially in high-reflection applications such as stamping part detection.
[0059] Sobel operator to calculate gradients: Use the Sobel operator to calculate the horizontal and vertical gradient maps of the image. The gradient map can reflect the edge information in the image and provide a basis for subsequent registration calculations.
[0060] Threshold processing of the gradient map: Perform threshold processing on the calculated gradient map to remove the noise part with low gradient values and improve the registration accuracy.
[0061] Morphological dilation operation: Perform morphological dilation operation on the gradient map to enhance the edge information, thereby improving the accuracy of subsequent registration.
[0062] Gaussian blur: Perform Gaussian blur on the dilated gradient map to further smooth the image and reduce the influence of noise.
[0063] Downsampling: Perform downsampling operation on the blurred gradient map to reduce the computational complexity and improve the processing speed.
[0064] The output of these preprocessing steps is a gradient map that has been denoised, edge-enhanced, and downsampled, which is used as the input for the next mask generation.
[0065] Step 2: Mask generation
[0066] The purpose of mask generation is to delimit the registration area, ensuring that the registration process is carried out only within the target area, thereby improving the efficiency and robustness of registration.
[0067] Workpiece contour annotation: Use the labelme tool to draw the contour of the workpiece in the template image and generate the corresponding JSON file, which contains the area in the image that needs to be registered.
[0068] Generate the mask image: Generate the mask image according to the JSON file. The mask image defines the area to be registered. The valid area in the mask image corresponds to the contour of the workpiece, and the other parts are invalid areas, avoiding background interference.
[0069] Mask image downsampling: Downsample the generated mask image to ensure that the size of the mask image is consistent with that of the gradient image for subsequent calculations.
[0070] The output of mask generation is the mask image, which delimits the registration area and is used as the restricted area in subsequent loss calculations.
[0071] Step 3: Construction of the loss function
[0072] The loss function is the core calculation part of this embodiment, defining the matching degree between the template image and the target image, and optimizing the translation parameters of image registration by minimizing the loss value.
[0073] Construction of the affine transformation matrix: According to the current translation parameters, construct a 2x3 affine transformation matrix. This matrix is used to perform a translation transformation on the gradient image of the template image.
[0074] Affine transformation: Use the affine transformation matrix to perform a translation transformation on the gradient image of the template image, and calculate the difference between the transformed gradient image and the gradient image of the target image.
[0075] Calculate the loss function: Calculate the error between the template image and the target image, restricting the calculation area to the valid area marked in the mask image. By calculating the differences in the horizontal and vertical gradient images, the total error is obtained.
[0076] Error normalization: Normalize the error to eliminate the influence of image size and area, making the values of the loss function consistent and comparable.
[0077] The output of this step is the value of the loss function, representing the matching degree between the template image and the target image under the current translation parameters.
[0078] Step 4: Optimization of translation parameters
[0079] Optimization of translation parameters is achieved by continuously adjusting the translation parameters to minimize the loss function, thereby realizing the precise registration of the image.
[0080] Step size adjustment: Based on the value of the loss function, calculate the Hessian matrix to adjust the step size, ensuring that the update of the translation parameters is neither too large nor too small. The process of step size adjustment uses the inverse Newton method to improve the convergence speed and accuracy.
[0081] Iterative optimization: Through multiple iterations, continuously adjust the translation parameters. Calculate the new loss function value after each iteration until the loss value reaches the predetermined minimum value or the maximum number of iterations (such as 40 times) is reached. In each iteration, the step size decay coefficient is 0.95 to ensure gradual convergence.
[0082] Optimal parameter output: After multiple iterations of optimization, output the optimal translation parameters, representing the best matching position between the template image and the target image.
[0083] The optimized translation parameters are the final result of image registration and will be used for actual image transformation.
[0084] Step 5: Image registration
[0085] Finally, use the optimal translation parameters to perform a translation transformation on the template image and accurately register the template image onto the target image.
[0086] Affine transformation application: Use the optimal translation parameters to transform the template image to the position of the target image through an affine transformation matrix to obtain the finally registered image.
[0087] Registration result output: Compare the finally registered image with the target image and output the registration effect diagram to ensure the matching accuracy and quality of the image.
[0088] Embodiment 2: This embodiment further limits the efficient image translation registration method provided in Embodiment 1. The preprocessing includes converting the image to a grayscale image, performing threshold processing to remove noise and enhance edges, using the Sobel operator to calculate the horizontal and vertical gradients of the image, and performing threshold processing and morphological dilation operations on the gradient image to strengthen edge information.
[0089] Embodiment 3: This embodiment further limits the efficient image translation registration method provided in Embodiment 1. The preprocessing includes using Gaussian blur to smooth the gradient image and downsampling the gradient image to improve the processing speed.
[0090] Embodiment 4: This embodiment further limits the efficient image translation registration method provided in Embodiment 1. Use the labelme tool to draw the workpiece contour and generate a JSON file to define the area to be registered and generate the mask image.
[0091] Embodiment 5. This embodiment further limits the efficient image translation registration method provided in Embodiment 1, and uses the inverse Newton method to calculate the Hessian matrix and adjust the step size coefficient.
[0092] Embodiment 6. This embodiment further limits the efficient image translation registration method provided in Embodiment 1. The maximum number of iterations is 40, the maximum loss value is 1e8, the step size coefficient is 0.1, and the step size decay rate is 0.95.
[0093] Embodiment 7. This embodiment provides an efficient image translation registration device, including:
[0094] A module for preprocessing the image to be registered;
[0095] A module for generating a mask image using an annotation file based on the image workpiece contour to define the registration area;
[0096] A module for constructing an affine transformation matrix based on the translation parameters and performing an affine transformation on the gradient maps of the template and the target image using this matrix;
[0097] A module for calculating the difference between the template gradient map after affine transformation and the target gradient map, restricting the calculation area to the pixels within the mask range, and normalizing the error;
[0098] A module for using CUDA parallel computing to accelerate the calculation process, optimizing the loss function, and iteratively updating the translation parameters by step size adjustment until the loss value reaches a predetermined threshold or the maximum number of iterations is reached;
[0099] A module for outputting the optimal translation parameters to achieve image translation registration.
[0100] Embodiment 8. This embodiment provides a computer storage medium for storing a calculation program. When the computer program is read by a computer, the computer executes the method provided in Embodiment 1.
[0101] Embodiment 9. This embodiment provides a computer, including a processor and a storage medium. When the processor reads the computer program stored in the storage medium, the computer executes the method provided in Embodiment 1.
[0102] Embodiment 10. This embodiment provides a computer program product. As a computer program, when the computer program is executed, it implements the method provided in Embodiment 1.
[0103] Embodiment 11. In combination with Figure 1-2 This embodiment is described. Through specific embodiments, the above-provided technical solutions are further described in detail. Specifically, it includes:
[0104] 1. Image preprocessing (applicable to template images and target images):
[0105] Convert the image to grayscale.
[0106] Perform thresholding, setting pixel values greater than the threshold to the threshold value to remove noise and enhance edges.
[0107] Preferred implementation: For the high-reflectivity problem in stamping part detection, pixel values greater than 72 are set to 72.
[0108] Use the Sobel operator to calculate the horizontal and vertical gradients of the image.
[0109] Perform thresholding on the gradient image to remove pixels less than the threshold, further reducing noise.
[0110] Use morphological dilation operations to strengthen edge information.
[0111] Perform Gaussian blur on the dilated image to smooth the gradient image.
[0112] Perform 2x downsampling on the gradient image to improve processing speed.
[0113] 2. Mask generation:
[0114] Use the labelme tool to draw the workpiece contour of the template image and generate a JSON file.
[0115] Generate a mask image based on the JSON file to define the registration area.
[0116] Perform downsampling on the mask image to generate a mask of the same scale as the gradient image.
[0117] 3. Loss function:
[0118] Construct a 2x3 affine transformation matrix based on the image translation parameters.
[0119] Use the affine transformation matrix to transform the horizontal and vertical gradient images of the mask and the template.
[0120] Calculate the difference between the template gradient image after affine transformation and the target gradient image, restricting the calculation area to the pixels within the mask range.
[0121] Take the absolute value of the horizontal and vertical gradient differences respectively and sum them to obtain the total error.
[0122] Normalize the error and standardize it according to the image area.
[0123] 4. Image registration:
[0124] Preferred implementation parameters: Set the maximum number of iterations to 40, the maximum loss value to 1e8, the minimum threshold norm length to 1, the step coefficient to 0.1, and the step decay rate to 0.95.
[0125] Calculate the loss value under the current translation parameters. If it is less than the maximum loss value, update the maximum loss value and the optimal parameters.
[0126] Make an incremental adjustment to the current translation parameters with an increment value of 1, and calculate the displacement losses in the horizontal and vertical directions respectively.
[0127] Solve the parameter transformation through the Jacobian matrix.
[0128] Use the inverse Newton method to calculate the Hessian matrix and adjust the step size to update the translation parameters.
[0129] Repeat the above steps until the predetermined number of iterations is reached or the norm length of the parameter update is less than the set threshold. Output the optimal parameters.
[0130] Result output:
[0131] Multiply the optimal parameters by 4 to construct a 2x3 affine transformation matrix.
[0132] If it is the registration of the template image to the target image, perform an affine transformation using the above transformation matrix. If it is the registration of the target image to the template image, perform an affine transformation using the inverse of the above matrix.
[0133] The above further describes the technical solutions provided by the present invention through several specific implementation manners to highlight the advantages and beneficial effects of the technical solutions provided by the present invention. However, the above several specific implementation manners are not used as limitations to the present invention. Any reasonable modifications and improvements, combinations of implementation manners, and equivalent replacements based on the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. An efficient image translation registration method, characterized in that: include: A step of preprocessing the image to be registered; The step of generating a mask image using an annotation file based on an image artifact contour to define a registration region; The step of constructing an affine transformation matrix based on the translation parameters, and using the matrix to perform affine transformation on the gradient maps of the template and the target image; The step of calculating the difference between the template gradient map after affine transformation and the target gradient map, limiting the calculation area to pixels within the mask range, and normalizing the error; Use CUDA parallel computing to accelerate the calculation process, optimize the loss function and iteratively update the translation parameters by adjusting the step size until the loss value reaches a predetermined threshold or reaches the maximum number of iterations; The step of outputting optimal translation parameters to achieve translation alignment of the images.
2. The efficient image translation registration method according to claim 1, characterized in that: The preprocessing includes converting the image into a grayscale image, performing threshold processing to remove noise and enhance edges, using the Sobel operator to calculate the horizontal and vertical gradients of the image, and performing threshold processing and morphological dilation operations on the gradient image to enhance edge information.
3. The efficient image translation registration method according to claim 1, characterized in that: The preprocessing includes smoothing the gradient map using Gaussian blur and downsampling the gradient map to increase processing speed.
4. The efficient image translation registration method according to claim 1, characterized in that: The labelme tool is used to draw the workpiece outline and generate a JSON file, thereby defining the area to be registered and generating the mask image.
5. The efficient image translation registration method according to claim 1, characterized in that: Use the inverse Newton method to calculate the Hessian matrix and adjust the step size coefficient.
6. The efficient image translation registration method according to claim 1, characterized in that: The maximum number of iterations is 40, the maximum loss value is 1e8, the step coefficient is 0.1, and the step decay rate is 0.
95.
7. An efficient image translation registration device, characterized in that: include: A module for preprocessing the image to be registered; A module that generates a mask image using an annotation file based on the contours of the image artifacts to define the registration area; A module that constructs an affine transformation matrix based on translation parameters and uses the matrix to perform affine transformation on the gradient maps of the template and target image; A module that calculates the difference between the template gradient map after affine transformation and the target gradient map, limits the calculation area to pixels within the mask range, and normalizes the error; A module that uses CUDA parallel computing to accelerate the calculation process, optimizes the loss function, and iteratively updates the translation parameters by adjusting the step size until the loss value reaches a predetermined threshold or the maximum number of iterations is reached; A module that outputs the optimal translation parameters to achieve translational registration of images.
8. A computer storage medium for storing a computing program, characterized in that: When the computer program is read by a computer, the computer executes the method of claim 1 .
9. A computer, comprising a processor and a storage medium, characterized in that: When the processor reads the computer program stored in the storage medium, the computer executes the method of claim 1 .
10. A computer program product, being a computer program, characterized in that When the computer program is executed, the method of claim 1 is implemented.
Citation Information
Patent Citations
GPU-based image fast registration method
CN108629798A
Digital image registration method based on transformation increment
CN109859252A
Mask generation method, image registration method, computer equipment and storage medium
CN114332223A
Automated Image Registration With Varied Amounts of a Priori Information Using a Minimum Entropy Method
US20130077891A1