Large-range geometric misalignment multi-focus image sequence registration and fusion method and system
Patent Information
- Application Number
- CN202610606213.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-06
- Publication Date
- 2026-09-01
AI Technical Summary
[0008]针对现有技术的不足,本发明的目的是提供一种大范围几何错位多焦点图像序列配准与融合方法及系统,用于解决现有技术中仅依赖后端算法时采集过程不可控、图像序列存在大位移、旋转、尺度变化及视场变化时难以准确生成全焦融合图像的问题;通过线性位移平台实现沿光轴方向的稳定步进采集,并通过多尺度级联配准、聚焦检测和分层融合生成全焦图像,从而提高系统整体的成像清晰度、重复性和适用性
[0052]1.本发明将可控位移采集装置(电动滑台、运动控制器、触发协同机制)与后端图像配准融合算法有机结合,通过稳定的步进采集保证源图像质量,通过预设稳定时间和隔离供电减少机械振动与电磁干扰,为后续配准提供高一致性的输入,实现了硬件采集与算法处理深度协同。
Smart Images

Figure CN122675652A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of image processing technology, specifically to a method and system for registration and fusion of multifocal image sequences with large-scale geometric misalignment. Background Technology
[0002] Multifocal image fusion is a common image processing technique in computational photography and machine vision. Its purpose is to fuse multiple source images acquired at different focal planes into a single, clear image in all-focus. Existing multifocal image fusion methods mainly include transform domain methods, spatial domain methods, and deep learning-based methods.
[0003] However, existing technologies have the following shortcomings:
[0004] First, the acquisition process is disconnected from the post-processing algorithm. Existing solutions typically focus on the image post-processing algorithm itself, with insufficient description of the acquisition device, displacement control method, and acquisition triggering coordination relationship for multifocal image sequences. When the motion accuracy, dwell stability, or step size consistency of the acquisition platform is insufficient, significant translation, rotation, scale changes, and field of view changes will occur between source images, thereby increasing the difficulty of subsequent registration and fusion.
[0005] Second, registration is difficult under large-scale geometric misalignment. In multi-focus image fusion, most features are located in the sharp focus area, while a few are in the blurred area, and there are only a few common feature pairs between the reference and target images. When the camera moves along the optical axis to acquire data, perspective breathing effect is inevitably introduced. As the camera or object moves closer, the object's scale changes dynamically, leading to complex geometric misalignment. Existing registration methods are prone to getting trapped in local extrema when the maximum pixel displacement far exceeds the linear capture range of conventional algorithms, making it difficult to achieve global convergence.
[0006] Third, the fusion results contain ghosting and artifacts. When processing image sequences with large-scale geometric misalignment, existing fusion algorithms are prone to blurring, ghosting, or discontinuity at the focus edges, and the transition at the boundary between the focused and defocused areas is unnatural, affecting the visual quality of the full-focus image.
[0007] Therefore, it is necessary to provide a systematic solution that combines a controllable displacement acquisition device with an image registration and fusion algorithm to solve the problems in the existing technology where the acquisition process is uncontrollable when relying solely on the back-end algorithm, and where it is difficult to accurately generate full-focus fused images when the image sequence has large displacement, rotation, scale changes, and field of view changes. Summary of the Invention
[0008] To address the shortcomings of existing technologies, the present invention aims to provide a method and system for registration and fusion of multifocal image sequences with large-scale geometric misalignment. This method solves the problems in existing technologies where the acquisition process is uncontrollable when relying solely on backend algorithms, and where it is difficult to accurately generate full-focus fused images when image sequences exhibit large displacement, rotation, scale changes, and field-of-view variations. By using a linear displacement platform to achieve stable step acquisition along the optical axis, and by generating full-focus images through multi-scale cascaded registration, focus detection, and layered fusion, the overall imaging clarity, repeatability, and applicability of the system are improved.
[0009] The technical solution adopted by this invention to solve its technical problem is:
[0010] A method for registration and fusion of multifocal image sequences with large-scale geometric misalignment is presented, with the following specific steps.
[0011] Step S1, Image Registration: Acquire a multifocal image sequence acquired along the optical axis, select the last frame image with the smallest field of view as the global reference image; construct a multi-scale image representation, estimate the geometric transformation matrix between adjacent frames sequentially from coarse to fine scale, and map all source images to the coordinate system of the global reference image through the concatenation of the transformation matrices of adjacent frames, and output the registered image sequence.
[0012] Step S2, Group coarse fusion: Divide the registered long sequence of images into several image groups, calculate the average image within each group, calculate the correlation coefficient between each source image within the group and the average image of the group based on the local cross-correlation method, determine the best focusing source image for each pixel position, and generate coarse fused images for each group.
[0013] Step S3, Dual Reference and Edge Refinement Fusion: Using the coarse fused image as input, an initial decision map is constructed based on the first reference image through local correlation measurement; edge-preserving filtering and image matting algorithms are used sequentially to perform consistency optimization and boundary refinement on the decision map; weighted fusion is performed on the input image based on the optimized decision map to obtain the initial full-focus image;
[0014] Step S4, Reference Image Update and Iterative Optimization: The initial full-focus image is used as the second reference image to replace the first reference image. Step S3 is repeated, and a smaller local correlation coefficient calculation window is used to generate the final full-focus fused image.
[0015] As a preferred embodiment, a further technical solution of the present invention is:
[0016] Preferably, in step S1, acquiring a multifocal image sequence along the optical axis specifically includes: driving the image acquisition device or the object under test to make a step-like relative displacement along the imaging optical axis by an electric slide; the motion controller controls the slider of the electric slide to move according to a preset step distance and dwell time, triggering the image acquisition device at each focal plane position, continuously acquiring multiple frames of images covering different depths of field, forming a large-scale geometric misalignment multifocal image sequence with translation, rotation, scale changes and field of view changes.
[0017] Preferably, in step S1, constructing a multi-scale image representation involves estimating the geometric transformation matrix between adjacent frames sequentially from coarse to fine scales. Specifically, this includes...
[0018] S101: Construct a multi-scale image pyramid containing multiple scale layers from the lowest resolution to the original resolution, with each scale layer including at least three scales;
[0019] S102: The registration process starts from the lowest resolution coarse-scale layer, at which the enhancement correlation coefficient between the reference image and the target image is calculated. The enhancement correlation coefficient is maximized through iterative optimization, and the affine transformation matrix between adjacent frames is estimated as the initial estimate for global alignment.
[0020] S103: Upsample the initially estimated transformation matrix layer by layer and pass it to the next higher resolution scale layer as the starting point for iterative optimization of that layer;
[0021] S104: Perform sub-pixel level fine-tuning of the transformation matrix at the highest resolution original resolution layer, and output the optimal geometric transformation matrix between adjacent frames.
[0022] Preferably, in step S3, using the coarsely fused image as input, an initial decision map is constructed based on the first reference image through local correlation measurement, specifically including:
[0023] S301: Select two adjacent coarse fusion images from the multiple coarse fusion images output in step S2 according to the acquisition order or spatial adjacency relationship and fuse them one by one.
[0024] S302: For the two currently selected adjacent coarsely fused images, calculate the pixel-wise grayscale average value of the two, construct an average value image, and use this average value image as the first reference image;
[0025] S303: Calculate the local cross-relationship graphs between the two coarsely fused images and the first reference image respectively;
[0026] S304: At the same pixel location, compare the local cross-correlation coefficients of the two coarsely fused images and determine the coarsely fused image with the larger coefficient as the source of focus at that pixel location;
[0027] S305: Construct an initial decision map for the two coarsely fused images based on the determination result. The initial decision map is a matrix with the same size as the image. The value of each pixel position is used to identify which coarsely fused image the pixel should be taken from: if the local cross-correlation coefficient of the first coarsely fused image is greater than that of the second image, the first identifier value is taken at the corresponding position in the initial decision map; otherwise, the second identifier value is taken.
[0028] Preferably, in step S3, edge-preserving filtering and image matting algorithms are sequentially used to perform consistency optimization and boundary refinement on the decision graph, specifically including:
[0029] S306: The initial decision map is spatially consistent with the source image by using guided filtering. The source image is used as the guide image to maintain the alignment between the decision map and the boundaries of salient objects in the source image while eliminating noise.
[0030] S307: Set a threshold for the optimized decision map to generate a three-valued Trimap. The Trimap divides the image region into foreground region, background region and unknown region.
[0031] S308: Input the Trimap and the source image into the closed matting algorithm, and solve for the optimal alpha mask by minimizing the energy function, which includes Laplacian matrix terms and diagonal matrix constraint terms;
[0032] S309: Threshold the continuous alpha values obtained from the solution to generate a refined decision map with accurate boundary positioning.
[0033] Preferably, in step S3, weighted fusion is performed on the input image based on the optimized decision map, specifically including:
[0034] The refined decision map obtained after edge-preserving filtering and image matting algorithm optimization is used as the fusion weight;
[0035] The two coarsely merged images are then weighted at the pixel level using the following formula to generate an initial full-focus image. :
[0036] ;
[0037] in, For refined decision-making diagrams; Source image one; Image 2 is the source image.
[0038] This invention also discloses a system for registering and fusing multifocal image sequences with large-scale geometric misalignment, comprising,
[0039] An electric slide is used to support an image acquisition device or the object being measured and to achieve linear motion along the optical axis.
[0040] The motion controller, electrically connected to the electric slide table, is used to send motion control signals to the electric slide table;
[0041] An image acquisition device is used to acquire images in response to a trigger signal;
[0042] The processor is communicatively connected to both the motion controller and the image acquisition device. The processor sends motion control commands to the motion controller and trigger signals to the image acquisition device. The processor integrates an image processing module, which is used to perform registration and full-focus fusion methods for large-scale geometrically misaligned multifocal image sequences and outputs a full-focus fused image.
[0043] The power supply module is electrically connected to the electric slide, motion controller, image acquisition device and processor respectively, and is used to supply power.
[0044] Preferably, the motion controller is a pulse direction type controller, and the electric slide includes a stepper motor, a lead screw transmission mechanism, a linear guide rail and a slider; the stepper motor is electrically connected to the motion controller via a driver;
[0045] The processor sends a pulse direction signal to the motion controller, and the motion controller outputs a microstepping drive signal to the driver according to the preset step size, which drives the stepper motor to rotate the lead screw, so that the slider moves along the linear guide.
[0046] After the slider reaches the preset acquisition position and remains stationary for a preset stabilization time, the processor sends a trigger signal to the image acquisition device; the preset stabilization time is greater than or equal to the sum of the exposure time of the image acquisition device and the mechanical vibration decay time.
[0047] The power module includes isolated power branches that supply power independently to the stepper motor driver and the image acquisition device, in order to reduce electromagnetic interference to the image acquisition device during the start and stop of the stepper motor.
[0048] The present invention also discloses a:
[0049] A computer program product, including a computer program or instructions, which, when executed by a processor, implement steps S1 to S4 in a method for registering and fusing multifocal image sequences with large-scale geometric misalignment.
[0050] A readable storage medium having a computer program or instructions stored thereon, which, when executed by a processor, implements steps S1 to S4 in the method for registration and fusion of multifocal image sequences with large-scale geometric misalignment.
[0051] The present invention, which adopts the above technical solution, has the following prominent features compared with the prior art:
[0052] 1. This invention organically combines a controllable displacement acquisition device (electric slide, motion controller, triggering coordination mechanism) with a back-end image registration and fusion algorithm. It ensures the quality of the source image through stable step acquisition, reduces mechanical vibration and electromagnetic interference through preset stabilization time and isolated power supply, and provides highly consistent input for subsequent registration, thus realizing deep collaboration between hardware acquisition and algorithm processing.
[0053] 2. This invention selects the last frame image with the smallest field of view as the global reference coordinate system to avoid ineffective filling of black border areas; by constructing a multi-scale image pyramid, the transformation matrix is passed layer by layer from coarse scale to fine scale, so that large-scale physical displacement falls into the algorithm's convergence domain at low resolution, and finally sub-pixel-level fine adjustment is completed at the original resolution, effectively solving the problem of large-scale geometric misalignment caused by perspective breathing effect; and solving the registration problem under large-scale geometric misalignment.
[0054] 3. This invention divides long sequence images into several image groups, uses the average image within the group as a reference for local cross-correlation focusing detection, extracts clear regions in each group and reduces the number of subsequent images to be fused, thus significantly reducing computational complexity while ensuring fusion quality.
[0055] 4. The present invention adopts a reference image step-by-step update mechanism. First, an initial full-focus image is generated using the average image as a reference, and then the full-focus image is used as the updated reference image for a second focus detection. At the same time, a smaller local correlation coefficient calculation window is used, which enables the focus detector to perceive more subtle structural features and significantly improves the accuracy of the focus area determination.
[0056] 5. This invention combines spatial consistency optimization of guided filtering with boundary refinement of closed matting. Guided filtering eliminates noise and keeps the decision map aligned with the object edge, while closed matting achieves pixel-level precision boundary positioning by solving the alpha mask, ultimately generating a full-focus image with natural edge transitions and no obvious visual artifacts. Attached Figure Description
[0057] Figure 1 This is a flowchart of the multi-scale federal registration and grouping optimization fusion process in an embodiment of the present invention;
[0058] Figure 2 This is a flowchart of the coarse fusion image registration process in an embodiment of the present invention;
[0059] Figure 3 This is a flowchart of pixel weighted fusion based on guided filtering in an embodiment of the present invention;
[0060] Figure 4 This is a result image of the high-resolution image sequence dataset of baseball bats used in the verification experiment of this invention;
[0061] Figure 5This is a result image of the high-resolution image sequence dataset of screws used in the verification experiment of this invention;
[0062] Figure 6 This is a graph showing the fusion result of the Lytro dataset;
[0063] Figure 7 This is a graph showing the fusion result of the grayscale dataset;
[0064] Figure 8 This is a schematic diagram of the focus detection results performed on the "standard baseball bat" image in this invention;
[0065] Wherein: (a) is the first source image (Source A), focusing on the foreground; (b) is the second source image (Source B), focusing on the background; (c) is the result image obtained by performing focus detection on source image (a) using the average image of (a) and (b) as a reference image; (d) is the result image obtained by performing focus detection on source image (a) using the full-focus fused image finally generated by this invention as a reference image.
[0066] Figure 9 The images show a comparison of the effects of registering (aligning) the standard "baseball bat" image sequence before and after the registration process in this invention. (a) Before registration, the two source images (e.g., near-focus and far-focus images) are directly superimposed, revealing obvious misalignment and ghosting of the sphere, support rod, and other targets. (b) Before registration, the two source images are superimposed using the red (R) and green (G) channels respectively, showing the RGB channel effect. The misaligned areas appear as yellow (red + green) artifacts, further highlighting the misalignment problem. (c) After registration using the method of this invention, the two aligned images are superimposed, showing complete overlap of the sphere, support rod, and other targets, with the ghosting disappearing. (d) After registration, the RGB channel superposition effect shows a uniform gray or single color, indicating that the two images have achieved precise spatial alignment at the pixel level, verifying the effectiveness and high precision of the registration method of this invention. Detailed Implementation
[0067] The present invention will be further illustrated below with reference to specific embodiments. The purpose of this illustration is solely to provide a better understanding of the invention. Therefore, the examples given do not limit the scope of protection of the present invention.
[0068] This embodiment presents a method for registration and fusion of multifocal image sequences with large-scale geometric misalignment. The specific steps are as follows.
[0069] Step S1, Image Registration: Acquire a multifocal image sequence acquired along the optical axis. Due to the perspective breathing effect, the field of view (FOV) in the sequence changes dynamically as the object distance decreases; to unify all source images to the same spatial reference and avoid invalid black borders, the last frame image with the smallest field of view is selected. As a global reference image.
[0070] Acquiring a multifocal image sequence along the optical axis specifically involves using an electric slide to drive the image acquisition device or the object under test to make a step-like relative displacement along the imaging optical axis; the motion controller controls the movement of the slider of the electric slide according to a preset step distance and dwell time, triggering the image acquisition device at each focal plane position, continuously acquiring multiple frames of images covering different depths of field, forming a large-scale geometric misalignment multifocal image sequence with translation, rotation, scale changes and field of view changes.
[0071] To address the large-scale translation, rotation, and scale variations between images, a cascaded registration method based on multi-scale image pyramids and enhanced correlation coefficients is employed; specifically, it includes the following sub-steps:
[0072] S101: For each pair of adjacent frame images ( and , Construct a multi-scale image pyramid containing multiple scale layers from the lowest resolution to the original resolution, with each scale layer including at least three scales; for example, set the scale factor set to... These correspond to 1 / 4, 1 / 2, and full size of the original image, respectively.
[0073] S102: The registration process starts at the lowest resolution coarse-scale layer. At this scale, the large-scale physical displacement between images caused by perspective effects is effectively reduced, effectively solving the problem of non-convex objective function and easy trapping in local extrema when optimizing directly at the original resolution, and providing a reliable initial estimate for global registration. For adjacent frames... and By iteratively maximizing the enhanced correlation coefficient between the reference image and the target image. Estimate the affine transformation matrix between adjacent frames. This serves as the initial estimate for global alignment. The enhanced correlation coefficient is defined as:
[0074] ;
[0075] in, These are pixel coordinates; for The value of the enhanced correlation coefficient; For the reference image in coordinates Pixel value at that location, The transformed source image in coordinates Pixel value at; In the current corresponding window Inside, reference image The mean value of pixels; In the current corresponding window Internal, transformed source image The mean value of pixels; For A local neighborhood window centered on the object.
[0076] Maximize through iterative optimization (such as gradient ascent) The transformation matrix at the coarse scale is obtained. (here) (Corresponding to the coarsest scale). This process utilizes the global grayscale information of the image, eliminates the need to calculate feature points, and is robust among multi-focus images with only a few common feature pairs.
[0077] Figure 8 This demonstrates the focus detection results performed on the "bat" dataset. (Comparison) Figure 8 As can be seen from (c) and (d), the detection result (d) with the fused full-focus image as a reference is clearer and more accurate than the result (c) with the average image as a reference, thus proving the superiority of the reference image selection in this invention.
[0078] S103: Transform the initially estimated matrix Upsampling is performed layer by layer (e.g., for affine transformations, the translation parameters are multiplied by 2, while rotation and scaling parameters remain relatively proportional), and then passed to the next higher resolution scale layer. This serves as the starting point for iterative optimization of this layer.
[0079] ;
[0080] in, This indicates an upsampling operation (such as magnifying the translation parameter by a factor of 2 to accommodate a higher resolution coordinate mapping).
[0081] Subsequently, in terms of scale Below, with Using the initial value, iterate again to maximize the enhanced correlation coefficient. This yields a more accurate transformation matrix. This process iterates layer by layer until the original resolution layer is reached. ,like Finally, a sub-pixel level transformation matrix is obtained. .
[0082] S104: Perform sub-pixel-level fine-tuning of the transformation matrix at the highest resolution original resolution layer, and output the optimal geometric transformation matrix between adjacent frames; specifically, for the first... Frame image ( ), to the global reference image The final transformation matrix It is obtained by concatenating the transformation matrices of the frame with those of adjacent frames:
[0083] ;
[0084] in, Indicates the first Frame to the The transformation matrix of the frame (i.e., obtained through S102-S103) ).
[0085] Through this transformation matrix For the source image Processing is performed to generate a geometrically aligned image.
[0086] ;
[0087] in, It is a pixel coordinate vector. This indicates that the pixel coordinates are spatially transformed according to the transformation parameters. The aligned image; Source image; These are the transformation parameters.
[0088] Step S2, Group Coarse Fusion: The registered long sequence of images is divided into several image groups. The average image within each group is calculated. Based on the local cross-correlation method, the correlation coefficient between each source image within a group and the average image of that group is calculated. The optimal focusing source image for each pixel position is determined, and coarse fused images for each group are generated. In this embodiment, a sequence containing N images is divided into M image groups, each containing p images. In this embodiment, p=10 is set. For the Mth image group, the average image of that group is calculated.
[0089] ;
[0090] in, The average image for this group. For the first in the group Zhang images, among which .
[0091] Subsequently, for each image within the group Calculate its relationship with the average image Correlation coefficient plot between ;
[0092] ;
[0093] in, and Representing images respectively The mean and standard deviation, and Average image The mean and standard deviation. The location of each pixel is determined by maximizing the correlation coefficient. Best source image index;
[0094] ;
[0095] And a coarsely fused image is constructed based on this index mapping. :
[0096] ;
[0097] Finally, M coarsely fused images are obtained. This grouping strategy aims to reduce computational complexity, minimize registration errors between adjacent images, and preserve sharp image information within each group.
[0098] Step S3, Dual Reference and Edge Refinement Fine Blending: (e.g.) Figure 1 , 2 Using the coarsely fused image as input, based on the first reference image An initial decision graph is constructed using local correlation measurement. The decision map is then subjected to consistency optimization and boundary refinement using edge-preserving filtering and image matting algorithms in sequence. Based on the optimized decision map, a weighted fusion process is performed on the input image to obtain the initial full-focus image. The specific process is as follows:
[0099] S301: Combine the multiple coarsely fused images output from step S2. Based on the acquisition order or spatial adjacency, two adjacent coarse fusion images are selected sequentially for fusion; for example, M coarse fusion images are paired sequentially and fused pairwise. , ...pair them up and fuse them one by one.
[0100] S302: For the two currently selected adjacent coarsely blended images, calculate the pixel-wise grayscale average value of the two, construct an average value image, and use this average value image as the first reference image.
[0101] S303: Calculate the local cross-correlation graphs between the two coarsely fused images and the first reference image respectively. .
[0102] S304: Compare the local cross-correlation coefficients of the two coarsely merged images at the same pixel location. The size determines the source of focus for a pixel location by identifying the coarsely fused image with larger coefficients.
[0103] S305: Construct an initial decision map of the two coarsely fused images based on the judgment results. The initial decision map is a matrix with the same size as the image. The value of each pixel is used to identify which coarse fused image that pixel should be taken from: if the local cross-correlation coefficient of the first coarse fused image is greater than that of the second, the corresponding position in the initial decision map takes the first identifier value; otherwise, the second identifier value is taken. Specifically, this can be represented as follows:
[0104] ;
[0105] in, =1 indicates that the focus is more accurate at this location (preserving the source image). ), =0 indicates that the location is more blurred (preserve the source image). ).
[0106] Preferably, in step S3, edge-preserving filtering and image matting algorithms are sequentially used to perform consistency optimization and boundary refinement on the decision graph, specifically including:
[0107] S306: Guided filtering is used to optimize the spatial consistency of the initial decision map, using the source image as the guide image. While eliminating noise, it maintains the alignment of the decision map with the boundaries of salient objects in the source image; the output of the guided filter can be expressed as:
[0108] ;
[0109] in, The filter radius is... This is the regularization parameter. This process eliminates noise and small irregularities in the decision map while maintaining the consistency between the decision map and the guidance image. Aligning salient object boundaries improves the spatial smoothness and consistency of the decision graph.
[0110] S307: Optimized decision graph A threshold is set to generate a three-valued Trimap, which divides the image region into foreground regions. ), background area ( ) and unknown areas ( );
[0111] ;
[0112] in, This is the threshold parameter (0.3 in this example).
[0113] S308: Input the Trimap and the source image into the closed matting algorithm, and solve for the optimal alpha mask by minimizing the energy function, which includes Laplacian matrix terms and diagonal matrix constraint terms;
[0114] ;
[0115] in, It is a Laplace matrix; The constraint vector comes from Trimap (foreground is 1, background is 0, unknown is 0.5). It is a diagonal matrix; The weighting parameters are used to calculate the continuous alpha values. Threshold processing (setting the alpha threshold) The final refined decision map is generated with a value of 0.55. The decision map has accurate boundary localization and good spatial consistency, laying the foundation for subsequent image fusion.
[0116] S309: Threshold the continuous alpha values obtained from the solution to generate a refined decision map with accurate boundary positioning.
[0117] S30A: Perform weighted fusion on the input image based on the optimized decision map, specifically including...
[0118] The refined decision map obtained after edge-preserving filtering and image matting algorithms. As the fusion weights, the two coarsely fused images are then weighted at the pixel level according to the following formula to generate the initial full-focus image. :
[0119] ;
[0120] in, For refined decision-making diagrams; Source image one; Image 2 is the source image.
[0121] Step S4, Reference image update and iterative optimization: such as Figure 3 The initial full-focus image is used as the second reference image, replacing the first reference image. Step S3 is repeated, and a smaller local correlation coefficient calculation window is used to generate the final full-focus fused image.
[0122] To avoid overfitting or underfitting, key parameters need to be adjusted during the iteration process:
[0123] Local correlation coefficient calculation window narrowing: In step S3, the calculation window for the local correlation coefficient is... (e.g., 11×11 pixels); during iteration, the window is reduced to... (e.g., 7×7 pixels) to improve the resolution of the focused area and capture details more precisely.
[0124] Optimization of guided filter parameters: In step S306, the regularization parameter e of the guided filter can be appropriately increased (e.g., adjusted from 0.01 to 0.05) to enhance the smoothing effect; the filter radius r can be moderately reduced (e.g., adjusted from 10 to 5) to avoid excessive blurring of edges.
[0125] Adjusting the parameters of the matting algorithm: In step S308, the weight of the energy function constraint term of the closed matting algorithm can be adjusted appropriately (such as increasing the weight of the diagonal matrix constraint term) to make the alpha mask fit the real boundary better.
[0126] Termination Conditions and Final Output: Iteration stops when one of the following conditions is met: the number of iterations reaches a preset value or the difference (e.g., mean square error, MSE) between two adjacent iterations of full-focus images is less than a set threshold. After iteration terminates, the final full-focus fused image is output, which is the high-precision fusion result after multiple rounds of reference image updates and decision map optimization.
[0127] This invention also discloses a system for registering and fusing multifocal image sequences with large-scale geometric misalignment. The system consists of a hardware acquisition subsystem and an image processing subsystem. Hardware acquisition subsystem
[0128] 1. The core function of the hardware acquisition subsystem is to generate high-quality multifocal image sequences with clearly defined spatial displacement relationships. Its specific structure and operation are as follows:
[0129] Precision displacement platform: This includes an electric slide, a motion controller, and related transmission mechanisms. The electric slide is supported by an aluminum profile base, on which high-precision linear guides and ball screws are arranged in parallel. The slider is mounted on the linear guides and connected to the nut of the ball screw. A stepper motor is connected to one end of the ball screw via a coupling. The motion controller (preferably a pulse-direction type controller) is electrically connected to the stepper motor via a stepper motor driver. When the processor sends control commands, the motion controller outputs microstepping drive signals, driving the stepper motor to rotate precisely, thereby driving the slider to perform precise linear motion along the linear guides via the ball screw transmission. The slider is equipped with a mounting base for fixing the image acquisition device (industrial camera and industrial lens) or the object being measured, allowing it to move along the imaging optical axis.
[0130] Image acquisition device: Typically an industrial camera and its matching industrial lens, used to acquire images under the control of a trigger signal. In a preferred embodiment, the processor sends a hardware trigger signal to the image acquisition device after the slider moves to a preset position and remains stationary for a preset stabilization time, ensuring that image acquisition occurs after mechanical vibration has decayed, thus eliminating motion blur.
[0131] Processor and control unit: Typically an industrial control computer, it connects to the motion controller and image acquisition device via communication interfaces (such as USB or Ethernet). The processor is responsible for sending motion parameters, controlling the acquisition timing, and running image processing algorithms.
[0132] Power Supply and Interference Suppression Design: The power module supplies power to the entire system. Its key design feature is the inclusion of two independent isolated DC-DC power branches, supplying power to the stepper motor driver and the image acquisition device respectively. This isolation design effectively blocks the large current transients and electromagnetic noise generated during stepper motor start-up, stopping, and commutation from crosstalking to the sensitive image sensor circuitry via the power ground line, thereby significantly reducing electromagnetic interference fringes in the image and ensuring image quality from the source.
[0133] 2. Image Processing Subsystem:
[0134] The image processing subsystem is the software implementation carrier of the method of this invention, integrated into the aforementioned processor (such as an industrial computer). This subsystem exists in the form of a software module or program, specifically including:
[0135] Image registration module: Used to perform step S1. This module reads the raw sequence acquired by the hardware, performs cascaded registration based on multi-scale enhanced correlation coefficient (ECC), and outputs a spatially aligned image sequence.
[0136] Grouped coarse fusion module: This module performs step S2. It groups the long sequence into groups, calculates the local cross-correlation coefficient between the images within a group and the average image, extracts the clearest pixels from each group, and generates a coarsely fused image set with a significantly reduced number of pixels.
[0137] Dual-reference fine fusion module: used to execute step S3. This module is the core fusion unit, which sequentially implements: constructing an initial decision map using the average image as the first reference, optimizing spatial consistency using guided filtering, refining boundaries using closed matting, and finally performing weighted fusion based on the refined decision map to generate an initial full-focus image.
[0138] Iterative optimization module: This module is used to execute step S4. It updates the initial full-focus image with the second reference image and repeats the fine fusion process with finer parameters (such as a smaller local computation window) to perform iterative optimization and output the final, high-quality full-focus fused image.
[0139] After the system is powered on and initialized, it will operate according to the preset procedure:
[0140] (1) Parameter settings: Users set the acquisition parameters through the human-computer interaction software on the processor, including: total displacement stroke L, single step distance Δd (e.g., 10μm), slider movement speed v, preset number of acquisition positions N, camera exposure parameters, etc.
[0141] (2) Motion and data acquisition coordination: The processor sends motion commands to the motion controller. The motion controller calculates the required number of pulses based on the preset step distance Δd and sends a microstepping drive signal to the driver to drive the stepper motor to move the slider precisely by one step distance. When the slider reaches the preset data acquisition position, the motion controller sends a "positioned" signal back to the processor.
[0142] (3) Stabilization and Triggering Mechanism: After receiving the "in place" signal, the processor starts the stabilization timer. The preset stabilization time T... stable It needs to be determined experimentally, and its value must be greater than or equal to the camera's exposure time T. exposure The time T required for the mechanical vibration of the slider to completely decay damping The sum, i.e., T stable ≥T exposure +T damping Only after the settling time has ended and the platform is completely still will the processor send a hardware trigger signal to the image acquisition device. Upon receiving the trigger signal, the camera acquires a frame of image at the current focal plane and transmits the image data back to the processor's memory or storage unit via an interface (such as GigEVision).
[0143] (4) Loop and end: Repeat steps (2) and (3) until all N locations have been acquired, thus obtaining a multifocal image sequence covering different depths of field.
[0144] The image processing module integrated within the processor starts automatically after image sequence acquisition is complete, or can be manually invoked by the user. This module executes steps S1 to S4 sequentially:
[0145] S1 (Image Registration): Reads the original sequence, calls the multi-scale ECC registration submodule, and outputs the registered sequence.
[0146] S2 (Grouped Coarse Fusion): Calls the submodule for calculating group averaging and local cross-correlation, and outputs M coarsely fused images.
[0147] S3 (Fine Fusion): Calls the guided filtering, closed matting, and weighted fusion sub-modules to generate the initial full-focus image.
[0148] S4 (Iterative Optimization): Calls the reference image update and secondary optimization submodule to generate the final full-focus image.
[0149] After processing, the image processing module saves the final full-focus fused image to the specified path and displays it on the software interface.
[0150] The present invention also discloses a:
[0151] The computer program product includes a computer program or instructions that, when executed by a processor, implement steps S1 to S4 in the method for registering and fusing multifocal image sequences with large-scale geometric misalignment. The computer program product may be an installation package stored on a USB flash drive, external hard drive, or downloaded via a network, containing a computer program or instructions that, when run on the processor's operating system (such as Windows or Linux), guide the processor to execute the image processing steps.
[0152] A readable storage medium stores a computer program or instructions. When executed by a processor, the computer program or instructions implement steps S1 to S4 of the method for registering and fusing multifocal image sequences with large-scale geometric misalignment. The readable storage medium can be a system-built-in solid-state drive (SSD), read-only memory (ROM), or a removable SD card, CF card, etc. When the computer program or instructions stored thereon are loaded and executed by the processor, the registration and fusion method is implemented, enabling operators without professional image processing knowledge to generate high-quality full-focus images with a single click by running the program.
[0153] Experimental verification
[0154] To comprehensively and objectively evaluate the performance of the proposed method for registration and fusion of large-scale geometrically misaligned multifocal image sequences, this embodiment conducted systematic experiments on self-made industrial datasets and publicly available standard datasets. The effectiveness, superiority, and generalization ability of the invention were verified from two dimensions: qualitative visual analysis and quantitative index evaluation.
[0155] 1. Experimental setup
[0156] (1) Dataset
[0157] To evaluate the algorithm's performance in challenging industrial macro photography scenarios, two self-made datasets were used: the "Basket" dataset and the "Screw" dataset, with an image resolution of 2600×2160. Unlike standard multifocus datasets, this sequence of images was obtained by controlling the camera to take step-by-step shots along the optical axis, ensuring continuous depth-of-field coverage, but inevitably introducing complex geometric misalignments (translation, rotation, scale, and field of view changes) caused by perspective breathing.
[0158] To verify the generalization performance of the method, further extended comparative experiments were conducted on two public multifocal datasets, "Lytro" and "Grayscale".
[0159] (2) Comparison methods and evaluation indicators
[0160] To ensure the fairness of the comparison, before inputting the source sequence into any comparison method, the multi-scale cascaded registration module proposed in step S1 of this invention is applied to pre-align all images to eliminate global geometric errors and provide a consistent and high-quality input benchmark for all comparison methods.
[0161] This study selected 13 representative fusion algorithms as comparison benchmarks, covering classical transform domain methods, spatial domain methods, and cutting-edge deep learning models. Specifically, these include: Classical transform domain methods such as Dual-Tree Complex Wavelet Transform (DTCWT) and Non-Subsampled Profilotype Transform with Sparse Representation (NSCT-SR); Classical spatial domain methods such as Gradient-Based Fusion (GFF), Multi-Scale Weighted Gradient Fusion (MWGF), Independent Component Analysis (ICA), Quadtree Decomposition, Dense Scale-Invariant Feature Transform (DSIFT), and Multi-Scale Image Structure Fusion (MISF); and Deep learning-based methods such as Convolutional Neural Network (CNN), Graph Attention Convolutional Network (GACN), U2Fusion model, Spatial-Spectral Feature Extraction-Based Fusion (SESF), and Cross-Channel Sparse Representation (CCSR).
[0162] Five recognized metrics were used to measure the quality of the fusion results:
[0163] Normalized mutual information (Q) MI ) and nonlinear related information entropy (Q NCIE ): This quantizes the degree to which the fused image inherits information from the source image and the correlation between the images. A higher value is better.
[0164] Gradient-based metrics (Q) G ): Evaluates the retention of edge details and sharpness. The higher the value, the better.
[0165] Based on the index of the human visual system (Q) CB and Q CV ): Simulates the subjective perception of the human eye from the perspectives of contrast sensitivity and visual difference, respectively. Q CB The higher the value, the better, Q CV The lower the value, the less visual distortion there is.
[0166] 2. Experimental Results and Analysis
[0167] (1) Qualitative visual analysis
[0168] Visual contrast results on the “bat” and “screw” datasets (corresponding to the attached) Figure 4 , 5The fused image generated by the method of this invention achieves clear and natural imaging in key areas such as spheres, support structures, screw threads, and nut details, with smooth edge transitions and no obvious ghosting, blurring, or unnatural visual artifacts.
[0169] In contrast, some of the comparison algorithms have obvious shortcomings: for example, the fusion effect of CNN and U2Fusion is not ideal, and it fails to effectively preserve foreground focus information; methods such as NSCT-SR, ICA, and GACN show varying degrees of blurring at the focus edges; Quadtree and DSIFT methods have significant artifacts in their results; and SESF and DTCWT show discontinuities at the edges.
[0170] Extended experiments on public datasets (“Lytro” and “Grayscale”) (see attached) Figure 6 , 7 This further verifies the generalization ability of the present invention. The method of the present invention maintains the highest clarity and visual naturalness in areas such as crease text and complex textures, effectively avoiding problems such as ghosting, blurring, distortion or block artifacts that occur in other methods.
[0171] (2) Quantitative indicator evaluation
[0172] Tables 1 to 4 show the objective evaluation results of each algorithm on the “Basketball,” “Screw,” “Lytro,” and “Grayscale” datasets, respectively (in the tables, “Ours” represents the method of this invention).
[0173] Table 1 Bat Evaluation Results
[0174]
[0175] Table 2 Screw Evaluation Results
[0176]
[0177] Table 3. Evaluation Results of the Lytro Dataset
[0178]
[0179] Table 4 Evaluation Results of Grayscale Dataset
[0180]
[0181] Table 1 shows that the method of this invention achieved optimal values in all five evaluation indicators, demonstrating its comprehensive performance advantages.
[0182] Table 2 shows the sharpness index Q of the method of the present invention. G and contrast Q CBMaintain top performance on key metrics. Although in Q MI The Q value is slightly lower than that of the CCSR algorithm, but the Q value of the CCSR algorithm is... G and Q CV The poor performance indicates that it preserves information at the expense of overall image clarity and introduces severe visual distortion, making its overall performance inferior to that of the present invention.
[0183] In Tables 3 and 4, the method of this invention remains leading or among the top in the vast majority of metrics. For example, on the "Lytro" dataset, Q... G Q CB All are optimal; Q is optimal on the "Grayscale" dataset. G Q CB It also performed exceptionally well. This fully demonstrates that the technical framework of this invention is not only adept at handling large-scale geometric misalignments in industrial scenarios, but also possesses outstanding comprehensive performance and wide applicability in conventional multi-focus image fusion tasks.
[0184] 3. Experimental Conclusions
[0185] Through the qualitative and quantitative experimental verification of the above system, the following conclusions can be drawn:
[0186] Registration effectiveness: The multi-scale cascaded registration method proposed in this invention can effectively correct large-scale geometric misalignment caused by camera step displacement, providing a foundation for high-quality fusion.
[0187] Superiority of Fusion: The "group coarse fusion" and "dual reference iterative fine fusion" strategies proposed in this invention, combined with guided filtering and closed matting edge optimization, can generate full-focus images with clear edges, natural transitions, and no significant visual artifacts.
[0188] Comprehensive performance: On both self-made and publicly available datasets, the method of this invention significantly outperforms or is comparable to a variety of mainstream and cutting-edge comparison algorithms in multiple evaluation metrics, demonstrating stable, comprehensive and leading fusion performance, and verifying the advanced nature, effectiveness and robustness of the technology of this invention.
[0189] The above description is merely a preferred embodiment of the present invention and is not intended to limit the scope of the invention. All equivalent changes made based on the description and drawings of the present invention are included within the scope of the present invention.
Claims
1. A method for registration and fusion of multifocal image sequences with large-scale geometric misalignment, characterized in that: The specific steps are as follows: Step S1, Image Registration: Acquire a multifocal image sequence acquired along the optical axis, select the last frame image with the smallest field of view as the global reference image; construct a multi-scale image representation, estimate the geometric transformation matrix between adjacent frames sequentially from coarse to fine scale, and map all source images to the coordinate system of the global reference image through the concatenation of the transformation matrices of adjacent frames, and output the registered image sequence. Step S2, Group coarse fusion: Divide the registered long sequence of images into several image groups, calculate the average image within each group, calculate the correlation coefficient between each source image within the group and the average image of the group based on the local cross-correlation method, determine the best focusing source image for each pixel position, and generate coarse fused images for each group. Step S3, Dual Reference and Edge Refinement Fusion: Using the coarse fused image as input, an initial decision map is constructed based on the first reference image through local correlation measurement; edge-preserving filtering and image matting algorithms are used sequentially to perform consistency optimization and boundary refinement on the decision map; weighted fusion is performed on the input image based on the optimized decision map to obtain the initial full-focus image; Step S4, Reference Image Update and Iterative Optimization: The initial full-focus image is used as the second reference image to replace the first reference image. Step S3 is repeated, and a smaller local correlation coefficient calculation window is used to generate the final full-focus fused image.
2. The method for registration and fusion of large-scale geometrically misaligned multifocal image sequences according to claim 1, characterized in that: In step S1, acquiring a multifocal image sequence along the optical axis specifically includes: driving the image acquisition device or the object under test to make a step-like relative displacement along the imaging optical axis by an electric slide; the motion controller controls the slider of the electric slide to move according to a preset step distance and dwell time, triggering the image acquisition device at each focal plane position, continuously acquiring multiple frames of images covering different depths of field, forming a large-scale geometric misalignment multifocal image sequence with translation, rotation, scale changes and field of view changes.
3. The method for registration and fusion of large-scale geometrically misaligned multifocal image sequences according to claim 1, characterized in that: In step S1, a multi-scale image representation is constructed, and the geometric transformation matrix between adjacent frames is estimated sequentially from coarse to fine scale. Specifically, this includes: S101: Construct a multi-scale image pyramid containing multiple scale layers from the lowest resolution to the original resolution, with each scale layer including at least three scales; S102: The registration process starts from the lowest resolution coarse-scale layer, at which the enhancement correlation coefficient between the reference image and the target image is calculated. The enhancement correlation coefficient is maximized through iterative optimization, and the affine transformation matrix between adjacent frames is estimated as the initial estimate for global alignment. S103: Upsample the initially estimated transformation matrix layer by layer and pass it to the next higher resolution scale layer as the starting point for iterative optimization of that layer; S104: Perform sub-pixel level fine-tuning of the transformation matrix at the highest resolution original resolution layer, and output the optimal geometric transformation matrix between adjacent frames.
4. The method for registration and fusion of large-scale geometrically misaligned multifocal image sequences according to claim 1, characterized in that: In step S3, using the coarsely fused image as input, an initial decision map is constructed based on the first reference image through local correlation measurement, specifically including: S301: Select two adjacent coarse fusion images from the multiple coarse fusion images output in step S2 according to the acquisition order or spatial adjacency relationship and fuse them one by one. S302: For the two currently selected adjacent coarsely fused images, calculate the pixel-wise grayscale average value of the two, construct an average value image, and use this average value image as the first reference image; S303: Calculate the local cross-relationship graphs between the two coarsely fused images and the first reference image respectively; S304: At the same pixel location, compare the local cross-correlation coefficients of the two coarsely fused images and determine the coarsely fused image with the larger coefficient as the source of focus at that pixel location; S305: Construct an initial decision map for the two coarsely fused images based on the determination result. The initial decision map is a matrix with the same size as the image. The value of each pixel position is used to identify which coarsely fused image the pixel should be taken from: if the local cross-correlation coefficient of the first coarsely fused image is greater than that of the second image, the first identifier value is taken at the corresponding position in the initial decision map; otherwise, the second identifier value is taken.
5. The method for registration and fusion of large-scale geometrically misaligned multifocal image sequences according to claim 4, characterized in that: In step S3, edge-preserving filtering and image matting algorithms are used sequentially to perform consistency optimization and boundary refinement on the decision graph, specifically including: S306: The initial decision map is spatially consistent with the source image by using guided filtering. The source image is used as the guide image to maintain the alignment between the decision map and the boundaries of salient objects in the source image while eliminating noise. S307: Set a threshold for the optimized decision map to generate a three-valued Trimap, which divides the image region into foreground region, background region and unknown region; S308: Input the Trimap and the source image into the closed matting algorithm, and solve for the optimal alpha mask by minimizing the energy function, which includes Laplacian matrix terms and diagonal matrix constraint terms; S309: Threshold the continuous alpha values obtained from the solution to generate a refined decision map with accurate boundary positioning.
6. The method for registration and fusion of large-scale geometrically misaligned multifocal image sequences according to claim 5, characterized in that: In step S3, weighted fusion is performed on the input image based on the optimized decision map, specifically including: The refined decision map obtained after edge-preserving filtering and image matting algorithm optimization is used as the fusion weight; The two coarsely merged images are then weighted at the pixel level using the following formula to generate an initial full-focus image. : ; in, For refined decision-making diagrams; Source image one; Image 2 is the source image.
7. A system for registration and fusion of multifocal image sequences with large-scale geometric misalignment, characterized in that: include, An electric slide is used to support an image acquisition device or the object being measured and to achieve linear motion along the optical axis. The motion controller, electrically connected to the electric slide table, is used to send motion control signals to the electric slide table; An image acquisition device is used to acquire images in response to a trigger signal; The processor is communicatively connected to both the motion controller and the image acquisition device. The processor sends motion control commands to the motion controller and trigger signals to the image acquisition device. An image processing module is integrated within the processor, which executes the registration and full-focus fusion method for large-range geometrically misaligned multifocal image sequences as described in any one of claims 1 to 9, and outputs a full-focus fused image. The power supply module is electrically connected to the electric slide, motion controller, image acquisition device and processor respectively, and is used to supply power.
8. The large-scale geometric misalignment multifocal image sequence registration and fusion system according to claim 7, characterized in that: The motion controller is a pulse direction type controller, and the electric slide includes a stepper motor, a lead screw transmission mechanism, a linear guide rail, and a slider; the stepper motor is electrically connected to the motion controller via a driver; The processor sends a pulse direction signal to the motion controller, and the motion controller outputs a microstepping drive signal to the driver according to the preset step size, which drives the stepper motor to rotate the lead screw, so that the slider moves along the linear guide. After the slider reaches the preset acquisition position and remains stationary for a preset stabilization time, the processor sends a trigger signal to the image acquisition device; the preset stabilization time is greater than or equal to the sum of the exposure time of the image acquisition device and the mechanical vibration decay time. The power module includes isolated power branches that supply power independently to the stepper motor driver and the image acquisition device, in order to reduce electromagnetic interference to the image acquisition device during the start and stop of the stepper motor.
9. A computer program product, comprising a computer program or instructions, characterized in that: When a computer program or instruction is executed by a processor, steps S1 to S4 of the method for registration and fusion of large-scale geometrically misaligned multifocal image sequences as described in any one of claims 1 to 6 are implemented.
10. A readable storage medium having a computer program or instructions stored thereon, characterized in that: When a computer program or instruction is executed by a processor, steps S1 to S4 of the method for registration and fusion of large-scale geometrically misaligned multifocal image sequences as described in any one of claims 1 to 6 are implemented.