An underwater range-gated dynamic imaging method based on meta-heuristic pose estimation
By employing a metaheuristic attitude estimation method, combined with a composite objective function of adaptive weighted pixel loss and rolling kernel consistency loss, and utilizing particle swarm optimization algorithm for six-degree-of-freedom pose estimation in underwater range-gated laser imaging, the spatial misalignment problem caused by 6D pose changes in underwater imaging is solved, achieving high-precision and robust image registration.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- OCEAN UNIV OF CHINA
- Filing Date
- 2026-01-15
- Publication Date
- 2026-04-21
AI Technical Summary
Existing underwater range-gated laser imaging technology suffers from spatial misalignment due to 6D pose changes in complex environments. Traditional methods rely on feature point matching, which lacks robustness and accuracy in sparse and noisy environments, making it difficult to meet the requirements of high-precision imaging.
A metaheuristic pose estimation method is adopted. By constructing a composite objective function of adaptive weighted pixel loss and rolling kernel consistency loss, and combining it with particle swarm optimization algorithm, six-degree-of-freedom pose estimation and image registration are performed. Gray-level consistency is used to construct the optimization objective to suppress interference and error accumulation in bright areas.
It achieves high-precision image registration in highly sparse and noisy environments, improves the spatial consistency and robustness of images, and is suitable for underwater laser scanning scenarios with scarce features.
Smart Images

Figure CN121522662B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of underwater imaging technology, specifically to an underwater range-gated dynamic imaging method based on metaheuristic attitude estimation. Background Technology
[0002] Marine range-gated laser imaging is widely used for underwater target identification and environmental perception because it can achieve long-range, high-resolution target detection in complex environments. However, its imaging process is affected by equipment movement, attitude disturbances and water medium, resulting in significant 6D pose changes between consecutive frames. This causes spatial misalignment in single-frame imaging results, which seriously affects subsequent reconstruction work. In engineering applications, inertial navigation systems (INS) are commonly used for pose measurement. INS can continuously output platform motion parameters, providing a guarantee for preliminary registration in dynamic scenes. However, its equipment cost is high, it has cumulative drift, and it is easily affected by changes in water velocity, making it difficult to meet the needs of high-precision underwater image fusion.
[0003] Therefore, how to achieve autonomous high-precision estimation of 6D pose in multi-frame imaging without expensive hardware has become a key scientific problem and research hotspot in the field of underwater range-gated laser imaging. Current 6D pose estimation methods are mainly divided into learning-based and non-learning-based methods. Learning-based methods, such as FFB6D and UWCL-6D, have achieved high-precision pose prediction under RGB-D or simulated data through end-to-end deep learning modeling, but their generalization ability is severely limited by the extreme scarcity of practically available underwater datasets. Non-learning methods mainly rely on the stable detection of feature points between images and... However, these non-learning methods face unprecedented challenges in underwater range-gated laser imaging scenarios. On the one hand, the data generated by the line laser imaging system is extremely sparse and mainly concentrated in narrow strip regions, resulting in highly limited available spatial and texture information between two frames. On the other hand, although the range gating mechanism effectively suppresses scattering noise, it further compresses the effective imaging area, making the image area outside the target depth almost devoid of useful data. These factors together make feature point extraction and matching extremely prone to failure, thus causing a significant decrease in the robustness and accuracy of traditional feature-based attitude estimation algorithms in such extreme scenarios.
[0004] Given the aforementioned limitations, there is an urgent need for a novel method that can break free from the dependence on feature points and training data and adapt to high-noise and extremely sparse structure environments. Therefore, we propose an underwater range-gated dynamic imaging method based on metaheuristic attitude estimation. By leveraging the global optimization capability of the metaheuristic algorithm (MA) in complex problems such as high-dimensional nonlinearity, lack of gradient information, and weak prior knowledge, this method offers a potential solution to such attitude estimation challenges. Summary of the Invention
[0005] To solve the above-mentioned technical problems, the present invention is implemented through the following technical solution: an underwater range-gated dynamic imaging method based on metaheuristic attitude estimation, comprising the following steps:
[0006] S1: Construct an underwater range-gated laser imaging system to acquire a continuous sequence of spatiotemporally matched raw scan images by synchronously controlling laser emission, scanning, and gated exposure.
[0007] S2: Select reference frames from the original scanned image sequence and determine the effective imaging area. Use six-degree-of-freedom pose parameters to model the rigid body motion between frames and construct a global optimization objective function based on the mean square error of grayscale to evaluate the alignment degree of adjacent frames.
[0008] S3: Design an adaptive weighted pixel loss function, which assigns weights to pixels by normalizing grayscale and gradient information, suppresses the influence of bright areas, enhances the contribution of structural edges in registration, and improves local alignment accuracy and robustness.
[0009] S4: Introduce the rolling kernel consistency loss function, construct the rolling accumulation kernel based on the sliding window historical frame information, constrain the consistency between the current frame and the historical structure mean, suppress error accumulation and global drift, and ensure timing stability;
[0010] S5: A composite objective function is constructed by fusing the adaptive weighted pixel loss function and the rolling kernel consistency loss function. The particle swarm optimization algorithm is used to find the global optimization in the six-dimensional pose space. By iteratively updating the particle position and velocity, the optimal six-degree-of-freedom pose matrix is obtained to achieve high-precision image registration.
[0011] S6: The optimized six-DOF pose matrix is applied to the full sequence of scanned images, uniformly mapped to the reference coordinate system, eliminating inter-frame spatial deviation and distortion, and performing grayscale accumulation, normalization and brightness enhancement processing on the registered sequence to generate underwater dynamic fusion imaging results with continuous structure and clear details.
[0012] Preferably, step S1 includes the following steps:
[0013] An underwater range-gated laser imaging system is constructed, comprising a pulsed laser, a gating camera, a scanning mirror, and a synchronous timing circuit. The pulsed laser emits a pulsed laser with a center wavelength of 532nm, which has excellent underwater penetration and effectively improves imaging depth and signal strength.
[0014] By controlling the synchronous timing circuit, the laser emission, scanning mirror deflection and time-gated exposure of the gating camera are kept synchronized. This wavelength has excellent penetration underwater, effectively improving imaging depth and signal strength.
[0015] After being deflected by the scanning mirror, the laser scans point by point within the camera's field of view, forming a distance-gated image at each scanning position. This ultimately constitutes a continuous sequence of original scanned images with spatiotemporal matching. Point-by-point scanning achieves high-resolution coverage and forms a spatiotemporally consistent sequence, facilitating subsequent registration and fusion.
[0016] Preferably, step S2 includes the following steps:
[0017] The first frame with clear imaging is selected from the original scanned image sequence as the reference frame, and an effective imaging area containing laser stripes is delineated on the reference frame to achieve a stable reference for subsequent registration and avoid initial deviations caused by inter-frame differences.
[0018] The variables to be optimized are represented as six-degree-of-freedom pose parameters. Homogeneous transformation matrices are used to uniformly model rotation and translation, describing the rigid body motion between adjacent frames during the scanning process, fully representing the actual motion, and providing a mathematical model basis for accurate spatial alignment.
[0019] Construct a global optimization objective function based on the pixel grayscale mean square error, and use the reference frame... With the frame to be registered The mean square error of the grayscale difference at the same pixel location is used as the registration cost. The alignment degree between the reference frame and the frame to be registered within the effective imaging area is quantized. Optimization is driven by quantization error to ensure that the registration result achieves high-precision alignment at the pixel level.
[0020] Preferably, the global optimization objective function is specifically:
[0021] Suppose there are N effective pixels in the effective imaging area, and the grayscale value of the reference frame is denoted as N. The frames to be registered undergo pose transformation The grayscale value after that ;
[0022] Using the sum of the squared grayscale differences at the same pixel position in two frames as the registration cost, a global optimization objective function is constructed. This function is highly sensitive to small spatial deviations within the effective imaging area and accurately reflects the alignment degree of laser stripes in adjacent frames. Its expression is:
[0023] ;
[0024] In the formula: Let the objective function (fitness function) represent the pose transformation. Registration error under the following conditions; Here is the pose transformation matrix (6 degrees of freedom); The total number of pixels within the effective imaging area (ROI); For pixel index; For the first The coordinates of one pixel; For the reference frame in coordinates The grayscale value at that location; To transform in the current pose Below, the frame to be registered is in coordinates The grayscale value at that location.
[0025] Preferably, step S3 includes the following steps:
[0026] The pixel grayscale of the current frame and the reference frame are normalized to eliminate grayscale distribution offset between different frames, effectively eliminate the differences in illumination and sensor response, improve the grayscale consistency between frames, and enhance the registration stability.
[0027] An exponentially decaying adaptive weight is constructed based on normalized gray values and gradient magnitudes to suppress the weight of bright areas and enhance the weight of edge areas. This significantly reduces the pulling effect of overly bright pixels on registration, highlights the contribution of edge and texture areas, and improves the structural alignment accuracy.
[0028] Adaptive weights are introduced into pixel-level loss calculation to construct an adaptive weighted pixel loss function. To improve the structural consistency and robustness of registration, the registration process focuses more on structural features rather than areas of abnormal brightness, effectively suppressing underwater scattering and reflection interference, and improving the robustness of the algorithm;
[0029] The The mathematical definition can be expressed as:
[0030] ;
[0031] ;
[0032] ;
[0033] ;
[0034] ;
[0035] ;
[0036] In the formula: It is an adaptive weighted pixel loss function; Pixel-level adaptive weights; , as well as These are constant parameters that control the influence of the weights; Normalized grayscale values are pixel values. Normalizing grayscale values can reduce grayscale distribution shifts and range differences between different frames, thereby maintaining numerical stability during the optimization process. This is the current grayscale value; This is the maximum grayscale value; The minimum grayscale value; It is a very small constant to prevent division by zero; The gradient is the pixel gradient magnitude. The gradient only focuses on the change in gray level and is not affected by the absolute brightness, which can improve the ability to resist brightness interference. The global gradient mean; It is the sigmoid function; and These represent the horizontal and vertical convolution kernels of the Sobel operator, respectively.
[0037] Preferably, step S4 includes the following steps:
[0038] Save recent data using a sliding window The frame registration results are used to calculate the average gray level as the historical structure reference for the current frame. Through the sliding window mechanism, short-term noise disturbances are effectively suppressed, and the structural stability and noise resistance of the historical reference are improved.
[0039] By constructing a rolling accumulation kernel through a weighted update method, the fusion ratio of historical information and the current frame is controlled, achieving a dynamic balance between long-term structural trends and recent changes, and enhancing the system's adaptability to changes in real-world scenarios.
[0040] Based on the grayscale difference between the scrolling accumulation kernel and the current frame, a scrolling kernel consistency loss function is constructed. To suppress the accumulation of multi-frame errors and global drift, by constraining temporal consistency, registration drift is significantly reduced, and the stability of multi-frame images in the temporal dimension is ensured;
[0041] The Inherited The calculation method can be expressed as:
[0042] ;
[0043] ;
[0044] ;
[0045] In the formula: This represents the value of the rolling kernel consistency loss function. Here is the pose transformation matrix; This refers to the number of effective pixels. Pixel-level adaptive weights; The current time is used to accumulate the kernel in pixels. The grayscale value at that location; For pose transformation Then, the frame to be registered is in pixels The grayscale value at that location; For the nearest within the sliding window Average gray level of the frame; The length of the sliding window (in frames); For the first A grayscale image of a frame; The frame index within the sliding window; The current cumulative kernel grayscale value; This is the cumulative kernel grayscale value from the previous time step; The kernel update rate control parameters.
[0046] Preferably, step S5 includes the following steps:
[0047] Adaptive weighted pixel loss function Consistency loss function with rolling kernel Perform linear weighted fusion to construct a composite objective function. This fusion mechanism combines local alignment with global stability, effectively suppressing drift and improving registration robustness.
[0048] Using six-degree-of-freedom pose parameters as the search variables for the particle swarm, a number of particles are randomly initialized within a preset feasible region. The random initialization covers multi-dimensional space, which avoids the algorithm getting stuck in local optima and improves the global search capability.
[0049] The particle swarm optimization algorithm is adopted with a composite objective function as the fitness function to perform global optimization in the six-dimensional pose space. The particle position and velocity are iteratively updated until convergence or the maximum number of iterations is reached to obtain the optimal six-degree-of-freedom pose matrix. Iterative optimization gradually improves the pose estimation accuracy and finally outputs stable and reliable spatial alignment parameters.
[0050] Preferably, the expression for the composite objective function is as follows:
[0051] ;
[0052] In the formula: It is a composite objective function; It is an adaptive weighted pixel loss function; The rolling kernel consistency loss function; Here is the pose transformation matrix; and These are weight parameters, which control the balance between local alignment accuracy and global temporal consistency, in order to achieve coordinated adjustment among different loss properties. Emphasis is placed on local alignment accuracy. This is beneficial for stability and drift control.
[0053] Preferably, the process of global optimization in the six-dimensional pose space is as follows:
[0054] Five initial population sizes of 15, 20, 25, 30 and 35 were set, and the number of iterations was uniformly set to 300. After comparative experiments, the convergence performance of different sizes was significantly different. Finally, 25 was selected to achieve the optimal balance between convergence speed and accuracy.
[0055] In each iteration, the pose transformation corresponding to each particle is calculated. The composite objective function value is used to update the individual optimal and global optimal solutions. Through an adaptive update strategy, the global search process is significantly accelerated, and the problem of getting trapped in local optima is avoided.
[0056] By gradually approximating the optimal pose that minimizes the composite objective function through velocity and position update formulas, high-precision pose estimation and image registration are achieved, significantly improving the spatial consistency of multi-frame registration.
[0057] Preferably, step S6 includes the following steps:
[0058] The optimized six-degree-of-freedom pose matrix is applied to each frame of the full sequence of scanned images, and each frame of images is uniformly mapped to the reference coordinate system to achieve geometric alignment of multiple frames of images, eliminate spatial deviation, and significantly improve the consistency of image superposition.
[0059] The grayscale values of the registered scanned image sequence are accumulated according to the pixel position to achieve pixel-level superposition enhancement of multiple scan observations. The accumulated result is divided by the number of frames involved in the fusion and normalized to restore the dynamic range of the image, enhance the signal-to-noise ratio of the target area, suppress background noise, avoid brightness saturation, and improve imaging contrast.
[0060] Brightness enhancement processing is performed on the normalized fused image to improve the visibility and detail representation of weak texture areas. Local contrast adjustment is then used to further optimize the image structure continuity and edge sharpness, ultimately obtaining a fused imaging result with continuous structure and clear details for subsequent recognition and analysis. This enhances edge and texture details, improves image recognition, and provides reliable input for subsequent analysis and reconstruction.
[0061] This invention provides an underwater range-gated dynamic imaging method based on metaheuristic attitude estimation. It has the following advantages:
[0062] (I) This underwater range-gated dynamic imaging method based on metaheuristic attitude estimation applies the metaheuristic optimization idea to the 6D pose estimation problem of underwater range-gated laser imaging. Based on PSO, a 6D pose optimization framework suitable for highly sparse and highly constrained scenarios is constructed, overcoming the fundamental challenges of feature point scarcity and insufficient dataset.
[0063] (ii) The underwater range-gated dynamic imaging method based on metaheuristic attitude estimation is designed with a composite objective function that integrates adaptive weighted pixel loss and rolling kernel consistency loss. The adaptive weighted pixel loss effectively improves the local alignment accuracy and suppresses brightness anomaly interference, while the rolling kernel consistency loss ensures the stability of the global structure and suppresses the drift caused by error accumulation.
[0064] (III) This underwater distance-gated dynamic imaging method based on metaheuristic attitude estimation utilizes the particle swarm optimization algorithm to perform global optimization in the six-degree-of-freedom pose space. By fusing adaptive weighted pixel loss and rolling kernel consistency loss through a composite objective function, it effectively suppresses the drift problem caused by interference from bright areas and error accumulation in underwater dynamic imaging, thereby significantly improving the accuracy and robustness of image registration. It is suitable for underwater laser scanning scenarios with sparse features and complex noise.
[0065] (iv) This is an underwater range-gated dynamic imaging method based on metaheuristic attitude estimation. Traditional methods rely on feature point extraction and matching, which are prone to failure in underwater range-gated imaging. This method constructs an optimization target based on grayscale consistency, without relying on feature points or training data. It directly performs 6D attitude estimation through pixel-level information, overcoming the inherent limitation of feature scarcity in extremely sparse environments.
[0066] (v) The underwater range-gated dynamic imaging method based on metaheuristic attitude estimation introduces a rolling kernel consistency loss function and uses a sliding window mechanism to construct a historical structural reference, constrains the consistency of statistical information between the current frame and multiple frames, effectively suppresses the temporal accumulation of single-frame errors, prevents global drift, and improves the structural stability of long-sequence dynamic imaging. Attached Figure Description
[0067] Figure 1 This is a schematic diagram illustrating the workflow of an underwater range-gated dynamic imaging method based on metaheuristic attitude estimation according to the present invention.
[0068] Figure 2 This is a schematic diagram of the underwater range-gated imaging system according to an embodiment of the present invention;
[0069] Figure 3 This is a schematic diagram of the underwater range-gated imaging scanning process described in an embodiment of the present invention;
[0070] Figure 4 This is a single-frame scan image of underwater range-gated imaging as described in an embodiment of the present invention;
[0071] Figure 5 This refers to the unregistered dynamic scan fusion image described in the embodiments of the present invention;
[0072] Figure 6Box plots of the final convergence performance of PSO under different parameter settings as described in the embodiments of the present invention;
[0073] Figure 7 This is the convergence curve when the group size is set to 25 according to the embodiment of the present invention;
[0074] Figure 8 This refers to the registered dynamic scan fusion image described in the embodiments of the present invention. Detailed Implementation
[0075] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0076] Example 1, please refer to Figures 1 to 8 This invention provides a technical solution: an underwater range-gated dynamic imaging method based on metaheuristic attitude estimation, comprising the following steps:
[0077] S1: Construct an underwater range-gated laser imaging system. By synchronously controlling laser emission, scanning, and gated exposure, a continuous sequence of spatiotemporally matched raw scan images is obtained. The system consists of a pulsed laser, a gated camera, a scanning mirror, and a synchronous timing circuit. The pulsed laser emits a pulsed laser with a center wavelength of 532nm, which has excellent underwater penetration, effectively improving imaging depth and signal strength. The synchronous timing circuit controls the laser emission, scanning mirror deflection, and time-gated exposure of the gated camera to remain synchronized. After being deflected by the scanning mirror, the laser scans point by point within the camera's field of view, forming a range-gated image at each scanning position. This ultimately constitutes a continuous sequence of spatiotemporally matched raw scan images. Point-by-point scanning achieves high-resolution coverage and forms a spatiotemporally consistent sequence, facilitating subsequent registration and fusion.
[0078] The specific work involves: First, constructing an underwater range-gated laser imaging system, including a pulsed laser, a gating camera, a scanning mirror, and a synchronization timing circuit. The pulsed laser generates short pulses with a center wavelength of 532 nanometers, which has good underwater penetration. The gating camera uses time-gating technology to control the exposure window and suppress water scattering noise. The scanning mirror achieves laser scanning coverage within the camera's field of view through high-speed deflection. All components are precisely synchronized through the synchronization timing circuit to ensure that the laser emission pulse, scanning mirror position switching, and camera exposure gating are aligned in time. A combination of hardware triggering and software calibration ensures that timing errors are controlled at the nanosecond level, guaranteeing complete spatiotemporal matching between image acquisition and laser illumination at each scanning position. After the system starts operating, the laser beam enters the underwater environment after being reflected by the scanning mirror. The camera scans point by point along a preset path within its field of view. Each scanning position corresponds to a specific spatial target area. A laser illuminates the area at that position and triggers a gated camera to expose, forming a range-gated image containing local target information. By continuously controlling the angle of the scanning mirror, the entire target area is covered sequentially, forming a series of image frames in the time dimension. The images are spatially connected and temporally continuous and ordered, forming a complete sequence of original scanned images. Each image contains the target reflected laser information at a specific scanning moment, and background scattering is effectively suppressed. The image sequence as a whole has spatiotemporal consistency. After the scanned image sequence is acquired, the raw data undergoes standardized preprocessing, including image format unification, timestamp alignment, and metadata annotation. All images are stored with the same resolution and coordinate reference system, along with corresponding scanning position, time information, and synchronization parameters.
[0079] S2: Select a reference frame from the original scanned image sequence and determine the effective imaging region. Model the rigid body motion between frames using six-DOF pose parameters. Construct a global optimization objective function based on the mean square error of grayscale to evaluate the alignment of adjacent frames. Select the first clearly imaged frame from the original scanned image sequence as the reference frame, and delineate the effective imaging region containing the laser stripe on the reference frame to achieve a stable benchmark for subsequent registration, avoiding initial deviations introduced by inter-frame differences. Represent the variables to be optimized as six-DOF pose parameters. Use a homogeneous transformation matrix to uniformly model rotation and translation, describing the rigid body motion between adjacent frames during scanning, fully representing the actual motion, and providing a mathematical model basis for accurate spatial alignment. Construct a global optimization objective function based on the mean square error of pixel grayscale, and use the reference frame... With the frame to be registered The mean square error of the grayscale difference at the same pixel position is used as the registration cost. The alignment degree between the reference frame and the frame to be registered in the effective imaging area is quantized. Optimization is driven by quantization error to ensure that the registration result achieves high-precision alignment at the pixel level.
[0080] The specific work involves selecting the first frame with the best imaging quality from the acquired scanned image sequence as the reference for spatial registration. The selection criteria for the reference frame include the overall signal-to-noise ratio of the image, the integrity of the laser strip, and the edge sharpness. Based on this reference frame, the effective imaging area (ROI) is delineated using a grayscale threshold (taking the top 10% of the image grayscale distribution as the boundary of the bright area) and morphological closing operations (3×3 circular structuring elements). This ROI completely includes the main area of the laser scan strip, while excluding invalid pixels at the image boundaries and isolated noise areas caused by water scattering. In practice, the ROI area is controlled within 60%-80% of the image center area to ensure that subsequent pose estimation is performed within the effective signal set. For the inter-frame relative motion caused by platform disturbance or water ripples during the scanning process, a six-degree-of-freedom rigid body motion model is used for parameterization. The three rotational degrees of freedom correspond to the rotation angles around the X, Y, and Z axes, respectively, and the three translational degrees of freedom correspond to the rotation angles along the three axes. The displacement of the coordinate axes is represented by a 4×4 homogeneous transformation matrix, which uniformly represents the rotation and translation operations. The rotation submatrix in the matrix is generated by the Rodriguez rotation formula, and the translation vector directly corresponds to the spatial displacement. In the actual system, the initial estimated values of the pose parameters can be obtained through the scanning mirror angle feedback and the data from the inertial measurement unit. Based on the principle of pixel grayscale consistency within the effective imaging area, a global optimization objective function based on mean square error is constructed. This function calculates the squared grayscale difference between each corresponding pixel in the ROI between the reference frame and the frame to be registered, and averages the differences of all pixels to obtain a quantized alignment error index. In the actual calculation, the pixel grayscale value needs to be linearly normalized first to map the original 12-bit grayscale data to the [0, 1] interval. The error function is calculated by using bilinear interpolation to obtain the grayscale value of non-integer pixel positions. The interpolation kernel size is 2×2 pixels. By optimizing this objective function, the laser stripes of adjacent frames achieve the best alignment in spatial position.
[0081] Furthermore, the global optimization objective function is as follows: Assume there are N effective pixels within the effective imaging region, and denote the grayscale value of the reference frame as... The frames to be registered undergo pose transformation The grayscale value after that The sum of the squared grayscale differences at the same pixel position in two frames is used as the registration cost to construct a global optimization objective function. This function is highly sensitive to small spatial deviations within the effective imaging area and accurately reflects the alignment degree of laser stripes in adjacent frames. Its expression is:
[0082] ;
[0083] In the formula: Let the objective function (fitness function) represent the pose transformation. Registration error under the following conditions; Here is the pose transformation matrix (6 degrees of freedom); The total number of pixels within the effective imaging area (ROI); For pixel index; For the first The coordinates of one pixel; For the reference frame in coordinates The grayscale value at that location; To transform in the current pose Below, the frame to be registered is in coordinates The grayscale value at that location;
[0084] The specific work content is as follows: During the implementation process, record the total number of ROIs. One valid pixel coordinate, For the first pixel coordinates To apply pose transformation to the current frame The grayscale value at that location then clearly defines the N effective pixels within the effective imaging area as the main body of the laser strip, free from background noise and invalid boundary regions, and uses the grayscale value of the reference frame. The grayscale value corresponding to the frame to be registered after pose transformation All data were normalized to the [0, 1] range using 12-bit grayscale data to ensure numerical stability and computational consistency. Meanwhile, bilinear interpolation was used to estimate the grayscale at non-integer pixel locations, with an interpolation kernel size of 2×2 pixels to ensure accurate reconstruction of the target image's grayscale distribution even under small-angle rotation or slight translation.
[0085] ;
[0086] in, It is a homogeneous transformation matrix; Let be the rotation vector, representing the rotation angle about the x, y, and z axes; Indicates transpose. Let the orthogonal rotation matrix be constructed from this rotation vector; Let be a spatial translation vector, representing the displacement in the x, y, z directions; Represents the zero row vector ; Represents a scalar 1, preserving homogeneous coordinates; given the three-dimensional coordinates of any point. After pose transformation, the following is obtained This enables spatial mapping and registration of point sets in a unified coordinate system. The mean of the sum of squared grayscale differences between all valid pixels in two frames is calculated as a quantitative indicator of registration quality, thus constructing a global optimization objective function. This function is highly sensitive to minute shifts in the laser stripe, effectively reflecting inter-frame misalignment caused by platform vibration or water flow. During optimization, pose transformation... The initial value can be set by the scanning mirror angle feedback or inertial measurement unit information through the parametric representation of the six-degree-of-freedom rigid body motion model, and then iteratively corrected by the numerical optimization algorithm until the objective function converges to the minimum value, so as to achieve high-precision spatial alignment between adjacent frames.
[0087] S3: Design an adaptive weighted pixel loss function. This function assigns weights to pixels based on normalized grayscale and gradient information, suppressing the influence of bright areas and enhancing the contribution of structural edges in registration. This improves local alignment accuracy and robustness. The pixel grayscale of the current frame and the reference frame is normalized to eliminate grayscale distribution shifts between frames, effectively mitigating differences in illumination and sensor response, improving inter-frame grayscale consistency, and enhancing registration stability. Based on normalized grayscale values and gradient magnitudes, an exponentially decaying adaptive weight is constructed to suppress the weight of bright areas and enhance the weight of edge areas. This significantly reduces the pulling effect of overly bright pixels on registration, highlighting the contribution of edges and textured regions, and improving structural alignment accuracy. The adaptive weight is then introduced into pixel-level loss calculation to construct the adaptive weighted pixel loss function. To improve the structural consistency and robustness of registration, the registration process focuses more on structural features rather than areas of abnormal brightness, effectively suppressing underwater scattering and reflection interference, and improving the robustness of the algorithm;
[0088] The specific work involves: performing pixel grayscale normalization on the current frame and the reference frame to eliminate inconsistencies in grayscale distribution between frames caused by changes in illumination, water attenuation, or differences in device response; after reading the original 12-bit grayscale image data, aligning the minimum grayscale value of each frame to 0 and the maximum grayscale value to 1 through linear mapping to achieve normalization of the grayscale distribution of the entire sequence; and introducing a minimum regularization term to prevent data instability caused by an excessively narrow grayscale range. Based on the normalization results, a pixel-level adaptive weighting mechanism is further constructed. The weight design comprehensively considers the normalized grayscale value and local gradient information: for high-brightness areas (direct laser illumination areas), an exponential decay function is used to reduce their weight contribution, avoiding overly bright pixels dominating the registration process; for edge and texture-rich areas, gradient magnitude enhancement is used. The weights emphasize the importance of structural features in registration. Gradient calculation uses the Sobel operator for horizontal and vertical convolution to obtain the gradient magnitude of each pixel. The final weights are composed of a weighted combination of gray-level attenuation and gradient enhancement terms, achieving adaptive adjustment of the contribution of different regions. The adaptive weights are integrated into the pixel-level loss calculation to form the final adaptive weighted pixel loss function. When calculating the difference between two frames, this function assigns a weight to each pixel that matches its structural importance, making the optimization process focus more on discriminative edge and texture regions rather than uniform regions affected by brightness changes. This allows the registration algorithm to maintain its sensitivity to structural alignment even when faced with complex situations such as uneven underwater lighting, scattering interference, or target reflection, significantly improving the robustness and structural consistency of the registration results.
[0089] The mathematical definition can be expressed as:
[0090] ;
[0091] ;
[0092] ;
[0093] ;
[0094] ;
[0095] ;
[0096] In the formula: It is an adaptive weighted pixel loss function; Pixel-level adaptive weights; , as well as These are constant parameters that control the influence of the weights; Normalized grayscale values are pixel values. Normalizing grayscale values can reduce grayscale distribution shifts and range differences between different frames, thereby maintaining numerical stability during the optimization process. This is the current grayscale value; This is the maximum grayscale value; The minimum grayscale value; It is a very small constant to prevent division by zero; The gradient is the pixel gradient magnitude. The gradient only focuses on the change in gray level and is not affected by the absolute brightness, which can improve the ability to resist brightness interference. The global gradient mean; It is the sigmoid function; and These represent the horizontal and vertical convolution kernels of the Sobel operator, respectively.
[0097] S4: Introducing a rolling kernel consistency loss function, a rolling accumulation kernel is constructed based on the historical frame information of the sliding window. This constrains the consistency between the current frame and the historical structure mean, suppressing error accumulation and global drift, ensuring temporal stability. A sliding window is used to store the most recent... The registration results of the frames are used, and their average grayscale is calculated as the historical structure reference for the current frame. A sliding window mechanism is employed to effectively suppress short-term noise disturbances, improving the structural stability and noise resistance of the historical reference. A rolling accumulation kernel is constructed using a weighted update method to control the fusion ratio between historical information and the current frame, achieving a dynamic balance between long-term structural trends and recent changes, thus enhancing the system's adaptability to changes in real-world scenes. Based on the grayscale difference between the rolling accumulation kernel and the current frame, a rolling kernel consistency loss function is constructed. To suppress the accumulation of multi-frame errors and global drift, by constraining temporal consistency, registration drift is significantly reduced, and the stability of multi-frame images in the temporal dimension is ensured;
[0098] The specific work involves: during image sequence registration, to maintain temporal consistency and suppress the accumulation of single-frame errors, a fixed length is used. The sliding window mechanism allows the window to cover the previous frame while moving back to the current frame. All registered images of a frame are used to ensure the local timeliness of historical information. Each registered frame is included in a window queue. When a new frame is added, the oldest frame is automatically removed, thus achieving dynamic scrolling updates of the window. Based on all images within the window, the grayscale mean is calculated pixel by pixel to form the historical structural reference image corresponding to the current frame. This historical structural reference image reflects the spatial statistical consistency characteristics of recent frames and has strong noise resistance and structural stability. Based on the historical structural reference image, a rolling accumulation kernel is further constructed. To achieve gradual integration of information over time, the accumulator kernel is updated using an exponentially weighted moving average method. The construction and maintenance of the rolling accumulator kernel involves two steps: firstly, in the... Save the most recent frame using a sliding window Frame sequence And calculate the average gray level based on the moving average frame. Then, the cumulative kernel is updated using a weighted method; based on the rolling cumulative kernel, a rolling kernel consistency loss function is designed. This function is used to apply temporal constraints in multi-frame registration. The rolling kernel consistency loss function calculates the grayscale difference between each pixel in the effective imaging area between the current frame to be registered and the accumulation kernel, and performs a weighted average with adaptive weights. The form is consistent with the adaptive weighted pixel loss function, but the comparison object is replaced by the accumulation kernel instead of the reference frame. By minimizing this loss, each frame is prompted to not only align with the reference frame during the registration process, but also to maintain consistency with the historical structural trend, thereby effectively suppressing the global drift phenomenon caused by the gradual accumulation of single-frame registration errors.
[0099] Inherited The calculation method can be expressed as:
[0100] ;
[0101] ;
[0102] ;
[0103] In the formula: This represents the value of the rolling kernel consistency loss function. Here is the pose transformation matrix; This refers to the number of effective pixels. Pixel-level adaptive weights; The current time is used to accumulate the kernel in pixels. The grayscale value at that location; For pose transformation Then, the frame to be registered is in pixels The grayscale value at that location; For the nearest within the sliding window Average gray level of the frame; The length of the sliding window (in frames); For the first A grayscale image of a frame; The frame index within the sliding window; The current cumulative kernel grayscale value; This is the cumulative kernel grayscale value from the previous time step; As a parameter controlling the kernel update rate, a smaller value causes the accumulator kernel to more closely resemble long-term historical memory, resulting in stronger suppression of single-frame errors. However, this also slows down the response to scene changes, leading to structural lag. On the other hand, a larger value makes the accumulator kernel more sensitive to changes in the most recent frame, which is beneficial for quickly adapting to real-world environmental changes. However, this reduces its tolerance for sudden errors, making it prone to registration drift. Adjusting... It can flexibly balance the system's needs for stability and adaptability, ensuring that the accumulator kernel can both track real changes and effectively smooth short-term disturbances;
[0104] S5: A composite objective function is constructed by fusing the adaptive weighted pixel loss function and the rolling kernel consistency loss function. A particle swarm optimization algorithm is used for global optimization in the six-dimensional pose space. By iteratively updating the particle positions and velocities, the optimal six-degree-of-freedom pose matrix is obtained, achieving high-precision image registration. The adaptive weighted pixel loss function... Consistency loss function with rolling kernel Perform linear weighted fusion to construct a composite objective function. This fusion mechanism integrates local alignment and global stability, effectively suppressing drift and improving registration robustness. Using six-degree-of-freedom pose parameters as the search variables for the particle swarm, several particles are randomly initialized within a preset feasible region. Random initialization covers multi-dimensional space, avoiding the algorithm from getting trapped in local optima and improving global search capability. The particle swarm optimization algorithm is adopted with a composite objective function as the fitness function to perform global optimization in the six-dimensional pose space, iteratively updating the particle position and velocity until convergence or reaching the maximum number of iterations, obtaining the optimal six-degree-of-freedom pose matrix. Iterative optimization gradually improves the pose estimation accuracy, and finally outputs stable and reliable spatial alignment parameters.
[0105] The specific work involves: in the implementation, using a linear weighted fusion method to incorporate the adaptive weighted pixel loss function. Consistency loss function with rolling kernel By combining these elements, a composite objective function is formed. This composite objective function simultaneously considers single-frame structure alignment and multi-frame historical consistency during the optimization process, effectively suppressing drift caused by error accumulation. The parameters can be dynamically adjusted according to the actual scenario. and The values of are selected to adapt to the requirements of registration stability and adaptability in different underwater environments. The six-degree-of-freedom pose parameters are used as search variables for the particle swarm, and the particle swarm is randomly initialized within the preset feasible region to achieve global optimization of the composite objective function. Each particle corresponds to a candidate pose, and its position vector consists of three rotation angles and three translations. During the iteration process, the composite objective function value is used as the fitness evaluation standard. The velocity and position of each particle are updated according to the individual historical best and the global best of the group. By setting inertia weight, learning factor and velocity limit, the algorithm is guaranteed to search efficiently in the six-dimensional pose space and avoid getting trapped in local optima. The optimization process continues until the convergence condition is met or the preset maximum number of iterations is reached. Finally, the optimal pose matrix that minimizes the objective function is output.
[0106] Furthermore, the expression for the composite objective function is as follows:
[0107] ;
[0108] In the formula: It is a composite objective function; It is an adaptive weighted pixel loss function; The rolling kernel consistency loss function; Here is the pose transformation matrix; and These are weight parameters, which control the balance between local alignment accuracy and global temporal consistency, in order to achieve coordinated adjustment among different loss properties. Emphasis is placed on local alignment accuracy. This is beneficial for stability and drift control;
[0109] Furthermore, the global optimization process in the six-dimensional pose space is as follows: five initial population sizes of 15, 20, 25, 30, and 35 are set, and the number of iterations is uniformly set to 300. Comparative experiments show significant differences in convergence performance under different sizes. Ultimately, 25 is chosen to achieve the optimal balance between convergence speed and accuracy. In each iteration, the pose transformation corresponding to each particle is calculated. The composite objective function value is updated to the individual optimal and global optimal solutions. Through an adaptive update strategy, the global search process is significantly accelerated, avoiding getting trapped in local optima. The optimal pose that minimizes the composite objective function is gradually approximated through the velocity and position update formulas, thus achieving high-precision pose estimation and image registration, and significantly improving the spatial consistency of multi-frame registration.
[0110] The specific work involved experimentally optimizing key parameters of the Particle Swarm Optimization (PSO) algorithm to achieve efficient convergence and stable performance. Five initial population sizes (15, 20, 25, 30, and 35) were used as comparative experimental groups. The number of iterations was uniformly set to 300 to balance optimization accuracy and computational efficiency. Each particle represented a six-DOF pose parameter vector, including three rotation angles (in radians) and three translation distances (in pixels). The feasible region for rotation angles was [-0.1, 0.1] radians, and the feasible region for translation distances was [-15, 15] pixels, aiming to cover the range of small pose changes commonly encountered in laser scanning. The initial particle velocity was limited to 20% of the position range to avoid initial randomization leading to search instability. Each independent experiment was repeated 31 times to ensure the repeatability of statistical results. Finally, 25 was selected as the optimal population size, achieving the best balance between convergence accuracy and stability. During each iteration, all particles were traversed, and the pose transformation matrix corresponding to the current particle was used to... The frame image to be registered is back-projected onto the reference coordinate system, and then based on the composite objective function... Calculate the fitness value of the particle, and incorporate adaptive weighted pixel loss into the objective function. and rolling kernel consistency loss Simultaneously determine the weight parameters and maintain time consistency to determine the optimal position of each individual particle. with the global optimal position of the group The fitness value is updated in real time, and the velocity update adopts the standard PSO formula, where the inertia weight is set to a linear decreasing mode (from 0.9 to 0.4), and the cognitive and social learning factors are fixed at 2.0. The updates of position and velocity are truncated according to preset boundaries to ensure that the search process always stays within the effective feasible region. The optimization process is continuously iterated until the convergence condition is met or the maximum number of iterations of 300 is reached. The convergence criterion is: the change in the global optimal fitness value in 10 consecutive iterations is less than 1 / 3. At this point, the algorithm is considered to have converged stably, and the final output is the optimal six-degree-of-freedom pose matrix that minimizes the composite objective function. This completes high-precision registration of the entire sequence of images. The optimal pose is used for subsequent image fusion and reconstruction, enabling stable and continuous imaging of underwater dynamic targets.
[0111] S6: The optimized six-DOF pose matrix is applied to the entire sequence of scanned images, uniformly mapped to the reference coordinate system, eliminating inter-frame spatial bias and distortion. The registered sequence is then subjected to grayscale accumulation, normalization, and brightness enhancement processing to generate underwater dynamic fusion imaging results with continuous structure and clear details. The optimized six-DOF pose matrix is applied to each frame of the entire sequence of scanned images, uniformly mapping each frame to the reference coordinate system to achieve geometric alignment of multiple frames, eliminate spatial bias, and significantly improve image superposition consistency. Grayscale accumulation is performed on the registered scanned image sequence according to pixel position to achieve multi-scan observation. The pixel-level superposition enhancement is measured, and the accumulated result is divided by the number of frames participating in the fusion and normalized to restore the dynamic range of the image, enhance the signal-to-noise ratio of the target area, suppress background noise, avoid brightness saturation, and improve imaging contrast. Brightness enhancement processing is performed on the normalized fused image to improve the visibility and detail expression of weak texture areas. The image structure continuity and edge sharpness are further optimized through local contrast adjustment. Finally, a fused imaging result with continuous structure and clear details is obtained for subsequent recognition and analysis, which enhances edge and texture details, improves image recognition, and provides reliable input for subsequent analysis and reconstruction.
[0112] The specific work involves applying the optimized six-DOF pose matrix frame-by-frame to the entire sequence of scanned images, achieving a unified mapping of each frame in a spatial reference frame. This operation is based on a rigid body motion model, applying rotation and translation transformations to each frame through a homogeneous transformation matrix, projecting it onto a common coordinate system with the reference frame as the base. During this process, bilinear interpolation is used to handle grayscale reconstruction at non-integer pixel positions, ensuring that the image maintains the continuity of grayscale distribution after geometric transformation, and effectively eliminating inter-frame spatial deviations and geometric distortions caused by platform disturbances or water flow, enabling high-precision alignment of multiple laser scan strips in space. After spatial alignment, pixel-level grayscale accumulation and fusion are performed on the registered scanned image sequence. Specifically, the grayscale values at the same pixel position in multiple frames are accumulated to achieve signal superposition from multiple scan observations, enhancing the signal-to-noise ratio and contrast of the target area. After accumulation, the result is divided by the number of frames participating in the fusion. A normalized average grayscale image is obtained to restore the original dynamic range of the image, avoiding brightness saturation or contrast compression caused by direct superposition. This fusion process is performed within the effective imaging area, eliminating interference from background noise and invalid boundary areas, ensuring that the fusion result is concentrated on the target structure area covered by the laser scanning strip. The normalized fused image is further subjected to brightness and contrast enhancement processing to improve the visibility and detail expression of weak texture areas. Adaptive histogram equalization or local contrast stretching methods are used to differentiate the enhancement of different regions of the image, highlighting structural edges and texture features. At the same time, unsharpened masks or gradient domain enhancement techniques are used to further enhance the minute details and contour information in the image, improving edge sharpness and structural continuity. The final fused imaging result has the characteristics of high spatial consistency, rich details and sharp edges, and is suitable for subsequent target recognition, 3D reconstruction or visual analysis applications, improving the reliability of underwater dynamic imaging systems.
[0113] Example 2, as Figures 1 to 8 As shown, based on Embodiment 1, the present invention provides a technical solution: constructing an underwater range-gated laser imaging system and acquiring original image sequences; in underwater dynamic imaging, the acquired single-frame scan images are as follows... Figure 4 As shown, in a single frame image, laser illumination forms a bright scanning strip, revealing the local structure of the target; while the background is significantly suppressed by range gating, water scattering noise is greatly reduced, and the overall image contrast is improved.
[0114] Unregistered scanned image sequences are fused at the pixel level to obtain dynamic fusion results for comparison. Specifically, distance-gated scanned image sequences acquired in chronological order are directly superimposed using a pixel-level fusion method without attitude compensation and geometric registration to demonstrate the impact of underwater dynamic disturbances on the quality of the fused imaging. First, all scanned images are accumulated at pixel positions to achieve pixel-level superposition enhancement of multiple scan observations. Second, the accumulated result is divided by the number of frames involved in the fusion to restore the grayscale values of the fused image to the original dynamic range, thereby avoiding brightness saturation caused by simple superposition. Finally, the normalized image is further processed to enhance overall brightness or local contrast to highlight the structural information of the target area.
[0115] The resulting dynamic scan fusion image, obtained without any registration, is as follows: Figure 5 As shown; from Figure 5 It can be seen that, due to the platform attitude change and environmental disturbance in underwater dynamic imaging, there are significant misalignments and distortions between the scanned image sequences. The final fused image shows obvious ghosting and blurring, the target contour edges are stretched, and the detailed structure is superimposed and smoothed, resulting in a decrease in resolution and difficulty in identifying details. This result shows that relying solely on distance gating to suppress background and noise is not enough to obtain high-quality dynamic fused images. It is necessary to further introduce inter-frame attitude estimation and accurate registration.
[0116] To eliminate ghosting and blurring issues in the images, a particle swarm optimization (PSO) algorithm is introduced to construct a metaheuristic pose estimation module for inter-frame registration of the entire scanned image sequence, specifically including:
[0117] Reference frame selection and ROI determination: The first distance-gated image with complete stripes and clear imaging is selected from the scanned image sequence as the reference frame and used as a common spatial reference system; then, based on the laser stripe position, brightness distribution and imaging range, an effective imaging region (ROI) containing the target area is selected on the reference frame for subsequent error measurement and optimization calculation; by limiting the ROI, the interference of large-area water background and invalid pixels on the registration accuracy can be effectively avoided.
[0118] 6D Pose Parameterization and Objective Function Construction: To characterize the rigid motion relationship between adjacent frames, this invention uses six-degree-of-freedom pose parameters to model the registration transformation, unifying 3D rotation and 3D translation into homogeneous transformation matrices. Among them, among them, Let be the rotation vector. Let be the orthogonal rotation matrix constructed from this rotation vector. Let be a spatial translation vector; given the three-dimensional coordinates of any point. After pose transformation, the following is obtained This enables spatial mapping and registration of point sets in a unified coordinate system; finally, the 6D pose parameters are used as search variables for the particle swarm. As a fitness function, within the predefined feasible region The system performs a global optimization, iteratively updating the particle's position and velocity. When the convergence condition is met or the maximum number of iterations is reached, the optimal value is obtained. The minimum optimal pose is used to achieve high-precision pose estimation and inter-frame image registration for range-gated image sequences;
[0119] Adaptive weighted pixel loss function Construction: Addressing the characteristics of range-gated imaging, namely "extremely sparse background grayscale and excessively bright and highly concentrated laser stripe regions," traditional MSE objective functions tend to overemphasize local bright stripes while neglecting overall contour and structure matching. Therefore, an adaptive weighted pixel loss function is proposed. By assigning adaptive weights to each pixel, the error weights in bright areas are suppressed, and the contribution of edges and low-to-medium gray areas in registration is increased, thereby improving structural consistency and robustness.
[0120] Specifically, the pixel intensities of the current frame and the reference frame are first normalized; let the grayscale value of a pixel be denoted as... The global maximum and minimum gray levels are respectively , Then the normalized grayscale representation is: ,in, To prevent division by zero of extremely small constants;
[0121] Then, an exponentially decaying adaptive weight is constructed by combining the pixel normalization value and gradient information. : ,in, , as well as For constant parameters, For the Sigmoid function, The global gradient mean; The pixel gradient magnitude, which reflects only grayscale changes and does not depend on absolute brightness, can be expressed as: ,in, and These represent the horizontal and vertical convolution kernels of the Sobel operator, respectively, and their specific forms are as follows: Therefore, it can be seen that It decreases exponentially with increasing brightness, but increases at abrupt gradient changes, thus suppressing areas with abnormally high brightness and giving higher weight to structural edges and textured regions.
[0122] Within the ROI, there are a total of There are 100 valid pixels, and the gray level of the reference frame is 100. The current frame will be in pose. Grayscale values are obtained by inverting the projection onto the reference coordinate system. The adaptive weighted pixel loss is defined as follows: ;
[0123] By introducing weights This can significantly reduce the "traction effect" of local bright stripes in a single frame on the overall optimization direction, making pose estimation more focused on contour alignment and structural consistency.
[0124] Rolling kernel consistency loss function Construction: When processing long-term sequences, even if single frames are separated by... Even with good alignment, residual small errors can still accumulate over multiple frames, causing a slow drift in the overall registration result. To constrain the global stability of pose estimation over time, this invention further designs a rolling kernel consistency loss function. By introducing a rolling cumulative kernel Apply smooth constraints to cross-frame structures;
[0125] Specifically, in the first When saving frames, a sliding window is used to store the most recent frame. Frame registration results And calculate its average gray level: Then update the rolling cumulative kernel in a weighted manner: ,in, For kernel update rate control parameters; smaller Making the cumulative kernel closer to the long-term historical average can more effectively suppress occasional mismatches, but it responds more slowly to scene changes; a larger kernel... This makes the cumulative check more sensitive to recent changes, which is beneficial for tracking real dynamics, but reduces the tolerance for noise and sudden errors; the two can be configured in a compromise manner according to the application scenario.
[0126] In terms of pixel-level loss form, continue The weighted structure is defined as: ,in, For the current scrolling accumulation kernel in pixels The grayscale value at the location; this item effectively suppresses the temporal accumulation of short-term mismatch by constraining the consistency between the current frame and the "historical structural mean", avoiding overall drift and structural distortion after multiple frames are superimposed;
[0127] Construction of the composite objective function and global optimization of particle swarm optimization: Combining local alignment accuracy and global stability, the two loss functions mentioned above are linearly combined to construct the final composite objective function. ,in, and These are weighting parameters, used to achieve coordinated adjustment among losses of different properties. Emphasis is placed on local alignment accuracy. This is beneficial for stability and drift control;
[0128] Based on the aforementioned composite objective function, this invention employs a particle swarm optimization algorithm to perform global optimization of the 6D pose; specifically, the six-degree-of-freedom pose parameters... As the particle position vector, several particles are randomly initialized within a preset feasible region, with each particle corresponding to a candidate pose. In each iteration, the objective function value of each particle is calculated separately. As a fitness parameter, the particle velocity and position are updated based on the individual optimal and global optimal solutions, gradually approaching the optimal solution. The optimal pose that yields the minimum value;
[0129] Spatial unification and strip alignment based on 6D pose estimation: In underwater dynamic imaging, platform attitude disturbances and environmental flow can lead to significant spatial deviations between adjacent scan frames, i.e., inter-frame pose differences. If inter-frame stacking is performed directly in the original coordinate system, it will cause strip position misalignment, structural stretching, and severe ghosting, thereby reducing the accuracy of image stitching and target recognition. The 6D pose matrix obtained by optimization based on a composite objective function achieves better results. Laser scanning data collected at different times are uniformly mapped to the same spatial reference system; for any target surface point or strip pixel in any frame, its three-dimensional position in the coordinate system of this frame is determined through pose transformation. It can be mapped to a reference coordinate system; by performing this operation frame by frame on the entire sequence, a spatial overlay effect of multiple frame stripes can be obtained in a unified coordinate system;
[0130] Optimizing PSO Configuration Based on Parameter Analysis and Completing Final Fusion Imaging: To ensure that the aforementioned composite objective function possesses both good global search capability and stable convergence performance within the PSO framework, this invention systematically analyzes key parameters such as the population size of the PSO. Specifically, five initial population sizes of 15, 20, 25, 30, and 35 were set, with a uniform iteration count of 300. To ensure the representativeness and statistical reliability of the results, 31 independent experiments were conducted for each parameter configuration, all using the defined composite objective function. As a fitness evaluation standard;
[0131] like Figure 6As shown in the box plots of the final convergence performance of PSO under different population sizes, the algorithm exhibits the best and most stable convergence characteristics in both resolution board and composite structure target scenarios when the population size is 25. On the one hand, its global optimum is concentrated and has a small variance, indicating that the convergence results of multiple independent runs are highly consistent. On the other hand, when the population size is too small, premature convergence is likely to occur, making it difficult to fully explore the 6D pose space. When the population size is further increased, although the search diversity is enhanced to some extent, additional computational overhead is introduced, and experimental results show that the accuracy decreases and the fluctuations increase. In summary, considering the factors of accuracy, stability and efficiency, the standard initial population size of PSO is determined to be 25.
[0132] Furthermore, such as Figure 7 As shown, with a population size of 25, this invention evaluated the convergence speed and convergence path of PSO in a resolution board and composite structure target scenario; the results show that in the early stage of iteration, the composite objective function... The ability to decrease rapidly and approach a stable value within a few iterations indicates that the pose estimation converges quickly. In the later stages of iteration, the objective function curve fluctuates less and no obvious oscillations are observed, indicating that the algorithm has good robustness and convergence stability when using composite objective functions.
[0133] After completing the parameter optimization and pose estimation, the "registered" scan image sequence is re-fused pixel-level. Specifically, all registered distance-gated images are summed by pixel position, then normalized by dividing by the number of frames involved in the fusion. Finally, combined with brightness or local contrast enhancement processing, a dynamic scan fusion image based on metaheuristic pose estimation is obtained, such as... Figure 8 As shown; with Figure 5 Compared to the results of unregistered fusion, Figure 8 The target structure achieved high-precision alignment, stripe ghosting was basically eliminated, the edge contour was clear and continuous, and small and medium scale details were completely and sharply restored. This proves that the overall scheme of "composite objective function + metaheuristic attitude estimation + range gating fusion" proposed in this invention can significantly improve the spatial consistency and imaging quality of range-gated laser imaging in complex underwater dynamic environments.
[0134] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0135] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. An underwater range-gated dynamic imaging method based on metaheuristic attitude estimation, characterized in that, Includes the following steps: S1: Construct an underwater range-gated laser imaging system to acquire a continuous sequence of spatiotemporally matched raw scan images by synchronously controlling laser emission, scanning, and gated exposure. S2: Select reference frames from the original scanned image sequence and determine the effective imaging area. Use six-degree-of-freedom pose parameters to model the rigid body motion between frames and construct a global optimization objective function based on the mean square error of gray level. S3: Design an adaptive weighted pixel loss function that assigns weights to pixels based on normalized grayscale and gradient information to suppress the influence of highlight areas. Specifically, this includes: The pixel grayscale values of the current frame and the reference frame are normalized to eliminate grayscale distribution offset between different frames. An exponentially decaying adaptive weight is constructed based on normalized gray values and gradient magnitude to suppress the weight of bright areas and enhance the weight of edge areas. Adaptive weights are introduced into pixel-level loss calculation to construct an adaptive weighted pixel loss function. ; The The mathematical definition can be expressed as: ; ; ; ; ; ; In the formula: It is an adaptive weighted pixel loss function; Pixel-level adaptive weights; , as well as These are constant parameters; Normalized grayscale values for pixels; This is the current grayscale value; This is the maximum grayscale value; The minimum grayscale value; It is a very small constant to prevent division by zero; This represents the pixel gradient magnitude. The global gradient mean; It is the sigmoid function; and These represent the horizontal and vertical convolution kernels of the Sobel operator, respectively. S4: Introducing a rolling kernel consistency loss function, a rolling accumulation kernel is constructed based on the historical frame information of the sliding window to constrain the consistency between the current frame and the historical structure mean, suppressing error accumulation and global drift. Specifically, it includes: Save recent data using a sliding window The registration results of the frames are used to calculate the average gray level as a historical structure reference for the current frame. A rolling accumulation kernel is constructed using a weighted update method to control the fusion ratio between historical information and the current frame; Based on the grayscale difference between the scrolling accumulation kernel and the current frame, a scrolling kernel consistency loss function is constructed. To suppress the accumulation of multi-frame errors and global drift; The Inherited The calculation method can be expressed as: ; ; ; In the formula: This represents the value of the rolling kernel consistency loss function. Here is the pose transformation matrix; This refers to the number of effective pixels. Pixel-level adaptive weights; The current time is used to accumulate the kernel in pixels. The grayscale value at that location; For pose transformation Then, the frame to be registered is in pixels The grayscale value at that location; For the nearest within the sliding window Average gray level of the frame; The length of the sliding window; For the first A grayscale image of a frame; The frame index within the sliding window; The current cumulative kernel grayscale value; This is the cumulative kernel grayscale value from the previous time step; For kernel update rate control parameters; S5: A composite objective function is constructed by fusing the adaptive weighted pixel loss function and the rolling kernel consistency loss function. The particle swarm optimization algorithm is used to find the global optimization in the six-dimensional pose space. By iteratively updating the particle position and velocity, the optimal six-degree-of-freedom pose matrix is obtained to achieve high-precision image registration. S6: The optimized six-DOF pose matrix is applied to the full sequence of scanned images, uniformly mapped to the reference coordinate system, and grayscale accumulation, normalization and brightness enhancement are performed on the registered sequence to generate underwater dynamic fusion imaging results.
2. The underwater range-gated dynamic imaging method based on metaheuristic attitude estimation according to claim 1, characterized in that: S1 includes the following steps: An underwater range-gated laser imaging system is constructed, comprising a pulsed laser, a gating camera, a scanning mirror, and a synchronization timing circuit. The pulsed laser emits pulsed laser light with a center wavelength of 532 nm. The laser emission, scanning mirror deflection, and time-gated exposure of the gating camera are kept synchronized by a synchronous timing circuit. After being deflected by the scanning mirror, the laser scans point by point within the camera's field of view, forming a distance-gated image at each scanning position, ultimately creating a continuous spatiotemporally matched sequence of original scan images.
3. The underwater range-gated dynamic imaging method based on metaheuristic attitude estimation according to claim 1, characterized in that: S2 includes the following steps: The first frame with clear imaging is selected from the original scanned image sequence as the reference frame, and an effective imaging area containing laser stripes is delineated on the reference frame. The variables to be optimized are represented as six-degree-of-freedom pose parameters. Homogeneous transformation matrices are used to uniformly model rotation and translation, describing the rigid body motion between adjacent frames during the scanning process. Construct a global optimization objective function based on the pixel grayscale mean square error, and use the reference frame... With the frame to be registered The mean square error of the grayscale difference at the same pixel location is used as the registration cost to quantize the alignment degree between the reference frame and the frame to be registered within the effective imaging area.
4. The underwater range-gated dynamic imaging method based on metaheuristic attitude estimation according to claim 3, characterized in that: The global optimization objective function is specifically as follows: Suppose there are N effective pixels in the effective imaging area, and the grayscale value of the reference frame is denoted as N. The frames to be registered undergo pose transformation The grayscale value after that ; Using the sum of the squared grayscale differences at the same pixel position in the two frames as the registration cost, a global optimization objective function is constructed, the expression of which is: ; In the formula: The objective function is... Here is the pose transformation matrix; The total number of pixels within the effective imaging area; For pixel index; For the first The coordinates of one pixel; For the reference frame in coordinates The grayscale value at that location; To transform in the current pose Below, the frame to be registered is in coordinates The grayscale value at that location.
5. The underwater range-gated dynamic imaging method based on metaheuristic attitude estimation according to claim 1, characterized in that: S5 includes the following steps: Adaptive weighted pixel loss function Consistency loss function with rolling kernel Perform linear weighted fusion to construct a composite objective function. : Using the six-degree-of-freedom pose parameters as the search variables for the particle swarm, a number of particles are randomly initialized within a preset feasible region. The particle swarm optimization algorithm is used with a composite objective function as the fitness function to perform global optimization in a six-dimensional pose space. The particle position and velocity are iteratively updated until convergence or the maximum number of iterations is reached to obtain the optimal six-degree-of-freedom pose matrix.
6. The underwater range-gated dynamic imaging method based on metaheuristic attitude estimation according to claim 5, characterized in that: The expression for the composite objective function is as follows: ; In the formula: It is a composite objective function; It is an adaptive weighted pixel loss function; The rolling kernel consistency loss function; Here is the pose transformation matrix; and These are weight parameters, which control the balance between local alignment accuracy and global timing consistency.
7. The underwater range-gated dynamic imaging method based on metaheuristic attitude estimation according to claim 5, characterized in that: The process of global optimization in the six-dimensional pose space is as follows: Five initial population sizes of 15, 20, 25, 30, and 35 were set, and the number of iterations was uniformly set to 300. In each iteration, the pose transformation corresponding to each particle is calculated. The composite objective function value is used to update the individual optimal and global optimal solutions; By gradually approximating the optimal pose that minimizes the composite objective function through velocity and position update formulas, high-precision pose estimation and image registration are achieved.
8. The underwater range-gated dynamic imaging method based on metaheuristic attitude estimation according to claim 1, characterized in that: S6 includes the following steps: The optimized six-DOF pose matrix is applied to each frame of the full sequence of scanned images, and each frame of the image is uniformly mapped to the reference coordinate system. The grayscale values of the registered scanned image sequence are accumulated according to the pixel position to achieve pixel-level superposition enhancement of multiple scan observations. The accumulated result is divided by the number of frames involved in the fusion and normalized to restore the dynamic range of the image. The normalized fused image is subjected to brightness enhancement processing, and the image structure continuity and edge sharpness are further optimized by local contrast adjustment, so as to obtain a fused imaging result with continuous structure and clear details.