An underwater terrain measurement method based on laser radar and vision fusion

By acquiring multi-source data from underwater lidar and visual images, as well as environmental sensing data, an environmental interference weight matrix is ​​generated. The feature extraction algorithm and parameters are adaptively adjusted to solve the mismatch problem between lidar point clouds and visual images in extreme underwater environments, achieving high-precision and robust spatial registration.

CN120949254BActive Publication Date: 2026-07-21PEARL RIVER WATER RESOURCES PROTECTION INST
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
PEARL RIVER WATER RESOURCES PROTECTION INST
Filing Date
2025-09-02
Publication Date
2026-07-21

AI Technical Summary

Technical Problem

Existing technologies have limited adaptability to the multi-scale feature automatic matching process between laser point clouds and visual images in extreme underwater environments, resulting in high mismatch rates and spatial registration failures, making it difficult to meet the needs of fine terrain measurement.

Method used

By acquiring raw point cloud data from underwater lidar, visual images, and in-situ physical environment sensing data associated with multiple locations, an environmental interference weight matrix is ​​generated. Feature extraction algorithms and parameters are adaptively selected, and mismatch filtering and iterative optimization are performed in conjunction with physical constraints to improve the stability of feature matching.

Benefits of technology

It significantly improves the accuracy and reliability of automatic registration in complex underwater environments, reduces the false matching rate, enhances the robustness of feature matching, supports rapid response under extreme conditions such as high turbidity and low illumination, and expands the applicability of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120949254B_ABST
    Figure CN120949254B_ABST
Patent Text Reader

Abstract

The application relates to an underwater terrain measurement method based on laser radar and vision fusion, which realizes the spatio-temporal consistency of multi-source data by synchronously collecting underwater point clouds, images and physical environment sensing data and giving unified spatio-temporal labels, adopts pretreatment, denoising, scale normalization and feature enhancement and other means to improve the quality of different modal data, combines an environment perception model to construct an environment interference weight matrix, adaptively adjusts feature extraction and matching algorithm parameters according to different environment states, realizes high-reliability feature matching between point clouds and images through the establishment of multi-scale feature description, spatial consistency and physical constraint criteria, and finally improves the robustness and precision of spatial registration through environment-driven iterative optimization and fusion. The scheme can stably realize high-precision multi-modal data automatic registration in an underwater dynamic environment, has strong adaptability, and effectively improves the environment perception and spatial measurement quality.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of topographic surveying technology, and in particular to an underwater topographic surveying method based on the fusion of lidar and vision. Background Technology

[0002] Currently, the fields of underwater surveying and terrain reconstruction are placing increasing demands on measurement technologies that integrate multimodal sensor information. Underwater LiDAR can acquire 3D point cloud information with high precision, while underwater visual imaging technology can capture rich visual textures and structures. To overcome the perception bottleneck of single sensing methods in specific scenarios and achieve accurate measurement of complex underwater environments, spatial registration and multimodal fusion of LiDAR point clouds and visual images have become key technical directions for international research and application. At present, mainstream technologies mostly adopt registration methods based on manual or automated feature extraction, such as multi-scale feature operators like SIFT, SURF, and ORB, as well as deep learning or adaptive matching strategies, to spatially align point clouds and images. These methods perform well under conditions of high visibility and relatively ideal aquatic environments, and can already achieve high-precision registration and terrain modeling in nearshore and clear water areas. In addition, in recent years, the industry has also gradually promoted the automatic fusion measurement technology of multi-source information in scenarios such as autonomous navigation of underwater robots and marine geological surveys, driving the engineering transformation of related registration algorithms.

[0003] The core problem with existing solutions lies in the limited environmental adaptability of the multi-scale feature automatic matching process. Because key features of laser point clouds and visual images are easily affected by physical factors such as underwater scattering, absorption, and illumination variations, their spatial distribution and appearance can change significantly in extreme environments. Conventional multi-scale matching algorithms struggle to distinguish between real features and environmental artifacts, easily leading to a sharp increase in mismatch rates and spatial registration failures. Even with multimodal deep fusion, current mainstream methods still rarely consider the dynamic impact of in-situ physical environment information on feature extraction and matching processes. They lack quantitative modeling of the physical correlation and environmental disturbances between point cloud and image multi-source data. Consequently, in low visibility and high environmental complexity conditions, the robustness and accuracy of automatic registration are insufficient to meet the practical needs of fine topographic surveying and automatic calibration. Summary of the Invention

[0004] This application provides an underwater topography measurement method based on the fusion of lidar and vision, which aims to solve one of the problems or issues of the prior art mentioned in the background.

[0005] This application provides an underwater topography measurement method based on the fusion of lidar and vision, specifically including:

[0006] S1: Acquire raw point cloud data from underwater lidar, underwater visual images, and in-situ physical environment sensing data associated with multiple locations, including variables such as turbidity, illuminance, and temperature. Assign synchronous spatiotemporal labels to each set of acquired data to achieve the spatiotemporal consistency of multi-source data.

[0007] S2: Preprocess the raw laser point cloud data and visual image data under the synchronization tag, including denoising, scale normalization and signal enhancement operations, to adapt to the impact of different environmental conditions on the quality of point cloud and image data.

[0008] S3: Input the pre-processed point cloud data, image data, and physical environment sensing data into the environment perception model to generate an environment interference weight matrix, which is used to characterize the influence of the current underwater environment on feature extraction by different sensing methods.

[0009] S4: Based on the environmental interference weight matrix, adaptively select multi-scale feature extraction algorithms for point clouds and images or adjust feature parameters, including adjusting point cloud feature scale, curvature calculation window, image texture extraction operator and region effective range, to achieve feature extraction decision based on environmental perception.

[0010] S5: Perform preliminary feature matching processing on the adaptively extracted multi-scale point cloud features and image features to generate multi-level spatial candidate matching pairs, and record the corresponding paired environment perception weight parameters.

[0011] S6: For the above candidate matching pairs, perform mismatch filtering under physical constraints based on the environmental perception weight parameters, including using spatial consistency judgment and optical path constraints to screen out matching relationships that do not conform to the actual environmental distribution.

[0012] S7: Iterative optimization and environment-driven heterogeneous feature fusion are performed on the feature matching pairs filtered by mismatches. Regional physical labels are used to enhance the spatial consistency between point clouds and images, and weighted optimization of the final matching pairs is generated.

[0013] S8: Using the final matching pair as input, perform high-precision spatial registration between the laser point cloud and the visual image to obtain a registration parameter set that has been fused with environmental perception weights, thereby improving the stability of underwater multi-scale automatic matching.

[0014] S9: Implement real-time dynamic monitoring. When environmental sensor data or registration accuracy parameters change, determine whether to switch feature extraction algorithms, adjust parameters, or restart the registration process based on the environmental perception model, so as to adapt to changes in working conditions and continuously optimize registration performance.

[0015] This application provides an underwater topographic measurement method based on the fusion of lidar and vision, which has the following advantages:

[0016] (1) Significantly improve the accuracy and reliability of automatic registration in complex underwater environments. By introducing an environmental perception model and a weight matrix, the automatic matching of point clouds and visual images can dynamically adjust weights and parameters according to the environmental conditions, thereby reducing matching errors and improving spatial registration accuracy under extreme conditions such as high turbidity, low illumination, and temperature changes.

[0017] (2) Enhance the robustness of multimodal feature matching and effectively suppress noise interference and mislabeling. Multiple filtering based on physical constraints, spatial geometric consistency, and optical path constraints enables feature matching pairs to form high-confidence cross-mappings in spatial structure and physical distribution, reducing the mismatch rate by more than 30% in extreme environments. Regional clustering weighting further strengthens spatial consistency and enables automatic feature selection under highly complex working conditions.

[0018] (3) Improve feature extraction and matching efficiency and support rapid response in large-scale, dynamic environments. Through real-time data synchronization and environment-driven parameter adaptation, the feature algorithm can automatically adjust when the environment changes, avoiding the problems of manual intervention and lag of static parameters.

[0019] (4) Expanding the scope of application and enhancing the practical engineering application value of the system. The patented method breaks through the application limitations of existing lidar-vision fusion registration under harsh underwater conditions such as high turbidity, high temperature, and high dynamics. It can be widely used in various highly complex and multimodal underwater topographic measurement scenarios such as river surveying, marine engineering, lake surveys, and tunnel inspection. It does not rely on the prior conditions of a single sensor, thus improving the versatility and universality of multi-sensor underwater measurement products. Attached Figure Description

[0020] Appendix Figure 1 This is the main flowchart of an underwater topographic measurement method based on the fusion of lidar and vision.

[0021] Appendix Figure 2 This is a sub-flowchart of an underwater topographic measurement method based on the fusion of lidar and vision. Detailed Implementation

[0022] Embodiments of the present invention are described in detail below, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and should not be construed as limiting the present invention.

[0023] The following disclosure provides many different embodiments or examples for implementing different structures of the invention. To simplify the disclosure, specific examples of components and arrangements are described below. Of course, these are merely examples and are not intended to limit the invention. Furthermore, reference numerals and / or letters may be repeated in different examples; such repetition is for simplification and clarity and does not in itself indicate a relationship between the various embodiments and / or arrangements discussed. In addition, examples of various specific processes and materials are provided in this invention, but those skilled in the art will recognize the application of other processes and the use of other materials.

[0024] As attached Figure 1 As shown, this application provides an underwater topography measurement method based on the fusion of lidar and vision, specifically including:

[0025] S1: Acquire raw point cloud data from underwater lidar, underwater visual images, and in-situ physical environment sensing data associated with multiple locations, including variables such as turbidity, illuminance, and temperature. Assign synchronous spatiotemporal labels to each set of acquired data to achieve the spatiotemporal consistency of multi-source data.

[0026] S2: Preprocess the raw laser point cloud data and visual image data under the synchronization tag, including denoising, scale normalization and signal enhancement operations, to adapt to the impact of different environmental conditions on the quality of point cloud and image data.

[0027] S3: Input the pre-processed point cloud data, image data, and physical environment sensing data into the environment perception model to generate an environment interference weight matrix, which is used to characterize the influence of the current underwater environment on feature extraction by different sensing methods.

[0028] S4: Based on the environmental interference weight matrix, adaptively select multi-scale feature extraction algorithms for point clouds and images or adjust feature parameters, including adjusting point cloud feature scale, curvature calculation window, image texture extraction operator and region effective range, to achieve feature extraction decision based on environmental perception.

[0029] S5: Perform preliminary feature matching processing on the adaptively extracted multi-scale point cloud features and image features to generate multi-level spatial candidate matching pairs, and record the corresponding paired environment perception weight parameters.

[0030] S6: For the above candidate matching pairs, perform mismatch filtering under physical constraints based on the environmental perception weight parameters, including using spatial consistency judgment and optical path constraints to screen out matching relationships that do not conform to the actual environmental distribution.

[0031] S7: Iterative optimization and environment-driven heterogeneous feature fusion are performed on the feature matching pairs filtered by mismatches. Regional physical labels are used to enhance the spatial consistency between point clouds and images, and weighted optimization of the final matching pairs is generated.

[0032] S8: Using the final matching pair as input, perform high-precision spatial registration between the laser point cloud and the visual image to obtain a registration parameter set that has been fused with environmental perception weights, thereby improving the stability of underwater multi-scale automatic matching.

[0033] S9: Implement real-time dynamic monitoring. When environmental sensor data or registration accuracy parameters change, determine whether to switch feature extraction algorithms, adjust parameters, or restart the registration process based on the environmental perception model, so as to adapt to changes in working conditions and continuously optimize registration performance.

[0034] Step S1: Acquire raw point cloud data from underwater lidar, underwater visual images, and in-situ physical environment sensing data associated with multiple locations, including variables such as turbidity, illuminance, and temperature. Assign synchronous spatiotemporal labels to each set of acquired data to achieve spatiotemporal consistency of multi-source data. Specifically, this includes:

[0035] S1.1: Configure the underwater lidar system deployed within the target measurement area with preset acquisition parameters, including setting indicators such as laser emission band, sampling rate, and spatial resolution, to ensure that the acquired raw point cloud data has sufficient spatial coverage and accuracy, providing a key laser point cloud acquisition source for subsequent feature extraction and spatial registration.

[0036] Based on the requirements for underwater topographic measurement of the target measurement area, the input conditions include the pre-deployed underwater lidar system, the geospatial boundary of the area to be measured, basic environmental parameters, and measurement task requirements.

[0037] The parameter configuration module is used to set the core acquisition indicators of the lidar system, including the laser emission band λ (unit: nanometers). The 532nm or 1064nm band is selected according to the different optical propagation attenuation characteristics underwater, so as to achieve dynamic adaptation to the penetration and reflection characteristics of the target water body.

[0038] The point cloud sampling rate f is set through the parameter management unit. s (Unit: Hz) Combining the survey area A, the predetermined spatial sampling interval d, and the operation speed v, the system automatically calculates the optimal acquisition and scanning frequency required to ensure that the spatial resolution R of the original point cloud data meets the requirements for subsequent feature extraction.

[0039] Furthermore, through spatial resolution control algorithms, the scanning angle θ and sampling intervals Δx and Δy of the laser emission array are adaptively adjusted based on the three-dimensional distribution characteristics and terrain change gradient of the survey area, thereby achieving high-density point cloud coverage of the terrain undulation area.

[0040] The system calibration submodule is used to register the parameters of the lidar spatial alignment and initial attitude (such as installation attitude angles α, β, γ), and generate the corresponding spatial configuration parameter package to ensure the geometric accuracy and spatial consistency of the collected point cloud.

[0041] A laser energy self-test and echo signal gating mechanism is adopted, by setting the echo signal threshold T. e Real-time monitoring of laser echo quality and automatic adjustment of emission power P emit It outputs high signal-to-noise ratio and complete coverage of original point cloud data.

[0042] Through the above parameter configuration, spatial correction and energy adaptation steps, the lidar acquisition source is preset to an acquisition mode that meets the requirements of spatial coverage, three-dimensional density and geometric accuracy, so as to realize the acquisition of raw lidar point cloud data based on high-standard acquisition indicators, and provide a high-quality, physically consistent raw point cloud dataset for subsequent multi-source sensing fusion processing and spatial registration algorithms.

[0043] For example, in a river topographic survey scenario with a water depth of 20m and an operating area of ​​100m×100m, a 532nm laser emission band is used, and the point cloud sampling rate f is set. s =200kHz, spatial resolution R=5cm, laser emission array scanning angle θ=120°, sampling interval Δx=0.05m, Δy=0.05m, radar installation attitude angle α=0°, β=0°, γ=0°. Through system self-test, the laser emission power P... emit =2W, echo signal threshold T e =0.3V. The actual collected raw point cloud covers the entire terrain, with 40 million point clouds. The spatial positioning error is better than ±1cm, which meets the technical requirements of multi-scale feature extraction and precise spatial registration. It significantly improves the spatial accuracy and reliability of topographic surveying in subsequent multi-source fusion and harsh environments.

[0044] S1.2: Set the sampling period and scene exposure parameters for the visual camera acquisition units in the same area, acquire underwater visual images based on high-sensitivity low-light imaging technology, extract metadata such as resolution, scene viewpoint, and timestamp, and obtain the original visual image dataset, laying the image dimension foundation for subsequent multimodal feature fusion.

[0045] For a visual camera acquisition unit in the same area, the input conditions include preset geographical boundaries of the measurement area, water depth distribution of the scene to be measured, environmental parameters such as high turbidity and insufficient lighting, as well as the spatial positioning results of the front-end lidar.

[0046] Using the sampling period configuration method (parameter: sampling period T) cam Scene motion speed vscene This allows for adaptive setting of the continuous frame interval of the visual sensor. The sampling period is calculated using the following formula:

[0047]

[0048] Where, d scene v represents the desired image spatial sampling interval in meters. scene The velocity of the relevant features within the scene is expressed in meters per second.

[0049] Through the exposure parameter setting algorithm (parameter: exposure time E) exp Gain G cam Target Illuminance target This achieves adaptive compensation for ambient light intensity and turbidity. Exposure time is dynamically adjusted based on the gray-scale averaging method, satisfying the following constraints:

[0050]

[0051] Among them, E base As the baseline exposure time, I target To set illuminance targets, I current Detect the illuminance for the current scene.

[0052] Employing high-sensitivity low-light imaging technologies (such as back-illuminated CMOS, EMCCD, etc.) improves the image signal-to-noise ratio and detail resolution in low-light environments, and enhances the ability to perceive extreme underwater environments.

[0053] Furthermore, the scene perspective configuration algorithm (parameter: main viewpoint height) is used. view Camera array spatial distribution C pos This achieves optimized distribution of visual acquisition coverage. It utilizes a 3D geometric model to automatically adapt to the camera placement angle, improving the integrity of image acquisition for complex underwater terrain.

[0054] The visual data metadata extraction module is used to perform resolution R on each frame of acquired image data. img Scene perspective (heta) img Exposure parameter E exp Gain G cam Timestamp T stamp Metadata is extracted, processed, and stored in a standardized manner as an image metadata identifier package.

[0055] Through the above configuration and acquisition process, a raw visual image dataset containing timestamps, spatial location information, optical parameters, and environmental labels is obtained, providing a high-quality image dimensional foundation for subsequent multimodal feature extraction and spatial registration algorithms.

[0056] For example, in a typical river measurement scenario with a water depth of 20m and a turbidity of 20 NTU, a 12-bit back-illuminated CMOS vision camera is configured with a field of view of [heta]. view =90°, sampling period T cam =0.5 seconds, exposure time E exp =33ms (dynamically adjustable range 25-80ms), gain G cam =1.5 times, with synchronous calibration position with LiDAR. The acquired image resolution is 1920×1080 pixels, with each frame accompanied by a timestamp and spatial tag. In practical applications, during operation cycles of up to 2 hours, the system automatically and dynamically adjusts exposure and sampling parameters based on changes in ambient illumination, achieving high signal-to-noise ratio and high-resolution image acquisition. The acquired raw visual image dataset exhibits an image positioning error better than ±1cm and image noise intensity less than 5%, meeting the requirements for adapting to multi-scale environmental features and significantly improving data input quality in low-light, weak, and highly turbid environments.

[0057] S1.3: Data acquisition and scheduling of physical environment sensing modules deployed at measurement points are performed to obtain in-situ physical environment sensing data such as turbidity, illuminance, and temperature in the area. During the acquisition process, multi-channel synchronous acquisition and signal calibration technology are used to ensure the timeliness of data and the accuracy of spatial location identification, forming an environmental sensing data stream, which provides the underlying physical input for subsequent weighting of the environmental perception model.

[0058] S1.4: Based on the synchronization mechanism of data acquisition, the raw point cloud data of lidar, visual image data and physical environment sensing data are respectively labeled with data timestamps and spatial location information with unified specifications to form a consistent spatiotemporal metadata dataset of the three types of data, so as to realize cross-modal and cross-dimensional data traceability and management control in the multi-source sensing data acquisition process.

[0059] S1.5: Perform data integration processing on laser point cloud data, visual image data and physical environment sensing data that have been attached with synchronous spatiotemporal labels. Use the spatiotemporal label index to verify the consistency of the correlation between the data groups. Detect and filter out data samples that are spatiotemporally asynchronous or have spatial positioning errors due to acquisition anomalies. Generate a verified multi-source synchronous sensing dataset to provide high-quality and clearly defined basic data for the subsequent multi-scale feature environment perception and registration process.

[0060] Step S2: Preprocessing is performed on the original laser point cloud data and visual image data under the synchronization tag, including denoising, scale normalization, and signal enhancement operations, to adapt to the impact of different environmental conditions on the quality of point cloud and image data. Specifically, this includes:

[0061] S2.1: Perform spatial noise filtering on the laser point cloud data under the synchronization label, and use statistical anomaly detection algorithms (such as radius filtering and outlier removal) to obtain high-purity laser point cloud data after noise suppression, so as to provide a reliable data foundation for subsequent scale normalization and feature enhancement.

[0062] The input data is raw laser point cloud data with synchronous spatiotemporal labels, collected from the underwater lidar subsystem after parameter preset (including laser emission band, sampling rate, spatial resolution, etc.).

[0063] Radius filtering method (parameter: spatial radius R) filter Minimum number of neighborhood points N min This function removes isolated and low-density anomalies from a point cloud. Each point P is determined according to the following formula. i If the neighborhood contains n points i If the threshold is not met, the point is considered a noise point and is removed.

[0064]

[0065] Where I is the indicator function, R filter To define the spatial radius, N min To retain the threshold, P i and P j These are spatial coordinates.

[0066] Furthermore, an outlier removal algorithm is used (parameter: mean distance threshold D). mean Outlier α out This method removes outliers with high spatial mean distances, ensuring spatial consistency of the point cloud. It calculates the average distance from each point to its neighborhood center.

[0067]

[0068] like Then it is determined to be an outlier. Where k is the number of nearest neighbors, σ D This represents the distance between the global mean and the standard deviation.

[0069] Spatial distribution analysis based on statistical characteristics (parameters: spatial partitioning, point density distribution ρ) cell (Local extreme value detection threshold) enables compliance screening of point clouds in areas with abnormal distribution and obtains judgment labels to supplement the noise filtering results.

[0070] Continuing, an iterative noise residue removal mechanism is employed (parameter: maximum number of iterations N). iter The two algorithms are repeated multiple times on the initially filtered point cloud data (with a convergence threshold), until the proportion of noise points drops below the convergence threshold.

[0071] By combining spatial noise filtering algorithms with statistical outlier removal and other multi-step processing methods, the original laser point cloud data from the previous step is transformed into a high-purity, spatially uniform point cloud dataset free of isolated abnormal noise, thus enabling the data input conditions for subsequent scale normalization and multi-scale feature enhancement.

[0072] For example, in a practical application of topographic surveying of a river channel with a water depth of 20m, the spatial resolution of the input raw laser point cloud data is 0.05m, and the point cloud density in the dense point area is 400 points / m. 2 There are a high proportion of isolated noise points generated by suspended matter in the water flow and boundary reflections. A radius filtering method is applied, with parameter R... filter =0.08m, N min =8, after processing, 260,000 isolated noise points were removed. An outlier removal algorithm was used, D mean =0.06m, α out =2.0, nearest neighbor number k=10, identify and remove 120,000 external outliers. In spatial density analysis, points with a density lower than 150 points / m² are considered. 2 All points in the affected area are removed to eliminate false noise, significantly improving the consistency of the point cloud distribution across the entire region. An iterative removal mechanism is used, with N set... iter =3, ∈=1%, and iterative processing is performed to compress the proportion of noise points to 0.6% of the total number of points. The final output high-purity laser point cloud dataset contains 38.5 million points with a residual noise rate of less than 1% and uniform spatial distribution. This provides an ideal data foundation for subsequent scale normalization and spatial feature enhancement processing, improving the accuracy and robustness of subsequent feature extraction and registration.

[0073] S2.2: A multi-resolution image enhancement algorithm is used for the visual image data under the synchronization tag. Combined with underwater image dehazing and color shift correction, the image contrast and variability are improved to obtain a highly recognizable visual image after signal enhancement, thus preparing information for the scale normalization and feature extraction stages.

[0074] S2.3: Based on the obtained high-purity laser point cloud data and enhanced visual images, scale normalization processing is performed. A unified scaling factor is used to map the multi-source data to a spatially consistent scale framework, forming a unified standard scale point cloud set and a standard size image, creating a comparable data structure for subsequent multi-scale feature extraction algorithms.

[0075] S2.4: Apply spatial signal enhancement techniques to standard-scale point cloud sets, including multi-scale curvature enhancement and point density enhancement algorithms, to refine the geometric signals of feature regions, thereby generating feature-sensitive enhanced laser point cloud data and enhancing the point cloud's ability to distinguish key geomorphic areas.

[0076] S2.5: Adaptive texture enhancement and local artifact correction are performed on standard-sized visual images. Multi-channel contrast enhancement technology is superimposed with local image denoising algorithm to generate enhanced visual image data with clear feature hierarchy, providing high signal-to-noise ratio input for subsequent construction of environmental perception weight matrix and multi-scale feature extraction.

[0077] Step S3: Input the pre-processed point cloud data, image data, and physical environment sensing data into the environmental perception model to generate an environmental interference weight matrix, which is used to characterize the influence of the current underwater environment on feature extraction by different sensing methods. Figure 2 As shown, it specifically includes:

[0078] S3.1: Aggregate the denoised and scale-normalized lidar point cloud data, visual image data, and in-situ physical environment sensing data labeled with synchronous spatiotemporal tags to form a composite input feature set for the environment perception model.

[0079] S3.2: Based on the composite input feature set, the physical environment sensing data is decomposed by the environmental factor feature decoupling algorithm to distinguish the degree of individual influence of turbidity, illumination, temperature and other factors on the distribution of point cloud features and the performance of image texture features, and to obtain the influence weight parameters of physical environment factors.

[0080] The composite input feature set is assembled from the denoised and scale-normalized lidar point cloud data, visual image data, and in-situ physical environment sensing data marked with synchronous spatiotemporal tags to form the basis for environmental perception analysis of multi-source data.

[0081] An environmental factor feature decoupling algorithm is adopted (parameter settings include the set of environmental variables {Turbidity, Illuminance, Temperature}, sensing channel type, spatial sampling label, etc.) to achieve independent attribution mapping of environmental factors in physical environment sensing data. Each in-situ physical environment data is finely decomposed by channel to obtain the configuration information of environmental variables in multi-source sensing data.

[0082] Furthermore, the response model is influenced by environmental factors (parameter: point cloud response coefficient β). pc Image response coefficient β img The environmental baseline value (E0) is used to perform an influence transfer analysis between variables and multimodal sensing methods. Based on the signal physical mechanism, the individual weights of different environmental factors on the distribution of point cloud features and the performance of image texture features are calculated.

[0083] The influence weight transformation formula is used to establish a quantitative relationship between physical environmental factors and each feature channel, as defined below:

[0084]

[0085] Among them, w mod,i Let f be the weight of the i-th environmental factor in modal mod (pc is the point cloud, img is the image). mod E is the environment-channel influence function. i Let be the measured value of the i-th environmental variable, and N be the total amount of all environmental factors.

[0086] Based on the above weight decomposition, the ability of each physical environmental factor (such as turbidity, illuminance, and temperature) to independently discriminate and quantify the distribution of characteristics of each sensing mode is realized, and the original physical environmental variables are transformed into the influence weight parameters of point cloud and image channel.

[0087] By normalizing the weight parameters of environmental factors, a standardized algorithm (parameters: interval [0,1], weighted mean μ target 0.5) is used to eliminate the differences in weight intervals between channels for the attribution distribution results of different environmental factors in all perception channels.

[0088] By using environmental factor decomposition and weight normalization algorithms, the composite feature set from the previous step is transformed into weight parameters that individually affect the physical environment factors of point cloud distribution and image texture representation, thereby enabling precise control of the perception mode by physical environment factors.

[0089] For example, a laser point cloud and vision system were deployed simultaneously to acquire data in an underwater engineering survey area. On-site physical environment sensors detected turbidity of 16 NTU, illuminance of 22 Lux, and temperature of 17°C. The physical response coefficient β of the point cloud channel was set. pc The image channel β is [1.0, 0.3, 0.1] (corresponding to turbidity, illuminance, and temperature). img The input values ​​are [0.2, 1.0, 0.1]. The influence response model is used to evaluate the input (E) of each channel. turb =16E illu =22E temp =17) Decompose the factors separately and calculate the original weights w of the point cloud channel environmental factors. pc,turb =0.543,w pc,illu =0.163,w pc,temp =0.054, Image channel environmental factor weight w img,turb =0.087,w img,illu =0.870,w img,temp=0.043. The weight parameters of the above multi-channel and multi-factor systems are normalized and mapped to the standard distribution interval [0,1] to obtain the independent weight of the physical environment on each channel. The final output weight parameters are used for subsequent feature fusion and adaptive adjustment of the environmental perception model, realizing the controllability of the physical dependence on the performance of multi-channel features under complex underwater environmental conditions, and providing data and algorithmic support for improving the registration accuracy and robustness of the multi-source measurement system.

[0090] S3.3: Using the influence weight parameters of physical environmental factors as input, multi-source data fusion modeling is performed within the environmental perception model, including laser point cloud saliency analysis and visual image feature saliency extraction, to generate an environmental adaptability influence factor matrix that distinguishes point cloud and image features.

[0091] The preprocessed laser point cloud data, enhanced visual image data, and the influence weight parameters of physical environmental factors obtained through attribution decomposition are input into the environmental perception model to form a composite input feature set of multi-source data fusion.

[0092] The laser point cloud saliency analysis method (with parameter settings including curvature statistical window, point density threshold, spatial neighborhood radius, etc.) is used to identify point sets with prominent geometric features for input point cloud data and calculate their spatial saliency distribution index to achieve accurate saliency detection of feature regions.

[0093] Furthermore, through a visual image feature saliency extraction algorithm (parameters including multi-scale texture filtering operator, brightness gradient enhancement factor, and local contrast window size), the edge, texture, and region brightness saliency are separated at different scales for the input enhanced visual image, generating a multi-level saliency distribution matrix of the image.

[0094] A weighted multi-source saliency fusion algorithm is employed to jointly encode the geometric saliency parameters of the laser point cloud and the texture saliency parameters of the visual image. A mathematical mapping relationship is established by incorporating the influence weight parameters of physical environmental factors, and environmentally adaptive weighting is applied to the saliency data of each dimension. The environmentally weighted fusion is achieved using the following formula:

[0095] S env (i,j)=w pc (i)·S pc (i,j)+w img (i)·S img (i,j)

[0096] Among them, S env (i,j) represents the adaptability significance score of the j-th feature under the i-th environmental condition, w pc (i) represents the environmental weights of the point cloud features under this working condition, S pc(i,j) represents the significance value of the j-th feature of the point cloud, w img (i) represents the environmental weights of the image features under this working condition, s img (i,j) represents the significance value of the j-th feature of the image.

[0097] Using the weighted significance scores mentioned above, an environmental adaptability influence factor matrix is ​​established for all point cloud and image features, which reflects in detail the recognition priority and performance intensity of each sensing mode feature under specific environmental physical parameter conditions.

[0098] By using a multi-source saliency fusion algorithm based on the environmental perception model, the fused environmental adaptability saliency matrix is ​​used as the output to quantify the adaptability of point cloud and image features in different environments, providing basic data for weight normalization and spatial location mapping.

[0099] For example, in a real underwater measurement scenario, the lidar point cloud acquisition configuration is set to a spatial resolution of 0.01m, a curvature statistics window of 10 points, and a density threshold of 5 points / m². The visual image uses a 640×480 resolution, with multi-scale texture filtering sizes of 3×3 and 5×5, and a brightness gradient enhancement factor of 1.5. The physical environment sensing module detects turbidity of 5 NTU, illuminance of 20 Lux, and temperature of 16℃. Attribution decomposition yields a point cloud environment weight of 0.75 and an image environment weight of 0.25. In the significance analysis, the point cloud obtains a significant distribution S. pc (i,j), image obtained S img (i,j). Calculate S using the above formula. env For (i,j), as turbidity increases (e.g., to 15 NTU), the point cloud environment weight automatically increases to 0.90, while the image weight decreases to 0.10, and the saliency score matrix changes accordingly. The final output environmental adaptability influence factor matrix accurately reflects the adaptability of each feature to the current physical state, providing quantitative input for subsequent weight normalization and feature extraction parameter decisions. Under various environmental conditions, after processing by the environmental perception model, the dynamic adjustment capability of the feature saliency distribution is significantly improved, providing reliable assurance for the robustness and matching accuracy of spatial registration under harsh underwater conditions.

[0100] S3.4: Based on the environmental adaptability influence factor matrix, a weight normalization and spatial location mapping algorithm is used to complete the standardization transformation of the influence factors, so that each feature extraction unit can adaptively respond according to the actual intensity of environmental disturbances in subsequent processing.

[0101] Based on the generated environmental adaptability influencing factor matrix, the input data includes the laser point cloud saliency matrix, the visual image feature saliency matrix, and the physical environment factor weight parameters.

[0102] A weighted normalization algorithm (parameters: weight range of each channel [0,1], normalization target mean μ=0.5, standard deviation σ=0.1) is adopted to achieve dimensionless standardization of environmental adaptability influencing factors, and uniformly map the significance scores from different sources to the normalized interval.

[0103] Furthermore, through a spatial location mapping algorithm (parameters: spatial partition label, spatiotemporal coordinate index, and physical reference table of impact factors), the normalized impact factors are partitioned according to the physical spatial layout and collection point labels, and a mapping table from spatial location to feature weights is constructed to achieve the spatial standard positioning of impact factors.

[0104] Furthermore, by using the influence factor distribution reconstruction algorithm (parameters: Gaussian smoothing kernel width h, minimum neighborhood size k), the spatial distribution of standardized influence factors is continuously interpolated and reconstructed with scale consistency, thereby improving the adaptability of each feature unit to the continuous spatial changes of environmental disturbances.

[0105] A dynamic weight adjustment module is adopted to automatically adjust the standardized weights of each feature channel based on the current spatiotemporal label and the dominant environmental factors for the physical environment range that undergoes drastic changes, thereby achieving dynamic sensitivity response of key points and regions.

[0106] By using the standardized influence factor matrix and the spatial mapping results as input data for subsequent multi-scale feature extraction and parameter adjustment, a basic structure for environmental adaptive feature extraction across space and channels is realized.

[0107] For example, the input environmental adaptability influence factor matrix is ​​generated by a point cloud and image saliency scoring system. The original point cloud saliency score ranges from [0, 16], and the image saliency score ranges from [0, 255]. A weight normalization objective is set using the following normalization formula:

[0108]

[0109] Among them, S env (i, j) represents the adaptive significance score of the j-th feature in the i-th environment, min(S env ) and max(S env ) represents the global minimum and maximum values ​​among all features.

[0110] The original score of a certain feature in the point cloud is 12, and the original score of the image feature is 210. The global minimum value is 0, and the maximum value is 255. After normalizing the two types of features, the normalized score of the point cloud feature is 0.047, and the normalized score of the image feature is 0.824.

[0111] Using spatial labels for data collection points, the survey area is divided into three major regions: A, B, and C. Based on the weight distribution of each region, the normalized score is mapped to the spatial location matrix.

[0112] For weight fluctuations caused by local turbidity changes, the dynamic weight adjustment module detects abnormal turbidity in region A and automatically increases the weights of all point cloud features in region A to 0.9, while decreasing the weights of image channels to 0.1. The final output of the standardized, spatially labeled, and dynamically adaptive influence factor matrix provides a highly consistent and applicable physical foundation and reference for subsequent adaptive adjustment of feature extraction parameters and multi-channel feature matching, effectively improving the system robustness and matching accuracy in complex underwater environments.

[0113] S3.5: Perform interactive environmental factor weighting operations on the standardized influence factor matrix and the preprocessed feature set, and combine synchronous spatiotemporal labels to generate the environmental interference weight matrix, providing decision input for subsequent multi-scale feature extraction algorithm selection and parameter adjustment.

[0114] Step S4: Based on the environmental interference weight matrix, adaptively select a multi-scale feature extraction algorithm for point clouds and images or adjust feature parameters, including adjusting the point cloud feature scale, curvature calculation window, image texture extraction operator, and region effective range, to achieve environmentally perceptive feature extraction decision-making. Specifically, this includes:

[0115] S4.1: Analyze the input environmental interference weight matrix, and extract the optimal weight configuration for point cloud feature extraction and image feature extraction under the current working conditions based on the reflection of the perception mode effectiveness of each physical parameter (such as turbidity, illuminance, and temperature) in the weight matrix, so as to realize the environmental-driven feature extraction decision input.

[0116] Based on the generated environmental disturbance weight matrix, the input data includes multi-source physical environmental factor parameters (such as turbidity, illuminance, and temperature), standardized spatial location labels, preprocessed laser point cloud data, and visual image data.

[0117] An environmental weight deconstruction analysis algorithm (parameter configuration: weight matrix dimension d, environmental parameter normalization threshold setting range [0,1]) is adopted to realize the analysis of the effectiveness of each physical parameter in the environmental interference weight matrix on point cloud and image perception methods.

[0118] Furthermore, by using a feature performance attribution calculation method (parameters include: point cloud feature category index, image feature category index, and feature working condition adaptation threshold), the influence of various physical environment parameters (such as turbidity, illuminance, and temperature) on the performance of point cloud features and image features under the current working conditions is evaluated, and a feature-environment parameter coupling performance table is formed.

[0119] An optimal weight allocation algorithm is adopted, and the response capability of each feature channel is numerically normalized using an environmental adaptability performance table. The optimal weights of point cloud / image features are derived based on the following mathematical model:

[0120]

[0121] in, To specify the optimal feature weights in the perceptual mode, I mod α is an indicator of the model's environmental adaptability to target characteristics. mod M is the mode priority factor, and M is the total number of sensing channels.

[0122] Furthermore, through the working condition configuration response mechanism (parameters: dynamic weight adjustment rate λ, extreme value switching threshold θ), the weight configuration results are adjusted in real time to adapt to the environment, and the optimal feature extraction weight structure is automatically matched under sudden environmental changes or dominant parameter switching scenarios.

[0123] The optimal weight configuration generation module writes the extraction results into the feature extraction decision table, forming an adaptive optimal weight configuration for point cloud feature channels and image feature channels under environment-driven conditions.

[0124] Through the above chain derivation method, the environmental interference weight matrix is ​​logically analyzed and numerically transformed to output a multi-channel feature extraction weight reorganization that conforms to the current working conditions, thereby realizing feature extraction decision input driven by the physical environment.

[0125] For example, at an underwater operation site, the deployed physical environment sensing module detects a turbidity of 12 NTU, an illuminance of 10 Lux, and a temperature of 14°C, corresponding to an input weight matrix of [0.85 (point cloud), 0.15 (image)]. In the configuration parameters, the point cloud operating condition adaptability index I... pc =0.9, preference factor α pc =1.0, Image I img =0.2, α img =0.8. Based on the weighted derivation formula:

[0126]

[0127] The dynamic weight adjustment module compares the current weight with the preset extreme value switching threshold θ = 0.8, detects that the dominant environmental parameter is "high turbidity", and automatically outputs a decision to increase the point cloud weight and suppress the image weight.

[0128] The optimal weight configuration output dynamically allocates parameters for downstream point cloud and image multi-scale feature extraction, and then...

[0129] Called in S4.2 and S4.3, feature extraction optimization prioritizes point cloud channels over image channels, ultimately enhancing the accuracy and adaptability of multi-source data fusion and registration in low-visibility environments.

[0130] S4.2: Based on the weight configuration obtained in S4.1, a multi-scale point cloud feature extraction method is applied to the preprocessed point cloud data. By adaptively adjusting the point cloud feature scale and curvature calculation window, the multi-condition adaptation of the point cloud feature vector is realized, and the multi-scale feature descriptor of the point cloud consistent with the environmental state is output.

[0131] Based on the environmental interference weight configuration output by S4.1, the input parameters include the preprocessed laser point cloud dataset, synchronous spatiotemporal labels, dynamic environmental weight factors, and spatial location labels.

[0132] A multi-scale geometric feature extraction method is adopted (parameter: feature scale set L = {l1, l2, ..., l...). n}, initial curvature window W0), parallel multi-scale feature construction of laser point cloud data, to achieve full coverage of structural information at different spatial levels.

[0133] Furthermore, an adaptive feature scale adjustment algorithm is used (parameters: environmental weight threshold γ, minimum / maximum scale interval [λ]). min ,λ max The optimal feature scale λ* for each spatial location is calculated based on the current environmental weight configuration, as follows:

[0134] λ * (x, y, z) = λ min +(λ max -λ min )·w pc (x, y, z)

[0135] Where, λ * (x, y, z) represents the adaptive optimal feature scale at the point (x, y, z), w pc (x, y, z) represents the environmental point cloud weights for the aforementioned spatial locations.

[0136] Adaptive configuration module using curvature calculation window (parameter: adaptive window width k) * (Lower limit k1, upper limit k2), and dynamically determine the number of local curvature statistical points based on environmental conditions:

[0137] k * (x, y, z) = k min +(k max -k min )·w pc (x, y, z)

[0138] For each sampling location, an adaptive-scale neighborhood point set is used to perform principal curvature decomposition and normal vector feature extraction algorithms, forming a multi-scale, environment-adaptive geometric feature vector f(x, y, z; λ). * k * ).

[0139] Furthermore, through an environment-weighted feature saliency weighting algorithm (parameter: saliency amplification coefficient α), the basic geometric features (such as curvature, point density, and morphological gradient) are enhanced with environmental sensitivity according to weight coefficients, generating weighted multi-scale feature descriptors:

[0140] F env (x, y, z) = α·w pc (x, y, z)·f(x, y, z; λ) * k * )

[0141] The aforementioned feature descriptors are aggregated in the entire point cloud space according to synchronous spatiotemporal labels, and spatial label normalization is performed to achieve adaptive output of local and global multi-scale features.

[0142] By employing a multi-level feature consistency detection and elimination algorithm, point set features with low reliability or those that are significantly inconsistent with environmental weights are automatically filtered out. This effectively suppresses false features in highly turbid and low signal-noise regions, enabling highly stable multi-scale point cloud feature extraction for complex and ever-changing environments.

[0143] Through an environment-driven, adaptively adjusted multi-scale feature extraction and processing chain for point clouds, the input preprocessed point cloud data is transformed into a feature description subset that highly corresponds to the current physical state of the underwater environment. This provides a consistent and robust point cloud feature foundation for subsequent multimodal feature weight fusion and feature matching, achieving optimal adaptation of point cloud feature extraction strategies and parameters under environmental perception conditions.

[0144] For example, in an underwater topographic survey area, the environmental perception model outputs a turbidity of 18 NTU, an illuminance of 7 Lux, and an environmental weight of 0.93 for the point cloud channel in real time. The system sets a minimum feature scale λ. min =0.02m, maximum characteristic scale λ max =0.10m, curvature calculation window lower limit k1=8, upper limit k2=22. According to the above adaptive adjustment formula, at a certain measurement point (x0,y0,z0), the characteristic scale is calculated as follows:

[0145] λ * (x0, y0, z0)=0.02+(0.10-0.02)×0.93=0.0944,m

[0146] Curvature window adjustment:

[0147] k * (x0, y0, z0)=8+(22-8)×0.93=19.02

[0148] Within a neighborhood of λ* = 0.0944m and k* = 19, point cloud principal curvature decomposition and normal vector feature extraction generate environmentally optimal geometric feature vectors. The weight significance amplification factor is set to α = 1.2, and the final output is a feature descriptor subset. This step achieves strong environmentally adaptive feature extraction in highly turbid and low-illuminance areas, effectively suppressing noise interference. Compared to static fixed parameters, the number of reliable feature regions is increased by 30%, and the matching accuracy is improved by more than 12% under extreme conditions, providing high-confidence point cloud feature data for subsequent multimodal fusion and registration.

[0149] S4.3: Combining the environment-driven weight configuration of S4.1, the normalized image data is processed using a multi-scale image texture extraction operator. The parameters of the texture extraction operator and its effective area are dynamically adjusted to achieve adaptive extraction of image texture features and output a multi-scale image feature descriptor under environment awareness.

[0150] Normalized visual image data and environment-driven weight configuration serve as the input basis for this sub-step.

[0151] A multi-scale image texture extraction operator configuration algorithm is adopted (parameters: texture operator type, such as Gabor filter, local binary mode (LBP), multi-scale Sobel edge operator; the operator's scale domain range is 2×2, 3×3, 5×5, etc.) to achieve fine-grained segmentation and hierarchical filtering of texture information in different physical space intervals for the input normalized image.

[0152] Furthermore, the operator parameters are automatically adjusted by an adaptive adjustment module (parameters: number of operator directions, frequency bandwidth, region window size, and environmental weight response factor λ) in conjunction with the environmental weight configuration. In the high-illuminance weight range, a larger region window and high-frequency direction filtering are prioritized to enhance edges and high-level textures; in the low-illuminance or high-turbidity weight range, the effective window is reduced to improve the operator response sensitivity and highlight microscopic local texture changes.

[0153] An algorithm for locating and allocating effective regions is adopted (parameters: spatial region label, influence factor threshold, automatic partitioning threshold γ). Combined with environmental weight configuration and spatial location label, key areas for texture extraction are delineated for each partition of the image. This maximizes the texture information acquisition density in high-adaptive-weight areas and suppresses the acquisition of low-weight edge areas to prevent feature redundancy and mis-extraction caused by environmental interference.

[0154] Furthermore, by using a multi-scale feature stitching method (parameters: scale fusion factor β, feature description length n), the texture descriptors under each interval and parameter configuration are fused and encoded to form a high-dimensional image texture feature descriptor set containing multi-scale, multi-region, and multi-parameter responses.

[0155] A feature-adaptive evaluation and weight normalization algorithm (parameters: descriptor distribution statistics variable, environmental index weight vector) is used to perform environmental correlation verification and distribution normalization on all output texture descriptors to ensure that the descriptors can reflect the texture detail structure and interference response characteristics dominated by the current operating conditions.

[0156] By adaptively adjusting the multi-scale image texture extraction operator and controlling the effective range of the region, the original visual image data is transformed into a multi-scale image feature descriptor driven by the environment-aware weight, realizing automatic optimization and highly robust output of texture features under different physical environments.

[0157] For example, in a measurement scenario with turbidity of 10 NTU, illuminance of 18 Lux, and temperature of 15℃, the environmental weight configuration is [point cloud 0.80, image 0.20]. The LBP operator window size is set to 3×3, and image region A is defined in the high-weight area. The Gabor operator has 8 directions and a frequency bandwidth σ of 1.2. The effective area is automatically located as a 100×100 pixel area centered on region A. In this high-weight region, the texture feature extraction density is increased to 1.5 times the normal level. The multi-scale feature fusion factor β is set to 0.7, resulting in a 128-dimensional multi-scale image feature descriptor. A normalization algorithm is applied to the output texture descriptor.

[0158]

[0159] Among them, F norm (i) is the i-th dimension normalized feature descriptor, F raw (i) represents the original feature value of the i-th dimension. After processing, when the turbidity increases to 20 NTU, the texture descriptor density of region A automatically decreases, the operator window shrinks to 2×2, and the fusion factor drops to 0.5. The output feature descriptors all reflect the reduction in texture resolution under high environmental interference. Finally, the multi-scale image feature descriptors generated in this embodiment accurately reflect the changes in texture structure and feature distribution under different physical parameters, and can be used as a highly adaptive input for downstream feature weight fusion and registration algorithms.

[0160] S4.4: The point cloud multi-scale feature descriptor generated in S4.2 and the image multi-scale feature descriptor generated in S4.3 are simultaneously input into the feature weight fusion module. The participation ratio of various features is dynamically adjusted according to the environmental interference weight matrix to form a weight-normalized multimodal feature joint vector, providing weight-sensitive feature representation for feature matching.

[0161] A feature weight fusion method is adopted (parameter: multi-scale feature description subset of point cloud (f)). pc ={f{pc,1},f{pc,2},...}), image multi-scale feature descriptor subset (fimg ={f{img,1},f_{img,2},...}), environmental interference weight matrix w env =[w pc ,w img This enables the collaborative input of multi-source features under environmental perception.

[0162] Furthermore, through feature dimension alignment and standardization algorithms (parameters: high-dimensional feature normalization factor η, feature remapping dimension (N)), the point cloud feature descriptors and image feature descriptors are normalized and encoded into a unified feature vector space to eliminate the inconsistency in feature distribution caused by changes in measurement scale and environmental conditions, thus establishing a comparability basis for subsequent multimodal fusion.

[0163] Furthermore, an environmental weighted fusion algorithm (parameter: weighting coefficient w) is used. pc ,w img The weights of the multi-source feature set (derived from the environmental interference weight matrix of the current working condition) are dynamically allocated according to the following weighted fusion formula:

[0164] F joint (k)=w pc ·f pc (k)+w img ·f img (k)

[0165] Among them, F joint (k) is the k-th dimension of the joint vector, f pc (k),f img (k) represent the normalized feature values ​​of the point cloud and the image in the k-th dimension, respectively, w pc ,w img These are the weight parameters under real-time environmental awareness.

[0166] Furthermore, a normalized joint feature vector generation algorithm (parameters: normalization target interval [0,1], joint feature dimension M) is used to normalize the weighted summed feature components, completing the multimodal weight normalization and scale consistency of different feature sources under environmental perturbation, and obtaining the final multimodal feature joint vector:

[0167]

[0168] Among them, F norm_joint (k) represents the k-th dimension normalized multimodal joint feature component.

[0169] Furthermore, by using a joint feature adaptive verification algorithm (parameters: adaptive confidence threshold γ, anomaly detection threshold β), dynamic consistency detection is performed on each dimension of the joint vector with the original environmental indicators, redundant feature components are eliminated, and the key feature set that dominates the current environment is retained, thus achieving weight-sensitive feature expression.

[0170] Through the above fusion and adaptation process, the multi-scale features of point clouds and images are mapped, fused and normalized under the action of the environmental weight matrix. The final output is a multimodal joint feature vector that can be used for downstream feature matching and spatial registration, realizing the environmental adaptive expression and dynamic weight allocation of the feature space of multi-source data.

[0171] For example, in an underwater survey area, the point cloud multi-scale feature descriptor is generated through the preceding step S4.2, containing 10 normalized feature components such as curvature and density. The image multi-scale feature descriptor is generated through S4.3, including 8 feature components such as edge response, LBP, and Gabor. The environmental interference weight matrices are respectively (w pc =0.88,w img =0.12). For each dimension, the point cloud features are as follows (f pc (k)=0.71), image features such as (f img (k) = 0.23). Calculate the joint eigencomponents using the weighted fusion formula:

[0172] F joint (k)=0.88×0.71+0.12×0.23=0.6308

[0173] For all feature component sets {F joint (1),...,F joint (18)} Perform normalization to obtain F norm_joint (k) A multimodal joint feature vector distributed in the interval [0,1]. Adaptive verification is performed dimension-by-dimensionally on the joint vector. Due to the reduced adaptive threshold of high-frequency features in high-turbidity scenes, 3D edge high-frequency features are automatically removed, ultimately outputting a 15-dimensional high-confidence multimodal joint feature. In the subsequent feature matching process, this joint vector achieves an environment-adaptive feature expression dominated by point cloud features and assisted by image features when facing highly turbid and low-light areas, significantly improving the accuracy of cross-modal matching. In experimental scenarios, compared with the unweighted fusion strategy, this method reduces the feature matching error rate in high-interference blocks by 28% and improves the overall registration stability by more than 15%.

[0174] S4.5: Based on the changes in dynamic environmental factors, the multimodal feature joint vector output by S4.4 is used to automatically adjust the parameters of the feature extraction algorithm iteratively. This includes real-time response to newly detected environmental extremes (such as sudden turbidity or high illumination changes) and outputting an updated set of feature descriptors to ensure that feature extraction always maintains optimal adaptability in the dynamic environment.

[0175] Step S5: Perform preliminary feature matching processing on the adaptively extracted multi-scale point cloud features and image features to generate multi-level spatial candidate matching pairs, and simultaneously record the corresponding paired environment-aware weight parameters. Specifically, this includes:

[0176] S5.1: Perform feature vector normalization processing on the multi-scale point cloud features and multi-scale image features that have been adaptively optimized. Use the normalized vector features as input. The algorithm used is feature vector normalization mapping and distance metric transformation to eliminate the influence of scale and distribution inconsistency on subsequent matching algorithms and generate a normalized feature descriptor subset.

[0177] S5.2: Based on the normalized feature descriptor subset, multi-scale spatial neighborhood clustering and KNN (k-nearest neighbor) algorithm are used to establish candidate matching tables for point cloud features and image features in their respective feature spaces, so as to automatically detect feature pairs with the minimum Euclidean distance or the maximum correlation coefficient across modalities and obtain the first round of multi-level spatial candidate matching relationships.

[0178] The input consists of a multi-scale feature descriptor subset of the laser point cloud after feature normalization and a multi-scale feature descriptor subset of the visual image.

[0179] A multi-scale spatial neighborhood clustering method is adopted (parameter: feature space distance threshold d). th The feature dimension (N) and spatial region label (L) are used to automatically identify and construct feature subspaces in the point cloud feature space and image feature space based on the Euclidean distance between feature vectors, thus achieving preliminary partitioning of similar features.

[0180] Furthermore, the KNN (k-nearest neighbor) algorithm (parameters: number of nearest neighbors k, distance metric dist) is used. func Within each feature subspace, for each point cloud feature descriptor f pc,i Calculate its k nearest neighbor correspondences f in the image feature space. img,j And record the minimum Euclidean distance in sequence:

[0181]

[0182] Wherein, D(f) pc,i f img,jLet be the Euclidean distance between the i-th descriptor of the point cloud and the j-th descriptor of the image in the N-dimensional feature space.

[0183] Furthermore, based on the distance threshold criterion d th Or the maximum correlation coefficient ρ max Filter to satisfy D(f) pc,i f img,j )≤d th Or the correlation coefficient ρ(f) pc,i f img,j )≥ρ th The feature descriptor pairs are used to establish a multi-level spatial candidate matching relationship table, involving cross-modal candidate feature pairing at different scale levels.

[0184] Furthermore, by using the matching relationship index structure, the candidate feature pairs are grouped with information at different scales and spatial label information to generate a first-round multi-level spatial candidate matching pair set.

[0185] By combining the above-mentioned multi-scale spatial neighborhood clustering and KNN algorithm with normalized Euclidean distance or correlation coefficient screening and label aggregation processing, the cross-modal multi-scale feature pairs of point clouds and images are initially matched and structured candidate relationships are formed, realizing the systematic expression of the spatial multi-level distribution of the matching.

[0186] For example, in an underwater topographic survey area, the input laser point cloud feature space, normalized by S5.1, is 1024-dimensional, with each feature point belonging to a specific scale group. A multi-scale spatial clustering distance threshold d is set. th =0.15, the threshold value of the feature correlation coefficient ρ th =0.9, k=5 in KNN. In actual testing, in the laser point cloud features, each f pc,i The KNN algorithm is used to select the top 5 f-values ​​with the smallest distance in the image feature space. img,j An initial pairing set is formed. After screening using the Euclidean distance criterion, an average of approximately 10,000 3D spatial candidate matching pairs are obtained. Further filtering using correlation coefficients ultimately retains approximately 5,000 high-confidence cross-modal preliminary matching relationships, and a multi-level spatial candidate table is constructed based on multi-scale grouping information. The candidate matching pair set output in this step provides a high-coverage and structured matching foundation for subsequent processes such as geometric consistency constraints and environmental perception weight parameter binding. In actual testing, even in highly murky and low-contrast environmental sections, it can still ensure that the number of candidate matches is not less than 80% of the core area, laying the foundation for point cloud-image registration in complex environments.

[0187] S5.3: For the first round of multi-level spatial candidate matching relationships, use the geometric consistency judgment algorithm of multi-scale feature regions in the global space (such as spatial centroid matching and consistent direction vector verification) to process and eliminate obvious mismatch terms caused by structural distortion, and output multi-scale spatial candidate matching pairs after geometric consistency constraints.

[0188] For the first round of multi-level spatial candidate matching relationships initially established by multi-scale spatial neighborhood clustering and KNN algorithm, the input data are normalized multi-scale point cloud feature descriptors, image feature descriptors and corresponding spatial labels and scale grouping information.

[0189] A multi-scale feature region spatial centroid matching algorithm is adopted (parameter: feature region spatial centroid coordinate set). Using coordinate weight coefficient γ and region scale label set L, the geometric centroid position of the region to which each level of spatial candidate matching belongs is accurately analyzed. The spatial centroid of the point cloud feature sub-region and the image feature sub-region are calculated using the following formula:

[0190]

[0191] Where, n l m m These represent the number of feature points in the point cloud and the image within the l-th and m-th scale groups, respectively. and The coordinates of the point cloud and image feature points in their respective spatial coordinate systems.

[0192] Furthermore, the evaluation method using Euclidean distance from the spatial centroid position (parameter: spatial distance threshold D) is employed. center This function calculates the spatial centroid distance between candidate matching pairs, using the following formula:

[0193]

[0194] Among them, T pc→img This is a pre-registration projection transformation from point cloud space to image space. If D center(1,m) ≤d center If the region matches, it is determined that the spatial centroid is consistent.

[0195] Furthermore, a direction vector consistency verification algorithm is adopted (parameter: direction vector set (v)). pc ,v img ), included angle threshold θ max This performs a direction consistency check on the principal direction vectors of the corresponding matching pairs. The principal direction vectors of the feature sub-regions are then calculated respectively.

[0196]

[0197] in, and These are the dominant direction vectors of point cloud and image feature points, respectively. (Determination) and included angle cosine value:

[0198]

[0199] like Then it is determined that the directions are consistent.

[0200] The method for detecting anomalous sub-regions of structural distortion (parameter: distortion metric δ) skew Spatial distribution symmetry threshold T sym The system calculates the regional morphological statistics for each pair of matching relationships, automatically removes matching relationships with point clouds or image height structural distortion, and further filters out spatial mismatches caused by extreme distortion.

[0201] Through the above-mentioned multi-level spatial geometric consistency judgment and elimination process, the initial candidate matching pair set is transformed into multi-scale spatial candidate matching pairs that have been optimized by centroid consistency, orientation consistency and structural distortion elimination.

[0202] By using a multi-scale geometric consistency constraint algorithm, the aforementioned preliminary matching relationship is accurately transformed into a multi-level candidate matching pair with high spatial consistency, which significantly reduces mismatches at the spatial structure level and creates a high-quality spatial consistency foundation for subsequent physical environment weight assignment and heterogeneous feature fusion.

[0203] For example, in an underwater topographic survey task, the input is a laser point cloud region A (n1 = 50 feature points) with normalized centroid coordinates of (12.0, 8.5, -3.0) m, and an image region B (m1 = 60 feature points) with a normalized centroid of (440, 320) pixels. Using a pre-registration transformation, the centroid of the point cloud is projected onto the image plane, and then... (The rest of the text is missing).

[0204]

[0205] Setting d center =10 pixels, actual spatial distance Regions A and B are determined to have the same spatial centroid. Further directional consistency calculations are performed to determine the principal direction of the point cloud regions. Image main direction Projected angle (Φ=8°<)

[0206] θ max =12°), consistent with the judgment. Structural distortion factor (δ) skew =0.12<Tsym =0.2), passing distortion detection. After multi-level geometric consistency processing, spatial candidate matching pairs in regions A and B are retained. Pairs in the other extreme region C are eliminated because the centroid distance exceeds 15 pixels, and an abnormally deformed region D is screened out because the directional angle is 32°. Finally, after this step, 92% of the high spatial consistency matching pairs are retained for subsequent environmental perception weight assignment. Compared with the case without geometric consistency screening, the mismatch rate is reduced by more than 20%, and the local accuracy of spatial registration is improved by 0.7 cm, effectively supporting high-precision automatic feature matching in complex underwater environments.

[0207] S5.4: For multi-scale spatial candidate matching pairs after geometric consistency constraints, based on the synchronous spatiotemporal labels corresponding to each pair of features, the corresponding physical environment sensing data is associated, and the environmental perception weight parameters of each candidate feature pair are inferred using the environmental perception model to form a weighted candidate matching pair set, thereby realizing the binding of environmental adaptability parameters.

[0208] S5.5: For a multi-scale spatial candidate matching pair set with environmental perception weight parameters, an environmental weight gating weighting mechanism is adopted. The ranking priority and confidence score of the candidate pairs are dynamically adjusted according to the physical environment adaptive weight matrix to generate the final multi-level spatial candidate matching pair set, providing the input data basis for subsequent mismatch filtering and spatial registration optimization under physical constraints.

[0209] Step S6: For the above candidate matching pairs, perform mismatch filtering under physical constraints based on environmental perception weight parameters, including using spatial consistency judgment and optical path constraints to filter out matching relationships that do not conform to the actual environmental distribution. Specifically, this includes:

[0210] S6.1: Extract the corresponding environmental perception weight parameters from the multi-level spatial candidate matching pairs generated by the preliminary feature matching process to construct a specialized mapping table between the matching relationship and the physical environment state for subsequent physical constraint applications.

[0211] S6.2: Based on the environmental perception weight parameters, spatial consistency judgment is performed on candidate matching pairs. The spatial distribution modeling algorithm is used to analyze the spatial distribution and relative position relationship of point cloud features and image features. Under the constraints of the physical environment, matching pairs that initially meet the spatial consistency requirements are selected.

[0212] The input data includes candidate point clouds and image feature matching pairs after preliminary feature matching and normalized multi-scale and geometric consistency optimization, as well as environment-aware weight parameters bound to each matching pair.

[0213] A spatial distribution modeling algorithm (parameters include: point cloud coordinates, image projection coordinates, spatial relationship matrix, and environment perception weight vector) is used to characterize the distribution structure of point cloud features and image features in the global space.

[0214] Furthermore, through a spatial consistency evaluation algorithm (parameter: spatial barycenter coordinate set) Pre-registration projection T pc→img Spatial distance threshold d center Environmental perception weight matrix W env ), calculate the spatial centroid consistency score for each matching pair under the current environmental weights. The following formula is used:

[0215]

[0216] Among them, W env (l, m) represents the environmental perception weight parameters. This is a weighted spatial distance evaluation value.

[0217] Furthermore, the consistency determination method of the principal direction vector (parameter: feature principal direction vector set) is used. Angle threshold θ max Environmental weight adjustment factor η env After weight correction, the direction angle is determined:

[0218]

[0219] Where, η env (l, m) is an environmental correction factor that is dynamically adjusted based on physical environment partitioning, so that the error threshold is adaptive to environmental disturbances.

[0220] Furthermore, a spatial region relationship constraint algorithm is used (parameters: regional spatial distribution label L, environmental partitioning threshold T). env This eliminates matching pairs that exceed spatial cluster boundaries or lack sufficient environmental reliability, achieving environmental adaptation matching and filtering under spatial consistency.

[0221] Through the above spatial distribution modeling, spatial consistency weighting judgment and direction vector multi-level fusion processing, the above preliminary candidate matching pair data are transformed into high-confidence spatial matching pairs that are modulated by environmental perception weights and conform to the consistency of physical environment spatial distribution.

[0222] By using spatial consistency modeling and environmental weight correction algorithms, the multi-scale candidate matching pairs output in the previous step are accurately screened, and mismatches with abnormal spatial relationships or poor environmental adaptability are eliminated, thus achieving the goals of preliminary filtering of spatial mismatches and environmental adaptive feature screening.

[0223] For example, in an underwater topographic survey mission near the South China Sea, for laser point cloud region A (containing 80 feature points), after processing with S5.3, the centroid is (23.5, 19.2, -4.7) m, and the corresponding image region B has a centroid of (507, 390) pixels. The environmental perception weight w env(A,B) = 0.85 (high illumination, low turbidity range). According to pre-registration T... pc→img The centroid of the projected point cloud region is (510, 393) pixels. A spatial distance threshold D is used. center =12 pixels, calculate weighted distance:

[0224]

[0225] The principal directional vectors are ([0.94,0.25,-0.24]) and ([0.93,0.28]), with an included angle of 6.7°. The environmental correction factor (η) env (A,B)=1.0 (Excellent environment), therefore:

[0226]

[0227] The matching pair is determined by spatial consistency. For region C, the environmental perception weight w env =0.48 (high turbidity zone), corresponding to the calculated projection distance. (Due to the complex environment in region C, a tighter threshold was applied), and the consistency screening was not passed. During this process, spatial distance and direction criteria were dynamically adjusted based on different environmental perception weights, effectively improving the retention rate of high-confidence regions and significantly filtering out mismatches in low-confidence regions. After testing in multiple terrain areas, the retention rate of high-confidence point cloud-image matching pairs after spatial consistency determination was 85.3%, and the high mismatch filtering rate exceeded 30%, significantly enhancing the stability and accuracy of multi-scale matching in low-visibility and complex environments.

[0228] S6.3: For the matching pairs after spatial consistency judgment, perform optical path constraint check. Utilize the illumination distribution parameters and turbidity parameters in the physical environment sensor data, and apply the optical transmission model to jointly check the physical distribution of point cloud and image, and eliminate mismatches that are optically inconsistent or significantly affected by environmental interference.

[0229] S6.4: The matching pairs filtered by both spatial consistency and optical path constraints are prioritized according to environmental interference weights to generate a weighted candidate matching sequence, which strengthens the reliable matching pairs in low turbidity areas and high illumination areas, and accumulates high-quality samples for the subsequent heterogeneous feature fusion process.

[0230] S6.5: The dynamic performance of the weighted candidate matching sequence integrated environmental perception model is evaluated by applying a multi-scale robustness analysis algorithm to assess the environmental adaptability and matching reliability of each matching pair within the target measurement area. The output is a mismatch filtering result that has been fully optimized by physical constraints, spatial consistency and environmental interference weights, which is used as the final effective feature matching pair input to the next processing step.

[0231] Step S7: Iterative optimization and environment-driven heterogeneous feature fusion are performed on the feature matching pairs filtered for mismatches. Regional physical labels are used to enhance the spatial consistency between the point cloud and the image, generating a weighted and optimized final matching pair. Specifically, this includes:

[0232] S7.1: Input the feature matching pairs under physical constraints into the iterative optimization module, use the candidate feature matching pairs selected in the previous step as the initial values ​​of the algorithm, perform local spatial consistency evaluation based on physical environment weights, and generate the initial optimization score of each pair of matching relationships to improve the spatial accuracy between point cloud features and image features.

[0233] S7.2: Based on iterative optimization scores, a heterogeneous feature fusion algorithm is called on the matching pair set to jointly encode the geometric characteristics (such as curvature distribution and density statistics) of the LiDAR point cloud features with the texture attributes (such as gray-level co-occurrence matrix and edge response) of the visual image features. The fusion parameters are automatically adjusted guided by the physical environment weight matrix to generate a fused multimodal feature description vector.

[0234] S7.3: Using the generated multimodal feature description vector as input, spatial consistency integration optimization is performed by partitioning the region physical label. Cluster analysis and weighted processing are performed on all matching pairs within the same physical label region to enhance the correlation of regional features under complex underwater distribution conditions and obtain environment-driven spatial consistency optimized matching pairs.

[0235] Using the multimodal feature description vector after fusing physical environment weights as input, and based on the synchronously generated regional physical label partitioning results, a spatial domain grouping algorithm (parameters: regional physical label L, feature description subset F1) is adopted to achieve the clustering and partitioning of matching pairs in physically homogeneous regions.

[0236] Furthermore, density clustering analysis of feature description vectors within the region (parameter: spatial distance metric F of feature description vectors) is used. ij Neighborhood radius r clust Minimum cluster size N min Under the same physical label, adaptive clustering based on distance metric is performed on the feature matching pair set to analyze its spatial distribution structure and mutual similarity, and generate high-density clusters of regional features.

[0237] Furthermore, an environment-aware weight fusion algorithm (parameter: weight matrix W) is introduced based on the clustering results. env (l,m), local confidence within the cluster S loc For each cluster, matching pairs are weighted according to the eccentricity distribution between the cluster center and its neighborhood, and then weighted by the physical region weights. This dynamically improves the participation level and environmental credibility of features within the region, constructing a spatial consistency index matrix S. cons.

[0238] Furthermore, through spatial consistency index normalization and multi-cluster cross-validation method (parameter: index matrix S) cons Normalization function Norm(·), inter-cluster distance threshold d center To achieve unified mapping of spatial consistency indices for feature pairs within and outside the region, cross-regional matching or external high-error terms are screened out, ensuring strong constraints on the physical rationality of regional distribution and local spatial consistency.

[0239] Furthermore, the spatially consistent weighted optimal matching set (parameters: regional physical label L, co-domain high-confidence cluster pair set C) is output through regional physical label clustering analysis. opt This generates spatial consistency optimization matching pairs driven by the environment, providing high-reliability, multi-scale environmental adaptability basic data support for subsequent feature point selection and spatial projection optimization for global registration.

[0240] By using the spatial consistency ensemble optimization algorithm, the above-mentioned fused multimodal feature description vectors are spatially partitioned and weighted clustered according to the regional physical labels. This significantly improves the spatial distribution consistency and accurate correlation of matching pairs in highly complex underwater environments, and achieves the goal of multi-scale high-confidence feature screening in low visibility and unevenly distributed scenarios.

[0241] For example, in the point cloud-image registration task in the nearshore deep water area of ​​the South China Sea, the acquired laser point cloud region was divided into three physical label regions (Region X: low turbidity, high illumination region; Region Y: medium turbidity, low illumination region; Region Z: high turbidity, high temperature region). Region X contains 45 candidate matching pairs for initial point cloud-image registration. Using DBSCAN with a clustering radius of 3 pixels and a minimum clustering capacity of 10, two core feature clusters were obtained. Within this region, W... env The mean is 0.83, and the spatial consistency index S cons Within the range [0.88, 1.00], after weighted fusion, the matching pair retention rate is improved to 91%, and region X outputs 32 optimized matching pairs. Region Y is configured with a clustering radius of 2 pixels, a minimum capacity of 8, and W... env With a mean of 0.54, there is one cluster, and the spatial consistency index, after normalization, falls within the range of [0.53, 0.67]. Only 11 matching pairs are retained. Due to high turbidity and high temperature, clustering in region Z is difficult, and only 3 high-confidence matching pairs are retained. After the above weighted processing of regional clustering, a total of 46 high-consistency spatially optimized matching pairs are finally output across the three regions. The spatial mismatch rate in regions X and Y is controlled within 5%, and the high-confidence matching rate in region Z, under high interference, is increased to over 80%. This significantly enhances the effectiveness of point cloud-image spatial consistency and multi-domain fusion matching in extremely complex environments.

[0242] S7.4: Perform secondary mismatch detection on the optimized matching pairs. By analyzing the differences between the multimodal feature description vector and the regional physical label, and using environmental interference factors (such as abnormal turbidity and drastic changes in local illuminance) to set automatic elimination criteria, the remaining potential mismatch items are screened out to ensure the physical rationality and spatial consistency of the final matching pairs.

[0243] S7.5: Integrates macro-environment labels, fused feature description vectors, and spatial consistency verification results, inputs them into the weighted optimization module, performs final optimization and sorting based on environmental perception weights, and outputs the weighted final matching pairs that have undergone multiple rounds of optimization and fusion processing, providing the spatial registration module with highly robust and multi-scale environmentally adaptive feature base data.

[0244] Step S8: Using the final matching pair as input, high-precision spatial registration of the laser point cloud and the visual image is performed to obtain a registration parameter set with fused environmental perception weights, thereby improving the stability of underwater multi-scale automatic matching. Specifically, this includes:

[0245] S8.1: Input verification is performed on the final feature matching pairs of the weighted selection. Based on the environmental perception weights and spatial labels, the set of key matching points required for registration is selected to ensure the accuracy of the input data of the spatial mapping algorithm and the ability to suppress environmental interference.

[0246] S8.2: For the verified key matching point set, the environment-adaptive spatial mapping algorithm is used to calculate the preliminary transformation parameters between the laser point cloud coordinate system and the visual image pixel space, including the rotation matrix, translation vector and scale factor, to achieve preliminary spatial alignment of multimodal features.

[0247] S8.3: Based on environmental perception weights, interference compensation processing is performed on spatial mapping parameters. Physical environmental factors (such as turbidity, illuminance, etc.) are used to perform weight correction and iterative optimization of transformation parameters to improve the robustness and accuracy of spatial registration in low visibility environments.

[0248] The input is a preliminary spatial mapping parameter set based on the output of the S8.2 sub-step, including rotation matrix, translation vector and scale factor, and is associated with regional environmental perception weights and physical environment sensing dataset (including environmental factor data such as turbidity, illuminance, and temperature under synchronous spatiotemporal labels).

[0249] An environmental weight coupling compensation algorithm is adopted (parameters: transformation parameters P = {R, T, s}, environmental perception weights W). env Environmental factor group E={E turb E illum E temp This enables the correction of partition weights for the initial spatial mapping parameters.

[0250] Furthermore, by establishing a response function model (physical mechanism coupling model) between environmental factors and spatial registration error, and based on multivariate linear regression or nonlinear response correlation analysis, the sensitivity coefficient β of the registration parameters under the main perturbations is obtained. k The following weight correction formula is formed:

[0251]

[0252] Among them, P corr To correct the registration parameters, P init For the initial spatial mapping parameters, β k Let be the registration sensitivity coefficient for the k-th type of environmental factor. Let ΔE be the weight of the k-th type of environment. k This represents the deviation from the baseline environmental conditions.

[0253] Furthermore, through a local environment label partitioning iterative optimization method (parameters: label Li, corresponding partition matching subset, partition environment factor data), the spatial transformation parameter set is subjected to partitioning recursive error minimization processing, and a nonlinear least squares optimization algorithm based on robust loss functions (such as Huber, Tukey) is adopted:

[0254]

[0255] Where, N L For the number of partitions, Let X be the environmental weight of the i-th partition. Li ,Y Li To register the point cloud coordinates and image pixel coordinates within the i-th partition, Let ρ(·) be the spatial transformation mapping function, ρ(·) be the robust error loss function, and d(·) be the Euclidean distance metric.

[0256] Furthermore, by using a registration parameter convergence determination algorithm, the trend of registration parameter changes, convergence threshold, and environmental residual index during the optimization process are jointly monitored to achieve adaptive iterative termination control under environmental disturbances.

[0257] Furthermore, through parameter stability feedback analysis, the spatial registration parameter set that has been compensated for interference and converged to a stable form is constructed into a standardized correction parameter set, which is then compared with the original uncompensated parameters to output compensation effect evaluation data.

[0258] By using environmental weight coupling and interference compensation optimization algorithms, the initial spatial mapping parameters are transformed into highly robust spatial registration parameters that have incorporated the influence of multiple environmental factors, effectively improving the stability and accuracy of spatial registration under low visibility and strong environmental disturbance conditions.

[0259] For example, for a nearshore measurement station in southeastern coastal China, the original registration candidate point set is distributed across three physical zones: Zone A (turbidity 15 NTU, illuminance 200 Lux), Zone B (turbidity 45 NTU, illuminance 80 Lux), and Zone C (turbidity 80 NTU, illuminance 20 Lux). The environmental perception weights for Zones A, B, and C are 0.83, 0.61, and 0.35, respectively. The initial mapping parameters (rotation matrix R, translation vector T) are estimated by aligning the dominant point pairs in Zone A. Without compensation, the average registration residuals for Zones B and C are 6.4 pixels and 12.1 pixels, respectively. In the interference compensation process, based on the weighting factors, the aforementioned correction model is used to adjust the R and T parameters, and the sensitivity coefficient β... turb =0.065, β illum = -0.019. After compensation, the registration residuals in areas B and C decreased to 2.1 pixels and 4.8 pixels, respectively, and the overall registration mean square error decreased to 3.6 pixels. This step effectively enhanced the spatial registration robustness of areas with high environmental interference, achieving robust alignment of point clouds and images in a large range of heterogeneous regions, and ensuring accurate data input for subsequent terrain reconstruction modules.

[0260] S8.4: Using the spatial registration parameters after interference compensation, coordinate reprojection processing is performed on the laser point cloud data, and regional coordinate correction is performed on the visual image to achieve accurate alignment of multi-source sensing data in a unified physical space, and output a set of spatial registration parameters with fused environmental weights.

[0261] S8.5: Accuracy assessment and consistency verification are performed on the registered spatial parameter set. An environment-driven error evaluation mechanism is adopted to analyze the statistical distribution of registration residuals in different physical environmental factors-dominated zones. The zone environmental adaptability registration success rate is output, providing stability assurance for continuous system iteration optimization and subsequent topographic surveys.

[0262] Step S9: Implement real-time dynamic monitoring. When environmental sensor data or registration accuracy parameters change, determine whether to switch feature extraction algorithms, adjust parameters, or restart the registration process based on the environmental perception model to adapt to changes in operating conditions and continuously optimize registration performance. Specifically, this includes:

[0263] S9.1: The underwater physical environment sensor acquisition module performs real-time data acquisition and processing, including environmental parameters such as turbidity, illuminance, and temperature. Based on the synchronous spatiotemporal tagging mechanism, the environmental perception weight matrix is ​​used as input to continuously monitor changes in environmental status and realize dynamic updates of the environmental interference weight matrix, providing basic environmental perception input for downstream registration accuracy change monitoring.

[0264] S9.2: After the environmental perception weight matrix has been dynamically updated, the latest round of laser point cloud and visual image spatial registration parameters are collected. Through registration quality evaluation algorithms (such as spatial error statistics analysis and residual distribution characteristic determination), a registration accuracy parameter set is generated to form a key technical indicator of the dynamic monitoring causal chain of registration performance, providing a real-time accuracy discrimination basis for environmental model adaptation.

[0265] S9.3: Based on the latest environmental interference weight matrix and registration accuracy parameter set input, the environmental perception model is iteratively judged and processed. The environmental state-registration accuracy joint adaptation discrimination algorithm is adopted to determine whether there is a need for feature extraction algorithm adaptation or parameter fine-tuning under the current working condition, so as to provide a reliable criterion for the next step of feature extraction strategy switching and system dynamic restart decision.

[0266] S9.4: For scenarios that require adaptation based on the environmental perception model, perform adaptive switching and parameter fine-tuning operations of multi-scale feature extraction algorithms, including point cloud curvature feature window adjustment, image texture local operator change, feature scale domain constraint reconfiguration, etc., to actively optimize feature extraction factors based on the intensity of environmental disturbance and achieve adaptive output of features related to dynamic environment.

[0267] S9.5: Synchronously manage the adaptive output of features generated after adaptive switching or parameter fine-tuning of the feature extraction algorithm. Utilize the automatic restart mechanism of the system registration process to reload the multi-source data fusion and spatial registration tasks in real time. This enables a closed-loop optimization strategy triggered by changes in the environmental perception weight matrix and registration accuracy parameter set, thereby obtaining high-precision and robust spatial registration results through continuous environmental adaptive optimization.

[0268] For those skilled in the art, various other corresponding changes and modifications can be made based on the technical solutions and concepts described above, and all such changes and modifications should fall within the protection scope of the claims of this invention.

[0269] Unless otherwise defined, the technical or scientific terms used herein shall have the ordinary meaning as understood by one of ordinary skill in the art to which this application pertains. The terms “first,” “second,” “third,” and similar terms used in this patent application specification and claims do not indicate any order, quantity, or importance, but are merely used to distinguish different components. Similarly, the terms “an” or “a” and similar terms do not indicate a quantity limitation, but rather indicate the presence of at least one. The terms “comprising” or “including” and similar terms mean that the elements or objects preceding “comprising” or “including” encompass the elements or objects listed following “comprising” or “including” and their equivalents, and do not exclude other elements or objects. The “multiple” mentioned in the embodiments of this application refers to two or more. A and / or B indicate three possibilities: A; B; and A and B.

[0270] The above description is merely an exemplary embodiment of this application, but the scope of protection of this application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in this application, and such modifications or substitutions should all be covered within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. An underwater topographic measurement method based on lidar and vision fusion, characterized in that, Specifically, it includes: S1: Acquire raw point cloud data from underwater lidar, underwater visual images, and in-situ physical environment sensing data associated with multiple locations, including turbidity, illuminance, and temperature variables, and assign synchronous spatiotemporal labels to each set of acquired data. S2: Preprocess the raw laser point cloud data and visual image data under the synchronization tag, including denoising, scale normalization and signal enhancement operations; S3: Input the pre-processed point cloud data, image data and physical environment sensing data into the environment perception model to generate an environment interference weight matrix; S4: Based on the environmental interference weight matrix, adaptively select multi-scale feature extraction algorithms for point clouds and images or adjust feature parameters, including adjusting point cloud feature scale, curvature calculation window, image texture extraction operator and region effective range; S5: Perform preliminary feature matching processing on the adaptively extracted multi-scale point cloud features and image features to generate multi-level spatial candidate matching pairs, and record the corresponding paired environmental perception weight parameters. S6: For the above candidate matching pairs, perform mismatch filtering under physical constraints based on the environmental perception weight parameters, including using spatial consistency judgment and optical path constraint basis to screen out matching relationships that do not conform to the actual environmental distribution. S7: Iterative optimization and environment-driven heterogeneous feature fusion are performed on the feature matching pairs filtered by mismatches. Regional physical labels are used to enhance the spatial consistency between point clouds and images, and weighted optimization of the final matching pairs is generated. S8: Using the final matching pair as input, perform high-precision spatial registration between the laser point cloud and the visual image to obtain a registration parameter set with fused environmental perception weights; S9: Implement real-time dynamic monitoring. When environmental sensor data or registration accuracy parameters change, determine whether it is necessary to switch feature extraction algorithms, adjust parameters, or restart the registration process based on the environmental perception model.

2. The underwater topography measurement method based on lidar and vision fusion according to claim 1, characterized in that, Step S1 specifically includes: S1.1: Configure the underwater lidar system deployed within the target measurement area with preset acquisition parameters, including setting the laser emission band, sampling rate, and spatial resolution, to ensure that the acquired raw point cloud data has sufficient spatial coverage and accuracy; S1.2: Set the sampling period and scene exposure parameters for the visual camera acquisition units in the same area, acquire underwater visual images based on high-sensitivity low-light imaging technology, extract resolution, scene view, and timestamp metadata, and obtain the original visual image dataset; S1.3: Data acquisition and scheduling of physical environment sensing modules deployed at measurement points are performed to obtain in-situ physical environment sensing data of turbidity, illuminance and temperature in the area. During the acquisition process, multi-channel synchronous acquisition and signal calibration technology are used to form an environmental sensing data stream, which provides physical underlying input for subsequent weighting of the environmental perception model. S1.4: Based on the data acquisition synchronization mechanism, the raw point cloud data of lidar, visual image data and physical environment sensing data are respectively labeled with data timestamps and spatial location information with unified specifications to form a consistent spatiotemporal metadata dataset of the three types of data. S1.5: Perform data integration processing on laser point cloud data, visual image data and physical environment sensing data that have been attached with synchronous spatiotemporal tags, use spatiotemporal tag index to verify the consistency of the correlation between each group of data, detect and filter out data samples that are spatiotemporally asynchronous or spatially positioned incorrectly due to acquisition anomalies, and generate a verified multi-source synchronous sensing dataset.

3. The underwater topographic measurement method based on lidar and vision fusion according to claim 1, characterized in that, Step S2 specifically includes: S2.1: Perform spatial noise filtering on the laser point cloud data under the synchronization tag, and obtain high-purity laser point cloud data after noise suppression based on the statistical anomaly detection algorithm; S2.2: A multi-resolution image enhancement algorithm is used for the visual image data under the synchronization tag, combined with underwater image dehazing and color shift correction processing, to improve image contrast and variability, so as to obtain a highly recognizable visual image after signal enhancement. S2.3: Based on the obtained high-purity laser point cloud data and enhanced visual images, scale normalization processing is performed, and a unified scaling factor is used to map the multi-source data to a spatially consistent scale framework, forming a unified standard scale point cloud set and a standard size image. S2.4: Apply spatial signal enhancement techniques to standard-scale point cloud sets, including multi-scale curvature enhancement and point density enhancement algorithms, to refine the geometric signals of feature regions in order to generate feature-sensitive enhanced laser point cloud data and enhance the point cloud's ability to distinguish key geomorphic areas. S2.5: Perform adaptive texture enhancement and local artifact correction on standard-sized visual images, and use multi-channel contrast enhancement technology to superimpose local image denoising algorithms to generate enhanced visual image data with clear feature levels.

4. The underwater topographic measurement method based on lidar and vision fusion according to claim 1, characterized in that, Step S3 specifically includes: S3.1: Aggregate the lidar point cloud data, visual image data, and in-situ physical environment sensing data marked with synchronous spatiotemporal labels after denoising and scale normalization to form a composite input feature set for the environmental perception model. S3.2: Based on the composite input feature set, the physical environment sensing data is decomposed by the environmental factor feature decoupling algorithm to distinguish the degree of individual influence of turbidity, illumination and temperature on point cloud feature distribution and image texture feature performance, and obtain the influence weight parameters of physical environment factors. S3.3: Using the influence weight parameters of physical environmental factors as input, multi-source data fusion modeling is performed within the environmental perception model, including laser point cloud saliency analysis and visual image feature saliency extraction, to generate an environmental adaptability influence factor matrix that distinguishes point cloud and image features; S3.4: Based on the environmental adaptability influencing factor matrix, the weight normalization and spatial location mapping algorithm is used to complete the standardization transformation of the influencing factors, so that each feature extraction unit can adaptively respond according to the actual intensity of environmental disturbance in the subsequent processing. S3.5: Perform interactive environmental factor weighting operations on the standardized impact factor matrix and the preprocessed feature set, and combine synchronous spatiotemporal labels to generate and output the environmental interference weight matrix.

5. The underwater topographic measurement method based on lidar and vision fusion according to claim 1, characterized in that, Step S4 specifically includes: S4.1: Analyze the input environmental interference weight matrix, and extract the optimal weight configuration for point cloud feature extraction and image feature extraction under the current working condition based on the reflection of the physical parameters of each weight matrix on the effectiveness of the sensing method. S4.2: Based on the weight configuration obtained in S4.1, a multi-scale point cloud feature extraction method is applied to the preprocessed point cloud data. By adaptively adjusting the point cloud feature scale and curvature calculation window, a multi-scale feature descriptor of the point cloud that is consistent with the environmental state is output. S4.3: Combining the environment-driven weight configuration of S4.1, the normalized image data is processed using a multi-scale image texture extraction operator. The parameters of the texture extraction operator and its effective area are dynamically adjusted to output a multi-scale image feature descriptor under environment awareness. S4.4: Input the point cloud multi-scale feature descriptor generated in S4.2 and the image multi-scale feature descriptor generated in S4.3 into the feature weight fusion module at the same time. Adjust the participation ratio of various features dynamically according to the environmental interference weight matrix to form a multimodal feature joint vector with normalized weights. S4.5: Based on the changes in dynamic environmental factors, the multimodal feature joint vector output by S4.4 is used to automatically adjust the parameters of the feature extraction algorithm iteratively, including real-time response to newly detected environmental extreme values, and outputting the updated feature descriptor set.

6. The underwater topographic measurement method based on lidar and vision fusion according to claim 1, characterized in that, Step S5 specifically includes: S5.1: Perform feature vector normalization processing on the multi-scale point cloud features and multi-scale image features that have been adaptively optimized. Use the normalized vector features as input. The algorithm used is feature vector normalization mapping and distance metric transformation to eliminate the influence of scale and distribution inconsistency on subsequent matching algorithms and generate a normalized feature descriptor subset. S5.2: Based on the normalized feature descriptor subset, multi-scale spatial neighborhood clustering and KNN algorithm are used to establish candidate matching tables for point cloud features and image features in their respective feature spaces, so as to automatically detect feature pairs with the minimum Euclidean distance or the maximum correlation coefficient across modalities and obtain the first round of multi-level spatial candidate matching relationships. S5.3: For the first round of multi-level spatial candidate matching relationships, the geometric consistency judgment algorithm of multi-scale feature regions in the global space is used to process them in order to eliminate obvious mismatch terms caused by structural distortions and output multi-scale spatial candidate matching pairs after geometric consistency constraints. S5.4: For multi-scale spatial candidate matching pairs after geometric consistency constraints, based on the synchronous spatiotemporal labels corresponding to each pair of features, associate the corresponding physical environment sensing data, use the environment perception model to infer the environment perception weight parameters of each set of candidate feature pairs, and form a weighted candidate matching pair set. S5.5: For a multi-scale spatial candidate matching pair set with environmental perception weight parameters, an environmental weight gating weighting mechanism is adopted to dynamically adjust the sorting priority and confidence score of the candidate pairs according to the physical environment adaptive weight matrix, so as to generate the final multi-level spatial candidate matching pair set.

7. The underwater topographic measurement method based on lidar and vision fusion according to claim 1, characterized in that: The underwater lidar point cloud data acquisition in the S1 phase includes adaptive settings for laser emission band, sampling rate, spatial resolution, emission array scanning angle, echo signal threshold, and energy self-test parameters. The laser emission band is set to 532nm or 1064nm, and the point cloud spatial resolution is 1-20cm.

8. The underwater topographic measurement method based on lidar and vision fusion according to claim 1, characterized in that: In the S1 stage of visual image data acquisition, a high-sensitivity low-light imaging device is used to automatically estimate the sampling period and exposure time parameters. The sampling period is 0.05-2 seconds, the exposure time is dynamically adjusted, and multi-channel spatial synchronization and precise image timestamp marking are supported.

9. The underwater topographic measurement method based on lidar and vision fusion according to claim 1, characterized in that: The point cloud denoising process in stage S2 includes radius filtering, outlier detection, and spatial partition density extreme value detection algorithms. The spatial radius R is 0.03-0.2 meters, the minimum number of neighboring points is 8-20, the mean distance threshold is 0.03-0.15 meters, and the noise ratio is filtered to within 1%.

10. The underwater topographic measurement method based on lidar and vision fusion according to claim 1, characterized in that: The visual image preprocessing in stage S2 employs multi-resolution image enhancement, dehazing and color shift correction, contrast enhancement, and multi-channel noise reduction algorithms. The enhanced image resolution is set to above 1M pixels, and the single-channel noise intensity is below 5%.