Underwater multi-mode intelligent sensing method for marine ranching
By employing a multimodal collaborative perception architecture that combines visual-sonar fusion and physical imaging models, the problem of insufficient detection accuracy and 3D reconstruction accuracy in underwater environments has been solved, enabling efficient environmental perception and resource assessment for marine ranches.
Patent Information
- Application Number
- CN202510946258.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-09
- Publication Date
- 2025-10-28
- Estimated Expiration
- 2045-07-09
AI Technical Summary
Existing underwater sensing technologies suffer from reduced detection accuracy in turbid waters or under low light conditions. SLAM systems lack sufficient environmental perception capabilities, and their 3D reconstruction accuracy and robustness are poor, making it difficult to support resource assessment and environmental monitoring in marine ranches.
Employing a multimodal collaborative perception architecture, combining target detection, semantic SLAM, and medium-compensated 3D rendering, and through visual-sonar fusion, dense mapping, and physical imaging models, multi-granularity and multimodal collaborative perception are achieved, thereby improving detection accuracy and 3D reconstruction quality.
Achieving adaptive weighted fusion in complex underwater environments improves the accuracy of visual-sonar detection, constructs dense maps containing marine product distribution, enhances the color fidelity and real-time performance of 3D reconstruction, and supports resource surveys and ecological monitoring of marine ranches.
Smart Images

Figure CN120852689A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the technical field of multimodal sensing methods, and more particularly to an underwater multimodal intelligent sensing method for marine ranches. Background Technology
[0002] With the rapid development of marine ranching, the demand for intelligent sensing of complex underwater environments has increased significantly. Key technological requirements include accurate detection of seafood and marine debris, real-time localization and mapping (SLAM) for underwater robots, and high-precision 3D environment modeling. However, existing underwater sensing technologies still face many challenges, primarily including: Limitations of underwater target detection. Current underwater target detection methods primarily rely on optical cameras, but visual information is severely degraded in turbid waters or low-light conditions, leading to a significant decrease in detection accuracy. While sonar provides supplementary information in low-visibility environments, existing methods lack adaptive multimodal complementarity mechanisms.
[0003] SLAM systems suffer from insufficient environmental perception capabilities. Traditional SLAM systems only construct geometric maps and cannot simultaneously identify and label the distribution of marine products, making it difficult to support marine resource assessment tasks. Furthermore, feature-point-based SLAM systems struggle to generate dense point clouds rich in scene structure information, while RGB-D or LiDAR-based solutions are susceptible to scattering interference underwater. Water absorption and scattering cause image blurring and color distortion, affecting feature extraction and matching accuracy, and consequently increasing the pose estimation error of visual SLAM.
[0004] The accuracy and robustness of 3D reconstruction are issues. Current underwater 3D reconstruction methods typically employ multi-view stereo vision or structured light scanning, but due to the influence of the water medium, the reconstruction results often suffer from defects such as color distortion, strong geometric noise, and insufficient real-time performance.
[0005] To address the aforementioned challenges, this invention employs a three-level intelligent perception architecture—target detection, semantic SLAM, and media-compensated 3DGS—to achieve multi-granularity and multi-modal collaborative perception, thereby improving detection accuracy, mapping efficiency, and 3D reconstruction quality, and solving the environmental perception problem in marine ranch scenarios. Summary of the Invention
[0006] In response to the technical problems mentioned in the background section, this invention provides an underwater multimodal intelligent sensing method for marine ranches. This invention constructs a multimodal, collaborative, and multi-granular progressive underwater intelligent sensing framework, organically integrating target detection, semantic SLAM, and 3D rendering with media compensation to form a comprehensive sensing capability encompassing everything from macroscopic marine product distribution to microscopic scene structure.
[0007] The technical means employed in this invention are as follows: An underwater multimodal intelligent sensing method for marine ranches includes the following steps: Step 1: Based on the YOLO underwater multimodal detection module, sonar sensing modes are introduced to construct a dynamically adaptive multimodal perception system; Step 2: Perform semantic SLAM and dense coarse mapping; the process of performing semantic SLAM and dense coarse mapping includes: preprocessing of visual-inertial data, initialization and joint optimization, dense mapping based on surface elements, and semantic fusion mapping; Step 3: Perform physical-driven 3DGS fine reconstruction, including: initial 3DGS modeling and underwater medium estimation and compensation; Step 4: Output a multi-scale semantic map.
[0008] Further, step 1 includes the following steps: Step 11: Evaluate the quality of the visual data and establish a visual quality evaluation system based on the average value of the brightness gradient amplitude; the formula for calculating the brightness gradient amplitude is: ; in, Indicates the image in coordinates Pixel intensity at that location and These represent the horizontal and vertical gradients calculated using the Sobel operator, respectively. Step 12, the quantitative evaluation indicators are: ; in, Indicates image resolution, Indicates the visual image quality score; Step 13: Determine if the visual image quality score is too low; if the visual score is too low, increase the weight of the sonar detection result, and vice versa. The final detection result is corrected by the visual data quality score.
[0009] Furthermore, the formula for calculating the visual and sonar detection weights is as follows: ; in, Indicates visual clarity score, and This represents the modal sensitivity coefficient.
[0010] Furthermore, in step 2, the preprocessing of the visual-inertial data is as follows: ; in, Represents the rotational pre-integral quantity, i.e., from time t. i arrivej The relative rotational change is obtained by integrating the continuous IMU angular velocity measurements; Indicates the k Angular velocity measured by each IMU; This represents the velocity pre-integral quantity, that is, in i The velocity change obtained by integration in the time-time coordinate system; Indicates the k Linear acceleration measured by an IMU; Indicates the position pre-integral quantity, that is, in i The change in displacement in the time coordinate system.
[0011] Furthermore, the initialization and joint optimization employ a loosely coupled approach, aligning the IMU's pre-integral values with the visual trajectory: ; in, s represents the scale factor, and g represents the direction of gravity; When using a sliding window to perform local optimization on the visual-inertial navigation system, the state variables are: ; in, Indicate pose, velocity, and IMU bias; then the optimized objective function is: ; in, This represents the IMU pre-integration residual. This indicates visual reprojection error. Represents Huber robust kernel coefficients; The visual bag-of-words model is used to detect loop closures during semantic SLAM and dense coarse-scale graph construction. Feature points in the current frame are matched with historical keyframes, and similarity scores are calculated. If similar frames are identified, they are marked as loop closures. If loop closures exist, global optimization is performed, i.e., loop closure constraints are added to the pose graph for optimization. ; in This represents the solution obtained by minimizing the objective function. i The optimal pose estimation for each keyframe is calculated; otherwise, the pose with 6 degrees of freedom is directly output.
[0012] Furthermore, each surface element in the surface-based dense mapping is represented by the following parameters: ; in, Indicates the center position of the face element. The normal vector of a surface element. Indicates the radius of the element. Indicates color information, The timestamp of the last update of the face element; First, pixels with similar colors, textures, and spatial locations are extracted from the visual image to form superpixels, which are then associated with surface elements. The method involves backprojecting each superpixel through a depth map to obtain a set of 3D points and fitting a local plane. A threshold for the distance between the center position and a threshold for the angle between the normal vectors are used to determine whether the superpixel can be successfully associated with existing surface elements. If the association is successful, the surface element information is updated through weighted fusion. ; in, and This represents the updated face element position and normal vector. Indicates the center position of the superpixel. The normal vector of a superpixel. and All represent confidence levels; Update the color and radius of the opposite element: ; in, Indicates the expansion coefficient; If association with an existing facet fails, the center point, normal vector, and color information of the superpixel fitting will be stored in the new facet. Modify the timestamp to the most recently updated timestamp; when the SLAM system detects a loop closure, all facets need to undergo rigid body transformation based on the pose graph optimization results; if there is no loop closure, no processing is required. ; in, Indicates from keyframe i arrive j The pose transformation matrix optimized by loop closure; Global consistency is ensured through pose graph optimization: ; in, This represents the set of edges constrained by loops; the final map is a set of all polygons. If there are navigation requirements or other situations that require a dense point cloud map, the polygon map can be converted into a dense point cloud map by only retaining the center point position parameter.
[0013] Furthermore, the semantic fusion mapping method in step 2 is as follows: Semantic information is integrated during the mapping process. During semantic SLAM and dense coarse mapping, the frames captured by the camera are sent in real time to the underwater multimodal detection module based on YOLO to obtain the semantic labels and location information of the target objects. Since the detection results are based on the camera coordinate system, they are transformed to the world coordinate system through pose transformation. The transformation formula is as follows: ; in, Represents a 3D point in the camera coordinate system. Represents the corresponding point in the world coordinate system. This represents a 3×3 rotation matrix, which is the pose transformation from the camera coordinate system to the world coordinate system. This represents a 3×1 translation vector, which is the position of the camera's optical center in the world coordinate system; Obtained directly from the output of the SLAM module This allows us to calculate the target's position in the world coordinate system. After time and spatial filtering, we can determine whether the target is stable and whether it has been recorded. If it has not been recorded, we record the target's label and its position in the world coordinate system. If it has been recorded, we skip this recording step. Temporal filtering is used to filter transient noise to ensure temporal consistency across multiple detection frames; sliding window detection is employed, with temporal stability based on the following: ; in, Indicates the length of the sliding window. Indicates the Frame detection confidence, and These represent the confidence threshold and the minimum pass rate, respectively. Finally, the recorded target locations and label information are represented in the map construction described above, thus completing the semantic fusion map construction.
[0014] Furthermore, in step 3, the physical imaging model is integrated into the 3DGS mapping process. By analyzing the depth map generated by 3DGS, the water attenuation coefficient and scattering parameters are estimated in real time, and an underwater optical compensation model is constructed accordingly. Given a viewpoint, an image can be rendered efficiently by sorting Gaussians from front to back, projecting them onto the camera plane, and performing alpha synthesis; for a pixel x, the color of pixel x is calculated as: ; in, and Indicates the i Given a 3D Gaussian color and density, the underwater image formation process can be represented as follows: ; in, Indicates the actual captured image. This represents a realistic underwater scene image after removing the influence of the medium. Indicates the attenuation coefficient. Represents the backscattering coefficient. Representing the backscattered color at infinity, it is assumed that... Used as the background color.
[0015] Furthermore, the method for embedding the underwater image physical formation model into 3DGS is as follows: First, the depth map of the scene is obtained from the initial few rounds of 3DGS rendering. ,pass Estimate attenuation coefficient Scattering coefficient and background color The effects of attenuation and scattering can then be obtained:
[0016]
[0017] Then, the estimated underwater real image is superimposed using the following formula. Up, the first round can Set it directly to the original image;
[0018] The resulting image will then be generated. By constructing a reconstruction loss model that closely approximates the original image, and through multiple iterations, an image close to the actual underwater image can be obtained. J .
[0019] Compared with the prior art, the present invention has the following advantages: 1. The dynamic weight fusion mechanism provided by this invention utilizes a visual quality assessment system ( Svis ) and sonar compensation coefficient ( β The synergistic effect of visual and sonar detection results enables adaptive weighted fusion in complex underwater environments. When the visual score... Svis When the value is ≤15, the system automatically increases the sonar weight to 0.27-0.88 (see the weight table in Example 1), which solves the problem of traditional single visual detection failing in turbid waters.
[0020] 2. A dense map containing semantic labels for seafood distribution was constructed through real-time linkage between surface parameterization and the YOLO detection module. The superpixel association mechanism improved map update efficiency, while pose graph optimization ensured global consistency during loop closure. 3. By embedding the physical imaging model into the 3DGS rendering pipeline and iteratively estimating the attenuation coefficient and scattering parameters, the color fidelity of the reconstructed scene is improved. The reconstruction loss between the compensated image and the original image is significantly reduced compared to traditional methods.
[0021] Based on the above reasons, this invention can be widely applied in fields such as marine ranch resource survey and ecological monitoring, underwater robot autonomous navigation and operation, shipwreck archaeology and three-dimensional digitization of coral reefs. Attached Figure Description
[0022] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0023] Figure 1 This is a schematic diagram of the workflow of the present invention. Detailed Implementation
[0024] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0025] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0026] like Figure 1 As shown, this invention provides an underwater multimodal intelligent sensing method for marine ranches, comprising the following steps: Step 1: To address the dual needs of marine ranching yield assessment and pollution monitoring, this study developed an intelligent detection system integrating visual and sonar data. The core of the system employs an underwater target detection module based on the YOLOv5 architecture, enabling simultaneous marine product yield statistics and marine debris monitoring. Furthermore, to overcome the visual limitations commonly encountered in underwater environments, such as low illumination and turbid water, this application incorporates sonar sensing modalities based on the YOLO underwater multimodal detection module to construct a dynamically adaptive multimodal perception system.
[0027] In a preferred embodiment, step 1 in this application includes the following steps: Step 11: Evaluate the quality of the visual data and establish a visual quality evaluation system based on the average value of the brightness gradient amplitude; the formula for calculating the brightness gradient amplitude is: ; in, Indicates the image in coordinates Pixel intensity at that location and These represent the horizontal and vertical gradients calculated using the Sobel operator, respectively. Step 12, the quantitative evaluation indicators are: ; in, Indicates image resolution, Indicates the visual image quality score; Step 13: Determine if the visual image quality score is too low; if the visual score is too low, increase the weight of the sonar detection result, and vice versa, using the visual data quality score to correct the final detection result. Preferably, the formula for calculating the visual and sonar detection weights is: ; in, Indicates visual clarity score, and This represents the modal sensitivity coefficient.
[0028] Furthermore, real-time acquisition of one's own position and surrounding environment information is crucial during underwater operations. Therefore, step 2 involves semantic SLAM and dense coarse mapping. Considering the low frequency of visual keyframes and the high sampling frequency of the inertial measurement unit (IMU), the system first performs pre-integration processing on the IMU data to optimize computational efficiency. This method converts the acceleration and angular velocity information provided by the IMU into relative displacement, velocity, and rotation between adjacent keyframes, effectively reducing the computational burden while ensuring the accuracy and real-time performance of positioning and mapping. The semantic SLAM and dense coarse mapping process includes: preprocessing of visual-inertial data, initialization and joint optimization, dense mapping based on surface elements, and semantic fusion mapping. In step 2, the preprocessing of the visual-inertial data is as follows: ; in, Represents the rotational pre-integral quantity, i.e., from time t. i arrive j The relative rotational change is obtained by integrating the continuous IMU angular velocity measurements; Indicates the k Angular velocity measured by each IMU; This represents the velocity pre-integral quantity, that is, in i The velocity change obtained by integration in the time-time coordinate system; Indicates the k Linear acceleration measured by an IMU; Indicates the position pre-integral quantity, that is, in i The change in displacement in the time coordinate system.
[0029] In a preferred embodiment, FAST corner detection and KLT optical flow tracing are performed on the visual image for feature extraction and tracking. Then, joint initialization of inertial and visual components is performed. This invention employs a loosely coupled approach, utilizing pure visual structure of motion (SfM) technology to estimate the poses of 3D map points and keyframes, and then aligning the IMU pre-integral values with the visual trajectory. The initialization and joint optimization employ a loosely coupled approach, aligning the IMU pre-integral values with the visual trajectory: ; in, s represents the scale factor, and g represents the direction of gravity; When using a sliding window to perform local optimization on the visual-inertial navigation system, the state variables are: ; in, Indicate pose, velocity, and IMU bias; then the optimized objective function is: ; in, This represents the IMU pre-integration residual. This indicates visual reprojection error. Represents Huber robust kernel coefficients; The visual bag-of-words model is used to detect loop closures during semantic SLAM and dense coarse-scale graph construction. Feature points in the current frame are matched with historical keyframes, and similarity scores are calculated. If similar frames are identified, they are marked as loop closures. If loop closures exist, global optimization is performed, i.e., loop closure constraints are added to the pose graph for optimization. ; in This represents the solution obtained by minimizing the objective function. i The optimal pose estimation for each keyframe is calculated; otherwise, the pose with 6 degrees of freedom is directly output.
[0030] Preferably, each element of the element-based dense mapping is represented by the following parameters: ; in, Indicates the center position of the face element. The normal vector of a surface element. Indicates the radius of the element. Indicates color information, The timestamp of the last update of the face element; First, pixels with similar colors, textures, and spatial locations are extracted from the visual image to form superpixels, which are then associated with surface elements. The method involves backprojecting each superpixel through a depth map to obtain a set of 3D points and fitting a local plane. A threshold for the distance between the center position and a threshold for the angle between the normal vectors are used to determine whether the superpixel can be successfully associated with existing surface elements. If the association is successful, the surface element information is updated through weighted fusion. ; in, and This represents the updated face element position and normal vector. Indicates the center position of the superpixel. The normal vector of a superpixel. and All represent confidence levels; Update the color and radius of the opposite element: ; in, Indicates the expansion coefficient; If association with an existing facet fails, the center point, normal vector, and color information of the superpixel fitting will be stored in the new facet. Modify the timestamp to the most recently updated timestamp; when the SLAM system detects a loop closure, all facets need to undergo rigid body transformation based on the pose graph optimization results; if there is no loop closure, no processing is required. ; in, Indicates from keyframe i arrive j The pose transformation matrix optimized by loop closure; Global consistency is ensured through pose graph optimization: ; in, This represents the set of edges constrained by loops; the final map is a set of all polygons. If there are navigation requirements or other situations that require a dense point cloud map, the polygon map can be converted into a dense point cloud map by only retaining the center point position parameter.
[0031] Preferably, the semantic fusion mapping method is as follows: Semantic information is integrated during the mapping process. During semantic SLAM and dense coarse mapping, the frames captured by the camera are sent in real time to the underwater multimodal detection module based on YOLO to obtain the semantic labels and location information of the target objects. To achieve intelligent environmental perception, this system integrates semantic information during the mapping process. During SLAM module operation, keyframes are fed into the YOLO detection module in real time to obtain the semantic labels and location information of target objects. Since the detection results are based on the camera coordinate system, they need to be transformed to the world coordinate system through pose transformation. The transformation formula is as follows: ; in, Represents a 3D point in the camera coordinate system. Represents the corresponding point in the world coordinate system. This represents a 3×3 rotation matrix, which is the pose transformation from the camera coordinate system to the world coordinate system. This represents a 3×1 translation vector, which is the position of the camera's optical center in the world coordinate system; Obtained directly from the output of the SLAM module This allows us to calculate the target's position in the world coordinate system. After time and spatial filtering, we can determine whether the target is stable and whether it has been recorded. If it has not been recorded, we record the target's label and its position in the world coordinate system. If it has been recorded, we skip this recording step. Temporal filtering is used to filter transient noise to ensure temporal consistency across multiple detection frames; sliding window detection is employed, with temporal stability based on the following: ; in, Indicates the length of the sliding window. Indicates the Frame detection confidence, and These represent the confidence threshold and the minimum pass rate, respectively. Finally, the recorded target locations and label information are represented in the map constructed above, thus completing the semantic fusion map. Spatial domain filtering technology is used to achieve target deduplication. When a target object is detected, the system calculates the target's three-dimensional position in the world coordinate system and compares it with the targets previously recorded in the target storage table. If the spatial distance between the two is less than a preset threshold and they have the same semantic label, the system determines that they are the same target and will not record them again; if the difference is large, the target is stored as a new entry in the target storage table.
[0032] In this way, the system can track the positional changes of target objects in real time during SLAM and determine which targets are stable. Based on this continuously updated target storage table, the system can accurately count the number of targets to estimate output. Simultaneously, the system integrates the positional information of these semantically tagged target objects into the environmental map, ultimately constructing a navigation map rich in semantic information.
[0033] Furthermore, to address the need for detailed underwater mapping in edge computing scenarios, this invention proposes a 3DGS mapping strategy based on an underwater physical imaging model. This scheme enables detailed modeling of specific areas. Considering the unique light attenuation and backscattering issues in the underwater environment, this method integrates the physical imaging model into the 3DGS mapping process: by analyzing the depth map generated by 3DGS, the water attenuation coefficient and scattering parameters are estimated in real time, and an underwater optical compensation model is constructed accordingly. This model uses an iterative optimization algorithm to superimpose the simulated attenuation and scattering effects onto the original image and compares it with the actual observed image, gradually correcting the reconstruction error, and finally obtaining a high-fidelity 3D reconstruction result that closely approximates the real underwater scene. This technology effectively solves the problem of insufficient modeling accuracy of traditional pixel methods in complex underwater environments. Traditional 3DGS uses a set of 3D Gaussians to parameterize the scene. Each Gaussian has attributes such as mean µ, covariance Σ derived from scale S and rotation R, opacity o derived from spherical harmonic coefficients, and color c. Given a viewpoint, an image can be rendered efficiently by sorting Gaussians from front to back, projecting them onto the camera plane, and performing alpha synthesis.
[0034] Therefore, step 3, performing physical-driven 3DGS fine reconstruction, includes: initial 3DGS modeling and underwater medium estimation and compensation; integrating the physical imaging model into the 3DGS mapping process, estimating the water attenuation coefficient and scattering parameters in real time by analyzing the depth map generated by 3DGS, and constructing an underwater optical compensation model accordingly. Given a viewpoint, an image can be rendered efficiently by sorting Gaussians from front to back, projecting them onto the camera plane, and performing alpha synthesis; for a pixel x, the color of pixel x is calculated as: ; in, and Indicates the i Given a 3D Gaussian color and density, the underwater image formation process can be represented as follows: ; in, Indicates the actual captured image. This represents a realistic underwater scene image after removing the influence of the medium. Indicates the attenuation coefficient. Represents the backscattering coefficient. Representing the backscattered color at infinity, it is assumed that... Used as the background color.
[0035] The method for embedding underwater image physical modeling models into 3DGS is as follows: First, the depth map of the scene is obtained from the initial few rounds of 3DGS rendering. ,pass Estimate attenuation coefficient Scattering coefficient and background color The effects of attenuation and scattering can then be obtained:
[0036]
[0037] Then, the estimated underwater real image is superimposed using the following formula. Up, the first round can Set it directly to the original image;
[0038] The resulting image will then be generated. By constructing a reconstruction loss model that closely approximates the original image, and through multiple iterations, an image close to the actual underwater image can be obtained. J .
[0039] Finally, step 4: output a multi-scale semantic map.
[0040] Example 1 This invention mounts the system on an underwater robot to perform real-time underwater sensing tasks: 1. Generate semantic information: Semantic Information Generation Method: This invention employs a multimodal data collaborative detection strategy. First, a YOLOv5 detection model is trained based on underwater visual and sonar image datasets, respectively. In actual deployment, visual and sonar data are simultaneously acquired using a ZED2i binocular camera and an M1200d sonar device. A dual-thread parallel processing mechanism is used to input the two modalities of data into the YOLOv5 model for real-time inference. To improve detection reliability, the system dynamically calculates image quality scores through a visual quality assessment module and adaptively weights and fuses the detection results from the visual and sonar modalities accordingly, ultimately outputting a semantic detection result with a confidence score.
[0041] ; Visual sensitivity coefficient Set to 5, sonar compensation intensity coefficient Set it to 3 and increase the visual score. Normalized to [0, 1], the output characteristics of the weight formula are as follows:
[0042] This allows us to assign weights to the inference results of the two modalities, and then use time-domain and spatial-domain filtering to determine whether the target object is stable. ; The sliding window length N in the temporal stability criterion is set to 10, the confidence threshold γ to 0.5, and the minimum pass rate α to 0.7. The spatial domain criterion for determining whether two inferences are close to the target is: ; The distance threshold is used, and then combined with the label, it can be determined whether they are the same target and the target storage table is updated.
[0043] 2. Real-time construction of coarse-grained maps: Visual and inertial data acquired by the ZED2i device are packaged and sent to the SLAM module via a topic through the ROS platform. The SLAM module first performs data preprocessing, including: The received image is histogram equalized, and mirror and tangential distortion corrections are performed based on pre-calibrated camera intrinsic parameters and distortion coefficients.
[0044] Feature points are extracted using FAST corner detection. The image grid is divided into blocks, with 20×20 blocks as the unit. The five points with the highest response values are retained in each block. The KLT optical flow method is used to track feature points in adjacent frames.
[0045] Epipolar correction is performed on adjacent frames of images. The fundamental matrix F is calculated using epipolar geometric constraints, and the essential matrix E can be obtained by combining the camera intrinsic parameters K.
[0046] ; Then, by performing SVD decomposition on the essential matrix E, the relative pose between the two frames can be solved. . Based on the relative pose, the matched feature points are linearly triangulated to calculate 3D coordinates, and an initial map point set is constructed.
[0047] The raw data of the IMU, including angular velocity and acceleration, over the time interval Pre-integration is performed to obtain attitude and position information.
[0048] ; ; The camera pose estimated by visual SfM is aligned with the IMU pre-integration result. Then, the visual reprojection error, IMU pre-integration error, and scale / gravity error are jointly optimized using the Gauss-Newton method, and the state is updated in real time.
[0049] The Visual Bag-of-Words (DBoW2) model is used to match the features of the current frame with historical keyframes and calculate the similarity score. If the frames are determined to be similar, the relative poses of the two frames are calculated and loop closure constraints are inserted as edges into the pose graph to connect the current frame and the candidate frames, thereby correcting the overall pose.
[0050] Pixels with similar colors, textures, and spatial locations are extracted from the visual image to form superpixels and construct surface elements. Each superpixel is then back-projected onto a depth map to obtain a set of 3D points, which are then fitted to a local plane. A threshold for the distance between the center position and a threshold for the angle between the normal vectors are used to determine if they can be successfully associated with existing surface elements. If the association is successful, the surface element information is updated through weighted fusion. When the inertial navigation system detects this, all surface elements undergo rigid body transformation based on the pose map optimization results. The final map is a collection of all surface elements, which can also be further converted into a dense point cloud.
[0051] 3. Offline construction of fine-grained maps Using the pose obtained from SfM above and the initial 3D points as input, we construct an initial 3D Gaussian scene together with the image. Then, we gradually refine the scene by replicating and splitting the Gaussian sphere. Since the underwater scene color is affected by multiple disturbances, we use an underwater image formation model to constrain the rendering process to restore the true color of the scene. The underwater image formation process can be represented by the following formula: ; First, a scene depth map is obtained from the initial 1000 rounds of 3DGS rendering. ,pass The attenuation coefficient is obtained based on the attenuation effect. Scattering coefficient and background color The effects of attenuation and scattering can then be obtained: ; ; The two effects mentioned above are superimposed on the estimated underwater real image. Up, the first round of training will directly... Simply set it to the original image, and the resulting image will be displayed. By constructing a reconstruction loss with the original image, the two can be made as close as possible. After multiple iterations, an image J that is close to the real underwater image and a scene with restored real colors can be obtained.
[0052] The sequence numbers of the above embodiments of the present invention are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments. In the above embodiments of the present invention, the descriptions of each embodiment have their own emphasis; parts not described in detail in a certain embodiment can be referred to in the relevant descriptions of other embodiments. It should be understood that the disclosed technical content in the several embodiments provided in this application can be implemented in other ways.
[0053] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for underwater multimodal intelligent sensing for marine ranches, characterized in that, The following steps are involved: Step 1: Based on the YOLO underwater multimodal detection module, sonar sensing modes are introduced to construct a dynamic adaptive multimodal perception system; Step 2: Perform semantic SLAM and dense coarse graph construction; The process of semantic SLAM and dense coarse mapping includes: preprocessing of visual-inertial data, initialization and joint optimization, dense mapping based on surface elements, and semantic fusion mapping. Step 3: Perform physical-driven 3DGS fine reconstruction, including: initial 3DGS modeling and underwater medium estimation and compensation; Step 4: Output a multi-scale semantic map.
2. The underwater multimodal intelligent sensing method for marine ranching according to claim 1, characterized in that, Step 1 includes the following steps: Step 11: Evaluate the quality of the visual data and establish a visual quality evaluation system based on the average value of the brightness gradient amplitude; the formula for calculating the brightness gradient amplitude is: ; in, Indicates the image in coordinates Pixel intensity at that location and These represent the horizontal and vertical gradients calculated using the Sobel operator, respectively. Step 12, the quantitative evaluation indicators are: ; in, Indicates image resolution, Indicates the visual image quality score; Step 13: Determine if the visual image quality score is too low; if the visual score is too low, increase the weight of the sonar detection result, and vice versa. The final detection result is corrected by the visual data quality score.
3. The underwater multimodal intelligent sensing method for marine ranches according to claim 2, characterized in that, The formula for calculating the visual and sonar detection weights is as follows: ; in, Indicates visual clarity score, and This represents the modal sensitivity coefficient.
4. The underwater multimodal intelligent sensing method for marine ranching according to claim 1, characterized in that, In step 2, the preprocessing of the visual-inertial data is as follows: ; in, Represents the rotational pre-integral quantity, i.e., from time t. i arrive j The relative rotational change is obtained by integrating the continuous IMU angular velocity measurements; Indicates the first k Angular velocity measured by each IMU; This represents the velocity pre-integral quantity, that is, in i The velocity change obtained by integration in the time-time coordinate system; Indicates the first k Linear acceleration measured by an IMU; Indicates the position pre-integral quantity, that is, in i The change in displacement in the time coordinate system.
5. The underwater multimodal intelligent sensing method for marine ranching according to claim 1, characterized in that, The initialization and joint optimization adopt a loosely coupled approach, aligning the IMU's pre-integral values with the visual trajectory: ; in, s represents the scale factor, and g represents the direction of gravity; When using a sliding window to perform local optimization on the visual-inertial navigation system, the state variables are: ; in, Indicate pose, velocity, and IMU bias; then the optimized objective function is: ; in, represents the IMU pre-integration residual, This indicates visual reprojection error. Represents Huber robust kernel coefficients; The visual bag-of-words model is used to detect loop closures during semantic SLAM and dense coarse-scale graph construction. Feature points in the current frame are matched with historical keyframes, and similarity scores are calculated. If similar frames are identified, they are marked as loop closures. If loop closures exist, global optimization is performed, i.e., loop closure constraints are added to the pose graph for optimization. ; in This represents the solution obtained by minimizing the objective function. i The optimal pose estimation for each keyframe is calculated; otherwise, the pose with 6 degrees of freedom is directly output.
6. The underwater multimodal intelligent sensing method for marine ranching according to claim 1, characterized in that, Each cell in the cell-based dense mapping is represented by the following parameters: ; in, Indicates the center position of the face element. The normal vector of a surface element. Indicates the radius of the element. Indicates color information, The timestamp of the last update of the face element; First, pixels with similar colors, textures, and spatial locations are extracted from the visual image to form superpixels, which are then associated with surface elements. The method involves backprojecting each superpixel through a depth map to obtain a set of 3D points and fitting a local plane. A threshold for the distance between the center position and a threshold for the angle between the normal vectors are used to determine whether the superpixel can be successfully associated with existing surface elements. If the association is successful, the surface element information is updated through weighted fusion. ; in, and This represents the updated face element position and normal vector. Indicates the center position of the superpixel. The normal vector of a superpixel. and All represent confidence levels; Update the color and radius of the opposite element: ; in, Indicates the expansion coefficient; If association with an existing facet fails, the center point, normal vector, and color information of the superpixel fitting will be stored in the new facet. Modify the timestamp to the most recently updated timestamp; when the SLAM system detects a loop closure, all facets need to undergo rigid body transformation based on the pose graph optimization results; if there is no loop closure, no processing is required. ; in, Indicates from keyframe i arrive j The pose transformation matrix optimized by loop closure; Global consistency is ensured through pose graph optimization: ; in, This represents the set of edges constrained by loops; the final map is a set of all polygons. If there are navigation requirements or other situations that require a dense point cloud map, the polygon map can be converted into a dense point cloud map by only retaining the center point position parameter.
7. The underwater multimodal intelligent sensing method for marine ranching according to claim 1, characterized in that, The semantic fusion mapping method in step 2 is as follows: Semantic information is integrated during the mapping process. During semantic SLAM and dense coarse mapping, the frames captured by the camera are sent in real time to the underwater multimodal detection module based on YOLO to obtain the semantic labels and location information of the target objects. Since the detection results are based on the camera coordinate system, they are transformed to the world coordinate system through pose transformation. The transformation formula is as follows: ; in, Represents a 3D point in the camera coordinate system. Represents the corresponding point in the world coordinate system. This represents a 3×3 rotation matrix, which is the pose transformation from the camera coordinate system to the world coordinate system. This represents a 3×1 translation vector, which is the position of the camera's optical center in the world coordinate system; Obtained directly from the output of the SLAM module This allows us to calculate the target's position in the world coordinate system. After time and spatial filtering, we can determine whether the target is stable and whether it has been recorded. If it has not been recorded, we record the target's label and its position in the world coordinate system. If it has been recorded, we skip this recording step. Temporal filtering is used to filter transient noise to ensure temporal consistency across multiple detection frames; sliding window detection is employed, with temporal stability based on the following criteria: ; in, Indicates the length of the sliding window. Indicates the first Frame detection confidence, and These represent the confidence threshold and the minimum pass rate, respectively. Finally, the recorded target locations and label information are represented in the map construction described above, thus completing the semantic fusion map construction.
8. The underwater multimodal intelligent sensing method for marine ranching according to claim 1, characterized in that, In step 3, the physical imaging model is integrated into the 3DGS mapping process. By analyzing the depth map generated by 3DGS, the water attenuation coefficient and scattering parameters are estimated in real time, and an underwater optical compensation model is constructed accordingly. Given a viewpoint, an image can be rendered efficiently by sorting Gaussians from front to back, projecting them onto the camera plane, and performing alpha synthesis; for a pixel x, the color of pixel x is calculated as: ; in, and Indicates the first i Given a 3D Gaussian color and density, the underwater image formation process can be represented as follows: ; in, Indicates the actual captured image. This represents a realistic underwater scene image after removing the influence of the medium. Indicates the attenuation coefficient. Represents the backscattering coefficient. Representing the backscattered color at infinity, it is assumed that... Used as the background color.
9. The underwater multimodal intelligent sensing method for marine ranching according to claim 8, characterized in that, The method for embedding underwater image physical modeling models into 3DGS is as follows: First, the depth map of the scene is obtained from the initial few rounds of 3DGS rendering. ,pass Estimate attenuation coefficient Scattering coefficient and background color The effects of attenuation and scattering can then be obtained: Then, the estimated underwater real image is superimposed using the following formula. Up, the first round can Set it directly to the original image; The resulting image will then be generated. By constructing a reconstruction loss model that closely approximates the original image, and through multiple iterations, an image close to the actual underwater image can be obtained. J .
Citation Information
Patent Citations
Self-adaptive underwater multi-beam synchronous positioning and mapping method
CN110726415A
Dynamic scene SLAM method based on YOLO algorithm and GMS feature matching
CN111161318A
Bionic polarization semantic SLAM method based on neural radiation field
CN119444857A
Structured scene visual slam method based on point line surface features
WO2023184968A1
Cited By
Dynamic refraction visual correction method for ultra-shallow water blue-green laser sounding
CN121739980A
Dam body defect composite sensing method based on sound-light thickness guidance
CN121789022A