Underwater target intelligent identification system based on multi-sensor fusion
Through a multi-sensing fusion system, combined with optical, acoustic, three-dimensional laser and water quality sensors, the accuracy and reliability problems of underwater target recognition in dynamic environments are solved, and stable identification in complex underwater environments is achieved.
Patent Information
- Application Number
- CN202510598513.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-09
- Publication Date
- 2025-08-15
AI Technical Summary
In the prior art, underwater target recognition relies on a single sensor, and it is difficult to adapt to dynamically changing underwater environments, resulting in a decrease in recognition accuracy and reliability.
A multi-sensing fusion system is adopted, including optical sensors, acoustic sensors, three-dimensional laser sensors and water quality environmental sensors. Through data preprocessing, multi-modal feature fusion and intelligent analysis, robust underwater target recognition results are generated.
It improves the robustness and accuracy of target recognition in complex underwater environments, can adapt to water quality changes, and provides reliable target information support.
Smart Images

Figure CN120495861A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of underwater target recognition, and in particular to an underwater target intelligent recognition system based on multi-sensor fusion. Background Art
[0002] Underwater robots, devices capable of performing a variety of tasks underwater, have been applied in a variety of fields, including marine exploration, resource development, environmental monitoring, underwater security, and scientific research. Improving the autonomous perception and information processing capabilities of underwater robots, particularly the accuracy and efficiency of underwater target recognition, is crucial for their effective mission execution.
[0003] Underwater target recognition technology involves underwater robots acquiring environmental data through onboard sensors and applying techniques such as signal processing, image analysis, and pattern recognition to detect, classify, and locate specific targets in the underwater environment. Characteristics of the underwater environment, such as light attenuation and scattering in water, can affect the quality of optical images, manifesting as color deviation, decreased contrast, and blurred details. Furthermore, suspended matter, currents, and changes in illumination can also affect sensor data acquisition. The propagation characteristics of acoustic signals in water also have limitations. Therefore, acquiring stable and reliable target information and achieving accurate recognition in complex underwater environments remains an ongoing research area in the field of underwater target recognition technology.
[0004] Chinese patent application publication number CN119559489A discloses a method for underwater target recognition for underwater intelligent robots. This method first enhances captured underwater images using a multi-scale enhancement algorithm with color restoration based on improved bilateral filtering to improve image clarity and color reproduction. Subsequently, a network model for underwater target recognition, based on an improved feature fusion attention network architecture, is used to output clear, fog-free images and identify underwater targets in an end-to-end manner. This invention aims to improve the feasibility, accuracy, and real-time performance of underwater target recognition algorithms deployed on mobile devices such as underwater intelligent robots. This technical solution primarily focuses on underwater target recognition using optical images, with emphasis on the design of image enhancement algorithms and lightweight neural network models. However, this solution has the following drawbacks: It primarily relies on optical images as the sole source of information for target recognition. In underwater environments, particularly in turbid water or poor lighting conditions, the quality of optical images can significantly degrade, directly impacting the accuracy and reliability of subsequent target recognition. For the complex task of underwater target recognition, relying solely on optical images may not be sufficient for all scenarios. Although this method uses image enhancement technology, its processing strategy lacks an adaptive adjustment mechanism for dynamic changes in the underwater environment (such as real-time changes in water quality). This means that when environmental conditions change, fixed image enhancement parameters and network models may not be able to maintain optimal recognition results.
[0005] Therefore, how to effectively integrate the information advantages of multiple sensors, overcome the limitations of a single sensor, and combine intelligent algorithms to achieve accurate and stable identification of underwater targets is an urgent problem that needs to be solved in the current underwater detection field. Summary of the Invention
[0006] In view of this, the present invention addresses the problems existing in current underwater target recognition technology, such as over-reliance on single sensor information and difficulty in adapting to dynamically changing underwater environments. It provides a system and method that can effectively integrate the information advantages of multiple sensors, overcome the limitations of a single sensor, and combine with intelligent algorithms to achieve accurate and stable identification of underwater targets.
[0007] The technical solution of the present invention is achieved as follows:
[0008] The present invention provides an underwater target intelligent recognition system based on multi-sensor fusion, comprising:
[0009] An underwater robot equipped with optical sensors, acoustic sensors, three-dimensional laser sensors, and water quality environment sensors;
[0010] Data acquisition module, used to synchronously obtain raw multimodal data and water quality parameters from the underwater robot;
[0011] The data preprocessing module is used to evaluate the water quality level of the current water area based on water quality parameters, and perform parameterized preprocessing on the original multimodal data according to the water quality level to generate preprocessed multimodal data;
[0012] The multimodal feature fusion module is used to extract the features of each modality from the preprocessed multimodal data, determine the weight of each modality feature based on the water quality and feature quality, and use attention operation and weight combination to fuse the features of each modality and their time series information to generate fused features;
[0013] The target recognition and positioning module uses a hierarchical target detection network optimized for underwater environments to process fused features and output the category, location, and confidence level of underwater targets.
[0014] The intelligent analysis output module is used to conduct comprehensive analysis of underwater target information and generate analysis results and visual data based on the preset knowledge base.
[0015] Preferably, the data acquisition module adopts a synchronization mechanism based on the network time protocol or the precision time protocol, combined with the trigger signals of each sensor, to control the alignment error of the multimodal data collected by the optical sensor, acoustic sensor, three-dimensional laser sensor and water quality environment sensor in the time dimension to the millisecond level.
[0016] Preferably, the processing process of the data preprocessing module includes:
[0017] Calculate the comprehensive water quality assessment index based on water quality parameters, including turbidity, salinity, temperature and light attenuation coefficient;
[0018] According to the value range of the comprehensive water quality assessment index, it is divided into multiple discrete water quality levels;
[0019] The water quality level is mapped to the preprocessing parameters of each modal data through a preset parameter mapping table. Each modal data is preprocessed based on the preprocessing parameters to obtain preprocessed multimodal data:
[0020] Perform descattering, color correction and enhancement processing on the optical sensor data, where the descattering intensity and contrast enhancement factor increase with the increase of water quality level;
[0021] performing noise suppression and echo enhancement processing on the acoustic sensor data, wherein the filter window size decreases with increasing water quality level and the enhancement threshold decreases with increasing water quality level;
[0022] Outlier removal and noise reduction are performed on 3D laser sensor data, where the outlier determination threshold decreases with the increase of water quality level.
[0023] Preferably, the calculation formula of the comprehensive water quality assessment index is as follows:
[0024]
[0025] Among them, P i represents the measured value of the i-th water quality parameter, N is the number of water quality parameters, g i (P i ) is a function that normalizes water quality parameters and maps them to the degree of influence of the parameters on a specific sensor, ω i is the importance weight of the i-th water quality parameter in the comprehensive assessment, satisfying ∑ω i =1.
[0026] Preferably, the multimodal feature fusion module includes:
[0027] A parallel feature extraction network is used to extract each initial modal feature from the preprocessed multimodal data;
[0028] A feature quality assessment network is used to calculate quality assessment indicators for each initial modal feature based on the current water quality level, feature clarity and uniformity;
[0029] The modal weight adaptive network is used to calculate the weight coefficient of each modal feature in the fusion process according to the water quality level and the quality assessment index of each modal feature, and use the weight coefficient to weight each modal feature to obtain the weighted feature;
[0030] The temporal feature integration network is used to integrate the weighted features of each modality at different time steps through the modality-time joint attention mechanism to generate the final fusion feature.
[0031] Preferably, in the target recognition and positioning module, the target detection network includes:
[0032] Feature extraction backbone network, used to extract shallow, medium and deep multi-scale feature maps from the fused features;
[0033] Multi-scale feature enhancement network, which is used to integrate feature maps of different scales output by the backbone network and enhance feature expression through cross-scale channel attention and spatial attention;
[0034] Hierarchical object detection head, including:
[0035] A coarse-grained classification head, used to generate a heat map of the center points of the target categories;
[0036] A refined classification head is used to predict subcategories for a specific large category;
[0037] The attribute regression head is used to predict the size, offset and other attribute parameters of the target.
[0038] Preferably, the loss function used by the target detection network during training is:
[0039] L total =λ1·L hier-reg +λ2·L SSFL-heatmap +λ3·L size +λ4·L offset
[0040]
[0041]
[0042] Where, L hier-reg is the class consistency regularization loss; L SSFL-heatmap is the scale-sensitive focused heatmap loss; L size is the size regression loss; L offset is the center point offset loss; λ1, λ2, λ3, and λ4 are weight coefficients for balancing the contributions of different loss terms; I(·) is the indicator function, n is the sample index, and C coarse,n is the rough classification prediction result of the nth sample, C fine,n is the prediction result of fine classification, k0 is the coarse classification category, SubClass is the set of fine subclasses corresponding to the coarse classification k0, P(·) is the prediction probability; N is the number of valid positions in the heat map, H xy is the predicted heat map value at position (x,y), is the true heat map value, S gt is the actual size of the target, α and δ are focusing parameters, and β<0 is a scale-sensitive factor; L1 smooth is the smooth L1 loss function, m is the target index, M is the total number of targets, s m is the predicted size of the m-th target, is the actual size, γ is the size balance coefficient; o m is the predicted center point offset of the mth target, is the actual offset, and η is the size correlation coefficient.
[0043] Preferably, the underwater targets include:
[0044] Artificial targets: including shipwrecks, underwater pipelines, underwater cables, artificial reefs and archaeological remains;
[0045] Natural targets: including reefs, seabed topographic features, and seabed mineral resources;
[0046] Underwater life: including large aquatic organisms, fish schools and coral reefs;
[0047] The target recognition and positioning module predicts the direction and length of linear targets, the coverage and volume estimation of area targets, and the category and activity status of biological targets.
[0048] Preferably, the intelligent analysis output module includes:
[0049] A structured target information generation unit, used to generate a detection report containing underwater target category, location, size, quantity and feature description;
[0050] An underwater map construction unit is used to combine the identified underwater target information with the positioning data to dynamically construct and update the underwater environment map;
[0051] A navigation assistance unit, which is used to provide obstacle avoidance and path planning solutions for the underwater robot based on the detected underwater targets;
[0052] The risk assessment unit is used to analyze the impact of identified underwater targets on navigation safety and trigger an alarm when high-risk targets are found.
[0053] Preferably, the intelligent analysis output module further includes:
[0054] Environmental change monitoring unit, used to analyze the changing trends of underwater targets and environment by comparing data collected at different times;
[0055] The mission planning optimization unit is used to adjust the underwater robot's cruising strategy and sensor configuration according to the distribution of underwater targets and water quality conditions.
[0056] The present invention has the following beneficial effects compared to the prior art:
[0057] (1) This invention improves the overall performance of underwater target recognition by constructing a system that integrates multimodal data acquisition, adaptive preprocessing based on water quality parameters, dynamic feature fusion, and hierarchical target detection and intelligent analysis output. Compared with existing technologies, this invention can more comprehensively perceive the underwater environment, effectively overcoming the limitations of single sensors in complex underwater conditions. Through real-time assessment of water quality conditions and dynamic adjustment of processing strategies, it ensures recognition robustness and accuracy in different hydrological environments, thereby providing reliable target information support for the autonomous operation of underwater robots.
[0058] (2) The present invention simultaneously collects optical, acoustic, 3D laser, and water quality data, and performs parameterized preprocessing on each modal data based on the real-time water quality assessment. This mechanism enables the system to optimize the information acquisition quality of each sensor according to the specific conditions of the current water area, effectively reducing the impact of water quality changes on subsequent recognition performance, thereby maintaining high recognition stability and accuracy in a diverse and dynamically changing underwater environment.
[0059] (3) In the multimodal feature fusion module, the present invention not only extracts the features of each modality, but also introduces a feature quality assessment network and a modal weight adaptive network. The system can dynamically calculate the weight coefficient of each modal feature in the fusion process based on the current water quality level and quality indicators such as the clarity and uniformity of each modal feature, and integrate the features in combination with the modal-temporal joint attention mechanism. This design enables the system to intelligently focus on the modality with higher information quality and greater contribution to target recognition under current conditions, suppress the interference of noise or low-quality information, thereby generating more discriminative fusion features and improving the accuracy of target recognition;
[0060] (4) The target recognition and positioning module of the present invention utilizes a deep learning network structure comprising a feature extraction backbone network, a multi-scale feature enhancement network, and a hierarchical target detection head. This hierarchical design enables a rough classification of targets followed by refined sub-category prediction and attribute regression. Combined with multi-scale feature enhancement, it effectively handles the challenges of large size variations and unclear features of underwater targets. It also outputs rich information such as the target's specific category, location, size, and offset, improving the level of refinement in recognition. BRIEF DESCRIPTION OF THE DRAWINGS
[0061] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0062] Figure 1 Schematic diagram of the system module of the present invention;
[0063] Figure 2 It is a schematic diagram of the technical implementation of the present invention;
[0064] Figure 3 This is a data processing flow chart of the multimodal feature fusion module of the present invention;
[0065] Figure 4 This is a data processing flow chart of the target identification and positioning module of the present invention. DETAILED DESCRIPTION
[0066] The following will be combined with the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0067] like Figure 1 As shown, the present invention provides an underwater target intelligent recognition system based on multi-sensor fusion, comprising:
[0068] An underwater robot equipped with optical sensors, acoustic sensors, three-dimensional laser sensors, and water quality environment sensors;
[0069] Data acquisition module, used to synchronously obtain raw multimodal data and water quality parameters from the underwater robot;
[0070] The data preprocessing module is used to evaluate the water quality level of the current water area based on water quality parameters, and perform parameterized preprocessing on the original multimodal data according to the water quality level to generate preprocessed multimodal data;
[0071] The multimodal feature fusion module is used to extract the features of each modality from the preprocessed multimodal data, determine the weight of each modality feature based on the water quality and feature quality, and use attention operation and weight combination to fuse the features of each modality and their time series information to generate fused features;
[0072] The target recognition and positioning module uses a hierarchical target detection network optimized for underwater environments to process fused features and output the category, location, and confidence level of underwater targets.
[0073] The intelligent analysis output module is used to conduct comprehensive analysis of underwater target information and generate analysis results and visual data based on the preset knowledge base.
[0074] like Figure 2 As shown in the figure, the technical implementation process of the present invention is as follows: first, multimodal sensor data such as optical, acoustic, and three-dimensional laser and real-time water quality environmental parameters are synchronously collected; then, each modal data is adaptively preprocessed according to the assessed water quality level to optimize the quality of the original data; then, a deep neural network is used to extract features from the preprocessed multimodal data, and an attention fusion mechanism that dynamically adjusts the modal weights based on feature quality and water quality conditions is used to generate robust and information-rich fusion features; finally, the fusion features are input into a hierarchical target detection network optimized for underwater environments to achieve accurate classification, positioning and attribute analysis of underwater targets, and the recognition results are used to generate structured reports, build environmental maps and other intelligent applications, thereby overcoming the limitations of single sensors in complex underwater environments and improving the accuracy, robustness and intelligence of target recognition.
[0075] The system of the present invention is a distributed system with a layered architecture, primarily consisting of a hardware layer and a software layer. The hardware layer primarily includes: an underwater robot platform, which serves as the system's carrier and is responsible for carrying various sensors while traveling and operating in an underwater environment; a multimodal sensor system, installed on the underwater robot and comprising an optical camera, acoustic equipment, and a three-dimensional laser scanner; a water quality sensor for measuring water quality parameters such as turbidity, salinity, temperature, and light attenuation coefficient; and an onboard computing unit, which performs preliminary data processing and time-sensitive calculations on the underwater robot. The software layer includes a data preprocessing module, a multimodal feature fusion module, a target identification and positioning module, and an intelligent analysis and output module. The software layer is primarily built into the backend data center, with a hierarchical communication strategy employed between the two, with processing locations and communication timing determined based on task urgency and data volume.
[0076] Specifically, in one embodiment of the present invention, the underwater robot can be a cable-operated underwater robot (ROV) or an autonomous underwater vehicle (AUV). To achieve multimodal perception, the deployment of sensors on the underwater robot platform comprehensively considers factors such as the field of view of each sensor, detection range, mutual interference avoidance, and operational requirements.
[0077] The optical sensor uses at least one high-resolution industrial camera, mounted, for example, on the front of the AUV or at an observation point within the operating area. The camera lens should be a dedicated underwater lens with an appropriate focal length and field of view. It should also be equipped with an array of LED fill lights. The layout of the fill lights should provide uniform illumination and minimize backscattering caused by suspended matter. The entire camera and fill light system should be enclosed in a high-pressure, waterproof enclosure.
[0078] The acoustic sensor uses a high-frequency imaging sonar, mounted on the front of the underwater robot. Its acoustic beam emission direction is roughly aligned with or overlaps with the optical camera's field of view. The sonar transducer array should be installed to avoid obstruction by the robot's main structure or other ancillary equipment to minimize acoustic blind spots. For applications requiring side scan or 3D imaging, a side scan or 3D scanning sonar can be added, with appropriate placement on the sides or bottom of the robot.
[0079] The 3D laser sensor uses an underwater 3D laser scanner, mounted alongside or adjacent to an optical camera, to obtain 3D point cloud data of the target using laser scanning. The laser transmitter and receiver are also enclosed in a pressure-resistant and waterproof housing, and care is taken to prevent direct laser light from reaching the optical camera lens.
[0080] The water quality environment sensor uses a multi-parameter water quality sensor probe, such as an integrated probe that can measure turbidity NTU, salinity, temperature, chlorophyll, dissolved oxygen and other indicators. It is installed on the underwater robot body in a position where it can be exposed to free water flow to avoid being directly affected by the robot's propeller or other components' emissions, so as to obtain representative water environment parameters.
[0081] All sensors are connected to the data acquisition and control system inside the robot through high-reliability underwater connectors and special cables.
[0082] Specifically, in one embodiment of the present invention, the data acquisition module adopts a synchronization mechanism based on the network time protocol or the precision time protocol, combined with the trigger signals of each sensor, to control the alignment error of the multimodal data collected by the optical sensor, acoustic sensor, three-dimensional laser sensor and water quality environment sensor in the time dimension to the millisecond level.
[0083] In this embodiment, each sensor transmits data to the main control system through an interface. In order to make the data collected by different sensors have a unified timestamp, the Network Time Protocol (NTP) or the Precision Time Protocol (PTP) can be used. The specific synchronization process is: the main control computer is configured as an NTP client or a PTP slave clock. The client / slave clock regularly synchronizes time with the master clock source to calibrate their respective local system clocks. The synchronization frequency is set according to the clock drift rate. When each acquisition unit (optical, acoustic, laser, water quality) collects a frame of data or a reading, the data acquisition software will immediately obtain the current time from its synchronized local system clock, and record and store this high-precision timestamp together with the data. For example, the shooting time is recorded in the EXIF information of the optical image, the Ping time is recorded in the acoustic data packet header, the scanning time is recorded in the metadata of the laser point cloud file, and the sampling time is recorded in the water quality data.
[0084] Specifically, in one embodiment of the present invention, the data preprocessing module includes a water quality environment assessment unit, an optical image preprocessing unit, an acoustic data preprocessing unit, and a 3D laser data preprocessing unit. The processing process of the data preprocessing module includes:
[0085] Calculate the comprehensive water quality assessment index based on water quality parameters, including turbidity, salinity, temperature and light attenuation coefficient;
[0086] According to the value range of the comprehensive water quality assessment index, it is divided into multiple discrete water quality levels;
[0087] The water quality level is mapped to the preprocessing parameters of each modal data through a preset parameter mapping table. Each modal data is preprocessed based on the preprocessing parameters to obtain preprocessed multimodal data:
[0088] Perform descattering, color correction and enhancement processing on the optical sensor data, where the descattering intensity and contrast enhancement factor increase with the increase of water quality level;
[0089] performing noise suppression and echo enhancement processing on the acoustic sensor data, wherein the filter window size decreases with increasing water quality level and the enhancement threshold decreases with increasing water quality level;
[0090] Outlier removal and noise reduction are performed on 3D laser sensor data, where the outlier determination threshold decreases with the increase of water quality level.
[0091] In this embodiment, the water quality environment assessment unit receives real-time water quality parameter data collected by water quality environment sensors, including but not limited to turbidity (NTU value), salinity (PSU), temperature (°C), and light attenuation coefficient (c). This unit first pre-processes the raw water quality sensor data, including calibration and smoothing, then calculates a comprehensive water quality assessment index, and ultimately outputs the water quality grade to other pre-processing units.
[0092] The calculation formula of the comprehensive water quality assessment index is as follows:
[0093]
[0094] Among them, P i represents the measured value of the i-th water quality parameter, N is the number of water quality parameters, g i (P i ) is a function that normalizes water quality parameters and maps them to the degree of influence of the parameters on a specific sensor, ω i is the importance weight of the i-th water quality parameter in the comprehensive assessment, satisfying ∑ω i =1.
[0095] According to the influence characteristics of different water quality parameters on each sensor, a corresponding mapping function is designed for each water quality parameter in this embodiment:
[0096] For the turbidity parameter P1, its mapping function is as follows:
[0097]
[0098] Where α1 is a parameter that controls the slope of the mapping curve. This function maps turbidity to the interval [0, 1]. Higher turbidity indicates a greater negative impact on the optical sensor, closer to 1.
[0099] For the temperature parameter P2, its mapping function is as follows:
[0100]
[0101] Among them, P 2,optis the optimal working temperature, P 2,range is the temperature range considered. This function reflects the degree to which the water temperature deviates from the optimal operating temperature. The greater the deviation, the greater the adverse effect on the acoustic sensor.
[0102] For the salinity parameter P3, its mapping function is as follows:
[0103]
[0104] Among them, P 3,opt is the best working salinity, P 3,range is the salinity range considered. This function reflects the degree to which the water temperature deviates from the optimal operating temperature. The greater the deviation, the greater the adverse effect on the acoustic sensor.
[0105] For the light attenuation coefficient parameter P4, its mapping function is as follows:
[0106]
[0107] Wherein, α4 is a parameter that controls the slope of the mapping curve, which affects the optical sensor.
[0108] According to Q E The range of values is divided into several discrete water quality levels: excellent (I level, Q E <0.2), good (grade II, 0.2≤Q E <0.4), general (level III, 0.4≤Q E <0.6), poor (IV level, 0.6≤Q E <0.8) and severe (V level, Q E ≥0.8).
[0109] The optical image preprocessing unit receives the raw optical image and water quality information and performs adaptive image enhancement. First, basic preprocessing independent of water quality is performed, including dark field correction, bad pixel repair, and lens distortion correction. Next, preprocessing parameters are dynamically adjusted based on the water quality level. The descattering intensity parameter increases with increasing water quality level. For example, in a Grade I (excellent) water quality environment, this parameter is set to a lower value for mild descattering, while in a Grade V (poor) water quality environment, it is set to a higher value for strong descattering. In actual processing, a dark channel prior algorithm is used to descatter the image. Finally, contrast enhancement is performed using a contrast-limited adaptive histogram equalization algorithm. The contrast enhancement coefficient also increases with increasing water quality level. In poor water quality conditions, the enhancement coefficient is larger, resulting in better detail in low-contrast images. After image enhancement, bilateral filtering is applied for noise suppression to balance any noise that may be introduced during the enhancement process, resulting in the final preprocessed optical image.
[0110] The acoustic data preprocessing unit performs adaptive processing on the raw acoustic data according to the water quality level. First, the acoustic data is subjected to time-varying gain compensation. Next, noise suppression processing is performed, and an adaptive filtering algorithm is used to reduce the noise level of the acoustic image. The filter window size is a key parameter, which decreases with increasing water quality level. For example, in a Class I water quality environment, the filter window is larger; while in a Class V water quality environment, the filter window is smaller. After noise suppression, the acoustic data is subjected to echo enhancement processing to improve the visibility of weak echo targets. The enhancement threshold is another key parameter, which decreases with increasing water quality level. Under poor water quality conditions, a lower threshold is used to enhance weaker echo signals to improve the detection capability of targets. Finally, geometric correction is performed based on the sound velocity, and the sound velocity value is calculated from the water temperature and salinity parameters.
[0111] The 3D laser data preprocessing unit processes the original laser point cloud data, removes noise points and optimizes the point cloud quality. The outlier judgment threshold is a key parameter and decreases with the increase of water quality level. When the water quality is good (level I), a looser threshold is used to retain more original points; when the water quality is poor (level V), a stricter threshold is used to remove more noise points that may be affected by scattering. In specific processing, the statistical outlier filtering algorithm is first used to calculate the average distance of each point to its neighboring points, and outliers are identified and removed based on the distance distribution and the outlier judgment threshold. Then, the point cloud after removing the outliers is subjected to denoising. The moving least squares smoothing algorithm can be used. The denoising parameters are also dynamically adjusted according to the water quality level. A stronger denoising intensity is used under poor water quality conditions. Finally, the processed point cloud data is converted to a unified reference coordinate system.
[0112] Specifically, if Figure 3 As shown, in one embodiment of the present invention, the multimodal feature fusion module includes:
[0113] A parallel feature extraction network is used to extract each initial modal feature from the preprocessed multimodal data.
[0114] Specifically, for optical image data, a pre-trained convolutional neural network (such as ResNet50) is used as a feature extractor. For acoustic data, considering the differences between sonar images and ordinary optical images, the network used contains convolutional layers and nonlinear activation functions optimized for acoustic image characteristics. For three-dimensional laser point cloud data, a feature extraction network based on the PointNet++ architecture is used. This network can effectively process irregular point cloud data and capture local and global geometric features of the point cloud through hierarchical sampling and grouping strategies. The three feature extraction networks work in parallel and output optical image features F respectively. O , sonar image feature F S , laser point cloud feature F LTo ensure that these feature vectors have the same dimension, each feature extraction network is followed by a feature mapping layer to convert the original features into a representation of uniform dimension.
[0115] A feature quality assessment network is used to calculate quality assessment indicators for each initial modal feature based on the current water quality level, feature clarity and uniformity.
[0116] Specifically, a set of feature quality evaluation indicators are designed, including the clarity score of optical images, the signal-to-noise ratio score of acoustic images, and the density uniformity score of point cloud data. These indicators are calculated by analyzing the intrinsic properties of each modal feature. For example, the clarity of optical images can be evaluated by analyzing the gradient distribution or high-frequency components of the image; the signal-to-noise ratio of acoustic images can be evaluated by analyzing the contrast and noise level of the image; and the density uniformity of point clouds can be evaluated by analyzing the spatial distribution of points. The feature quality evaluation network outputs the quality evaluation indicator q k , including three evaluation indicators: q O ,q S ,q L , which represent the quality of optical features, acoustic features and laser point cloud features respectively.
[0117] The modal weight adaptive network is used to calculate the weight coefficient of each modal feature in the fusion process according to the water quality grade and the quality assessment index of each modal feature, and use the weight coefficient to weight each modal feature to obtain the weighted feature.
[0118] Specifically, the Modality Weight Adaptation Network (MWAN) is a small multi-layer perceptron (MLP) whose input is the current water quality level Q level And the quality evaluation index q of each modal feature itself k , the output is the attention weight A of each modality k , the calculation process is as follows:
[0119] A k =Softmax(f MWAN (Q level ,q O ,q S ,q L )) k
[0120] Among them A k is the attention weight of the kth modality, f MWANIt is the forward calculation function of the modal weight adaptive network. The Softmax operation ensures that the sum of all modal weights is 1. For example, in the case of good water quality, the optical feature may get a higher weight, while in the case of poor water quality, the weights of the acoustic feature and the laser point cloud feature will increase accordingly. The weighted features
[0121] In this embodiment, the modal weight adaptation network not only assigns different weights to different modalities, but also introduces a cross-modal guided channel attention mechanism within the modality, further improving the expressiveness of features. Taking optical features as an example, the calculation of their channel attention weights is guided by acoustic features:
[0122]
[0123] Where σ is the Sigmoid function, MLP is the multi-layer perceptron, Pool is the global average pooling, and W SO is a learnable linear transformation matrix that maps the sonar feature space to a space compatible with the optical signature. This allows the potential target region detected by the sonar to enhance the response of the corresponding channel in the optical signature. Similar cross-modal guidance mechanisms exist between other modalities, forming a mutually reinforcing feature network.
[0124] The temporal feature integration network is used to integrate the weighted features of each modality at different time steps through the modality-time joint attention mechanism to generate the final fusion feature.
[0125] Specifically, the temporal feature integration network uses a Transformer-based architecture that can effectively model long-range dependencies and temporal context. Specifically, the weighted features extracted from different modalities on the time series are concatenated or mapped through an embedding layer and then input into a multimodal Transformer encoder.
[0126] The Transformer encoder designs a Modality-Temporal Joint Attention (MTJA) mechanism. Building on the standard Transformer self-attention mechanism, MTJA not only considers the temporal dimension when generating queries, keys, and values, but also explicitly encodes modality information.
[0127] Assume that the input sequence is X=[x1,…,x T ],in Query Q i , key K j , value V j When calculating the attention score, the modality embedding M can be introduced i ,M jand time embedding T i ,T j :
[0128]
[0129] Where W Q ,W K is the weight matrix, d k is the square root of the feature dimension and is used to scale the attention score. This approach enables the attention mechanism to learn the complex associations between a specific modality at a specific time step and other modalities and other time steps. The output F fused-seq It is a timing fusion feature enhanced by MTJA.
[0130] The multimodal feature fusion module is trained in a phased approach. First, each feature extraction network is pre-trained. For the optical feature extraction network, it is pre-trained on a large-scale image classification dataset (such as ImageNet) and then fine-tuned on an underwater image dataset. For the acoustic feature extraction network, it is either trained from scratch or fine-tuned on a sonar image dataset. For the laser point cloud feature extraction network, it is pre-trained on a large point cloud classification dataset.
[0131] The pre-trained feature extraction network is fixed, and the modal weight adaptation network and temporal feature integration network are trained. In this phase, a multimodally annotated underwater target dataset is used to train the network through supervised learning. This allows the network to dynamically adjust modal weights based on water quality and feature quality, while effectively integrating temporal information. Cross-entropy loss is used during training, and network parameters are optimized through backpropagation.
[0132] During actual operation, the workflow of the multimodal feature fusion module is as follows: first, the preprocessed multimodal data and water quality grade information are received; then, the initial features of each modality are extracted through a parallel feature extraction network; then, the quality of each feature is evaluated through a feature quality assessment network; after that, the modal weight adaptation network calculates the weight of each modality according to the water quality grade and feature quality, and applies a cross-modal guided channel attention mechanism to enhance feature expression; finally, the temporal feature integration network fuses the multimodal weighted features at different time steps through a modal-temporal joint attention mechanism to generate the final fusion feature, which is passed to the subsequent target recognition and positioning module.
[0133] Specifically, if Figure 4 As shown, in one embodiment of the present invention, in the target recognition and positioning module, the target detection network includes:
[0134] The feature extraction backbone network is used to extract shallow, medium and deep multi-scale feature maps from the fused features.
[0135] Specifically, taking into account the particularity of underwater images, this embodiment replaces the original backbone network of CenterNet with a lightweight network HRNet that is more suitable for processing the low contrast, blurred details, and color distortion characteristics of underwater images, and combines it with an underwater image enhancement module. The advantage of HRNet is that it can maintain high-resolution feature maps throughout the entire network process and better preserve spatial detail information. In addition, color correction and defogging layers are integrated at the front end of the network to further enhance the quality of input features. The backbone network extracts multi-scale feature maps from shallow to deep layers. Shallow features contain rich texture and edge information, while deep features have stronger semantic expression capabilities.
[0136] The multi-scale feature enhancement network is used to integrate the feature maps of different scales output by the backbone network and enhance the feature expression through cross-scale channel attention and spatial attention.
[0137] Specifically, this embodiment introduces a multi-scale feature-guided composite attention module (MFG-CA) into the multi-scale feature enhancement network. This module uses channel attention and spatial attention in parallel, and the generation of its attention map is jointly guided by feature maps of different scales from the backbone network, enabling it to focus on global context and local details at the same time, especially enhancing the response to small-size, low signal-to-noise ratio targets.
[0138] For a certain layer of feature map F in , its channel attention M c and spatial attention M s In the calculation of , the semantic features F from higher layers will be incorporated high and lower-level detail features F low Information:
[0139] M c (F in ∣∣F high ,F low )=σ(MLP c (Pool(F in ))+W ch Pool(F high )+W cl Pool(F low ))
[0140] M s (F in ∣∣F high ,F low )=σ(Conv s ([AvgPool(F in );MaxPool(F in )])+Conv sh (F′ high )+Convsl (F′ low ))
[0141] Among them, σ is the Sigmoid function, MLP c is a multi-layer perceptron for channel attention, Pool is a pooling operation, W ch and W cl Is a learnable weight matrix used to map high-level and low-level features to the channel attention space. Conv s 、Conv sh 、Conv sl is a convolution operation, AvgPool is an average pooling operation, MaxPool is a maximum pooling operation, F′ high and F′ low is adjusted to the same value as F through convolution and up / down sampling. in High- and low-level features of the same size. This guided composite attention can focus on the target area more accurately.
[0142] Hierarchical object detection head, including:
[0143] A coarse-grained classification head, used to generate a heat map of the center points of the target categories;
[0144] A refined classification head is used to predict subcategories for a specific large category;
[0145] The attribute regression head is used to predict the size, offset and other attribute parameters of the target.
[0146] Specifically, based on the improved CenterNet architecture, the present invention adopts an anchor-free center point representation method to detect targets. Compared with the traditional bounding box regression method, this representation is more suitable for processing underwater targets with complex shapes and blurred edges.
[0147] The coarse-grained classification head is used to generate heat maps of the center points of target categories. Underwater targets are divided into three categories: artificial targets, including shipwrecks, underwater pipelines, underwater cables, artificial reefs, and archaeological remains; natural targets, including reefs, seabed topography features, and seabed mineral resources; and underwater organisms, including large aquatic organisms, fish schools, and coral reefs. Using large-category classification can take advantage of the obvious feature differences between large categories and improve the stability of the initial classification. The coarse-grained classification head outputs K heat maps (K is the number of large categories), each of which represents the probability distribution of the center point location of the corresponding large category target.
[0148] The refined classification head predicts subcategories for a specific broad category. This two-level "broad category-subcategory" classification architecture effectively addresses the complexity and data imbalance of underwater objects. For example, the broad category of artificial objects is further subdivided into subcategories such as shipwrecks, underwater pipelines, and underwater cables; while the broad category of underwater organisms is further subdivided into specific species, such as different types of fish and corals. This hierarchical recognition approach makes more efficient use of limited training data.
[0149] The attribute regression head is used to predict various physical attribute parameters of the target. For linear targets (such as underwater pipes and cables), its direction and length are predicted; for area targets (such as coral reefs and seabed topography), its coverage and volume estimation are predicted; and for biological targets, its activity status is predicted. These attributes are implemented through additional regression branches, including size regression (predicting the width and height of the target), center point offset regression (refining the center point location), and direction regression (predicting the target's main direction angle). For specific types of targets, specialized attribute regression branches can also be added, such as a biological activity status classification branch.
[0150] The target detection network of the present invention includes a pre-training process, and the loss function used in the training is as follows:
[0151] L total =λ1·L hier-reg +λ2·L SSFL-heatmap +λ3·L size +λ4·L offset
[0152]
[0153] Where, L hier-reg is the class consistency regularization loss; L SSFL-heatmap is the scale-sensitive focused heatmap loss; L size is the size regression loss; L offset is the center point offset loss; λ1, λ2, λ3, and λ4 are weight coefficients for balancing the contributions of different loss terms; I(·) is the indicator function, n is the sample index, and C coarse,n is the rough classification prediction result of the nth sample, C fine,n is the prediction result of fine classification, k0 is the coarse classification category, SubClass is the set of fine subclasses corresponding to the coarse classification k0, P(·) is the prediction probability; N is the number of valid positions in the heat map, H xy is the predicted heat map value at position (x,y), is the true heat map value, S gt is the actual size of the target, α and δ are focusing parameters, and β<0 is a scale-sensitive factor; L1 smooth is the smooth L1 loss function, m is the target index, M is the total number of targets, sm is the predicted size of the m-th target, is the actual size, γ is the size balance coefficient; o m is the predicted center point offset of the mth target, is the actual offset, and η is the size correlation coefficient.
[0154] The object detection network is trained using a multi-stage strategy. First, the backbone network is pre-trained using a large-scale general-purpose object detection dataset (such as COCO) to equip it with basic feature extraction capabilities. Then, the entire network is fine-tuned using underwater target datasets. These datasets include real-world underwater images, sonar data, and laser point cloud data, along with corresponding annotation information (such as target category, location, and size). To address the difficulty of acquiring underwater data, synthetic data augmentation techniques can also be used, such as synthesizing ordinary target images into underwater backgrounds or using physical models to simulate the image degradation effects under different water quality conditions.
[0155] During the training process, a curriculum learning strategy can also be adopted to first train the network under simpler water conditions (such as clear water bodies), and then gradually introduce more complex environments (such as turbid water bodies).
[0156] When training an object detection network, you can optionally connect the multimodal feature fusion module for end-to-end joint training or fine-tuning. This end-to-end training approach allows the losses of the object recognition and localization module to be back-propagated to the multimodal feature fusion module, optimizing the generation of fused features and making them more conducive to the final object recognition task. In practice, you can first pre-train the multimodal feature fusion module and the object recognition and localization module independently, and then connect them together for end-to-end joint fine-tuning.
[0157] Specifically, in one embodiment of the present invention, the intelligent analysis output module includes:
[0158] The structured target information generation unit generates an inspection report containing the underwater target's category, location, size, quantity, and characteristic descriptions. This unit first receives basic information such as the target's category, location, size, and quantity, and then extracts specific attribute parameters based on the target type. Specifically, using a predefined report template, this information is organized into a hierarchical structure to generate a comprehensive inspection report containing text descriptions, data tables, and image annotations.
[0159] The underwater map construction unit is used to combine the identified underwater target information with the positioning data to dynamically construct and update the underwater environment map. Specifically, a coordinate reference system is first established to map all identified targets into three-dimensional space according to their geographic location and attribute information. The map construction process adopts an incremental update strategy. New recognition results will be continuously integrated into the existing map, and possible data conflicts will be resolved based on timestamps and location information. In terms of specific implementation, this unit uses a point cloud map representation method, which can simultaneously express the geometric shape and semantic information of the underwater scene. For known fixed targets (such as underwater pipelines, reefs, etc.), the system will continuously refine their position and shape information over multiple observations; for mobile targets (such as schools of fish), their activity trajectories and frequency of occurrence are recorded. The constructed underwater map supports multi-level browsing. Users can choose to view a global overview or detailed information of a specific area, or they can choose to display specific types of targets.
[0160] The navigation assistance unit is used to provide obstacle avoidance and path planning solutions for underwater robots based on detected underwater targets. Specifically, the identified obstacles (such as large aquatic organisms, shipwrecks, reefs, etc.) are first marked as obstacle avoidance areas, and the safety distance is dynamically calculated based on the type, size and position of the target. For high-value detection targets (such as underwater pipelines, archaeological remains, etc.), they are marked as points of interest and the optimal approach path is calculated. By comprehensively analyzing the underwater environment and target distribution, the unit can generate a variety of navigation solutions: global optimal path, obstacle avoidance priority path, energy-optimal path, etc., and recommend the most suitable solution based on the current mission requirements and underwater conditions.
[0161] The risk assessment unit is used to analyze the impact of the identified underwater targets on navigation safety and trigger an alarm when a high-risk target is found. Specifically, a risk assessment model is first established for different types of underwater targets, taking into account factors such as the type of target (such as large active organisms have higher uncertainty than static shipwrecks), size, speed, distance from the underwater robot, relative motion trend, etc. Based on these factors, the system calculates a comprehensive risk index and divides the risk level into three levels: low, medium, and high. For high-risk targets (such as rapidly approaching large aquatic organisms, unstable suspended debris, etc.), the system will immediately trigger an alarm and recommend emergency avoidance measures; for medium-risk targets, the system provides early warning information and recommends adjusting the operation plan; for low-risk targets, the system only marks its location on the navigation map for reference.
[0162] as well as:
[0163] The environmental change monitoring unit is used to analyze the changing trends of underwater targets and the environment by comparing data collected at different times. Specifically, a time series database is first established to store the underwater target information, water quality parameters, and environmental conditions collected by the system at different times. Through data mining and time series analysis techniques, the unit can detect various change patterns: including slow and gradual changes (such as the growth of coral reefs and the accumulation of seabed sediments), periodic changes (such as tidal fluctuations in water quality parameters and seasonal changes in biological activity), and sudden changes (such as new obstacles and sudden changes in water quality). For important changes detected, the system will generate a change report, including the type, scope, rate, and possible cause analysis.
[0164] The mission planning and optimization unit is used to adjust the underwater robot's cruise strategy and sensor configuration based on the distribution of underwater targets and water quality conditions. Specifically, based on the mission objectives (such as pipeline inspection, seabed mapping, biological monitoring, etc.) and current environmental conditions, the underwater robot's cruise route, speed, depth, and posture are optimized to maximize information acquisition efficiency and minimize energy consumption. In response to different water quality conditions, this unit will recommend adjusting the sensor configuration, such as increasing the weight of acoustic sensors in turbid water bodies and prioritizing the use of optical sensors in clear water bodies.
[0165] In this embodiment, the intelligent analysis output module provides a variety of visual interfaces, including: a real-time monitoring interface that displays the currently identified underwater targets and environmental status; a three-dimensional interactive map interface that allows operators to view the underwater environment from different angles and scales; a target details interface that provides in-depth analysis of specific targets; a historical data playback interface for reviewing and analyzing historical operation processes; and a customizable report generation interface to meet the data presentation needs of different users.
[0166] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. The underwater target intelligent recognition system based on multi-sensor fusion is characterized by: include: An underwater robot equipped with optical sensors, acoustic sensors, three-dimensional laser sensors, and water quality environment sensors; Data acquisition module, used to synchronously obtain raw multimodal data and water quality parameters from the underwater robot; The data preprocessing module is used to evaluate the water quality level of the current water area based on water quality parameters, and perform parameterized preprocessing on the original multimodal data according to the water quality level to generate preprocessed multimodal data; The multimodal feature fusion module is used to extract the features of each modality from the preprocessed multimodal data, determine the weight of each modality feature based on the water quality and feature quality, and use attention operation and weight combination to fuse the features of each modality and their time series information to generate fused features; The target recognition and positioning module uses a hierarchical target detection network optimized for underwater environments to process fused features and output the category, location, and confidence level of underwater targets. The intelligent analysis output module is used to conduct comprehensive analysis of underwater target information and generate analysis results and visual data based on the preset knowledge base.
2. The underwater target intelligent recognition system based on multi-sensor fusion according to claim 1 is characterized in that: The data acquisition module adopts a synchronization mechanism based on the Network Time Protocol or the Precision Time Protocol, combined with the trigger signals of each sensor, to control the alignment error of the multimodal data collected by optical sensors, acoustic sensors, three-dimensional laser sensors and water quality environment sensors in the time dimension to the millisecond level.
3. The underwater target intelligent recognition system based on multi-sensor fusion according to claim 1 is characterized in that: The processing of the data preprocessing module includes: Calculate the comprehensive water quality assessment index based on water quality parameters, including turbidity, salinity, temperature and light attenuation coefficient; According to the value range of the comprehensive water quality assessment index, it is divided into multiple discrete water quality levels; The water quality level is mapped to the preprocessing parameters of each modal data through a preset parameter mapping table. Each modal data is preprocessed based on the preprocessing parameters to obtain preprocessed multimodal data: Perform descattering, color correction and enhancement processing on the optical sensor data, where the descattering intensity and contrast enhancement factor increase with the increase of water quality level; performing noise suppression and echo enhancement processing on the acoustic sensor data, wherein the filter window size decreases with increasing water quality level and the enhancement threshold decreases with increasing water quality level; Outlier removal and noise reduction are performed on 3D laser sensor data, where the outlier determination threshold decreases with the increase of water quality level.
4. The underwater target intelligent recognition system based on multi-sensor fusion according to claim 3 is characterized in that: The calculation formula of the comprehensive water quality assessment index is as follows: Among them, P i represents the measured value of the i-th water quality parameter, N is the number of water quality parameters, g i (P i ) is a function that normalizes water quality parameters and maps them to the degree of influence of the parameters on a specific sensor, ω i is the importance weight of the i-th water quality parameter in the comprehensive assessment, satisfying ∑ω i =1.
5. The underwater target intelligent recognition system based on multi-sensor fusion according to claim 1 is characterized in that: The multimodal feature fusion module includes: A parallel feature extraction network is used to extract each initial modal feature from the preprocessed multimodal data; A feature quality assessment network is used to calculate quality assessment indicators for each initial modal feature based on the current water quality level, feature clarity and uniformity; The modal weight adaptive network is used to calculate the weight coefficient of each modal feature in the fusion process according to the water quality level and the quality assessment index of each modal feature, and use the weight coefficient to weight each modal feature to obtain the weighted feature; The temporal feature integration network is used to integrate the weighted features of each modality at different time steps through the modality-time joint attention mechanism to generate the final fusion feature.
6. The underwater target intelligent recognition system based on multi-sensor fusion according to claim 1 is characterized in that: In the target recognition and positioning module, the target detection network includes: Feature extraction backbone network, used to extract shallow, medium and deep multi-scale feature maps from the fused features; Multi-scale feature enhancement network, which is used to integrate feature maps of different scales output by the backbone network and enhance feature expression through cross-scale channel attention and spatial attention; Hierarchical object detection head, including: A coarse-grained classification head, used to generate a heat map of the center points of the target categories; A refined classification head is used to predict subcategories for a specific large category; The attribute regression head is used to predict the size, offset and other attribute parameters of the target.
7. The underwater target intelligent recognition system based on multi-sensor fusion according to claim 6 is characterized in that: The loss function used by the target detection network during training is: L total =λ1·L hier-reg +λ2·L SSFL-heatmap +λ3·L size +λ4·L offset Where, L hier-reg is the class consistency regularization loss; L SSFL-heatmap is the scale-sensitive focused heatmap loss; L size is the size regression loss; L offset is the center point offset loss; λ1, λ2, λ3, and λ4 are weight coefficients for balancing the contributions of different loss terms; I(·) is the indicator function, n is the sample index, and C coarse,n is the rough classification prediction result of the nth sample, C fine,n is the prediction result of fine classification, k0 is the coarse classification category, SubClass is the set of fine subclasses corresponding to the coarse classification k0, P(·) is the prediction probability; N is the number of valid positions in the heat map, H xy is the predicted heat map value at position (x,y), is the true heat map value, S gt is the real size of the target, α and δ are focusing parameters, and β<0 is a scale-sensitive factor; L1 smooth is the smooth L1 loss function, m is the target index, M is the total number of targets, s m is the predicted size of the m-th target, is the actual size, γ is the size balance coefficient; m is the predicted center point offset of the mth target, is the actual offset, and η is the size correlation coefficient.
8. The underwater target intelligent recognition system based on multi-sensor fusion according to claim 1 is characterized in that: Underwater targets include: Artificial targets: including shipwrecks, underwater pipelines, underwater cables, artificial reefs and archaeological remains; Natural targets: including reefs, seabed topographic features, and seabed mineral resources; Underwater life: including large aquatic organisms, fish schools and coral reefs; The target recognition and positioning module predicts the direction and length of linear targets, the coverage and volume estimation of area targets, and the category and activity status of biological targets.
9. The underwater target intelligent recognition system based on multi-sensor fusion according to claim 1 is characterized in that: Intelligent analysis output modules include: A structured target information generation unit, used to generate a detection report containing underwater target category, location, size, quantity and feature description; An underwater map construction unit is used to combine the identified underwater target information with the positioning data to dynamically construct and update the underwater environment map; A navigation assistance unit, which is used to provide obstacle avoidance and path planning solutions for the underwater robot based on the detected underwater targets; The risk assessment unit is used to analyze the impact of identified underwater targets on navigation safety and trigger an alarm when high-risk targets are found.
10. The underwater target intelligent recognition system based on multi-sensor fusion according to claim 9 is characterized in that: The intelligent analysis output module also includes: Environmental change monitoring unit, used to analyze the changing trends of underwater targets and environment by comparing data collected at different times; The mission planning optimization unit is used to adjust the underwater robot's cruising strategy and sensor configuration according to the distribution of underwater targets and water quality conditions.
Citation Information
Patent Citations
Underwater target identification method for underwater intelligent robot
CN119559489A
Cited By
Underwater target identification method and system based on multi-source sensor data fusion
CN121167371A
Fiber bundle automatic segmentation and quantitative labeling method for white matter abnormality of Parkinson's disease
CN121169856A
Underwater structure defect dynamic grading method based on multi-modal perception fusion
CN121234286A
Underwater object identification method, device, equipment, medium and product
CN121392561A