A multi-modal fusion underwater target detection method based on reinforcement learning
Patent Information
- Application Number
- CN202611125427.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-28
- Publication Date
- 2026-09-22
- Estimated Expiration
- 2046-07-28
AI Technical Summary
这些传统方法往往无法根据水下环境的变化自动优化融合策略,从而导致在复杂环境中的性能不稳定
1)本发明采用强化学习技术,能够根据水下环境的光照强度、水质浑浊度和目标距离等实时变化的环境参数,动态调整可见光和点云数据的融合策略。与传统依赖固定规则的融合方法相比,本发明能够根据不同环境自动优化融合权重,从而在复杂水下环境中提供更高的目标检测精度和鲁棒性;
Smart Images

Figure CN122637185B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of underwater target detection technology, and in particular to a multimodal fusion underwater target detection method based on reinforcement learning. Background Technology
[0002] With the continuous development of underwater unmanned detection technology, unmanned underwater vehicles (UUVs) are increasingly being used in underwater environmental monitoring, resource exploration, marine scientific research, and emergency rescue. The underwater environment is typically complex and variable, with uncertainties such as currents, waves, obstacles, and noise, all of which can affect the detection accuracy of the detectors. Therefore, improving the detection capabilities of underwater unmanned detectors has become one of the current research hotspots.
[0003] Traditional underwater target detection methods typically rely on single-modal data, such as visible light images, point cloud data, or infrared images. However, the limitations of single-modal data significantly impact the accuracy and robustness of target detection in complex underwater environments. Specifically: visible light images are often unable to provide sufficiently clear target information due to insufficient underwater lighting and turbidity. Under low-light conditions, the quality of underwater visible light images is poor, making it difficult to accurately capture targets. Interference from factors such as water flow and temperature stratification leads to low point cloud spatial resolution, low signal-to-noise ratio, and blurred target details. Infrared images can provide reliable heat source information in low-light environments, but their application in underwater environments is also limited by the heat absorption properties of water.
[0004] Therefore, single-modal data often fails to provide complete target detection information, leading to a decrease in detection accuracy in complex underwater environments. To address this issue, multimodal data fusion has become an important technical approach. Existing single optical image detection generally employs methods such as data augmentation, network improvement, and image enhancement models, while single point cloud image recognition typically utilizes preprocessing, target segmentation, and recognition models for detection and identification.
[0005] Current multimodal data fusion methods mostly rely on fixed rules or preset models for data fusion, lacking the ability to adapt to changes in different underwater environments. These traditional methods often fail to automatically optimize the fusion strategy according to changes in the underwater environment, resulting in unstable performance in complex environments. Therefore, how to adjust the multimodal data fusion strategy according to dynamic changes in the environment in underwater target detection tasks is an urgent problem to be solved.
[0006] Therefore, there is an urgent need to propose a method that integrates information from multiple modalities, such as visible light and point cloud data, to effectively improve the accuracy and robustness of target detection. Summary of the Invention
[0007] The technical problem to be solved by the present invention is to provide a multimodal fusion underwater target detection method based on reinforcement learning, which addresses the shortcomings of the existing technology.
[0008] The technical solution adopted by this invention to solve its technical problem is: This invention provides a multimodal fusion underwater target detection method based on reinforcement learning, which includes the following steps: Step 1: Acquire visible light images and point cloud data of the underwater environment using a visible light camera and sonar, respectively. Perform feature extraction after preprocessing the visible light images and point cloud data. Step 2: Obtain the state parameters of the underwater environment, including light intensity, water turbidity and target distance. Use an improved deep Q network as a reinforcement learning control module to dynamically analyze the state parameters of the underwater environment and output actions based on the analysis results to generate dynamic fusion weights for visible light images and point cloud data. Step 3: Based on the dynamic fusion weights, the features of the visible light image and point cloud data are adaptively fused and detected using a pixel-by-pixel weighted method to generate detection results; Step 4: Calculate the reward based on the detection results, and update the policy of the improved deep Q network in the reinforcement learning control module based on the reward.
[0009] Furthermore, the method for preprocessing visible light images and point cloud data in step 1 of the present invention includes: For visible light images, Gaussian filtering is used to denoise the acquired images; By combining point cloud data with visible light images for coordinate mapping, each point in the 3D point cloud is mapped to the image coordinate system through projection, and each point in the point cloud corresponds to a pixel position in the 2D image.
[0010] Furthermore, the feature extraction method in step 1 of the present invention includes: Feature extraction is performed on the preprocessed visible light images and point cloud data; Feature extraction of visible light images is performed using a convolutional neural network, and the obtained image information includes: texture, edges, and color. Deep learning methods are used to extract geometric features from point cloud data, including: normals, reflection intensity, and spatial distribution. The extracted features are standardized or normalized.
[0011] Furthermore, the state parameters of the underwater environment in step 2 of the present invention include: State parameters of the underwater environment are dynamically extracted using sensors and image analysis to construct a multidimensional state vector:
[0012] in, To determine the light intensity, the mean grayscale histogram of the visible light image is calculated and mapped to the actual light intensity calibrated by the underwater photometer, and finally normalized to a value between 0 and 1. The turbidity of water is determined by analyzing the gradient amplitude variance of point cloud data in the near-sonar band, quantifying the intensity of scattered noise, and correlating it with the turbidity level calibrated in the laboratory. Finally, it is normalized to a continuous value from 0 to 1, representing clear water to extreme turbidity. The target distance is obtained through binocular visual parallax or sonar ranging. After normalization, it represents the relative distance between the target and the camera, with 0 being the closest and 1 being the farthest.
[0013] Furthermore, the improved deep Q-network in step 2 of this invention is specifically as follows: The improved deep Q-network has a built-in state embedding layer and uses a three-layer fully connected structure to map the multidimensional environment vector composed of light intensity, water turbidity and target distance into a high-dimensional implicit representation. In terms of architecture design, a dual-network mode is adopted, in which the main network and the target network are independent of each other and are periodically synchronized. At the same time, a priority experience replay mechanism is introduced, which guides the network to perform non-uniform sampling by calculating the learning error priority of the samples. Combined with an adaptive strategy that dynamically decays the initial exploration rate as the training process progresses, the network can search the environment during training to obtain the optimal fusion path and achieve stable convergence, thereby realizing intelligent adaptive fusion of underwater multimodal data.
[0014] Furthermore, the method for generating dynamic fusion weights of visible light images and point cloud data based on the analysis results in step 2 of the present invention includes: The input information of the improved deep Q network is Simultaneously, dynamic fusion weights from the previous time step are introduced to maintain the historical continuity of decision-making, and combined with the signal-to-noise ratios of each mode estimated based on the feature map gradient response variance, to jointly characterize the complexity of the current detection environment; the output is the dynamic fusion weights: visible light weights. Sonar weight .
[0015] Furthermore, the adaptive fusion method in step 3 of the present invention includes: Visible light image feature maps are processed using a pixel-by-pixel weighting method. and point cloud data feature map To perform the fusion, for each pixel, the two weighted feature maps will be added together:
[0016] in, It is a feature map extracted from a visible light image. It is a feature map extracted from point cloud data. Represents visible light weight , representing the feature map of a visible light image Contribution in the final fused feature map; Represents sonar weight , representing the feature map of point cloud data Contribution in the final fused feature map; obtained This is the feature map after weighted fusion; The fused feature map is linearly normalized so that all pixel values fall within a standard range:
[0017] in, It is the fused feature map after linear normalization.
[0018] Furthermore, the method for generating the detection result in step 3 of the present invention includes: The normalized fused feature map is input into the YOLO network, and the output is as follows: Assumption Indicates the first The first grid cell and the first Predicted information for each bounding box. Yes, that's fine. This is a column, and the output for each bounding box is:
[0019] in, It is the offset of the bounding box center relative to the grid cell. These are the width and height of the bounding box. It is the confidence score of the bounding box, representing the product of the probability that the bounding box contains the target and the intersection-union ratio (IU / IU) of the bounding box and the target. Each category The probability indicates that the bounding box belongs to the category. The probability of.
[0020] Furthermore, step 4 of the present invention includes: With target detection accuracy as the optimization guide, a multi-dimensional reward system is designed in conjunction with the stability of the fusion strategy: Target localization accuracy and location reward: The Intersection over Union (IoU) is used to measure the overlap between the predicted bounding box and the ground truth target bounding box; the larger the IoU, the more accurate the target localization, and the higher the reward; the location reward is:
[0021] in, It is the bounding box predicted by the detection results. It is the actual bounding box; the closer the IoU value is to 1, the greater the reward value. Different reward weights are set for low light and high turbidity conditions; When the light intensity At the same time, encourage sonar weighting:
[0022] When water turbidity At the same time, visible light weighting is encouraged:
[0023] The total reward function is:
[0024] in, This represents a reward adjustment factor, used to balance the magnitude of positional rewards and environmental constraint rewards; This represents the visible light fusion weights output by the reinforcement learning control module at the current moment, based on the total reward function. Update and improve the parameters of the deep Q-network.
[0025] This invention provides a multimodal fusion underwater target detection system based on reinforcement learning, comprising: Memory, used to store executable computer programs; The processor, when executing an executable computer program stored in memory, implements the aforementioned reinforcement learning-based multimodal fusion underwater target detection method.
[0026] The beneficial effects of this invention are: 1) This invention employs reinforcement learning technology to dynamically adjust the fusion strategy of visible light and point cloud data based on real-time changing environmental parameters such as underwater light intensity, water turbidity, and target distance. Compared to traditional fusion methods that rely on fixed rules, this invention can automatically optimize the fusion weights according to different environments, thereby providing higher target detection accuracy and robustness in complex underwater environments. 2) By combining the advantages of visible light images and point cloud data, and utilizing a reinforcement learning control module to adaptively adjust the fusion weights, this invention effectively enhances target detection capabilities in complex environments such as low light and turbid water. Compared to single-modal data processing, this invention significantly improves the environmental adaptability and detection performance of underwater target detection systems. Attached Figure Description
[0027] The present invention will be further described below with reference to the accompanying drawings and embodiments. In the accompanying drawings: Figure 1This is a flowchart of the multimodal fusion underwater target detection method based on reinforcement learning according to an embodiment of the present invention; Figure 2 This is a block diagram of a multimodal fusion underwater target detection method based on reinforcement learning according to an embodiment of the present invention; Figure 3 This is a flowchart of the data preprocessing process according to an embodiment of the present invention. Detailed Implementation
[0028] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.
[0029] like Figure 1 The diagram shown is a flowchart of a multimodal fusion underwater target detection method based on reinforcement learning, according to an embodiment of the present invention. This embodiment proposes an adaptive fusion method based on reinforcement learning, which can dynamically adjust the fusion strategy of visible light and point cloud data, automatically adapting to different underwater environments, thereby significantly improving the accuracy and environmental adaptability of target detection. Specifically, it includes: ① Acquire visible light images and point cloud data of the underwater environment using a visible light camera and sonar, respectively; ②Use reinforcement learning control modules to dynamically analyze the state parameters of the underwater environment, including light intensity, water turbidity, and target distance; ③ Generate dynamic fusion weights for visible light images and point cloud data based on the analysis results, and adaptively fuse and detect the features of visible light images and point cloud data based on the fusion weights to generate detection results; ④ The adjusted samples are input into the reinforcement learning optimization algorithm for policy updates, in order to improve the learning efficiency and decision-making performance of the underwater detector.
[0030] like Figure 2 As shown, in another preferred embodiment of the present invention, the underwater target detection method based on reinforcement learning and visible light and point cloud data fusion includes the following steps: Step 1: Acquire visible light images and point cloud data of the underwater environment using a visible light camera and sonar, respectively.
[0031] The main task of this invention is to dynamically adjust the fusion weights of visible light and point cloud data through a pre-trained reinforcement learning model in order to obtain better target detection results. First, data of different modalities should be acquired through visible light cameras and sonar to provide raw data for subsequent fusion and target detection.
[0032] A high-resolution underwater visible light camera was selected, capable of providing clear images within a certain underwater depth range. Simultaneously, a suitable sonar was chosen to effectively acquire point cloud data under low-light and complex underwater environmental conditions.
[0033] like Figure 3 As shown, after acquiring the image, Gaussian filtering is used to denoise the acquired image to reduce the impact of underwater noise (such as water waves, plankton, etc.) on image quality. Each point in the 3D point cloud is mapped to the image coordinate system through projection, and each point in the point cloud corresponds to a pixel position in the 2D image, thus obtaining the position and other features of each point on the image.
[0034] Feature extraction is performed on data from visible light and point cloud data for subsequent fusion processing. During extraction, a convolutional neural network (CNN) is used to extract features from the image, obtaining information such as texture, edges, and color. Deep learning methods (such as PointNet) are used to extract geometric features of the point cloud (such as normals, reflection intensity, and spatial distribution). Since the feature distributions of point clouds and images can differ significantly, the extracted features typically need to be standardized or normalized to ensure that the data from both modalities can be processed at the same scale.
[0035] Step 2: Use the reinforcement learning control module to dynamically analyze the state parameters of the underwater environment, including light intensity, water turbidity and target distance, to obtain fused parameters.
[0036] An improved Deep Q-Network (DQN) is employed as the core framework for reinforcement learning, enhancing learning stability and convergence efficiency in complex underwater environments through several targeted optimizations. First, the network incorporates a state embedding layer, utilizing a three-layer fully connected structure to map a multi-dimensional environment vector composed of light intensity, water turbidity, and target distance into a high-dimensional latent representation, effectively addressing the mismatch between the original state space and the network input dimension. In terms of architecture design, a dual-network mode is adopted, where the main network and the target network are independent and periodically synchronized, significantly reducing target value oscillations during training. Simultaneously, a priority experience replay mechanism is introduced, guiding the network to perform non-uniform sampling by calculating the learning error priority of samples, enabling it to focus more on and learn more challenging key samples. Finally, an adaptive strategy that dynamically decays the initial exploration rate as training progresses ensures that the model can fully search the environment to obtain the optimal fusion path in the early stages of training and achieve stable convergence in the later stages, thereby realizing intelligent adaptive fusion of underwater multimodal data. This addresses the instability problem of traditional Q-learning in high-dimensional state spaces.
[0037] The network input is The improved deep Q-network achieves dynamic weight decision-making by integrating multi-source information. The input layer of this network not only receives a real-time environmental state vector including light intensity, water turbidity, and target distance, but also introduces the fusion weights from the previous time step to maintain the historical continuity of the decision. Combined with the signal-to-noise ratio of each mode estimated based on the feature map gradient response variance, it jointly characterizes the complexity of the current detection environment.
[0038] In the specific processing flow, the system first concatenates and encodes the above input information to form a unified state tensor input to the main network. The network then calculates and outputs a value estimate for a preset discrete weighted action set. This action space is formed by discretizing the visible light weights within the range of 0 to 1, resulting in multiple candidate actions. Subsequently, based on a preset exploration and utilization strategy, the system balances global search and optimal value utilization, selecting the optimal action index for the current moment. Finally, this action is directly mapped to the dynamic fusion weights of the visible light image, and the corresponding weights of the point cloud data are automatically determined, thereby achieving intelligent real-time response to changes in complex underwater environments. The final output is the dynamic fusion weights: visible light weights. Sonar weight Specifically: (1) State-space design This invention dynamically extracts the following environmental features through sensors and image analysis to construct a multidimensional state vector, with the state space set as follows:
[0039] Illumination intensity is a core indicator of visible light image quality and directly affects the effectiveness of visible light modes. By calculating the mean grayscale histogram of the visible light image (range 0-255) and establishing a mapping relationship with the actual illumination intensity (lux) calibrated by the underwater photometer, it is finally normalized to a value between 0 and 1.
[0040] Water turbidity reflects the scattering effect of water on light and is a key parameter determining the ability of visible light to penetrate. By analyzing the gradient amplitude variance of point cloud data in the near-sonar band (700-1000nm), the intensity of scattering noise is quantified and correlated with the turbidity level (NTU value) calibrated in the laboratory, and finally normalized to a continuous value from 0 (clear water) to 1 (extreme turbidity).
[0041] The target distance is obtained through binocular visual parallax or sonar ranging, and after normalization, it represents the relative distance between the target and the camera (0 is the closest and 1 is the farthest).
[0042] (2) Action space design The reinforcement learning module of this invention dynamically adjusts the fusion of visible light and point cloud data by selecting dynamic fusion weights for visible light and point cloud data, and the action space is set as follows: Action space: (Visible light weight), sonar weight is ; Simultaneously perform discretization processing, Divide into 10 discrete actions (step size 0.1), for example: ; Step 3: Generate dynamic fusion weights for visible light images and point cloud data based on the analysis results, and adaptively fuse and detect the features of visible light images and point cloud data based on the fusion weights to generate detection results.
[0043] After obtaining the fusion weights, this invention uses a pixel-by-pixel weighting method to fuse the feature maps of the two modalities (visible light image and point cloud data). Specifically: Feature maps extracted from visible light images using convolutional neural networks (CNNs) contain visual information about the target (such as texture and edges). Feature maps of point cloud data, such as normals, reflection intensity, and spatial distribution, are extracted using PointNet. For each pixel, the two weighted feature maps will be added together:
[0044]
[0045] in, It is a feature map extracted from a visible light image. It is a feature map extracted from point cloud data. It is a weighting coefficient representing the feature map of a visible light image. Contribution in the final fused feature map. It is a weighting coefficient representing the feature map of point cloud data. Contribution in the final fused feature map.
[0046] Received This is the weighted fusion feature map, which contains information from visible light and point cloud data, and can adapt to the target detection needs in different underwater environments.
[0047] The fused feature maps may have inconsistent numerical ranges due to weighting operations, affecting the processing performance of subsequent networks. Therefore, the fused feature maps can be normalized: linear normalization is performed to ensure that all pixel values fall within a standard range (e.g., between [0, 1]).
[0048] This helps object detection networks better learn and understand image features, avoiding the negative impact of numerical range on subsequent calculations.
[0049] The normalized result is input into the YOLO network, and the output is as follows: Assumption Indicates the first grid cells ( Yes, that's fine. (is a list) and the first The predicted information for each bounding box. The output for each bounding box is:
[0050] in: It is the offset (normalized) of the bounding box center relative to the grid cell. These are the width and height of the bounding box, which are usually normalized relative to the entire image. It is the confidence score of the bounding box, which represents the product of the probability that the bounding box contains the target and the intersection-union ratio (IOU) between the box and the real target. Each category The probability indicates that the bounding box belongs to the category. The probability of.
[0051] Step 4: Input the result samples and rewards into the reinforcement learning optimization algorithm for policy updates, so as to improve the learning efficiency and decision-making performance of the underwater detector.
[0052] This invention prioritizes target detection accuracy and incorporates a multi-dimensional reward design based on the stability of the fusion strategy. Target localization accuracy (location reward): The Intersection over Union (IoU) ratio is used to measure the overlap between the predicted bounding box and the ground truth target bounding box. A higher IoU indicates more accurate target localization, and therefore a higher reward is given.
[0053]
[0054] in, It is the bounding box predicted by the model. It is the actual bounding box. The closer the IoU value is to 1, the greater the reward value.
[0055] Furthermore, to simplify the training process, different reward weights were set for different lighting conditions and high turbidity conditions.
[0056] When the light intensity At the same time, encourage sonar weighting:
[0057] When water turbidity At the same time, visible light weighting is encouraged:
[0058] The total reward function is:
[0059] This represents a reward adjustment factor, used to balance the magnitude of positional rewards and environmental constraint rewards; This represents the visible light fusion weight action value output by the reinforcement learning module at the current moment; based on the total reward function. The target value is calculated using a target Q-network. By minimizing the mean squared error loss function between the target value and the current evaluation network output value, the weight parameters of the deep Q-network are updated using the stochastic gradient descent algorithm. Meanwhile, the parameters of the evaluation network are periodically synchronized to the target network, and a small batch of data is randomly drawn from the historical sample pool for training through an experience replay mechanism to break the correlation between samples, ultimately achieving smooth convergence and optimization of the network strategy.
[0060] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0061] It should be understood that those skilled in the art can make improvements or modifications based on the above description, and all such improvements and modifications should fall within the protection scope of the appended claims.
Claims
1. A multimodal fusion underwater target detection method based on reinforcement learning, characterized in that, The method includes the following steps: Step 1: Acquire visible light images and point cloud data of the underwater environment using a visible light camera and sonar, respectively. Perform feature extraction after preprocessing the visible light images and point cloud data. Step 2: Obtain the state parameters of the underwater environment, including light intensity, water turbidity and target distance. Use an improved deep Q network as a reinforcement learning control module to dynamically analyze the state parameters of the underwater environment and output actions based on the analysis results to generate dynamic fusion weights for visible light images and point cloud data. Step 3: Based on the dynamic fusion weights, the features of the visible light image and point cloud data are adaptively fused and detected using a pixel-by-pixel weighted method to generate detection results; Step 4: Calculate the reward based on the detection results, and update the policy of the improved deep Q-network in the reinforcement learning control module based on the reward; The underwater environment state parameters in step 2 include: State parameters of the underwater environment are dynamically extracted using sensors and image analysis to construct a multidimensional state vector: in, To determine the light intensity, the mean grayscale histogram of the visible light image is calculated and mapped to the actual light intensity calibrated by the underwater photometer, and finally normalized to a value between 0 and 1. The turbidity of water is determined by analyzing the gradient amplitude variance of point cloud data in the near-sonar band, quantifying the intensity of scattered noise, and correlating it with the turbidity level calibrated in the laboratory. Finally, it is normalized to a continuous value from 0 to 1, representing clear water to extreme turbidity. The target distance is obtained through binocular visual parallax or sonar ranging. After normalization, it represents the relative distance between the target and the camera, with 0 being the closest and 1 being the farthest. The improved deep Q-network in step 2 is specifically as follows: The improved deep Q-network has a built-in state embedding layer and uses a three-layer fully connected structure to map the multidimensional environment vector composed of light intensity, water turbidity and target distance into a high-dimensional implicit representation. In terms of architecture design, a dual-network mode is adopted, in which the main network and the target network are independent of each other and are periodically synchronized. At the same time, a priority experience replay mechanism is introduced, which guides the network to perform non-uniform sampling by calculating the learning error priority of the samples. Combined with an adaptive strategy that dynamically decays the initial exploration rate as the training process progresses, the network can search the environment during training to obtain the optimal fusion path and achieve stable convergence, thereby realizing intelligent adaptive fusion of underwater multimodal data. The method for generating dynamic fusion weights for visible light images and point cloud data based on the analysis results in step 2 includes: The input information of the improved deep Q network is Simultaneously, dynamic fusion weights from the previous time step are introduced to maintain the historical continuity of decision-making, and combined with the signal-to-noise ratios of each mode estimated based on the feature map gradient response variance, to jointly characterize the complexity of the current detection environment; the output is the dynamic fusion weights: visible light weights. Sonar weight ; The adaptive fusion method in step 3 includes: Visible light image feature maps are processed using a pixel-by-pixel weighting method. and point cloud data feature map To perform the fusion, for each pixel, the two weighted feature maps will be added together: in, It is a feature map extracted from a visible light image. It is a feature map extracted from point cloud data. Represents visible light weight , representing the feature map of a visible light image Contribution in the final fused feature map; Represents sonar weight , representing the feature map of point cloud data Contribution in the final fused feature map; obtained This is the feature map after weighted fusion; The fused feature map is linearly normalized so that all pixel values fall within a standard range: in, It is a fused feature map after linear normalization; The method for generating the detection results in step 3 includes: The normalized fused feature map is input into the YOLO network, and the output is as follows: Assumption Indicates the first The first grid cell and the first Predicted information for each bounding box. Yes, This is a column, and the output for each bounding box is: in, It is the offset of the bounding box center relative to the grid cell. These are the width and height of the bounding box. It is the confidence score of the bounding box, representing the product of the probability that the bounding box contains the target and the intersection-union ratio (IUU) of the bounding box and the target. Each category The probability indicates that the bounding box belongs to the category. The probability of.
2. The multimodal fusion underwater target detection method based on reinforcement learning according to claim 1, characterized in that, The method for preprocessing visible light images and point cloud data in step 1 includes: For visible light images, Gaussian filtering is used to denoise the acquired images; By combining point cloud data with visible light images for coordinate mapping, each point in the 3D point cloud is mapped to the image coordinate system through projection, and each point in the point cloud corresponds to a pixel position in the 2D image.
3. The multimodal fusion underwater target detection method based on reinforcement learning according to claim 2, characterized in that, The method for feature extraction in step 1 includes: Feature extraction is performed on the preprocessed visible light images and point cloud data; Feature extraction of visible light images is performed using a convolutional neural network, and the obtained image information includes: texture, edges, and color. Deep learning methods are used to extract geometric features from point cloud data, including: normals, reflection intensity, and spatial distribution. The extracted features are standardized or normalized.
4. The multimodal fusion underwater target detection method based on reinforcement learning according to claim 1, characterized in that, The method in step 4 includes: With target detection accuracy as the optimization guide, a multi-dimensional reward system is designed in conjunction with the stability of the fusion strategy: Target localization accuracy and location reward: The Intersection over Union (IoU) is used to measure the overlap between the predicted bounding box and the ground truth target bounding box; the larger the IoU, the more accurate the target localization, and the higher the reward; the location reward is: in, It is the bounding box predicted by the detection results. It is the actual bounding box; the closer the IoU value is to 1, the greater the reward value. Different reward weights are set for low light and high turbidity conditions; When the light intensity At the same time, encourage sonar weighting: When water turbidity At the same time, visible light weighting is encouraged: The total reward function is: in, This represents a reward adjustment factor, used to balance the magnitude of positional rewards and environmental constraint rewards; This represents the visible light weights output by the reinforcement learning control module at the current moment, based on the total reward function. Update and improve the parameters of the deep Q-network.
5. A multimodal fusion underwater target detection system based on reinforcement learning, characterized in that, include: Memory, used to store executable computer programs; A processor, when executing an executable computer program stored in a memory, implements the reinforcement learning-based multimodal fusion underwater target detection method as described in any one of claims 1 to 4.
Citation Information
Patent Citations
Semi-supervised target detection method for visible light-infrared multi-mode fusion scene
CN121213883A
Underwater image target detection method and system based on multi-modal fusion
CN121482380A