A multi-sensor-based underwater target detection method
By using multi-sensor components and data fusion technology, the problems of weak detail capture capability and poor robustness in traditional marine target detection have been solved, achieving a more efficient underwater target detection effect.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- NAVAL UNIV OF ENG PLA
- Filing Date
- 2026-03-05
- Publication Date
- 2026-06-02
AI Technical Summary
Traditional single-sensor-based marine target detection suffers from weak detail capture capabilities, susceptibility to environmental interference, and poor robustness. Multi-sensor fusion schemes fail to fully exploit the correlations between data, making it difficult to improve the effectiveness of detection results.
The system employs a multi-sensor assembly including a visual sensor, a sonar sensor, a magnetic field sensor, an electric field sensor, and a temperature sensor. Through data preprocessing, spatiotemporal matching, and multi-sensor data fusion, it utilizes a dual-branch network and an attention fusion network to extract and weightedly fuse feature vectors, thereby achieving target recognition and localization.
It improves the effectiveness and accuracy of underwater target detection data, enhances robustness to environmental interference, and improves the detection effect of multi-sensor data fusion.
Smart Images

Figure CN122131419A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of underwater target detection technology, and in particular relates to an underwater target detection method based on multiple sensors. Background Technology
[0002] The ocean is a precious treasure of the Earth, and its development is crucial to human survival and development. With increasing ocean exploitation, higher demands are being placed on marine environmental detection. Traditional single-sensor-based marine target detection suffers from weak detail capture capabilities, susceptibility to environmental interference, and poor robustness. To address the shortcomings of single sensors, existing technologies have proposed multi-sensor simultaneous detection schemes. However, these often employ simple feature stitching or weighting, failing to fully explore the correlations between multi-sensor data and failing to effectively combine the data, thus hindering the improvement of the effectiveness of underwater target detection results through multi-sensor data fusion. Summary of the Invention
[0003] The purpose of this invention is to overcome the shortcomings of the prior art and provide a multi-sensor-based underwater target detection method to further improve the effectiveness of underwater target detection data fused from multiple sensors.
[0004] To achieve the above objectives, the present invention adopts the following technical solution.
[0005] A multi-sensor-based underwater target detection method includes the following steps:
[0006] Step 1, Multi-sensor data acquisition, refers to the acquisition of raw data of underwater targets or detection areas based on multi-sensor components; the multi-sensor components include visual sensors, sonar sensors, magnetic field sensors, electric field sensors, and temperature sensors;
[0007] Step 2: Spatiotemporal matching of sensor data, specifically including steps b1 to b3;
[0008] b1 Data preprocessing: Preprocesses the image data acquired by the vision sensor and the coordinates collected by the sonar, performs target detection through the image recognition model, and marks the area of the target object to be detected.
[0009] b2 Timing Matching: Adds timestamps to the data from the magnetic field sensor, electric field sensor, and temperature sensor, based on 3... Outliers are removed in principle, and missing values are filled in using linear interpolation before timing alignment is performed. At the same time, signals with different sampling rates can be resampled to ensure timing consistency.
[0010] b3 Spatial matching, which is mainly based on the position of the target's true coordinates in the corresponding data coordinates; specifically including:
[0011] Establish a sonar coordinate system and a visual coordinate system. Based on the rotation and translation matrices of the visual coordinate system and the environmental coordinate system, obtain the position expression of the underwater target in the visual coordinate system. According to the imaging principle of the visual sensor, first transform to obtain the coordinate data on the imaging plane of the visual sensor. Establish an image coordinate system and establish the transformation relationship between the pixel coordinates on the image acquired by the visual sensor and the environmental coordinates.
[0012] Based on the spatial matching results, the target data obtained by sonar is mapped onto the image obtained by the visual sensor. The correlation algorithm is used to determine the matching degree between the image recognition target and the sonar detection target, and the joint target detection data is output.
[0013] Step 3: Multi-sensor data fusion detection, including the construction of a dual-branch network, comprising a first branch network, a second branch network, and an attention fusion network;
[0014] The first branch network takes image and sonar information as input and extracts visual features, while the second branch network takes electric field, magnetic field and temperature field data as input and extracts distance and direction related information. The attention fusion network uses the attention module to weight and fuse the features of the two branches to obtain a fused feature vector. The feature vector obtained by fusing multi-sensor data is then input into the underwater target detection and recognition model to complete target recognition and target localization.
[0015] In a further improvement or preferred embodiment of the aforementioned underwater target detection method based on multiple sensors, b1 specifically includes: preprocessing the image data acquired by the visual sensor by means filtering to reduce noise, and performing size adjustment and format conversion as described in the previous steps; performing threshold judgment on the data acquired by the sonar to remove invalid false data; performing image enhancement preprocessing on the image frames acquired by the camera, and then performing target detection through an image recognition model to mark the required target area.
[0016] In a further improvement or preferred embodiment of the aforementioned underwater target detection method based on multiple sensors, the b2 timing matching specifically includes: adding timestamps to all analog voltage data, removing outliers according to the 3 principle, filling in missing values using a linear interpolation method, and then aligning the data timing; at the same time, signals with different sampling rates can be resampled to ensure timing consistency.
[0017] In a further improvement or preferred embodiment of the aforementioned multi-sensor-based underwater target detection method, step b3, spatial matching, specifically includes:
[0018] First, establish the location of the sonar sensor based on the position of the sonar sensor. Detecting the direction of movement of the platform The direction pointing towards the center of the earth. perpendicular to To the left Establish a sonar coordinate system ,by The point where the direction intersects with the sea level is the origin coordinate system. , to pass through the origin Furthermore, establish an environmental coordinate system on each plane that is parallel or perpendicular to the sonar coordinate system. Then, for the radial distance in the sonar sensor data, it is... yaw angle is The coordinates of the target p in the environment coordinate system are ;
[0019] in This represents the distance of the sonar sensor from the sea level.
[0020] Secondly, the location of the vision sensor is taken as the origin of the coordinate system. Establish a visual coordinate system with all coordinate axes aligned with the sonar coordinate system. Then, based on the rotation matrix between the visual coordinate system and the environment coordinate system and translation matrix This allows us to obtain the position representation of the underwater target p in the visual coordinate system. ;
[0021] Then, based on the imaging principle of the vision sensor, the coordinate data on the imaging plane of the vision sensor are first converted to obtain the coordinate data. ,in Let be the distance between the imaging plane and the optical center of the vision sensor; establish an image coordinate system with the top left corner of the image acquired by the vision sensor as the origin, and the horizontal and vertical axes as the x and y coordinates, respectively. Then, the point on the imaging plane coordinate system... Coordinate data in the image coordinate system ;in This refers to the coordinates of a point in an image. The size of a single pixel in the image;
[0022] Based on the transformation relationships established in the preceding steps, the transformation relationship between pixel coordinates and environmental coordinates in the image acquired by the visual sensor can be simultaneously established as follows: ;
[0023] Finally, based on the spatial matching results, the target data obtained by the sonar is mapped onto the image obtained by the visual sensor to generate the detection box corresponding to the sonar. The matching degree of the detection results of the image recognition target and the sonar detection target is judged by the association algorithm. When the matching degree of the target recognition result meets the accuracy requirements, the joint target detection data obtained by the sonar sensor and the visual sensor is output. If the accuracy requirements are not met, the target is eliminated by decision or the best matching result is output.
[0024] In a further improvement or preferred embodiment of the aforementioned underwater target detection method based on multiple sensors, in step b3, when the matching degree is much lower than the normal range, the validity of the sonar sensor data is first checked. If the data acquired by the sonar sensor is valid, the output of the joint target detection data is corrected based on the sonar sensor data. If the validity of the sonar sensor data cannot be determined, the target data acquired by the sonar sensor and the visual sensor are cross-verified to determine the valid data. When the matching degree is within the normal range, the output of the joint target detection data is corrected based on the distance data acquired by the sonar sensor and the position and category data acquired by the visual sensor, respectively.
[0025] In a further improvement or preferred embodiment of the aforementioned underwater target detection method based on multiple sensors, step b3 involves offset correction when using a morphing lens with a large field of view. The offset correction includes edge magnification offset correction and mirror non-parallelism offset correction. The relevant parameters for offset correction are determined by imaging calibration of the visual sensor lens.
[0026] In a further improvement or preferred embodiment of the aforementioned underwater target detection method based on multiple sensors, step b3 establishes a mechanism for the synchronous activation and acquisition of sonar sensors and visual sensors during the target detection process.
[0027] A further improvement or preferred embodiment of the aforementioned multi-sensor-based underwater target detection method, step three, multi-sensor data fusion detection, specifically includes:
[0028] Construct a dual-branch network, including a first branch network, a second branch network, and an attention fusion network;
[0029] The first branch network takes image and sonar information as input and extracts visual features, while the second branch network takes electric field, magnetic field and temperature field data as input and extracts distance and direction related information. The attention fusion network uses the attention module to weight and fuse the features of the two branches to obtain a fused feature vector.
[0030] The first branch network includes a two-layer convolutional pooling unit, a global pooling unit, a two-layer fully connected unit, and a fully connected output unit. Based on the joint target detection data obtained in the preceding steps, the first branch network extracts the image features acquired by the visual sensor, performs image dimensionality reduction using the two-layer convolutional pooling unit, compresses the low-dimensional image into a one-dimensional feature vector using the global pooling unit, extracts sonar coordinate data, and maps the original coordinate values into a one-dimensional feature vector using the two-layer fully connected unit. The two one-dimensional vectors are concatenated and the output of the fully connected layer is used to obtain the output features of the first branch network.
[0031] The second branch network includes a one-dimensional convolutional pooling unit, an average pooling unit, and a fully connected output unit. The second branch network extracts temporal feature data from the three-dimensional time-series data related to the magnetic field signal, electric field signal, and temperature field signal through the one-dimensional convolutional pooling unit, and outputs it as a one-dimensional feature vector by global average pooling. Then, the fully connected output unit maps it to a feature vector with the same dimension as the output of the first branch network.
[0032] The attention fusion network includes an average pooling unit, a two-layer fully connected unit, and a weighted fusion unit. The attention fusion network uses the average pooling unit to average the output features of the two branch networks, uses the two-layer fully connected unit to output attention weights, and uses the attention weights to perform weighted fusion of the feature outputs of the two branch networks. The feature vector obtained by fusing multi-sensor data is input into the underwater target detection and recognition model to complete target recognition and target localization, providing a basis for subsequent decision-making.
[0033] In a further improved or preferred embodiment of the aforementioned multi-sensor-based underwater target detection method, the visual sensor is used to acquire image data of the underwater target or detection area; the sonar sensor is used to acquire sonar signals related to the distance and position information of the underwater target or detection area; the magnetic field sensor is used to acquire magnetic field signals of the underwater target or detection area; the electric field sensor is used to acquire electric field signals of the underwater target or detection area; and the temperature sensor is used to acquire temperature field signals of the underwater target or detection area. The sonar signal, magnetic field signal, electric field signal, and temperature field signal are analog voltage signals. Attached Figure Description
[0034] Figure 1 This is a schematic diagram of the preprocessing and timing synchronization of raw data from multiple sensors in this application;
[0035] Figure 2 This is a schematic diagram of a two-branch network process. Detailed Implementation
[0036] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention. Furthermore, the technical features involved in the various embodiments of this invention described below can be combined with each other as long as they do not conflict with each other.
[0037] This invention relates to a multi-sensor-based underwater target detection method, which is mainly used to provide a way to better acquire underwater target feature data hidden in multi-sensor data while ensuring consistency of spatiotemporal attributes, so as to provide richer and more complete underwater target detection methods and improve the effect of underwater target detection.
[0038] The multi-sensor data in this application is based on the acquisition of raw data of underwater targets or detection areas using a multi-sensor assembly. Based on current technology, the multi-sensor assembly includes: a visual sensor for acquiring image data of the underwater target or detection area; a sonar sensor for acquiring sonar signals related to the distance and position information of the underwater target or detection area; a magnetic field sensor for acquiring magnetic field signals of the underwater target or detection area; an electric field sensor for acquiring electric field signals of the underwater target or detection area; and a temperature sensor for acquiring temperature field signals of the underwater target or detection area.
[0039] In the specific implementation process, the sonar signal, magnetic field signal, electric field signal, and temperature field signal are analog voltage signals;
[0040] In practical implementation, the visual sensor in a multi-sensor assembly is affected by real-time underwater environment and lighting conditions, resulting in significant variations in its effective image acquisition range, which is generally lower than that of other sensor devices such as sonar sensors. Furthermore, the time and efficiency of different sensor elements in acquiring relevant signals vary. Therefore, ensuring the spatiotemporal consistency of sensor data is crucial.
[0041] The sensor data output in the form of analog voltage signals can be accurately unified through electronic components or methods such as timers. Therefore, the spatiotemporal consistency between the analog voltage signals and the data acquired by the vision sensor should be primarily ensured; specifically including:
[0042] b1 Data Preprocessing
[0043] like Figure 1 As shown, for image data acquired by a visual sensor, preprocessing is performed such as mean filtering to reduce noise, as well as the size adjustment and format conversion steps mentioned above; necessary processing such as threshold judgment is performed on the data acquired by sonar to remove invalid false data; image frames acquired by a camera are first preprocessed with image enhancement, and then target detection is performed through an image recognition model to mark the required target area;
[0044] b2 Timing Matching: Analog voltage signals have high-speed time sensitivity and are easy to synchronize. By adding timestamps to all analog voltage data, according to 3... Outliers are removed in principle, and missing values are filled in using linear interpolation, which makes it easy to align data timing. At the same time, signals with different sampling rates can be resampled to ensure timing consistency.
[0045] Among them, the data content of the sonar sensor, which is the main method of collecting data on the target or regional terrain structure, has the highest intuitive matching degree with the data collected by the visual sensor, and has complete positioning coordinates and position and distance algorithms. Therefore, this application uses the data obtained by the sonar sensor to perform spatial matching on the data collected by the visual sensor to achieve correction and unification. The following is a detailed explanation.
[0046] b3 Spatial matching: Spatial matching is mainly based on the position of the target's real coordinates in the corresponding data coordinates.
[0047] First, establish the location of the sonar sensor based on the position of the sonar sensor. Detecting the direction of movement of the platform The direction pointing towards the center of the earth. perpendicular to To the left Establish a sonar coordinate system ,by The point where the direction intersects with the sea level is the origin coordinate system. , to pass through the origin Furthermore, establish an environmental coordinate system on each plane that is parallel or perpendicular to the sonar coordinate system. Then, for the radial distance in the sonar sensor data, it is... yaw angle is The coordinates of the target p in the environment coordinate system are ;
[0048] in This represents the distance of the sonar sensor from the sea level.
[0049] Secondly, the location of the vision sensor is taken as the origin of the coordinate system. Establish a visual coordinate system with all coordinate axes aligned with the sonar coordinate system. Then, based on the rotation matrix between the visual coordinate system and the environment coordinate system and translation matrix This allows us to obtain the position representation of the underwater target p in the visual coordinate system. ;
[0050] Since the image pixel data in the visual sensor is generated starting from the top-left pixel and following the horizontal and vertical coordinates of the image, the data acquisition method also adopts this order to improve the data input efficiency of the visual sensor. To ensure the spatiotemporal consistency of data acquisition, mapping based on the transformation relationship between the camera coordinate system and the image coordinate system is required. Specifically:
[0051] Then, based on the imaging principle of the vision sensor, the coordinate data on the imaging plane of the vision sensor are first converted to obtain the coordinate data. ,in Let be the distance between the imaging plane and the optical center of the vision sensor; establish an image coordinate system with the top left corner of the image acquired by the vision sensor as the origin, and the horizontal and vertical axes as the x and y coordinates, respectively. Then, the point on the imaging plane coordinate system... Coordinate data in the image coordinate system ;in This refers to the coordinates of a point in an image. The size of a single pixel in the image;
[0052] Based on the transformation relationships established in the preceding steps, the transformation relationship between pixel coordinates and environmental coordinates in the image acquired by the visual sensor can be simultaneously established as follows: ;
[0053] Finally, based on the spatial matching results, the target data obtained by the sonar is mapped onto the image obtained by the visual sensor to generate the detection box corresponding to the sonar. The matching degree of the detection results of the image recognition target and the sonar detection target is judged by the association algorithm. When the matching degree of the target recognition result meets the accuracy requirements, the joint target detection data obtained by the sonar sensor and the visual sensor is output. If the accuracy requirements are not met, the target is eliminated by decision or the best matching result is output.
[0054] In particular, during the matching degree and progress analysis process described above, the accuracy and effectiveness of the output data can be further optimized by combining the data accuracy and effectiveness of the visual sensor and sonar sensor themselves. Specifically:
[0055] When the matching degree is much lower than the normal range, it usually means that there is an error or interference in the data from one of the sensors. Since the stability and effectiveness of visual sensors are usually much higher than those of sonar sensors, the effectiveness of the sonar sensor data is checked first. If the data acquired by the sonar sensor is valid, the output of the joint target detection data will be corrected based on the sonar sensor data. If the effectiveness of the sonar sensor data cannot be determined, the target data acquired by the sonar sensor and the visual sensor will be cross-verified to determine the valid data.
[0056] When the matching degree is within the normal range, the output of the joint target detection data is corrected based on the distance data obtained by the sonar sensor and the position and category data obtained by the visual sensor, respectively.
[0057] In particular, due to energy management needs and the need to optimize prior data on target location, sonar or visual sensors may not remain on continuously during target detection. Therefore, in order to achieve spatiotemporal matching of sonar and visual data, a mechanism for synchronous power-on acquisition of sonar and visual sensors should be established when necessary.
[0058] In particular, in some cases, in order to obtain a larger field of view and improve data acquisition efficiency, a wide field of view anamorphic lens is used. Since the target image obtained by this type of lens has pixel position offset, it is necessary to perform offset correction in order to obtain accurate target position. Offset correction includes edge magnification offset correction and mirror non-parallelism offset correction. In actual use, the relevant parameters of offset correction are determined by imaging calibration of the vision sensor lens.
[0059] Image signals, sonar signals, magnetic field signals, electric field signals, and temperature field signals acquired by sensor components each contain feature attributes of underwater targets in different dimensions. To optimize underwater target detection data using multi-sensor data, it is necessary to extract and generate multi-sensor fusion features through methods such as feature fusion. Specifically:
[0060] like Figure 2 As shown, a dual-branch network is constructed, including a first branch network, a second branch network, and an attention fusion network;
[0061] The first branch network takes image and sonar information as input and extracts visual features, while the second branch network takes electric field, magnetic field and temperature field data as input and extracts distance and direction related information. The attention fusion network uses the attention module to weight and fuse the features of the two branches to obtain a fused feature vector.
[0062] The first branch network includes a two-layer convolutional pooling unit, a global pooling unit, a two-layer fully connected unit, and a fully connected output unit. Based on the joint target detection data obtained in the preceding steps, the first branch network extracts image features acquired by the visual sensor, performs image dimensionality reduction using the two-layer convolutional pooling unit, compresses the low-dimensional image into a one-dimensional feature vector (typically 128 dimensions) using the global pooling unit, extracts sonar coordinate data, and maps the original coordinate values into a one-dimensional feature vector (typically 128 dimensions) using the two-layer fully connected unit. After concatenating the two one-dimensional vectors, the output of the fully connected layer is used to obtain the output features of the first branch network (typically 256 dimensions).
[0063] The second branch network includes a one-dimensional convolutional pooling unit, an average pooling unit, and a fully connected output unit. The second branch network extracts temporal feature data from the three-dimensional time-series data related to the magnetic field signal, electric field signal, and temperature field signal through the one-dimensional convolutional pooling unit, and outputs it as a one-dimensional feature vector (generally 128 dimensions) by global average pooling. Then, the fully connected output unit maps it to a feature vector with the same dimensions as the output of the first branch network.
[0064] The attention fusion network includes an average pooling unit, a two-layer fully connected unit, and a weighted fusion unit. The attention fusion network uses the average pooling unit to average the output features of the two branch networks, uses the two-layer fully connected unit to output attention weights, and uses these attention weights to perform a weighted fusion of the feature outputs of the two branch networks.
[0065] The feature vector obtained by fusing multi-sensor data is input into the underwater target detection and recognition model to complete target recognition and target localization, providing a basis for subsequent decision-making.
[0066] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit the scope of protection of the present invention. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the essence and scope of the technical solutions of the present invention.
Claims
1. A multi-sensor-based underwater target detection method, characterized in that, Includes the following steps: Step 1, Multi-sensor data acquisition, refers to the acquisition of raw data of underwater targets or detection areas based on multi-sensor components; the multi-sensor components include visual sensors, sonar sensors, magnetic field sensors, electric field sensors, and temperature sensors; Step 2: Spatiotemporal matching of sensor data, specifically including steps b1 to b3; b1 Data preprocessing: Preprocesses the image data acquired by the vision sensor and the coordinates collected by the sonar, performs target detection through the image recognition model, and marks the area of the target object to be detected. b2 Timing Matching: Adds timestamps to the data from the magnetic field sensor, electric field sensor, and temperature sensor, based on 3... Outliers are removed in principle, and missing values are filled in using linear interpolation before timing alignment is performed; at the same time, signals with different sampling rates are resampled to ensure timing consistency. b3 Spatial matching, which is mainly based on the position of the target's true coordinates in the corresponding data coordinates; specifically including: Establish a sonar coordinate system and a visual coordinate system. Based on the rotation and translation matrices of the visual coordinate system and the environmental coordinate system, obtain the position expression of the underwater target in the visual coordinate system. According to the imaging principle of the visual sensor, first transform to obtain the coordinate data on the imaging plane of the visual sensor. Establish an image coordinate system and establish the transformation relationship between the pixel coordinates on the image acquired by the visual sensor and the environmental coordinates. Based on the spatial matching results, the target data obtained by sonar is mapped onto the image obtained by the visual sensor. The correlation algorithm is used to determine the matching degree between the image recognition target and the sonar detection target, and the joint target detection data is output. Step 3: Multi-sensor data fusion detection, including the construction of a dual-branch network, comprising a first branch network, a second branch network, and an attention fusion network; The first branch network takes image and sonar information as input and extracts visual features, while the second branch network takes electric field, magnetic field and temperature field data as input and extracts distance and direction related information. The attention fusion network uses the attention module to weight and fuse the features of the two branches to obtain a fused feature vector. The feature vector obtained by fusing multi-sensor data is then input into the underwater target detection and recognition model to complete target recognition and target localization.
2. The underwater target detection method based on multiple sensors according to claim 1, characterized in that, Specifically, b1 includes: for image data acquired by the visual sensor, preprocessing with mean filtering to reduce noise and performing size adjustment and format conversion as described above; thresholding the data acquired by the sonar to remove invalid false data; performing image enhancement preprocessing on the image frames acquired by the camera, and then performing target detection through an image recognition model to mark the required target area.
3. The underwater target detection method based on multiple sensors according to claim 1, characterized in that, The b2 timing matching specifically includes: adding timestamps to all analog voltage data, removing outliers according to the 3 principles, filling in missing values using linear interpolation, and then aligning the data timing; at the same time, resampling signals with different sampling rates to ensure timing consistency.
4. The underwater target detection method based on multiple sensors according to claim 1, characterized in that, The spatial matching step b3 specifically includes: First, establish the location of the sonar sensor based on the position of the sonar sensor. Detecting the direction of movement of the platform The direction pointing towards the center of the earth. perpendicular to To the left Establish a sonar coordinate system ,by The point where the direction intersects with the sea level is the origin coordinate system. , to pass through the origin Furthermore, establish an environmental coordinate system on each plane that is parallel or perpendicular to the sonar coordinate system. Then, for the radial distance in the sonar sensor data, it is... yaw angle is The coordinates of the target p in the environment coordinate system are ; in This represents the distance of the sonar sensor from the sea level. Secondly, the location of the vision sensor is taken as the origin of the coordinate system. Establish a visual coordinate system with all coordinate axes aligned with the sonar coordinate system. Then, based on the rotation matrix between the visual coordinate system and the environment coordinate system and translation matrix This yields the position representation of the underwater target p in the visual coordinate system. ; Then, based on the imaging principle of the vision sensor, the coordinate data on the imaging plane of the vision sensor are first converted to obtain the coordinate data. ,in Let be the distance between the imaging plane and the optical center of the vision sensor; establish an image coordinate system with the top left corner of the image acquired by the vision sensor as the origin, and the horizontal and vertical axes as the x and y coordinates, respectively. Then, the point on the imaging plane coordinate system... Coordinate data in the image coordinate system ;in This refers to the coordinates of a point in an image. The size of a single pixel in the image; Based on the transformation relationships established in the preceding steps, the transformation relationship between pixel coordinates and environmental coordinates in the image acquired by the visual sensor can be simultaneously established as follows: ; Finally, based on the spatial matching results, the target data obtained by the sonar is mapped onto the image obtained by the visual sensor to generate the detection box corresponding to the sonar. The matching degree of the detection results of the image recognition target and the sonar detection target is judged by the association algorithm. When the matching degree of the target recognition result meets the accuracy requirements, the joint target detection data obtained by the sonar sensor and the visual sensor is output. If the accuracy requirements are not met, the target is eliminated by decision or the best matching result is output.
5. The underwater target detection method based on multiple sensors according to claim 1, characterized in that, In step b3, when the matching degree is much lower than the normal range, the validity of the sonar sensor data is first checked. If the data acquired by the sonar sensor is valid, the output of the joint target detection data will be corrected based on the sonar sensor data. If the validity of the sonar sensor data cannot be determined, the target data acquired by the sonar sensor and the vision sensor will be cross-verified to determine the valid data. When the matching degree is within the normal range, the output of the joint target detection data will be corrected based on the distance data acquired by the sonar sensor and the position and category data acquired by the vision sensor, respectively.
6. The underwater target detection method based on multiple sensors according to claim 1, characterized in that, In step b3, when using a large field-of-view anamorphic lens, offset correction is performed. Offset correction includes edge magnification offset correction and mirror non-parallelism offset correction. The relevant parameters for offset correction are determined by imaging calibration of the visual sensor lens.
7. The underwater target detection method based on multiple sensors according to claim 1, characterized in that, In step b3, a mechanism is established for the synchronous power-on and data acquisition of the sonar sensor and the visual sensor during the target detection process.
8. The underwater target detection method based on multiple sensors according to claim 1, characterized in that, Step three, multi-sensor data fusion detection, specifically includes: Construct a dual-branch network, including a first branch network, a second branch network, and an attention fusion network; The first branch network takes image and sonar information as input and extracts visual features, while the second branch network takes electric field, magnetic field and temperature field data as input and extracts distance and direction related information. The attention fusion network uses the attention module to weight and fuse the features of the two branches to obtain a fused feature vector. The first branch network includes a two-layer convolutional pooling unit, a global pooling unit, a two-layer fully connected unit, and a fully connected output unit. Based on the joint target detection data obtained in the preceding steps, the first branch network extracts the image features acquired by the visual sensor, performs image dimensionality reduction using the two-layer convolutional pooling unit, compresses the low-dimensional image into a one-dimensional feature vector using the global pooling unit, extracts sonar coordinate data, and maps the original coordinate values into a one-dimensional feature vector using the two-layer fully connected unit. The two one-dimensional vectors are concatenated and the output of the fully connected layer is used to obtain the output features of the first branch network. The second branch network includes a one-dimensional convolutional pooling unit, an average pooling unit, and a fully connected output unit. The second branch network extracts temporal feature data from the three-dimensional time-series data related to the magnetic field signal, electric field signal, and temperature field signal through the one-dimensional convolutional pooling unit, and outputs it as a one-dimensional feature vector by global average pooling. Then, the fully connected output unit maps it to a feature vector with the same dimension as the output of the first branch network. The attention fusion network includes an average pooling unit, a two-layer fully connected unit, and a weighted fusion unit. The attention fusion network uses the average pooling unit to average the output features of the two branch networks, uses the two-layer fully connected unit to output attention weights, and uses the attention weights to perform weighted fusion of the feature outputs of the two branch networks. The feature vector obtained by fusing multi-sensor data is input into the underwater target detection and recognition model to complete target recognition and target localization, providing a basis for subsequent decision-making.
9. The underwater target detection method based on multiple sensors according to claim 1, characterized in that, The visual sensor is used to acquire image data of underwater targets or detection areas; the sonar sensor is used to acquire sonar signals related to the distance and position information of underwater targets or detection areas; the magnetic field sensor is used to acquire magnetic field signals of underwater targets or detection areas; the electric field sensor is used to acquire electric field signals of underwater targets or detection areas; the temperature sensor is used to acquire temperature field signals of underwater targets or detection areas; the sonar signals, magnetic field signals, electric field signals, and temperature field signals are analog voltage signals.