Target identification method and system based on deep learning
By performing beamforming and multi-frame image stitching on the active sonar array element domain signal to form a beam-dimensional cumulative time-flow image, the problem of low recognition performance of underwater acoustic targets in complex marine environments is solved, and efficient target recognition and model stability are achieved in strong interference environments.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-04
- Publication Date
- 2026-04-10
AI Technical Summary
Existing underwater acoustic target recognition methods have low recognition performance in complex marine environments. Single-cycle data is difficult to provide stable features, deep learning models rely heavily on small sample datasets, and single-frame image input methods cannot fully describe the target motion characteristics, resulting in insufficient recognition accuracy, especially in environments with strong interference.
By converting the active sonar array element domain signal into a two-dimensional echo image, stitching together multiple consecutive frames of images to form a beam accumulation image, extracting the target region to form a beam accumulation time-flow image, and inputting it into a deep learning model for target recognition, combined with target feature representation under multi-period conditions, it is suitable for deep learning models.
It maintains high recognition accuracy and robustness in complex marine environments and under small sample conditions, can separate targets from interference in strong interference environments, enhances the interpretability of deep learning model input, reduces dependence on the amount of dataset, improves recognition accuracy and feature discrimination, and has good practicality.
Smart Images

Figure CN121831748A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of target recognition, specifically to a deep learning-based target recognition method and system, and more particularly to a deep learning-based deep learning recognition method and system for slow, small targets based on active sonar time-stream images. Background Technology
[0002] Acoustic methods are the most effective means of underwater target detection. However, existing underwater acoustic target recognition methods exhibit low performance in complex marine environments, primarily because single-cycle data is insufficient to provide stable features, and deep learning models rely heavily on small sample datasets. Current methods largely focus on improving network structures, with limited research on input data representation methods. Single-frame image input cannot adequately describe target motion characteristics, leading to insufficient recognition accuracy, especially under conditions of strong interference. Therefore, a new method is needed that can jointly represent target features under multi-cycle conditions and is suitable for deep learning recognition models. Summary of the Invention
[0003] In view of the shortcomings of the prior art, the purpose of this invention is to provide a target recognition method and system based on deep learning.
[0004] A deep learning-based target recognition method provided by the present invention includes: Step S1: Convert the active sonar array element domain signal into a two-dimensional echo image after beamforming and matched filtering; Step S2: Stitch together multiple consecutive frames of two-dimensional echo images to obtain a beam accumulation image; By cropping the target region from the beam accumulation image, a beam dimension accumulation time stream image is obtained; Step S3: Input the beam dimension cumulative time stream image into the deep learning model and output the target recognition result.
[0005] Preferably, the two-dimensional echo image uses distance and orientation as coordinate axes and includes spatial distribution information of the target echo.
[0006] Preferably, step S1 further includes preprocessing the two-dimensional echo image; the preprocessing includes normalization, noise reduction, and enhancement.
[0007] Preferably, step S2 includes: Selecting N frames means generating a beam accumulation image from N consecutive frames of images; Assuming the height and width of the two-dimensional echo image are h and w, the i-th beam of each frame of the two-dimensional echo image is stitched together to obtain a beam accumulation image with a height of h and a width of w*N. Using the target's azimuth and distance as the center, take (1+N) / 2 frames of beam accumulation image and extract a rectangular frame of size S*S to obtain the beam accumulation time-flow image of the target area.
[0008] Preferably, by stitching together N consecutive frames of data with the same beam, the time information is spread out horizontally, and the target's trajectory is represented as a tilted straight line; the direction of the straight line indicates whether the target is moving away from or towards the detection point, and the slope of the straight line corresponds to the target's speed.
[0009] A deep learning-based target recognition system according to the present invention includes: Module M1: Converts the active sonar array element domain signal into a two-dimensional echo image after beamforming and matched filtering; Module M2: Stitches together multiple consecutive frames of two-dimensional echo images to obtain a beamforming image; By cropping the target region from the beam accumulation image, a beam dimension accumulation time stream image is obtained; Module M3: Inputs the beam dimension cumulative time-stream image into the deep learning model and outputs the target recognition result.
[0010] Preferably, the two-dimensional echo image uses distance and orientation as coordinate axes and includes spatial distribution information of the target echo.
[0011] Preferably, the module M1 further includes preprocessing of the two-dimensional echo image; the preprocessing includes normalization, noise reduction and enhancement.
[0012] Preferably, the module M2 includes: Selecting N frames means generating a beam accumulation image from N consecutive frames of images; Assuming the height and width of the two-dimensional echo image are h and w, the i-th beam of each frame of the two-dimensional echo image is stitched together to obtain a beam accumulation image with a height of h and a width of w*N. Using the target's azimuth and distance as the center, take (1+N) / 2 frames of beam accumulation image and extract a rectangular frame of size S*S to obtain the beam accumulation time-flow image of the target area.
[0013] Preferably, by stitching together N consecutive frames of data with the same beam, the time information is spread out horizontally, and the target's trajectory is represented as a tilted straight line; the direction of the straight line indicates whether the target is moving away from or towards the detection point, and the slope of the straight line corresponds to the target's speed.
[0014] Compared with the prior art, the present invention has the following beneficial effects: 1. This invention explicitly embeds the target's motion information into the input image through multi-period joint accumulation of beam dimensions, enabling the deep learning model to efficiently extract physically relevant features, thereby maintaining high recognition accuracy and robustness even in complex marine environments and under small sample conditions.
[0015] 2. The beam dimension accumulation in this invention can separate the target from the interference in a strong interference environment, so that the target trajectory can be expressed with obvious geometric features. By using multi-period joint representation of time-flow images based on active sonar, the interpretability of the input of the deep learning model is enhanced, the dependence of the model on the amount of dataset is reduced, and the optimal performance of the model is achieved in the case of small sample size.
[0016] 3. This invention combines multiple features such as target intensity, scale, and velocity into a single image, enhancing the interpretability of the deep learning model input; the beam dimension accumulation method can avoid interference enhancement and highlight the dynamic features of the target.
[0017] 4. This invention can still improve the recognition accuracy and feature discrimination even with small sample sizes; and it can directly use pre-trained weights from large-scale optical image datasets for transfer learning, enhancing the model training stability and generalization ability, thus having good practicality. Attached Figure Description
[0018] Other features, objects, and advantages of the present invention will become more apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings: Figure 1 This is a single-frame distance-azimuth image in this invention.
[0019] Figure 2 This is a schematic diagram of the cumulative time flow of the beam dimension in the target area.
[0020] Figure 3 Accumulated time-flow image of beam dimension for the target region.
[0021] Figure 4 A diagram illustrating the steps involved in network design and training.
[0022] Figure 5 The curves showing the change in accuracy of target recognition as a function of cumulative frames.
[0023] Figure 6 This is a flowchart of the method of the present invention. Detailed Implementation
[0024] The present invention will now be described in detail with reference to specific embodiments. These embodiments will help those skilled in the art to further understand the present invention, but do not limit the invention in any way. It should be noted that those skilled in the art can make several changes and improvements without departing from the concept of the present invention. These all fall within the protection scope of the present invention.
[0025] Reference Figure 6 As shown, a deep learning-based target recognition method includes: Step S1: Convert the active sonar array element domain signal into a two-dimensional echo image after beamforming and matched filtering; Step S2: Stitch together multiple consecutive frames of two-dimensional echo images to obtain a beam accumulation image; By cropping the target region from the beam accumulation image, a beam dimension accumulation time stream image is obtained; Step S3: Input the beam dimension cumulative time stream image into the deep learning model and output the target recognition result.
[0026] In one embodiment, the specific steps include: Step 1: Data Preprocessing The active sonar array element domain signal is converted into a range-azimuth image through beamforming and matched filtering, and then normalized. After beamforming and matched filtering, the active sonar array element domain signal yields a two-dimensional echo image with range and azimuth as coordinate axes. This image contains spatial distribution information of the target echo, but a single frame cannot fully reflect the target's motion characteristics in the temporal dimension. Optionally, preprocessing operations such as normalization, noise reduction, and enhancement are performed on the image.
[0027] Step 2: Construction of the time-stream image Selecting N frames signifies generating a beam accumulation image from N consecutive frames. Assuming the original active sonar echo image has height and width h and w, the i-th beam from each frame is stitched together to obtain a beam accumulation image with height h and width w*N. Taking the target's azimuth and distance as the center of (1+N) / 2 frames, a rectangle of size S*S is cropped to obtain the beam accumulation time-flow image of the heavily interfered target region: by stitching together N consecutive frames of the same beam data, the time information is expanded horizontally, and the target's trajectory is represented as a sloping straight line. The direction of the line indicates whether the target is moving away from or towards the detection point, and the slope of the line corresponds to the target's velocity. This method avoids the cumulative effect of static interference and maintains the target's prominent features against a strong interference background.
[0028] Step 3: Network Design and Training 1. Create a dataset: such as Figure 4As shown, the time-stream image constructed in step two is used as the input data for the deep learning model. To ensure the effectiveness of the training process and the reliability of the results, this embodiment divides the time-stream image into a training set and a test set, and ensures the consistency of their data distribution, thereby avoiding the deviation between the training and testing phases from affecting the model's generalization ability. 2. Network Design: The deep learning model can be a two-dimensional convolutional neural network, a three-dimensional convolutional neural network, a residual network, an attention network, or a combination thereof; the model extracts local features through convolution operations and converges them layer by layer to form a global representation; 3. Network Training: The model employs transfer learning, initializing with pre-trained weights from a large-scale optical image dataset. Through transfer learning, the model can quickly adapt to sonar data using these pre-trained weights, thereby improving training stability under small sample conditions.
[0029] Step 4: Target Recognition 1. Classification Output: The deep learning model extracts and infers features from the input time-stream image, and finally outputs a classification result of the target category, used to determine which category the target belongs to. Through this classification output, automatic identification of different types of underwater targets can be achieved.
[0030] 2. Confidence Assessment and Post-processing: Upon obtaining the classification results, the model can output the probability value of each category as a confidence score. This embodiment can set a threshold; when the confidence score is higher than the threshold, the recognition result is output; when the confidence score is lower than the threshold, the sample can be marked as an "uncertain target" for manual review or further processing. Optionally, post-processing of the recognition results for consecutive frames can be performed using majority voting or time window smoothing methods to reduce single-frame errors and improve overall recognition stability.
[0031] 3. Interpretability Analysis: To improve the interpretability and credibility of the recognition results, this embodiment may optionally employ the following methods: Grad-CAM (Gradient-Weighted Class Activation Mapping): This method visualizes the feature maps of the model's last convolutional layer, generating heatmaps to show the key regions the model focuses on during recognition. Analyzing the heatmaps can verify whether the model is focusing on the target region rather than background noise.
[0032] t-SNE feature visualization: The high-dimensional features extracted by the model are reduced in dimensionality and mapped to display the distribution of various targets in a two-dimensional plane. Different categories of targets show obvious clustering effects on the mapped plane, thus verifying the separability of the features and the rationality of the model's recognition.
[0033] Example 1 This embodiment provides an active sonar target recognition method based on deep learning. To verify the effectiveness of the invention, network training is performed using data from a certain sea trial for further illustration.
[0034] Step 1: Data Preprocessing The active sonar array element domain signal is converted into a range-azimuth image through beamforming and matched filtering, and then normalized. After beamforming and matched filtering, the active sonar array element domain signal yields a two-dimensional echo image with range and azimuth as coordinate axes, such as... Figure 1 As shown.
[0035] Step 2: Construction of the time-stream image Selecting 10 frames means generating a beamforming image from 10 consecutive frames. The i-th beam of each frame is then processed as follows: Figure 2 The beam accumulation image is obtained by stitching the images together. Taking the target's azimuth and range as the center, a 224*224 rectangle is cropped from the 5th frame image to obtain the beam accumulation time-flow image of the strong interference target region, as shown below. Figure 3 As shown.
[0036] Step 3: Network Design and Training 1. Dataset establishment: The time-stream images obtained in step 2 are labeled to form 6 training samples for each flight. The samples are divided according to the target motion state, and the ratio of training set to test set is 4:1 to ensure data distribution consistency.
[0037] 2. Network Design: In this embodiment, ResNet18 is selected as the deep learning recognition model, and its structure includes: Input layer: Performs convolution and pooling operations on the input time-stream image to initially extract low-level features and reduce the feature map size; Residual module: It is composed of multiple basic residual blocks stacked together. Each residual block consists of two convolutional layers and introduces an identity mapping between the input and output to achieve cross-layer feature fusion, thereby alleviating the degradation problem in deep network training. Global average pooling layer: Performs global feature aggregation after the residual module to obtain the overall feature representation of the target; Fully connected layer and classification layer: Map global features to the target category space and output the final classification result.
[0038] Based on the above structure, the ResNet18 model in this embodiment can fully extract the spatial features and local dynamic features of the sonar time-flow image, and combine the residual connection mechanism to ensure training stability, thereby achieving high-precision underwater target recognition.
[0039] 3. Transfer Learning: During training, this embodiment uses weights pre-trained on a large-scale image dataset (e.g., ImageNet) as the initialization parameters of ResNet18, and then fine-tunes them on a sonar time-flow image dataset. Through transfer learning, the convergence speed and recognition accuracy under small sample conditions can be effectively improved.
[0040] Step 4: Target Recognition After training, the deep learning model was used to conduct target recognition experiments. By inputting a beam accumulation time-flow image, the model outputs classification results to determine the target category. Furthermore, this embodiment compares the recognition performance under different beam accumulation methods, and the results are as follows: Figure 5 As shown in the figure. Experiments show that the beam-dimensional accumulation method can effectively highlight the dynamic features of targets in strong interference backgrounds, significantly improving the accuracy of model recognition. After using the beam-dimensional accumulation method, the recognition accuracy of various targets is improved. Specifically, when the number of accumulated frames is 20, the recognition accuracy of target 1 is 92.34%, the recognition accuracy of target 2 is 98.52%, the recognition accuracy of target 3 is 100%, and the recognition accuracy of target 4 is 99.21%.
[0041] The present invention also provides a target recognition system based on deep learning, which can be implemented by executing the process steps of the target recognition method based on deep learning. That is, those skilled in the art can understand the target recognition method based on deep learning as a preferred embodiment of the target recognition system based on deep learning.
[0042] A deep learning-based target recognition system includes: Module M1: Converts the active sonar array element domain signal into a two-dimensional echo image after beamforming and matched filtering; Module M2: Stitches together multiple consecutive frames of two-dimensional echo images to obtain a beamforming image; By cropping the target region from the beam accumulation image, a beam dimension accumulation time stream image is obtained; Module M3: Inputs the beam dimension cumulative time-stream image into the deep learning model and outputs the target recognition result.
[0043] The two-dimensional echo image uses distance and orientation as coordinate axes and includes spatial distribution information of the target echo.
[0044] The module M1 also includes preprocessing of the two-dimensional echo image; the preprocessing includes normalization, noise reduction and enhancement.
[0045] The module M2 includes: Selecting N frames means generating a beam accumulation image from N consecutive frames of images; Assuming the height and width of the two-dimensional echo image are h and w, the i-th beam of each frame of the two-dimensional echo image is stitched together to obtain a beam accumulation image with a height of h and a width of w*N. Using the target's azimuth and distance as the center, take (1+N) / 2 frames of beam accumulation image and extract a rectangular frame of size S*S to obtain the beam accumulation time-flow image of the target area.
[0046] By stitching together N consecutive frames of data from the same beam, the time information is expanded horizontally, and the target's trajectory is represented as a slanted straight line; the direction of the line indicates whether the target is moving away from or towards the detection point, and the slope of the line corresponds to the target's speed.
[0047] Those skilled in the art will understand that, besides implementing the system and its various devices, modules, and units provided by this invention in the form of purely computer-readable program code, the same functions can be achieved entirely through logical programming of the method steps, making the system and its various devices, modules, and units of this invention function in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers. Therefore, the system and its various devices, modules, and units provided by this invention can be considered as a hardware component, and the devices, modules, and units included therein for implementing various functions can also be considered as structures within the hardware component; alternatively, the devices, modules, and units for implementing various functions can be considered as both software modules implementing the method and structures within the hardware component.
[0048] Specific embodiments of the present invention have been described above. It should be understood that the present invention is not limited to the specific embodiments described above, and those skilled in the art can make various changes or modifications within the scope of the claims, which do not affect the essence of the present invention. Unless otherwise specified, the embodiments and features described in this application can be arbitrarily combined with each other.
Claims
1. A target recognition method based on deep learning, characterized in that, include: Step S1: Convert the active sonar array element domain signal into a two-dimensional echo image after beamforming and matched filtering; Step S2: Stitch together multiple consecutive frames of two-dimensional echo images to obtain a beam accumulation image; By cropping the target region from the beam accumulation image, a beam dimension accumulation time stream image is obtained; Step S3: Input the beam dimension cumulative time stream image into the deep learning model and output the target recognition result.
2. The target recognition method based on deep learning according to claim 1, characterized in that, The two-dimensional echo image uses distance and orientation as coordinate axes and includes spatial distribution information of the target echo.
3. The target recognition method based on deep learning according to claim 1, characterized in that, Step S1 further includes preprocessing the two-dimensional echo image; the preprocessing includes normalization, noise reduction and enhancement.
4. The target recognition method based on deep learning according to claim 1, characterized in that, Step S2 includes: Selecting N frames means generating a beam accumulation image from N consecutive frames of images; Assuming the height and width of the two-dimensional echo image are h and w, the i-th beam of each frame of the two-dimensional echo image is stitched together to obtain a beam accumulation image with a height of h and a width of w*N. Using the target's azimuth and distance as the center, take (1+N) / 2 frames of beam accumulation image and extract a rectangular frame of size S*S to obtain the beam accumulation time-flow image of the target area.
5. The target recognition method based on deep learning according to claim 4, characterized in that, By stitching together N consecutive frames of data from the same beam, the time information is expanded horizontally, and the target's trajectory is represented as a slanted straight line; the direction of the line indicates whether the target is moving away from or towards the detection point, and the slope of the line corresponds to the target's speed.
6. A target recognition system based on deep learning, characterized in that, include: Module M1: Converts the active sonar array element domain signal into a two-dimensional echo image after beamforming and matched filtering; Module M2: Stitches together multiple consecutive frames of two-dimensional echo images to obtain a beamforming image; By cropping the target region from the beam accumulation image, a beam dimension accumulation time stream image is obtained; Module M3: Inputs the beam dimension cumulative time-stream image into the deep learning model and outputs the target recognition result.
7. The deep learning-based target recognition system according to claim 6, characterized in that, The two-dimensional echo image uses distance and orientation as coordinate axes and includes spatial distribution information of the target echo.
8. The target recognition system based on deep learning according to claim 6, characterized in that, The module M1 also includes preprocessing of the two-dimensional echo image; the preprocessing includes normalization, noise reduction and enhancement.
9. The target recognition system based on deep learning according to claim 6, characterized in that, The module M2 includes: Selecting N frames means generating a beam accumulation image from N consecutive frames of images; Assuming the height and width of the two-dimensional echo image are h and w, the i-th beam of each frame of the two-dimensional echo image is stitched together to obtain a beam accumulation image with a height of h and a width of w*N. Using the target's azimuth and distance as the center, take (1+N) / 2 frames of beam accumulation image and extract a rectangular frame of size S*S to obtain the beam accumulation time-flow image of the target area.
10. The deep learning-based target recognition system according to claim 9, characterized in that, By stitching together N consecutive frames of data from the same beam, the time information is expanded horizontally, and the target's trajectory is represented as a slanted straight line; the direction of the line indicates whether the target is moving away from or towards the detection point, and the slope of the line corresponds to the target's speed.