Underwater active sonar echo time-frequency tensor feature fusion method based on sub-beam filling
By combining sub-beam filling with multiple time-frequency analysis methods, the problems of spatial information loss and insufficient single time-frequency transformation in traditional beamforming technology are solved, and the accuracy and robustness of underwater target recognition are improved.
Patent Information
- Application Number
- CN202510917883.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-03
- Publication Date
- 2025-10-17
AI Technical Summary
Traditional beamforming technology has problems in underwater target recognition, such as spatial information loss and insufficient single time-frequency transformation method, resulting in insufficient target recognition accuracy and robustness.
A sub-beam filling strategy is adopted to divide the receiving array element into several sub-beams, and beamforming is performed separately. The time-frequency feature tensor is generated by short-time Fourier transform, continuous wavelet transform and smoothed pseudo-Wigner-Ville distribution analysis, and then a fused feature tensor is formed by combining Gaussian distribution weighting and frequency domain splicing.
It significantly enhances the spatial information retention and time-frequency feature expression capabilities of target echoes, improves the accuracy and robustness of underwater target recognition, and adapts to complex underwater environments.
Smart Images

Figure CN120805045A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of sonar echo processing, and particularly relates to a method for fusing underwater active sonar echo time-frequency tensor features based on sub-beam filling. BACKGROUND
[0002] Active sonar technology has a wide range of applications in underwater detection and target recognition. It achieves underwater target detection and recognition by emitting sound waves and receiving the echo signals reflected by the target. However, in practical applications, the echo signals of active sonar are often disturbed by complex underwater environments, such as noise and reverberation, which greatly increase the difficulty of target feature extraction.
[0003] Currently, the active sonar target recognition method mainly relies on traditional beamforming technology, which enhances the target signal through coherent accumulation of signals received by multiple array elements, thereby improving the signal-to-noise ratio (SNR). However, the method of extracting target features based on traditional beamforming has significant shortcomings. Specifically, on the one hand, the beamforming process enhances the target signal while losing the spatial information provided by the array, which affects the accurate recognition of the target; on the other hand, the time-frequency image after traditional beamforming only contains the information of the synthesized signal, lacking the ability to express the details of the target echo, which further affects the accuracy of target recognition.
[0004] In addition, existing methods also have obvious limitations in feature fusion, usually only using a single time-frequency transform method such as short-time Fourier transform (STFT), continuous wavelet transform (CWT) or Wigner-Ville distribution (WVD), failing to fully utilize the complementarity of multiple time-frequency transform methods, resulting in insufficient target feature expression ability and difficulty in meeting the target recognition needs in complex underwater environments. SUMMARY
[0005] In view of the shortcomings of the existing underwater target echo feature extraction method based on beamforming, the present application proposes a method for fusing underwater active sonar echo time-frequency tensor features based on sub-beam filling (SBF) to enhance the spatial information expression ability of the echo signal and fully exploit the time-frequency features of the target signal, providing support for improving the accuracy and robustness of underwater target recognition.
[0006] In order to achieve the above technical purpose, the following technical solutions are adopted in the present application:
[0007] In one aspect of the present application, a method for fusing underwater active sonar echo time-frequency tensor features based on sub-beam filling is provided, comprising the following steps:
[0008] S1: divide the receiving elements into several sub-beams, each sub-beam is independently beamformed, and the delay difference between the sub-beams is calculated for time alignment;
[0009] S2: the echo signals of each sub-beam are respectively subjected to short-time Fourier transform, continuous wavelet transform and smooth pseudo Wigner-Ville distribution analysis to generate time-frequency feature tensors corresponding to the sub-beams;
[0010] S3: the time-frequency features are weighted in the frequency domain by using a Gaussian distribution weighting function;
[0011] S4: the weighted time-frequency features are normalized, and the feature tensors of the short-time Fourier transform, the continuous wavelet transform and the smooth pseudo Wigner-Ville distribution are spliced in the frequency dimension to form a fusion feature tensor.
[0012] In one embodiment, according to the numbering order of the elements in the array, the elements are divided into left (L), middle (C) and right (R) sub-beams;
[0013] wherein the left sub-beam element number is the middle sub-beam element number is the right sub-beam element number is N is the total number of array elements.
[0014] In one embodiment, the output of the kth sub-beam is:
[0015]
[0016] wherein x n (t) is the time-domain signal received by the nth element, h n (t) is the delay compensation impulse response corresponding to the nth element, is a convolution operation, G k is the element set corresponding to the kth sub-beam.
[0017] In one embodiment, the delay difference between adjacent sub-beams is:
[0018]
[0019] wherein θ is the incident angle of the target echo, c is the sound speed in water, d is the element spacing, and N is the total number of array elements.
[0020] In one embodiment, the short-time Fourier transform is:
[0021]
[0022] wherein w(t) is a window function, s sub_beamis the beam domain signal output by the subarray beam, and t is the time after delay compensation.
[0023] In an embodiment, the continuous wavelet transform is:
[0024]
[0025] wherein a is a scale, b is a translation, ψ(t) is a mother wavelet, and s sub_beam is the beam domain signal output by the subarray beam, and t is the time after delay compensation.
[0026] In an embodiment, the smoothed pseudo Wigner-Ville distribution is:
[0027]
[0028] wherein g(ξ) is a smoothing function, s sub_beam is the beam domain signal output by the subarray beam, and t is the time after delay compensation.
[0029] In an embodiment, the expression of the weighting weight using a Gaussian distribution function is:
[0030]
[0031] wherein f is each frequency point in the frequency band, f0 is the center frequency of the echo; σ is a standard deviation, and controls the frequency weighting width.
[0032] In an embodiment, the normalized processing of the weighted time-frequency features is:
[0033]
[0034] wherein T(t, f) represents a time-frequency matrix of size t x f after frequency weighting, t represents the number of time points, and f represents the number of frequency points.
[0035] The beneficial effects of the present application are:
[0036] The present application proposes an innovative underwater target echo feature extraction method, the core of which is to use a sub-beam filling strategy and the combination of multiple time-frequency analysis methods. Specifically, the present application significantly enhances the spatial information retention capability of the target echo through the sub-beam filling strategy, so that the extracted features can more comprehensively reflect the spatial characteristics of the target. At the same time, the present application combines three time-frequency analysis methods of short-time Fourier transform (STFT), continuous wavelet transform (CWT) and smoothed pseudo Wigner-Ville distribution (SPWVD), greatly enriches the time-frequency feature expression ability of the target echo, and thus more accurately describes the time-frequency characteristics of the target echo.
[0037] In addition, the application also introduces a frequency domain weighted fusion technology, which not only effectively enhances the saliency of target features, but also significantly improves the robustness and adaptability of the method in a low signal-to-noise ratio environment. The extracted feature tensor can be directly used as the input of a deep learning network, thereby improving the classification accuracy of underwater target recognition. Experimental results show that the method proposed in the application achieves excellent results.
[0038] The application effectively solves the problem of spatial information loss in the conventional underwater target echo feature extraction method based on beam forming, and the deficiency of single time-frequency transformation method in representing target echo separable features, has a broad application prospect, and can be directly applied to an actual system. BRIEF DESCRIPTION OF DRAWINGS
[0039] Figure 1 is a flowchart of the underwater active sonar echo time-frequency tensor feature fusion method of the embodiment of the application;
[0040] Figure 2 is a schematic diagram of array sub-beam division of the embodiment of the application;
[0041] Figure 3 is a schematic diagram of underwater moving target time domain echo of the embodiment of the application;
[0042] Figure 4 is a time-frequency feature map of three types of targets after beam forming (BF) of the embodiment of the application;
[0043] Figure 5 is a time-frequency feature map of three types of targets after sub-beam filling (SBF) of the embodiment of the application;
[0044] Figure 6 is a time-frequency tensor fusion feature map of three types of targets based on sub-beam filling of the embodiment of the application. DETAILED DESCRIPTION
[0045] The technical solutions of the application will be described clearly and completely in combination with specific embodiments, but those skilled in the art will understand that the following described embodiments are part of the embodiments of the application, not all the embodiments, and are only used to illustrate the application, and should not be regarded as limiting the scope of the application. Based on the embodiments in the application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the application.
[0046] The application forms a sub-beam filled fusion time-frequency tensor by sub-beam forming the signals received by the array, fusing the different time-frequency methods of the sub-beams, and filling them into a tensor.
[0047] In one embodiment, with reference to Figure 1As shown, a sub-beam filling based underwater active sonar echo time-frequency tensor feature fusion method, the specific steps are as follows:
[0048] S1: Divide the receiving elements into several sub-beams, each sub-beam is independently beamformed, and the delay difference between sub-beams is calculated for time alignment.
[0049] In this embodiment, the receiving elements are divided into left (L), middle (C) and right (R) sub-beams, each sub-beam is independently beamformed to retain spatial information.
[0050] Specifically, referring to Figure 2 As shown, array sub-beam division:
[0051] Suppose the receiving array is composed of N elements, and the element spacing is d. The elements are evenly divided into three sub-beams according to the number: left sub-beam (L-subbeam), element number Middle sub-beam (C-subbeam), element number Right sub-beam (R-subbeam), element number For each sub-beam, independent beamforming processing is performed.
[0052] Wherein, the output of the kth sub-beam is:
[0053]
[0054] Wherein, x n (t) is the time domain signal received by the nth element, h n (t) is the delay compensation impulse response corresponding to the nth element, is the convolution operation, G k is the element set corresponding to the kth sub-array, k∈{L,C,R}.
[0055] In order to ensure the time consistency of different sub-beam outputs, delay compensation is needed. The delay difference formula is:
[0056]
[0057] Wherein, θ is the incident angle of the target echo; c is the sound speed in water, taking 1500m / s.
[0058] S2: Perform short-time Fourier transform, continuous wavelet transform and smooth pseudo Wigner-Ville distribution analysis on the echo signal of each sub-beam respectively, and generate the time-frequency feature tensor corresponding to the sub-beam.
[0059] For each sub-beam L, C, R, extract three kinds of time-frequency features respectively to obtain the corresponding time-frequency tensor, the size of the time-frequency tensor of each sub-array beam is [T tftX F tft X 1, T tft F is the time dimension of the selected time-frequency transform method tft F is the frequency dimension of the selected time-frequency transform method, and the number of channels is 1.
[0060] Specifically, the STFT is:
[0061]
[0062] Where w(t) is a window function, s sub_beam is the beam domain signal of the subarray beam output, and t is the time after delay compensation.
[0063] Specifically, the CWT is:
[0064]
[0065] Where a is the scale, b is the translation, ψ(t) is the mother wavelet, s sub_beam is the beam domain signal of the subarray beam output, and t is the time after delay compensation.
[0066] Specifically, the SPWVD is:
[0067]
[0068] Where g(ξ) is a smoothing function, s sub_beam is the beam domain signal of the subarray beam output, and t is the time after delay compensation.
[0069] S3: A Gaussian distribution weighting function is used to weight the time-frequency features in the frequency domain. This is achieved by element-wise multiplication of each time-frequency matrix and the frequency weight matrix.
[0070] To highlight the characteristics of the target frequency band, the application uses a Gaussian distribution function as the frequency weighting mechanism, and the weighting weight expression is set as follows:
[0071]
[0072] Where f is each frequency point in the frequency band, f0 is the center frequency of the echo; σ is the standard deviation, which controls the frequency weighting width.
[0073] S4: The weighted time-frequency features are normalized, and the feature tensors of short-time Fourier transform, continuous wavelet transform and smooth pseudo Wigner-Ville distribution are spliced in the frequency dimension to form a fusion feature tensor.
[0074] Specifically, the normalized weighted time-frequency feature map is:
[0075]
[0076] T(t,f) represents a time-frequency matrix of size t x f after frequency weighting, t represents the number of time points, and f represents the number of frequency points.
[0077] In the frequency dimension, the characteristic tensors of the short-time Fourier transform, the continuous wavelet transform, and the smooth pseudo Wigner-Ville distribution are spliced to form a fusion characteristic tensor:
[0078]
[0079] Where c represents the number of sub-beam channels.
[0080] S5: Feature recognition and verification. A ResNet-34 network is used as a classifier, the input is the fusion characteristic tensor, and the training set and test set are divided by 70% / 15% / 15%. The network outputs the classification accuracy and F1-score indicators to evaluate the fusion feature effect.
[0081] Embodiment
[0082] Taking pool test data as an example, the time-domain waveforms of the three types of targets are shown in Figure 3 The STFT, CWT, and SPWVD time-frequency diagrams obtained after beamforming are shown in Figure 4 The STFT, CWT, and SPWVD after sub-beam filling are shown in Figure 5 As can be clearly seen from the external appearance of the two groups of images, the pseudo-color image generated by the traditional beamforming (BF) method has obvious deficiencies in overall clarity and structural expression. The pseudo-color image processed by the STFT and CWT methods has low brightness, unclear spectral region, and difficulty in distinguishing the differences between targets, especially in the CWT image, where the three types of targets only show low-frequency weak response near the bottom, and almost no recognizable structural features can be formed. Although the SPWVD retains some energy concentration areas in the pseudo-color image, the overall performance is still limited by the spatial information compression and main lobe shift interference caused by BF, resulting in serious loss of feature details and low recognition degree.
[0083] In contrast, the color image generated by the sub-beam filling (SBF) strategy has significant advantages in image expression. Whether it is STFT, CWT, or SPWVD, the color image can clearly present the overall structure and local details of the target echo spectrum, the target main energy band is located in the medium and high frequency region, the frequency band width and intensity distribution law are obvious, and there are significant differences between different targets in energy trajectory, frequency trend, and highlight area form. Among them, STFT provides a stable spectral reference, CWT shows certain multi-scale features in local details, and SPWVD realizes accurate depiction of high-resolution transient structure, each of which has its own advantages, forming a complementary system of time-frequency expression.
[0084] The SBF method enhances the expression capability of the target spectral structure by maintaining the spatial multi-channel characteristics of the array, and is a key basis for realizing fine feature extraction of underwater targets. Meanwhile, the three time-frequency analysis methods have significant complementarity in the spectral structure dimension.
[0085] Figure 6 The time-frequency tensor feature images of three types of sea trial targets under the condition of sub-beam filling (SBF) are shown in FIG. 6. From the visual performance, the fused feature map is significantly better than the single time-frequency method in terms of spectral integrity and detail level. The overall presents strong structure separation and color level distribution, and different targets exhibit unique "energy trajectory" and color band pattern, reflecting the multi-dimensional information of frequency trend, amplitude change and time distribution. The upper layer of the fused image presents the X-shaped structure formed by the high-resolution features of SPWVD, which reflects the stable symmetry of the target echo in the instantaneous frequency dimension; the middle color band is dominated by the STFT features, representing the main energy distribution of the overall spectrum of the target, with good frequency band coverage and robustness; and the lower layer mainly presents the multi-scale structure stripes of CWT, enhancing the ability to capture slowly varying and non-stationary features. The sequential arrangement of the three in the frequency dimension and the channel fusion make the image have global reference, local detail and high-resolution transient response at the same time, thereby realizing structured, multi-scale and anti-interference feature expression.
[0086] To further quantitatively verify the performance of the method, ResNet-34 is used as the classification network to test the recognition of the data sets constructed by different time-frequency features under BF and SBF. The results are shown in Table 1. The SBF strategy is better than the traditional BF method in all single features, especially in CWT (from 77.4% to 81.6%). The final fused feature further improves the recognition accuracy based on SBF, reaching an accuracy of 85.3%, which fully illustrates that the sub-beam filling and multi-time-frequency feature fusion mechanism proposed in the present application has significant effect in enhancing the target recognition and robustness.
[0087] Table 1 Recognition effect of different feature extraction methods
[0088] Dataset Classification accuracy F1 -score BF + STFT 82.1% 80.9% BF + CWT 77.4% 75.6% BF + SPWVD 79.3% 77.7% SBF + STFT 84.4% 83.1% SBF + CWT 81.6% 80.1% SBF + SPWVD 79.9% 78.3% SBF + FUSION 85.3% 84.0%
[0089] Although the embodiments of the present application are described above in combination with the drawings, the present application is not limited to the above specific embodiments and application fields, and the above specific embodiments are only illustrative and guiding, but not limiting. Those skilled in the art can make many forms under the inspiration of the present application and without departing from the scope protected by the claims of the present application, which all belong to the protection of the present application.
Claims
1. A method for fusion of underwater active sonar echo time-frequency tensor features based on sub-beam filling, characterized by: include: The receiving array element is divided into several sub-beams, each sub-beam is independently beamformed, and the delay difference between the sub-beams is calculated for time alignment; The echo signal of each sub-beam is subjected to short-time Fourier transform, continuous wavelet transform and smoothed pseudo-Wigner-Ville distribution analysis to generate the time-frequency feature tensor corresponding to the sub-beam; Gaussian distribution weighting function is used to perform frequency domain weighting on time-frequency features; The weighted time-frequency features are normalized, and the feature tensors of short-time Fourier transform, continuous wavelet transform and smoothed pseudo-Wigner-Ville distribution are spliced in the frequency dimension to form a fused feature tensor.
2. The underwater active sonar echo time-frequency tensor feature fusion method according to claim 1, characterized in that: According to the numbering order of the array elements in the array, the array elements are divided into three sub-beams: left (L), center (C), and right (R); Among them, the left sub-beam array element number is Neutron beam array element number is The right sub-beam element number is N is the total number of array elements.
3. The underwater active sonar echo time-frequency tensor feature fusion method according to claim 1, characterized in that: The kth sub-beam output is: Among them, x n (t) is the time domain signal received by the nth array element, h n (t) is the delay compensation impulse response corresponding to the nth array element, is the convolution operation, G k is the array element set corresponding to the kth sub-beam.
4. The method for fusing time-frequency tensor features of underwater active sonar echoes according to claim 1, characterized in that: The delay difference between adjacent sub-beams is: Where θ is the incident angle of the target echo, c is the speed of sound in water, d is the array element spacing, and N is the total number of array elements.
5. The method for fusing time-frequency tensor features of underwater active sonar echoes according to claim 1, characterized in that: The short-time Fourier transform is: Among them, w(t) is the window function, s sub_beam is the beam domain signal output by the subarray beam, and t is the time after delay compensation.
6. The underwater active sonar echo time-frequency tensor feature fusion method according to claim 1, characterized in that: The continuous wavelet transform is: Among them, a is the scale, b is the translation, ψ(t) is the mother wavelet, s sub_beam is the beam domain signal output by the subarray beam, and t is the time after delay compensation.
7. The underwater active sonar echo time-frequency tensor feature fusion method according to claim 1, characterized in that: The smoothed pseudo-Wigner-Ville distribution is: Among them, g(ξ) is a smooth function, s sub_beam is the beam domain signal output by the subarray beam, and t is the time after delay compensation.
8. The method for fusing time-frequency tensor features of underwater active sonar echoes according to claim 1, characterized in that: The weight expression using Gaussian distribution function is: Where f is each frequency point in the frequency band, f0 is the echo center frequency, and σ is the standard deviation, which controls the frequency weighting width.
9. The method for fusing time-frequency tensor features of underwater active sonar echoes according to claim 1, characterized in that: The weighted time-frequency features are normalized as follows: Where T(t,f) represents the time-frequency matrix of size t×f after frequency weighting, t represents the number of time points, and f represents the number of frequency points.