Detection, segmentation and identification method for underwater acoustic communication signal
By constructing a Gaussian mixture model and short-time fractional Fourier transform combined with adaptive morphological filtering and Hough transform, and combining it with the multi-head self-attention mechanism of a lightweight neural network, the problem of low signal segmentation and recognition efficiency in underwater acoustic communication signal detection and recognition is solved, and efficient signal detection and recognition is achieved.
Patent Information
- Application Number
- CN202510828777.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-20
- Publication Date
- 2025-10-03
AI Technical Summary
Existing underwater acoustic communication signal detection and recognition methods are inefficient in complex ocean environments, lack effective signal segmentation methods, and suffer from severe multipath effects and noise interference, resulting in degraded recognition performance.
A Gaussian mixture model is constructed, and the time-frequency features are obtained using short-time fractional Fourier transform. A detection and segmentation model combining adaptive morphological filtering and Hough transform is designed. The cyclic spectrum and square power spectrum features are extracted, and a lightweight neural network is combined to perform multi-scale fusion multi-head self-attention mechanism recognition.
It improves the detection capability of underwater acoustic communication signals under low signal-to-noise ratio, realizes efficient signal segmentation and recognition, and improves recognition accuracy.
Smart Images

Figure CN120744358A_ABST
Abstract
Description
Technical Field
[0001] The invention relates to a detection, segmentation and recognition method for underwater acoustic communication signals, and belongs to the technical field of signal processing. Background Art
[0002] With the continuous deepening of human development of the ocean, collecting and analyzing large amounts of ocean data in complex marine environments has become a research hotspot in the field of underwater acoustics. Among them, the detection and recognition of underwater acoustic communication signals plays an important role in underwater acoustic reconnaissance.
[0003] Existing detection methods primarily include matched filtering and energy detection. The matched filtering method is computationally simple but relies on prior information, while the energy detection method is easy to implement but sensitive to noise. Existing recognition methods primarily include likelihood ratio-based and feature extraction-based recognition methods. The former relies on assumed prior knowledge of the received signal and has limited applicability. The latter has lower computational complexity, but single features are significantly affected by the environment, resulting in a low overall recognition rate in practical applications. However, existing recognition methods lack the ability to determine whether an underwater acoustic communication signal is present, resulting in low detection efficiency. Furthermore, most methods assume that the signal is segmented, but in practical applications, there is a lack of effective signal segmentation methods, resulting in reduced recognition performance. Furthermore, multipath effects and noise interference in the underwater acoustic channel severely impact detection and recognition performance. These challenges require the design of an efficient method for detecting and recognizing multiple underwater acoustic communication signals, combining advanced signal processing techniques with neural network methods, to address issues such as signal presence detection, signal segmentation, and validity.
[0004] Aiming at the problem of underwater acoustic communication signal detection and recognition in an underwater acoustic channel environment, the present invention proposes a detection, segmentation and recognition method for underwater acoustic communication signals. A Gaussian mixture model is constructed to determine whether the acquired underwater acoustic signal contains an underwater acoustic communication signal. The short-time fractional-order Fourier transform is used to obtain the short-time fractional-order time-frequency features of different underwater acoustic data. A detection and segmentation model combining adaptive morphological filtering and Hough transform is designed. For the segmented underwater acoustic communication signal, its cyclic spectrum contour map and square power spectrum features are extracted and used together with the short-time fractional-order time-frequency features as the input of the modulation recognition model. A lightweight network model with a multi-head self-attention mechanism of multi-scale fusion is designed to realize the recognition of the modulation mode of underwater acoustic communication signals under low signal-to-noise ratio conditions. Summary of the Invention
[0005] The purpose of the present invention is to provide a method for detecting, segmenting and recognizing underwater acoustic communication signals. First, a Gaussian mixture model is constructed and trained using historical data or simulation data to generate a Gaussian mixture prediction model. The Kullback-Leibler divergence is obtained to determine whether the acquired underwater acoustic signal contains an underwater acoustic communication signal. Second, the short-time fractional-order time-frequency features of different underwater acoustic data are obtained using a short-time fractional-order Fourier transform. Third, an adaptive morphological filtering method is designed to perform morphological filtering on the obtained short-time fractional-order time-frequency features of the underwater acoustic data to improve the quality of the feature map. The Hough transform method is then used to perform morphological detection and segment the underwater acoustic communication signal. Then, for the segmented underwater acoustic communication signal, its cyclic spectrum contour map and square power spectrum features are extracted and used together with the short-time fractional-order time-frequency features as the input of a modulation recognition model. Finally, based on a lightweight neural network model, a multi-head self-attention mechanism with multi-scale fusion is designed to form a lightweight recognition network for the modulation mode of the underwater acoustic communication signal. The recognition model is then trained using historical data to realize the recognition of the modulation mode of the underwater acoustic communication signal. Compared with the existing technology, the detection capability of underwater acoustic communication signals under low signal-to-noise ratio is improved, and efficient detection, segmentation and recognition are achieved.
[0006] To achieve the above-mentioned purpose, the present invention adopts the following technical solution: a method for detecting, segmenting and identifying underwater acoustic communication signals, comprising the following steps:
[0007] Step S1: The underwater acoustic data includes underwater acoustic communication signals and non-underwater acoustic communication signals. After frame processing, Gaussian mixture models are constructed for each signal to obtain the Kullback-Leibler (KL) divergence of the two probability distributions, which is used to determine whether the acquired underwater acoustic signals contain underwater acoustic communication signals.
[0008] Step S2: normalize the underwater acoustic communication signal data obtained in step S1, and obtain short-time fractional-order time-frequency features using short-time fractional-order Fourier transform.
[0009] Step S3: Design an adaptive morphological filtering method to perform morphological filtering on the short-time fractional-order time-frequency features of the underwater acoustic data obtained in step S2 to improve the quality of the feature map, and use the Hough transform method to perform morphological detection to segment the underwater acoustic communication signal.
[0010] Step S4: For the underwater acoustic communication signal segmented in step S3, extract its cyclic spectrum contour map and square power spectrum features, and use them together with the short-time fractional-order time-frequency features as inputs to the modulation recognition model.
[0011] Step S5: Based on the lightweight neural network, a lightweight network with a multi-head self-attention mechanism based on multi-scale fusion is designed. The short-time fractional-order time-frequency diagram, cyclic spectrum contour diagram, and square power spectrum feature training obtained in step S4 are used to generate a modulation recognition model to realize the recognition of the modulation mode of the underwater acoustic communication signal.
[0012] Preferably, in step S1, the underwater acoustic signal at the receiving end can be expressed by the following formula.
[0013] S sig (t) = x(t)*h(t)+n(t)(1)
[0014] Where: x(t) represents the transmitting end signal, including linear frequency modulation signal (LFM), binary frequency shift keying signal (2FSK), quaternary frequency shift keying signal (4FSK), binary phase shift keying signal (BPSK), quadrature phase shift keying signal (QPSK), direct spread spectrum signal (DSSS), orthogonal frequency division multiplexing signal (OFDM) and other signals; h(t) represents the impulse response of the underwater acoustic channel, which can be expressed by the following formula.
[0015]
[0016] Where: A n represents the signal amplitude arriving along the nth path, where there are N paths, τ n is the path delay.
[0017] The acquired underwater acoustic data is framed and processed. A Gaussian mixture model is constructed to obtain a Gaussian probability density function. The KL divergence between the Gaussian distributions of underwater acoustic communication signals and non-underwater acoustic communication signals is calculated to determine whether the acquired underwater acoustic data contains underwater acoustic communication signals. The two Gaussian density functions are shown in the following equations.
[0018]
[0019] Where: p1 and p2 represent the Gaussian density functions of two signal data, s i represents a random variable, π k represents the weight of the kth mixture component, μ k and σ k Represents the mean vector and covariance matrix of the kth mixture component, N(s i |μ k ,σ k ) represents the probability distribution function of the kth mixture component, as shown below.
[0020]
[0021] The KL divergence between two Gaussian mixture models can be expressed as follows.
[0022]
[0023] Preferably, in step S2, the short-time fractional Fourier transform of the signal uses a sliding window to perform spectrum truncation calculation on the signal to obtain the relationship between time and frequency. The calculation formula is shown below.
[0024]
[0025] Where: t represents the time variable, u represents the variable in the fractional Fourier transform domain, α represents the fractional order, ω(t) represents the window function, K α (u,β) represents the kernel function. The kernel function is:
[0026]
[0027] Where: Φ = απ / 2 represents the rotation angle.
[0028] Preferably, in step S3, an adaptive morphological filtering method is designed to obtain a new feature enhancement method for underwater acoustic communication signal detection. The short-time fractional-order time-frequency features of the underwater acoustic data obtained in step S2 are subjected to adaptive morphological processing, and the morphological processing steps are as follows:
[0029] Step S31: Convert the short-time fractional-order time-frequency feature S(f,τ) into a two-dimensional image matrix V(f,τ). The format of the two-dimensional image matrix V(f,τ) is:
[0030]
[0031] Where: s fL (t0) indicates that the frequency at time t0 is f L The signal strength of the communication signal.
[0032] Step S32: Calculate the local density in the neighborhood of the pixel points of the two-dimensional image matrix V(f, t), and obtain the optimal scale parameter through mapping according to the maximum scale of the structure element.
[0033] Step S33: Perform a morphological opening operation on the two-dimensional image matrix V(f,τ) according to the optimal scale parameter. The morphological opening operation expression is:
[0034]
[0035] Where: B represents the structural element, that is, the core in the image corrosion and expansion; Θ represents the corrosion operation; represents the dilation operation; Indicates the opening operation.
[0036] Step S34: Perform a morphological closing operation on the matrix V'(f,τ). The result of the morphological closing operation is:
[0037]
[0038] Where: V"(f,τ) represents the matrix after morphological closing operation, which is used to obtain the filtered time-frequency map.
[0039] Preferably, in step S3, the Hough transform method is used for morphological detection, and the Canny operator can be used to extract edges from the feature map output by the adaptive morphological filter to obtain edge pixels in the image, and then the edge points are converted to the Hough space according to the Hough transform. The transformation relationship is expressed as:
[0040] ρ=xcosθ+ysinθ (12)
[0041] Where: ρ is the vertical distance from the origin to the straight line, and θ is the counterclockwise angle between the perpendicular line of the straight line and the horizontal axis.
[0042] Quantize the polar coordinate θ-P space into M×N L The small grid forms the accumulator space. The accumulator space is expressed as:
[0043]
[0044] Where: x max and y max Respectively represent the maximum coordinate values of the pixel points in the image in the x and y directions, M represents the number of discrete points in the θ-P space along the θ direction, and N L Represents the number of discrete points along the P direction in the θ-P space. Based on the coordinates (x, y) of each point in the rectangular coordinate system, the P value is calculated in small grid steps within the range θ = 0° to 180°. Whenever the resulting value falls within a grid, the accumulator for that grid increments by 1. After all the edge pixels in the rectangular coordinate system have been transformed, the (θP) value corresponding to the maximum value in the accumulator is selected as the desired line.
[0045] Finally, the minimum length and maximum gap of the line segment are set, and the starting and ending points of the underwater acoustic communication signal are calculated after eliminating short line segments and noise, thus completing the detection and segmentation of the underwater acoustic communication signal.
[0046] Preferably, in step S4, the cyclic spectrum contour map and square power spectrum features of the underwater acoustic communication signal segmented in step S3 are extracted and used together with the short-time fractional-order time-frequency features as inputs to the modulation recognition model. The steps of extracting features are as follows:
[0047] Step S41: Use the Welch method to obtain the square power spectrum of the signal. The square power spectrum is expressed as:
[0048]
[0049] Where: J represents the number of segments used to divide the signal sequence, The power spectrum of the j-th segment signal can be expressed as follows:
[0050]
[0051] Where: L j Indicates the length of the j-th segment signal, L sig represents the length of the signal, w(i) represents the window function, and U represents the normalization factor.
[0052] Step S42: Perform Fourier transform on the cyclic autocorrelation function to obtain the cyclic spectrum density function, generate a three-dimensional cyclic spectrum, and then convert it into a contour map to obtain the cyclic spectrum contour features of the underwater acoustic communication signal. The cyclic spectrum can be expressed as:
[0053]
[0054] Where: represents the cyclic autocorrelation function, as shown below:
[0055]
[0056] Where: ζ represents the cycle frequency, t1 represents the time, and Δt represents the time interval.
[0057] Preferably, in step S5, based on a lightweight neural network, a lightweight network based on multi-scale fusion and multi-head self-attention mechanism (MF-MHSA-LNet) is designed, and the short-time fractional-order time-frequency map, cyclic spectrum contour map, and square power spectrum feature obtained in step S4 are used to train and generate a recognition network to realize the recognition of the modulation mode of underwater acoustic communication signals. The MF-MHSA-LNet model is specifically as follows:
[0058] Step S51: Based on the lightweight neural network SqueezeNet, the feature maps output by Fire2, Fire5, and Fire9 are extracted. For an input size of 227×227×3, the output feature maps are sized 56×56×128, 28×28×256, and 14×14×512, corresponding to the shallow features, intermediate transition features, and deep features of the image, respectively, and serve as the multi-scale feature inputs for the multi-head self-attention module. The SqueezeNet model consists of two convolutional layers, eight Fire modules, and a Softmax layer. The Fire module consists of compression and expansion layers, using a large number of 1×1 convolution kernels and a mixture of 1×1 and 3×3 convolution kernels instead of 3×3 convolution kernels. Ignoring the bias, the expansion layer reduces the number of network parameters by 9 times. The number of parameters in the expansion layer is shown in the following formula.
[0059] Y Squeeze =Con×Num×3×3(18)
[0060] Where: Con represents the number of convolution kernels, and Num represents the number of input channels.
[0061] Step S52: The multi-head self-attention module uses m with different projection matrices h Head, input feature X MHSA Mapped to different vector spaces, transformed into m h Group query vector Key Vector Sum vector The m obtained by the splicing operation is sent to the attention pool in parallel for calculation. h Different attention outputs. The results of the attention output are:
[0062]
[0063] Where: Represents m h The learnable projection matrix of each head; d a =d / m h Indicates the size of each attention head.
[0064] Using learnable linear projections right Perform dimension transformation to obtain attention representation As shown in the following formula.
[0065]
[0066] Where: Concat[.] represents the concatenation operation.
[0067] Step S53: The attention representation of the MHSA module Input the fully connected classifier for recognition, thereby completing the recognition of the modulation mode of the underwater acoustic communication signal.
[0068] Beneficial effects of the present invention: The present invention proposes a method for detecting, segmenting and identifying underwater acoustic communication signals. First, a Gaussian mixture model is constructed, and the KL distance is used to realize the identification of underwater acoustic communication signals and non-underwater acoustic communication signals. Secondly, the short-time fractional Fourier transform is used to obtain the time-frequency characteristics of different underwater acoustic data, and an adaptive morphological filtering method is designed for morphological filtering to enhance the signal characteristics. The Hough transform method is used for morphological detection, which can quickly detect and segment underwater acoustic communication signals and achieve high-precision detection of target signals. Then, a lightweight network model with a multi-head self-attention mechanism of multi-scale fusion is constructed, and the short-time fractional time-frequency graph, cyclic spectrum contour graph, and square power spectrum features are used to realize efficient recognition of underwater acoustic communication signals with various modulation modes. This is close to engineering practice and easy to apply in engineering. BRIEF DESCRIPTION OF THE DRAWINGS
[0069] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are merely embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying any creative work.
[0070] Figure 1 It is an overall flow chart of the method of the present invention;
[0071] Figure 2 is a schematic diagram of a Gaussian mixture model in an embodiment of the present invention;
[0072] Figure 3 is a flow chart of adaptive morphological filtering in an embodiment of the present invention;
[0073] Figure 4 is a flow chart of the Hough transform in an embodiment of the present invention;
[0074] Figure 5 This is the architecture of the MF-MHSA-LNet model in the embodiment of the present invention;
[0075] Figure 6 It is the underwater acoustic communication signal recognition result in the embodiment of the present invention. DETAILED DESCRIPTION
[0076] The technical solution of the present invention is further described below with reference to the accompanying drawings and using actual data. Obviously, the embodiments described are only some embodiments of the present invention, not all embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts are within the scope of protection of the present invention.
[0077] The process of the method of the present invention is as follows Figure 1 As shown, the following steps are included:
[0078] Step S1: In this embodiment, the types of underwater acoustic signal data include LFM, 2FSK, 4FSK, BPSK,
[0079] QPSK, DSSS), OFDM and other signals. The underwater acoustic signal at the receiving end can be expressed as follows.
[0080] S sig (t)=x(t)*h(t)+n(t)(22)
[0081] Where: x(t) represents the transmitted signal; h(t) represents the impulse response of the underwater acoustic channel, which can be expressed by the following formula.
[0082]
[0083] Where: A n represents the signal amplitude arriving along the nth path, where there are N paths, τ n is the path delay.
[0084] Considering the differences in the sampling rates of received signals, the received signals were resampled before the test to obtain the same sampling rate, and a data set of 1000 sea trial signals was obtained through framing.
[0085] In this embodiment, for the acquired underwater acoustic data, a Gaussian mixture model is constructed to obtain a Gaussian probability density function, the component is selected to maximize the probability, and the KL distribution between the Gaussian distribution of the underwater acoustic communication signal and the non-underwater acoustic communication signal is calculated.
[0086] Divergence is used to calculate the empirical threshold to determine whether the underwater acoustic communication signal is included. Figure 2 A schematic diagram of a Gaussian mixture model is shown. Specifically, the Gaussian density function is shown below.
[0087]
[0088] The KL divergence between two Gaussian mixture models can be expressed as follows.
[0089]
[0090] Step S2: Normalize the underwater acoustic communication signal data obtained in step S1, and use short-time fractional Fourier transform to obtain time-frequency features.
[0091] In this embodiment, the short-time fractional Fourier transform S(t,u,α) is expressed as follows.
[0092]
[0093] Where t represents the time variable, u represents the variable in the fractional Fourier transform domain, α represents the fractional order, ω(t) represents the window function, and K α (u,β) represents the kernel function. The kernel function can be expressed as:
[0094]
[0095] Wherein, Φ=απ / 2 represents the rotation angle.
[0096] Step S3: performing adaptive morphological processing on the time-frequency map constructed in step S2 to obtain a filtered time-frequency map. Figure 3 The flowchart of the adaptive morphological filtering is shown. Subsequently, the Hough transform method based on polar coordinate transformation space is used to detect lines, obtaining the start and end points of the signal. The start and end times of the underwater acoustic communication signal are calculated based on the start and end points of the signal, achieving a detection probability of 94% for this model. Figure 4 shows a flowchart of the Hough transform.
[0097] Step S4: For the underwater acoustic communication signal segmented in step S3, extract its cyclic spectrum contour map and square power spectrum features, and use them together with the short-time fractional-order time-frequency features as inputs to the modulation recognition model.
[0098] In this embodiment, first, the square power spectrum of the signal is obtained using the Welch method, as shown in the following formula.
[0099]
[0100] Where: J represents the number of segments used to divide the signal sequence, The power spectrum of the j-th segment signal can be expressed as follows:
[0101]
[0102] Where: L j Indicates the length of the j-th segment signal, L sig represents the length of the signal, w(i) represents the window function, and U represents the normalization factor.
[0103] Secondly, the cyclic autocorrelation function is Fourier transformed to obtain the cyclic spectrum density function, which is then converted into a cyclic spectrum contour map. The cyclic spectrum can be expressed as:
[0104]
[0105] Where: represents the cyclic autocorrelation function, as shown below:
[0106]
[0107] Where: represents the cycle frequency, t1 represents the time, and Δt represents the time interval.
[0108] Step S5: Using the short-time fractional time-frequency graph, cyclic spectrum contour graph, and square power spectrum features of the underwater acoustic communication signal data as learning features, a lightweight network based on multi-scale fusion and multi-head self-attention mechanism (MF-MHSA-LNet) is designed. Figure 5 shows the architecture of the MF-MHSA-LNet model.
[0109] The lightweight neural network SqueezeNet model consists of two convolutional layers, eight Fire modules, and a Softmax layer. The Fire module consists of compression and expansion layers. It uses a large number of 1×1 convolution kernels and a mixture of 1×1 and 3×3 convolution kernels instead of 3×3 convolution kernels, reducing the number of network parameters. Based on SqueezeNet, the feature maps output by Fire2, Fire5, and Fire9 are extracted. For an input size of 227×227×3, the output feature maps are sized 56×56×128, 28×28×256, and 14×14×512, corresponding to the shallow features, intermediate transition features, and deep features of the image, respectively. These serve as the multi-scale feature input for the multi-head self-attention module.
[0110] The multi-head self-attention module uses m with different projection matrices h heads, mapping the input features to different vector spaces and transforming them into m h The query vector, key vector and value vector are sent to the attention pool in parallel for calculation. The m obtained by the splicing operation h The attention output is transformed by learnable linear projection to obtain the attention representation, which is then input into a fully connected classifier for recognition, and finally the underwater acoustic communication signal recognition model is obtained.
[0111] In this embodiment, the underwater acoustic communication signal recognition model is applied to the signal's short-time fractional-order time-frequency diagram, cyclic spectrum contour diagram, and square power spectrum features for testing to complete the recognition of the modulation mode and obtain the recognition accuracy of the model. The confusion matrix of the recognition result is as follows: Figure 6 As shown in the figure, the average recognition accuracy of the model is about 97%, which shows that the model can achieve accurate modulation pattern recognition and verifies the effectiveness of the model.
[0112] The present invention has been described in detail. To avoid obscuring the concept of the present invention, some details well known in the art have not been described. Based on the above description, those skilled in the art can fully understand how to implement the technical solutions disclosed in the present invention.
[0113] The specific implementation methods described above further illustrate the objectives, technical solutions and beneficial effects of the present invention in detail. It should be understood that the above description is only a specific implementation method of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A method for detecting, segmenting and identifying underwater acoustic communication signals, characterized in that It includes the following steps: Step S1: constructing a Gaussian mixture model based on the characteristic differences between underwater acoustic communication signals and non-underwater acoustic communication signals, and using historical data to train and generate a Gaussian mixture prediction model to determine whether the acquired underwater acoustic signal contains an underwater acoustic communication signal; Step S2: for each underwater acoustic data obtained in step S1, a short-time fractional-order Fourier transform is used to obtain a short-time fractional-order time-frequency feature; Step S3: Design an adaptive morphological filtering method to perform morphological filtering on the short-time fractional-order time-frequency features of the underwater acoustic data obtained in step S2 to improve the quality of the feature map, and use the Hough transform method to perform morphological detection to segment the underwater acoustic communication signal; Step S4: extracting the cyclic spectrum contour map and square power spectrum features of the underwater acoustic communication signal segmented in step S3, and using them together with the short-time fractional-order time-frequency features as inputs to the modulation recognition model; Step S5: Based on the lightweight neural network, a lightweight network with a multi-head self-attention mechanism based on multi-scale fusion is designed. The short-time fractional-order time-frequency diagram, cyclic spectrum contour diagram, and square power spectrum feature training obtained in step S4 are used to generate a modulation recognition model to realize the recognition of the modulation mode of the underwater acoustic communication signal.
2. The method for detecting, segmenting and identifying underwater acoustic communication signals according to claim 1, wherein: The specific steps of step S1 are as follows: The underwater acoustic signal at the receiving end can be expressed as follows: S sig (t)=x(t)*h(t)+n(t)(1) Where: x(t) represents the transmitting end signal, including linear frequency modulation signal, binary frequency shift keying signal, quaternary frequency shift keying signal, binary phase shift keying signal, orthogonal phase shift keying signal, direct spread spectrum signal, orthogonal frequency division multiplexing signal and other signals; h(t) represents the impulse response of the underwater acoustic channel Where: A n represents the signal amplitude arriving along the nth path, where there are N paths, τ n is the path delay; The acquired underwater acoustic data is framed and processed, a Gaussian mixture model is constructed to obtain a Gaussian probability density function, and the KL divergence between the Gaussian distributions of the underwater acoustic communication signal and the non-underwater acoustic communication signal is calculated to determine whether the acquired underwater acoustic signal contains an underwater acoustic communication signal. The two Gaussian density functions are shown in the following formulas: Where: p1 and p2 represent the Gaussian density functions of two signal data, s i represents a random variable, π k represents the weight of the kth mixture component, μ k and σ k Represents the mean vector and covariance matrix of the kth mixture component, N(s i |μ k ,σ k ) represents the probability distribution function of the kth mixture component, as shown below; The KL divergence between two Gaussian mixture models can be expressed as follows; 3. The method for detecting, segmenting and identifying underwater acoustic communication signals according to claim 1, wherein: In step S2, the short-time fractional Fourier transform of the signal uses a sliding window to perform spectrum truncation calculation on the signal to obtain the relationship between time and frequency, which is calculated as follows: Where: t represents the time variable, u represents the variable in the fractional Fourier transform domain, α represents the fractional order, ω(t) represents the window function, K α (u,β) represents the kernel function; the kernel function is: Where: Φ = απ / 2 represents the rotation angle.
4. The method for detecting, segmenting and identifying underwater acoustic communication signals according to claim 1, wherein: The morphological processing flow of step S3 is as follows: Step S31: Convert the short-time fractional-order time-frequency feature S(f,τ) into a two-dimensional image matrix V(f,τ). The format of the two-dimensional image matrix V(f,τ) is: Where: Indicates that the frequency at time t0 is f L Signal strength of the communication signal; Step S32: Calculate the local density in the neighborhood of the pixel point of the two-dimensional image matrix V(f, t), and obtain the optimal scale parameter through mapping according to the maximum scale of the structure element; Step S33: Perform a morphological opening operation on the two-dimensional image matrix V(f,τ) according to the optimal scale parameter. The morphological opening operation expression is: Where: B represents the structural element, that is, the core in the image corrosion and expansion; Θ represents the corrosion operation; represents the dilation operation; Indicates open operation; Step S34: Perform a morphological closing operation on the matrix V'(f,τ). The result of the morphological closing operation is: Where: V"(f,τ) represents the matrix after morphological closing operation, which is used to obtain the filtered time-frequency map; Preferably, in step S3, the Hough transform method is used for morphological detection, and the Canny operator can be used to extract edges from the feature map output by the adaptive morphological filter to obtain edge pixels in the image. Then, the edge points are converted to the Hough space according to the Hough transform. The transformation relationship is expressed as: ρ=x cosθ+y sinθ (12) Where: ρ represents the vertical distance from the origin to the straight line, θ represents the angle between the perpendicular line of the straight line and the horizontal axis in the counterclockwise direction; Quantize the polar coordinate θ-P space into M×N L The small grid forms an accumulator space; the accumulator space is expressed as: Where: x max and y max Respectively represent the maximum coordinate values of the pixel points in the image in the x and y directions, M represents the number of discrete points in the θ-P space along the θ direction, and N L Represents the number of discrete points along the P direction in the θ-P space; based on the coordinates (x, y) of each point in the rectangular coordinates, calculate each P value in the range of θ = 0° to 180° with a small grid step size. If the obtained value falls within a certain small grid, the accumulator counter of the small grid is increased by 1; after all the edge pixels in the rectangular coordinates are transformed, the (θP) value corresponding to the maximum value in the accumulator is selected as the desired line; Finally, the minimum length and maximum gap of the line segment are set, and the starting and ending points of the underwater acoustic communication signal are calculated after eliminating short line segments and noise, thus completing the detection and segmentation of the underwater acoustic communication signal.
5. The method for detecting, segmenting and identifying underwater acoustic communication signals according to claim 1, wherein: In step S4, the steps of extracting features are as follows: Step S41: Obtain the square power spectrum of the signal using the Welch method; the square power spectrum is expressed as: Where: J represents the number of segments for dividing the signal sequence, The power spectrum of the j-th segment signal can be expressed as follows: Where: L j Indicates the length of the j-th segment signal, L sig represents the length of the signal, w(i) represents the window function, and U represents the normalization factor; Step S42: Perform Fourier transform on the cyclic autocorrelation function to obtain a cyclic spectrum density function, generate a three-dimensional cyclic spectrum, and then convert it into a contour map to obtain the cyclic spectrum contour features of the underwater acoustic communication signal; the cyclic spectrum can be expressed as: Where: represents the cyclic autocorrelation function, as shown below: Where: represents the cycle frequency, t1 represents the time, and Δt represents the time interval.
6. The method for detecting, segmenting and identifying underwater acoustic communication signals according to claim 1, characterized in that: In step S5, the MF-MHSA-LNet model is specifically as follows: Step S51: Based on the lightweight neural network SqueezeNet, the feature maps output by Fire2, Fire5, and Fire9 are extracted. For an input size of 227×227×3, the output feature maps are sized 56×56×128, 28×28×256, and 14×14×512, corresponding to the shallow features, intermediate transition features, and deep features of the image, respectively. These serve as the multi-scale feature inputs for the multi-head self-attention module. The SqueezeNet model consists of two convolutional layers, eight Fire modules, and one Softmax layer. The Fire module consists of compression and expansion layers, using a large number of 1×1 convolution kernels and a mixture of 1×1 and 3×3 convolution kernels instead of 3×3 convolution kernels. Ignoring the bias, the expansion layer reduces the network parameter count by 9 times. The parameter count of the expansion layer is shown in the following formula: Y Squeeze =Con×Num×3×3(18) Where: Con represents the number of convolution kernels, Num represents the number of input channels; Step S52: The multi-head self-attention module uses m with different projection matrices h Head, input feature X MHSA Mapped to different vector spaces, transformed into m h Group query vector Key Vector Sum value vector The m obtained by the splicing operation is sent to the attention pool in parallel for calculation. h Different attention outputs. The results of the attention output are: Where: Represents m h The learnable projection matrix of each head; d a =d / m h Indicates the size of each attention head; Using learnable linear projections right Perform dimension transformation to obtain attention representation As shown in the following formula: Where: Concat[.] represents the concatenation operation; Step S53: The attention representation of the MHSA module Input the fully connected classifier for recognition, thereby completing the recognition of the modulation mode of the underwater acoustic communication signal.