A training method based on time-frequency cross information fusion neural network
By constructing a time-frequency cross-information fusion neural network and combining simulation and experimental data, the accuracy and robustness problems of traditional underwater target detection in low signal-to-noise ratio environments are solved, and efficient underwater target identification is achieved.
Patent Information
- Application Number
- CN202510334887.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-20
- Publication Date
- 2025-11-21
- Estimated Expiration
- 2045-03-20
AI Technical Summary
Traditional underwater target detection methods suffer from low accuracy and poor robustness in low signal-to-noise ratio environments. Existing deep learning models lack time-frequency interaction mechanisms, resulting in insufficient ability to express weak line spectrum features. Furthermore, dataset construction is costly and difficult to generalize.
A neural network based on time-frequency cross-information fusion is constructed, including time-domain and frequency-domain feature extraction branches. Through cross-convolution and feature fusion classification branches, the network is trained using simulated and measured data to extract time-frequency features and perform probability classification.
Achieving high accuracy in underwater target detection under low signal-to-noise ratio conditions improves the model's generalization ability to non-stationary noise, accurately identifies sparsely distributed harmonic components, reduces reliance on artificial features, adapts to background noise in different sea areas, and enables transfer learning.
Smart Images

Figure CN120297349B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application belongs to the technical field of underwater detection, and more particularly relates to a training method based on a time-frequency cross information fusion neural network. BACKGROUND
[0002] Traditional underwater target detection mainly relies on acoustic technology, but the seawater environment is complex and the target noise reduction technology is increasingly upgraded, so the acoustic detection effect is limited. The shaft frequency magnetic field signal has the characteristics of low frequency propagation slow attenuation, strong anti-interference, etc., and becomes an important research direction of non-acoustic detection. However, the traditional underwater target shaft frequency magnetic field detection method has significant limitations.
[0003] At the algorithm level, the traditional algorithm (such as matched filter, adaptive spectral line enhancement) relies on Gaussian noise assumption and accurate parameter tuning, and the performance decreases sharply in non-stationary, low signal-to-noise ratio environment. At the feature extraction level, the traditional frequency domain method (such as FFT) has low resolution for short-time non-stationary signals, and the harmonics are easily covered by noise. Although the wavelet transform can provide time-frequency localization information, it is restricted by the selection of basis functions and has high computational complexity. The traditional time domain method is sensitive to waveform distortion, and it is difficult to distinguish signal and noise transient fluctuations by statistical features. The existing deep learning model mainly uses single time domain or frequency domain input, lacks time-frequency interaction mechanism, and the expression ability of weak line spectrum features is insufficient. At the data set construction level, the cost of collecting real measured data is high, and it is difficult to construct a large-scale data set. The pure simulation data has deviation in propagation attenuation, noise distribution, etc. compared with the real signal, which leads to insufficient robustness and generalization ability of the existing method in low signal-to-noise ratio environment.
[0004] In view of the above problems, it is urgent to design a target detection network with high accuracy and strong robustness in low signal-to-noise ratio environment. SUMMARY
[0005] In view of the above defects or improvement needs of the prior art, the present application provides a training method based on a time-frequency cross information fusion neural network, which aims to solve the technical problems of low accuracy and poor robustness of existing underwater target detection.
[0006] To achieve the above purpose, according to one aspect of the present application, a training method based on a time-frequency cross information fusion neural network is provided, comprising:
[0007] S1: Construct a mixed data set including a positive sample set and a negative sample set; the positive sample set includes simulated non-stationary non-Gaussian noise data b in shallow water environment and real-time measured background magnetic field data c without target; the negative sample set includes: simulated non-stationary non-Gaussian noise data b in shallow water environment and simulated shaft frequency magnetic field data a when underwater target exists;
[0008] S2: training a time-frequency cross information fusion neural network with samples in the mixed dataset until convergence; wherein the time-frequency cross information fusion neural network comprises:
[0009] a time domain feature extraction branch, configured to extract a time domain feature TF of an input sample;
[0010] a frequency domain feature extraction branch, configured to perform an FFT operation on the input sample to obtain a frequency domain coding feature GF, and perform coding and convolution operations on the frequency domain coding feature GF to obtain a frequency domain feature FF;
[0011] a feature cross fusion classification branch, connected with the time domain feature extraction branch and the frequency domain feature extraction branch, configured to perform cross convolution operations on the time domain feature TF and the frequency domain feature FF to obtain a fusion feature CF; and then map the fusion feature CF to a low-dimensional feature space through a linear layer, and perform probability classification on the fusion feature CF in the low-dimensional space to obtain a signal category.
[0012] In one of the embodiments, the time domain feature extraction branch comprises, in sequence: an integer module, a residual convolution structure, and an average pooling layer;
[0013] The integer module is configured to perform format processing on an input sample to obtain a feature tensor;
[0014] The residual convolution structure is configured to perform feature extraction on the feature tensor while ensuring the dimension, to obtain a local-global collaborative feature;
[0015] The average pooling layer is configured to perform average pooling on the local-global collaborative feature to obtain the time domain feature TF.
[0016] In one of the embodiments, the residual convolution structure comprises, in sequence: a convolution module and a residual network constructed by a plurality of cascaded FFC basic modules.
[0017] In one of the embodiments, the frequency domain feature extraction branch comprises, in sequence:
[0018] an FFC module, configured to perform a time domain to frequency domain operation on an input sample to obtain a frequency domain feature sequence F;
[0019] a block module, configured to split the frequency domain feature sequence F to obtain a frequency domain feature sequence block P;
[0020] an encoding module, configured to encode the frequency domain feature sequence block P to obtain a frequency domain coding feature GF;
[0021] a convolution module, configured to perform convolution on the frequency domain coding feature GF to obtain a frequency domain spatial feature SF;
[0022] an average pooling layer, configured to compress and aggregate the frequency domain spatial feature SF to obtain the frequency domain feature FF.
[0023] In one of the embodiments, the encoding module sequentially comprises a position encoding module, a multi-head attention mechanism, a residual connection & layer normalization, a feedforward network, and a residual connection & layer normalization.
[0024] The position encoding module is configured to perform position encoding on the frequency domain feature sequence block P to obtain an initial frequency domain feature with position information.
[0025] The multi-head attention mechanism is configured to obtain a context-aware feature by dynamically assigning weights to different frequency bands of the initial frequency domain feature, so as to capture frequency domain time sequence dependency.
[0026] The residual connection & layer normalization is configured to process the context-aware feature to obtain a frequency domain global feature.
[0027] The feedforward network is configured to perform nonlinear feature enhancement feature expression on the frequency domain global feature to obtain a nonlinear feature.
[0028] The residual connection & layer normalization is configured to output the nonlinear feature as a final frequency domain encoding feature GF.
[0029] In one of the embodiments, the feature cross-fusion classification branch comprises:
[0030] A first convolution module is configured to perform convolution on the frequency domain feature FF to obtain a frequency domain refined feature.
[0031] A first GELU activation layer is connected with the first convolution module and configured to activate the frequency domain refined feature to obtain a frequency domain activated feature.
[0032] A second convolution module is configured to perform convolution on the time domain feature TF to obtain a time domain refined feature.
[0033] A second GELU activation layer is connected with the second convolution module and configured to activate the time domain refined feature to obtain a time domain activated feature.
[0034] A first Hadamard product calculator is connected with the first GELU activation layer and the second convolution module and configured to calculate a Hadamard product of the frequency domain activated feature and the time domain refined feature to obtain a first cross feature.
[0035] A second Hadamard product calculator is connected with the second GELU activation layer and the first convolution module and configured to calculate a Hadamard product of the time domain activated feature and the frequency domain refined feature to obtain a second cross feature.
[0036] a third convolution module, connected with the first Hadamard product calculator and the second Hadamard product calculator, configured to perform convolution on the superposition of the first cross feature and the second cross feature to obtain the fusion feature CF;
[0037] a classification module, connected with the third convolution module, configured to map the fusion feature CF to a low-dimensional feature space through a linear layer, and perform probability classification on the fusion feature CF in the low-dimensional feature space to obtain a signal category.
[0038] According to another aspect of the present application, a method for detecting underwater targets based on a time-frequency cross information fusion neural network is provided, comprising: inputting an axial frequency magnetic field signal corresponding to a current underwater environment into a time-frequency cross information fusion neural network trained to obtain a current category to determine whether there is a target underwater.
[0039] According to another aspect of the present application, a device for detecting underwater targets based on a time-frequency cross information fusion neural network is provided, comprising:
[0040] a construction module configured to construct a mixed data set comprising a positive sample set and a negative sample set; the positive sample set comprises simulated non-stationary non-Gaussian noise data b in a shallow sea environment and background magnetic field data c measured in real time when there is no target; the negative sample set comprises simulated non-stationary non-Gaussian noise data b in a shallow sea environment and simulated axial frequency magnetic field data a when there is an underwater target;
[0041] a training module configured to train a time-frequency cross information fusion neural network using samples in the mixed data set until convergence; wherein the time-frequency cross information fusion neural network comprises: a time domain feature extraction branch configured to extract time domain features TF of input samples; a frequency domain feature extraction branch configured to perform FFT operation on input samples to obtain frequency domain coding features GF, and perform coding and convolution operations on the frequency domain coding features GF to obtain frequency domain features FF; a feature cross fusion classification branch connected with the time domain feature extraction branch and the frequency domain feature extraction branch, configured to perform cross convolution operation on the time domain features TF and the frequency domain features FF to obtain fusion features CF; and then map the fusion features CF to a low-dimensional feature space through a linear layer, and perform probability classification on the fusion features CF in the low-dimensional feature space to obtain a signal category;
[0042] a recognition module configured to input an axial frequency magnetic field signal corresponding to a current underwater environment into a time-frequency cross information fusion neural network trained to obtain a current category to determine whether there is a target device underwater.
[0043] According to another aspect of the present application, there is provided an underwater target detection device based on a time-frequency cross information fusion neural network, comprising a memory and a processor, the memory storing a computer program, and the processor implementing the steps of the method when executing the computer program.
[0044] According to another aspect of the present application, there is provided a computer readable storage medium having a computer program stored thereon, the computer program being executed by a processor to implement the steps of the method.
[0045] Overall, compared with the prior art, the above technical solutions conceived by the present application can achieve the following beneficial effects:
[0046] (1) The present application provides a training method based on a time-frequency cross information fusion neural network. First, a positive sample set is constructed using simulated non-stationary non-Gaussian noise data b and real-time measured background magnetic field data c in a shallow water environment, and a negative sample set is constructed using simulated non-stationary non-Gaussian noise data b and simulated axial frequency magnetic field data a when an underwater target exists in a shallow water environment. This design has the advantage that Gaussian weighted noise simulates the low-frequency energy concentration characteristics of shallow water geomagnetic noise, forcing the model to learn to separate the narrowband harmonic of the target from the low-frequency interference, and the introduction of real-time noise improves the model's generalization ability to non-stationary noise, making the training data more similar to the non-Gaussian noise characteristics of the real scene. In addition, the use of a time-harmonic magnetic dipole to simulate the axial frequency magnetic field data solves the problem of a small number of data samples and difficulty in obtaining axial frequency magnetic field data. The time-frequency cross information fusion neural network includes a time-domain feature extraction branch, a frequency-domain feature extraction branch, and a feature cross fusion classification branch. This design has the advantage that the time-domain feature extraction branch extracts the local fluctuation changes of the time-domain envelope. The frequency-domain feature extraction branch extracts spatial features, enabling the model to accurately identify the sparse distribution of harmonic components in a noisy background. The feature cross fusion classification branch enhances the complementarity of the two types of features through bidirectional interaction, and this fusion mechanism ensures that the model only determines the target signal when the time-domain envelope and the frequency-domain harmonic exist simultaneously. The periodic weight provided by the time-domain branch enhances the frequency-domain harmonic feature, and the reverse verification of the frequency-domain branch excludes the time-domain artifacts caused by noise. Finally, the model achieves high-precision detection through bidirectional feature complementarity.
[0047] (2) The application provides a kind of underwater target detection method based on time-frequency cross information fusion neural network, first, training is obtained based on time-frequency cross information fusion neural network;Then, the axis frequency magnetic field signal corresponding to current underwater environment is input into the based time-frequency cross information fusion neural network obtained by training, to obtain current category, to determine whether there is target equipment under water.Such design, the advantage is, axis frequency magnetic field signal contains both transient time domain characteristics (waveform change), and stable frequency domain characteristics (axis frequency fundamental frequency and harmonic).Traditional method is often analyzed separately in time domain or frequency domain information, and time-frequency cross fusion can capture the global characteristics and local details of signal.Underwater magnetic field signal is susceptible to geomagnetic background noise, marine biological activity or artificial interference (such as cable).Time-frequency fusion model can filter out high-frequency random noise by using frequency domain analysis, and identify continuous target signal by combining time domain waveform stability.End-to-end network adaptive learning can reduce the dependence of artificial features, and automatic feature extraction avoids human bias, and migration learning can also be realized by adding background noise of different sea areas. BRIEF DESCRIPTION OF DRAWINGS
[0048] Figure 1 The time-frequency information cross fusion neural network provided for embodiment 1 of the application is provided with a whole flow chart;
[0049] Figure 2 The time-frequency information cross fusion neural network provided for embodiment 1 of the application is provided with a whole flow chart;
[0050] Figure 3 The time-frequency information cross fusion neural network provided for embodiment 1 of the application is provided with a whole flow chart;
[0051] Figure 4 The time-frequency information cross fusion neural network provided for embodiment 1 of the application is provided with a whole flow chart;
[0052] Figure 5 The time-frequency information cross fusion neural network provided for embodiment 1 of the application is provided with a whole flow chart;
[0053] Figure 6 The time-frequency information cross fusion neural network provided for embodiment 1 of the application is provided with a whole flow chart;
[0054] Figure 7 The time-frequency information cross fusion neural network provided for embodiment 1 of the application is provided with a whole flow chart;
[0055] Figure 8 The time-frequency information cross fusion neural network provided for embodiment 1 of the application is provided with a whole flow chart;
[0056] Figure 9A comparison chart of the results of TFSRM-Net provided for embodiment 1 of the present application and the traditional algorithm Duffing;
[0057] Figure 10 A comparison chart of the results of TFSRM-Net provided for embodiment 1 of the present application and other deep learning methods in different signal-to-noise ratio simulation test sets;
[0058] Figure 11 A comparison chart of the results of TFSRM-Net provided for embodiment 1 of the present application and other deep learning methods in different data sets;
[0059] Figure 12 A comparison table of the ablation experiment results of the simulation test set and the measured data set provided for embodiment 1 of the present application. DETAILED DESCRIPTION
[0060] In order to make the purpose, technical scheme and advantages of the present application clearer and more comprehensible, the present application will be further described in detail below in combination with the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and do not limit the present application. In addition, the technical features involved in each embodiment of the present application described below can be combined with each other as long as they do not conflict with each other.
[0061] Embodiment 1
[0062] The present embodiment provides a training method based on a time-frequency cross information fusion neural network, comprising: S1: constructing a mixed data set comprising a positive sample set and a negative sample set; the positive sample set comprises simulated non-stationary non-Gaussian noise data b in shallow water environment and background magnetic field data c measured in real time without target. The negative sample set comprises: simulated non-stationary non-Gaussian noise data b in shallow water environment and simulated axial frequency magnetic field data a when an underwater target exists. S2: training the time-frequency cross information fusion neural network using the samples in the mixed data set until convergence.
[0063] Since the underwater target axial frequency signal is obtained by using the simulation signal method, and the geomagnetic environmental noise without target is obtained by using the simulation combined with a small amount of measured method. In order to express more conveniently subsequently, the underwater target axial frequency signal contained in the geomagnetic environmental noise is called target signal, and the pure geomagnetic environmental noise signal is called noise signal. It is worth mentioning that the signals in the final input network are all normalized, and the noise signals have been operated such as detrending. The label of the target signal is set to 0, and the label of the noise signal is set to 1. In order to obtain a large number of training samples required by the network, a sufficient number of target signals and noise signals need to be obtained first.
[0064] The target signal refers to an underwater target axis frequency signal containing geomagnetic background noise. Considering the difficulty of obtaining a large number of real target signals, the target signals are simulated. The axis frequency magnetic field signal of the underwater target decays faster, the speed is low, and the navigation depth is deep. Assuming that the underwater target is a single propeller mode, the frequency spectrum characteristics are set to the fundamental frequency and its multiples within 1-7 Hz. A large number of target sample signals are generated by setting different magnetic moment sizes, target motion speeds, target depths, target frequencies, and multiple numbers. Figure 2 Fig. 1 shows a time-frequency domain image of a raw simulated axis frequency magnetic field target signal provided according to an embodiment of the present embodiment. Figure 2 The parameter settings of the given signal are as follows: sampling rate 2000 Hz, magnetic moment size 30-200 A·m 2 , target speed 2-10 m / s, target depth 50-300 m, target fundamental frequency 1-7 Hz, target multiple number 3-7, and CPA 100 m. The raw simulated axis frequency magnetic field target signal obtained by simulation is shown in Figure 2 , assuming that the fundamental frequency is 1.8 Hz, the multiple number is 5, the depth is 50 m, the motion speed is 8 m / s, and the total magnetic moment is set to 100 A·m 2 .
[0065] Figure 3 Fig. 2 shows a time-frequency domain image of weighted Gaussian noise provided according to an embodiment of the present embodiment. Most of the noise signals are obtained by Gaussian weighted simulation, and the remaining small part (about 2.2%) is obtained by collecting the geomagnetic background field in the real world 5 days before the test, and then using a band-pass filter to filter the trend item. The magnetic field noise in the actual environment, such as geomagnetic noise, often presents strong non-Gaussian characteristics and is mostly colored noise. The present embodiment uses weighted Gaussian noise (Weighted Gaussian Noise) to obtain simulated environmental background magnetic field noise, as shown in Figure 3 .
[0066] Figure 4 Fig. 3 shows the time domain graph and frequency spectrum graph of the simulated target signal after adding 0 dB noise and -25 dB noise to the raw simulated axis frequency magnetic field target signal, and the wavelet transform result thereof. The real axis frequency magnetic field data is collected from an inductive magnetic sensor, which contains various noises, especially large low-frequency noise. Therefore, the noise signal simulated by the Gaussian weighted algorithm is superimposed on the simulated target signal, which can be used to simulate the real situation under different signal-to-noise ratios. The calculation formula of the signal-to-noise ratio is shown in the following formula (1), wherein s(i) and n(i) represent the sampling values of the signal and the noise, respectively:
[0067]
[0068] Figure 4The frequency component results in (a) indicate that the fundamental frequency is 1.8 Hz and the waveform contains a 5th harmonic component. However, from... Figure 4 In (b) of the original waveform and frequency analysis, it is difficult to distinguish detailed harmonic components. Wavelet transform is a multi-resolution analysis method suitable for processing non-stationary signals. Compared with time-domain or frequency-domain plots, wavelet transform can provide both time and frequency information simultaneously, and can enhance weak signals by selecting appropriate wavelet bases. Figure 4 (c) Figure 4 (d) in the text represents the corresponding Figure 4 (a) and Figure 4 The wavelet transform diagram of the signal in (b) shows the frequency components of the signal at the corresponding time and the intensity of that frequency. The brighter the color, the higher the signal intensity at the corresponding frequency. It can be seen that after adding -25dB of noise, the energy of the target signal is so low that it cannot be distinguished by the naked eye.
[0069] By randomly changing the values of various parameters in the simulated axial frequency magnetic field signal, 10,000 target signal samples and 10,000 noise signal samples were generated, totaling 20,000 data points. The original 30,000 sampling points were downsampled, resulting in each sample having a length of 5,000 points. Noise was then superimposed on the normalized signal to ensure the signal-to-noise ratio (SNR) of the noise-laden signal was uniformly distributed between 0 dB and -25 dB. The 20,000 data points were split into training, validation, and test sets in a 7:2:1 ratio. Since the target signal frequency increased in steps of 0.1 during network training, a separate test set (sets 2-7) with a single SNR was established. The target signal frequency changed in steps of 0.03, and all noise signals were simulated, thus simultaneously testing the network's robustness under different SNRs.
[0070] The neural network based on time-frequency cross-information fusion includes: a time-domain feature extraction branch, used to extract the time-domain features TF of the input sample; a frequency-domain feature extraction branch, used to perform FFT operation on the input sample to obtain the frequency-domain encoded features GF, and to encode and convolve the frequency-domain encoded features GF to obtain the frequency-domain features FF; a feature cross-fusion classification branch, connected to the time-domain feature extraction branch and the frequency-domain feature extraction branch, used to perform cross-convolution operation on the time-domain features TF and the frequency-domain features FF to obtain the fusion features CF; then the fusion features CF are mapped to a low-dimensional feature space through a linear layer, and the fusion features CF in the low-dimensional space are probabilistically classified to obtain the signal category.
[0071] The time domain feature extraction branch uses a residual convolution structure instead of repeated stacking of convolution layers, and the FFC basic module therein not only uses traditional convolution to extract local information, but also can obtain global context information of the axis frequency magnetic field time domain signal through fast Fourier convolution (FFC), thereby enhancing the learning ability. The input of the time domain feature extraction branch is an axis frequency magnetic field or background magnetic field signal sequence B of length L after simple preprocessing such as normalization of the collected signal total , the length is 5000 points, in order to facilitate calculation, 4096 points are intercepted, and m patches are divided, each patch has a length of m and a channel number of 1, and the feature tensor T is transformed to [batch, channels, m, m], T ∈ R B×C×m×m , thereby reducing the computational complexity and retaining the local time domain correlation.
[0072] The frequency domain feature extraction branch uses a combination of Transformer and CNN. In the frequency domain, the time sequence correlation of the attention frequency points is embedded in multiple blocks, the Transformer dynamically allocates weights to different frequency bands, and the CNN extracts the spatial correlation of the frequency sequence. For example, the local time domain features and the global frequency domain features are complemented through parallel processing. A 3×3 standard convolution kernel is used to extract local waveform details, and a batch normalization layer BN and a ReLU activation function are used to enhance the non-linear expression ability. Meanwhile, the input signal is converted to the frequency domain by using a Spectral Transform module, the real and imaginary parts are respectively subjected to convolution operation in the frequency domain, the global periodicity features of the signal are captured, the time domain signal is reconstructed through inverse Fourier transform, and finally the local and global branch outputs are weighted and fused according to the learnable weights, so as to realize the joint representation of local details and global outlines.
[0073] The process involves time-frequency domain feature fusion and classification: the feature results extracted from the two branches are cross-fused and input into a classifier, which predicts the category by mapping the fused features to the label space. For example, a deep feature extraction network is constructed based on the ResNet18 residual network architecture, with its core consisting of cascaded FFC basic modules 1-4. The first layer of FFCResNet extracts primary features through 3×3 convolutions, BN, and ReLU. The second layer further refines the features while retaining the input dimension. The input signal passes sequentially through these four FFC basic modules, gradually abstracting the Local-Global Co-operational Feature (LGF). Each FFC basic module contains two FFC modules. The LGF output from the last residual block is subjected to spatial dimension average pooling to compress redundant information and retain key features, yielding the temporal feature (TF). A linear layer then maps the high-dimensional feature vector to a low-dimensional space, finally outputting the temporal branch feature (TF) for subsequent cross-fusion modules.
[0074] As an optional implementation, the temporal feature extraction branch includes, in sequence, an integer module, a residual convolutional structure, and an average pooling layer. The integer module is used to format the input samples to obtain a feature tensor; the residual convolutional structure is used to extract features while maintaining the dimensionality of the feature tensor, obtaining local and global collaborative features; the average pooling layer is used to perform average pooling on the local and global collaborative features to obtain temporal features (TF).
[0075] The input to the time-domain feature extraction branch is a sequence of axis frequency magnetic field or background magnetic field signals of length L, which has undergone only simple preprocessing such as normalization. total The feature tensor has a length of 5000 points. For ease of calculation, 4096 points are truncated and divided into m patches. Each patch has a length of m and 1 channel. The feature tensor T is transformed to have dimensions [batch, channels, m, m], where T ∈ R. B×C×m×m This reduces computational complexity while preserving local temporal correlations. The baseline of this branch uses the ResNet residual network structure, combined with Fast Fourier Convolution to form FFCResNet. The network's input features are divided into local and global parts. Standard convolutions are used to process local features, while Spectral Transform is used to process global features. A fusion operation then integrates the local and global features. Residual connections are used to ensure the training stability of deep networks, resulting in powerful feature representation capabilities.
[0076] Figure 5The diagram shown is a BasicBlock structure diagram of FCCResNet in the temporal feature extraction branch provided by the embodiment of this invention. FCCResNet is a deep neural network composed of multiple modules organized in layers, such as... Figure 1 As shown, the reshaped temporal features T are first subjected to a 3×3 two-dimensional convolution (including BN and ReLU), and then fed into the residual unit, the core of which is the BasicBlock. The residual block composed of the BasicBlock structure is as follows: Figure 5 As shown, each BasicBlock contains two core FFC-based convolutional modules that perform local-to-global feature transformation and fusion. It also includes a batch normalization layer (BatchNorm) and a ReLU activation function, a downsampling layer that matches the input and fuses residual information, and the final activation functions are used for the outputs of the local and global branches, respectively. The SpectralTransform module, following Conv, performs frequency domain transformation, containing convolutional layers and FFT and iFFT operations. It's primarily a network module combining Fourier transform and convolution, enhancing the network's receptive field and high-frequency information capture capability by fusing global and local features. The FU (FourierUnit) and LFU (LocalFourierUnit) modules perform Fast Fourier Transform followed by convolution to process the real and imaginary features, then transform the processed frequency domain features back to the spatial domain using Inverse Fourier Transform (iFFT). LFU performs frequency domain transformation on small local regions, suitable for detailed features. After average pooling and reshaping, the output is the temporal features extracted by that branch. Where T dim This represents the number of filters in the last convolutional layer.
[0077] As an optional implementation, the residual convolutional structure includes a residual network constructed by a convolutional module and multiple cascaded FFC base modules connected in sequence.
[0078] As an optional implementation, the frequency domain feature extraction branch includes the following components connected in sequence: an FFC module, used to perform time-domain to frequency-domain transformation on the input samples to obtain a frequency domain feature sequence F; a block segmentation module, used to segment the frequency domain feature sequence F to obtain a frequency domain feature sequence block P; an encoding module, used to encode the frequency domain feature sequence block P to obtain a frequency domain encoded feature GF; a convolution module, used to convolve the frequency domain encoded feature GF to obtain a frequency domain spatial feature SF; and an average pooling layer, used to compress and aggregate the frequency domain spatial feature SF to obtain frequency domain features.
[0079] As an optional implementation, the Transformer encoding module includes, in sequence: a position encoding module, a multi-head attention mechanism, residual connections & layer normalization, a feedforward network, and residual connections & layer normalization. The position encoding module is used to perform position encoding on the frequency domain feature sequence block P to obtain initial frequency domain features with positional information; the multi-head attention mechanism is used to obtain context-aware features by dynamically assigning weights to different frequency bands of the initial frequency domain features to capture frequency domain temporal dependencies; the residual connections & layer normalization is used to process the context-aware features to obtain global frequency domain features; the feedforward network is used to perform nonlinear feature enhancement on the global frequency domain features to obtain nonlinear features; and the residual connections & layer normalization is used to output the final frequency domain encoded feature GF from the nonlinear features.
[0080] The first step in the frequency domain feature extraction branch is to perform a Fast Fourier Transform (FFT) on the sequence. The Discrete Fourier Transform (DFT) is the discrete Fourier transform of a finite-length sequence, and the Fast Fourier Transform (FFT) is an efficient and fast algorithm for the DFT. Given a discrete-time sequence x(n) of length N, where 0 ≤ n ≤ N-1, the DFT can convert the sequence into its frequency domain representation.
[0081]
[0082] Where j represents the imaginary unit, this formula is obtained by discretizing the sequence x(n) in the time and frequency domains based on the continuous Fourier transform. k The spectrum at k / N is denoted by X(k), and only the first N points need to be considered.
[0083] Let W N =e^(-j(2π / N)), Fast Fourier Transform (FFT) utilizes The symmetry and periodicity of the DFT reduce the computational cost from O(N) to O(N). 2 The optimization is reduced to O(NlogN). We obtain the frequency domain representation X(k) of the time series by performing FFT, as shown in Equation 2. Similarly, given a sequence B of length L after normalization and other preprocessing... total Its frequency domain expression is calculated as follows:
[0084]
[0085] In the formula, This represents a one-dimensional FFT operation, where L' is the length of the transformed frequency sequence, L' = L / 2 + 1. Each channel of the time series in each batch is transformed independently to obtain the frequency domain features F of the original time series.
[0086] The frequency domain feature sequence F is divided into M groups: patterns {P1, P2, ..., P...} M}, each block P i A segment F containing no overlap, where the length of each patch is predefined as S, such that each patch... Obtain spectral sequence blocks
[0087] This embodiment employs the Transformer Encoder architecture. Considering the signal variations in the frequency domain data, the Transformer Encoder can dynamically allocate weights. The segmented spectral sequence block P is fed into the Transformer Encoder, and the resulting output is the encoded feature GF∈R. B×M×S The sequence information of the frequency domain data is extracted. Then, a one-dimensional convolutional module is used to extract the spatial information of the frequency domain data. This module contains two stacked CNN modules, denoted as Conv1 and Conv2. Each CNN module consists of two one-dimensional convolutional layers, one max-pooling layer, and batch normalization, similar to the VGG network. Figure 6 As shown. Figure 6 The diagram shows the CNN module structure in the TransformerEncoder within the frequency domain feature extraction branch provided in this embodiment. Each convolutional layer uses the same stride, kernel size, and ReLU activation function, but the number of filters increases with depth, with the last convolutional layer having F filters. dim Finally, sequence average pooling and reshape operations are used to output the frequency domain features extracted by this branch.
[0088] As an optional implementation, the feature cross-fusion classification branch includes: a first convolutional module for convolving the frequency domain feature TF to obtain a frequency domain refined feature; a first GELU activation layer connected to the first convolutional module for activating the frequency domain refined feature to obtain a frequency domain activation feature; a second convolutional module for convolving the time domain feature CF to obtain a time domain refined feature; a second GELU activation layer connected to the second convolutional module for activating the time domain refined feature to obtain a time domain activation feature; and a first Hadamard product calculator connected to the first GELU activation layer and the second convolutional module for calculating the Hadamard product of the frequency domain activation feature and the time domain refined feature. The first cross feature is obtained by multiplying the first cross feature; the second Hadamard product calculator, connected to the second GELU activation layer and the first convolution module, is used to calculate the Hadamard product of the temporal activation feature and the frequency domain refinement feature to obtain the second cross feature; the third convolution module, connected to the first Hadamard product calculator and the second Hadamard product calculator, is used to convolve the superposition of the first cross feature and the second cross feature to obtain the fused feature CF; the classification module, connected to the third convolution module, is used to map the fused feature CF to a low-dimensional feature space through a linear layer, and to perform probability classification on the fused feature CF in the low-dimensional space to obtain the signal category.
[0089] The output of the temporal feature extraction branch is TF, and the output of the frequency domain feature extraction branch is FF, which serve as the input to the feature fusion module. The intermediate process variables are the cross features CF1 and CF2, and the calculation process is as follows, where Conv1(·) and Conv2(·) are two one-dimensional convolutional layers, and φ is the GELU activation function:
[0090] CF1=φ(Conv1(TF))⊙Conv2(FF) (4)
[0091] CF2=φ(Conv2(FF))⊙Conv1(TF) (5)
[0092] The activated features are summed and then restored to the original summed size by the last convolutional layer, Conv3(·).
[0093] CF = Conv3(CF1 + CF2) (6)
[0094] The output CF of the feature fusion module is the cross-fused feature.
[0095] Does the classification and prediction target exist in S42? The output CF of the feature fusion module is the cross-fused feature, which is also the input of the last linear classifier layer of the network:
[0096] predict = CF × W + bias (7)
[0097] Here, `predict` represents the label predicted by the network, `W` is the parameter the model wants to learn, and `bias` is the vector bias. During the training phase, the `LabelSmoothingcross-entropy` function is used to update the parameters.
[0098]
[0099] In the formula p(x i The expression represents the probability of the predicted target's existence after processing with the softmax function, and N represents the number of samples. LabelSmooth is a regularization technique used to prevent the model from becoming overconfident in classification problems and improve its generalization ability. It introduces noise by converting one-hot labels to soft labels, making the model learn more smoothly and thus reducing overfitting.
[0100]
[0101] In the formula y i ε is the result of LabelSmoothing the true label of the i-th decision window, where ε is a very small constant representing the value of smoothing the label, n represents the number of categories, and target represents the current category.
[0102] Example 2
[0103] This embodiment provides an underwater target detection method based on a time-frequency cross-information fusion neural network, including: inputting the axis frequency magnetic field signal corresponding to the current underwater environment into the trained time-frequency cross-information fusion neural network to obtain the current category, so as to determine whether there is a target underwater.
[0104] This invention discloses an underwater target detection method based on a time-frequency cross-information fusion neural network, which can accurately detect axial frequency magnetic field target signals from samples with low signal-to-noise ratios. First, to address the problem of limited and difficult-to-obtain data samples, a magnetic dipole simulation is used in MATLAB to generate axial frequency magnetic field signals from underwater targets. A Gaussian weighted algorithm is used to simulate the geomagnetic background field, and a small number of measured geomagnetic background field signals from the same location at different times are combined to construct a reliable axial frequency magnetic field dataset for deep learning. Then, the proposed time-frequency cross-information fusion axial frequency magnetic field detection network (TFSRM-Net) is used to extract the time-frequency fusion features of the signal through a neural network, thereby detecting the target's axial frequency magnetic field signal under low signal-to-noise ratio conditions. This invention solves the technical bottlenecks of difficult data acquisition and weak feature extraction at low signal-to-noise ratios in underwater target axial frequency magnetic field detection, providing an efficient and reliable solution for marine safety monitoring and underwater target identification.
[0105] To demonstrate the superior performance of this invention in detecting the axial frequency magnetic field of underwater targets, the performance of this method is presented and analyzed using field test data. Figure 7 The diagram shown is a flowchart of the training and testing process for an experiment provided according to an embodiment of the present invention.
[0106] To verify the proposed network's detection capability in a real-world environment, field experiments were conducted to obtain measured axial frequency magnetic field signals, thus creating a measured dataset. The experiment used a low-frequency copper wire winding coil to simulate the axial frequency signal of an underwater target. A test boat was used to transport the coil at different speeds and distances to simulate the axial frequency magnetic field signal generated by the underwater target during navigation. Background geomagnetic field data was collected on December 22, 2023, and incorporated into the dataset used for network training. The formal axial frequency magnetic field detection experiment was conducted from December 26th to 27th, producing the measured dataset. For the measured signals, the original measured time series data was first cleaned, removing outliers and subtracting the mean to remove DC signal interference. A bandpass filter (0.001–50 Hz) was then used to filter out diurnal variations in the geomagnetic field, geological noise, power frequency noise, and other unwanted electromagnetic interference. The data was then divided into non-overlapping segments with 5000 sequence points each. Finally, the data was normalized, with each segment representing a signal sample.
[0107] Next is the network training setup. For our proposed network TFSRM-Net, the He method is used to initialize some network parameters, Adam is used as the optimizer, a smooth loss function is used, the learning rate is set to 0.00001, and the network is trained for 50 epochs with early stopping mechanism. The batch size for each epoch is 32. The model with the best performance in the validation set is selected and saved after sufficient training for subsequent testing and comparison.
[0108] To facilitate the description of performance metrics, four machine learning terms are introduced for detection results: (a) Truepositives (TP), meaning the sample data contains a target signal and is correctly classified; (b) Truenegatives (TN), meaning the sample data contains only noise and is correctly classified; (c) Falsepositives (FP), meaning the sample data contains only noise and is incorrectly classified as a target signal; and (d) Falsenegatives (FN), meaning the sample data contains a target signal but is incorrectly classified as noise.
[0109] Therefore, detection accuracy (DA) is used to evaluate the detection performance of a signal. It is defined as the proportion of correctly detected target signals out of all target signal samples. Its value is the same as Recall. The higher this value, the better the network performance.
[0110]
[0111] The phenomenon of mistaking noise signals for target signals can be defined using the Probability of False Alarm (PFA) as the ratio of the number of signals incorrectly identified as target signals to the total number of noise signals. A lower PFA indicates better network performance.
[0112]
[0113] In addition, the accuracy obtained from the classification report represents the ratio of correctly classified target signals and noise signals to all samples, the precision represents the proportion of correctly detected target signals to the total number of samples classified as target signals by the network, and the F1 score obtained from the classification report is the harmonic mean of Recall and Precision.
[0114] The feature vectors output by the network are visualized. A t-distributed stochastic neighbor embedding (t-SNE) plot is used to compress the feature vectors, and the differences between the mapping results of the target signal and the noise signal are compared to verify whether the feature extraction results of different types of signals are effectively clustered in the target space. The visualization results are as follows: Figure 8 (a) and Figure 8 As shown in (b) in the figure, this is a classification result diagram of the simulation test set and the actual test set before and after network training according to an embodiment of the present invention. Figure 8 The t-SNE method was used to plot and reduce the dimensionality of the classification results on the simulation test set before and after network training. Figure 8 (c) and Figure 8 In section (d), the t-SNE method is used to plot and reduce the dimensionality of the classification results on the experimental test set before and after network training, proving that the network can effectively separate the two types of signals in the target space. Therefore, the axial frequency magnetic field target signal can be identified based on the obvious difference in the feature vectors.
[0115] To demonstrate the superior performance of this invention in detecting the axial frequency magnetic field of underwater targets, a comparative experiment with other detection methods is conducted to showcase and analyze the performance of this method.
[0116] The trained TFSRM-Net model is compared with the traditional Duffing detection algorithm. The public test set for the comparison experiment is a test set of simulation datasets with different signal-to-noise ratios. The best-performing TFSRM-Net trained model and the Duffing algorithm are tested on the test set data respectively, and the detection results of TFSRM-Net and Duffing are obtained, as shown below.Figure 9 As shown. Figure 9 The figure shown is a comparison of the results of TFSRM-Net provided by the embodiment of the present invention and the traditional algorithm Duffing. The results show that TFSRM-Net is significantly better than the traditional Duffing detection method.
[0117] To verify the robustness of the proposed method, this section conducts comparative experiments with different mainstream time series analysis models. The performance of the proposed method is compared with classic methods such as CNN, CNN variants MC-DCNN, CNN-LSTM, ResNet, and InceptionTime. Each network was trained for 50 epochs and achieved its optimal performance for that model. The performance of each network model on the simulation test set is shown below. Figure 10 As shown, Figure 10 The figure shows a comparison of simulation test results between TFSRM-Net and other deep learning methods under different signal-to-noise ratios (SNRs) according to embodiments of the present invention. The results show that ResNet performs the worst under the same SNR, while the TFSRM-Net proposed in this paper performs the best. Next, the impact of different SNRs on the network is explored. It can be seen that at low SNRs, due to the diminished signal features and limited deep learning capabilities, the detection performance of other networks drops rapidly at -25dB. The detection model proposed in this paper still performs well, maintaining a high accuracy rate.
[0118] The performance of each network model on the simulation test set and the experimental dataset is as follows: Figure 11 As shown, Figure 11 The table shown is a comparison of the results of TFSRM-Net provided according to embodiments of the present invention with other deep learning methods on different datasets. Acc-1, DA-1, and PFA-1 represent the accuracy, detection accuracy, and false alarm rate on the simulation test set, respectively. Similarly, Acc-2, DA-2, and PFA-2 represent the accuracy, detection accuracy, and false alarm rate on the actual test dataset, respectively. Higher accuracy and detection accuracy, and lower false alarm rate indicate better detection performance of the network. It can be seen that TFSRM-Net proposed in this paper is generally superior to other deep learning methods.
[0119] Ablation experiments were conducted on both the simulation test set and the experimental dataset to demonstrate the effectiveness of the time- and frequency-domain fusion features used by the network and the effectiveness of the feature fusion scheme. This section presents probing experiments on both the simulation test set and the experimental dataset from the field. Figure 12 The table shown is a comparison of ablation experiment results between the simulation test set and the actual test set provided according to an embodiment of the present invention. Figure 12It can be seen that when using a single feature for detection, the detection accuracy of the time-domain feature module classifier is lower than that of the frequency-domain feature module-based classifier. However, mixing the two features can improve the detection accuracy of the model. Furthermore, it can be seen that using the cross-fusion method improves the detection accuracy by 6.25% compared to not using it. In summary, the results of the ablation experiments confirm that our time- and frequency-domain cross-fusion features can bring a stable and significant improvement to the detection of axial frequency magnetic field target signals.
[0120] In summary, this invention can effectively achieve underwater axial frequency magnetic field detection through time-frequency cross-information fusion neural network. This algorithm effectively solves the problems of lack of datasets in deep learning algorithms for underwater target axial frequency magnetic field detection and the complexity of parameter tuning in traditional algorithms. TFSRM-Net has good detection performance in both low signal-to-noise ratio synthesized signals and field measured signals, achieving an accuracy of 97.92% on field measured signals.
[0121] Example 3
[0122] This embodiment provides an underwater target detection device based on a time-frequency cross-information fusion neural network, including a construction module and a training module. The construction module is used to construct a hybrid dataset including a positive sample set and a negative sample set; the positive sample set includes simulated non-stationary, non-Gaussian noise data b in a shallow sea environment and measured background magnetic field data c when there is no target; the negative sample set includes simulated non-stationary, non-Gaussian noise data b in a shallow sea environment and simulated axial frequency magnetic field data a when an underwater target is present.
[0123] The training module is used to train the time-frequency cross-information fusion neural network using samples from the mixed dataset until convergence. The time-frequency cross-information fusion neural network includes: a time-domain feature extraction branch, used to extract the time-domain features (TF) of the input samples; a frequency-domain feature extraction branch, used to perform FFT operations on the input samples to obtain frequency-domain encoded features (GF), and then encode and convolve the frequency-domain encoded features (GF) to obtain frequency-domain features (FF); a feature cross-fusion classification branch, connected to the time-domain feature extraction branch and the frequency-domain feature extraction branch, used to perform cross-convolution operations on the time-domain features (TF) and the frequency-domain features (FF) to obtain fusion features (CF); then, the fusion features (CF) are mapped to a low-dimensional feature space through a linear layer, and the fusion features (CF) in the low-dimensional space are probabilistically classified to obtain the signal category.
[0124] The identification module is used to input the shaft frequency magnetic field signal corresponding to the current underwater environment into the trained time-frequency cross-information fusion neural network to obtain the current category, so as to determine whether there is a target device underwater.
[0125] Example 4
[0126] This embodiment provides an underwater target detection device based on a time-frequency cross-information fusion neural network, including a memory and a processor. The memory stores a computer program, and the processor executes the steps of the method implemented by the computer program.
[0127] Example 5
[0128] This embodiment provides a computer-readable storage medium storing a computer program thereon, the steps of a method implemented when the computer program is executed by a processor.
[0129] Those skilled in the art will readily understand that the above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A training method of a time-frequency cross information fusion neural network, characterized in that, Comprise: S1: Construct a mixed data set comprising a positive sample set and a negative sample set; the positive sample set comprises simulated non-stationary non-Gaussian noise data b in shallow water environment and real-time measured background magnetic field data c without target; The negative sample set comprises: simulated non-stationary non-Gaussian noise data b in shallow water environment and simulated axial frequency magnetic field data a when underwater target exists; S2: Train the time-frequency cross information fusion neural network based on the samples in the mixed data set until convergence; wherein the time-frequency cross information fusion neural network comprises: A time domain feature extraction branch for extracting time domain features TF of input samples; A frequency domain feature extraction branch for performing FFT operation on input samples to obtain frequency domain coding features GF, and performing coding and convolution operations on the frequency domain coding features GF to obtain frequency domain features FF; A feature cross fusion classification branch comprising: A first convolution module for convolving the frequency domain features FF to obtain frequency domain refined features; A first GELU activation layer connected with the first convolution module for activating the frequency domain refined features to obtain frequency domain activation features; A second convolution module for convolving the time domain features TF to obtain time domain refined features; A second GELU activation layer connected with the second convolution module for activating the time domain refined features to obtain time domain activation features; A first Hadamard product calculator connected with the first GELU activation layer and the second convolution module for calculating the Hadamard product of the frequency domain activation features and the time domain refined features to obtain first cross features; A second Hadamard product calculator connected with the second GELU activation layer and the first convolution module for calculating the Hadamard product of the time domain activation features and the frequency domain refined features to obtain second cross features; A third convolution module connected with the first Hadamard product calculator and the second Hadamard product calculator for convolving the superposition of the first cross features and the second cross features to obtain fusion features CF; A classification module connected with the third convolution module for mapping the fusion features CF to a low-dimensional feature space through a linear layer, and performing probability classification on the fusion features CF in the low-dimensional space to obtain signal categories.
2. The training method of a time-frequency cross information fusion neural network according to claim 1, wherein, The time domain feature extraction branch comprises, in sequence: an integer module, a residual convolution structure and an average pooling layer; The integer module is used for format processing of input samples to obtain a feature tensor; The residual convolution structure is used for feature extraction of the feature tensor while ensuring the dimension to obtain local-global collaborative features; The average pooling layer is used for average pooling of the local-global collaborative features to obtain the time domain features TF.
3. The training method of a time-frequency cross information fusion neural network according to claim 2, wherein, The residual convolution structure comprises, in sequence: a convolution module and a residual network constructed by a plurality of cascaded FFC basic modules. 4.The training method of the time-frequency cross information fusion neural network according to claim 1, wherein, The frequency domain feature extraction branch comprises, in sequence: An FFT module for time domain to frequency domain operation on input samples to obtain a frequency domain feature sequence F; A block module is configured to split the frequency domain feature sequence F to obtain a frequency domain feature sequence block P; An encoding module is configured to encode the frequency domain feature sequence block P to obtain a frequency domain encoded feature GF; A convolution module is configured to convolve the frequency domain encoded feature GF to obtain a frequency domain spatial feature SF; An average pooling layer is configured to compress and aggregate the frequency domain spatial feature SF to obtain the frequency domain feature FF.
5. The training method of a time-frequency cross information fusion neural network according to claim 4, wherein, The encoding module sequentially comprises a position encoding module, a multi-head attention mechanism, residual connection & layer normalization, a feedforward network, and residual connection & layer normalization; The position encoding module is configured to encode the frequency domain feature sequence block P to obtain an initial frequency domain feature with position information; The multi-head attention mechanism is configured to dynamically assign weights to different frequency bands of the initial frequency domain feature to obtain a context-aware feature, so as to capture frequency domain time sequence dependency; The residual connection & layer normalization is configured to process the context-aware feature to obtain a frequency domain global feature; The feedforward network is configured to perform nonlinear feature enhancement feature expression on the frequency domain global feature to obtain a nonlinear feature; The residual connection & layer normalization is configured to output the nonlinear feature as a final frequency domain encoded feature GF.
6. An underwater target detection method based on a time-frequency cross information fusion neural network, characterized in that, It comprises: Inputting an axis frequency magnetic field signal corresponding to a current underwater environment into the time-frequency cross information fusion neural network trained in any one of claims 1-5 to obtain a current category, so as to determine whether there is a target device underwater.
7. An underwater target detection device based on a time-frequency cross information fusion neural network, characterized in that, It comprises: A construction module, a training module, and an identification module; The construction module is configured to construct a mixed data set comprising a positive sample set and a negative sample set; The positive sample set comprises simulated non-stationary non-Gaussian noise data b in a shallow sea environment and background magnetic field data c measured in real time without a target; The negative sample set comprises simulated non-stationary non-Gaussian noise data b in a shallow sea environment and simulated axis frequency magnetic field data a when an underwater target exists; The training module is configured to train the time-frequency cross information fusion neural network based on the samples in the mixed dataset until convergence; wherein the time-frequency cross information fusion neural network comprises: a time-domain feature extraction branch configured to extract time-domain features TF of an input sample; a frequency-domain feature extraction branch configured to perform FFT operation on the input sample to obtain frequency-domain coded features GF, and perform coding and convolution operation on the frequency-domain coded features GF to obtain frequency-domain features FF; and a feature cross fusion classification branch comprising: a first convolution module configured to perform convolution on the frequency-domain features FF to obtain frequency-domain refined features; a first GELU activation layer connected to the first convolution module and configured to activate the frequency-domain refined features to obtain frequency-domain activated features; a second convolution module configured to perform convolution on the time-domain features TF to obtain time-domain refined features; a second GELU activation layer connected to the second convolution module and configured to activate the time-domain refined features to obtain time-domain activated features; a first Hadamard product calculator connected to the first GELU activation layer and the second convolution module and configured to calculate a Hadamard product of the frequency-domain activated features and the time-domain refined features to obtain first cross features; a second Hadamard product calculator connected to the second GELU activation layer and the first convolution module and configured to calculate a Hadamard product of the time-domain activated features and the frequency-domain refined features to obtain second cross features; a third convolution module connected to the first Hadamard product calculator and the second Hadamard product calculator and configured to perform convolution on a superposition of the first cross features and the second cross features to obtain fusion features CF; and a classification module connected to the third convolution module and configured to map the fusion features CF to a low-dimensional feature space through a linear layer, and perform probability classification on the fusion features CF in the low-dimensional space to obtain a signal category. The identification module is configured to input an axis frequency magnetic field signal corresponding to a current underwater environment into the trained time-frequency cross information fusion neural network to obtain a current category, so as to determine whether a target device exists underwater.
8. An underwater target detection device based on a time-frequency cross information fusion neural network, comprising a memory and a processor, the memory storing a computer program, characterized in that, The processor executes the computer program to implement the steps of the method of claim 6.
9. A computer readable storage medium having stored thereon a computer program, characterized in that, The computer program is executed by the processor to implement the steps of the method of any one of claims 1 to 6.