Deep learning broadband signal detection method based on time-frequency feature fusion

By adopting a deep learning method of time-frequency feature fusion in broadband signal detection, integrating historical and current frequency domain features, the problem of insufficient reliability and real-time performance of broadband signal detection in the prior art is solved, and the detection effect of low signal-to-noise ratio signals and adaptability to complex signals is significantly improved.

CN119989270APending Publication Date: 2025-05-13UNIV OF ELECTRONICS SCI & TECH OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510075623.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-17
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The existing broadband signal detection methods show insufficient reliability and timeliness when facing complex modulation methods, variable center frequency, multiple noise and interference factors, and real-time requirements, especially under low signal-to-noise ratio conditions.

Method used

The deep learning broadband signal detection method based on time-frequency feature fusion is adopted. By fusing the frequency domain characteristics of the previous frame and the current frame, the detection effect of the current frame is enhanced and corrected using historical information, a network model including Backbone, Neck and Heads is built, and the network hyperparameters and loss functions are optimized to improve detection performance.

Benefits of technology

The detection effect of low signal-to-noise ratio signals is significantly improved, the detection performance of variable frequency signals and multi-carrier modulated signals is optimized, and the training speed and prediction accuracy are ensured.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119989270A_ABST
    Figure CN119989270A_ABST
Patent Text Reader

Abstract

The invention discloses a deep learning broadband signal detection method based on time-frequency feature fusion, and the method comprises the following steps: 1, simulating a general communication link structure through matlab simulation software, and generating a noisy sample data set of various communication modulation signals; 2, preprocessing the sample: carrying out short-time Fourier transform on the noisy sample data; step 3, constructing a broadband signal detection network model based on time-frequency feature fusion, wherein the network model is composed of three parts, namely a Backbone network, a Neck network and a Heads network; and step 4, training a broadband signal detection network model based on time-frequency feature fusion, and applying the trained broadband signal detection network model to detection of unknown signals. According to the invention, the time domain and frequency domain characteristics of the signal are fully combined, the detection capability of the broadband signal detection network is improved through feature fusion, and the detection of the signal with the variable center frequency is optimized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of signal processing, and specifically relates to a deep learning broadband signal detection method based on time-frequency feature fusion. Background Art

[0002] Broadband signal detection is a technology that uses the time domain and frequency domain characteristics of the receiving broadband range to detect whether there is a signal and detect the corresponding basic parameters. This method is widely used in wireless communication systems, radar reconnaissance, electronic warfare, radio spectrum management and monitoring, intelligent transportation systems and other fields. In wireless communication, it can be used to identify and detect various communication signals to ensure the stability and security of the communication link; in radar reconnaissance, it is used to identify and classify different types of radar signals to improve the accuracy of target detection and tracking; in electronic warfare, it can identify enemy signal sources and implement effective interference and suppression strategies; in radio spectrum management and monitoring, it detects and classifies various types of radio signals in real time to optimize the utilization of spectrum resources; in intelligent transportation systems, it is used to detect and identify vehicle communication signals to improve traffic management and safety. At present, broadband signal detection faces serious difficulties and challenges. The main reasons are as follows: First, the continuous advancement of communication and modulation technology has brought more and more complex modulation methods, making its signal characteristics complex and diverse, and the center frequency flexible and changeable, which brings challenges to the reliability and timeliness of broadband carrier detection. Secondly, the noise and interference factors in complex electromagnetic environments are of various types and change rapidly, and the existing methods are insufficient in anti-interference ability and robustness. Finally, broadband signal detection has high real-time requirements. Existing deep learning models have bottlenecks in processing efficiency and are difficult to meet the needs of real-time detection. Therefore, solving the impact of these problems on broadband signal detection has become a technical problem that needs to be solved urgently and without delay.

[0003] Among the existing methods, CN112784690 A discloses a broadband signal parameter estimation method based on deep learning, which generates a grayscale time-frequency diagram by segmenting the time series data and performing normalization processing, and then uses the YOLOv4 network for signal detection. Although this method is effective, it lacks the use of the time series information of the signal, which limits the detection effect of low signal-to-noise ratio, center frequency changes, and multi-carrier modulation signals. Summary of the invention

[0004] The purpose of the present invention is to overcome the shortcomings of the prior art and provide a method using time-frequency feature fusion, which uses the frequency domain features of the previous frame and the frequency domain features of the current frame to fuse, and uses historical information to enhance and correct the detection effect of the current frame, so as to enhance the detection effect of low signal-to-noise ratio signals. A deep learning broadband signal detection method based on time-frequency feature fusion is proposed.

[0005] The objective of the present invention is achieved through the following technical solution: a deep learning broadband signal detection method based on time-frequency feature fusion, comprising the following steps:

[0006] Step 1: Use MATLAB simulation software to simulate the general communication link structure and generate noisy sample data sets of various communication modulation signals;

[0007] Step 2: Preprocess the samples: perform short-time Fourier transform on the noisy sample data, and then divide it into a training sample set and a validation sample set;

[0008] Step 3: Construct a broadband signal detection network model based on time-frequency feature fusion, including the following sub-steps:

[0009] Step 31, construct a broadband signal detection network model based on time-frequency feature fusion, the network model consists of three parts: Backbone network, Neck network, and Heads network;

[0010] (1) The Backbone network consists of a seven-layer structure, the specific structure is as follows:

[0011] The first layer of the network is the input layer. The input data first passes through a convolution block consisting of a convolution layer with a convolution kernel of 1x3 and a stride of 1x4, a batch normalization layer, and a Hardswish activation function in sequence; then passes through a convolution layer with a convolution kernel of 1x3 and a stride of 1x1, a batch normalization layer, and a ReLU activation function in sequence, and finally passes through a convolution layer with a convolution kernel of 1x1 and a stride of 1x1 and a batch normalization layer in sequence;

[0012] The second layer consists of two inverted residual blocks with the same structure. The inverted residual block includes three sequential convolution blocks:

[0013] 1) The first convolutional block includes a convolutional layer with a convolution kernel of 1x1 and a stride of 1x1, a batch normalization layer, and a ReLU or Hardswish activation function;

[0014] 2) The second convolutional block includes a convolutional layer with a kernel of 1x3 and a stride of 1x1, a batch normalization layer, and a ReLU activation function;

[0015] 3) The third convolutional block includes a convolutional layer with a convolution kernel of 1x1 and a stride of 1x1 and a batch normalization layer in sequence;

[0016] The third layer consists of three inverted residual blocks with the same structure connected in sequence; the inverted residual block includes three convolution blocks connected in sequence, where the structures of the first and third convolution blocks are the same as those of the first and third convolution blocks in the second layer; the second convolution block includes a convolution layer with a convolution kernel of 1x5 and a stride of 1x1, a batch normalization layer, and a ReLU activation function in sequence;

[0017] The fourth layer consists of six inverted residual blocks with the same structure, each of which includes three concatenated convolution blocks. The structures of the first and third convolution blocks are the same as those in the second layer. The second convolution block includes a convolution layer with a kernel of 1x3 and a stride of 1x1, a batch normalization layer, and a Hardswish activation function.

[0018] The fifth layer consists of three inverted residual blocks with the same structure connected in sequence. The inverted residual block includes three convolution blocks connected in sequence. The structures of the first and third convolution blocks are the same as those in the second layer. The second convolution block includes a convolution layer with a convolution kernel of 1x5 and a stride of 1x1, a batch normalization layer, and a Hardswish activation function.

[0019] The sixth and seventh layers are composed of two inverted residual blocks with the same structure. The structure of the inverted residual block is the same as that of the inverted residual block in the fifth layer.

[0020] The output features of each layer are used as a set of time series feature data. A total of seven sets of time series features f_1~f_7 are obtained in the seven layers, and all features are stored in the feature memory;

[0021] A time series feature fusion module is set after the Backbone network, which is used to extract the time series features extracted from the previous frame from the feature memory and fuse them with the time series features extracted from the current frame to obtain fused features f'_1~f'_7; the fused time series features f'_1~f'_7 are input into the Neck network; if there is no historical frame data in the feature memory, the current time series features are input into the Neck network;

[0022] (2) The Neck network includes a seven-layer structure, wherein the first to sixth layers include a 1x1 lateral convolution layer, a smoothing layer consisting of a 1x3 convolution, an upsampling layer A and an upsampling layer B, and the seventh layer includes a 1x1 lateral convolution layer and a smoothing layer consisting of a 1x3 convolution;

[0023] The seven groups of temporal features f'_1~f'_7 output by the temporal feature fusion module are respectively input into the lateral convolution layers of the first to seventh layers of the Neck network; the outputs of the lateral convolution layers of the first to sixth layers are respectively input into the smoothing layer and the upsampling layer A, and the output of the lateral convolution layer of the seventh layer is input into the smoothing layer; the outputs of the smoothing layers of the first to sixth layers are respectively input into the upsampling layer B; the output of the upsampling layer A is added to the output of the next lateral convolution layer and used as the input of the next smoothing layer; the output of the upsampling layer B is added to the output of the next smoothing layer and used as the input of the next smoothing layer; finally, the output of the seventh smoothing layer and the output of the sixth upsampling layer B are added as the output of the Neck network and input into the Heads network;

[0024] (3) The Heads network includes three different heads: reg_head, bw_head, and off_head, which are connected in parallel; bw_head is composed of an adaptive pooling layer, a convolution layer with a convolution kernel of 1x3 and a stride of 1x1, a convolution layer with a convolution kernel of 1x1 and a stride of 1x1, a ReLU activation function, and a convolution layer with a convolution kernel of 1x1 and a stride of 1x1; the activation function of reg_head and off_head is changed to sigmoid, and the other structures are the same as bw_head;

[0025] Step 32: setting the hyper parameters of the broadband signal detection network model based on time-frequency feature fusion, including learning rate and number of iterations; the optimization algorithm of the network model is the Adam algorithm;

[0026] Step 33: Design the loss function to be used in the iterative training process of the network model;

[0027] Step 4: Train a broadband signal detection network model based on time-frequency feature fusion, and use the trained broadband signal detection network model for the detection of unknown signals.

[0028] The beneficial effects of the present invention are:

[0029] First, the present invention adopts a method of time-frequency feature fusion, which fuses the frequency domain features of the previous frame with the frequency domain features of the current frame, and uses historical information to enhance and correct the detection effect of the current frame, thereby enhancing the detection effect of low signal-to-noise ratio signals.

[0030] Second, the present invention introduces timing information by fusing the features of the previous frame, and uses the correlation between consecutive frames to continuously track changes in signal parameters over time. While ensuring training speed and prediction accuracy, it will also be able to optimize the detection effect of variable frequency signals and multi-carrier modulation signals. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] Figure 1 is a flow chart of broadband signal detection of the present invention;

[0032] Figure 2 It is a schematic diagram of the network structure of the present invention;

[0033] Figure 3 This is a detection effect diagram of the present invention when the threshold is 0.8. DETAILED DESCRIPTION

[0034] The specific idea of ​​the present invention is to obtain a broadband signal feature with more comprehensive information by fusing the features extracted by the feature extraction network in the time dimension based on the target detection network, so as to enhance the existing target detection effect. The algorithm uses the characteristics and timing information of the signal frequency domain to train the feature extraction network for signals of different modulation modes under different signal-to-noise ratio conditions, and then fuses and regresses the extracted features to form a complete detection network model, thereby ensuring the detection effect of broadband carrier signals under various conditions. The technical solution of the present invention is further explained below in conjunction with the accompanying drawings.

[0035] like Figure 1 As shown, a deep learning broadband signal detection method based on time-frequency feature fusion of the present invention comprises the following steps:

[0036] Step 1: Use MATLAB simulation software to simulate the general communication link structure and generate noisy sample data sets of various communication modulation signals; including the following steps:

[0037] Step 11, when using MATLAB to simulate and generate a communication modulation signal sample data set, the sampling frequency fs of each modulation signal is set to 3.2MHz, the number of carriers and the frequency are random within the range, the added noise is Gaussian white noise, the data sampling length sample_num is set to 50000 sampling points, the total data sampling time sample_time is set to 0.08 seconds, and the data is grouped according to the length of 0.001 seconds, and a 0.001 second signal sample is input into the model each time;

[0038] Step 12. According to different modulation methods of communication signals, 16 types of modulation signals are generated, namely 2FSK, 4FSK, 8FSK, 8PSK, 8QAM, BPSK, OQPSK, 16QAM, 16QPSK, 32QAM, 32QPS, PI4QPSK, OFDM16, OFDM39, LFM, and FM. Each type of modulation method includes 10 signal-to-noise ratios, which are 1dB-10dB respectively; each type of modulation signal sample includes 10 groups of each signal-to-noise ratio, each group has 80 frames of signal samples, and each type of modulation signal produces 8000 signal samples, and finally 128000 modulation signal samples are obtained.

[0039] Step 2: Preprocess the samples: Perform short-time Fourier transform on the noisy sample data, and then divide it into a training sample set and a verification sample set; Short-time Fourier transform (STFT) can perform time-frequency analysis on the time domain signal sample set to extract the time-frequency characteristics of the signal; the calculation formula of short-time Fourier transform is as follows:

[0040]

[0041] Among them, X i (m, k) represents the STFT transformation result of the i-th signal sample at time m and frequency k, x i [n] is the discrete time domain signal of the i-th signal sample, w[nm] is the window function, and N is the total number of sampling points of a single sample data. By applying STFT, the time domain signal can be converted into a time-frequency domain representation, which helps to analyze the frequency components of the signal in different time periods, thereby more effectively identifying and processing signal features.

[0042] Step 3: Construct a broadband signal detection network model based on time-frequency feature fusion, including the following sub-steps:

[0043] Step 31: Construct a broadband signal detection network model based on time-frequency feature fusion. The network model consists of three parts: Backbone network, Neck network, and Heads network. The specific structure is as follows: Figure 2 As shown;

[0044] The Backbone network includes a seven-layer structure with many inverted residual blocks. The inverted residual block is generally composed of three convolutional blocks stacked together with residual connections. The specific structure of the Backbone network is as follows:

[0045] The first network layer is the input layer, which accepts input data of size (1,16384). The input data first passes through a convolutional block consisting of a convolutional layer with a convolution kernel of 1x3 and a stride of 1x4, a batch normalization layer, and a Hardswish activation function in sequence; then passes through a convolutional layer with a convolution kernel of 1x3 and a stride of 1x1, a batch normalization layer, and a ReLU activation function in sequence, and finally passes through a convolutional layer with a convolution kernel of 1x1 and a stride of 1x1 and a batch normalization layer in sequence;

[0046] The second layer consists of two inverted residual blocks with the same structure. The inverted residual block includes three sequential convolution blocks:

[0047] 1) The first convolutional block includes a convolutional layer with a convolution kernel of 1x1 and a stride of 1x1, a batch normalization layer, and a ReLU or Hardswish activation function;

[0048] 2) The second convolutional block includes a convolutional layer with a kernel of 1x3 and a stride of 1x1, a batch normalization layer, and a ReLU activation function;

[0049] 3) The third convolutional block includes a convolutional layer with a convolution kernel of 1x1 and a stride of 1x1 and a batch normalization layer in sequence;

[0050] The third layer consists of three inverted residual blocks with the same structure connected in sequence; the inverted residual block includes three convolution blocks connected in sequence, where the structures of the first and third convolution blocks are the same as those of the first and third convolution blocks in the second layer; the second convolution block includes a convolution layer with a convolution kernel of 1x5 and a stride of 1x1, a batch normalization layer, and a ReLU activation function in sequence;

[0051] The fourth layer consists of six inverted residual blocks with the same structure, each of which includes three concatenated convolution blocks. The structures of the first and third convolution blocks are the same as those in the second layer. The second convolution block includes a convolution layer with a kernel of 1x3 and a stride of 1x1, a batch normalization layer, and a Hardswish activation function.

[0052] The fifth layer consists of three inverted residual blocks with the same structure connected in sequence. The inverted residual block includes three convolution blocks connected in sequence. The structures of the first and third convolution blocks are the same as those in the second layer. The second convolution block includes a convolution layer with a convolution kernel of 1x5 and a stride of 1x1, a batch normalization layer, and a Hardswish activation function.

[0053] The sixth and seventh layers are composed of two inverted residual blocks with the same structure. The structure of the inverted residual block is the same as that of the inverted residual block in the fifth layer.

[0054] The output features of each layer are respectively used as a set of time series feature data. Seven sets of time series features f_1~f_7 are obtained in seven layers, and all features are stored in the feature memory.

[0055] A time series feature fusion module is set after the Backbone network, which is used to extract the time series features extracted from the previous frame from the feature memory and fuse them with the time series features extracted from the current frame to obtain fused features f'_1~f'_7; the fused time series features f'_1~f'_7 are input into the Neck network. If there is no historical frame data in the feature memory, that is, the current frame is the first frame, the current time series features are directly input into the Neck network.

[0056] The Neck network is used to process feature maps of different scales, generate high-level features, and smooth and further process feature maps of different scales. The Neck network uses an FPN (Feature Pyramid Network) module, including a seven-layer structure, wherein the first to sixth layers include a 1x1 lateral convolution layer, a smoothing layer consisting of 1x3 convolutions, an upsampling layer A and an upsampling layer B, and the seventh layer includes a 1x1 lateral convolution layer and a smoothing layer consisting of 1x3 convolutions;

[0057] The seven groups of temporal features f'_1~f'_7 output by the temporal feature fusion module are respectively input into the lateral convolution layers of the first to seventh layers of the Neck network (in the figure, the Backbone network is the 1st to 7th layers from bottom to top, and the Neck network is the 1st to 7th layers from top to bottom); the outputs of the lateral convolution layers of the first to sixth layers are respectively input into the smoothing layer and the upsampling layer A, and the output of the lateral convolution layer of the seventh layer is input into the smoothing layer; the outputs of the smoothing layers of the first to sixth layers are respectively input into the upsampling layer B; the output of the upsampling layer A is added to the output of the next lateral convolution layer as the input of the next smoothing layer; the output of the upsampling layer B is added to the output of the next smoothing layer as the input of the next smoothing layer; finally, the output of the seventh smoothing layer and the output of the sixth upsampling layer B are added as the output of the Neck network and input into the Heads network;

[0058] The lateral convolutional layer converts feature maps of different scales into 64 channels through 1x1 convolution. There are 7 convolutional layers in total, and the input channel numbers of the first to seventh layers are 160, 160, 160, 112, 40, 24, and 16.

[0059] The Heads network contains three different heads: reg_head, bw_head, and off_head, which are connected in parallel. The structures of these three heads are similar: bw_head consists of an adaptive pooling layer, a convolution layer with a convolution kernel of 1x3 and a stride of 1x1, a convolution layer with a convolution kernel of 1x1 and a stride of 1x1, a ReLU activation function, and a convolution layer with a convolution kernel of 1x1 and a stride of 1x1. The activation functions of reg_head and off_head are changed to sigmoid, and the other structures are the same as bw_head.

[0060] Step 32: setting the hyper parameters of the broadband signal detection network model based on time-frequency feature fusion, including learning rate and number of iterations; the optimization algorithm of the network model is the Adam algorithm;

[0061] Step 33: Design the loss function to be used in the iterative training process of the network model; the defined loss functions include three:

[0062] Loss function 1 uses Focal Loss, which is expressed as follows:

[0063] L1=FL(p t )=-α t (1-p t ) γ log(p t )

[0064] Among them, p t is the predicted value, α tis a balancing factor used to adjust the weights of positive and negative samples; γ is a focus parameter used to adjust the weight ratio of difficult samples to easy-to-classify samples;

[0065] Both loss function 2 and loss function 3 use absolute error loss, which is expressed as follows:

[0066]

[0067] Among them, y pred,i is the predicted value of the model, y true,i is the true label, indicating the target value.

[0068] Step 4: training a broadband signal detection network model based on time-frequency feature fusion, and using the trained broadband signal detection network model for detecting unknown signals; including the following sub-steps:

[0069] Step 41, training a broadband signal detection network model based on time-frequency feature fusion, includes the following sub-steps:

[0070] Step 41-1, disrupt the arrangement order of all samples in the training sample set and the verification sample set, input each group of signals into the feature extraction network Backbone in 0.001 second time division, obtain and save the feature information of the current frame, if the feature information of the previous frame exists in the feature memory, perform feature fusion on the signal features of the previous frame and the signal features of the current frame, introduce the timing information of the continuous frames, and then continue to input into the subsequent network;

[0071] Step 41-2, iteratively optimize the loss functions L1, L2 and L3 through the Adam optimization algorithm to train the entire network;

[0072] Step 41-3, setting the learning rate cosine annealing mechanism and early stopping mechanism, after the iterative optimization is completed, a trained network model is obtained;

[0073] Step 42: Use the trained network model to detect unknown signals.

[0074] Figure 3 This is a statistical diagram of the detection effect of the present invention when the threshold is 0.8 (the prediction is considered correct only if the iou overlap between the detection value and the actual label is greater than 0.8). The F1 Score is an indicator used in statistics to measure the accuracy of the model. It takes into account both the precision and recall of the classification model. The F1 score can be regarded as a harmonic average of the model's precision and recall, with a maximum value of 1 and a minimum value of 0. It can be seen from the figure that there is also a good detection effect in a low signal-to-noise ratio environment (the F1 score is also higher than 0.8), and the effect gradually improves with the improvement of the carrier signal-to-noise ratio.

[0075] Those skilled in the art will appreciate that the embodiments described herein are intended to help readers understand the principles of the present invention, and should be understood that the protection scope of the present invention is not limited to such specific statements and embodiments. Those skilled in the art can make various other specific variations and combinations that do not deviate from the essence of the present invention based on the technical revelations disclosed by the present invention, and these variations and combinations are still within the protection scope of the present invention.

Claims

1. A deep learning broadband signal detection method based on time-frequency feature fusion, characterized in that: The steps include: Step 1: Use MATLAB simulation software to simulate the general communication link structure and generate noisy sample data sets of various communication modulation signals; Step 2: Preprocess the samples: perform short-time Fourier transform on the noisy sample data; Step 3: Construct a broadband signal detection network model based on time-frequency feature fusion, including the following sub-steps: Step 31, construct a broadband signal detection network model based on time-frequency feature fusion, the network model consists of three parts: Backbone network, Neck network, and Heads network; (1) The Backbone network consists of a seven-layer structure, the specific structure is as follows: The first layer of the network is the input layer. The input data first passes through a convolution block consisting of a convolution layer with a convolution kernel of 1x3 and a stride of 1x4, a batch normalization layer, and a Hardswish activation function in sequence; then passes through a convolution layer with a convolution kernel of 1x3 and a stride of 1x1, a batch normalization layer, and a ReLU activation function in sequence, and finally passes through a convolution layer with a convolution kernel of 1x1 and a stride of 1x1 and a batch normalization layer in sequence; The second layer consists of two inverted residual blocks with the same structure. The inverted residual block includes three sequential convolution blocks: 1) The first convolutional block includes a convolutional layer with a convolution kernel of 1x1 and a stride of 1x1, a batch normalization layer, and a ReLU or Hardswish activation function; 2) The second convolutional block includes a convolutional layer with a kernel of 1x3 and a stride of 1x1, a batch normalization layer, and a ReLU activation function; 3) The third convolutional block includes a convolutional layer with a convolution kernel of 1x1 and a stride of 1x1 and a batch normalization layer in sequence; The third layer consists of three inverted residual blocks with the same structure connected in sequence; the inverted residual block includes three convolution blocks connected in sequence, where the structures of the first and third convolution blocks are the same as those of the first and third convolution blocks in the second layer; the second convolution block includes a convolution layer with a convolution kernel of 1x5 and a stride of 1x1, a batch normalization layer, and a ReLU activation function in sequence; The fourth layer consists of six inverted residual blocks with the same structure, each of which includes three concatenated convolution blocks. The structures of the first and third convolution blocks are the same as those in the second layer. The second convolution block includes a convolution layer with a kernel of 1x3 and a stride of 1x1, a batch normalization layer, and a Hardswish activation function. The fifth layer consists of three inverted residual blocks with the same structure connected in sequence. The inverted residual block includes three convolution blocks connected in sequence. The structures of the first and third convolution blocks are the same as those in the second layer. The second convolution block includes a convolution layer with a convolution kernel of 1x5 and a stride of 1x1, a batch normalization layer, and a Hardswish activation function. The sixth and seventh layers are composed of two inverted residual blocks with the same structure. The structure of the inverted residual block is the same as that of the inverted residual block in the fifth layer. The output features of each layer are used as a set of time series feature data. A total of seven sets of time series features f_1~f_7 are obtained in the seven layers, and all features are stored in the feature memory; A time series feature fusion module is set after the Backbone network, which is used to extract the time series features extracted from the previous frame from the feature memory and fuse them with the time series features extracted from the current frame to obtain fused features f'_1~f'_7; the fused time series features f'_1~f'_7 are input into the Neck network; if there is no historical frame data in the feature memory, the current time series features are input into the Neck network; (2) The Neck network includes a seven-layer structure, wherein the first to sixth layers include a 1x1 lateral convolution layer, a smoothing layer consisting of a 1x3 convolution, an upsampling layer A and an upsampling layer B, and the seventh layer includes a 1x1 lateral convolution layer and a smoothing layer consisting of a 1x3 convolution; The seven groups of temporal features f'_1~f'_7 output by the temporal feature fusion module are respectively input into the lateral convolution layers of the first to seventh layers of the Neck network; the outputs of the lateral convolution layers of the first to sixth layers are respectively input into the smoothing layer and the upsampling layer A, and the output of the lateral convolution layer of the seventh layer is input into the smoothing layer; the outputs of the smoothing layers of the first to sixth layers are respectively input into the upsampling layer B; the output of the upsampling layer A is added to the output of the next lateral convolution layer and used as the input of the next smoothing layer; The output of the upsampling layer B is added to the output of the next smooth layer and used as the input of the next smooth layer. Finally, the output of the seventh smooth layer and the output of the sixth upsampling layer B are added as the output of the Neck network and input into the Heads network. (3) The Heads network includes three different heads: reg_head, bw_head, and off_head, which are connected in parallel; bw_head is composed of an adaptive pooling layer, a convolution layer with a convolution kernel of 1x3 and a stride of 1x1, a convolution layer with a convolution kernel of 1x1 and a stride of 1x1, a ReLU activation function, and a convolution layer with a convolution kernel of 1x1 and a stride of 1x1; the activation function of reg_head and off_head is changed to sigmoid, and the other structures are the same as bw_head; Step 32: setting the hyper parameters of the broadband signal detection network model based on time-frequency feature fusion, including learning rate and number of iterations; the optimization algorithm of the network model is the Adam algorithm; Step 33: Design the loss function to be used in the iterative training process of the network model; Step 4: Train a broadband signal detection network model based on time-frequency feature fusion, and use the trained broadband signal detection network model for the detection of unknown signals.

2. The method for detecting broadband signals by deep learning based on time-frequency feature fusion according to claim 1, characterized in that: The step 1 comprises the following steps: Step 11, when using MATLAB to simulate and generate a communication modulation signal sample data set, the sampling frequency fs of each modulation signal is set to 3.2MHz, the number of carriers and the frequency are random within the range, the added noise is Gaussian white noise, the data sampling length sample_num is set to 50000 sampling points, the total data sampling time sample_time is set to 0.08 seconds, and the data is grouped according to the length of 0.001 seconds, and a 0.001 second signal sample is input into the model each time; Step 12. According to different modulation methods of communication signals, 16 types of modulation signals are generated, namely 2FSK, 4FSK, 8FSK, 8PSK, 8QAM, BPSK, OQPSK, 16QAM, 16QPSK, 32QAM, 32QPS, PI4QPSK, OFDM16, OFDM39, LFM, and FM. Each type of modulation method includes 10 signal-to-noise ratios, which are 1dB-10dB respectively; each type of modulation signal sample includes 10 groups of each signal-to-noise ratio, each group has 80 frames of signal samples, and each type of modulation signal produces 8000 signal samples, and finally 128000 modulation signal samples are obtained.

3. The deep learning broadband signal detection method based on time-frequency feature fusion according to claim 1 is characterized in that: The loss functions defined in step 33 include three: Loss function 1 uses Focal Loss, which is expressed as follows: L1=FL(p t )−α t (1-p t ) γ log(p t ) Among them, p t is the predicted value, α t is a balancing factor used to adjust the weights of positive and negative samples; γ is a focus parameter used to adjust the weight ratio of difficult samples to easy-to-classify samples; Both loss function 2 and loss function 3 use absolute error loss, which is expressed as follows: Among them, y pred,i is the predicted value of the model, y true,i is the true label, indicating the target value.