Time-frequency mode decomposition method and system based on deep learning, terminal and storage medium

By employing a deep learning-based time-frequency mode decomposition method, high-resolution time-frequency analysis and mode segmentation are performed using convolutional neural networks and Transformer decoding layers. This solves the problems of time-frequency resolution and mode separation in existing technologies, and enables accurate extraction and separation of complex time-domain signals.

CN122045595APending Publication Date: 2026-05-15深圳开鸿数字产业发展有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
深圳开鸿数字产业发展有限公司
Filing Date
2025-12-26
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

Existing time-frequency analysis and mode decomposition methods cannot effectively achieve high-resolution time-frequency representation and accurate mode separation, especially when dealing with non-stationary signals, they suffer from time-frequency resolution trade-offs, cross-term interference, and noise sensitivity.

Method used

A deep learning-based time-frequency mode decomposition method is adopted, which performs high-resolution time-frequency analysis, multi-scale time-frequency representation segmentation and signal reconstruction through a convolutional neural network model. The multi-scale feature extractor and Transformer decoding layer of the convolutional neural network are used for mode segmentation, and the demodulation frequency function is optimized by combining the alternating direction multiplier method to achieve accurate extraction of signal components.

Benefits of technology

It achieves efficient separation and trajectory prediction of aliased components in complex time-domain signals, improves the resolution of time-frequency analysis and the accuracy of mode decomposition, and can effectively handle the multi-mode coupling characteristics in complex non-stationary signals.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122045595A_ABST
    Figure CN122045595A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of time-frequency analysis and modal decomposition, and discloses a time-frequency modal decomposition method and system based on deep learning, a terminal and a storage medium, and the method comprises the steps: obtaining an input complex number of time-domain signals, analyzing the complex number of time-domain signals through a high-resolution time-frequency analysis module of a convolutional neural network model, and obtaining a complex number of time-domain signals; obtaining multi-scale time-frequency representation; performing instance segmentation on the multi-scale time-frequency representation through a time-frequency mode segmentation module of a convolutional neural network model to obtain a time-frequency mode; and performing trajectory prediction on the time-frequency mode and the complex number time-domain signal through a signal reconstruction module of a convolutional neural network model to obtain a signal component. According to high-resolution time-frequency analysis and time-frequency modal segmentation, separation and trajectory prediction are carried out on aliasing components in complex time-domain signals, and accurate extraction of signal components is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of time-frequency analysis and mode decomposition technology, and in particular to a time-frequency mode decomposition method, system, terminal and computer-readable storage medium based on deep learning. Background Technology

[0002] Non-stationary signals in natural scenes, such as frequency-hopping signals in communication, micro-Doppler echoes in radar, and mechanical vibration signals, are characterized by time-varying instantaneous frequency (IF), multimodal coupling, and strong noise interference. How to extract time-varying features and separate mixed-mode information to assist in channel equalization, fault diagnosis, and target recognition has become a research hotspot in the field of signal processing and information acquisition. Traditional time-frequency analysis (TFA) techniques struggle to handle the inherent trade-offs between time and frequency resolution and cross-term interference, while traditional mode decomposition methods suffer from endpoint effects and sensitivity to noise.

[0003] Existing time-frequency analysis and mode decomposition methods include: time-frequency analysis methods based on synchronous squeezing transform, mode decomposition methods based on variational models, and parameter optimization methods based on traditional signal processing. These traditional methods are limited by the inherent trade-offs in time-frequency resolution, cross-term interference, and sensitivity to noise, and cannot effectively achieve high-resolution time-frequency representation and accurate mode separation. Therefore, existing time-frequency analysis and mode decomposition methods still need to be improved and optimized. Summary of the Invention

[0004] The main objective of this invention is to provide a time-frequency mode decomposition method, system, terminal, and computer-readable storage medium based on deep learning, aiming to solve the problem that existing time-frequency analysis and mode decomposition methods cannot effectively achieve high-resolution time-frequency representation and accurate mode separation.

[0005] To achieve the above objectives, the present invention provides a time-frequency mode decomposition method based on deep learning, the method comprising the following steps: The input complex time-domain signal is acquired, and the complex time-domain signal is analyzed by the high-resolution time-frequency analysis module of the convolutional neural network model to obtain a multi-scale time-frequency representation; The multi-scale time-frequency representation is segmented into instances using the time-frequency mode segmentation module of the convolutional neural network model to obtain the time-frequency modes; The signal reconstruction module of the convolutional neural network model performs trajectory prediction on the time-frequency mode and the complex time-domain signal to obtain signal components.

[0006] Optionally, the deep learning-based time-frequency mode decomposition method, wherein acquiring the input complex time-domain signal and analyzing the complex time-domain signal through a high-resolution time-frequency analysis module of a convolutional neural network model to obtain a multi-scale time-frequency representation specifically includes: The input complex time-domain signal is acquired, and the complex time-domain signal is initialized by multiple complex convolutions of the convolutional neural network model and the short-time filtering framework of STFT to obtain the initial time-frequency representation. The initial time-frequency representation is weighted by the channel attention of a convolutional neural network model to obtain time-frequency features; The time-frequency features are subjected to hierarchical transformation by a multi-scale feature extractor of a convolutional neural network model to obtain a multi-scale time-frequency representation.

[0007] Optionally, in the deep learning-based time-frequency mode decomposition method, the complex time-domain signal includes radar signals and communication signals.

[0008] Optionally, in the deep learning-based time-frequency mode decomposition method, the hierarchical transformation includes hierarchical encoding and hierarchical decoding. The process of performing a hierarchical transformation on the time-frequency features using a multi-scale feature extractor based on a convolutional neural network model to obtain a multi-scale time-frequency representation specifically includes: The time-frequency features are hierarchically encoded using the residual connection network of the multi-scale feature extractor to obtain a globally defined representation; The global definition representation is decoded hierarchically through the residual connection network to obtain a multi-scale time-frequency representation.

[0009] Optionally, in the deep learning-based time-frequency mode decomposition method, the time-frequency mode segmentation module includes a pixel decoding layer and a Transformer decoding layer; The step of segmenting the multi-scale time-frequency representation using a time-frequency mode segmentation module based on a convolutional neural network model to obtain time-frequency modes specifically includes: The pixel decoding layer transforms the structural and semantic information of the multi-scale time-frequency representation to obtain a high-dimensional pixel representation; The time-frequency mode is obtained by embedding the high-dimensional pixel representation through the Transformer decoding layer.

[0010] Optionally, the deep learning-based time-frequency mode decomposition method, wherein the step of converting the structural and semantic information of the multi-scale time-frequency representation through the pixel decoding layer to obtain a high-dimensional pixel representation specifically includes: Based on the U-Net submodule in the TFA module, the structural and semantic information of the multi-scale time-frequency representation is transformed to obtain intermediate results; The intermediate results are parsed by the pixel decoding layer to obtain a high-dimensional pixel representation.

[0011] Optionally, the deep learning-based time-frequency mode decomposition method, wherein the signal reconstruction module using a convolutional neural network model performs trajectory prediction on the time-frequency mode and the complex time-domain signal to obtain signal components, specifically includes: The signal reconstruction module of the convolutional neural network model constructs the trajectory of the time-frequency mode to obtain the initial demodulation frequency function. The complex time-domain signal is inversely deconstructed based on the initial demodulation frequency function to obtain the signal components.

[0012] Optionally, the deep learning-based time-frequency mode decomposition method, wherein the step of constructing the trajectory of the time-frequency mode using the signal reconstruction module of the convolutional neural network model to obtain the initial demodulation frequency function specifically includes: The peak value of the time-frequency mode is located by the signal reconstruction module of the convolutional neural network model, and multiple trajectory points are obtained. By constructing trajectories from all trajectory points, the initial demodulation frequency function is obtained.

[0013] Optionally, the deep learning-based time-frequency mode decomposition method, wherein the step of inversely deconstructing the complex time-domain signal according to the initial demodulation frequency function to obtain signal components specifically includes: Based on the signal reconstruction module, the complex time-domain signal is analyzed to obtain the original time series; The original time series is deconstructed inversely based on the initial demodulation frequency function to obtain the signal components.

[0014] Optionally, the deep learning-based time-frequency mode decomposition method further includes, after performing inverse decomposition of the complex time-domain signal according to the initial demodulation frequency function to obtain signal components: Based on the alternating direction multiplier method, the initial demodulation frequency function is repeatedly updated to obtain the optimal initial demodulation frequency function. The signal components are reconstructed using the alternating direction multiplier method to obtain the optimal signal components.

[0015] Furthermore, to achieve the above objectives, the present invention also provides a time-frequency mode decomposition system based on deep learning, wherein the time-frequency mode decomposition system based on deep learning: The complex time-domain signal analysis module is used to acquire the input complex time-domain signal and analyze the complex time-domain signal through the high-resolution time-frequency analysis module of the convolutional neural network model to obtain a multi-scale time-frequency representation; The multi-scale time-frequency representation segmentation module is used to perform instance segmentation on the multi-scale time-frequency representation through the time-frequency mode segmentation module of the convolutional neural network model to obtain the time-frequency mode; The signal component construction module is used to predict the trajectory of the time-frequency mode and the complex time-domain signal through the signal reconstruction module of the convolutional neural network model to obtain the signal components.

[0016] Optionally, in the deep learning-based time-frequency mode decomposition system, the complex time-domain signal analysis module includes: The complex time-domain signal unit is used to acquire the input complex time-domain signal, and to initialize the complex time-domain signal through multiple complex convolutions of the convolutional neural network model and the short-time filtering framework of STFT to obtain the initial time-frequency representation; The initial time-frequency representation weighting unit is used to weight the initial time-frequency representation through the channel attention of the convolutional neural network model to obtain time-frequency features; The hierarchical transformation unit is used to perform hierarchical transformation on the time-frequency features through the multi-scale feature extractor of the convolutional neural network model to obtain a multi-scale time-frequency representation.

[0017] Optionally, in the deep learning-based time-frequency mode decomposition system, the hierarchical transformation unit includes: The hierarchical coding subunit is used to hierarchically encode the time-frequency features through the residual connection network of the multi-scale feature extractor to obtain a globally defined representation; The global definition representation decoding subunit is used to perform hierarchical decoding of the global definition representation through the residual connection network to obtain a multi-scale time-frequency representation.

[0018] Optionally, in the deep learning-based time-frequency mode decomposition system, the multi-scale time-frequency representation segmentation module includes: The semantic conversion unit is used to convert the structural and semantic information of the multi-scale time-frequency representation through the pixel decoding layer of the time-frequency modality segmentation module to obtain a high-dimensional pixel representation; An embedded interaction unit is used to perform embedded interaction on the high-dimensional pixel representation through the Transformer decoding layer of the time-frequency modality segmentation module to obtain the time-frequency modality.

[0019] Optionally, in the deep learning-based time-frequency mode decomposition system, the semantic transformation unit includes: The multi-scale time-frequency representation conversion subunit is used to convert the structural and semantic information of the multi-scale time-frequency representation based on the U-Net submodule in the TFA module to obtain intermediate results; The intermediate result parsing subunit is used to parse the intermediate result through the pixel decoding layer to obtain a high-dimensional pixel representation.

[0020] Optionally, in the deep learning-based time-frequency mode decomposition system, the signal component construction module includes: The time-frequency mode reconstruction unit is used to construct the trajectory of the time-frequency mode through the signal reconstruction module of the convolutional neural network model to obtain the initial demodulation frequency function; The complex time-domain signal inverse deconstruction unit is used to inversely deconstruct the complex time-domain signal according to the initial demodulation frequency function to obtain signal components.

[0021] Optionally, in the deep learning-based time-frequency mode decomposition system, the time-frequency mode reconstruction unit comprises: The trajectory point localization subunit is used to locate the peak of the time-frequency mode through the signal reconstruction module of the convolutional neural network model to obtain multiple trajectory points; The trajectory construction sub-unit is used to construct trajectories for all trajectory points to obtain the initial demodulation frequency function.

[0022] Optionally, in the deep learning-based time-frequency mode decomposition system, the complex time-domain signal inverse decomposition unit comprises: The complex time-domain signal analysis subunit is used to analyze the complex time-domain signal based on the signal reconstruction module to obtain the original time series; The original time series inverse deconstruction subunit is used to inversely deconstruct the original time series according to the initial demodulation frequency function to obtain signal components.

[0023] Furthermore, to achieve the above objectives, the present invention also provides a computer-readable storage medium, wherein the computer-readable storage medium stores a deep learning-based time-frequency mode decomposition program, which, when executed by a processor, implements the steps of the deep learning-based time-frequency mode decomposition method as described above.

[0024] In this invention, an input complex time-domain signal is acquired, and the complex time-domain signal is analyzed by a high-resolution time-frequency analysis module of a convolutional neural network model to obtain a multi-scale time-frequency representation. The multi-scale time-frequency representation is then segmented into time-frequency modes by a time-frequency mode segmentation module of the convolutional neural network model. Finally, the time-frequency modes and the complex time-domain signal are used for trajectory prediction by a signal reconstruction module of the convolutional neural network model to obtain signal components. This invention separates and predicts the trajectory of aliased components in the complex time-domain signal based on high-resolution time-frequency analysis and time-frequency mode segmentation, achieving accurate extraction of signal components. Attached Figure Description

[0025] Figure 1 This is a flowchart of a preferred embodiment of the time-frequency mode decomposition method based on deep learning of the present invention; Figure 2 This is a flowchart of a time-frequency mode decomposition system, which is a preferred embodiment of the signal parameter prediction method based on deep learning of the present invention. Figure 3 This is a flowchart illustrating the specific implementation process of step S10 in a preferred embodiment of the signal parameter prediction method based on deep learning of the present invention. Figure 4 This is a flowchart illustrating the specific implementation process of step S13 in a preferred embodiment of the signal parameter prediction method based on deep learning of the present invention. Figure 5 This is a flowchart of step S20 of a preferred embodiment of the signal parameter prediction method based on deep learning of the present invention; Figure 6 This is a flowchart of step S22 of a preferred embodiment of the signal parameter prediction method based on deep learning of the present invention; Figure 7 This is a flowchart of step S30 of a preferred embodiment of the signal parameter prediction method based on deep learning of the present invention; Figure 8 This is a flowchart of the alternating direction multiplier method processing, a preferred embodiment of the signal parameter prediction method based on deep learning of the present invention. Figure 9 This is a flowchart of step S31 of a preferred embodiment of the signal parameter prediction method based on deep learning of the present invention; Figure 10 This is a schematic diagram illustrating the TFR performance of a preferred embodiment of the deep learning-based signal parameter prediction method of the present invention; Figure 11 This is a schematic diagram of the time-frequency mode decomposition results of a preferred embodiment of the signal parameter prediction method based on deep learning of the present invention; Figure 12 This is a flowchart of step S32 of a preferred embodiment of the signal parameter prediction method based on deep learning of the present invention; Figure 13 This is a schematic diagram of a time-frequency mode decomposition system based on deep learning according to the present invention; Figure 14 This is another schematic diagram of the time-frequency mode decomposition system based on deep learning according to the present invention; Figure 15 This is a structural diagram of a preferred embodiment of the terminal of the device of the present invention. Detailed Implementation

[0026] To make the objectives, technical solutions, and advantages of this invention clearer and more explicit, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.

[0027] Currently, non-stationary signals in natural scenes, such as communication frequency-hopping signals, radar micro-Doppler echoes, and mechanical vibration signals, are characterized by time-varying instantaneous frequencies, multi-mode coupling, and strong noise interference. How to extract time-varying features and separate mixed-mode information to assist in channel equalization, fault diagnosis, and target recognition has become a research hotspot in the field of signal processing and information acquisition. Traditional time-frequency analysis techniques struggle to handle the inherent trade-offs between time-frequency resolution and cross-term interference, while traditional mode decomposition methods suffer from endpoint effects and sensitivity to noise.

[0028] In recent years, research on improving time-frequency resolution has often utilized synchronous squeezing transform. This method sharpens the energy distribution by aligning the energy distribution in the TF plane with the true IF mode obtained from the phase information extracted from the time-frequency representation. Based on SST, researchers have developed methods such as time redistribution and synchronous squeezing, synchronous extraction transform, and multi-synchronous squeezing transform. With the development of deep learning, some TFA networks have been proposed, further improving the energy concentration of TFR. However, the aforementioned TFA methods only focus on energy focusing in the TFR and do not consider mode separation.

[0029] In mode decomposition, variational nonlinear chirped mode decomposition transforms the separation process into an optimization problem involving signal demodulation. Given a suitable initial IF estimate, this method can progressively refine the IF measurement and recover all waveform components. Therefore, researchers have developed enhanced initialization methods, such as adaptive chirped mode decomposition techniques, regrouping plus intrinsic chirped component decomposition, and improved variational generalized nonlinear mode decomposition. However, these traditional signal processing steps require extensive empirical parameter tuning to obtain satisfactory mode decomposition results.

[0030] Existing time-frequency analysis and mode decomposition methods include: time-frequency analysis methods based on synchronous squeezing transform, mode decomposition methods based on variational models, and parameter optimization methods based on traditional signal processing. These traditional methods are limited by the inherent trade-offs of time-frequency resolution, cross-term interference, and sensitivity to noise, and cannot effectively achieve high-resolution time-frequency representation and accurate mode separation. Therefore, a deep learning-based time-frequency mode decomposition method is needed to separate and predict the trajectory of aliased components in complex time-domain signals based on high-resolution time-frequency analysis and time-frequency mode segmentation, thereby achieving accurate extraction of signal components and avoiding the problem of not being able to effectively achieve high-resolution time-frequency representation and accurate mode separation.

[0031] The time-frequency mode decomposition method based on deep learning described in the preferred embodiment of the present invention, such as... Figure 1 and Figure 2 As shown, the deep learning-based time-frequency mode decomposition method includes the following steps: like Figure 3 As shown, in step S10, the input complex time-domain signal is acquired, and the complex time-domain signal is analyzed by the high-resolution time-frequency analysis module of the convolutional neural network model to obtain a multi-scale time-frequency representation.

[0032] Step S10 includes: Step S11: Obtain the input complex time-domain signal, and initialize the complex time-domain signal through multiple complex convolutions of the convolutional neural network model and the short-time filtering framework of STFT to obtain the initial time-frequency representation; Step S12: Weight the initial time-frequency representation using the channel attention of the convolutional neural network model to obtain time-frequency features; Step S13: Perform hierarchical transformation on the time-frequency features using the multi-scale feature extractor of the convolutional neural network model to obtain a multi-scale time-frequency representation.

[0033] Specifically, the input complex time-domain signal (such as radar signals and communication signals) is acquired, and initialized using multiple complex convolutions of a convolutional neural network model (the TFA module (convolutional neural network model) consists of three sets of complex convolutions with different kernel lengths, generating multi-resolution time-frequency representations in a learnable manner) and the short-time filtering framework of the STFT, to obtain an initial time-frequency representation (the structure of the initial time-frequency representation corresponds to the short-time filtering framework of the STFT). The initial time-frequency representation is then weighted using the channel attention of the convolutional neural network model to obtain time-frequency features (the convolutional output is weighted by channel attention to form time-frequency features in multiple subspaces). Finally, the time-frequency features are subjected to hierarchical transformation by the multi-scale feature extractor of the convolutional neural network model to obtain a multi-scale time-frequency representation (the output of this module is updated by the network, such as...). Figure 2The high-resolution TFR shown in the upper left corner includes radar signals and communication signals.

[0034] In this embodiment, as Figure 2 As shown, the high-resolution TFA (Convolutional Neural Network Model) module uses multi-scale time-frequency transform blocks and conventional blocks of residual U-Net shape to extract time-frequency features and generate high-resolution TFRs. The high-resolution time-frequency analysis module is based on the query-based structure of Mask2Former: first, multi-scale features from the TFA module are input into a pixel decoder with multi-scale deformable attention to obtain unified high-dimensional pixel embeddings; then, a set of learnable queries interacts with these embeddings through the Transformer decoder to generate corresponding time-frequency masks (time-frequency modes) to segment different IF trajectories. The model-based signal reconstruction module constructs an initial demodulation frequency function based on the segmented IF trajectories. By minimizing the bandwidth of the demodulated signal as the objective, the alternating direction method of multipliers (ADMM) is used to repeatedly update the demodulation frequency function and the reconstructed waveform, gradually converging the demodulated signal to the optimal solution (optimal initial demodulation frequency function and optimal signal components) that satisfies the narrowest bandwidth constraint.

[0035] As an example, assume the input signal is a complex time-domain signal of mechanical vibration with a sampling frequency of 10 kHz and a length of 10,000 points. First, this signal is input into a designed convolutional neural network model containing three sets of complex convolutional kernels of different lengths, corresponding to filters with different window lengths in the Short-Time Fourier Transform (STFT): the first set of kernels has a length of 64, capturing high-frequency details within a shorter time window; the second set has a length of 128, balancing time and frequency resolution; and the third set has a length of 256, focusing on the stable extraction of low-frequency components. Each set of kernels performs a one-dimensional complex convolution operation on the input signal, simulating the short-time filtering process of the STFT, and outputs the corresponding complex time-frequency representation. Through parallel processing of multiple sets of convolutional kernels, the network achieves multi-resolution time-frequency feature extraction, obtaining an initial multi-channel time-frequency representation tensor.

[0036] Furthermore, for the multi-channel initial time-frequency representation obtained in step S11, the network introduces a channel attention mechanism (such as an SE module or a CBAM module) to weight the features of each channel: First, global average pooling is performed on the time-frequency representation of each channel to obtain the channel descriptor; the weight distribution between channels is learned through a two-layer fully connected network, with ReLU and Sigmoid activation functions; the learned weight coefficients are multiplied by the time-frequency features of the corresponding channel to enhance important frequency components and time periods, and suppress noise and irrelevant information. For example, for the 500 Hz mode in a mechanical vibration signal, the channel attention mechanism will automatically increase the response intensity of the corresponding channel, making subsequent processing more focused on the time-frequency features of this mode.

[0037] In this embodiment, the weighted time-frequency features are input into a multi-scale feature extractor, which employs a U-Net structure with residual connections for hierarchical encoding and decoding: the encoder extracts high-level semantic information of the time-frequency features through multiple convolutional and pooling operations, while simultaneously reducing the time-frequency resolution; the decoder restores the time-frequency resolution and fuses features at different scales through upsampling and skip connections, enhancing detail representation; residual connections ensure smooth information transmission, avoid gradient vanishing, and improve training stability. Finally, the network outputs a multi-scale time-frequency representation tensor, containing both fine-grained high-resolution time-frequency information and global modal structure features, providing rich input for the subsequent time-frequency modality segmentation module.

[0038] like Figure 4 As shown, step S13 includes: Step S131: The time-frequency features are hierarchically encoded through the residual connection network of the multi-scale feature extractor to obtain a globally defined representation; Step S132: Perform hierarchical decoding on the global definition representation through the residual connection network to obtain a multi-scale time-frequency representation.

[0039] Specifically, the time-frequency features are hierarchically encoded through the residual connection network of the multi-scale feature extractor to obtain a global definition representation. The global definition representation is then hierarchically decoded through the residual connection network to obtain a multi-scale time-frequency representation (a multi-scale time-frequency representation refined layer by layer is obtained through hierarchical encoding and decoding using a U-Net with residual connections). The hierarchical transformation includes hierarchical encoding and hierarchical decoding.

[0040] In this embodiment, the original complex time-domain signal is first subjected to multi-scale convolutional transformation using a multi-scale feature extractor to extract time-frequency features at different time and frequency resolutions. Subsequently, these multi-scale time-frequency features are hierarchically encoded using an encoding network with residual connections. This encoding process gradually aggregates local time-frequency information through layer-by-layer downsampling and convolution operations, capturing the global structure and semantic features of the signal, thereby generating a high-dimensional representation with global definition capabilities. The residual connection design in the encoding network effectively alleviates the gradient vanishing problem in deep network training, ensuring effective information transmission and stable feature extraction. Simultaneously, the residual structure enhances the network's ability to preserve signal details and edge information, which is helpful for subsequent modality segmentation and signal reconstruction. After obtaining the globally defined representation, a corresponding residual connection decoding network is used for hierarchical decoding. The decoding process gradually restores the spatial resolution and detailed information of the time-frequency features through layer-by-layer upsampling and convolution operations, combined with multi-scale feature skip connections from the encoding stage. This decoding network also employs a residual connection structure to ensure smooth information transmission and feature fusion during the decoding process, avoiding information loss and ambiguity. Overall, the encoding and decoding networks constitute a U-Net structure with residual connections. Through hierarchical encoding and decoding transformations, multi-scale fusion and layer-by-layer refinement of time-frequency features are achieved. This structure not only improves the resolution and accuracy of time-frequency representation but also enhances the ability to express multi-modal coupling features in complex non-stationary signals. The final output multi-scale time-frequency representation retains both the global structural information of the signal and rich local details, providing high-quality input features for subsequent time-frequency mode segmentation modules. In summary, this invention effectively achieves multi-scale representation and layer-by-layer refinement of time-frequency features through a multi-scale feature extractor and a hierarchical encoding and decoding network with residual connections, significantly improving the performance of time-frequency analysis and the accuracy of mode decomposition.

[0041] As an example, the input complex time-domain signal is first processed through a preliminary time-frequency transform block to obtain a preliminary time-frequency feature map. In order to better extract the multi-scale time-frequency information of the signal, the system adopts a multi-scale feature extraction network with residual connections. The network structure is similar to the classic U-Net, but residual connections are specially designed to alleviate the gradient vanishing problem in deep network training and enhance the stability of feature propagation.

[0042] Specifically, the hierarchical encoding process starts with the input time-frequency features and gradually extracts more abstract and global time-frequency features through multiple convolution and pooling operations. Each encoding layer not only extracts features at the current scale but also directly passes feature information from the previous layer to subsequent layers through residual connections, ensuring the preservation of detailed information and effective gradient propagation. After multiple encoding layers, the network obtains a globally defined time-frequency representation that integrates information from different scales and time-frequency resolutions, providing a more comprehensive reflection of the signal's time-frequency structure. Subsequently, the hierarchical decoding process restores this global representation layer by layer. The decoding network gradually restores the spatial resolution of the time-frequency features through upsampling and convolution operations, while combining the residual connection features of the corresponding layers in the encoding stage to achieve feature fusion and refinement. In this way, the decoder not only restores the details of the time-frequency map but also enhances the signal's local and global feature representation capabilities. Finally, the decoder outputs multi-scale time-frequency representations that meticulously characterize the signal's time-frequency structure at different scales, facilitating subsequent time-frequency mode segmentation and signal reconstruction. For example, suppose the input signal, after initial time-frequency transformation, yields a 256×256 time-frequency feature map. The encoder first compresses this to 128×128, 64×64, or even lower resolution feature maps, preserving key information from the previous layer at each step through residual connections. The decoder operates in reverse, progressively enlarging the low-resolution global features back to 256×256 while fusing corresponding features from the encoding stage, ultimately outputting a set of multi-scale time-frequency representations, such as feature maps at 256×256, 128×128, and 64×64 scales. These multi-scale feature maps capture the detailed changes and overall trends of the signal, providing rich time-frequency information support for subsequent modality segmentation. Through this hierarchical encoding and decoding design, combined with the advantages of residual connections, the system not only improves the expressive power of time-frequency features but also enhances the network's training stability and generalization ability, thereby achieving high-quality multi-scale time-frequency representations and laying a solid foundation for end-to-end time-frequency mode decomposition.

[0043] Step S20: The multi-scale time-frequency representation is segmented into instances using the time-frequency mode segmentation module of the convolutional neural network model to obtain the time-frequency modes.

[0044] like Figure 5 As shown, step S20 includes: Step S21: The structure and semantic information of the multi-scale time-frequency representation are transformed through the pixel decoding layer to obtain a high-dimensional pixel representation; Step S22: Embed the high-dimensional pixel representation through the Transformer decoding layer to obtain the time-frequency mode.

[0045] Specifically, the pixel decoding layer transforms the structural and semantic information of the multi-scale time-frequency representation to obtain a high-dimensional pixel representation (preserving structural and semantic information). The high-dimensional pixel representation is then embedded and interacted with through the Transformer decoding layer to obtain a time-frequency modality (using a set of learnable queries to embed and interact with the high-dimensional pixel representation through the Transformer decoder). The time-frequency modality segmentation module includes a pixel decoding layer and a Transformer decoding layer.

[0046] In this embodiment, a pixel decoding layer is first used to fuse and transform the multi-scale time-frequency features from the high-resolution TFA module. This pixel decoding layer, based on a multi-scale deformable attention mechanism, effectively integrates time-frequency information at different scales, fully capturing the local details and global structural features of the signal. Through this layer's processing, the structural information (such as the continuity and shape of frequency trajectories) and semantic information (such as the distinguishing features between modes) in the original multi-scale time-frequency representation are preserved and transformed into a unified high-dimensional pixel representation. This high-dimensional representation not only contains rich time-frequency features but also possesses good expressive power, facilitating subsequent mode segmentation tasks. Subsequently, this high-dimensional pixel representation is input into the Transformer decoding layer for deep embedding interaction. The Transformer decoding layer utilizes self-attention and multi-head attention mechanisms to dynamically capture long-distance dependencies and complex interaction patterns between time-frequency features. In particular, the decoding layer introduces a set of learnable query vectors. These queries, as representatives of the modes, gradually focus on and separate the individual time-frequency modes through interaction with the high-dimensional pixel representation. Each query vector corresponds to a time-frequency mode. Through iterative updates of a multi-layer Transformer decoder, the corresponding IF trajectory and its time-frequency energy distribution can be accurately located and segmented. This design enables the time-frequency mode segmentation module to automatically learn and extract the time-frequency masks of each mode from complex multi-modal time-frequency representations in an end-to-end manner, achieving accurate separation of multi-modal components in non-stationary signals. Compared with traditional segmentation methods based on thresholds or manual features, this module fully utilizes the expressive power of deep learning and the global modeling advantages of Transformer, effectively improving the accuracy and robustness of segmentation. In summary, the time-frequency mode segmentation module, through the synergistic effect of the pixel decoding layer and the Transformer decoding layer, achieves efficient fusion and mode segmentation of multi-scale time-frequency features, becoming a key component of the end-to-end time-frequency mode decomposition method of this invention.

[0047] As an example, multi-scale time-frequency features from the high-resolution TFA module are first input into the pixel decoding layer. This pixel decoding layer employs a multi-scale deformable attention mechanism, which effectively fuses time-frequency features at different scales, fully utilizing the local structure and global semantic information contained in the time-frequency image. Specifically, the pixel decoding layer transforms the multi-scale features into a unified high-dimensional pixel representation tensor by spatially and channel-weighted integration. This high-dimensional representation not only preserves the spatial structure of the input features (such as the time-frequency distribution of the signal) but also integrates semantic information (such as the feature differences between different modes), providing a rich and accurate feature foundation for subsequent mode segmentation. Subsequently, this high-dimensional pixel representation is input into the Transformer decoding layer. The Transformer decoding layer contains a set of learnable query vectors, which represent different time-frequency modes to be segmented. Through multi-head self-attention and cross-attention mechanisms, the query vectors interact with the high-dimensional pixel representation, dynamically capturing the feature distribution and interrelationships of each mode in the time-frequency image. Specifically, the query vectors focus on the salient time-frequency regions of the corresponding mode through attention calculation with the pixel representation, gradually generating the corresponding time-frequency mask. These masks can accurately segment different instantaneous frequency trajectories and modal components, achieving effective separation of complex non-stationary signals. For example, assuming the input multi-scale time-frequency features contain feature maps of three scales: 256×256, 128×128, and 64×64, the pixel decoding layer first upsamples and fuses these three sets of features to obtain a unified 256×256×C high-dimensional pixel representation (C being the number of channels). This representation includes both fine-grained local time-frequency details and coarse-grained global semantic information. Next, the Transformer decoding layer sets N learnable query vectors (e.g., N=4, representing 4 modalities). Each query interacts with the high-dimensional pixel representation through multi-head attention to generate a corresponding time-frequency mask. The final output mask clearly identifies the distribution of each modality on the time-frequency plane, achieving end-to-end time-frequency modal segmentation. By working in concert with the pixel decoding layer and the Transformer decoding layer, the time-frequency mode segmentation module of this invention can make full use of the structural and semantic information of multi-scale time-frequency features to achieve accurate segmentation of multi-modal features in complex signals, significantly improving the accuracy and robustness of mode separation.

[0048] like Figure 6 As shown, step S21 includes: Step S211: Based on the U-Net submodule in the TFA module, the structural and semantic information of the multi-scale time-frequency representation is transformed to obtain intermediate results; Step S212: The intermediate result is parsed through the pixel decoding layer to obtain a high-dimensional pixel representation.

[0049] Specifically, based on the U-Net submodule in the TFA module, the structural and semantic information of the multi-scale time-frequency representation is transformed to obtain intermediate results (considering that the U-Net submodule in the TFA module retains the structural and semantic information, its intermediate results are used as multi-scale features and input into the Mask2Former decoder). The intermediate results are then parsed by the pixel decoding layer to obtain a high-dimensional pixel representation.

[0050] In this embodiment, the time-frequency analysis module employs a residual U-Net structure to extract multi-scale features from the input complex time-domain signal. This U-Net submodule, through an encoder-decoder architecture, captures local details and global semantic information in the time-frequency representation layer by layer, achieving a hierarchical expression of time-frequency features. Specifically, the encoder extracts time-frequency features at different scales through multi-layer convolution and downsampling operations, capturing fine-grained structural information of the signal; the decoder, through upsampling and skip connections, fuses shallow spatial details and deep semantic features to recover a high-resolution time-frequency representation. The intermediate output of this U-Net submodule not only contains rich structural information (such as edges and textures of time-frequency energy distribution) but also implies the semantic hierarchy of the signal (such as time-frequency patterns of different modes), thus providing strong feature support for subsequent modality segmentation. Considering this, this invention uses the multi-scale intermediate features of the U-Net submodule as input, passing them to the Mask2Former segmentation module based on the Transformer architecture. In Mask2Former, the multi-scale time-frequency features output by U-Net are first fused and parsed through a pixel decoding layer with a multi-scale deformable attention mechanism. This pixel decoding layer can dynamically aggregate feature information at different scales to generate a unified and high-dimensional pixel embedding representation. In this way, the spatial structure and semantic information in the time-frequency representation are effectively encoded into high-dimensional feature vectors, facilitating the subsequent query mechanism for modality instance segmentation. In summary, the multi-scale time-frequency features extracted by the U-Net submodule in the TFA module not only preserve the details and semantic information of the signal but also transform them into high-dimensional pixel representations through the pixel decoding layer. This provides rich and structured input features for the Mask2Former Transformer decoder, significantly improving the accuracy and robustness of time-frequency modality segmentation.

[0051] As an example, in the time-frequency mode decomposition process of this invention, the TFA module uses a residual U-Net structure to extract multi-scale time-frequency features from the input complex time-domain signal. This U-Net structure achieves layer-by-layer refinement of the time-frequency representation through hierarchical encoding and decoding: Hierarchical encoding: The input time-frequency representation first undergoes multiple convolution and pooling operations to progressively extract low-dimensional but semantically rich feature representations. Each encoding layer not only captures local time-frequency structure information but also retains the detailed features of the previous layer through residual connections, ensuring continuous information transmission. Through multi-scale convolution kernel design, the encoder can simultaneously perceive signal features of different time windows and frequency bandwidths, enhancing its adaptability to non-stationary signals. Hierarchical decoding: The multi-scale features obtained through encoding are progressively upsampled and fused through the corresponding decoding layers to restore the spatial resolution of the time-frequency representation. During decoding, skip connections fuse the detailed features of the encoding layers with the decoding layers, enhancing the ability to reconstruct the time-frequency structure. This process achieves layer-by-layer reconstruction from coarse-grained semantics to fine-grained structure, forming a multi-scale and semantically rich intermediate time-frequency representation. The intermediate result is the multi-scale feature map of U-Net, which contains both structural information of the signal (such as the shape of the instantaneous frequency trajectory) and semantic information (such as the distinguishing features of different modes). This invention uses this multi-scale feature map as input to the pixel decoding layer of Mask2Former. In the pixel decoding layer, a multi-scale deformable attention mechanism is used to fuse and parse the multi-scale features output by U-Net, generating a unified high-dimensional pixel representation. Specifically, this includes: spatial alignment and weighted fusion of feature maps at different scales to ensure effective integration of information at different resolutions; and dynamically focusing on key regions in the time-frequency plane (such as modal peak trajectories) using the deformable attention mechanism to enhance the perception of complex modal structures. The generated high-dimensional pixel representation retains both fine-grained time-frequency structure and strong semantic distinguishing ability, providing rich contextual information for subsequent Transformer query decoders. For example, for a mechanical vibration signal, the U-Net encoder first extracts multi-scale time-frequency features, capturing modal information at different frequency bandwidths; the decoder then recovers the spatial details of the time-frequency map layer by layer, forming the intermediate result of multi-scale fusion. After the result is input into the pixel decoding layer, a high-dimensional pixel representation is generated through hierarchical fusion and attention mechanisms, which accurately reflects the time-frequency distribution characteristics of each modality and provides a solid foundation for the query-based instance segmentation of Mask2Former.

[0052] In summary, the hierarchical encoding and decoding based on the U-Net submodule in the TFA module not only realizes the multi-scale structure and semantic information conversion of time-frequency representation, but also generates high-dimensional and semantically rich pixel representation through hierarchical fusion of the pixel decoding layer, which greatly improves the accuracy and robustness of time-frequency modality segmentation.

[0053] like Figure 7As shown, in step S30, the signal reconstruction module of the convolutional neural network model performs trajectory prediction on the time-frequency mode and the complex time-domain signal to obtain signal components.

[0054] Step S30 includes: Step S31: Construct the trajectory of the time-frequency mode using the signal reconstruction module of the convolutional neural network model to obtain the initial demodulation frequency function; Step S32: Perform inverse deconstruction on the complex time-domain signal according to the initial demodulation frequency function to obtain the signal components.

[0055] Specifically, such as Figure 8 As shown, after step S32, the method further includes repeatedly updating the initial demodulation frequency function based on the alternating direction multiplier method to obtain the optimal initial demodulation frequency function; reconstructing the signal components based on the alternating direction multiplier method to obtain the optimal signal components (by repeatedly updating the demodulation frequency function and reconstructing the waveform (trajectory construction) using the alternating direction multiplier method); constructing the trajectory of the time-frequency mode through the signal reconstruction module of the convolutional neural network model to obtain the initial demodulation frequency function; and inversely deconstructing the complex time-domain signal according to the initial demodulation frequency function to obtain the signal components.

[0056] In this embodiment, firstly, a convolutional neural network model is used to construct trajectories for the segmented time-frequency modes, generating an initial demodulation frequency function. This frequency function reflects the instantaneous frequency change trend of each mode, providing accurate prior information for subsequent signal demodulation and reconstruction. Through the learning capabilities of deep networks, the dynamic frequency characteristics in complex non-stationary signals can be effectively captured, improving the accuracy and robustness of the initial frequency estimation. Subsequently, based on this initial demodulation frequency function, the alternating direction multiplier method is used to iteratively update the demodulation frequency function. ADMM decomposes the complex non-convex optimization problem into multiple easily solvable sub-problems, alternately optimizing the demodulation frequency function and signal component waveforms, gradually approaching the global optimum. In each iteration, the bandwidth minimization target constrains the spectral concentration of the signal components, ensuring that the demodulated signal satisfies the narrowest band condition, thereby achieving high-quality mode separation. During the iteration process, the updates of the demodulation frequency function and signal components complement each other: fine adjustment of the frequency function promotes accurate reconstruction of the signal waveform, while feedback from the reconstructed waveform further optimizes the estimation of the frequency function. This collaborative optimization mechanism effectively suppresses noise interference and cross-modal influences, significantly improving the purity and separation effect of signal components. Finally, after multiple rounds of ADMM iteration, the convergent optimal demodulation frequency function and corresponding signal components are obtained, achieving inverse deconstruction and accurate reconstruction of complex time-domain signals. This method not only inherits the advantages of the physical model of traditional VNCMD but also integrates the powerful expressive capabilities of deep learning, greatly improving the accuracy and stability of non-stationary signal mode decomposition. In summary, this invention, by combining trajectory construction of convolutional neural networks with iterative optimization based on the alternating direction multiplier method, achieves end-to-end closed-loop processing from time-frequency mode segmentation to signal component reconstruction, effectively solving the problems of complex parameter adjustment and limited decomposition accuracy in traditional methods, and possessing broad engineering application prospects.

[0057] As an example, trajectory construction: A convolutional neural network is used to locate the temporal peaks of the time-frequency mask, extracting the instantaneous frequency points corresponding to each moment to form the initial IF trajectory. This network enhances its robustness to noise and cross-terms through multi-layer convolution and pooling operations, ensuring the extracted trajectory is smooth and continuous. Simultaneously, the network combines the spatial structure information of the time-frequency modes to automatically correct abnormal jumps and breaks, obtaining high-quality frequency function curves. Initial demodulation frequency function generation: The extracted IF trajectory is mapped to a continuous frequency function as the initial estimate for signal demodulation. This frequency function reflects the instantaneous frequency change pattern of each mode and is a key input for subsequent signal reconstruction. Inverse deconstruction of the complex time-domain signal: Based on the initial demodulation frequency function, the original complex time-domain signal is inversely transformed using signal demodulation principles. Specifically, the alternating direction multiplier method (ADMM) is used to iteratively optimize the demodulation frequency function and signal waveform, ensuring the demodulated signal satisfies the narrowest band constraint and gradually approximates the true modal components. This process includes constructing a bandwidth minimization objective function, achieving accurate recovery of signal components by constraining the spectral concentration of the demodulated signal. Signal component output: After multiple rounds of iterative optimization, the time-domain signal components corresponding to each mode are obtained. These components have high time-frequency concentration and low noise interference, facilitating subsequent signal analysis and applications. For example, for a mechanical vibration signal, the signal reconstruction module first extracts the IF trajectories of each mode using a convolutional neural network to generate an initial demodulation frequency function. Subsequently, based on this frequency function, the original signal is inversely deconstructed, and the ADMM optimization algorithm is used for iterative updates to finally recover the time-domain waveforms of each mode. This method effectively overcomes the problems of strong dependence on initial frequency estimation and noise sensitivity in traditional mode decomposition, achieving high-precision end-to-end signal component reconstruction.

[0058] In summary, this invention achieves efficient conversion from time-frequency modes to time-domain signal components through convolutional neural network-assisted trajectory construction and model-based inverse decomposition, significantly improving the accuracy and robustness of mode decomposition.

[0059] like Figure 9 As shown, step S31 includes: Step S311: The peak value of the time-frequency mode is located by the signal reconstruction module of the convolutional neural network model to obtain multiple trajectory points; Step S312: Construct the trajectory of all trajectory points to obtain the initial demodulation frequency function.

[0060] Specifically, the peak value of the time-frequency mode is located by the signal reconstruction module of the convolutional neural network model to obtain multiple trajectory points. Trajectory is constructed from all trajectory points to obtain the initial demodulation frequency function (after obtaining an accurate IF trajectory estimate by locating the peak value of each separated time-frequency mode, the signal reconstruction module based on VNCMD takes the original time series as input and deconstructs the waveform of each signal component according to the signal demodulation principle).

[0061] In this embodiment, as Figure 10 and Figure 11 As shown, Figure 10 The table compares the TFR performance of different methods. The first row, from left to right, shows IVGNMD (Improved Variational Generalized Nonlinear Mode Decomposition), ACMD (Adaptive Chirp Mode Decomposition), RPRG+ICCD (a joint algorithm of RPRG and ICCD), and VNCMD (Variational Nonlinear Chirp Mode Decomposition). The second row, from left to right, shows SST (Synchronous Compression Transform), SET (Synchronous Extraction Transform, Time-Frequency Analysis Network), TFA-Net (Time-Frequency Analysis Network), and the TFD-Net (Convolutional Neural Network model) method proposed in this invention. The vertical axis represents frequency, and the horizontal axis represents time. Figure 11 The first row represents the signal components, and the second row represents the reconstruction error.

[0062] As an example, a convolutional neural network is first used to perform peak detection on the multimodal time-frequency mask output by the time-frequency mode segmentation module, accurately locating the instantaneous frequency trajectory points of each mode. Specifically, the network extracts local and global time-frequency information of the modes by encoding multi-scale time-frequency features layer by layer, and combined with a residual connection structure, enhances the ability to identify modal peaks in complex signals. Subsequently, based on a hierarchical decoding process, the encoded multi-scale features are restored layer by layer, refining the time-frequency resolution of peak location and ensuring the accuracy and continuity of trajectory points. Through hierarchical fusion and smoothing of all trajectory points, an initial demodulation frequency function is constructed. This function serves as a key input for signal reconstruction, reflecting the instantaneous frequency change trend of each mode. Next, based on the variational nonlinear chirped mode decomposition method, this initial demodulation frequency function is used to demodulate the original time-domain signal. The specific process includes: hierarchical encoding: multi-level feature extraction of the original time-domain signal, capturing the non-stationary characteristics and multimodal coupling information of the signal, providing rich contextual information for subsequent demodulation. Trajectory Construction and Demodulation Frequency Function Generation: Combining coding features and peak localization results, a continuous and smooth demodulation frequency function is constructed as the basis for signal demodulation. Hierarchical Decoding: Through layer-by-layer decoding, the demodulation frequency function is mapped back to each modal component of the time-domain signal, gradually restoring the signal's details and structural features. Iterative Optimization: The alternating direction multiplier method is used to repeatedly optimize the demodulation frequency function and reconstructed waveform in the hierarchical structure, gradually converging to the optimal solution that satisfies the narrowest band constraint. For example, for a mechanical vibration signal containing multimodal coupling, the signal reconstruction module first extracts multi-scale time-frequency features through hierarchical coding of a convolutional network, locating the IF peak trajectory points of each mode. Subsequently, the hierarchical decoding process refines the time-frequency resolution of the trajectory points, constructing a smooth demodulation frequency function. Based on this function, the VNCMD method demodulates the original signal, recovering the waveforms of each mode layer by layer, ultimately achieving high-precision mode separation and signal reconstruction.

[0063] In summary, this invention, by combining the hierarchical encoding and decoding structure of convolutional neural networks with a model-based VNCMD signal reconstruction method, achieves an end-to-end process from time-frequency mode peak localization to accurate demodulation frequency function construction, and then to high-quality signal component recovery, significantly improving the accuracy and robustness of non-stationary signal mode decomposition.

[0064] like Figure 12 As shown, step S31 includes: Step S311: Based on the signal reconstruction module, the complex time-domain signal is analyzed to obtain the original time series; Step S312: Perform inverse deconstruction on the original time series according to the initial demodulation frequency function to obtain the signal components.

[0065] Specifically, based on the signal reconstruction module, the complex time-domain signal is analyzed to obtain the original time series (the signal reconstruction module based on VNCMD takes the original time series as input and deconstructs the waveforms of each signal component according to the signal demodulation principle), and the original time series is deconstructed according to the initial demodulation frequency function to obtain the signal components.

[0066] In this embodiment, the signal reconstruction module, as a key step in the time-frequency mode decomposition process, is responsible for converting the separated time-frequency mode information into specific time-domain signal components. This module takes the complex time-domain signal and its corresponding initial demodulation frequency function as input, and systematically analyzes and reconstructs each mode component of the signal using a variational nonlinear chirped mode decomposition method.

[0067] In this embodiment, the signal reconstruction module first receives the initial demodulation frequency function output by the time-frequency mode segmentation module. This function reflects the instantaneous frequency trajectory of each mode and is crucial prior information for mode separation. Subsequently, the module uses the original complex time-domain signal as a basis and performs inverse processing on it using signal demodulation principles. This process includes frequency compensation and bandwidth compression of the signal using the initial demodulation frequency function, gradually stripping away the frequency components of each mode to achieve inverse signal deconstruction. Based on this, the module employs the Alternating Direction Multiplier Method (ADMM) optimization strategy to iteratively update the demodulation frequency function and signal component waveforms. By minimizing the bandwidth of the demodulated signal, it ensures that each component meets the narrowest-band constraint, thereby improving the purity and separation accuracy of the components. This iterative process not only enhances the time-frequency concentration of the mode components but also effectively suppresses noise and cross-interference, ensuring the stability and accuracy of signal reconstruction. Finally, the signal reconstruction module outputs individual time-domain signal components, which correspond to different mode components in the original complex signal. Through the analysis and reconstruction of this module, the entire time-frequency mode decomposition method realizes the end-to-end conversion from complex aliased signals to clearly separated modes, which greatly improves the performance and application value of non-stationary signal processing.

[0068] As an example, the input complex time-domain signal is first processed to convert it into the original time-series signal. This step is equivalent to restoring the complex time-frequency information output by the deep learning model back to the waveform in the time domain, ensuring that subsequent mode decomposition can be based on the real time signal. Next, using the variational nonlinear chirped mode decomposition (VNCMD) method, the initial demodulation frequency function obtained from the previous time-frequency mode segmentation module is used as a reference. This frequency function reflects the instantaneous frequency change trend of each mode signal. The signal reconstruction module performs "inverse deconstruction" processing on the original time series according to the principle of signal demodulation, that is, by adjusting the frequency components of the signal, the frequency features of each mode are extracted separately. Specifically, for each mode signal, the system uses its corresponding initial frequency information to "flatten" the frequency components of that mode in the mixed signal to near the baseband, facilitating subsequent filtering and separation. Then, through repeated iterative optimization, the frequency function and signal waveform are gradually adjusted to minimize the frequency bandwidth of each mode signal, achieving a clearer separation effect. Ultimately, through multiple iterations and optimizations, the system can accurately separate the time waveforms of each mode from the original complex time-domain signal. Each separated signal component corresponds to an independent mode, possesses clear time-frequency characteristics, and noise interference is effectively suppressed, thus achieving high-quality mode decomposition and signal recovery. For example, suppose the input signal is composed of multiple signals whose frequencies vary over time and contains noise. Through the above signal reconstruction process, the system can gradually extract the signal components corresponding to each frequency change trajectory based on the initial frequency estimate, recovering a clear single-mode waveform for convenient subsequent analysis.

[0069] Furthermore, such as Figure 13 As shown, based on the above-described deep learning-based time-frequency mode decomposition method, this invention also provides a deep learning-based time-frequency mode decomposition system, wherein the deep learning-based time-frequency mode decomposition system includes: The complex time-domain signal analysis module 50 is used to acquire the input complex time-domain signal and analyze the complex time-domain signal through the high-resolution time-frequency analysis module of the convolutional neural network model to obtain a multi-scale time-frequency representation; The multi-scale time-frequency representation segmentation module 60 is used to perform instance segmentation on the multi-scale time-frequency representation through the time-frequency mode segmentation module of the convolutional neural network model to obtain time-frequency modes; The signal component construction module 70 is used to predict the trajectory of the time-frequency mode and the complex time-domain signal through the signal reconstruction module of the convolutional neural network model to obtain the signal components.

[0070] like Figure 14As shown in this embodiment of the time-frequency mode decomposition system based on deep learning, in this embodiment, the complex time-domain signal analysis module 50 includes: The complex time-domain signal unit 501 is used to acquire the input complex time-domain signal, and to initialize the complex time-domain signal through multiple complex convolutions of the convolutional neural network model and the short-time filtering framework of STFT to obtain an initial time-frequency representation; The initial time-frequency representation weighting unit 502 is used to weight the initial time-frequency representation through the channel attention of the convolutional neural network model to obtain time-frequency features; The hierarchical transformation unit 503 is used to perform hierarchical transformation on the time-frequency features through the multi-scale feature extractor of the convolutional neural network model to obtain a multi-scale time-frequency representation.

[0071] In this embodiment, the hierarchical transformation unit 503 includes: The hierarchical coding subunit 5031 is used to hierarchically encode the time-frequency features through the residual connection network of the multi-scale feature extractor to obtain a globally defined representation; The global definition representation decoding subunit 5032 is used to perform hierarchical decoding of the global definition representation through the residual connection network to obtain a multi-scale time-frequency representation.

[0072] In this embodiment, the multi-scale time-frequency representation segmentation module includes: The semantic conversion unit 601 is used to convert the structure and semantic information of the multi-scale time-frequency representation through the pixel decoding layer of the time-frequency modality segmentation module to obtain a high-dimensional pixel representation; The embedded interaction unit 602 is used to perform embedded interaction on the high-dimensional pixel representation through the Transformer decoding layer of the time-frequency modality segmentation module to obtain the time-frequency modality.

[0073] In this embodiment, the semantic conversion unit 601 includes: The multi-scale time-frequency representation conversion subunit 6011 is used to convert the structural and semantic information of the multi-scale time-frequency representation based on the U-Net submodule in the TFA module to obtain intermediate results; The intermediate result parsing subunit 6012 is used to parse the intermediate result through the pixel decoding layer to obtain a high-dimensional pixel representation.

[0074] In this embodiment, the signal component construction module 70 includes: The time-frequency mode reconstruction unit 701 is used to construct the trajectory of the time-frequency mode through the signal reconstruction module of the convolutional neural network model to obtain the initial demodulation frequency function. The complex time-domain signal inverse deconstruction unit 702 is used to inversely deconstruct the complex time-domain signal according to the initial demodulation frequency function to obtain signal components.

[0075] In this embodiment, the time-frequency mode reconstruction unit 701 includes: The trajectory point localization subunit 7011 is used to locate the peak of the time-frequency mode through the signal reconstruction module of the convolutional neural network model to obtain multiple trajectory points; The trajectory construction subunit 7012 is used to construct trajectories for all trajectory points to obtain the initial demodulation frequency function.

[0076] In this embodiment, the complex time-domain signal inverse deconstruction unit 702 includes: The complex time-domain signal analysis subunit 7021 is used to analyze the complex time-domain signal based on the signal reconstruction module to obtain the original time series; The original time series inverse deconstruction subunit 7022 is used to inversely deconstruct the original time series according to the initial demodulation frequency function to obtain signal components.

[0077] Furthermore, such as Figure 15 As shown, based on the above-mentioned deep learning-based time-frequency mode decomposition method and system, the present invention also provides a terminal, which includes a processor 10, a memory 20 and a display 30. Figure 15 Only some of the terminal components are shown; however, it should be understood that it is not required to implement all of the components shown, and more or fewer components may be implemented instead.

[0078] In some embodiments, the memory 20 may be an internal storage unit of the terminal, such as a hard disk or memory. In other embodiments, the memory 20 may be an external storage device of the terminal, such as a plug-in hard disk, smart media card (SMC), secure digital card (SD), flash card, etc. Further, the memory 20 may include both internal and external storage devices. The memory 20 is used to store application software and various types of data installed on the terminal, such as the program code installed on the terminal. The memory 20 can also be used to temporarily store data that has been output or will be output. In one embodiment, the memory 20 stores a deep learning-based time-frequency mode decomposition program 40, which can be executed by the processor 10 to implement the deep learning-based time-frequency mode decomposition method of this application.

[0079] In some embodiments, the processor 10 may be a central processing unit (CPU), a microprocessor, or other data processing chip, used to run program code stored in the memory 20 or process data, such as executing the deep learning-based time-frequency mode decomposition method.

[0080] In some embodiments, the display 30 may be an LED display, a liquid crystal display, a touch-sensitive liquid crystal display, or an OLED (Organic Light-Emitting Diode) touchscreen. The display 30 is used to display information on the terminal and to display a visual user interface. The terminals communicate with each other via a system bus.

[0081] In one embodiment, when the processor 10 executes the deep learning-based time-frequency mode decomposition program 40 in the memory 20, the following steps are performed: The input complex time-domain signal is acquired, and the complex time-domain signal is analyzed by the high-resolution time-frequency analysis module of the convolutional neural network model to obtain a multi-scale time-frequency representation; The multi-scale time-frequency representation is segmented into instances using the time-frequency mode segmentation module of the convolutional neural network model to obtain the time-frequency modes; The signal reconstruction module of the convolutional neural network model performs trajectory prediction on the time-frequency mode and the complex time-domain signal to obtain signal components.

[0082] Specifically, the acquisition of the input complex time-domain signal, and the analysis of the complex time-domain signal by the high-resolution time-frequency analysis module of the convolutional neural network model to obtain a multi-scale time-frequency representation, includes: The input complex time-domain signal is acquired, and the complex time-domain signal is initialized by multiple complex convolutions of the convolutional neural network model and the short-time filtering framework of STFT to obtain the initial time-frequency representation. The initial time-frequency representation is weighted by the channel attention of a convolutional neural network model to obtain time-frequency features; The time-frequency features are subjected to hierarchical transformation by a multi-scale feature extractor of a convolutional neural network model to obtain a multi-scale time-frequency representation.

[0083] The complex time-domain signals include radar signals and communication signals.

[0084] The hierarchical transformation includes hierarchical encoding and hierarchical decoding; The process of performing a hierarchical transformation on the time-frequency features using a multi-scale feature extractor based on a convolutional neural network model to obtain a multi-scale time-frequency representation specifically includes: The time-frequency features are hierarchically encoded using the residual connection network of the multi-scale feature extractor to obtain a globally defined representation; The global definition representation is decoded hierarchically through the residual connection network to obtain a multi-scale time-frequency representation.

[0085] The time-frequency mode segmentation module includes a pixel decoding layer and a Transformer decoding layer; The step of segmenting the multi-scale time-frequency representation using a time-frequency mode segmentation module based on a convolutional neural network model to obtain time-frequency modes specifically includes: The pixel decoding layer transforms the structural and semantic information of the multi-scale time-frequency representation to obtain a high-dimensional pixel representation; The time-frequency mode is obtained by embedding the high-dimensional pixel representation through the Transformer decoding layer.

[0086] Specifically, the step of converting the structural and semantic information of the multi-scale time-frequency representation through the pixel decoding layer to obtain a high-dimensional pixel representation includes: Based on the U-Net submodule in the TFA module, the structural and semantic information of the multi-scale time-frequency representation is transformed to obtain intermediate results; The intermediate results are parsed by the pixel decoding layer to obtain a high-dimensional pixel representation.

[0087] Specifically, the signal reconstruction module using a convolutional neural network model performs trajectory prediction on the time-frequency mode and the complex time-domain signal to obtain signal components, which includes: The signal reconstruction module of the convolutional neural network model constructs the trajectory of the time-frequency mode to obtain the initial demodulation frequency function. The complex time-domain signal is inversely deconstructed based on the initial demodulation frequency function to obtain the signal components.

[0088] Specifically, the signal reconstruction module using the convolutional neural network model constructs the trajectory of the time-frequency mode to obtain the initial demodulation frequency function, which includes: The peak value of the time-frequency mode is located by the signal reconstruction module of the convolutional neural network model, and multiple trajectory points are obtained. By constructing trajectories from all trajectory points, the initial demodulation frequency function is obtained.

[0089] Specifically, the step of inversely deconstructing the complex time-domain signal according to the initial demodulation frequency function to obtain signal components includes: Based on the signal reconstruction module, the complex time-domain signal is analyzed to obtain the original time series; The original time series is deconstructed inversely based on the initial demodulation frequency function to obtain the signal components.

[0090] The step of inversely deconstructing the complex time-domain signal according to the initial demodulation frequency function to obtain signal components further includes: Based on the alternating direction multiplier method, the initial demodulation frequency function is repeatedly updated to obtain the optimal initial demodulation frequency function. The signal components are reconstructed using the alternating direction multiplier method to obtain the optimal signal components.

[0091] The present invention also provides a computer-readable storage medium, wherein the computer-readable storage medium stores a deep learning-based time-frequency mode decomposition program, which, when executed by a processor, implements the steps of the deep learning-based time-frequency mode decomposition method as described above.

[0092] In summary, this invention provides a time-frequency mode decomposition method, system, terminal, and storage medium based on deep learning. The method includes: acquiring an input complex time-domain signal; analyzing the complex time-domain signal using a high-resolution time-frequency analysis module of a convolutional neural network model to obtain a multi-scale time-frequency representation; performing instance segmentation on the multi-scale time-frequency representation using a time-frequency mode segmentation module of the convolutional neural network model to obtain time-frequency modes; and performing trajectory prediction on the time-frequency modes and the complex time-domain signal using a signal reconstruction module of the convolutional neural network model to obtain signal components. This invention separates and predicts the trajectory of aliased components in a complex time-domain signal based on high-resolution time-frequency analysis and time-frequency mode segmentation, achieving accurate extraction of signal components.

[0093] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or terminal system that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or terminal system. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or terminal system that includes that element.

[0094] Of course, those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware (such as a processor, controller, etc.). The program can be stored in a computer-readable storage medium, and when executed, it can include the processes described in the above method embodiments. The computer-readable storage medium can be a memory, magnetic disk, optical disk, etc.

[0095] It should be understood that the application of the present invention is not limited to the examples above. Those skilled in the art can make improvements or modifications based on the above description, and all such improvements and modifications should fall within the protection scope of the appended claims.

Claims

1. A time-frequency mode decomposition method based on deep learning, characterized in that, The deep learning-based time-frequency mode decomposition method includes: The input complex time-domain signal is acquired, and the complex time-domain signal is analyzed by the high-resolution time-frequency analysis module of the convolutional neural network model to obtain a multi-scale time-frequency representation; The multi-scale time-frequency representation is segmented into instances using the time-frequency mode segmentation module of the convolutional neural network model to obtain the time-frequency modes; The signal reconstruction module of the convolutional neural network model performs trajectory prediction on the time-frequency mode and the complex time-domain signal to obtain signal components.

2. The time-frequency mode decomposition method based on deep learning according to claim 1, characterized in that, The process of acquiring the input complex time-domain signal and analyzing it using a high-resolution time-frequency analysis module of a convolutional neural network model to obtain a multi-scale time-frequency representation specifically includes: The input complex time-domain signal is acquired, and the complex time-domain signal is initialized by multiple complex convolutions of the convolutional neural network model and the short-time filtering framework of STFT to obtain the initial time-frequency representation. The initial time-frequency representation is weighted by the channel attention of a convolutional neural network model to obtain time-frequency features; The time-frequency features are subjected to hierarchical transformation by a multi-scale feature extractor of a convolutional neural network model to obtain a multi-scale time-frequency representation.

3. The time-frequency mode decomposition method based on deep learning according to claim 2, characterized in that, The complex time-domain signals include radar signals and communication signals.

4. The time-frequency mode decomposition method based on deep learning according to claim 2, characterized in that, The hierarchical transformation includes hierarchical encoding and hierarchical decoding; The process of performing a hierarchical transformation on the time-frequency features using a multi-scale feature extractor based on a convolutional neural network model to obtain a multi-scale time-frequency representation specifically includes: The time-frequency features are hierarchically encoded using the residual connection network of the multi-scale feature extractor to obtain a globally defined representation; The global definition representation is decoded hierarchically through the residual connection network to obtain a multi-scale time-frequency representation.

5. The time-frequency mode decomposition method based on deep learning according to claim 1, characterized in that, The time-frequency mode segmentation module includes a pixel decoding layer and a Transformer decoding layer; The step of segmenting the multi-scale time-frequency representation using a time-frequency mode segmentation module based on a convolutional neural network model to obtain time-frequency modes specifically includes: The pixel decoding layer transforms the structural and semantic information of the multi-scale time-frequency representation to obtain a high-dimensional pixel representation; The time-frequency mode is obtained by embedding the high-dimensional pixel representation through the Transformer decoding layer.

6. The time-frequency mode decomposition method based on deep learning according to claim 5, characterized in that, The step of converting the structural and semantic information of the multi-scale time-frequency representation through the pixel decoding layer to obtain a high-dimensional pixel representation specifically includes: Based on the U-Net submodule in the TFA module, the structural and semantic information of the multi-scale time-frequency representation is transformed to obtain intermediate results; The intermediate results are parsed by the pixel decoding layer to obtain a high-dimensional pixel representation.

7. The time-frequency mode decomposition method based on deep learning according to claim 1, characterized in that, The signal reconstruction module using a convolutional neural network model performs trajectory prediction on the time-frequency mode and the complex time-domain signal to obtain signal components, specifically including: The signal reconstruction module of the convolutional neural network model constructs the trajectory of the time-frequency mode to obtain the initial demodulation frequency function. The complex time-domain signal is inversely deconstructed based on the initial demodulation frequency function to obtain the signal components.

8. The time-frequency mode decomposition method based on deep learning according to claim 7, characterized in that, The signal reconstruction module using a convolutional neural network model constructs the trajectory of the time-frequency mode to obtain the initial demodulation frequency function, specifically including: The peak value of the time-frequency mode is located by the signal reconstruction module of the convolutional neural network model, and multiple trajectory points are obtained. By constructing trajectories from all trajectory points, the initial demodulation frequency function is obtained.

9. The time-frequency mode decomposition method based on deep learning according to claim 7, characterized in that, The step of inversely deconstructing the complex time-domain signal according to the initial demodulation frequency function to obtain signal components specifically includes: Based on the signal reconstruction module, the complex time-domain signal is analyzed to obtain the original time series; The original time series is deconstructed inversely based on the initial demodulation frequency function to obtain the signal components.

10. The time-frequency mode decomposition method based on deep learning according to claim 8, characterized in that, The step of inversely deconstructing the complex time-domain signal according to the initial demodulation frequency function to obtain signal components further includes: Based on the alternating direction multiplier method, the initial demodulation frequency function is repeatedly updated to obtain the optimal initial demodulation frequency function. The signal components are reconstructed using the alternating direction multiplier method to obtain the optimal signal components.

11. A time-frequency mode decomposition system based on deep learning, characterized in that, The deep learning-based time-frequency mode decomposition system includes: The complex time-domain signal analysis module is used to acquire the input complex time-domain signal and analyze the complex time-domain signal through the high-resolution time-frequency analysis module of the convolutional neural network model to obtain a multi-scale time-frequency representation; The multi-scale time-frequency representation segmentation module is used to perform instance segmentation on the multi-scale time-frequency representation through the time-frequency mode segmentation module of the convolutional neural network model to obtain the time-frequency mode; The signal component construction module is used to predict the trajectory of the time-frequency mode and the complex time-domain signal through the signal reconstruction module of the convolutional neural network model to obtain the signal components.

12. The time-frequency mode decomposition system based on deep learning according to claim 11, characterized in that, The complex time-domain signal analysis module includes: The complex time-domain signal unit is used to acquire the input complex time-domain signal, and to initialize the complex time-domain signal through multiple complex convolutions of the convolutional neural network model and the short-time filtering framework of STFT to obtain the initial time-frequency representation; The initial time-frequency representation weighting unit is used to weight the initial time-frequency representation through the channel attention of the convolutional neural network model to obtain time-frequency features; The hierarchical transformation unit is used to perform hierarchical transformation on the time-frequency features through the multi-scale feature extractor of the convolutional neural network model to obtain a multi-scale time-frequency representation.

13. The time-frequency mode decomposition system based on deep learning according to claim 12, characterized in that, The hierarchical transformation unit includes: The hierarchical coding subunit is used to hierarchically encode the time-frequency features through the residual connection network of the multi-scale feature extractor to obtain a globally defined representation; The global definition representation decoding subunit is used to perform hierarchical decoding of the global definition representation through the residual connection network to obtain a multi-scale time-frequency representation.

14. The time-frequency mode decomposition system based on deep learning according to claim 11, characterized in that, The multi-scale time-frequency representation segmentation module includes: The semantic conversion unit is used to convert the structural and semantic information of the multi-scale time-frequency representation through the pixel decoding layer of the time-frequency modality segmentation module to obtain a high-dimensional pixel representation; An embedded interaction unit is used to perform embedded interaction on the high-dimensional pixel representation through the Transformer decoding layer of the time-frequency modality segmentation module to obtain the time-frequency modality.

15. The time-frequency mode decomposition system based on deep learning according to claim 14, characterized in that, The semantic conversion unit includes: The multi-scale time-frequency representation conversion subunit is used to convert the structural and semantic information of the multi-scale time-frequency representation based on the U-Net submodule in the TFA module to obtain intermediate results; The intermediate result parsing subunit is used to parse the intermediate result through the pixel decoding layer to obtain a high-dimensional pixel representation.

16. The time-frequency mode decomposition system based on deep learning according to claim 11, characterized in that, The signal component construction module includes: The time-frequency mode reconstruction unit is used to construct the trajectory of the time-frequency mode through the signal reconstruction module of the convolutional neural network model to obtain the initial demodulation frequency function; The complex time-domain signal inverse deconstruction unit is used to inversely deconstruct the complex time-domain signal according to the initial demodulation frequency function to obtain signal components.

17. The time-frequency mode decomposition system based on deep learning according to claim 16, characterized in that, The time-frequency mode reconstruction unit includes: The trajectory point localization subunit is used to locate the peak of the time-frequency mode through the signal reconstruction module of the convolutional neural network model to obtain multiple trajectory points; The trajectory construction sub-unit is used to construct trajectories for all trajectory points to obtain the initial demodulation frequency function.

18. The time-frequency mode decomposition system based on deep learning according to claim 16, characterized in that, The complex time-domain signal inverse deconstruction unit includes: The complex time-domain signal analysis subunit is used to analyze the complex time-domain signal based on the signal reconstruction module to obtain the original time series; The original time series inverse deconstruction subunit is used to inversely deconstruct the original time series according to the initial demodulation frequency function to obtain signal components.

19. A terminal, characterized in that, The terminal includes: a memory, a processor, and a deep learning-based time-frequency mode decomposition program stored in the memory and executable on the processor. When the deep learning-based time-frequency mode decomposition program is executed by the processor, it implements the steps of the deep learning-based time-frequency mode decomposition method as described in any one of claims 1-10.

20. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a deep learning-based time-frequency mode decomposition program, which, when executed by a processor, implements the steps of the deep learning-based time-frequency mode decomposition method as described in any one of claims 1-10.