Intelligent sound signal sensing method and system based on time-frequency feature fusion

By designing a dual-branch network architecture that fuses time and frequency features, the problem of high cost relying on large-scale pre-training and single-domain modeling in existing technologies is solved, achieving efficient audio recognition on resource-constrained devices and improving the accuracy of audio scene recognition.

CN120895051AActive Publication Date: 2025-11-04ZHEJIANG UNIV +1

Patent Information

Application Number
CN202511437990.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-10
Publication Date
2025-11-04
Estimated Expiration
2045-10-10

AI Technical Summary

Technical Problem

Existing high-performance audio signal models rely on large-scale pre-training, which is costly, and single-domain modeling methods cannot fully utilize the complementary information of audio signals, especially the phase information in the frequency domain, which limits the improvement of model performance.

Method used

The design employs a dual-branch network architecture that fuses time and frequency features. By processing in parallel in the frequency and time domains, it extracts deep features in the frequency and time domains respectively, and then splices and fuses them along the channel dimension. Finally, it inputs the data into a fully connected layer for audio scene recognition.

Benefits of technology

It significantly improves the accuracy of audio recognition in end-to-end training scenarios, reduces the dependence on large-scale pre-training data, is suitable for resource-constrained devices, and improves the adaptability and efficiency of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120895051A_ABST
    Figure CN120895051A_ABST
Patent Text Reader

Abstract

The invention discloses an intelligent sound signal sensing method and system based on time-frequency feature fusion, and relates to the technical field of intelligent sound signal processing, and the method comprises the steps: obtaining an original audio signal; performing short-time Fourier transform on the original audio signal to obtain a plurality of frequency spectrum features; inputting the plurality of frequency spectrum features into a plurality of neural networks for processing, and extracting frequency domain depth features; processing the original audio signal by using a one-dimensional convolutional neural network, and extracting time domain depth features; splicing and fusing the frequency domain depth feature and the time domain depth feature in a channel dimension to obtain a time-frequency fusion feature; and inputting the time-frequency fusion feature into a full connection layer classifier to obtain a final audio scene recognition result. According to the invention, by designing a double-branch network architecture of time domain and frequency domain parallel processing, time domain waveform information and frequency domain amplitude and phase information are effectively fused, and in an end-to-end training scene, the accuracy of audio recognition is remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of intelligent sound signal processing, and more particularly to an intelligent sound signal perception method and system based on time-frequency feature fusion. BACKGROUND

[0002] Acoustic scene recognition (ASR) is one of the important tasks of environmental sound understanding. In recent years, multi-modal or self-supervised models based on large-scale dataset pre-training have achieved remarkable results on audio classification tasks. For example, the method based on multi-modal data and pre-training proposed by Siddharth et al. achieves near-perfect classification accuracy on the standard ESC-50 dataset. However, the success of such models is heavily dependent on massive training data and huge computing resources, and the high training and deployment costs limit their research and application in resource-constrained scenarios, such as edge computing devices or mobile terminals. Therefore, exploring how to design more efficient and powerful models in the setting of end-to-end training only using the target task dataset without pre-training still has important academic value and practical significance.

[0003] Traditional audio modeling methods usually focus on single-dimensional features. One class of methods is based on frequency domain features, which can effectively capture the spectral structure of sound, but usually ignores the phase information, while the phase information is crucial for distinguishing certain fine acoustic scenes. Another class of methods is to model the original time-domain waveform directly, which can theoretically preserve the most complete information. However, it is necessary to learn multi-scale acoustic features from the original waveform, which puts higher requirements on network structure design and optimization, and may be difficult to directly learn the intuitive harmonic structure in the frequency domain.

[0004] It can be seen that the existing technology mainly has the following problems: the existing high-performance model is heavily dependent on large-scale pre-training, with high cost, and the model parameter quantity is usually large, which is difficult to apply on resource-constrained devices; the modeling method of a single domain (time domain or frequency domain) cannot fully utilize the complementary information contained in the audio signal, especially the phase information in the frequency domain is often ignored, which limits the further improvement of the model performance.

[0005] Therefore, how to provide an intelligent sound signal perception method and system based on time-frequency feature fusion, efficiently fuse time domain and frequency domain information, and without pre-training, to improve the accuracy of audio recognition is a problem that needs to be solved by those skilled in the art. SUMMARY

[0006] Therefore, the application provides an intelligent sound signal perception method and system based on time-frequency feature fusion.

[0007] To achieve the above object, the application adopts the following technical scheme: an intelligent sound signal perception method based on time-frequency feature fusion, comprising: The original audio signal is subjected to frequency domain feature extraction and time domain feature extraction respectively to obtain frequency domain deep features and time domain deep features. The frequency domain feature extraction of the original audio signal comprises: The original audio signal is subjected to short-time Fourier transform to obtain complex spectrum features. The complex spectrum features are input into a complex neural network for processing to extract frequency domain deep features. The time domain feature extraction of the original audio signal comprises: The original audio signal is processed by a one-dimensional convolutional neural network to extract time domain deep features. The frequency domain deep features and the time domain deep features are spliced and fused in the channel dimension to obtain time-frequency fusion features. The time-frequency fusion features are input into a fully connected layer classifier to obtain the final audio scene recognition result.

[0008] Preferably, the complex neural network is stacked by a plurality of complex residual modules, and the complex residual module comprises a complex convolutional layer, a complex batch normalization layer, a complex activation function and a complex attention module connected in sequence.

[0009] Preferably, the method further comprises: After obtaining the original audio signal, the original audio signal is resampled to obtain sampling features.

[0010] Preferably, the short-time Fourier transform of the original audio signal to obtain complex spectrum features comprises: The original audio signal is subjected to short-time Fourier transform by using a preset window function, Fourier transform point number and frame shift length to obtain a complex tensor containing real and imaginary parts. The real and imaginary parts of the complex tensor are subjected to global normalization processing by using pre-calculated global mean and standard deviation to obtain complex spectrum features.

[0011] Preferably, the complex feature map is input into a complex convolutional layer for processing, which is represented as: ; in, and Represents complex convolution kernel The real and imaginary parts; and Representing the first The characteristics of the real and imaginary parts output by each complex residual module. and They represent and Input the real and imaginary features of the output of the complex convolutional layer; The complex batch normalization layer and Normalization is performed to obtain normalized features. and ; The complex activation function calculates the modulus of the normalized feature. and scaling factor The calculated scaling factor is then applied to both the real and imaginary parts of the input features, resulting in: ; ; in, and express and Real and imaginary part features after complex convolutional layers, complex normalization layers, and complex activation functions; The complex attention module is based on and The modulus is used to obtain the complex channel weights: ; ; in, This is the first The input feature maps of each complex residual module are dynamically generated, with complex weights for each channel. and It consists of real and imaginary parts; and express and Real and imaginary features after complex convolutional layers, complex normalization layers, complex activation layers, and complex attention layers; Output features Add each term to the residual term and perform pooling to obtain the first term. Complex feature map after processing by each module .

[0012] Preferably, the complex feature map output by the last complex residual module is subjected to global average pooling to obtain real part and imaginary part vectors and ; The spatial dimension is processed to obtain flattened real part and imaginary part vectors and ; The flattened real part and imaginary part vectors are spliced in the channel dimension, and the spliced feature vector is input into a fully connected layer to obtain a final frequency domain deep feature vector.

[0013] Preferably, the one-dimensional convolutional neural network comprises a one-dimensional convolutional layer, an anti-aliasing pooling layer, a one-dimensional residual module and a time series aggregation module connected in sequence.

[0014] Preferably, the one-dimensional convolutional layer is used for preliminary feature extraction of the input waveform. The anti-aliasing pooling layer is used for down-sampling. The one-dimensional residual module is used for extracting deep time domain features. The time series aggregation module is used for aggregating sequence features and generating time domain deep features.

[0015] Preferably, an intelligent sound signal perception system based on time-frequency feature fusion comprises: A signal acquisition module is configured to acquire an original audio signal. A feature extraction module is configured to perform frequency domain feature extraction and time domain feature extraction on the original audio signal respectively to obtain frequency domain deep features and time domain deep features. A feature fusion module is configured to splice and fuse the frequency domain deep features and the time domain deep features in the channel dimension to obtain time-frequency fusion features. A scene recognition module is configured to input the time-frequency fusion features into a fully connected layer classifier to obtain a final audio scene recognition result.

[0016] According to the above technical solution, compared with the prior art, the present disclosure provides an intelligent sound signal perception method and system based on time-frequency feature fusion, which has the following beneficial effects: (1) Through the architecture of parallel processing of time domain and frequency domain double branches, the complementarity of original waveform information and complex spectrum (including amplitude and phase) information is fully utilized, and the comprehensive understanding ability of the model for audio signals is enhanced.

[0017] (2) The frequency domain branch adopts a full complex neural network design, including complex convolution, complex normalization, complex activation and complex attention mechanism, which can directly process and model the phase information in the complex spectrum, and improve the recognition degree of the fine structure of the audio.

[0018] (3) The application focuses on an end-to-end scene, is different from a path relying on a large-scale pre-training model, gets rid of the dependence on large-scale pre-training data, significantly reduces the resource and data cost of model training, makes it more adaptive in resource-limited research and application scenes, effectively improves the classification accuracy in the end-to-end scene, and has higher efficiency and practicability. BRIEF DESCRIPTION OF DRAWINGS

[0019] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings needed to be used in the embodiments or prior art description will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor on the basis of the provided drawings.

[0020] Figure 1 A flow chart of an intelligent sound signal perception method based on time-frequency feature fusion is provided for the present application.

[0021] Figure 2 The overall architecture schematic diagram of a time-frequency fusion neural network (FusionNet) provided for the embodiments of the present application.

[0022] Figure 3 The internal structure schematic diagram of a complex residual module (ComplexResBlock) in the frequency domain branch provided for the embodiments of the present application. DETAILED DESCRIPTION

[0023] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, not all. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.

[0024] The embodiments of the present application disclose an intelligent sound signal perception method based on time-frequency feature fusion, as shown in Figure 1 The method comprises the following steps: S100, acquiring an original audio signal; S200, respectively performing frequency domain feature extraction and time domain feature extraction on the original audio signal to obtain frequency domain deep features and time domain deep features; The frequency domain feature extraction on the original audio signal comprises: Performing short-time Fourier transform on the original audio signal to obtain complex spectrum features; Inputting the complex spectrum features into a complex neural network for processing to extract frequency domain deep features; The original audio signal is subjected to time domain feature extraction, comprising: The original audio signal is processed by using a one-dimensional convolutional neural network to extract time domain deep features. S300, the frequency domain deep features and the time domain deep features are spliced and fused in the channel dimension to obtain time-frequency fusion features. S400, the time-frequency fusion features are input into a full connection layer classifier to obtain a final audio scene recognition result.

[0025] The embodiment of the application parallelly establishes a time domain feature extraction branch and a frequency domain feature extraction branch for the input original audio waveform signal; in the frequency domain branch, the original audio signal is subjected to short-time Fourier transform (STFT) to obtain complex spectrum features, and the complex spectrum features are input into a complex neural network for processing to extract frequency domain deep features; in the time domain branch, a one-dimensional convolutional neural network is used to process the original audio signal to extract time domain deep features.

[0026] The embodiment of the application designs a parallel double-branch neural network architecture, respectively models the original waveform and complex spectrum of the audio in depth, and effectively fuses the features of the two, aiming to improve the accuracy and robustness of audio scene recognition without relying on large-scale pre-training.

[0027] Specifically, an original audio waveform signal is obtained and resampled to a specific frequency to obtain a sampling feature .

[0028] Specifically, in the frequency domain network, the original audio signal is subjected to short-time Fourier transform to obtain complex spectrum features, comprising: The original audio signal is subjected to short-time Fourier transform by using a preset window function, Fourier transform point number and frame shift length to obtain a complex tensor containing real and imaginary parts; The real and imaginary parts of the complex tensor are respectively subjected to global normalization processing by using pre-computed global mean and standard deviation to obtain complex spectrum features.

[0029] Specifically, the sampling feature is subjected to short-time Fourier transform to obtain a complex spectrum , the formula is as follows: ; Wherein, The sampling feature is represented as: is a complex matrix, represented as: ; is an imaginary unit, and These are the real and imaginary part matrices, respectively; Global normalization is performed on the real and imaginary matrices respectively to obtain the complex spectral characteristics.

[0030] Global normalization is performed on the real and imaginary parts separately to stabilize model training: ; in, , This represents the global mean and standard deviation of the real part of the STFT, which are pre-computed over the entire training dataset. , This represents the global mean and standard deviation of the imaginary part of the STFT, pre-computed across the entire training dataset. The input to the frequency domain network is a complex feature map. : ; First, we pass through the main part of the frequency domain network, namely... k A continuous complex residual module.

[0031] Specifically, the complex neural network is composed of multiple complex residual modules (ComplexResBlock), each of which includes a complex convolutional layer (ComplexConv2d), a complex batch normalization layer (ComplexBatchNorm2d), a complex activation function (ComplexModReLU), and a complex attention module (ComplexSEBlock) connected in sequence.

[0032] Each complex residual module receives the output of the previous complex residual module and generates a new, more abstract complex feature map. This embodiment uses the first... k This will be illustrated using a complex residual module as an example.

[0033] No. Complex feature maps input from each complex residual module Its internal calculation process consists of four steps. Residual connection: Input It goes through a shortcut connection. If the number of input and output channels is different, the shortcut will go through a... The complex convolutional layer (ComplexConv2d) is used to match the dimensions to obtain the residual terms: ; Main Path: Input The following sub-modules are used sequentially: a. Complex Convolutional Layer (ComplexConv2d): The complex feature map is input into the complex convolutional layer for processing, represented as: ; where, and are standard real-valued 2D convolution kernels. Conceptually, they represent the real and imaginary parts of the complex-valued convolution kernel ; and represent the real and imaginary parts of the output of the th complex residual block, and represent the real and imaginary parts of the input to the complex convolutional layer. and ;

[0034] b. Complex Batch Normalization layer, which normalizes and to obtain normalized features and ; c. Complex ModReLU, which computes the modulus of the normalized features and a scaling factor . This activation function implements a nonlinear transformation through a scaling factor that acts on the modulus of a complex number. First, the modulus of the input feature map is computed : ; Then, the scaling factor is computed : ; where b is a learnable activation bias, is a very small positive number to ensure numerical stability; Finally, the computed scaling factor is applied to both the real and imaginary parts of the input features to obtain ; ; where and represent the real and imaginary parts of and after the complex convolutional, normalization, and activation functions.

[0035] d. Complex Attention module, which computes the complex channel weights and : ; ; in, This is the first The input feature maps of each complex residual module are dynamically generated, with complex weights for each channel. and It consists of real and imaginary parts; and Indicating the residual main path, and Real and imaginary features after complex convolutional layers, complex normalization layers, complex activation layers, and complex attention layers.

[0036] These features extend the attention mechanism beyond amplitude adjustment in the real domain to the complex domain, enabling adaptive modulation of both feature intensity and phase relationship. This is one of the key design features of the frequency domain branch in this invention, allowing it to effectively process and utilize phase information.

[0037] Output features Add each term to the residual term and perform pooling to obtain the first term. Complex feature map after processing by each module .

[0038] Specifically, the main path output features are added to the residual terms: ; .

[0039] Specifically, the complex feature map output by the last complex residual module is subjected to global average pooling to obtain the real and imaginary part vectors. and ; By processing the spatial dimensions, we obtain the flattened real and imaginary vectors. and ; The flattened real and imaginary vectors are concatenated along the channel dimension, and the concatenated feature vector is input into a fully connected layer to obtain the final frequency domain depth feature vector.

[0040] Specifically, pooling: The output of some residual modules is connected to a max pooling layer (MaxPool2d), which is applied independently to each module. and To reduce the spatial dimension of the feature map, the first... k Complex feature map after processing by each module The complex feature map output by the last complex residual module (ComplexResBlock) is obtained after global average pooling. and Flatten each part, remove the 1x1 spatial dimension, and obtain and The flattened real and imaginary vectors are concatenated along the channel dimension, and the concatenated feature vector is input into a fully connected layer to obtain the final frequency domain depth feature vector. .

[0041] Specifically, the one-dimensional convolutional neural network includes a one-dimensional convolutional layer, an anti-aliasing pooling layer (AADownsample), multiple one-dimensional residual modules (ResBlock1dTF), and a temporal aggregation module (TAggregate) based on a Transformer encoder, connected in sequence.

[0042] Specifically, the one-dimensional convolutional layer is used to perform preliminary feature extraction on the input waveform; The anti-overlapping pooling layer is used for downsampling; The one-dimensional residual module is used to extract deep temporal features; The temporal aggregation module is used to aggregate sequence features and generate temporal deep features.

[0043] Specifically, in the time-domain processing branch, Input to time domain network The time-domain features are denoted as , where the function This represents a deep one-dimensional convolutional neural network, which internally contains, in sequence, a one-dimensional convolutional layer for feature extraction, an anti-aliasing pooling layer for downsampling, a one-dimensional residual module for deepening the network, and a temporal aggregation module for aggregating global temporal information.

[0044] Specifically, the classifier is a fully connected layer (Linear), whose input dimension is the sum of the dimensions of the frequency domain depth features and the time domain depth features, and whose output dimension is the preset number of audio categories.

[0045] Specifically, the frequency domain depth features and time domain depth features are concatenated and fused along the channel dimension to obtain a time-frequency fusion feature vector. The fused features are input into a fully connected layer classifier to obtain audio scene recognition results. : ; and It is the weight matrix and bias vector that the classifier can learn.

[0046] In one specific embodiment of the present invention, an intelligent sound signal perception system based on time-frequency feature fusion includes: The signal acquisition module is used to acquire the raw audio signal; a feature extraction module configured to perform frequency domain feature extraction and time domain feature extraction on the original audio signal respectively to obtain frequency domain deep features and time domain deep features; a feature fusion module configured to perform channel dimension splicing fusion on the frequency domain deep features and the time domain deep features to obtain time-frequency fusion features; a scene recognition module configured to input the time-frequency fusion features into a full connection layer classifier to obtain a final audio scene recognition result.

[0047] In one specific embodiment of the present application, an intelligent sound signal perception method based on time-frequency feature fusion, as shown in Figure 2 includes parallel branch establishment: For an input original audio waveform signal (usually a one-dimensional tensor), first logically send it into two parallel processing branches: a time domain processing branch (TimeNet) and a frequency domain processing branch (FrequencyNet). The two branches will independently extract features from the signal.

[0048] In the frequency domain processing branch, the goal is to extract deep frequency domain features containing amplitude and phase information: first, perform short-time Fourier transform (STFT) on the original audio waveform. For example, the Fourier transform point number n_fft=512, the frame shift length hop_length=256, and the Hanning window hann_window can be set. The result of STFT is a complex tensor with dimensions [batch size, frequency points, time frames]. The complex tensor is separated into two real tensors, the real part and the imaginary part.

[0049] Next, to stabilize training, an STFTNormalizer module is used to globally normalize the real part and the imaginary part of the STFT, i.e., subtract the pre-calculated mean on the entire training set and divide by the standard deviation.

[0050] The normalized real and imaginary part tensors are stacked into a four-dimensional tensor [batch size, 2, frequency points, time frames] as the input of the subsequent complex neural network.

[0051] The complex neural network FrequencyNet is mainly composed of a series of stacked complex residual modules ComplexResBlock (as shown in Figure 3 Each ComplexResBlock contains: ComplexConv2d: Complex two-dimensional convolution layer. It simulates complex multiplication (a+bi) (c + di) = (ac - bd) + i(ad + bc), which processes the real and imaginary parts of the input separately through two parallel real convolution kernels, and then combines the real and imaginary parts of the output.

[0052] ComplexBatchNorm2d: Complex batch normalization layer. It concatenates the real and imaginary parts of the input along the channel dimension, and then feeds it into a standard BatchNorm2d layer for normalization. Finally, it separates the real and imaginary parts of the output.

[0053] ComplexModReLU: Complex activation function. It calculates a scaling factor based on the magnitude of the input complex number and a learnable bias, and then multiplies both the real and imaginary parts by the factor to achieve nonlinear activation.

[0054] ComplexSEBlock: Complex channel attention module. It learns the amplitude weight and phase weight of the channel dimension by performing Squeeze-and-Excitation operation on the magnitude of the input feature, thereby generating a complex weight to adaptively recalibrate different channels.

[0055] After multiple ComplexResBlock stacking processes, the feature map is processed through a global adaptive average pooling layer, then flattened and projected through a linear projection layer, finally outputting a fixed-dimensional frequency domain deep feature vector.

[0056] The time domain processing branch TimeNet (SoundNetRaw in this embodiment) directly processes one-dimensional raw audio waveforms.

[0057] The network first extracts preliminary features from the input waveform through a convolutional layer.

[0058] Then, the signal is alternately stacked through multiple down-sampling modules Down and one-dimensional residual modules ResBlock1dTF. The down-sampling module Down contains an AADownsample layer for anti-aliasing to preserve effective information while reducing the time resolution. ResBlock1dTF is used to deepen the network and extract more complex time domain features.

[0059] At the deep layer of the network, the feature sequence is fed into a TAggregate module. This module is essentially a standard Transformer encoder that aggregates information from the entire time sequence through self-attention mechanism. By introducing a learnable classification token (cls_token) and concatenating it with the sequence features, the final global time domain deep feature vector representing the entire audio segment is extracted from the output corresponding to the token.

[0060] Feature fusion and classification: After extracting the frequency domain deep features (feat_freq) and the time domain deep features (feat_time) in the two branches respectively, the two feature vectors are concatenated in the channel dimension (dim=1) to form a longer and more informative fused feature vector fused_feat.

[0061] Finally, the fused feature vector is fed into a final fully connected classifier fc. The classifier maps the fused features to the pre-defined number of audio categories and outputs the scores (logits) for each category, thus completing the audio scene recognition task.

[0062] In the training phase, the whole FusionNet model is optimized in an end-to-end manner. The difference between the predicted scores and the true labels is calculated using loss functions such as CrossEntropyLoss or LabelSmoothCrossEntropyLoss, and all learnable parameters in the network are updated using the AdamW optimizer through the backpropagation algorithm.

[0063] The various embodiments described in this specification are progressive in nature, and each embodiment focuses on the differences from other embodiments. The same or similar parts between embodiments can be referred to each other. For the apparatus disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the method part.

[0064] The above description of disclosed embodiments enables a person skilled in the art to implement or use the present application. Various modifications to the embodiments will be apparent to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to the embodiments shown herein, but will conform to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for intelligent sound signal perception based on time-frequency feature fusion, characterized in that, include: Acquire the raw audio signal; Frequency domain feature extraction and time domain feature extraction are performed on the original audio signal respectively to obtain frequency domain depth features and time domain depth features; The process of extracting frequency domain features from the original audio signal includes: Perform a short-time Fourier transform on the original audio signal to obtain complex spectral features; The complex spectral features are input into a complex neural network for processing to extract frequency domain depth features; Temporal feature extraction of the original audio signal includes: The original audio signal is processed using a one-dimensional convolutional neural network to extract temporal depth features; The frequency domain depth features and the time domain depth features are concatenated and fused along the channel dimension to obtain the time-frequency fusion features; The time-frequency fusion features are input into a fully connected layer classifier to obtain the final audio scene recognition result.

2. The intelligent sound signal perception method based on time-frequency feature fusion according to claim 1, characterized in that, The complex neural network is composed of multiple complex residual modules stacked together. Each complex residual module includes a complex convolutional layer, a complex batch normalization layer, a complex activation function, and a complex attention module connected in sequence.

3. The intelligent sound signal perception method based on time-frequency feature fusion according to claim 1, characterized in that, Also includes: After acquiring the original audio signal, the original audio signal is resampled to obtain the sampling features.

4. The intelligent sound signal perception method based on time-frequency feature fusion according to claim 2, characterized in that, Perform a short-time Fourier transform on the original audio signal to obtain complex spectral features, including: A short-time Fourier transform is performed on the original audio signal using a preset window function, number of Fourier transform points, and frame shift length to obtain a complex tensor containing real and imaginary parts; Using the pre-calculated global mean and standard deviation, the real and imaginary parts of the complex tensor are globally normalized to obtain the complex spectral features.

5. The intelligent sound signal perception method based on time-frequency feature fusion according to claim 2, characterized in that, The complex feature map is input into a complex convolutional layer for processing, as shown below: ; in, and Represents complex convolution kernel The real and imaginary parts; and Representing the first The characteristics of the real and imaginary parts output by each complex residual module. and They represent and Input the real and imaginary features of the output of the complex convolutional layer; The complex batch normalization layer and Normalization is performed to obtain normalized features. and ; The complex activation function calculates the modulus of the normalized feature. and scaling factor The calculated scaling factor is then applied to both the real and imaginary parts of the input features, resulting in: ; ; in, and express and Real and imaginary part features after complex convolutional layers, complex normalization layers, and complex activation functions; The complex attention module is based on and The modulus is used to obtain the complex channel weights: ; ; in, This is the first The input feature maps of each complex residual module are dynamically generated, with complex weights for each channel. and It consists of real and imaginary parts; and express and Real and imaginary features after complex convolutional layers, complex normalization layers, complex activation layers, and complex attention layers; Output features Add each term to the residual term and perform pooling to obtain the first term. Complex feature map after processing by each module .

6. The intelligent sound signal perception method based on time-frequency feature fusion according to claim 5, characterized in that, The complex feature map output by the last complex residual module is then subjected to global average pooling to obtain the real and imaginary vectors. and ; By processing the spatial dimensions, we obtain the flattened real and imaginary vectors. and ; The flattened real and imaginary vectors are concatenated along the channel dimension, and the concatenated feature vector is input into a fully connected layer to obtain the final frequency domain depth feature vector.

7. The intelligent sound signal perception method based on time-frequency feature fusion according to claim 1, characterized in that, The one-dimensional convolutional neural network includes a one-dimensional convolutional layer, an anti-aliasing pooling layer, a one-dimensional residual module, and a temporal aggregation module connected in sequence.

8. The intelligent sound signal perception method based on time-frequency feature fusion according to claim 7, characterized in that, The one-dimensional convolutional layer is used to perform preliminary feature extraction on the input waveform; The anti-overlapping pooling layer is used for downsampling; The one-dimensional residual module is used to extract deep temporal features; The temporal aggregation module is used to aggregate sequence features and generate temporal deep features.

9. An intelligent sound signal perception system based on time-frequency feature fusion, employing the intelligent sound signal perception method based on time-frequency feature fusion as described in any one of claims 1-8, characterized in that, include: The signal acquisition module is used to acquire the raw audio signal; The feature extraction module is used to extract frequency domain features and time domain features from the original audio signal to obtain frequency domain depth features and time domain depth features. The feature fusion module is used to concatenate and fuse the frequency domain depth features and the time domain depth features in the channel dimension to obtain time-frequency fusion features; The scene recognition module is used to input the time-frequency fusion features into the fully connected layer classifier to obtain the final audio scene recognition result.

Citation Information

Patent Citations

  • One-dimensional convolution acceleration device and method for complex neural network

    CN111626412A

  • Audio scene classification method and device, terminal equipment and storage medium

    CN114186094A

  • Plural convolutional neural network speech enhancement method and system based on attention

    CN115938377A

  • Speech enhancement method, system and equipment based on attention of complex coordinates, and medium

    CN118762705A

  • Voice emotion recognition model and method based on cross-space-time fusion attention network

    CN120496583A

Cited By

  • Intelligent water leakage detection method and system for water supply network

    CN121188581A