Sound classification method, device, electronic device, storage medium and computer program product
By introducing the pulse residual module and the attention feature fusion module into the sound classification method, the problem of high power consumption and low efficiency of the sound classification method in the prior art is solved, and efficient and accurate sound classification is achieved.
Patent Information
- Application Number
- CN202411523922.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-29
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2044-10-29
AI Technical Summary
The existing sound classification methods have problems of high power consumption and low efficiency in processing timing information and complex features.
The pulse residual module and attention feature fusion module are used to trigger the LIF neuron processing to extract pulse features through leakage integration in the pulse form, and the characteristics are fused through residual processing and attention mechanism to achieve efficient classification of sound signals.
It significantly reduces system power consumption, improves the computing efficiency and accuracy of sound classification, and can process complex timing data more effectively.
Smart Images

Figure CN119360893B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer technology, and more particularly, to a sound classification method, device, electronic device, storage medium, and computer program product. Background Art
[0002] Sound classification has a wide range of applications in speech recognition, environmental sound recognition, and acoustic event detection. In related technologies, the main sound classification methods mainly rely on Convolutional Neural Networks (CNN), Recurrent Neural Networks (RNN), and Residual Neural Networks (ResNet).
[0003] Convolutional neural networks (CNN) are mainly used to extract features from images, but their powerful feature extraction capabilities are also used in sound classification. Recurrent neural networks (RNN) can effectively process the temporal features in sound signals, and are particularly suitable for tasks such as speech recognition that require capturing long-term dependencies. Residual neural networks (ResNet) can better extract and fuse multi-level features in sound classification tasks, thereby ensuring the accuracy of sound classification.
[0004] However, the above-mentioned sound classification methods in the related arts still have certain limitations in processing time series information and complex feature fusion, and generally have the defects of high power consumption and low efficiency. Summary of the invention
[0005] The present disclosure provides a sound classification method, device, electronic device, storage medium and computer program product to at least solve the problem that the sound classification method in the above-mentioned related technologies generally has the defects of high power consumption and low efficiency.
[0006] According to a first aspect of an embodiment of the present disclosure, a sound classification method is provided, comprising: extracting audio features of a sound signal to be classified; inputting the audio features into a pulse residual module to obtain a first pulse residual feature; inputting the first pulse residual feature into at least one of the pulse residual modules to obtain a second pulse residual feature; inputting the second pulse residual feature and the first pulse residual feature after downsampling into an attention feature fusion module to obtain a first attention fusion feature; classifying the sound signal to be classified based on the first attention fusion feature; wherein the pulse residual module is configured to: perform pulse-form leaky integral triggered LIF neuron processing based on the input features input to the pulse residual module to obtain a pulse feature; perform residual processing based on the pulse feature and the input feature to obtain a pulse residual feature as an output.
[0007] Optionally, the leaky integral triggering LIF neuron processing in pulse form based on the input features input to the pulse residual module to obtain the pulse features includes: performing LIF neuron processing and segmentation processing on the input features to obtain multiple segmented parts; for a first segmented part among the multiple segmented parts, performing LIF neuron processing on the first segmented part to obtain a segmented part pulse feature corresponding to the first segmented part; for each segmented part among the other segmented parts among the multiple segmented parts, adding the segmented part pulse features corresponding to each segmented part and the previous segmented part, and performing LIF neuron processing on the added features to obtain the segmented part pulse features corresponding to each segmented part; splicing the segmented part pulse features corresponding to each segmented part among the multiple segmented parts to obtain the pulse features.
[0008] Optionally, before classifying the sound signal to be classified based on the first attention fusion feature, it also includes: inputting the second pulse residual feature into at least one pulse residual fusion module to obtain a residual fusion feature; classifying the sound signal to be classified based on the first attention fusion feature, including: downsampling the first attention fusion feature to obtain a downsampling result; inputting the residual fusion feature and the downsampling result into an attention feature fusion module to obtain a second attention fusion feature; classifying the sound signal to be classified based on the second attention fusion feature; wherein the pulse residual fusion module is configured to: input the input feature to the pulse residual fusion module Perform LIF neuron processing and segmentation processing to obtain multiple segmented parts; for a first segmented part among the multiple segmented parts, perform LIF neuron processing on the first segmented part, and output the segmented part pulse feature corresponding to the first segmented part; for each segmented part among the other segmented parts among the multiple segmented parts, input the output features corresponding to each segmented part and the previous segmented part into the attention feature fusion module, and output the attention fusion feature corresponding to each segmented part; splice the segmented part pulse feature corresponding to the first segmented part and the attention fusion feature corresponding to each segmented part among the other segmented parts to obtain a spliced feature; based on the spliced feature, obtain a residual fusion feature.
[0009] Optionally, the second pulse residual feature and the first pulse residual feature after downsampling are input into the attention feature fusion module to obtain the first attention fusion feature, including: performing splicing processing on the second pulse residual feature and the first pulse residual feature after downsampling to obtain a splicing result; using the local attention mechanism contained in the attention feature fusion module to calculate the first attention weight corresponding to the second pulse residual feature and the second attention weight corresponding to the first pulse residual feature after downsampling based on the splicing result; performing nonlinear transformation processing on the first attention weight and the second attention weight respectively to obtain a first nonlinear attention weight corresponding to the first attention weight and a second nonlinear attention weight corresponding to the second attention weight; performing weighted summation based on the second pulse residual feature, the first pulse residual feature after downsampling, the first nonlinear attention weight and the second nonlinear attention weight to obtain the first attention fusion feature.
[0010] Optionally, at least one of the pulse residual modules is cascade connected.
[0011] Optionally, the at least one pulse residual fusion module is cascade connected.
[0012] According to a second aspect of an embodiment of the present disclosure, a sound classification device is provided, comprising: an audio feature extraction module, configured to extract audio features of a sound signal to be classified; a first pulse residual feature acquisition module, configured to input the audio feature into a pulse residual module to obtain a first pulse residual feature; a second pulse residual feature acquisition module, configured to input the first pulse residual feature into at least one of the pulse residual modules to obtain a second pulse residual feature; an attention fusion feature acquisition module, configured to input the second pulse residual feature and the first pulse residual feature after downsampling into an attention feature fusion module to obtain a first attention fusion feature; a classification module, configured to classify the sound signal to be classified based on the first attention fusion feature; wherein the pulse residual module is configured to: perform pulse-form leaky integral triggered LIF neuron processing based on the input feature input to the pulse residual module to obtain a pulse feature; and perform residual processing based on the pulse feature and the input feature to obtain a pulse residual feature as an output.
[0013] Optionally, the pulse residual module is configured to: perform LIF neuron processing and segmentation processing on the input features to obtain multiple segmented parts; for a first segmented part among the multiple segmented parts, perform LIF neuron processing on the first segmented part to obtain a segmented part pulse feature corresponding to the first segmented part; for each segmented part among the other segmented parts among the multiple segmented parts, add the segmented part pulse features corresponding to each segmented part and the previous segmented part, and perform LIF neuron processing on the added features to obtain a segmented part pulse feature corresponding to each segmented part; splice the segmented part pulse features corresponding to each segmented part among the multiple segmented parts to obtain the pulse feature.
[0014] Optionally, the sound classification device further includes: a residual fusion feature acquisition module, configured to input the second pulse residual feature into at least one pulse residual fusion module to obtain a residual fusion feature; the classification module is configured to: downsample the first attention fusion feature to obtain a downsampling result; input the residual fusion feature and the downsampling result into an attention feature fusion module to obtain a second attention fusion feature; based on the second attention fusion feature, classify the sound signal to be classified; wherein the pulse residual fusion module is configured to: perform LIF neuron processing and segmentation processing on the input features input to the pulse residual fusion module, A plurality of segmented parts are obtained; for a first segmented part among the plurality of segmented parts, LIF neuron processing is performed on the first segmented part, and a segmented part pulse feature corresponding to the first segmented part is output; for each segmented part among the other segmented parts among the plurality of segmented parts, the output features corresponding to each segmented part and the previous segmented part are input into an attention feature fusion module, and an attention fusion feature corresponding to each segmented part is output; the segmented part pulse feature corresponding to the first segmented part and the attention fusion feature corresponding to each segmented part among the other segmented parts are spliced to obtain a spliced feature; based on the spliced feature, a residual fusion feature is obtained.
[0015] Optionally, the attention fusion feature acquisition module is configured to: perform splicing processing on the second pulse residual feature and the first pulse residual feature after downsampling to obtain a splicing result; use the local attention mechanism contained in the attention feature fusion module to calculate the first attention weight corresponding to the second pulse residual feature and the second attention weight corresponding to the first pulse residual feature after downsampling based on the splicing result; perform nonlinear transformation processing on the first attention weight and the second attention weight respectively to obtain the first nonlinear attention weight corresponding to the first attention weight and the second nonlinear attention weight corresponding to the second attention weight; perform weighted summation based on the second pulse residual feature, the first pulse residual feature after downsampling, the first nonlinear attention weight and the second nonlinear attention weight to obtain the first attention fusion feature.
[0016] Optionally, at least one of the pulse residual modules is cascade connected.
[0017] Optionally, the at least one pulse residual fusion module is cascade connected.
[0018] According to a third aspect of an embodiment of the present disclosure, an electronic device is provided, comprising: a processor; and a memory for storing instructions executable by the processor; wherein the processor is configured to execute the instructions to implement the sound classification method according to the present disclosure.
[0019] According to a fourth aspect of an embodiment of the present disclosure, a computer-readable storage medium is provided. When instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is enabled to perform the sound classification method according to the present disclosure.
[0020] According to a fifth aspect of an embodiment of the present disclosure, a computer program product is provided, including a computer program, and when the computer program is executed by a processor, the sound classification method according to the present disclosure is implemented.
[0021] The technical solution provided by the embodiments of the present disclosure brings at least the following beneficial effects:
[0022] In the present disclosure, a spiking neural network (SNN) is a neural network inspired by biological neurons, which calculates and transmits information through discrete pulses. Since the spiking neural network (SNN) only calculates when a pulse occurs, it has higher computational efficiency and lower energy consumption when classifying sounds compared to traditional artificial neural networks (ANNs). In addition, by using a residual neural network for feature extraction, the accuracy of sound recognition can be guaranteed. It can be seen that the present disclosure can make full use of the advantages of spiking neural networks (SNNs) and residual neural networks, can achieve efficient and accurate sound classification, and can significantly reduce system power consumption.
[0023] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] The drawings herein are incorporated into and constitute a part of the specification, illustrate embodiments consistent with the present disclosure, and together with the description are used to explain the principles of the present disclosure, and do not constitute improper limitations on the present disclosure.
[0025] Figure 1 is a flowchart illustrating a sound classification method according to an exemplary embodiment of the present disclosure;
[0026] Figure 2 is a schematic diagram showing the structure of a pulse residual module (SpikR) according to an exemplary embodiment of the present disclosure;
[0027] Figure 3 is a schematic diagram showing the structure of an attention feature fusion module (AFF) according to an exemplary embodiment of the present disclosure;
[0028] Figure 4 is a schematic diagram showing the structure of a pulse residual fusion module (SpikRF) according to an exemplary embodiment of the present disclosure;
[0029] Figure 5 is a schematic diagram showing a connection relationship between a pulse residual module, a pulse residual fusion module, and an attention feature fusion module according to an exemplary embodiment of the present disclosure;
[0030] Figure 6 is a schematic diagram showing a cascaded pulse residual module and a cascaded pulse residual fusion module according to an exemplary embodiment of the present disclosure;
[0031] Figure 7 is a block diagram illustrating a sound classification apparatus according to an exemplary embodiment of the present disclosure;
[0032] Figure 8 is a block diagram illustrating an electronic device according to an exemplary embodiment of the present disclosure. DETAILED DESCRIPTION
[0033] In order to enable ordinary persons in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the accompanying drawings.
[0034] It should be noted that the terms "first", "second", etc. in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. The implementation methods described in the following examples do not represent all implementation methods consistent with the present disclosure. Instead, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the attached claims.
[0035] It should be noted that the phrase "at least one of the items" in the present disclosure includes three types of parallel situations: "any one of the items", "a combination of any number of the items", and "all of the items". For example, "including at least one of A and B" includes the following three parallel situations: (1) including A; (2) including B; (3) including A and B. Another example is "executing at least one of step 1 and step 2" which means the following three parallel situations: (1) executing step 1; (2) executing step 2; (3) executing step 1 and step 2.
[0036] Sound classification has a wide range of applications in speech recognition, environmental sound recognition, and acoustic event detection. Traditional sound classification methods mainly rely on Convolutional Neural Networks (CNN), Recurrent Neural Networks (RNN), and Residual Neural Networks (ResNet), but these methods have certain limitations in processing time series information and complex feature fusion.
[0037] Convolutional Neural Network (CNN): Convolutional neural network is mainly used to extract features from images, and its powerful feature extraction ability is also applied to sound classification. Generally, the sound signal in the time domain is first converted into a feature map in the frequency domain, such as a Mel-spectrogram, through methods such as short-time Fourier transform (STFT). The feature map is then input into the convolutional neural network, and high-level features are extracted through multiple convolutional layers and pooling layers. Finally, the feature map output by the convolutional layer is expanded into a vector and classified through a fully connected layer.
[0038] Recurrent Neural Network (RNN): Due to its advantages in processing time series data, it is widely used in sound classification tasks. RNN can memorize and process dependencies in time series data through its internal loop structure. In order to overcome the gradient vanishing problem of traditional RNN, LSTM and GRU were introduced, which can better capture long-term dependencies. The sound signal is divided into several time segments, each time segment is processed by RNN, and finally classified by combining the features of all time segments. RNN can effectively process the time series features in sound signals, and is especially suitable for tasks such as speech recognition that need to capture long-term dependencies.
[0039] Residual Neural Network (ResNet): Introduced shortcut connections to solve the gradient vanishing problem of deep networks. By introducing shortcut connections, the network can directly learn the residual between input and output, thereby deepening the number of layers of the network without causing gradient vanishing. ResNet can extract features from different scales and enhance feature extraction capabilities by stacking multiple layers of residual modules. Classification is performed through fully connected layers and softmax functions. In the task of sound classification, ResNet can better extract and fuse multi-level features and improve classification accuracy. However, the above-mentioned sound classification methods in related technologies still have certain limitations in processing time series information and complex feature fusion, and generally have the defects of high power consumption and low efficiency.
[0040] In order to solve the above-mentioned problems existing in the related art, the present disclosure provides a sound classification method, device, electronic device, storage medium and computer program product. The pulse neural network (SNN) is a neural network inspired by biological neurons, which calculates and transmits information through discrete pulses. Since the pulse neural network (SNN) only calculates when a pulse occurs, it has higher computing efficiency and lower energy consumption when performing sound classification compared to the traditional artificial neural network (ANN). In addition, by using the residual neural network for feature extraction, the accuracy of sound recognition can be guaranteed. It can be seen that the present disclosure can make full use of the advantages of the pulse neural network (SNN) and the residual neural network, can achieve efficient and accurate sound classification, and can significantly reduce system power consumption.
[0041] Figure 1 is a flowchart illustrating a sound classification method according to an exemplary embodiment of the present disclosure.
[0042] Reference Figure 1In step 101, the audio features of the sound signal to be classified can be extracted. It should be noted that the sound signal to be classified can be preprocessed first, for example, the sound signal to be classified can be standardized to remove background noise. Then, feature extraction can be performed on the preprocessed sound signal, including but not limited to: Mel Frequency Cepstral Coefficient (MFCC) and Short Time Fourier Transform (STFT) and the like.
[0043] Specifically, the sound signal to be classified can be subjected to noise reduction, normalization and other processing, the sound signal to be classified can be divided into frames with a length of 10 milliseconds, and a window function can be applied to each frame to reduce spectrum leakage, and each frame can be converted from the time domain to the frequency domain. Then, the energy of a specific frequency band can be extracted by a Mel frequency filter bank. Next, the energy output by the filter can be logarithmically transformed. Then, the signal after logarithmic transformation can be normalized.
[0044] In step 102, the audio feature can be input into a spiking residual module (SpikR) to obtain a first spiking residual feature, wherein the spiking residual module can be configured to perform leaky integrate-and-fire module (LIF) in a pulse form based on the input feature input into the spiking residual module to obtain a spiking feature. Then, residual processing can be performed based on the spiking feature and the input feature to obtain a spiking residual feature as an output. That is, in the present disclosure, the pre-processed sound data can be input into a multi-scale spiking residual module to extract spiking residual features.
[0045] It should be noted that the spiking neural network (SNN), as a neural network inspired by biological neurons, calculates and transmits information through discrete pulses, making it excellent in processing high-speed, dynamic and high-dimensional data. It has the following specific advantages:
[0046] Advantages of time dynamic processing: SNN can capture the time dynamic characteristics of sound signals and is suitable for processing complex time series data; Advantages of energy saving and high efficiency: Since SNN only performs calculations when pulses occur, it has higher computing efficiency and lower energy consumption than traditional artificial neural networks (ANN); Advantages of strong robustness: SNN has good robustness against noise and signal interference, and can effectively classify sounds in noisy environments; Advantages of event-driven: SNN uses event-driven mechanism to quickly respond to emergencies in sound signals and improve the real-time performance of sound classification.
[0047] Based on the above advantages, pulse neural networks have shown great potential in sound classification tasks. In addition, by combining multi-scale fusion technology, the performance of sound feature extraction and classification can be further improved, thereby achieving more accurate and efficient sound classification.
[0048] Figure 2 is a schematic diagram showing the structure of a pulse residual module (SpikR) according to an exemplary embodiment of the present disclosure. Figure 2 The pulse residual module may also be called a multi-scale pulse residual module, and its input may be a data tensor, and its size may be: (batch_size, channels, height, width). The pulse residual module may include multiple residual modules, each of which may be composed of several layers, and the audio features may be input into the multiple residual modules of the multi-scale pulse residual module for feature extraction.
[0049] According to an exemplary embodiment of the present disclosure, referring to Figure 2 , the input features can be processed by LIF neurons and segmented to obtain multiple segmented parts, that is, n segmented parts can be obtained. Specifically, the number of input channels of the input features can be adjusted by the first layer of 1x1 convolution (Conv1×1), and batch normalization operation (Batch Norm) can be performed. Then, the input features after the batch normalization operation can be sequentially processed by pulse-form leaky integral triggering neurons and segmentation processing, and multiple segmented parts can be obtained. Exemplarily, the feature map obtained by LIF processing can be divided into 3 segmented parts according to width, namely: x1, x2 and x3.
[0050] Next, for each of the n segmented parts, 3x3 convolution (Conv3×3) and batch normalization processing can be performed. Then, LIF processing can be performed on each segmented part after batch normalization processing.
[0051] Specifically, for the first segmented part x1 among the above-mentioned multiple segmented parts, 3x3 convolution, batch normalization processing and LIF processing can be performed on the first segmented part x1 in sequence, and then the segmented part pulse feature corresponding to the first segmented part x1 can be obtained.
[0052] For each of the other segmented parts in the above-mentioned multiple segmented parts, the segmented part pulse features corresponding to each segmented part and the previous segmented part can be added, and the added features can be sequentially performed with 3x3 convolution, batch normalization processing and LIF processing, so as to obtain the segmented part pulse features corresponding to each segmented part.
[0053] Exemplarily, for the second segmented part x2, the second segmented part x2 and the segmented part pulse characteristics corresponding to the previous segmented part of the second segmented part x2, i.e., the first segmented part x1, can be added, and the added characteristics can be sequentially subjected to 3x3 convolution, batch normalization processing and LIF processing, so as to obtain the segmented part pulse characteristics corresponding to the second segmented part x2; and, for the third segmented part x3, the third segmented part x3 and the segmented part pulse characteristics corresponding to the previous segmented part of the third segmented part x3, i.e., the second segmented part x2, can be added, and the added characteristics can be sequentially subjected to 3x3 convolution, batch normalization processing and LIF processing, so as to obtain the segmented part pulse characteristics corresponding to the third segmented part x3.
[0054] Then, the pulse features of the segmented parts corresponding to each segmented part in the above-mentioned multiple segmented parts can be concatenated (Concatenation) to obtain the pulse feature. Next, 1x1 convolution and batch normalization can be performed on the pulse feature to restore the number of channels. Next, the input features of the input pulse residual module can be added to the output feature map of the restored channel number through a shortcut connection to form a residual connection. Then, LIF processing can be performed on the residual connection result to obtain the pulse residual feature as the output.
[0055] In step 103, the first pulse residual feature may be input into at least one pulse residual module to obtain a second pulse residual feature.
[0056] According to an exemplary embodiment of the present disclosure, the at least one pulse residual module mentioned above can be cascaded, that is, when there are multiple pulse residual modules, these multiple pulse residual modules can be first-connected in sequence to form a cascade.
[0057] In step 104, the second pulse residual feature and the downsampled first pulse residual feature may be input into an attention feature fusion module (AFF) to obtain a first attention fusion feature.
[0058] According to an exemplary embodiment of the present disclosure, Figure 3 is a schematic diagram showing the structure of an attention feature fusion module (AFF) according to an exemplary embodiment of the present disclosure. Figure 3 , a local attention mechanism can be constructed through a 1x1 convolution layer, a batch normalization layer, and a SiLU activation function.
[0059] First, the second pulse residual feature and the downsampled first pulse residual feature can be spliced to obtain a splicing result. That is, the second pulse residual feature and the downsampled first pulse residual feature can be spliced in the channel dimension to form a new feature map.
[0060] Then, the above-mentioned local attention mechanism included in the attention feature fusion module can be used to calculate the first attention weight corresponding to the second pulse residual feature and the second attention weight corresponding to the downsampled first pulse residual feature based on the splicing result.
[0061] Next, 1x1 convolution, batch normalization and nonlinear transformation (tanh) can be performed in sequence on the first attention weight and the second attention weight to obtain a first nonlinear attention weight corresponding to the first attention weight and a second nonlinear attention weight corresponding to the second attention weight.
[0062] Then, a first attention fusion feature can be obtained by performing a weighted sum (Broadcasting Addition) based on the second pulse residual feature, the downsampled first pulse residual feature, the first nonlinear attention weight and the second nonlinear attention weight.
[0063] In step 105, the sound signal to be classified can be classified based on the first attention fusion feature. Specifically, the first attention fusion feature can be pooled using temporal statistical pooling (TSTP) to extract the pooled feature statistics. Next, the pooled feature statistics can be input into the embedding layer for feature mapping, that is, the pooled feature statistics can be mapped to the embedding space through the linear layer. Then, the mapped features can be input into the classification layer for classification, that is, the embedded features can be input into the classification layer for classification.
[0064] It should be noted that, in the present disclosure, the result of classifying the sound signal may be, but is not limited to: various environmental sounds, animal calls, human voices, and the like.
[0065] In addition, the present disclosure does not limit the hardware on which the sound classification method is specifically run. Exemplarily, it is recommended to use the Python programming language to implement the sound classification method of the present disclosure. The present disclosure can use a GPU with a 3.2 GHz central processor and 16 GB of memory and at least one 12 GB video memory, and can use the Python 3.11 programming language and the Pytorch 2.0.1 version deep learning framework to implement the sound classification method of the present disclosure.
[0066] It should be noted that the temporal dynamic processing advantages and energy-saving and efficient characteristics of the pulse neural network enable it to perform well in processing sound signals, and is particularly suitable for sound classification tasks in high-speed, dynamic and complex environments. Therefore, the sound classification method provided by the present disclosure can more efficiently extract and fuse the time-frequency features in the sound data by combining the pulse neural network and the multi-scale fusion technology, and significantly improve the accuracy of classification. In addition, the parameter amount can be effectively reduced to reduce energy consumption.
[0067] According to an exemplary embodiment of the present disclosure, the second pulse residual feature can also be input into at least one pulse residual fusion module (SpikRF) to obtain a residual fusion feature. Moreover, when there are multiple pulse residual fusion modules, these multiple pulse residual fusion modules can be cascaded by first connection. Moreover, the first attention fusion feature can also be downsampled to obtain a downsampled result.
[0068] Next, the above residual fusion features and downsampling results can be input into the attention feature fusion module to obtain the second attention fusion feature. Then, the sound signal to be classified can be classified based on the second attention fusion feature. In addition, the implementation process of "classifying the sound signal to be classified based on the second attention fusion feature" refers to the previous description of "classifying the sound signal to be classified based on the first attention fusion feature", which will not be repeated here.
[0069] According to an exemplary embodiment of the present disclosure, when the number of at least one pulse residual fusion module (SpikRF) is plural, these multiple pulse residual fusion modules can be first connected in sequence to form a cascade connection.
[0070] According to an exemplary embodiment of the present disclosure, Figure 4 is a schematic diagram showing the structure of a pulse residual fusion module (SpikRF) according to an exemplary embodiment of the present disclosure. Figure 4 , the pulse residual fusion module can also be called a multi-scale pulse residual fusion module. The input of the multi-scale pulse residual fusion module can be the feature of the multi-scale pulse residual module output. In addition, the multi-scale pulse residual fusion module can include a convolution layer, a batch normalization layer, a LIF neuron, and the like.
[0071] The above pulse residual fusion module can be configured as:
[0072] The input features input to the pulse residual fusion module can be processed by LIF neurons and segmented to obtain multiple segmented parts. Specifically, for the input features, the number of channels of the input features can be adjusted by 1x1 convolution, and batch normalization can be performed on the input features with adjusted channel numbers. Then, the leaky integral triggered neuron processing and segmentation processing in the form of pulses can be performed in sequence on the results of batch normalization processing, and m segmented parts can be obtained. Exemplarily, the feature map obtained by LIF processing can be segmented according to width, and three segmented parts can be obtained, namely: y1, y2 and y3.
[0073] Next, multi-scale feature fusion can be performed through the attention feature fusion module (AFF). Specifically, for the first segmentation y1 among the multiple segmentations, 3x3 convolution, batch normalization, and pulse-form leaky integral triggering neuron processing can be performed on the first segmentation y in sequence, and then the segmentation pulse feature corresponding to the first segmentation y1 can be output.
[0074] Then, for each segmented part in the other segmented parts among the above-mentioned multiple segmented parts, the output features corresponding to each segmented part and the previous segmented part of the segmented part can be input into the attention feature fusion module, and then the attention fusion features corresponding to each segmented part can be output.
[0075] Specifically, for the second segmented part y2 among the above m segmented parts, the segmented part pulse features corresponding to the second segmented part y2 and the first segmented part y1 can be input into the attention feature fusion module to obtain the attention fusion features corresponding to the second segmented part y2.
[0076] For each of the third to mth segments among the m segments, the attention fusion features corresponding to each segment and the previous segment of the segment can be input into the attention feature fusion module, and then the attention fusion features corresponding to each segment can be obtained. Exemplarily, for the third segment y3, the attention fusion features corresponding to the third segment y3 and the previous segment of the third segment y3, that is, the second segment y2, can be input into the attention feature fusion module, and then the attention fusion features corresponding to the third segment y3 can be obtained.
[0077] Next, the segmentation pulse feature corresponding to the first segmentation and the attention fusion feature corresponding to each segmentation in the other segmentations may be concatenated to obtain a concatenated feature. Specifically, the segmentation pulse feature corresponding to the first segmentation and the attention fusion features corresponding to the second to mth segmentations may be concatenated to obtain a concatenated feature.
[0078] Then, the residual fusion feature can be obtained based on the above-mentioned splicing feature. Specifically, the splicing feature can be subjected to 1x1 convolution and batch normalization processing to restore the number of channels. Next, the input feature of the input pulse residual fusion module can be added to the output feature map of the restored channel number through a shortcut connection to form a residual connection. Then, the residual connection result can be subjected to LIF processing to obtain the above-mentioned residual fusion feature.
[0079] Figure 5 is a schematic diagram showing the connection relationship between the pulse residual module, the pulse residual fusion module, and the attention feature fusion module according to an exemplary embodiment of the present disclosure. Figure 5 , a total of 3 pulse residual modules, 4 pulse residual fusion modules and 3 attention feature fusion modules are shown. Moreover, the 3 pulse residual modules include a single pulse residual module (SpikR×1) and a cascaded pulse residual module combination (SpikR×2) consisting of 2 pulse residual modules; the 4 pulse residual fusion modules include a single pulse residual fusion module (SpikRF×1) and a cascaded pulse residual fusion module combination (SpikRF×3) consisting of 3 pulse residual fusion modules. In addition, Figure 5 The three attention feature fusion modules shown in are fuse_mode12, fuse_mode123 and fuse_mode1234 respectively.
[0080] Figure 6 is a schematic diagram showing a cascaded pulse residual module and a cascaded pulse residual fusion module according to an exemplary embodiment of the present disclosure. Figure 6 The cascaded pulse residual module combination (SpikR×2) includes a total of two pulse residual modules, and these two pulse residual modules are connected head to tail in sequence, that is, the output of the first pulse residual module will be used as the input of the second pulse residual module; the cascaded pulse residual fusion module combination (SpikRF×3) includes a total of three pulse residual fusion modules, and these three pulse residual fusion modules are connected head to tail in sequence, that is, the output of the first pulse residual fusion module will be used as the input of the second pulse residual fusion module, and the output of the second pulse residual fusion module will be used as the input of the third pulse residual fusion module.
[0081] Return to reference Figure 5 After preprocessing the audio signal to be classified to obtain audio features, the audio features can be input into the first layer consisting of a pulse residual module (SpikR×1) to obtain output out1. Then, out1 can be input into the second layer (SpikR×2) consisting of two connected pulse residual modules to obtain output out2. Next, out1 can be downsampled to obtain out1_downsample. Then, out2 and out1_downsample can be input into the attention feature fusion module (fuse_mode12) to obtain the fused output fuse_out12.
[0082] Next, out2 can be input to the third layer (SpikRF×3) composed of three connected pulse residual fusion modules to obtain output out3. Then, the above fuse_out12 can be downsampled to obtain fuse_out12_downsample. Next, out3 and fuse_out12_downsample can be input to the attention feature fusion module (fuse_mode123) to obtain the fused output fuse_out123.
[0083] Then, out3 can be input to the 4th layer (SpikRF×1) composed of 1 pulse residual fusion module for processing to obtain output out4. Next, the above fuse_out123 can be downsampled to obtain fuse_out123_downsample. Then, out4 and fuse_out123_downsample can be input to the attention feature fusion module (fuse_mode1234) to obtain the final fusion output fuse_out1234.
[0084] Finally, the final fusion output fuse_out1234 can be pooled to extract feature statistics, and then it can enter the fully connected layer (FC) for sound classification.
[0085] The sound classification method, device, electronic device, storage medium and computer program product provided by the present disclosure combine the advantages of pulse neural networks in processing sound timing information and the advantages of residual neural networks in feature extraction, so that features can be better extracted from sound signals. Specifically, by using pulse neural networks to construct data streams, their advantages in low power consumption and efficient processing of timing data can be fully utilized; in addition, by using residual neural networks for feature extraction, the accuracy of sound recognition can be further improved. It can be seen that the present disclosure can fully combine and utilize the advantages of these two neural networks, thereby achieving efficient and accurate sound classification, and can also significantly reduce system power consumption.
[0086] Figure 7 is a block diagram illustrating a sound classification apparatus according to an exemplary embodiment of the present disclosure.
[0087] Reference Figure 7 The sound classification device 700 may include an audio feature extraction module 701, a first pulse residual feature acquisition module 702, a second pulse residual feature acquisition module 703, an attention fusion feature acquisition module 704 and a classification module 705.
[0088] The audio feature extraction module 701 can extract the audio features of the sound signal to be classified. It should be noted that the sound signal to be classified can be preprocessed first, for example, the sound signal to be classified can be standardized to remove background noise. Then, feature extraction can be performed on the preprocessed sound signal, including but not limited to: Mel frequency cepstral coefficient (MFCC) and short-time Fourier transform (STFT) and the like.
[0089] The first pulse residual feature acquisition module 702 can input the audio feature into the pulse residual module (SpikR) to obtain the first pulse residual feature, wherein the pulse residual module can be configured to: perform leaky integral triggered neuron processing (LIF) in the form of pulses based on the input features input to the pulse residual module to obtain the pulse feature. Then, residual processing can be performed based on the pulse feature and the input feature to obtain the pulse residual feature as an output. That is, in the present disclosure, the sound data after preprocessing can be input into the multi-scale pulse residual module for pulse residual feature extraction.
[0090] According to an exemplary embodiment of the present disclosure, the pulse residual module may be configured as follows:
[0091] The input features are processed by LIF neurons and segmented to obtain multiple segmented parts. Then, for a first segmented part among the multiple segmented parts, LIF neuron processing can be performed on the first segmented part to obtain a segmented part pulse feature corresponding to the first segmented part. Next, for each segmented part among the other segmented parts among the multiple segmented parts, the segmented part pulse features corresponding to each segmented part and the previous segmented part can be added, and the added features can be processed by LIF neurons to obtain the segmented part pulse features corresponding to each segmented part. Then, the segmented part pulse features corresponding to each segmented part among the multiple segmented parts can be spliced to obtain the pulse features.
[0092] The second pulse residual feature acquisition module 703 may input the first pulse residual feature into at least one pulse residual module to obtain a second pulse residual feature.
[0093] According to an exemplary embodiment of the present disclosure, the at least one pulse residual module mentioned above can be cascaded, that is, when there are multiple pulse residual modules, these multiple pulse residual modules can be first-connected in sequence to form a cascade.
[0094] The attention fusion feature acquisition module 704 can input the second pulse residual feature and the downsampled first pulse residual feature into the attention feature fusion module (AFF) to obtain the first attention fusion feature.
[0095] According to an exemplary embodiment of the present disclosure, first, the attention fusion feature acquisition module 704 can perform splicing processing on the second pulse residual feature and the first pulse residual feature after downsampling to obtain a splicing result. That is, the second pulse residual feature and the first pulse residual feature after downsampling can be first spliced in the channel dimension to form a new feature map.
[0096] Then, the attention fusion feature acquisition module 704 can use the above-mentioned local attention mechanism contained in the attention feature fusion module to calculate the first attention weight corresponding to the second pulse residual feature and the second attention weight corresponding to the first pulse residual feature after downsampling based on the splicing result.
[0097] Next, the attention fusion feature acquisition module 704 can perform 1x1 convolution, batch normalization processing and nonlinear transformation processing (tanh) on the first attention weight and the second attention weight in sequence to obtain a first nonlinear attention weight corresponding to the first attention weight and a second nonlinear attention weight corresponding to the second attention weight.
[0098] Then, the attention fusion feature acquisition module 704 can perform weighted summation (Broadcasting Addition) based on the second pulse residual feature, the first pulse residual feature after downsampling, the first nonlinear attention weight and the second nonlinear attention weight to obtain the first attention fusion feature.
[0099] The classification module 705 can classify the sound signal to be classified based on the first attention fusion feature. Specifically, the first attention fusion feature can be pooled using temporal statistical pooling (TSTP) to extract the pooled feature statistics. Next, the pooled feature statistics can be input into the embedding layer for feature mapping, that is, the pooled feature statistics can be mapped to the embedding space through the linear layer. Then, the mapped features can be input into the classification layer for classification, that is, the embedded features can be input into the classification layer for classification.
[0100] It should be noted that, in the present disclosure, the result of classifying the sound signal may be, but is not limited to: various environmental sounds, animal calls, human voices, and the like.
[0101] The advantages of pulse neural network in time dynamic processing and energy-saving and high-efficiency make it perform well in processing sound signals, and it is particularly suitable for sound classification tasks in high-speed, dynamic and complex environments. Therefore, the sound classification method provided by the present disclosure can more efficiently extract and fuse the time-frequency features in sound data by combining pulse neural network and multi-scale fusion technology, significantly improving the accuracy of classification. In addition, it can also effectively reduce the number of parameters to reduce energy consumption.
[0102] According to an exemplary embodiment of the present disclosure, the sound classification device 700 may further include a residual fusion feature acquisition module.
[0103] The residual fusion feature acquisition module can also input the second pulse residual feature into at least one pulse residual fusion module (SpikRF) to obtain the residual fusion feature. In addition, when there are multiple pulse residual fusion modules, these multiple pulse residual fusion modules can be cascaded by first connection.
[0104] The classification module 705 can downsample the above-mentioned first attention fusion feature to obtain a downsampling result. Next, the classification module 705 can input the above-mentioned residual fusion feature and the downsampling result into the attention feature fusion module to obtain a second attention fusion feature. Then, the classification module 705 can classify the sound signal to be classified based on the second attention fusion feature. In addition, the implementation process of "classifying the sound signal to be classified based on the second attention fusion feature" refers to the previous description of "classifying the sound signal to be classified based on the first attention fusion feature", which will not be repeated here.
[0105] According to an exemplary embodiment of the present disclosure, when the number of at least one pulse residual fusion module (SpikRF) is plural, these multiple pulse residual fusion modules can be first connected in sequence to form a cascade connection.
[0106] According to an exemplary embodiment of the present disclosure, the pulse residual fusion module may be configured as follows:
[0107] The input features input to the pulse residual fusion module can be processed by LIF neurons and segmented to obtain multiple segmented parts. Then, for the first segmented part among the multiple segmented parts, the LIF neuron processing can be performed on the first segmented part, and then the segmented part pulse feature corresponding to the first segmented part can be output. Next, for each segmented part in the other segmented parts among the multiple segmented parts, the output features corresponding to each segmented part and the previous segmented part can be input into the attention feature fusion module, and then the attention fusion feature corresponding to each segmented part can be output. Then, the segmented part pulse feature corresponding to the first segmented part and the attention fusion feature corresponding to each segmented part in the other segmented parts can be spliced to obtain a spliced feature. Next, the residual fusion feature can be obtained based on the spliced feature.
[0108] The sound classification method, device, electronic device, storage medium and computer program product provided by the present disclosure combine the advantages of pulse neural networks in processing sound timing information and the advantages of residual neural networks in feature extraction, so that features can be better extracted from sound signals. Specifically, by using pulse neural networks to construct data streams, their advantages in low power consumption and efficient processing of timing data can be fully utilized; in addition, by using residual neural networks for feature extraction, the accuracy of sound recognition can be further improved. It can be seen that the present disclosure can fully combine and utilize the advantages of these two neural networks, thereby achieving efficient and accurate sound classification, and can also significantly reduce system power consumption.
[0109] Figure 8 is a block diagram illustrating an electronic device 800 according to an exemplary embodiment of the present disclosure.
[0110] Reference Figure 8 The electronic device 800 includes at least one memory 801 and at least one processor 802. The at least one memory 801 stores instructions. When the instructions are executed by the at least one processor 802, the sound classification method according to the exemplary embodiment of the present disclosure is executed.
[0111] As an example, the electronic device 800 may be a PC, a tablet device, a personal digital assistant, a smart phone, or other device capable of executing the above instructions. Here, the electronic device 800 is not necessarily a single electronic device, but may also be any device or circuit capable of executing the above instructions (or instruction sets) individually or in combination. The electronic device 800 may also be part of an integrated control system or system manager, or may be configured as a portable electronic device interconnected with a local or remote (e.g., via wireless transmission) interface.
[0112] In the electronic device 800, the processor 802 may include a central processing unit (CPU), a graphics processing unit (GPU), a programmable logic device, a dedicated processor system, a microcontroller or a microprocessor. As an example and not limitation, the processor may also include an analog processor, a digital processor, a microprocessor, a multi-core processor, a processor array, a network processor, etc.
[0113] The processor 802 may execute instructions or codes stored in the memory 801, wherein the memory 801 may also store data. Instructions and data may also be sent and received over a network via a network interface device, wherein the network interface device may employ any known transmission protocol.
[0114] The memory 801 may be integrated with the processor 802, for example, by placing RAM or flash memory within an integrated circuit microprocessor or the like. In addition, the memory 801 may include a separate device, such as an external disk drive, a storage array, or any other storage device that can be used by a database system. The memory 801 and the processor 802 may be operatively coupled, or may communicate with each other, such as through an I / O port, a network connection, etc., so that the processor 802 can read files stored in the memory.
[0115] In addition, the electronic device 800 may also include a video display (such as a liquid crystal display) and a user interaction interface (such as a keyboard, a mouse, a touch input device, etc.) All components of the electronic device 800 may be connected to each other via a bus and / or a network.
[0116] According to an exemplary embodiment of the present disclosure, a computer-readable storage medium may also be provided, and when the instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is enabled to perform the above-mentioned sound classification method. Examples of computer-readable storage media here include: read-only memory (ROM), random access programmable read-only memory (PROM), electrically erasable programmable read-only memory (EEPROM), random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), flash memory, non-volatile memory, CD-ROM, CD-R, CD+R, CD-RW, CD+RW, DVD-ROM, DVD-R, DVD+R, DVD-RW, DVD+RW, DVD-RAM, BD-ROM, BD-R, BD-R LTH, BD-RE, Blu-ray or optical disk storage, hard disk drive (HDD), solid state drive (SSD), card storage (such as, multimedia card, secure digital (SD) card or extreme digital (XD) card), magnetic tape, floppy disk, magneto-optical data storage device, optical data storage device, hard disk, solid state disk and any other device, any other device is configured to store computer programs and any associated data, data files and data structures in a non-transitory manner and provide the computer programs and any associated data, data files and data structures to a processor or computer so that the processor or computer can execute the computer program. The computer program in the above-mentioned computer-readable storage medium can be run in an environment deployed in a computer device such as a client, a host, an agent device, a server, etc. In addition, in one example, the computer program and any associated data, data files and data structures are distributed on a networked computer system, so that the computer program and any associated data, data files and data structures are stored, accessed and executed in a distributed manner by one or more processors or computers.
[0117] According to an exemplary embodiment of the present disclosure, a computer program product may also be provided, including a computer program, and when the computer program is executed by a processor, the sound classification method according to the present disclosure is implemented.
[0118] According to the sound classification method, device, electronic device, storage medium and computer program product disclosed in the present invention, the pulse neural network (SNN) is a neural network inspired by biological neurons, which calculates and transmits information through discrete pulses. Since the pulse neural network (SNN) only calculates when the pulse occurs, it has higher computing efficiency and lower energy consumption when performing sound classification compared to the traditional artificial neural network (ANN). In addition, by using the residual neural network for feature extraction, the accuracy of sound recognition can be guaranteed. It can be seen that the present invention can make full use of the advantages of the pulse neural network (SNN) and the residual neural network, can achieve efficient and accurate sound classification, and can significantly reduce system power consumption.
[0119] Those skilled in the art will readily appreciate other embodiments of the present disclosure after considering the specification and practicing the invention disclosed herein. The present disclosure is intended to cover any variations, uses or adaptations of the present disclosure that follow the general principles of the present disclosure and include common knowledge or customary techniques in the art that are not disclosed in the present disclosure. The description and examples are to be considered exemplary only, and the true scope and spirit of the present disclosure are indicated by the following claims.
[0120] It should be understood that the present disclosure is not limited to the exact structures that have been described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present disclosure is limited only by the appended claims.
Claims
1. A sound classification method, characterized in that: include: Extracting audio features of the sound signal to be classified; Inputting the audio feature into a pulse residual module to obtain a first pulse residual feature; Inputting the first pulse residual feature into at least one of the pulse residual modules to obtain a second pulse residual feature; Inputting the second pulse residual feature and the first pulse residual feature after downsampling into an attention feature fusion module to obtain a first attention fusion feature; classifying the sound signal to be classified based on the first attention fusion feature; Wherein, the pulse residual module is configured as follows: Based on the input features input to the pulse residual module, leaky integration in the form of pulses is triggered by LIF neuron processing to obtain pulse features; Performing residual processing based on the pulse feature and the input feature to obtain a pulse residual feature as an output; The step of performing leaky integration triggering LIF neuron processing in the form of pulses based on the input features input to the pulse residual module to obtain pulse features includes: Performing LIF neuron processing and segmentation processing on the input features to obtain multiple segmented parts; For a first segmented part among the multiple segmented parts, perform LIF neuron processing on the first segmented part to obtain a segmented part pulse feature corresponding to the first segmented part; For each segmented part in other segmented parts among the multiple segmented parts, adding the segmented part pulse features corresponding to each segmented part and the previous segmented part, and performing LIF neuron processing on the added features to obtain the segmented part pulse features corresponding to each segmented part; The pulse features of the segments corresponding to each segment of the multiple segments are spliced together to obtain the pulse features.
2. The sound classification method according to claim 1, characterized in that: Before classifying the sound signal to be classified based on the first attention fusion feature, the method further includes: Inputting the second pulse residual feature into at least one pulse residual fusion module to obtain a residual fusion feature; The classifying the sound signal to be classified based on the first attention fusion feature includes: Downsampling the first attention fusion feature to obtain a downsampling result; Inputting the residual fusion feature and the downsampling result into an attention feature fusion module to obtain a second attention fusion feature; classifying the sound signal to be classified based on the second attention fusion feature; Wherein, the pulse residual fusion module is configured as follows: Performing LIF neuron processing and segmentation processing on the input features input to the pulse residual fusion module to obtain multiple segmented parts; For a first segmented part among the multiple segmented parts, perform LIF neuron processing on the first segmented part, and output a segmented part pulse feature corresponding to the first segmented part; For each segmented part in other segmented parts among the multiple segmented parts, input the output features corresponding to each segmented part and the previous segmented part into the attention feature fusion module, and output the attention fusion features corresponding to each segmented part; splicing the segmented portion pulse feature corresponding to the first segmented portion and the attention fusion feature corresponding to each segmented portion in the other segmented portions to obtain a spliced feature; Based on the splicing features, residual fusion features are obtained.
3. The sound classification method according to claim 1, characterized in that: The step of inputting the second pulse residual feature and the first pulse residual feature after downsampling into an attention feature fusion module to obtain a first attention fusion feature includes: Performing a splicing process on the second pulse residual feature and the downsampled first pulse residual feature to obtain a splicing result; Calculate, based on the splicing result, a first attention weight corresponding to the second pulse residual feature and a second attention weight corresponding to the downsampled first pulse residual feature using a local attention mechanism included in the attention feature fusion module; Performing nonlinear transformation processing on the first attention weight and the second attention weight respectively to obtain a first nonlinear attention weight corresponding to the first attention weight and a second nonlinear attention weight corresponding to the second attention weight; The first attention fusion feature is obtained by weighted summing up the second pulse residual feature, the downsampled first pulse residual feature, the first nonlinear attention weight and the second nonlinear attention weight.
4. The sound classification method according to claim 1, characterized in that: When there are multiple pulse residual modules, the multiple pulse residual modules are cascade-connected.
5. The sound classification method according to claim 2, characterized in that: When there are multiple pulse residual fusion modules, the multiple pulse residual fusion modules are cascade-connected.
6. A sound classification device, characterized in that: include: An audio feature extraction module, configured to extract audio features of a sound signal to be classified; A first pulse residual feature acquisition module is configured to input the audio feature into a pulse residual module to obtain a first pulse residual feature; A second pulse residual feature acquisition module is configured to input the first pulse residual feature into at least one of the pulse residual modules to obtain a second pulse residual feature; an attention fusion feature acquisition module, configured to input the second pulse residual feature and the first pulse residual feature after downsampling into an attention feature fusion module to obtain a first attention fusion feature; a classification module, configured to classify the sound signal to be classified based on the first attention fusion feature; Wherein, the pulse residual module is configured as follows: Based on the input features input to the pulse residual module, leaky integration in the form of pulses is triggered by LIF neuron processing to obtain pulse features; Performing residual processing based on the pulse feature and the input feature to obtain a pulse residual feature as an output; Wherein, the pulse residual module is specifically configured as follows: Performing LIF neuron processing and segmentation processing on the input features to obtain multiple segmented parts; For a first segmented part among the multiple segmented parts, perform LIF neuron processing on the first segmented part to obtain a segmented part pulse feature corresponding to the first segmented part; For each segmented part in other segmented parts among the multiple segmented parts, adding the segmented part pulse features corresponding to each segmented part and the previous segmented part, and performing LIF neuron processing on the added features to obtain the segmented part pulse features corresponding to each segmented part; The pulse features of the segments corresponding to each segment of the multiple segments are spliced together to obtain the pulse features.
7. An electronic device, characterized in that: include: processor; a memory for storing instructions executable by the processor; The processor is configured to execute the instructions to implement the sound classification method as described in any one of claims 1 to 5.
8. A computer-readable storage medium, characterized in that: When the instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is enabled to perform the sound classification method according to any one of claims 1 to 5.
9. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the sound classification method according to any one of claims 1 to 5 is implemented.
Citation Information
Patent Citations
Voice extraction method based on supervised learning and auditory attention, system and device thereof
CN109448749A
Pump machine equipment fault detection method based on attention and integrated learning mechanism
CN114893390A