Audio recognition method, device and system and storage medium
Through artificial neural networks that mimic brain auditory pathways, traditional audio signal processing methods are solved inadequate interpretation of deep-level regular capture and deep learning models, achieving more efficient audio recognition and transparency.
Patent Information
- Application Number
- CN202510995950.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-18
- Publication Date
- 2025-08-29
AI Technical Summary
Traditional audio signal processing methods are difficult to capture the deep rules of music, and deep learning models have problems such as insufficient explanatory and high demand for computing resources.
An artificial neural network that mimics the brain's auditory pathways, including primary auditory modules, Belt and PB modules, mid-temporal and superior temporal modules, and classification decoders, is used to simulate human auditory processing paths through shallow network design to improve audio recognition accuracy and transparency.
It improves the accuracy and computing efficiency of audio recognition, solves the problems of limited generalization capabilities of traditional methods and poor interpretability of deep learning models, and provides a more transparent decision-making process.
Smart Images

Figure CN120561697A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of audio signal processing, and in particular relates to an audio recognition method and device, a system, and a storage medium. Background Art
[0002] Traditional methods for music genre classification rely on feature engineering and hand-crafted classifiers. Classification is achieved by extracting low-level features from audio signals, such as instrument timbre, rhythmic patterns, harmonic structure, and spectral characteristics. These methods often struggle to capture the underlying patterns in music when faced with complex audio signals. The complexity and diversity of audio signals make traditional features incapable of fully capturing the full spectrum of music. Furthermore, the musical styles of each genre vary, necessitating manual feature tuning by domain experts. This reliance on manual design is not only labor-intensive but also limited by the designer's experience, constraining the model's generalization capabilities.
[0003] With the development of deep learning, the ability to extract features automatically has significantly improved, and traditional frameworks have gradually been replaced by deep neural networks. Deep learning can automatically learn richer feature representations from raw audio data through multi-layer neural networks, eliminating the need for manual feature design. However, deep neural network models are often very complex, with numerous parameters, and require a large amount of data and computing resources to train. In addition, the "black box" nature of deep learning models leads to a lack of transparency and difficulty in explaining the model's decision-making process, which poses a challenge in some practical applications. Although deep learning has achieved significant progress in classification accuracy, its lack of interpretability and complexity still pose challenges in optimization and application. Summary of the Invention
[0004] The technical problem to be solved by the present invention is to provide an audio recognition method and device, system, and storage medium, which optimize the processing and classification performance of audio sequences by imitating the structure and working mechanism of the brain's auditory pathway.
[0005] To achieve the above object, the present invention adopts the following technical solutions:
[0006] An audio recognition method, comprising:
[0007] Step S1, obtaining audio data;
[0008] Step S2: obtaining an auditory signal based on the audio data;
[0009] Step S3: Input the auditory signal into an artificial neural network based on the brain's auditory neural pathway for audio recognition; wherein the artificial neural network includes: a primary auditory module, a Belt and PB module, a middle temporal and superior temporal module, and a classification decoder connected in sequence; the primary auditory module, the Belt and PB module, and the middle temporal and superior temporal module are respectively used to imitate the auditory signal processing of the A1, Belt, PB, and T2 / T3 areas in the brain's auditory recognition pathway; a linear classifier is used to convert the auditory signal processing results output by the middle temporal and superior temporal modules into behavioral decisions.
[0010] The present invention also provides an audio recognition device, comprising:
[0011] A first processing module is used to obtain audio data;
[0012] A second processing module is used to obtain an auditory signal based on the audio data;
[0013] The third processing module is used to input the auditory signal into an artificial neural network based on the brain's auditory neural pathway for audio recognition; wherein, the artificial neural network includes: a primary auditory module, Belt and PB modules, middle temporal and superior temporal modules, and a classification decoder connected in sequence; the primary auditory module, Belt and PB modules, middle temporal and superior temporal modules are used to imitate the auditory signal processing of the A1, Belt, PB and T2 / T3 areas in the brain's auditory recognition pathway, respectively; the linear classifier is used to convert the auditory signal processing results output by the middle temporal and superior temporal modules into behavioral decisions.
[0014] The present invention also provides an audio recognition system, comprising: a memory and a processor, wherein the memory stores a computer program to be run by the processor, and the computer program executes the audio recognition method when run by the processor.
[0015] The present invention also provides a storage medium, on which a computer program is stored, and the computer program executes the audio recognition method when running.
[0016] The neural network architecture of this invention simulates the neuroanatomical structure of the human auditory cortex and adopts a shallow network design that better aligns with the human auditory processing pathway. By mimicking the processing methods of the auditory cortex in the human brain, the network can more effectively capture key features in audio signals while improving the network's interpretability, allowing the activation patterns of neurons in each layer to clearly correspond to audio features. This brain-like structure not only improves the model's accuracy in audio recognition but also, to a certain extent, addresses the lack of interpretability of traditional deep learning models, thereby providing a more transparent decision-making process.
[0017] In addition, the use of shallow network design also helps to improve computing efficiency and reduce the consumption of computing resources, making the architecture more efficient in practical applications. Not only does it improve the accuracy of audio recognition, but it also makes the practical application of the model in complex environments more feasible. Therefore, by imitating the structure and function of the human auditory system, the present invention not only effectively improves the effect of audio recognition, but also solves the problems of poor interpretability and low computational efficiency of traditional deep neural networks. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are merely embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying any creative work.
[0019] Figure 1 This is a flow chart of an audio recognition method according to an embodiment of the present invention;
[0020] Figure 2 Schematic diagram of auditory signal processing based on an artificial neural network of the brain's auditory neural pathway according to an embodiment of the present invention;
[0021] Figure 3 Schematic diagram of audio recognition results. DETAILED DESCRIPTION
[0022] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0023] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.
[0024] Example 1:
[0025] like Figure 1 、 2 As shown, an embodiment of the present invention provides an audio recognition method, including:
[0026] Step S1, obtaining audio data;
[0027] Step S2: obtaining an auditory signal based on the audio data;
[0028] Step S3: Input the auditory signal into an artificial neural network based on the brain's auditory neural pathway for audio recognition; wherein the artificial neural network includes: a primary auditory module, a Belt and PB module, a middle temporal and superior temporal module, and a classification decoder connected in sequence; the primary auditory module, the Belt and PB module, and the middle temporal and superior temporal module are respectively used to imitate the auditory signal processing of the A1, Belt, PB, and T2 / T3 areas in the brain's auditory recognition pathway; a linear classifier is used to convert the auditory signal processing results output by the middle temporal and superior temporal modules into behavioral decisions.
[0029] As an implementation of the embodiment of the present invention, in step S2, based on the audio data F t Computing decoding c from cochlear fields t , extracting the ventral pathway of audio features; based on the cochlear response c t Get the auditory signal map m t .
[0030] As an implementation method of an embodiment of the present invention, in step S3, artificial neural networks (ANNs) are used as the basic architecture, because their neurons are the basic units of information processing, and the response of each neuron in the network can be directly mapped to the activation pattern of the cerebral cortex. Given that audio sequences are time-dependent, recursive neural networks (RNNs) are particularly suitable for auditory recognition tasks due to their ability to naturally adapt to time series data. Specifically, the neural responses in the auditory ventral pathway exhibit strong temporal characteristics, so the brain-based auditory recognition network (BAN) is designed to gradually generate activation responses over time to more accurately simulate the brain's auditory processing mechanism.
[0031] In terms of predictive performance, models aligned with the brain's neuroanatomical structure and neural response patterns exhibit more accurate intermediate layer and final output behavior. This alignment mechanism enables the proposed artificial neural network to effectively infer brain activation patterns and accurately perform category selection. As a result, the artificial neural network maintains high recognition performance while also simulating, to a certain extent, the brain's mechanisms and decision-making processes when processing audio signals.
[0032] The goal of this invention is to achieve a high degree of similarity between artificial neural networks and the human brain in auditory recognition tasks. Specifically, in the human brain's auditory recognition pathway, the primary auditory cortex (A1) performs the initial processing of input signals, the Belt and PB areas integrate sound signals across spatial dimensions, and the T2 / T3 areas generate predictive auditory labels. The design of the neural network fully considers the implications of this biological process, aiming to achieve auditory recognition capabilities that are closer to the brain and provide a more biologically meaningful model for neural information processing.
[0033] Drawing on the workings of the brain's auditory recognition pathway, this paper constructs a neuroanatomical map that maps cortical regions to different layers in an artificial neural network (ANN). To facilitate comparison of neural network structures, this map is created by identifying the ANN layer that best corresponds to activation in a specific cortical region. Ideally, these responses should be automatically inferred by the neural network without the need for additional parameters. The specific inputs, operations, and outputs of each neural network layer are shown in Table 1.
[0034] Table 1
[0035]
[0036]
[0037] Primary Auditory Module A1: The processed auditory signal reaches the primary auditory module A1 of the network, which is crucial for processing basic sound features such as pitch and volume. The input is a data matrix represented by the time-frequency representation m t , the input is a mel-spectrogram. Each module is designed with a specific neural circuit, in which neurons perform basic computational tasks such as convolution, addition, normalization, nonlinear processing, or pooling. Except for the primary auditory module, the circuit structure of other auditory modules remains consistent, but the total number of neurons in each module varies. To manage high computational costs, the primary auditory module A1 applies a 7×7 convolution with a stride of 2, followed by a 3×3 max pooling, and then a 3×3 convolution.
[0038] Furthermore, when the signal enters the primary auditory module A1, a 7×7 convolution operation is first performed. The convolution kernel is in the Mel spectrum data matrix m t The calculation is performed on the local area data, and the dot product is performed and summed to obtain a new feature map. This feature map highlights the local time-frequency features in the input data. Different convolution kernel weight settings determine the type of features extracted. Since the stride is 2, the feature map is downsampled in the spatial dimension, reducing the amount of subsequent calculations. Then, 3×3 maximum pooling is performed. Within the 3×3 window, the maximum value is selected as the output, and the feature map is further downsampled in the spatial dimension, while retaining the most significant feature information, suppressing some unimportant details, and enhancing feature selectivity. Finally, a 3×3 convolution is performed, and the downsampled and feature-extracted data is convolved again to further explore the complex time-frequency patterns in the data, enabling the module to learn more advanced auditory features.
[0039] Belt module and PB module: The belt cortex (Belt) and parabelt cortex (PB) are key areas in the auditory cortex, mainly responsible for processing complex auditory information. The belt module is located around the primary auditory module A1, representing the first-level auditory cortical module beyond the primary auditory module A1, and directly receives input from the primary auditory module v t As a secondary auditory processing module, the Belt module is crucial for analyzing more complex sound features than the primary auditory module A1. It integrates information from the primary auditory module A1 and provides a more refined understanding of the sound, such as recognizing complex patterns and the spatial location of the sound source. The PB module is located next to the Belt module and represents a higher level of auditory processing. It receives input from the Belt module. t The PB module plays a key role in high-level auditory scene analysis, such as distinguishing multiple sounds in a noisy environment or interpreting modulated sounds like music and speech. It also integrates auditory information with other sensory inputs, helping to form a unified perception of the surrounding environment.
[0040] The output v of the primary auditory module A1 t With the gate signal g t Performs calculations and controls the signals entering the Belt module. The gating mechanism regulates the flow of information and determines which feature information can enter the subsequent processing stage.
[0041] The signal after gated processing enters the Belt module and undergoes two convolution operations, and the output is b t This module performs more complex feature extraction on the signal and mines more advanced sound features than the primary auditory module A1.
[0042] The PB module receives the output b of the Belt module t , and the same convolution operation is performed twice, and the output is p t The PB module further processes the sound information, focusing on advanced auditory scene analysis, such as distinguishing sounds in complex environments.
[0043] The Belt module performs two convolution operations on the gated v t The PB module receives the output of the Belt module b and processes the related signals. Based on the basic features extracted by the primary auditory module A1, it further mines the complex sound features. t , through two more convolution operations, we continue to deepen the analysis of sound features, process more advanced auditory scene related features, and output p t .
[0044] The middle and superior temporal modules (T2 / T3 modules): Areas T2 and T3 are important brain regions involved in multiple aspects of sensory processing, particularly auditory recognition. Located in the temporal lobe, these areas are crucial for interpreting and understanding complex auditory signals, such as music. Area T2 is primarily known for its role in auditory processing, particularly in perceptual motion and integrating audiovisual information. During auditory recognition, T2 cortex becomes active when visual cues need to be integrated with auditory signals. The superior temporal gyrus (STG) within area T3 plays a direct role in auditory processing and is crucial for recognizing complex sounds, such as music. Within the STG, Hirschl's gyrus (HG) is the first brain region to receive audio input, while surrounding areas, including the planum temporale (PT), are involved in processing higher-order sound features, such as speech comprehension. The posterior portion of the STG and the nearby superior temporal sulcus contribute to the analysis of more complex sound attributes, such as intonation and the emotional content of speech. T3 also has extensive connections with other brain regions, supporting its role in integrating auditory information with other sensory modalities and cognitive functions.
[0045] During auditory recognition, the Belt, PB, and T2 / T3 modules work together to jointly perform the tasks of sound decoding and interpretation, enabling individuals to accurately recognize and respond to various auditory stimuli, such as distinguishing speech sounds, understanding language content, and appreciating music. Specifically, the Belt, PB, and T2 / T3 modules each perform two 1×1 convolutions, followed by a bottleneck 3×3 convolution (which quadruples the feature set), and finally complete feature extraction with a single 1×1 convolution. This recursive operation is achieved by feeding the output of a processing module back to the module multiple times. For example, after processing the initial input features, the Belt module processes the results as new inputs for further processing. The Belt and PB modules are repeated twice, respectively, while the T2 / T3 modules are repeated four times. This configuration achieves optimal model performance with a small number of layers. This architecture draws on the principles of ResNet, with batch normalization and ReLU activation functions following each convolutional module to enhance the model's stability and nonlinear representation capabilities.
[0046] Furthermore, the middle and superior temporal cortices are located in the temporal lobe. The T2 area performs outstandingly in auditory processing, especially in perceiving movement and integrating audiovisual information. The superior temporal gyrus (STG) in the T3 area plays a direct role in auditory processing, the Hirschl's gyrus (HG) receives input audio, and the surrounding planum temporale (PT) participates in high-order sound feature processing, such as language comprehension. The posterior STG and the superior temporal sulcus are responsible for analyzing complex sound attributes such as intonation and emotion, and the T3 area is extensively connected with many other brain areas.
[0047] Similar to the Belt and PB modules, the T2 / T3 modules implement a specific convolutional process. They first perform two 1×1 convolutions to compress or transform feature dimensions. Next, they perform a bottleneck 3×3 convolution to quadruple the feature set and deeply mine feature information. Finally, a 1×1 convolution completes feature extraction and integrates optimized features.
[0048] The T2 / T3 module feeds the output of the processing module back to the module for recursive operation four times. This method optimizes the model performance with fewer layers.
[0049] Batch normalization is added after each convolution module to make the data distribution more stable and accelerate model convergence. It is also followed by the ReLU activation function to introduce nonlinearity and enhance the model's ability to express complex patterns.
[0050] T2 / T3 module receives PB module output p t After four convolution operations, the sound features are deeply integrated and abstracted, gradually transitioning from primary features to high-level, abstract feature representations to support the final sound recognition decision, and the output is d t .
[0051] Classification decoder: A linear classifier is used to perform the computation using a weighted sum, with each object label corresponding to a separate weighted sum. To reduce the number of neural activations fed into the classifier, the responses of each feature map are aggregated by averaging, effectively reducing computational complexity.
[0052] Furthermore, the mid-temporal and superior temporal modules (T2 / T3 modules) process the input audio and output a series of feature maps. To reduce the number of neural activations input to the classifier, the framework requires that the responses of each feature map be averaged and aggregated. The information from multiple feature maps is then combined into a single feature vector, which serves as the input for the subsequent weighted sum calculation.
[0053] Clearly, there are multiple different object labels, such as blues, reggae, classical, and rock. A set of weights is assigned to each object label, constructing a weight matrix. Each element in the matrix represents the importance a particular object label places on a particular feature dimension. These weights are typically learned and adjusted based on large amounts of data using an optimization algorithm during model training.
[0054] For each object label, its corresponding weight vector is multiplied by the corresponding element of the processed feature vector, and then all products are added together to obtain the weighted sum corresponding to the object label. This calculation is performed independently for each object label to obtain its own weighted sum.
[0055] After calculating the weighted sums of all object labels, the weighted sums are compared. The object label with the largest weighted sum is usually selected as the final classification result to achieve audio signal classification.
[0056] Artificial neural networks are trained by optimizing a combination of loss functions, including recognition loss and auxiliary loss. The total network model loss L b The definition is as follows:
[0057] L b =L r +L u
[0058] Among them, L r is the recognition loss, L u It is auxiliary loss.
[0059] Recognition loss L r for:
[0060]
[0061] Among them, L r is the true label d i With the predicted label p i The cross entropy loss between , N is the number of samples. Auxiliary loss L u for:
[0062]
[0063] Among them, φ t is the parameter that changes at different time t, and θ is the network parameter.
[0064] This example uses the GTZAN dataset, one of the most widely used datasets for music genre identification. This dataset contains 30-second audio files from 10 different genres: blues, reggae, classical, rock, country, disco, jazz, pop, metal, and hip-hop. From the original collection, this example randomly selected 54 pieces of music from each genre, resulting in a total of 540 pieces of music for research. All audio clips were normalized based on their root mean square value.
[0065] The fMRI experiment consisted of 12 training runs and 6 test runs, for a total of 18 runs. Each run lasted 10 minutes and contained 40 music excerpts. During the training phase, 480 music excerpts were used, while the remaining 60 excerpts were used in the test runs. In each test run, a set of 10 music excerpts was presented four times in the same order. No excerpts were repeated during the training runs.
[0066] To evaluate the similarities between artificial neural networks (ANNs) and human auditory recognition, embodiments of the present invention extract task-related activations and identify corresponding brain regions for comparison.
[0067] fMRI data processing: Motion correction was performed for each experimental run using the Statistical Parametric Mapping Toolbox (SPM 12), and all volumetric images were aligned to the first image of each participant. The present embodiment used a filter with a 240-second window to remove low-frequency drift. To improve the accuracy of the method, the present embodiment normalized the activation of each voxel by subtracting the mean and variance. The cortical surface was identified using FreeSurfer, a tool that aligns anatomical data with functional voxels. During analysis, only cortical voxels were used as targets, and the present embodiment focused on the voxels identified within the cortex for each participant.
[0068] Genre representation regions: In order to obtain reliable estimates of human brain regions associated with music genre, the present embodiment adopts the following steps: First, all activation records are randomly divided into a training set (75%) and a test set (25%). Using the optimal weights of the genre label model, the present embodiment uses the training data to fit an encoding model and uses the test set to evaluate the accuracy of the model. Parameter fitting is performed using a generalized linear model. Random resampling is performed 100 times, and voxels with a prediction accuracy exceeding 75% are selected as regions of interest. Therefore, the present embodiment induced 473 voxels of interest for each participant, 468 for participant sub-01, 581 for sub-02, 1593 for sub-03, and 529 for sub-05. All subsequent analyses used these extracted regions of interest.
[0069] The present invention found that the existing data is insufficient to effectively train a convolutional neural network (CNN). To improve this problem, the present invention uses data augmentation technology to generate new synthetic samples by fine-tuning the original data, thereby expanding the data set. The purpose of data augmentation is to improve the robustness of the model in the face of these changes and enhance its generalization ability. To ensure the effectiveness of this method, the added perturbations need to keep the labels of the original samples unchanged. Common enhancement methods include adding noise, adjusting the starting position of the audio, changing the playback speed, and changing the pitch.
[0070] This embodiment of the present invention uses Mel-Frequency Cepstral Coefficients (MFCCs) to extract features related to different musical styles. For timbre and loudness features, this embodiment sets the duration of each frame to 25 milliseconds, with a 50% overlap between adjacent frames. For pitch and rhythm information, this embodiment sets each frame duration to 3 seconds, with a 33% overlap, and averages each feature within a 1.5-second time window. To extract MFCC features, this embodiment of the present invention also utilizes the MIR toolbox.
[0071] In terms of model architecture, this embodiment of the present invention uses an artificial neural network (ANN) to represent each auditory cortex region, where the neural network performs common operations such as convolution. These modules correspond to the auditory cortex regions, and this embodiment of the present invention optimizes the model by adjusting the number of neurons in each region. Considering the computational requirements, this embodiment of the present invention uses the Python programming language and the PyTorch framework to implement this model.
[0072] In order to verify whether the participants' brain activities and the network model proposed by the present invention can effectively distinguish different music labels in the experiment, the embodiment of the present invention uses a decoding model based on brain activity and neural network for music recognition. Figure 3 As shown, an embodiment of the present invention evaluated the confusion matrix and its diagonal elements (i.e., recognition accuracy) by analyzing brain responses within regions of interest associated with specific music types. The recognition results showed that the classification performance between different music labels varied, with the classification accuracy of classical music always reaching 100%, while the classification accuracy of rock music was lower among different participants, at only 40%. In addition, participants often misclassified reggae music as rock (with a confusion rate of 33.3%) and rock as country music (with a confusion rate of 28.6%). By analyzing the data between participants, the resulting confusion matrix showed a high degree of consistency (Pearman correlation coefficient ρ = 0.562 ± 0.087; p-values for all participant combinations were less than 0.001).
[0073] Example 2:
[0074] An embodiment of the present invention further provides an audio recognition device, comprising:
[0075] A first processing module is used to obtain audio data;
[0076] A second processing module is used to obtain an auditory signal based on the audio data;
[0077] The third processing module is used to input the auditory signal into an artificial neural network based on the brain's auditory neural pathway for audio recognition; wherein, the artificial neural network includes: a primary auditory module, Belt and PB modules, middle temporal and superior temporal modules, and a classification decoder connected in sequence; the primary auditory module, Belt and PB modules, middle temporal and superior temporal modules are used to imitate the auditory signal processing of the A1, Belt, PB and T2 / T3 areas in the brain's auditory recognition pathway, respectively; the linear classifier is used to convert the auditory signal processing results output by the middle temporal and superior temporal modules into behavioral decisions.
[0078] Example 3:
[0079] An embodiment of the present invention further provides an audio recognition system, comprising: a memory and a processor, wherein the memory stores a computer program executed by the processor, and the computer program executes an audio recognition method when executed by the processor.
[0080] Example 4:
[0081] An embodiment of the present invention further provides a storage medium, wherein a computer program is stored on the storage medium, and the computer program executes the audio recognition method when running.
[0082] The embodiments described above are merely descriptions of preferred embodiments of the present invention and are not intended to limit the scope of the present invention. Without departing from the spirit of the present invention, various modifications and improvements made to the technical solutions of the present invention by persons skilled in the art should fall within the scope of protection defined by the claims of the present invention.
Claims
1. An audio recognition method, characterized in that: include: Step S1, obtaining audio data; Step S2: obtaining an auditory signal based on the audio data; Step S3: Input the auditory signal into an artificial neural network based on the brain's auditory neural pathway for audio recognition; wherein the artificial neural network includes: a primary auditory module, a Belt and PB module, a middle temporal and superior temporal module, and a classification decoder connected in sequence; the primary auditory module, the Belt and PB module, and the middle temporal and superior temporal module are respectively used to imitate the auditory signal processing of the A1, Belt, PB, and T2 / T3 areas in the brain's auditory recognition pathway; a linear classifier is used to convert the auditory signal processing results output by the middle temporal and superior temporal modules into behavioral decisions.
2. An audio recognition device, characterized in that: include: A first processing module is used to obtain audio data; A second processing module is used to obtain an auditory signal based on the audio data; The third processing module is used to input the auditory signal into an artificial neural network based on the brain's auditory neural pathway for audio recognition; wherein, the artificial neural network includes: a primary auditory module, Belt and PB modules, middle temporal and superior temporal modules, and a classification decoder connected in sequence; the primary auditory module, Belt and PB modules, middle temporal and superior temporal modules are used to imitate the auditory signal processing of the A1, Belt, PB and T2 / T3 areas in the brain's auditory recognition pathway, respectively; the linear classifier is used to convert the auditory signal processing results output by the middle temporal and superior temporal modules into behavioral decisions.
3. An audio recognition system, characterized in that: include: A memory and a processor, wherein the memory stores a computer program executed by the processor, and when the computer program is executed by the processor, the audio recognition method according to claim 1 is executed.
4. A storage medium, characterized in that The storage medium stores a computer program, which executes the audio recognition method according to claim 1 when running.