Vocal music feature decoupling identification method and system based on mutual information minimization
By using a dual-stream decoupled deep neural network based on mutual information minimization and physiological bias correction technology, the problem of coupling between sound source features and vocal tract features in vocal feature recognition is solved, achieving accurate quantification and feedback of nasal sounds, and improving the accuracy of vocal teaching and self-learning efficiency.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HUAZHONG UNIV OF SCI & TECH
- Filing Date
- 2026-03-04
- Publication Date
- 2026-05-08
AI Technical Summary
Existing vocal feature recognition methods suffer from a high degree of coupling between sound source features and vocal tract features, making it difficult to provide effective guidance when pitch is correct but vocalization is incorrect. Furthermore, they lack the ability to dynamically adjust evaluation criteria according to different artistic styles, resulting in evaluation results that lack physical interpretability and accuracy. Users also find it difficult to understand the feedback mechanism, leading to low self-learning efficiency.
A dual-stream decoupled deep neural network based on mutual information minimization is employed. Through constant Q-transform and physiological bias correction, sound source features and vocal tract features are separated. Combined with individual acoustic feature benchmarks and dynamic nasalization coefficients, accurate quantification and feedback of nasal sounds are achieved.
It achieves precise decoupling and recognition of vocal features, provides physical-level feature decoupling, improves the accuracy and self-learning efficiency of nasal sound recognition, adapts to evaluation standards of different artistic styles, and provides intuitive physiological feedback.
Smart Images

Figure CN121999804A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the interdisciplinary field of audio signal processing and artificial intelligence-assisted education, and more specifically, relates to a method and system for decoupling and recognizing vocal features based on minimizing mutual information. Background Technology
[0002] Vocal feature decoupling and identification refers to the process of separating and identifying independent acoustic features generated by different physiological structures (such as vocal cords, oral cavity, and nasal cavity) from singing audio signals. Accurate decoupling is a crucial prerequisite for objectively evaluating vocal techniques, especially for timbre control (such as nasal tone management). However, in actual acoustic signals, the sound source features (such as pitch and loudness) generated by vocal cord vibration are highly coupled with the vocal tract features (such as formants) generated by oral and nasal cavity modulation. This poses a significant challenge to refined vocal teaching based on non-contact audio.
[0003] Currently, methods for identifying and evaluating vocal features include the following: The first is acoustic feature analysis based on traditional signal processing, which detects nasal sounds by calculating the energy ratio of specific frequency bands, or uses fundamental frequency extraction algorithms to evaluate pitch and rhythm; the second is an end-to-end black-box recognition method based on deep learning models, which trains neural networks with a large amount of labeled data to directly map the nonlinear relationship between audio and evaluation results; the third is a method that relies on dedicated hardware devices, which uses contact or invasive sensors such as nasal flow meters and accelerometers to directly measure nasal cavity vibration or airflow.
[0004] However, all of the above-mentioned existing methods have some significant drawbacks: First, the acoustic feature analysis method based on traditional signal processing is easily affected by the change in fundamental harmonic energy when a singer sings high notes, which leads to serious misjudgment of the key formant that characterizes nasal sounds and makes it difficult to provide effective guidance in complex situations where the pitch is correct but the vocalization is wrong. Second, the aforementioned end-to-end black-box recognition method based on deep learning models suffers from a serious feature coupling problem. The model often fails to distinguish between "obstructive nasal sounds" caused by physiological structures such as colds and rhinitis and "functional nasal sounds" caused by improper vocal techniques, resulting in a lack of physical interpretability and accuracy in the evaluation results. Third, the acoustic feature analysis methods based on traditional signal processing have inherent limitations in the low-frequency band due to insufficient frequency resolution. They are difficult to clearly capture the fine acoustic structures used to distinguish dense low-frequency harmonics from nasal cavity anti-resonance points, which brings physical obstacles to the accurate quantitative assessment of nasal tone. Fourth, the aforementioned methods that rely on dedicated hardware devices usually use abstract waveforms or spectrograms as feedback, which are difficult for ordinary users to understand and associate with their own physiological movements (such as the rise and fall of the soft palate). This non-intuitive feedback mechanism makes it difficult for users to establish correct muscle memory, which greatly limits the efficiency of self-learning. Fifth, the acoustic feature analysis methods based on traditional signal processing and the end-to-end black box recognition methods based on deep learning models mentioned above usually use fixed acoustic thresholds for judgment, lacking the ability to dynamically adjust evaluation criteria according to different art styles such as bel canto and pop, which can easily lead to "false corrections" of artistic expression. Summary of the Invention
[0005] To address the aforementioned deficiencies or improvement needs of existing technologies, this invention provides a vocal feature decoupling and recognition method and system based on mutual information minimization. Its purpose is to solve the technical problems of existing acoustic feature analysis methods based on traditional signal processing, which are highly susceptible to interference from changes in fundamental harmonic energy during high-note singing, leading to serious misjudgments in the identification of key formants characterizing nasal sounds and making it difficult to provide effective guidance in complex situations where pitch is correct but vocalization is incorrect; and the serious feature coupling problem in existing end-to-end black-box recognition methods based on deep learning models, where the models often cannot distinguish between "obstructive nasal sounds" caused by physiological structures such as colds and rhinitis and "functional nasal sounds" caused by improper vocal techniques, resulting in a lack of physical interpretability and accuracy in evaluation results. The existing acoustic feature analysis methods suffer from inherent limitations in the low-frequency band due to insufficient frequency resolution. This makes it difficult to clearly capture the fine acoustic structures used to distinguish dense low-frequency harmonics from nasal cavity anti-resonance points, creating a physical obstacle to the accurate quantitative assessment of nasal tone. Furthermore, existing methods relying on dedicated hardware typically use abstract waveforms or spectrograms as feedback, which are difficult for ordinary users to understand and relate to their own physiological movements, hindering the establishment of correct muscle memory and significantly limiting self-learning efficiency. Additionally, existing acoustic feature analysis methods based on traditional signal processing and end-to-end black-box recognition methods based on deep learning models often use fixed acoustic thresholds for judgment, lacking the ability to dynamically adjust evaluation criteria according to different artistic styles such as bel canto and pop, which can easily lead to miscorrection of artistic expression.
[0006] To achieve the above objectives, according to one aspect of the present invention, a nasal sound recognition method based on acoustic feature decoupling is provided, comprising the following steps: (1) Acquire the real-time singing audio signal to be tested, and perform microphone frequency response compensation on the real-time singing audio signal to obtain a standard digital audio stream; (2) Perform constant Q-transform (CQT) on the standard digital audio stream obtained in step (1) to obtain the time-frequency tensor; (3) Input the time-frequency tensor obtained in step (2) into the pre-trained dual-stream decoupled deep neural network to obtain the sound source feature vector and the vocal tract transmission feature vector; (4) Obtain the individual acoustic feature reference of the singer corresponding to the real-time singing audio signal obtained in step (1), and use the sound source feature vector obtained in step (3) to remove the fundamental frequency interference generated by vocal cord vibration in the individual acoustic feature reference, thereby obtaining the processed individual acoustic feature reference, and use the processed individual acoustic feature reference to perform physiological bias correction on the vocal tract transmission feature vector obtained in step (3) to obtain the normalized dynamic nasalization coefficient. (5) Obtain the final nasal sound recognition result based on the normalized dynamic nasalization coefficient obtained in step (4).
[0007] Preferably, the identification result specifically refers to whether the singer corresponding to the real-time singing audio signal has normal nasal tone, high nasal tone defect, low nasal tone blockage, or stylized nasal tone.
[0008] Preferably, if the dynamic nasalization coefficient is between 0.2 and 0.5, it indicates that the singer corresponding to the real-time singing audio signal has normal nasal tone, indicating that the ratio of oral and nasal resonance is in an acoustically balanced state, which conforms to standard vocal science. If the dynamic nasalization coefficient is between 0.7 and 1.0, it indicates that the singer corresponding to the real-time singing audio signal has a high nasal tone defect, indicating that the soft palate is drooping or not fully closed, resulting in too much airflow entering the nasal cavity and producing a noticeable "noisy" feeling. If the dynamic nasalization coefficient is between 0 and 0.2, it indicates that the singer corresponding to the real-time singing audio signal has low nasal tone blockage, indicating that their nasal passage is not smooth, resulting in the absence of necessary nasal resonance frequencies in the sound signal. If the dynamic nasalization coefficient is between 0.5 and 0.7, it indicates that the singer corresponding to the real-time singing audio signal has stylized nasal technique, indicating that they are using a specific singing technique.
[0009] Preferably, the two-stream decoupled deep neural network adopts a 12-layer structure, wherein: Layer 1 is a shared feature extraction layer, whose input is a dimension of... The layer performs a convolutional normalization activation operation on the single-channel CQT time-frequency tensor to obtain a dimension of The feature map is generated and output; the parameters of the convolutional normalized activation operation are: 64 output channels, 3x3 kernel size, stride 1, and padding value 1. This represents the total number of frequency points on the frequency axis of the single-channel CQT time-frequency tensor, which is used to characterize the audio resolution of the single-channel CQT time-frequency tensor. This indicates the number of time frames of the single-channel CQT time-frequency tensor on the time axis, which is used to characterize the duration of the single-channel CQT time-frequency tensor. The second layer is the feature distribution layer, whose input is the dimension of the output of the shared feature extraction layer. The feature map is simultaneously output to the first source coding branch layer and the first channel coding branch layer. Layer 3 is the first source coding branch layer. Its input is the feature map output from the feature distribution layer. This layer performs a one-dimensional depthwise separable convolution operation on this feature map to obtain a dimension of... The feature map is generated and output, where the parameters of the one-dimensional depthwise separable convolution operation are: 128 output channels, 5 convolution kernels, and 1 stride. The fourth layer is the second source coding branch layer. Its input is the feature map output by the first source coding branch layer. This layer performs instance normalization on the feature map and performs feature reconstruction on the normalized feature map to obtain source coding features and output them.
[0010] Preferably, the fifth layer is the first channel coding branch layer, whose input is the feature map output by the feature distribution layer. This layer performs global average pooling and linear mapping processing on the feature map in sequence to obtain the spectral weighting coefficients for the first resonance peak (F1 band and 2kHz-4kHz nasal cavity resonance sensitive area), and uses the spectral weighting coefficients to perform weighted enhancement processing to obtain the spectrally enhanced channel feature map and output it. Layer 6 is the second channel coding branch layer, and its input is the dimension of the output of the first channel coding branch layer. The layer performs multi-scale convolution on the spectral-enhanced vocal tract feature map to obtain multiple multi-dimensional resonant peak envelope features reflecting the physical morphology of the oral and nasal cavities. All multi-dimensional resonant peak envelope features are then concatenated to obtain a dimension-... The vocal tract morphology feature map is generated and output.
[0011] Layer 7 is the mutual information adversarial discrimination layer, whose input is the dimension of the output of the second channel coding branch layer. The layer first performs a three-level fully connected mapping process on the vocal tract morphology feature map to obtain the probability distribution result used to predict sound source information. Then, the layer introduces the real pitch label corresponding to the standard digital audio stream, calculates the cross-entropy loss between the predicted probability distribution result and the real pitch label, and outputs the cross-entropy loss value as the mutual information loss value. The 8th layer is a gradient inversion layer, whose input is the dimension of the output of the second channel coding branch layer. The layer generates a 128-channel morphological feature map; during the forward propagation phase, it performs an identity mapping operation on the morphological feature map; and during the backward propagation phase, it performs an inversion process on the gradient returned by the mutual information adversarial discrimination layer to obtain the inverse gradient and passes it back to the second channel coding branch layer.
[0012] Preferably, the 9th layer is a global pooling layer, whose inputs are the dimensions of the output of the second source coding branch layer. The source coding feature map and the dimension of the output of the second channel coding branch layer are... The layer performs global average pooling on the two feature maps, respectively, to reduce their spatial dimensions. Compress to 1 to obtain and output sound source feature vectors and vocal tract feature vectors, each with a dimension of 128. The 10th layer is a two-stream independent mapping layer. Its input is the source feature vector and the vocal tract feature vector output by the global pooling layer. This layer performs fully connected linear projection processing on the source feature vector and the vocal tract feature vector respectively to obtain source semantic vector and vocal tract semantic vector with a dimension of 64 and output them respectively. Layer 11 is the dimensionality reduction feature output layer. The input to this layer is the source semantic vector and the vocal tract semantic vector output from the dual-stream independent mapping layer. This layer performs fully connected dimensionality reduction processing on the source semantic vector and the vocal tract semantic vector respectively to obtain compact source feature vectors with a dimension of 32. and vocal tract feature vectors ; The 12th layer is the output interface layer. Its input is the compact sound source feature vector and the vocal tract transmission feature vector output by the dimensionality reduction feature output layer. This layer performs output mapping processing on these two feature vectors, that is, the vocal tract transmission feature vector is output as a physical feature representing the resonance state of the oral cavity and nasal cavity, and the compact sound source feature vector is output as a physical feature representing the vibration state of the vocal cords, thereby obtaining the final acoustic feature decoupling recognition result.
[0013] Preferably, the two-stream decoupled deep neural network is trained through the following steps: (1-1) Obtain the original singing audio sample set and divide the original singing audio sample set into a training set and a test set in an 8:2 ratio; (1-2) Perform a constant Q transformation on the training set obtained in step (1-1) to obtain the transformed training set; (1-3) Perform data augmentation on the transformed training set obtained in step (1-2) to obtain the augmented training set; (1-4) For each dimension of the augmented training set obtained in step (1-3), the following is... For the singing audio sample, the singing audio sample is input into the shared feature extraction layer of a two-stream decoupled deep neural network and subjected to convolutional normalization activation processing to obtain the corresponding feature sample with dimension [missing information]. Shallow feature map; (1-5) For each singing audio sample in the enhanced training set obtained in step (1-3), the dimension of the singing audio sample corresponding to that sample obtained in step (1-4) is... The shallow feature map is input into the feature distribution layer of the two-stream decoupled deep neural network for distribution processing, so as to obtain two identical feature maps with the same dimension. Feature map; (1-6) For each singing audio sample in the enhanced training set obtained in step (1-3), the feature map corresponding to the singing audio sample obtained in step (1-5) is input into the first source coding branch layer of the dual-stream decoupled deep neural network for one-dimensional depthwise separable convolution processing to obtain the feature map corresponding to the singing audio sample with dimension [missing information]. The intermediate feature map of the sound source; (1-7) For each singing audio sample in the enhanced training set obtained in step (1-3), the dimension of the singing audio sample corresponding to that sample obtained in step (1-6) is... The intermediate feature map of the sound source is input into the second sound source coding branch layer of the dual-stream decoupled deep neural network for instance normalization and feature reconstruction to obtain the dimension corresponding to the singing audio sample. The source coding feature map; (1-8) For each singing audio sample in the training set after enhancement processing obtained in step (1-3), the other feature map corresponding to the singing audio sample obtained in step (1-5) is input into the first channel coding branch layer of the dual-stream decoupled deep neural network for prediction value calculation and weighted enhancement processing, so as to obtain the corresponding feature map of the singing audio sample with dimension . Spectral enhancement acoustic spectral characteristics; (1-9) For each singing audio sample in the enhanced training set obtained in step (1-3), the dimension of the singing audio sample corresponding to that sample obtained in step (1-8) is... The spectral enhancement tract feature map is input into the second tract coding branch layer of the dual-stream decoupled deep neural network and subjected to multi-scale convolution and concatenation processing to obtain the dimensional data corresponding to the singing audio sample. A diagram showing the morphological characteristics of the vocal tract; (1-10) For each singing audio sample in the enhanced training set obtained in step (1-3), the dimension of the singing audio sample corresponding to that sample obtained in step (1-9) is... The vocal tract morphology feature map is input into the mutual information adversarial discrimination layer of the dual-stream decoupled deep neural network and processed by a three-level fully connected mapping process to obtain the prediction probability distribution result corresponding to the singing audio sample, which is used to predict the sound source information. (1-11) For each singing audio sample in the training set after enhancement processing obtained in step (1-3), calculate the cross-entropy loss value based on the predicted probability distribution result corresponding to the singing audio sample obtained in step (1-10) and the real pitch label corresponding to the singing audio sample, and use it as the mutual information loss corresponding to the singing audio sample. ; (1-12) For each singing audio sample in the training set after enhancement processing obtained in step (1-3), the mutual information loss corresponding to the singing audio sample obtained in step (1-11) is updated in reverse using the gradient reversal layer and gradient descent method to obtain the reverse gradient, and the weight of the sound source coding branch is obtained using the reverse gradient. (1-13) For each singing audio sample in the enhanced training set obtained in step (1-3), the source coding feature map corresponding to the singing audio sample obtained in step (1-7) and the vocal tract morphology feature map corresponding to the singing audio sample obtained in step (1-9) are respectively input into the global pooling layer of the dual-stream decoupled deep neural network for global average pooling processing to obtain the two-dimensional feature map corresponding to the singing audio sample. A fixed-length eigenvector; (1-14) For each singing audio sample in the training set after enhancement processing obtained in step (1-3), the two fixed-length feature vectors obtained in step (1-13) are input into the independent mapping layer of the two-stream decoupled deep neural network for independent fully connected linear projection processing to obtain the two-dimensional feature vectors corresponding to the singing audio sample. semantic vector; (1-15) For each singing audio sample in the training set after enhancement processing obtained in step (1-3), the two semantic vectors corresponding to the singing audio sample obtained in step (1-14) are input into the dimensionality reduction feature output layer of the dual-stream decoupled deep neural network for dimensionality reduction processing, so as to obtain the final dimension of the singing audio sample. Compact sound source feature vector and channel transmission feature vector; (1-16) Repeat steps (1-4) to (1-15) until the mutual information loss obtained in step (1-11) converges or reaches the preset number of iterations, so as to obtain a pre-trained two-stream decoupled deep neural network. (1-17) Use the test set obtained in step (1-1) to fine-tune and calibrate the preliminarily trained dual-stream decoupled deep neural network obtained in step (1-16) until the cumulative error of the dynamic quantization coefficient of the dual-stream decoupled deep neural network is lower than the preset threshold, so as to obtain the final trained dual-stream decoupled deep neural network.
[0014] Preferably, the original singing audio sample set in step (1-1) includes bel canto, pop, folk singing, as well as the vocal range from C2 to C6 in the three octaves, and various singing audio samples with different nasal states. The data augmentation process in steps (1-3) includes one or any combination of randomized small pitch shifts, time stretching, and adding Gaussian white noise.
[0015] According to another aspect of the present invention, a nasal sound recognition system based on acoustic feature decoupling is provided, comprising the following modules: The first module is used to acquire the real-time singing audio signal to be tested and to perform microphone frequency response compensation on the real-time singing audio signal to obtain a standard digital audio stream. The second module is used to perform constant Q-transform (CQT) on the standard digital audio stream obtained by the first module to obtain the time-frequency tensor. The third module is used to input the time-frequency tensor obtained from the second module into a pre-trained dual-stream decoupled deep neural network to obtain the sound source feature vector and the vocal tract transmission feature vector. The fourth module is used to obtain the individual acoustic feature benchmark of the singer corresponding to the real-time singing audio signal obtained by the first module. The fundamental frequency interference caused by vocal cord vibration in the individual acoustic feature benchmark is removed by the sound source feature vector obtained by the third module, thereby obtaining the processed individual acoustic feature benchmark. The physiological bias correction of the vocal tract transmission feature vector obtained by the third module is performed by the processed individual acoustic feature benchmark to obtain the normalized dynamic nasalization coefficient. The fifth module is used to obtain the final nasal sound recognition result based on the normalized dynamic nasalization coefficient obtained from the fourth module.
[0016] In summary, compared with the prior art, the above-described technical solutions conceived by this invention can achieve the following beneficial effects: (1) This invention adopts the dual-stream decoupled deep neural network designed in step (3) and introduces the mutual information minimization adversarial training mechanism in steps (1-10) to (1-12). By forcibly stripping the sound source information from the vocal tract feature vector, physical-level feature decoupling is achieved. Therefore, it can solve the technical problem of serious "feature coupling" in the existing end-to-end black box recognition method based on deep learning model, which cannot effectively distinguish between physiological and skillful nasal sounds. (2) Since the present invention uses constant Q transformation (CQT) to perform on the audio signal in step (2), it utilizes its high frequency resolution characteristics in the low frequency band to clearly capture the acoustic zero point and anti-resonance point that characterize the nasal sound. Therefore, it can solve the technical problem that the traditional acoustic feature analysis method based on traditional signal processing has insufficient resolution in the low frequency band and is difficult to accurately quantify the degree of nasal sound. (3) Since the present invention adopts the physiological bias correction based on individual acoustic feature benchmark in step (4), by subtracting the user's inherent nasal background noise, it can accurately capture the dynamic nasalization coefficient caused by soft palate movement. Therefore, it can solve the technical problem that the acoustic feature analysis method based on traditional signal processing is easily affected by fundamental frequency harmonic interference when singing high notes, resulting in low accuracy of nasal feature recognition.
[0017] (4) Since the present invention uses the interval quantization of the dynamic nasalization coefficient in step (5) and combines the decoupled vocal tract features in the whole process to provide physical feedback basis directly related to physiological actions (such as soft palate lifting and lowering), it can solve the technical problem that the existing method feedback is abstract and difficult for users to understand and associate with their own physiological actions, resulting in low self-learning efficiency. (5) In this invention, when performing interval quantification of the dynamic nasalization coefficient in step (5), the preset art style label is used to determine the nasal sound (0.5-0.7) in a specific interval as a normal skill performance rather than a vocal defect. Therefore, this invention can solve the technical problem that the existing technology lacks the ability to dynamically adjust the evaluation standard according to different art styles such as bel canto and pop, which easily causes "false correction" of artistic expression. (6) In addition, the present invention constructs training data covering various singing styles, vocal ranges and nasal states in step (1-1), and performs data augmentation processing in step (1-3), so that the model has good robustness and generalization ability and wide applicability; (7) The implementation of this invention is based on mature deep learning framework and signal processing technology. The modular design is clear and easy to deploy and implement on different hardware platforms (such as smartphones and personal computers). The implementation method is simple and efficient. Attached Figure Description
[0018] Figure 1 This is a network architecture diagram of the dual-stream decoupled deep neural network of the present invention; Figure 2 This is a flowchart of the nasal sound recognition method based on acoustic feature decoupling of the present invention. Detailed Implementation
[0019] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention. Furthermore, the technical features involved in the various embodiments of this invention described below can be combined with each other as long as they do not conflict with each other.
[0020] The basic idea of this invention is to construct a two-stream decoupled deep neural network and introduce a mutual information minimization adversarial training strategy to forcibly separate the sound source features (vocal cord vibration) and vocal tract features (oral and nasal cavity modulation) from a single audio signal. Based on this, the high-resolution characteristics of the CQT transform are used to accurately extract nasal-related acoustic features, and physiological bias correction is performed by combining personal acoustic fingerprints. Ultimately, this achieves accurate quantification and identification of different nasal states (normal, defective, skillful), thus providing objective, intuitive, and artistically adaptable feedback for vocal music teaching.
[0021] This embodiment proposes a nasal sound recognition method and system based on acoustic feature decoupling. The physical carrier can be a smartphone, tablet, or personal computer equipped with a microphone array. During the system startup initialization phase, the processor first calls the audio acquisition interface, configuring the sampling rate to 44.1kHz or 48kHz, the quantization precision to 16bit or 24bit, and allocates a dual-ring buffer in memory for real-time storage of the audio data stream. To eliminate interference from frequency response differences in the hardware devices themselves on nasal sound recognition, the system performs a "device fingerprint removal" operation. This involves reading a preset microphone frequency response curve and using a deconvolution filter to perform spectral flattening calibration on the input audio, ensuring that the signal subsequently input to the neural network reflects the user's true acoustic characteristics, rather than the tinted features of the hardware.
[0022] After acquiring the calibrated digital audio signal, the system employs Constant Q Transform (CQT) as the core algorithm for feature extraction. The processor segments the continuous audio stream into overlapping frames, sets the time frame shift to 512 sampling points, and configures the center frequency of the CQT filter to cover the commonly used vocal range from C2 (Greater Oscillator) to C6 (Lesser Oscillator). For each frequency band, the system calculates its logarithmic amplitude spectrum and performs normalization to generate a high-dimensional time-frequency tensor. This tensor, while preserving the time resolution of the high-frequency portion, significantly improves the frequency resolution of the low-frequency portion, thus providing a high-quality data foundation for subsequently distinguishing dense low-frequency harmonics from nasal cavity anti-resonance points.
[0023] The core of this system lies in a pre-trained two-stream decoupled deep neural network. When the CQT feature tensor is input into this network, the data stream is distributed to two parallel encoder branches. The first is the source encoding branch, which employs a multi-layer one-dimensional convolutional neural network (1D-CNN) combined with instance normalization layers to extract the fundamental frequency and harmonic structure related to vocal cord vibration. Its output latent vector contains only spectral information such as pitch, duration, and energy intensity. The second is the vocal tract encoding branch, which focuses on extracting the formant envelope determined by the shape of the oral cavity and the opening and closing of the nasal cavity. To prevent the strong energy of the fundamental frequency from masking subtle nasal features, this branch introduces a spectral attention mechanism, assigning higher weights to the vicinity of the first formant (F1) and the 2kHz to 4kHz frequency band, as these bands contain key acoustic zeros characterizing the degree of nasalization.
[0024] To ensure that the features extracted by the two branches do not interfere with each other, a mutual information minimization strategy and an adversarial training mechanism are introduced during the model training phase. Specifically, the system attempts to infer pitch information from the "vocal tract feature vector" through a discriminator network. If the discriminator cannot accurately determine the pitch, it means that the vocal tract feature vector has successfully separated the sound source information, achieving true physical-level decoupling.
[0025] After obtaining the decoupled vocal tract feature vectors, the system enters a quantitative evaluation phase based on physiological calibration. The system reads samples of closed nasal sounds (e.g., / b / , / d / ) and open nasal sounds (e.g., / m / , / n / ) recorded by the user during registration to construct a baseline for the user's personal acoustic fingerprint. During real-time inference, the system subtracts the user's inherent nasal background noise from the nasalization probability of the current frame, thereby calculating the dynamic nasalization coefficient caused solely by soft palate movement. This processing method ensures that the system can accurately capture technical vocal errors, avoiding misjudgments caused by physiological structural differences such as nasal septum deviation or turbinate hypertrophy.
[0026] like Figure 2 As shown, this invention provides a nasal sound recognition method based on acoustic feature decoupling, comprising the following steps: (1) Acquire the real-time singing audio signal to be tested, and perform microphone frequency response compensation on the real-time singing audio signal to obtain a standard digital audio stream; The advantage of this step (1) is that by eliminating the frequency response difference of the hardware device itself through the deconvolution filter, the signal processed afterward reflects the user's real acoustic characteristics, thus improving the universality and accuracy of the method. (2) Perform a constant Q-transformation (CQT) on the standard digital audio stream obtained in step (1) to obtain the time-frequency tensor; The advantage of this step (2) is that by utilizing the high frequency resolution characteristics of CQT in the low frequency band, it is possible to clearly capture the acoustic zero point and anti-resonance point that characterize the nasal sound, laying a physical foundation for subsequent feature stripping and precise quantization. (3) Input the time-frequency tensor obtained in step (2) into the pre-trained dual-stream decoupled deep neural network to obtain the sound source feature vector and the vocal tract transmission feature vector; (4) Obtain the individual acoustic feature reference of the singer corresponding to the real-time singing audio signal obtained in step (1), and use the sound source feature vector obtained in step (3) to remove the fundamental frequency interference generated by vocal cord vibration in the individual acoustic feature reference, thereby obtaining the processed individual acoustic feature reference, and use the processed individual acoustic feature reference to perform physiological bias correction on the vocal tract transmission feature vector obtained in step (3) to obtain the normalized dynamic nasalization coefficient.
[0027] The advantages of steps (3) and (4) above are: by correcting the individual acoustic reference through the decoupled sound source features, the inherent nasal background noise of the user can be accurately reduced, thereby effectively suppressing the fundamental harmonic interference when singing high notes, and solving the problem of low nasal sound recognition accuracy caused by feature coupling in traditional methods. (5) Obtain the final nasal sound recognition result based on the normalized dynamic nasalization coefficient obtained in step (4); Specifically, the identification results include whether the singer corresponding to the real-time singing audio signal has normal nasal tone, high nasal tone defect, low nasal tone blockage, or stylized nasal tone.
[0028] More specifically, this step determines the vocal state by performing interval quantization of the dynamic nasalization coefficient: If the dynamic nasalization coefficient is between 0.2 and 0.5, it indicates that the singer corresponding to the real-time singing audio signal has a normal nasal tone, indicating that the ratio of oral and nasal resonance is in an acoustically balanced state, which conforms to standard vocal science. If the dynamic nasalization coefficient is between 0.7 and 1.0, it indicates that the singer corresponding to the real-time singing audio signal has a high nasal tone defect, indicating that the soft palate is drooping or not fully closed, resulting in too much airflow entering the nasal cavity and producing a noticeable "noisy" feeling. If the dynamic nasalization coefficient is between 0 and 0.2, it indicates that the singer corresponding to the real-time singing audio signal has low nasal tone blockage, indicating that their nasal passage is not clear (such as a cold or pathological blockage), resulting in the absence of necessary nasal resonance frequencies in the sound signal. If the dynamic nasalization coefficient is between 0.5 and 0.7, it indicates that the singer corresponding to the real-time singing audio signal has stylized nasal technique, which means that he is using a specific singing technique (such as certain specific resonances in pop singing or forward vocalization in folk singing). At this time, the system will combine the preset art style label to determine it as a normal technical performance rather than a vocal defect.
[0029] like Figure 1 As shown, the dual-stream decoupled deep neural network of this invention adopts a 12-layer structure, and its specific model structure is as follows: Layer 1 is a shared feature extraction layer, whose input is a dimension of... The layer performs a convolutional normalization activation operation (Conv+BN+ReLU) on the single-channel CQT time-frequency tensor to obtain a dimension of The feature map is generated and output; the parameters of the convolutional normalized activation operation are: 64 output channels, 3x3 kernel size, stride 1, and padding value 1. This represents the total number of frequency points on the frequency axis of the single-channel CQT time-frequency tensor, which is used to characterize the audio resolution of the single-channel CQT time-frequency tensor. This indicates the number of time frames of the single-channel CQT time-frequency tensor on the time axis, which is used to characterize the duration of the single-channel CQT time-frequency tensor.
[0030] The second layer is the feature distribution layer, whose input is the dimension of the output of the shared feature extraction layer. The feature map is simultaneously output to the first source coding branch layer and the first channel coding branch layer. Layer 3 is the first source coding branch layer. Its input is the feature map output from the feature distribution layer. This layer performs a one-dimensional depthwise separable convolution operation on this feature map to obtain a dimension of... The feature map is generated and output (aimed at locking the fundamental harmonic features along the frequency axis), where the parameters of the one-dimensional depthwise separable convolution operation are: 128 output channels, 5 convolution kernels, and 1 stride. The fourth layer is the second source coding branch layer. Its input is the feature map output by the first source coding branch layer. This layer performs instance normalization on the feature map (to remove global timbre and individual performance style in the audio), and performs feature reconstruction on the normalized feature map to obtain source coding features (which only retain the original texture of vocal cord vibration) and output them. The fifth layer is the first channel coding branch layer. Its input is the feature map output by the feature distribution layer. This layer performs global average pooling and linear mapping on the feature map to obtain the spectral weight coefficients for the first resonance peak (F1 band and 2kHz-4kHz nasal resonance sensitive area). The spectral weight coefficients are then used to perform weighted enhancement processing to obtain the spectrally enhanced channel feature map and output it. Layer 6 is the second channel coding branch layer, and its input is the dimension of the output of the first channel coding branch layer. The layer performs multi-scale convolution processing on the spectrally enhanced vocal tract feature map (with kernels of 3×3, 5×5, and 7×7) to obtain multiple multi-dimensional resonant peak envelope features reflecting the physical morphology of the oral and nasal cavities. All multi-dimensional resonant peak envelope features are then concatenated to obtain a dimensionless array. The vocal tract morphology feature map is generated and output.
[0031] Layer 7 is the mutual information adversarial discrimination layer, whose input is the dimension of the output of the second channel coding branch layer. The layer first performs a three-level fully connected mapping process on the vocal tract morphology feature map (with 256, 128, and 64 neurons respectively) to obtain the probability distribution result for predicting sound source information (such as pitch labels). Then, the layer introduces the real pitch labels corresponding to the standard digital audio stream, calculates the cross-entropy loss between the predicted probability distribution result and the real pitch labels, and outputs the cross-entropy loss value as the mutual information loss value (used to measure the degree to which the vocal tract features contain sound source information in adversarial training).
[0032] The 8th layer is a gradient inversion layer, whose input is the dimension of the output of the second channel coding branch layer. The vocal tract morphology feature map is 128. During the forward propagation phase, this layer performs an identity mapping operation on the vocal tract morphology feature map (i.e., without changing the data values, directly converting the dimension to 128). The feature map is output to the mutual information adversarial discrimination layer); and the gradient back from the mutual information adversarial discrimination layer is inverted during the backpropagation stage (that is, the gradient is multiplied by a negative coefficient) to obtain the inverse gradient and pass it back to the second channel coding branch layer (the purpose is to use adversarial training to force the model to remove the interference information related to the sound source from the channel features). Layer 9 is a global pooling layer, whose inputs are the dimensions of the outputs from the second source coding branch layer. The source coding feature map and the dimension of the output of the second channel coding branch layer are... The layer performs global average pooling on both feature maps to reduce their spatial dimensions. Compress to 1 to obtain and output the sound source feature vector and the vocal tract feature vector, both of which have a dimension of 128.
[0033] The 10th layer is a two-stream independent mapping layer. Its input is the source feature vector and the vocal tract feature vector output by the global pooling layer. This layer performs fully connected linear projection processing on the source feature vector and the vocal tract feature vector respectively (that is, maps the two to a semantic space that does not interfere with each other) to obtain source semantic vector and vocal tract semantic vector with a dimension of 64 and output them respectively. Layer 11 is the dimensionality reduction feature output layer. The input to this layer is the source semantic vector and the vocal tract semantic vector output from the dual-stream independent mapping layer. This layer performs fully connected dimensionality reduction processing on the source semantic vector and the vocal tract semantic vector respectively (to remove redundant information) to obtain compact source feature vectors with a dimension of 32. and vocal tract feature vectors ; The 12th layer is the output interface layer. Its input is the compact sound source feature vector and the vocal tract transmission feature vector output by the dimensionality reduction feature output layer. This layer performs output mapping processing on these two feature vectors, that is, the vocal tract transmission feature vector is output as a physical feature representing the resonance state of the oral cavity and nasal cavity, and the compact sound source feature vector is output as a physical feature representing the vibration state of the vocal cords, thereby obtaining the final acoustic feature decoupling recognition result.
[0034] The dual-stream decoupling deep neural network of the present invention is obtained through the following steps: (1-1) Obtain the original singing audio sample set and divide the original singing audio sample set into a training set and a test set in an 8:2 ratio; Specifically, the original singing audio sample set includes various singing styles (bel canto, pop, folk), multiple vocal ranges (C2 to C6), and various singing audio samples with different nasal sounds.
[0035] The advantage of this step is that by collecting samples covering various singing styles (bel canto, pop, folk), multiple vocal ranges (C2 to C6), and different nasal states, it ensures the diversity of data sources and provides a foundation for the model to learn the common nasal features under different vocal styles.
[0036] (1-2) Perform a constant Q transformation on the training set obtained in step (1-1) to obtain the transformed training set (i.e., the set of time frequency tensors). Specifically, this step involves performing a constant Q-transformation on each singing audio sample in the training set, with the generation dimension being... The single-channel time-frequency tensor; The advantage of this step is that it utilizes the high frequency resolution of CQT in the low-frequency band to clearly capture the acoustic zero point and anti-resonance point that characterize nasal sounds, laying a physical foundation for subsequent feature stripping.
[0037] (1-3) Perform data augmentation on the transformed training set obtained in step (1-2) to obtain the augmented training set; Specifically, data augmentation processing includes one or any combination of random small pitch shifts, time stretching, and adding Gaussian white noise. The advantage of this step is that it expands the distribution of the training set, simulates a real recording environment, and improves the robustness of the model.
[0038] (1-4) For each singing audio sample in the enhanced training set obtained in step (1-3) (its dimension is...) For example, the singing audio sample is input into the shared feature extraction layer of a two-stream decoupled deep neural network and subjected to convolutional normalization activation processing to obtain the corresponding feature sample with dimension [missing information]. Shallow feature map; (1-5) For each singing audio sample in the enhanced training set obtained in step (1-3), the dimension of the singing audio sample corresponding to that sample obtained in step (1-4) is... The shallow feature map is input into the feature distribution layer of the two-stream decoupled deep neural network for distribution processing, so as to obtain two identical feature maps with the same dimension. Feature map; (1-6) For each singing audio sample in the enhanced training set obtained in step (1-3), the feature map corresponding to the singing audio sample obtained in step (1-5) is input into the first source coding branch layer of the dual-stream decoupled deep neural network for one-dimensional depthwise separable convolution processing to obtain the feature map corresponding to the singing audio sample with dimension [missing information]. The intermediate feature map of the sound source; (1-7) For each singing audio sample in the enhanced training set obtained in step (1-3), the dimension of the singing audio sample corresponding to that sample obtained in step (1-6) is... The intermediate feature map of the sound source is input into the second sound source coding branch layer of the dual-stream decoupled deep neural network for instance normalization and feature reconstruction to obtain the dimension corresponding to the singing audio sample. The source coding feature map; (1-8) For each singing audio sample in the training set after enhancement processing obtained in step (1-3), the other feature map corresponding to the singing audio sample obtained in step (1-5) is input into the first channel coding branch layer of the dual-stream decoupled deep neural network for prediction value calculation and weighted enhancement processing, so as to obtain the corresponding feature map of the singing audio sample with dimension . Spectral enhancement acoustic spectral characteristics; (1-9) For each singing audio sample in the enhanced training set obtained in step (1-3), the dimension of the singing audio sample corresponding to that sample obtained in step (1-8) is... The spectral enhancement tract feature map is input into the second tract coding branch layer of the dual-stream decoupled deep neural network and subjected to multi-scale convolution and concatenation processing to obtain the dimensional data corresponding to the singing audio sample. A diagram showing the morphological characteristics of the vocal tract; The advantage of steps (1-4) to (1-9) above is that, through hierarchical extraction and a dual-stream parallel architecture, the fundamental frequency texture of vocal cord vibration and the resonance envelope of the oral cavity and nasal cavity are locked respectively, thus achieving preliminary separation of the physical dimension in the feature extraction stage.
[0039] (1-10) For each singing audio sample in the enhanced training set obtained in step (1-3), the dimension of the singing audio sample corresponding to that sample obtained in step (1-9) is... The vocal tract morphology feature map is input into the mutual information adversarial discrimination layer of the dual-stream decoupled deep neural network and processed by a three-level fully connected mapping to obtain the prediction probability distribution result corresponding to the singing audio sample, which is used to predict the sound source information (pitch label). (1-11) For each singing audio sample in the training set after enhancement processing obtained in step (1-3), calculate the cross-entropy loss value based on the predicted probability distribution result corresponding to the singing audio sample obtained in step (1-10) and the real pitch label corresponding to the singing audio sample, and use it as the mutual information loss corresponding to the singing audio sample. ; (1-12) For each singing audio sample in the training set after enhancement processing obtained in step (1-3), the mutual information loss corresponding to the singing audio sample obtained in step (1-11) is updated by a reverse gradient (multiplied by a negative coefficient) using a gradient reversal layer and gradient descent method to obtain the reverse gradient, and the weight of the sound source coding branch is obtained using the reverse gradient. The advantage of steps (1-10) to (1-12) is that by using the "minimal-maximal" adversarial game mechanism, the model is forced to strip away the residual pitch information in the vocal tract features, thereby achieving the minimization of mutual information at the mathematical level and solving the problem of high-frequency harmonic interference in nasal sound recognition.
[0040] (1-13) For each singing audio sample in the enhanced training set obtained in step (1-3), the source coding feature map corresponding to the singing audio sample obtained in step (1-7) and the vocal tract morphology feature map corresponding to the singing audio sample obtained in step (1-9) are respectively input into the 9th layer (global pooling layer) of the dual-stream decoupled deep neural network for global average pooling processing to obtain the two-dimensional feature map corresponding to the singing audio sample. A fixed-length eigenvector; (1-14) For each singing audio sample in the enhanced training set obtained in step (1-3), the two fixed-length feature vectors obtained in step (1-13) are input into the 10th layer (two-stream independent mapping layer) of the two-stream decoupled deep neural network for independent fully connected linear projection processing to obtain the two-dimensional feature vectors corresponding to the singing audio sample. semantic vector; (1-15) For each singing audio sample in the training set after enhancement processing obtained in step (1-3), the two semantic vectors corresponding to the singing audio sample obtained in step (1-14) are input into the 11th layer (dimensionality reduction feature output layer) for dimensionality reduction processing to obtain the final dimension of the singing audio sample. Compact sound source feature vector and channel transmission feature vector; (1-16) Repeat steps (1-4) to (1-15) until the mutual information loss obtained in step (1-11) converges or reaches the preset number of iterations (the value ranges from 5,000 to 20,000 times, preferably 10,000 times) to obtain a pre-trained two-stream decoupled deep neural network. (1-17) Use the test set obtained in step (1-1) to perform fine-tuning tests and physiological benchmark calibration on the initially trained dual-stream decoupled deep neural network obtained in step (1-16) until the cumulative error of the dynamic quantization coefficient of the dual-stream decoupled deep neural network is lower than the preset threshold (the value range is 0.01 to 0.05, preferably 0.02) to obtain the finally trained dual-stream decoupled deep neural network.
[0041] Those skilled in the art will readily understand that the above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A nasal sound recognition method based on acoustic feature decoupling, characterized in that, Includes the following steps: (1) Acquire the real-time singing audio signal to be tested, and perform microphone frequency response compensation on the real-time singing audio signal to obtain a standard digital audio stream; (2) Perform constant Q-transform (CQT) on the standard digital audio stream obtained in step (1) to obtain the time-frequency tensor; (3) Input the time-frequency tensor obtained in step (2) into the pre-trained dual-stream decoupled deep neural network to obtain the sound source feature vector and the vocal tract transmission feature vector; (4) Obtain the individual acoustic feature reference of the singer corresponding to the real-time singing audio signal obtained in step (1), and use the sound source feature vector obtained in step (3) to remove the fundamental frequency interference generated by vocal cord vibration in the individual acoustic feature reference, thereby obtaining the processed individual acoustic feature reference, and use the processed individual acoustic feature reference to perform physiological bias correction on the vocal tract transmission feature vector obtained in step (3) to obtain the normalized dynamic nasalization coefficient. (5) Obtain the final nasal sound recognition result based on the normalized dynamic nasalization coefficient obtained in step (4).
2. The nasal sound recognition method based on acoustic feature decoupling according to claim 1, characterized in that, The specific identification result is whether the singer corresponding to the real-time singing audio signal has normal nasal tone, high nasal tone defect, low nasal tone blockage, or stylized nasal tone.
3. The nasal sound recognition method based on acoustic feature decoupling according to claim 1 or 2, characterized in that, If the dynamic nasalization coefficient is between 0.2 and 0.5, it indicates that the singer corresponding to the real-time singing audio signal has a normal nasal tone, indicating that the ratio of oral and nasal resonance is in an acoustically balanced state, which conforms to standard vocal science. If the dynamic nasalization coefficient is between 0.7 and 1.0, it indicates that the singer corresponding to the real-time singing audio signal has a high nasal tone defect, indicating that the soft palate is drooping or not fully closed, resulting in too much airflow entering the nasal cavity and producing a noticeable "noisy" feeling. If the dynamic nasalization coefficient is between 0 and 0.2, it indicates that the singer corresponding to the real-time singing audio signal has low nasal tone blockage, indicating that their nasal passage is not smooth, resulting in the absence of necessary nasal resonance frequencies in the sound signal. If the dynamic nasalization coefficient is between 0.5 and 0.7, it indicates that the singer corresponding to the real-time singing audio signal has stylized nasal technique, indicating that they are using a specific singing technique.
4. The nasal sound recognition method based on acoustic feature decoupling according to any one of claims 1 to 3, characterized in that, The two-stream decoupled deep neural network uses a 12-layer structure, in which: Layer 1 is a shared feature extraction layer, whose input is a dimension of... The layer performs a convolutional normalization activation operation on the single-channel CQT time-frequency tensor to obtain a dimension of The feature map is generated and output; the parameters of the convolutional normalized activation operation are: 64 output channels, 3x3 kernel size, stride 1, and padding value 1. This represents the total number of frequency points on the frequency axis of the single-channel CQT time-frequency tensor, which is used to characterize the audio resolution of the single-channel CQT time-frequency tensor. This indicates the number of time frames of the single-channel CQT time-frequency tensor on the time axis, which is used to characterize the duration of the single-channel CQT time-frequency tensor. The second layer is the feature distribution layer, whose input is the dimension of the output of the shared feature extraction layer. The feature map is simultaneously output to the first source coding branch layer and the first channel coding branch layer. Layer 3 is the first source coding branch layer. Its input is the feature map output from the feature distribution layer. This layer performs a one-dimensional depthwise separable convolution operation on this feature map to obtain a dimension of... The feature map is generated and output, where the parameters of the one-dimensional depthwise separable convolution operation are: 128 output channels, 5 convolution kernels, and 1 stride. The fourth layer is the second source coding branch layer. Its input is the feature map output by the first source coding branch layer. This layer performs instance normalization on the feature map and performs feature reconstruction on the normalized feature map to obtain source coding features and output them.
5. The nasal sound recognition method based on acoustic feature decoupling according to claim 4, characterized in that, The fifth layer is the first channel coding branch layer. Its input is the feature map output by the feature distribution layer. This layer performs global average pooling and linear mapping on the feature map to obtain the spectral weight coefficients for the first resonance peak (F1 band and 2kHz-4kHz nasal resonance sensitive area). The spectral weight coefficients are then used to perform weighted enhancement processing to obtain the spectrally enhanced channel feature map and output it. Layer 6 is the second channel coding branch layer, and its input is the dimension of the output of the first channel coding branch layer. The layer performs multi-scale convolution on the spectral-enhanced vocal tract feature map to obtain multiple multi-dimensional resonant peak envelope features reflecting the physical morphology of the oral and nasal cavities. All multi-dimensional resonant peak envelope features are then concatenated to obtain a dimension-... The vocal tract morphology feature map is generated and output. Layer 7 is the mutual information adversarial discrimination layer, whose input is the dimension of the output of the second channel coding branch layer. The layer first performs a three-level fully connected mapping process on the vocal tract morphology feature map to obtain the probability distribution result used to predict sound source information. Then, the layer introduces the real pitch label corresponding to the standard digital audio stream, calculates the cross-entropy loss between the predicted probability distribution result and the real pitch label, and outputs the cross-entropy loss value as the mutual information loss value. The 8th layer is a gradient inversion layer, whose input is the dimension of the output of the second channel coding branch layer. A diagram showing the morphological characteristics of the vocal tract of a 128-channel tract; During the forward propagation phase, this layer performs an identity mapping operation on the morphological feature map of the vocal tract; and during the backward propagation phase, it performs an inversion process on the gradient returned by the mutual information adversarial discrimination layer to obtain the inverse gradient and propagates it back to the second vocal tract coding branch layer.
6. The nasal sound recognition method based on acoustic feature decoupling according to claim 5, characterized in that, Layer 9 is a global pooling layer, whose inputs are the dimensions of the outputs from the second source coding branch layer. The source coding feature map and the dimension of the output of the second channel coding branch layer are... The layer performs global average pooling on the two feature maps, respectively, to reduce their spatial dimensions. Compress to 1 to obtain and output sound source feature vectors and vocal tract feature vectors, each with a dimension of 128. The 10th layer is a two-stream independent mapping layer. Its input is the source feature vector and the vocal tract feature vector output by the global pooling layer. This layer performs fully connected linear projection processing on the source feature vector and the vocal tract feature vector respectively to obtain source semantic vector and vocal tract semantic vector with a dimension of 64 and output them respectively. Layer 11 is the dimensionality reduction feature output layer. The input to this layer is the source semantic vector and the vocal tract semantic vector output from the dual-stream independent mapping layer. This layer performs fully connected dimensionality reduction processing on the source semantic vector and the vocal tract semantic vector respectively to obtain compact source feature vectors with a dimension of 32. and vocal tract feature vectors ; The 12th layer is the output interface layer. Its input is the compact sound source feature vector and the vocal tract transmission feature vector output by the dimensionality reduction feature output layer. This layer performs output mapping processing on these two feature vectors, that is, the vocal tract transmission feature vector is output as a physical feature representing the resonance state of the oral cavity and nasal cavity, and the compact sound source feature vector is output as a physical feature representing the vibration state of the vocal cords, thereby obtaining the final acoustic feature decoupling recognition result.
7. The nasal sound recognition method based on acoustic feature decoupling according to claim 6, characterized in that, The two-stream decoupling deep neural network is trained through the following steps: (1-1) Obtain the original singing audio sample set and divide the original singing audio sample set into a training set and a test set in an 8:2 ratio; (1-2) Perform a constant Q transformation on the training set obtained in step (1-1) to obtain the transformed training set; (1-3) Perform data augmentation on the transformed training set obtained in step (1-2) to obtain the augmented training set; (1-4) For each dimension of the augmented training set obtained in step (1-3), the following is... For the singing audio sample, the singing audio sample is input into the shared feature extraction layer of a two-stream decoupled deep neural network and subjected to convolutional normalization activation processing to obtain the corresponding feature sample with dimension [missing information]. Shallow feature map; (1-5) For each singing audio sample in the enhanced training set obtained in step (1-3), the dimension of the singing audio sample corresponding to that sample obtained in step (1-4) is... The shallow feature map is input into the feature distribution layer of the two-stream decoupled deep neural network for distribution processing, so as to obtain two identical feature maps with the same dimension. Feature map; (1-6) For each singing audio sample in the enhanced training set obtained in step (1-3), the feature map corresponding to the singing audio sample obtained in step (1-5) is input into the first source coding branch layer of the dual-stream decoupled deep neural network for one-dimensional depthwise separable convolution processing to obtain the feature map corresponding to the singing audio sample with dimension [missing information]. The intermediate feature map of the sound source; (1-7) For each singing audio sample in the enhanced training set obtained in step (1-3), the dimension of the singing audio sample corresponding to that sample obtained in step (1-6) is... The intermediate feature map of the sound source is input into the second sound source coding branch layer of the dual-stream decoupled deep neural network for instance normalization and feature reconstruction to obtain the dimension corresponding to the singing audio sample. The source coding feature map; (1-8) For each singing audio sample in the training set after enhancement processing obtained in step (1-3), the other feature map corresponding to the singing audio sample obtained in step (1-5) is input into the first channel coding branch layer of the dual-stream decoupled deep neural network for prediction value calculation and weighted enhancement processing, so as to obtain the corresponding feature map of the singing audio sample with dimension . Spectral enhancement acoustic spectral characteristics; (1-9) For each singing audio sample in the enhanced training set obtained in step (1-3), the dimension of the singing audio sample corresponding to that sample obtained in step (1-8) is... The spectral enhancement tract feature map is input into the second tract coding branch layer of the dual-stream decoupled deep neural network and subjected to multi-scale convolution and concatenation processing to obtain the dimensional data corresponding to the singing audio sample. A diagram showing the morphological characteristics of the vocal tract; (1-10) For each singing audio sample in the enhanced training set obtained in step (1-3), the dimension of the singing audio sample corresponding to that sample obtained in step (1-9) is... The vocal tract morphology feature map is input into the mutual information adversarial discrimination layer of the dual-stream decoupled deep neural network and processed by a three-level fully connected mapping process to obtain the prediction probability distribution result corresponding to the singing audio sample, which is used to predict the sound source information. (1-11) For each singing audio sample in the training set after enhancement processing obtained in step (1-3), calculate the cross-entropy loss value based on the predicted probability distribution result corresponding to the singing audio sample obtained in step (1-10) and the real pitch label corresponding to the singing audio sample, and use it as the mutual information loss corresponding to the singing audio sample. ; (1-12) For each singing audio sample in the training set after enhancement processing obtained in step (1-3), the mutual information loss corresponding to the singing audio sample obtained in step (1-11) is updated in reverse using the gradient reversal layer and gradient descent method to obtain the reverse gradient, and the weight of the sound source coding branch is obtained using the reverse gradient. (1-13) For each singing audio sample in the enhanced training set obtained in step (1-3), the source coding feature map corresponding to the singing audio sample obtained in step (1-7) and the vocal tract morphology feature map corresponding to the singing audio sample obtained in step (1-9) are respectively input into the global pooling layer of the dual-stream decoupled deep neural network for global average pooling processing to obtain the two-dimensional feature map corresponding to the singing audio sample. A fixed-length eigenvector; (1-14) For each singing audio sample in the training set after enhancement processing obtained in step (1-3), the two fixed-length feature vectors obtained in step (1-13) are input into the independent mapping layer of the two-stream decoupled deep neural network for independent fully connected linear projection processing to obtain the two-dimensional feature vectors corresponding to the singing audio sample. semantic vector; (1-15) For each singing audio sample in the training set after enhancement processing obtained in step (1-3), the two semantic vectors corresponding to the singing audio sample obtained in step (1-14) are input into the dimensionality reduction feature output layer of the dual-stream decoupled deep neural network for dimensionality reduction processing, so as to obtain the final dimension of the singing audio sample. Compact sound source feature vector and channel transmission feature vector; (1-16) Repeat steps (1-4) to (1-15) until the mutual information loss obtained in step (1-11) converges or reaches the preset number of iterations, so as to obtain a pre-trained two-stream decoupled deep neural network. (1-17) Use the test set obtained in step (1-1) to fine-tune and calibrate the preliminarily trained dual-stream decoupled deep neural network obtained in step (1-16) until the cumulative error of the dynamic quantization coefficient of the dual-stream decoupled deep neural network is lower than the preset threshold, so as to obtain the final trained dual-stream decoupled deep neural network.
8. The nasal sound recognition method based on acoustic feature decoupling according to claim 7, characterized in that, The original singing audio sample set in step (1-1) includes bel canto, pop, folk singing, as well as the vocal range from C2 to C6 in the three octaves, and various singing audio samples with different nasal sounds. The data augmentation process in steps (1-3) includes one or any combination of randomized small pitch shifts, time stretching, and adding Gaussian white noise.
9. A nasal sound recognition system based on acoustic feature decoupling, characterized in that, Includes the following modules: The first module is used to acquire the real-time singing audio signal to be tested and to perform microphone frequency response compensation on the real-time singing audio signal to obtain a standard digital audio stream. The second module is used to perform constant Q-transform (CQT) on the standard digital audio stream obtained by the first module to obtain the time-frequency tensor. The third module is used to input the time-frequency tensor obtained from the second module into a pre-trained dual-stream decoupled deep neural network to obtain the sound source feature vector and the vocal tract transmission feature vector. The fourth module is used to obtain the individual acoustic feature benchmark of the singer corresponding to the real-time singing audio signal obtained by the first module. The fundamental frequency interference caused by vocal cord vibration in the individual acoustic feature benchmark is removed by the sound source feature vector obtained by the third module, thereby obtaining the processed individual acoustic feature benchmark. The physiological bias correction of the vocal tract transmission feature vector obtained by the third module is performed by the processed individual acoustic feature benchmark to obtain the normalized dynamic nasalization coefficient. The fifth module is used to obtain the final nasal sound recognition result based on the normalized dynamic nasalization coefficient obtained from the fourth module.