Method, system, medium and program product for recognizing continuous voice segments of human voice
Through variational modal decomposition and gamma pass frequency cepspectral coefficient filter combined with neural network, the continuous voice segment of children's voice is identified, solving the problem of children's voice detection in complex environments and improving the performance of speech recognition and evaluation systems.
Patent Information
- Application Number
- CN202510803825.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-17
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-06-17
AI Technical Summary
In the context of complex environments, it is difficult for the prior art to effectively detect continuous voice segments of children's voices, resulting in a decrease in the efficiency of speech recognition and voice evaluation systems. Especially when parents accompany their children to study, the existence of parental voice leads to low detection efficiency.
Variable mode decomposition and gamma pass frequency cepspectral coefficient filter are used to extract specific features related to children's voices, combined with speech category classification neural network, and continuously speech segments are identified through Markov state transfer relationships to enhance the distinction between children's voices and background noise.
It improves the ability of children's voice detection in the context of complex noise, ensures the stability and efficiency of the speech recognition and evaluation system, and can accurately identify the continuous voice segments of children's voices.
Smart Images

Figure CN120356475A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of voice detection, and more specifically, to a method, system, non-transitory storage medium, and computer program product for identifying continuous voice segments of human voices. Background Art
[0002] Nowadays, with the development of artificial intelligence (AI) technology, the combination of AI technology and education has become increasingly close. Whether it is pre-class preview, in-class learning, or after-class homework, audio technologies such as speech recognition and oral evaluation have been fully practiced. However, in specific environments, especially in environments with complex background noise, the efficiency of subsequent tasks for children's human voices, such as speech recognition, voiceprint recognition, and speech evaluation, will drop significantly, seriously affecting the product and overall efficiency. At the same time, in actual usage scenarios, parents often accompany their children to practice together, and the voices of parents also appear in the recordings, resulting in low efficiency in detecting continuous voice segments of children's human voices from these recordings. Summary of the Invention
[0003] To solve the problem of detecting children's human voice in a complex environmental background, this proposal mainly presents a voice detection solution specifically for detecting children's human voices, solves the problem of detecting continuous voice segments of children's human voices, and effectively improves the operating efficiency of downstream audio processing.
[0004] According to one aspect of the present application, there is provided a method for identifying continuous voice segments of human voices, including: extracting a plurality of to-be-identified features regarding the simulated basilar membrane induction information of human voices from a plurality of temporally continuous to-be-identified audio frames through variational mode decomposition and gamma-tone frequency cepstral coefficient filters; inputting the extracted plurality of to-be-identified features into a voice category classification neural network to determine a plurality of posterior probabilities of having human voices in the plurality of to-be-identified audio frames from the plurality of to-be-identified audio frames; and identifying one or more continuous voice segments of human voices according to the plurality of posterior probabilities of having human voices determined in the plurality of to-be-identified audio frames.
[0005] According to one aspect of the present application, there is provided a device for identifying continuous voice segments of human voices, including: an extraction device configured to extract a plurality of to-be-identified features regarding the simulated basilar membrane induction information of human voices from a plurality of temporally continuous to-be-identified audio frames through variational mode decomposition and gamma-tone frequency cepstral coefficient filters; a determination device configured to input the extracted plurality of to-be-identified features into a voice category classification neural network to determine a plurality of posterior probabilities of having human voices in the plurality of to-be-identified audio frames from the plurality of to-be-identified audio frames; and an identification device configured to identify one or more continuous voice segments of human voices according to the plurality of posterior probabilities of having human voices determined in the plurality of to-be-identified audio frames.
[0006] According to another aspect of the present application, there is provided a device for identifying continuous speech segments of human voices, including: a memory for storing instructions; and a processor for executing the instructions in the memory and performing the method according to at least one embodiment of the present disclosure.
[0007] According to another aspect of the present application, there is provided a non-transitory storage medium storing instructions thereon, wherein when the instructions are executed by a processor, the processor is caused to perform the method according to at least one embodiment of the present disclosure.
[0008] According to another aspect of the present application, there is provided a computer program product including computer instructions, wherein when the instructions are executed by a processor, the processor is caused to perform the method according to at least one embodiment of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present disclosure, and for those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0010] Figure 1 The flowchart of a method for identifying continuous speech segments of human voices according to at least one embodiment of the present disclosure is shown.
[0011] Figure 2 The flowchart of steps for extracting features of simulated human cochlear basilar membrane induction information about human voices from training audio frames through variational mode decomposition and gamma-tone frequency cepstral coefficient filters according to at least one embodiment of the present disclosure is shown.
[0012] Figure 3 The process of generating a training set of training audio files with noise and reverberation according to at least one embodiment of the present disclosure is shown.
[0013] Figure 4 The schematic diagram of the start and end points of a continuous speech segment according to at least one embodiment of the present disclosure is shown.
[0014] Figure 5 The schematic diagram of the start and end points of two continuous speech segments according to at least one embodiment of the present disclosure is shown.
[0015] Figure 6 The Markov state transition relationship diagram of each speech state according to at least one embodiment of the present disclosure is shown.
[0016] Figure 7A schematic diagram showing the training process of a voice transition segment classification neural network according to at least one embodiment of the present disclosure.
[0017] Figure 8 An example diagram showing the left 12 frames and the right 12 frames at a specific time point being temporarily stored through a delay memory slot with a length L2 of 24 frames according to at least one embodiment of the present disclosure.
[0018] Figure 9 A block diagram of a device for identifying a continuous speech segment of a human voice according to at least one embodiment of the present disclosure.
[0019] Figure 10 A block diagram of an exemplary device for identifying a continuous speech segment of a human voice suitable for implementing at least one embodiment of the present disclosure.
[0020] Figure 11 A schematic diagram of an application scenario according to at least one embodiment of the present disclosure. Detailed implementation manners
[0021] Now, specific embodiments of the present application will be described in detail. Examples of the present application are illustrated in the accompanying drawings. Although the present application will be described in conjunction with specific embodiments, it will be understood that it is not intended to limit the present application to the described embodiments. On the contrary, it is intended to cover modifications, variations, and equivalents included within the spirit and scope of the present application as defined by the appended claims. It should be noted that the method steps described herein can all be implemented by any functional block or functional arrangement, and any functional block or functional arrangement can be implemented as a physical entity or a logical entity, or a combination of both.
[0022] Figure 1 A flowchart of a method 100 for identifying a continuous speech segment of a human voice according to at least one embodiment of the present disclosure.
[0023] In the actual usage scenario of identifying a human voice, the environmental background may be relatively complex. Therefore, the recorded sounds collected may include the child's own voice, adult voices, environmental noises, etc., resulting in a low efficiency in accurately identifying the continuous speech segment of the child's human voice from these recordings. Figure 11 A schematic diagram of an application scenario according to at least one embodiment of the present disclosure.
[0024] To solve the problem of continuous speech segment detection of children's voices (or children's human voices, which are also a type of human voice) in complex environmental backgrounds, in the prior art, it mainly relies on methods based on signal-to-noise ratio threshold judgment using energy and methods based on basic feature training and neural network training. Among them, the method based on signal-to-noise ratio threshold judgment using energy calculates the ratio of speech energy to noise energy (including other silences, etc.), and then uses the ratio of likelihood functions and refers to the threshold for judgment; for the method using feature training, it mainly relies on basic speech features such as MFCC (Mel Frequency Cepstral Coefficient), PLP (Perceptual linear predictive), and FBANK (FilterBank), and is trained through neural networks such as FNN (Feedforward Neural Network), RNN (Recurrent Neural Network), CNN (Convolutional Neural Network), and LSTM (Long Short-Term Memory Network), and judges whether there is speech in each frame of the speech stream through hard decision-making.
[0025] However, the main drawbacks of these methods in the prior art are as follows: The energy method based on signal-to-noise ratio threshold judgment using energy has very poor anti-noise ability. If the noise energy exceeds the energy of the training audio frame, it is impossible to correctly distinguish whether it is speech or noise; for the detection system based on neural networks and features such as MFCC / PLP / FBANK, due to the poor anti-noise ability of the features in this application, it also has the drawback of being unable to correctly distinguish whether it is speech or noise.
[0026] The main strategies of this solution include considering the characteristics of children's voices, using variational mode decomposition and Gammatone (GT) frequency cepstral coefficient filters to extract specific features related to children's voices, enhancing the distinguishability between children's voices and background noise, and accurately detecting children's voices. The background noise can include adult voices, as well as environmental noises such as indoor reverberation, echo, and single-frequency noise. This solution also uses the Markov state transition relationship of speech states to determine the continuous speech segments of children's voices from before speech to during speech to the end of speech in terms of time, which also helps to further verify whether the previously identified children's voices are correct or eliminate overly short continuous speech segments, further ensuring the accuracy of the recognition results. In this way, it is possible to improve the detection ability of continuous speech segments of children's voices in complex noise backgrounds and ensure the stability of the recognition performance and evaluation performance of speech recognition and speech evaluation systems in educational scenarios, for example.
[0027] As Figure 1 shown, the method 100 for identifying continuous speech segments of human voices includes step 110, step 120, and step 130.
[0028] In step 110, multiple features to be recognized regarding the simulated cochlear basilar membrane induction information of human voices are extracted from multiple consecutive audio frames to be recognized through Variational Mode Decomposition (VMD) and Gammatone Filter Cepstral Coefficient (GFCC) filters.
[0029] In step 120, the multiple features to be recognized are input into a speech category classification neural network to determine multiple posterior probabilities of having human voices in the multiple audio frames to be recognized from the multiple audio frames to be recognized.
[0030] In step 130, one or more continuous speech segments of human voices are recognized according to the multiple posterior probabilities of having human voices determined in the multiple audio frames to be recognized.
[0031] In this way, by creatively considering the characteristics of children's human voices, Variational Mode Decomposition and Gammatone Filter Cepstral Coefficient filters are used to extract specific features related to children's human voices, enhancing the distinguishability between children's human voices and background noise and enabling more accurate detection of children's human voices.
[0032] In step 110, each of the multiple audio frames to be recognized of human voices can be decomposed into multiple intrinsic mode function components through Variational Mode Decomposition; the features to be recognized of the simulated cochlear basilar membrane induction information are extracted from each of the multiple intrinsic mode function components through a Gammatone Filter Cepstral Coefficient filter.
[0033] Since after an external speech signal enters the cochlear basilar membrane, it will be decomposed according to frequency and generate traveling wave vibrations, thus stimulating auditory receptor cells. And the Gammatone filter is a set of filter models used to simulate the frequency decomposition characteristics of the cochlea. It can be used for the decomposition of audio signals to facilitate the extraction of the features of the simulated cochlear basilar membrane induction information. The Gammatone filter bank has similarities with the cochlear basilar membrane in terms of impulse response, amplitude-frequency characteristics, etc. Therefore, a set of Gammatone filter banks with logarithmically evenly distributed center frequencies can be used to simulate the cochlear basilar membrane.
[0034] Decomposing each audio frame to be recognized into multiple intrinsic mode function components through variational mode decomposition may include: framing and preprocessing each audio frame to be recognized; calculating the power spectral density of the framed and preprocessed audio frames to be recognized; decomposing each audio frame to be recognized into multiple intrinsic mode function components based on the power spectral density; wherein, extracting the characteristics of the simulated human cochlear basilar membrane induction information for each intrinsic mode function component among the multiple intrinsic mode function components includes: based on each intrinsic mode function component among the multiple intrinsic mode function components, using a gamma-tone frequency cepstral coefficient filter to extract the characteristics as the characteristics of the simulated human cochlear basilar membrane induction information regarding human voices. Later, it will also be combined with Figure 2 to describe specific examples of extracting multiple features to be recognized of the simulated human cochlear basilar membrane induction information regarding human voices from multiple temporally continuous audio frames to be recognized through variational mode decomposition and gamma-tone frequency cepstral coefficient filters.
[0035] The multiple features to be recognized of the simulated human cochlear basilar membrane induction information regarding human voices extracted in this way are particularly suitable for detecting children's voices because the dispersion of each intrinsic mode function component of adult voices is more obvious, and the energy regions of the component frequency bands are more dispersed. The energy of each intrinsic mode function component of children's voices is more concentrated. Therefore, through each intrinsic mode function component and the corresponding feature extraction, it is more able to significantly distinguish children's voices, thus achieving an accurate classification effect in the subsequent classification of the neural network.
[0036] In step 120, the extracted multiple features to be recognized are input into a speech category classification neural network to determine the multiple posterior probabilities of having human voices among the multiple audio frames to be recognized from the multiple audio frames to be recognized. The speech category classification neural network can be trained through the following steps: extracting the characteristics of the simulated human cochlear basilar membrane induction information regarding human voices from the training audio frames through variational mode decomposition and gamma-tone frequency cepstral coefficient filters; training the speech category classification neural network with the extracted characteristics and the labels corresponding to the training audio frames.
[0037] Here, the labels corresponding to the training audio frames can be determined in the following way.
[0038] Perform label definition, including multiple categories such as children's voices (i.e., the predetermined voice is children's voice), non-children's voices (such as adult voices), other (non-human voice) noises, silence, etc. (determined according to the actual situation).
[0039] In at least one embodiment of the present disclosure, a child is defined as a person within a certain age range, such as 3 - 6 years old, but not limited thereto. Since the voices of infants under 3 years old are not considered in at least one embodiment of the present disclosure, and for the consideration of the language learning scenario of foreign languages, non - child human voices are regarded as adult human voices to distinguish them from child human voices. However, by using a similar method described in at least one embodiment of the present disclosure and changing the parameters of the features associated with the dispersion of the component frequency - band energy regions, etc., it is actually possible to identify human voices such as infant voices and adult human voices. For the purpose clearly described in at least one embodiment of the present disclosure, the specific implementation process is described by taking the detection of child human voices as an example of such human voices in at least one embodiment of the present disclosure.
[0040] Generate labels corresponding to the training audio frames: Add at least one of one or more human voices, noises, and reverberations to the speech sample data to generate speech with added noise and reverberation. Among them, the speech category labels at least include predetermined human voices, non - predetermined human voices, non - human noises, and silences. Specifically, use child human voices, adult human voices, non - human noises, reverberations, etc., and mix them in proportion to obtain various training audio frames, so that whether the various training audio frames are adult human voices or child human voices, they all carry other human voices, non - human noises, reverberations, etc., and set labels that conform to the above - mentioned label definitions for the various training audio frames after mixing.
[0041] Next, describe the specific process of extracting features of the simulated basilar membrane induction information of human voices from the training audio frames through variational mode decomposition and gamma - tone frequency cepstral coefficient filters.
[0042] First, extracting features of the simulated basilar membrane induction information of human voices from the training audio frames through variational mode decomposition and gamma - tone frequency cepstral coefficient filters may include: decomposing the training audio frames into multiple intrinsic mode function components through variational mode decomposition; using gamma - tone frequency cepstral coefficient filters to extract features of the simulated basilar membrane induction information for each of the multiple intrinsic mode function components.
[0043] Decomposing the training audio frames into multiple intrinsic mode function components through variational mode decomposition may include: framing and pre - processing the training speech audio; calculating the power spectral density of the framed and pre - processed training audio frames; decomposing the training audio frames into multiple intrinsic mode function components based on the power spectral density.
[0044] Specifically, Figure 2 Shows a flowchart of the steps of extracting features of the simulated basilar membrane induction information of human voices from the training audio frames through variational mode decomposition and gamma - tone frequency cepstral coefficient filters according to at least one embodiment of the present disclosure.
[0045] First, the training audio is framed and preprocessed. In step 201, the speech waveform signal (for example, various training audios are obtained by mixing children's voices, adult voices, non-human voices, reverberations, etc. in proportion as before) is framed for the first time. Specifically, the frame length of each frame can be set to 360 ms (expressed in time), and each time it is moved by 360 ms to obtain multiple training audio frames with a frame length of 360 ms for each frame. This setting is mainly obtained through experiments for this special transformation of variational mode decomposition. If it is too long or too short, a good separation effect cannot be obtained, and there is no overlap between each frame. It can correspond to 2880 sampling points, and the center frequency of each frame can be selected according to the position of the maximum energy in the energy spectrum. Framing is to process a long training audio frame in segments. However, if the training audio frame is not long, framing can also be considered not necessary. Therefore, step 201 is optional.
[0046] In step 202, speech pre-emphasis is performed on the single-frame speech waveform of each training audio frame.
[0047] Since speech needs to be transmitted through a medium such as air, energy will be consumed. The higher the frequency of the training audio frame, the more serious the loss of speech energy by the medium. Therefore, pre-emphasis is used to enhance the high-frequency components of the training audio frame to compensate for the attenuation of the high-frequency components during transmission. Of course, this step is optional.
[0048] In step 203, the first windowing is performed on the pre-emphasized training audio frame. The window function used in at least one embodiment of the present disclosure can be the Hamming window, and the Hamming window function is determined as follows.
[0049]
[0050]
[0051] In the window function, t represents time, and M represents the window length, which generally takes the frame length of the framing. The frame length is calculated according to the frame length in the above formula. Formula calculation.
[0052] In step 204, the windowed speech waveform is obtained through the Hamming window function, that is, the training audio frame in at least one embodiment of the present disclosure.
[0053] After that, variational mode decomposition is to be performed on the windowed speech waveform f that has been framed and preprocessed, that is, the training audio frame f in at least one embodiment of the present disclosure.
[0054] The principle of variational mode decomposition lies in the assumption that a signal is composed of the superposition of sub-signals dominated by different frequencies. Therefore, its purpose is to decompose the signal into sub-signals of different frequencies (referred to as finite bandwidth intrinsic mode function (IMF) components or simply mode components).
[0055] The total number K of mode components and the penalty coefficient α for variational mode decomposition can be set.
[0056] In one embodiment, the total number K of mode components and the penalty coefficient α can be set based on the dispersion difference and energy concentration of the component frequency band energy regions of non-child voices and child voices.
[0057] Specifically, the setting of the K value can be based on the following observations. In at least one embodiment of the present disclosure, a comparative analysis was respectively performed on the components in the high-frequency region of adult male and female voices (i.e., non-child voices) and the voices of children (e.g., 3 - 6 years old). The analysis found that there are differences in the dispersion of the component frequency band energy regions between non-child voices and child voices: the dispersion of each component of adults is more obvious, and the component frequency band energy region is more dispersed, that is, the dispersion is larger; the energy of child voices is more concentrated, that is, the dispersion is smaller.
[0058] At the same time, in at least one embodiment of the present disclosure, setting the penalty factor to 5000 can satisfy the acquisition of the main distinguishing components of child voice signals. Child voices show very strong energy concentration information below 5000 Hz.
[0059] Therefore, in one embodiment, the total number K of mode components and the penalty coefficient α are set respectively as . During this experiment, in the process of variational mode decomposition, the preset K value determines the number of finite bandwidth intrinsic mode function components or simply mode components. In at least one embodiment of the present disclosure, select . The larger the value of K, the greater the computational consumption. At the same time, it will cause the introduction of some signal components outside the voice frequency range, which is not beneficial to the overall effect. The penalty coefficient α determines the bandwidth of the mode components. The smaller the penalty coefficient α, the larger the bandwidth of each mode component, but too large a bandwidth will cause some components to contain other component signals; the larger the value of α, the smaller the bandwidth of each mode component, and too small a bandwidth will cause some signals in the decomposed signal to be lost. So the total number K of mode components and the penalty coefficient α are respectively will be appropriate. Of course, K and α can also take other values. For example, the value range of K is between 1 - 10, and the value range of α is between 1000 - 10000, etc. And preferably, the value range of K is between 3 - 8, and the value range of α is between 3000 - 8000, etc. This is not limited here. For example 6000, etc.
[0060] The training audio frame f(t) can be described by the following formula:
[0061] (1)
[0062] In expression (1), K represents the number of modal components, and t represents time. It can be seen that during the variational mode decomposition process, it is necessary to decompose each modal component u k (t).
[0063] The total energy value of the training audio frame f(t) is:
[0064] (2)
[0065] The model for variational mode decomposition is defined as follows:
[0066] (3)
[0067] Where respectively represent the k-th modal component and its corresponding center frequency. The total number of modal components is K, is the unit impulse function, j represents the imaginary unit; * represents the convolution operation; is the partial derivative operation; f(t) is the target signal. Introduce the penalty coefficient (or penalty factor) , and the Lagrange multiplier to solve the variational constraint problem. Therefore, the augmented Lagrangian expression is as follows:
[0068] (4)
[0069] (5)
[0070] (6)
[0071] (7)
[0072] Since the above Lagrangian augmented function has multiple variables, the Alternating Direction Method of Multipliers (ADMM) can be used to update and iterate to solve the saddle point of formula (7), and iterate and update in the frequency domain to decompose the signal into K modal components.
[0073] Specifically, the variational mode decomposition decomposes the signal f(t) into modal components, and the steps are as follows:
[0074] (1) Initialization
[0075] (2) They are iteratively updated separately by step (3).
[0076] (3)
[0077]
[0078]
[0079] Among them, is the fidelity coefficient; represents the Fourier transform; n is the number of iterations;
[0080] Set the iteration termination condition: , is the discrimination accuracy. After the iteration terminates, the training audio frame is decomposed into K modal components.
[0081] The above process can be described by steps 205, 206, 207, 208, 209.
[0082] Specifically, in step 205, calculate the power spectral density of the training audio frame f(t), and find the frequency corresponding to the maximum power spectral density as the initial center frequency ω used for the first decomposition i , i = 1.
[0083] In step 206, perform the first variational mode decomposition to decompose the 1st (k = 1) modal component u1(t) from the signal f(t). In step 207, judge whether the iteration termination condition is satisfied. If the judgment is yes, the algorithm terminates in step 208. Here, since 5 modal components have not been obtained yet, the judgment is no, so subtract the 1st (k = 1) modal component u1(t) from the signal f(t) to obtain the remaining signal.
[0084] Then, in step 205, calculate the power spectral density of the remaining signal, and find the frequency corresponding to the maximum power spectral density as the initial center frequency ω used for the second decomposition i, i = 2. Then, in step 206, perform the second variational mode decomposition to decompose the second (k = 2) modal component u2(t) from the remaining signal. In step 207, determine whether the iteration termination condition is satisfied. If the determination is yes, the algorithm terminates in step 208. Here, since 5 modal components have not been obtained yet, the determination is no. Then, subtract the second (k = 2) modal component u2(t) from the remaining signal to obtain the remaining signal. Then, the remaining signal obtained again is passed through steps 205 - 207 and iterated until the decomposition of 5 modal components is completed. In step 208, the algorithm stops and proceeds to step 209 to obtain the speech waveform signal with K = 5 intrinsic modal components.
[0085] In step 210, perform a second frame segmentation on each modal component. For example, the frame length is 45 ms and the frame shift is 15 ms.
[0086] In the first frame segmentation, after (360 ms) VMD decomposition, K finite - bandwidth intrinsic modal components can be obtained on the VMD modal axis. In the second frame segmentation, 24 audio frames can be segmented from each intrinsic modal component on the time axis.
[0087] In step 211, perform a second windowing on each audio frame after the second frame segmentation. A Hann window can also be used.
[0088] In step 212, use a gammatone frequency cepstral coefficient feature extractor composed of gammatone filters to extract the features simulating the basilar membrane induction information of the human ear for each audio frame after the second windowing.
[0089] Specifically, for each audio frame after the second time, use a predetermined number of gammatone filter banks with logarithmically distributed center frequencies to extract the gammatone frequency cepstral coefficients of a predetermined number of dimensions as the features simulating the basilar membrane induction information of the human ear, so as to obtain two - dimensional features for each audio frame.
[0090] Taking the predetermined number of center frequencies as 128 as an example, that is, for each audio frame of each modal component, use 128 gammatone filter banks with logarithmically distributed center frequencies to extract 128 - dimensional gammatone frequency cepstral coefficients. The reason for choosing 128 filters is that for the human speech bandwidth, the resolution of 128 is sufficient, and increasing it further will not bring more benefits. Here, the value of 128 is only an example, not a limitation. Of course, a value less than 128 can also be chosen, such as 80, 64, etc., and the computational amount will also be reduced accordingly. The gammatone frequency cepstral coefficient feature can simulate the basilar membrane induction information of the human ear. Finally, two - dimensional features are obtained , where the letter represents a real number. In at least one embodiment of the present disclosure, this feature is used as the input feature of the speech category classification neural network.
[0091] Note that during the training process and the inference process of the speech category classification neural network, the above-mentioned feature extraction is performed on the input audio frames (whether they are training audio frames or audio frames to be recognized) as the input of the speech category classification neural network. Here, the above-mentioned feature extraction process will not be described in detail again for the two processes.
[0092] In some embodiments, before inputting the features extracted in the above steps into the speech category classification neural network (for training or inference), feature preprocessing can be performed: normalizing the features extracted in the above steps using mean and variance normalization, etc.
[0093] In order to further highlight the specificity of the children's spectrogram, since the energy of the children's spectrogram has a higher contrast compared to the adult spectrogram. Therefore, the binaryzation method can better highlight the sharp structure of this spectrogram. The binaryzation method is as follows. In at least one embodiment of the present disclosure, the features are binaryzation processed using the global mean. (1) Accumulate all the values of all spectrogram points of each frame of data of each speech in the training set to obtain the cumulative sum of the spectral values of all data frames in the training set, and at the same time record the sum of all frames in the training set. Use the cumulative sum / frame sum to calculate the global mean of the global spectral value ; (2) According to the mean calculated in (1) , re-assign each spectral value point in the spectrogram according to the following formula.
[0094]
[0095] In the formula respectively represent the binaryzation feature point value and the original feature point value in the spectrogram feature. Here, the purpose of performing mean and variance normalization and binaryzation preprocessing is to train the neural network more uniformly, but in other embodiments, the features may not necessarily be preprocessed.
[0096] As mentioned above, the label generation corresponding to the training audio frames can be pre-performed. First, use children's human voices, adult human voices, non-human voices, reverberation, etc., and mix them in proportion to obtain various training audio frames of the voice files with added noise and reverberation. Then, corresponding labels need to be added to each training audio frame to indicate what kind of speech state the training audio frame has, such as children's human voice, non-children's human voice, silence, or other noises, etc. However, this application is not limited to this. For specific recognition requirements, other types and other numbers of labels can also be generated.
[0097] Figure 3 The flowchart of generating a training set of training audio files with noise and reverberation according to at least one embodiment of the present disclosure is shown.
[0098] As Figure 3As shown, in terms of generating training audio files with noise and reverberation, a noise addition step 301 and a reverberation addition step 302 are performed on the speech sample data to obtain an audio file with added noise and reverberation as a training sample, so as to simulate the sounds in various real noisy environments.
[0099] In step 301, common background noises such as home study, classroom study, and road study, including home background noise, road background noise, and common background noise on buses and subways, can be added. The purpose is to simulate the speech signals contaminated by environmental noise after passing through these environments.
[0100] In step 302, reverberation information in common learning spaces such as home space environment, classroom space environment, buses, subways, etc. can be added to obtain reverberated speech, aiming to simulate the speech signals contaminated by reverberation after passing through these spaces.
[0101] In this way, speech with added noise and reverberation is generated based on the speech sample data.
[0102] In terms of obtaining labels, first, an intrinsic mode decomposition feature extraction step is performed on the speech sample data (this feature extraction step can be the same as steps 201 - 212 of Figure 2 , that is, sample features about the simulated human ear basilar membrane induction information of the human voice are extracted from the speech sample data through variational mode decomposition and gamma-tone frequency cepstral coefficient filters) (step 303) to obtain gamma-tone frequency cepstral coefficient features (GFCC features). The combination (GMM + HMM) model training of the "speech state" granularity Gaussian Mixture Model (Gaussian Mixture Model, GMM) and Hidden Markov Model (Hidden Markov Model, HMM) is completed using the GFCC features (step 304). The Baum-Welch algorithm can be used for training to obtain the GMM + HMM model. In step 305, time information of the speech state granularity related to the speech sample data is generated through the sample features.
[0103] Specifically, when using the GMM + HMM model to complete the time information annotation at the "voice state" granularity for the voice training file, the time information annotation at the "voice state" granularity refers to annotating the start timestamp and end timestamp of each voice state. For example, the voice training file includes a segment of children's voice, two segments of non - children's voice, three segments of silence, five segments of other noises, etc. Then, the start timestamp of a segment of children's voice can be obtained as AM 11:30, and the end timestamp as AM 11:32. The start timestamps of the two segments of non - children's voice are AM 11:35 and AM 11:37, and the end timestamps are AM 11:37 and AM 11:38 respectively. The start and end timestamps of the three segments of silence (not listed one by one here), the start and end timestamps of the five segments of other noises (not listed one by one here). Based on these timestamps, voice states can be obtained. For example, before AM11:30 is the "before voice" state, from AM 11:30 to AM 11:32 is the "during voice" state, AM 11:32 is the "end of voice" state, from AM 11:32 to AM 11:35 is the "before voice" state, from AM 11:35 to AM 11:38 is the "during voice" state, and so on (not listed one by one here). In this example, the time information annotation for the "before voice" state is before AM 11:30, and for the "during voice" state is from AM 11:30 to AM 11:32, and so on. In addition, the voice state can also include other states, which will be described in detail later. The voice category labels can include 4 labels: children's voice, non - children's voice, silence, other noises. Of course, the labels are not limited to this and can be changed according to the actual situation.
[0104] Since it is known that the voice with added noise and reverberation is obtained from the voice sample data, the time correspondence relationship between the voice sample data and the voice with added noise and reverberation can be known. And the above - mentioned completion of the time information annotation at the "voice state" granularity for the voice training file obtains the start time and end timestamp of each voice state at the voice state granularity. Therefore, the voice with added noise and reverberation can be cut into multiple training audio frames lasting for a period of time (each training audio frame can correspond to each time period in the voice sample data in terms of time). Based on the start time and end timestamp of the voice segment of each state obtained for the voice sample data, the time information of the voice category corresponding to each training audio frame (such as the start time and end time of the "before voice" state, the start time and end time of the "during voice" state, the start time and end time of the "end of voice" state, etc.) and the corresponding voice category (such as children's voice, non - children's voice, noise, silence, etc.) can be obtained.
[0105] Thus, in step 306, multiple training audio frames and corresponding speech categories can be used as a training set to train a speech category classification neural network.
[0106] Next, it is described specifically how to use multiple training audio frames and corresponding speech categories as a training set to train a speech category classification neural network.
[0107] Specifically, the training audio frames are decomposed into multiple intrinsic mode function components through variational mode decomposition (including framing (if necessary) and preprocessing the training audio frames; calculating the power spectral density of the framed and preprocessed training audio frames; decomposing the framed and preprocessed training audio frames into multiple intrinsic mode function components based on the power spectral density); and the gamma-tone frequency cepstral coefficient filter is used to extract the features of simulating the information sensed by the basilar membrane of the human ear for each of the multiple intrinsic mode function components. The above method using variational mode decomposition and gamma-tone frequency cepstral coefficient filter is similar to that described previously for Figure 2 which can be trimmed and added according to the actual situation, and will not be repeated here.
[0108] The extracted features are fed into the speech category classification neural network as input data, and the neural network is trained in combination with the output labels. In at least one embodiment of the present disclosure, the output labels are speech categories, that is, it can include 4 labels: children's human voice, non-children's human voice, silence, and other noises.
[0109] Note that the output of the speech category classification neural network is the posterior probability that a speech frame to be recognized is respectively a children's human voice, a non-children's human voice, silence, and other noises. For example, when inputting a speech frame to be recognized, the output of the neural network is 80% for children's human voice, 10% for non-children's human voice, 5% for silence, and 5% for other noises. Finally, based on this output, it can be obtained that the speech frame to be recognized (with high probability) is a children's human voice. Of course, in subsequent processes, the posterior probabilities of the 4 outputs can be used as the output of the speech category classification neural network for description, but in other embodiments, only the children's human voice with the highest probability can also be used as the output of the speech category classification neural network.
[0110] The output layer of the speech category classification neural network can use the softmax function, and the number of layers, structure activation function, and the number of neurons in each layer of the speech category classification neural network can be determined through experiments or experience according to the duration of the actual data set. In the case where 360 ms is used as a time window range including 24 speech frames to be recognized, the output dimension of the speech category classification neural network can be: Posterior probabilities, where the 4 speech category recognition results respectively correspond to 4 posterior probabilities, and 24 frames will generate posterior probabilities, which can form a matrix.
[0111] Among them, the network layer structure of the voice category classification neural network can be adjusted according to the actual data volume, and the output labels can also be adjusted according to the categories to be modeled.
[0112] In this way, by creatively using variational mode decomposition and Gammatone (GT) frequency cepstral coefficient filters to extract specific features related to children's voices, the distinguishability between children's voices and background noise is enhanced to accurately detect children's voices.
[0113] Note that the above signal operations and neural network operations are all for detecting whether a frame of speech signal is a children's voice, a non-children's voice, silence, or other noise, and are based on one frame. These operations are some special optimizations for the children's speech scenario and are effective for this specific type of children's speech. After determining whether a frame of speech signal is a children's voice, a non-children's voice, silence, or other noise, it is necessary to verify based on the temporally continuous speech signals frame by frame (i.e., the speech signal sequence, or speech stream), and extract continuous long speech segments (or continuous speech segments, that is, continuous speech segments without interruption, such as a child reading a sentence continuously. This is more helpful for subsequent speech recognition and reading quality scoring of the continuous speech of a sentence. The following describes the operations of verifying the speech signal sequence composed of multiple audio frames to be recognized and extracting continuous long speech segments.
[0114] The reason for verifying the temporally continuous speech signals frame by frame here is that after determining whether a frame is a children's voice, a non-children's voice, silence, or other noise, a judgment also needs to be made on the speech stream. This is because the judgment of a neural network on a single frame may actually have certain errors. For example: for a speech stream, use 1 to represent that a frame is speech and 0 to represent that a frame is non-speech (assuming including non-children's voices, silence, and other noises). Then the output of the neural network may be: 11110011111111. But in fact, for example, the correct answer should be 11111111111111, all 1s. That is to say, the judgment that the middle 2 frames of signals are non-children's voices may be incorrect. Therefore, a verification and judgment logic is designed as follows to solve the problem of incorrect speech category judgment. For example, it can be inferred that the correct speech stream in the above example actually has no non-speech, that is, no 0s, so as to more accurately determine the continuous speech segment of the real children's voice.
[0115] First, define the voice transition segment. A transition segment means that from the generation to the end of the voice, it will go through, for example, 4 transition stages. The voice transfers from before voice (when the voice has not been generated yet) to before voice (when the voice has not been generated yet), from before voice (when the voice has not been generated yet or the voice starting point) to during voice (when the voice is being generated normally) (or the inflection point before voice), from during voice to during voice, and from during voice to the end of voice (or the inflection point at the end of voice). The voice states include before voice, during voice, and end of voice. All continuous voice segments must go through 4 consecutive transition segments in the correct order: before voice -> before voice, before voice -> during voice, during voice -> during voice, and during voice -> end of voice.
[0116] Figure 4 The figure shows a schematic diagram of the start and end points of a continuous voice segment according to at least one embodiment of the present disclosure. Figure 5 The figure shows a schematic diagram of the start and end points of two continuous voice segments according to at least one embodiment of the present disclosure. It can be seen that a complete continuous voice segment generally includes before voice (starting point) to during voice to end of voice in the correct order. Therefore, according to the voice transition segments in the correct order, it can assist in judging the continuous voice segments of real children's voices.
[0117] As mentioned before, a plurality of posterior probabilities with human voices have been determined from a plurality of consecutive audio frames to be recognized in time output by the voice category classification neural network. To determine the voice transition segments, it is still necessary to input the posterior probabilities output by the above-mentioned voice category classification neural network into a voice transition segment classification neural network to recognize the voice transition segments.
[0118] What the voice transition segment classification neural network needs to learn are 4 labels: before voice transfers to before voice, before voice transfers to during voice, during voice transfers to during voice, and during voice transfers to end of voice. Therefore, the transition segments of each voice state of the training audio file with added noise and reverberation can be labeled first to obtain the labels for training the voice transition segment classification neural network. Here, the before voice and the end part can be labeled according to the time interval of the clean voice, and then the middle part is the during voice state. This labeling step can be completed manually.
[0119] After the labeling is completed, 4 labels are obtained: before voice transfers to before voice, before voice transfers to during voice, during voice transfers to during voice, and during voice transfers to end of voice. The input of the voice transition segment classification neural network is the posterior probability of the voice category, such as 80% for children's voices, 10% for non - children's voices, 5% for silence, and 5% for other noises. For a 360 - ms time window, the input is the posterior probabilities as mentioned before.
[0120] The purpose of training this voice transition segment classification neural network is to determine two turning points, namely the starting point and the ending point, in the voice stream, that is, the two turning points of the voice transition segments "before voice -> during voice" and "during voice -> end of voice". These two turning points are opposite to each other. That is, the turning at the starting point is the voice changing from non-existent to existent, while the turning at the ending point is the voice changing from existent to non-existent. Therefore, in order to unify the input of the voice transition segment classification neural network, it can be considered to reverse the two turning points to the same turning direction to obtain a unified inflection point, so that it has the same prediction method and training method, in order to use the trained voice transition segment classification neural network to predict this similar turning.
[0121] Other state transitions, such as before voice -> before voice, during voice -> during voice, and those mentioned later such as too long voice silence, too long voice, too short voice, etc., can be judged based on some time thresholds when the voice transition segment classification neural network does not predict the two turning points of the voice transition segments "before voice -> during voice" and "during voice -> end of voice".
[0122] For example, in at least one embodiment of the present disclosure, the posterior probability of the voice category in the "during voice -> end of voice" is horizontally flipped at the starting rising position and the ending falling position of the voice data, so that it has the same arrangement method, prediction method and training method as the posterior probability of the input voice category of "before voice -> during voice". However, "during voice -> during voice" maintains the same input arrangement method as "before voice -> during voice", and other state transitions occur when these two state transitions are not realized and a certain threshold position is reached. In at least one embodiment of the present disclosure, the architecture of this voice transition segment classification neural network can adopt the use of Long Short-Term Memory-Convolutional Neural Networks (LSTM-CNN). At the same time, considering that the voice start inflection point (i.e., from before voice to during voice) and the voice end inflection point (i.e., from during voice to end of voice) are symmetric to each other, in order to capture the voice start inflection point and the voice end inflection point, the output of this voice transition segment classification neural network can use the softmax function, and the number of layers and width of this neural network can be adjusted according to the scale of the data set.
[0123] Figure 7 shows a schematic diagram of the training process of the voice transition segment classification neural network according to at least one embodiment of the present disclosure.
[0124] As Figure 7 shown, the features to be recognized (such as 128 5 The 24-dimensional GFCC feature matrix) is input into a speech category classification neural network to obtain 4 24-dimensional posterior probabilities of speech categories. The 4 24-dimensional posterior probabilities of speech categories are input into a speech transition segment classification neural network, and the speech transition segment labels manually annotated for the training audio frames mentioned above are used, such as 4 labels: pre-speech to pre-speech, pre-speech to mid-speech, mid-speech to mid-speech, mid-speech to end-of-speech, to train the speech transition segment classification neural network.
[0125] Note that the input of the posterior probability of each frame of the speech transition segment classification neural network can include: 12 frames on the left (previous in time) of a specific time point, the posterior probability of the speech category at the specific time point, and 12 frames (or other quantities) on the right (subsequent in time) of the specific time point, i.e., the posterior probability of each frame of the speech category. Figure 8 An example diagram is shown of temporarily storing 12 frames on the left and 12 frames on the right of a specific time point through a delay memory slot with a length L2 of 24 frames according to at least one embodiment of the present disclosure. That is, the detection sliding window length is maintained as frames (where 24 frames are only an example and not a limitation, and other frame numbers can also be taken. The smaller the window length, the higher the detection accuracy). The posterior probability matrix output from the middle frame of the detection sliding window added to the training audio frame is obtained in real time as the input of the speech transition segment classification neural network, and the dimension of the posterior probability matrix is 24 4, and the probability values of the children's voices and other human voice parts are combined. The 24 4 matrices are combined into a 24 3 matrix.
[0126] The purpose of doing this is that only multiple frames that are continuous in time can have the determination of the speech transition segment. For example, 12 frames on the left and 12 frames on the right of a specific time point (or the context frames of the specific time point, denoted by frames) can be spliced to form a speech transition segment from pre-speech to mid-speech. Then, the output label of this input of the speech transition segment classification neural network is "pre-speech to mid-speech", and the speech transition segment classification neural network is trained accordingly.
[0127] As described above, when in the state of "during speech -> end of speech", the start rising position and the end falling position of the speech data can be horizontally flipped so that they have the same arrangement, prediction method, and training method as the posterior probability of the input speech category of "before speech -> during speech". Therefore, during the training process, when the label of the current frame transitions from the state of "before speech" to the state of "during speech", the input to the speech transition segment classification neural network is the context frame input with the order swapped (i.e., the previous 12 frames at a specific time point and the subsequent 12 frames at a specific time point are swapped in order, and in the time sequence from front to back, they are the subsequent 12 frames at a specific time point and the previous 12 frames at a specific time point). And when the label of the current frame transitions from the state of "during speech" to the state of "end of speech", the input to the speech transition segment classification neural network is the context frame input in the normal order (i.e., in the time sequence from front to back, they are the previous 12 frames at a specific time point and the subsequent 12 frames at a specific time point). Therefore, the speech transition segment classification neural network trained in this way can easily and more accurately determine only the speech transition segment from no speech to speech.
[0128] Specifically, according to the frame-level judgment result that the current frame is in the state of "before speech", the context is used as to perform the flipping process. That is, for the speech segment of "before speech", an appropriate start turning frame position is manually selected , and for the other frames on both sides of the i-th frame, the following method is used to complete the frame position swapping:
[0129]
[0130] The input frames after completing the frame position swapping, that is, in the time sequence from front to back, the previous 12 frames at a specific time point and the subsequent 12 frames at a specific time point, are used as the input to train the speech transition segment classification neural network. For the speech segments of "during speech" and "end of speech", no special processing is performed, and they are directly used as the input for training the speech transition segment classification neural network.
[0131] The speech transition segment classification neural network trained in this way can also be used to reason, predict, and judge speech inflection points.
[0132] After training the speech transition segment classification neural network, during the inference process, the features to be recognized can also be extracted from the audio frames to be recognized through variational mode decomposition and gamma-tone frequency cepstral coefficient filters and input into the speech category classification neural network to obtain the posterior probability of the speech category. The posterior probability of the speech category is input into the speech transition segment classification neural network to output the posterior probability of the speech transition segment, such as the posterior probability of each label among the four labels of before speech -> before speech, before speech -> during speech, during speech -> during speech, and during speech -> end of speech. Note that in at least one embodiment of the present disclosure, the symbol "->" can represent "transfer to".
[0133] Similarly, during the inference process, when the current frame is in the "before speech" state, the voice transition segment classification neural network uses the context frames in the swapped order for input. When the current is in the "during speech" state, the voice transition segment classification neural network uses the context frames in the normal order for input.
[0134] As before, since the classification judgment of a neural network for a frame may actually have certain errors. For example: for a voice stream, use 1 to represent a frame as voice and 0 to represent a frame as non-voice (assuming including non-child human voices (adult human voices), silence, and other noises). Then the output of the neural network may be: 11110011111111. But in fact, for example, the correct answer should be 11111111111111, all 1s. That is to say, it may be incorrect that the middle 2 frame signals are judged as non-child human voices.
[0135] To solve the above problem of misclassification, in some embodiments, step 130 of identifying one or more continuous voice segments of a human voice according to the posterior probabilities of multiple human voices determined from multiple audio frames to be recognized may include: determining one or more continuous voice segments of a human voice by using the Markov state transition relationship according to the posterior probabilities of the voice transition segments of the determined multiple audio frames to be recognized.
[0136] Figure 6 The Markov state transition relationship diagram of each voice state according to at least one embodiment of the present disclosure is shown.
[0137] As Figure 6 shown, the states included in the Markov state transition process are: before speech, too long silence before speech, during speech, end of speech, too long speech, too short speech. Among them, the three states of before speech, during speech, and end of speech are the ultimate feedback states. Transferring to these three states in chronological order means completing a state acquisition and immediately jumping out of the Markov state transition process judgment. That is, the complete voice segments of before speech, during speech, and end of speech in chronological order are one continuous voice segment of the human voice to be determined. And if the states of the Markov state are too long silence before speech, too long speech, or too short speech, since the audio is a linear time series, introducing these three states is not the ultimate state. The state transition relationship is specifically, for example, as Figure 6 shown.
[0138] Figure 6 The Markov state transition relationship of
[0139] Table 1 Transfer matrix adjacency table definition
[0140]
[0141] As Figure 6 shown in Table 1, the pre - speech state can transition to the pre - speech state, in - speech state, or pre - speech silence - too - long state, but should not transition to speech - too - long, speech - too - short, or speech - ended, because the in - speech state has not been entered yet. That is to say, if the conclusion based on the posterior probability of the speech transition segment obtained from the speech transition segment classification neural network is that the pre - speech state is followed by the speech - ended state, then such an output is incorrect.
[0142] As previously trained, the speech transition segment classification neural network can determine the inflection points of speech. Next, based on the speech detection sliding window as shown in Figure 8 (for example, a detection sliding window of frames before and after), the speech frames to be recognized can be input into the trained speech transition segment classification neural network to detect the speech state, and then state transitions can be made in the Markov state transition relationship as shown in Figure 6 . Thus, speech segments that do not conform to the Markov state transition relationship can be automatically removed, and the temporally continuous speech segments that conform to the Markov state transition relationship can be automatically extracted as a complete speech segment of the child's voice, so as to perform subsequent speech content detection, pronunciation scoring, etc.
[0143] To make state transitions on the Markov state transition relationship diagram as shown in Figure 6 , define the state variable tuple . represents the minimum threshold of speech duration and the maximum threshold of speech duration. represents the relaxation factor for pre - speech and speech - ended. represents the maximum pre - speech silence threshold.
[0144] During the calculation of pre - speech state transitions, first initialize all variables. The current state is in the "pre - speech" state, including re - transitioning to the "pre - speech" state.
[0145] Therefore, according to the posterior probability of the speech transition segments of the determined multiple audio frames to be recognized, using the Markov state transition relationship to determine one or more continuous speech segments of the human voice includes the following steps.
[0146] Starting from the "pre - speech" state, judge whether the "pre - speech" state transitions to the "pre - speech silence - too - long" state through the following conditions: the current state is in the "pre - speech" state, the speech transition segment classification neural network does not detect the transition of the "pre - speech" state to the "in - speech" state, and the cumulative time of the "pre - speech" state satisfies being greater than or equal to the first threshold, then the state directly transitions to the "pre - speech silence - too - long" state, where the state is reset to the "pre - speech" state.
[0147] The "before speech" state transfers to the "before speech" state under the following conditions: The current state is the "before speech" state, and the speech transition segment classification neural network detects whether there is an inflection point where the "before speech" state transfers to the "in speech" state. If the cumulative time of the "before speech" state does not meet the requirement of being greater than or equal to the first threshold, and the speech transition segment classification neural network does not detect an inflection point where the "before speech" state transfers to the "in speech" state, then the state transfers to the "before speech" state.
[0148] The "before speech" state transfers to the "in speech" state under the following conditions: The current state is the "before speech" state, and the speech transition segment classification neural network detects an inflection point where the "before speech" state transfers to the "in speech" state, then the state transfers to the "in speech" state.
[0149] The "in speech" state transfers to the "speech too long" state under the following conditions: The current state is the "in speech" state, and the cumulative time of the "in speech" state meets the requirement of being greater than or equal to the second threshold and still has not transferred to the "end of speech" state, then the state transfers to the "speech too long" state. The speech segment from the "before speech" state to the second threshold is intercepted and output as a continuous speech segment of the human voice. Among them, the state is reset to the "before speech" state.
[0150] The "in speech" state transfers to the "end of speech" state under the following conditions: The current state is the "in speech" state, and the speech transition segment classification neural network detects an inflection point where the "in speech" state transfers to the "end of speech" state, then the state transfers to the "end of speech" state.
[0151] The "end of speech" state transfers to the "speech too short" state under the following conditions: The current state is the "end of speech" state, and the cumulative time of the "in speech" state before the "end of speech" state meets the requirement of being less than or equal to the third threshold, then the state transfers to the "speech too short" state. Among them, the speech segment with a duration of the third threshold starting from the "before speech" state is discarded and output, and the state is reset back to the "before speech" state.
[0152] The "in speech" state transfers to the "in speech" state under the following conditions: The current state is the "in speech" state, the speech transition segment classification neural network does not detect an inflection point where the "in speech" state transfers to the "end of speech" state, and the cumulative time of the "in speech" state does not meet the requirement of being less than or equal to the third threshold, then the state transfers back to the "in speech" state.
[0153] After judging the state and the state transitions for multiple consecutive audio frames to be recognized in time, a speech segment labeled as from the "before speech" state to the "in speech" state to the "end of speech" state in the training audio frames is extracted as a continuous speech segment of the human voice.
[0154] That is, specifically in combination with the threshold formula, it is judged that the "before speech" state transitions to the "too long silence before speech" state through the following conditions: the current state is in the "before speech" state, the voice transition segment classification neural network does not detect an inflection point where the "before speech" state transitions to the "during speech" state, and the cumulative time of the "before speech" state satisfies , and it directly transitions to the "too long silence before speech" state. Then reset and re-transition to the calculation process of the before speech state transition.
[0155] It is judged that the "before speech" state transitions to the "before speech" state through the following conditions: the current state is in the "before speech" state, and the voice transition segment classification neural network detects frame by frame whether there is an inflection point where the "before speech" state transitions to the "during speech" state. If the cumulative time of the "before speech" state does not satisfy , and no inflection point is detected, then it transitions to the "before speech" state.
[0156] It is judged that the "before speech" state transitions to the "during speech" state through the following conditions: the current state is in the "before speech" state, and the voice transition segment classification neural network detects an inflection point where the "before speech" state transitions to the "during speech" state, and it transitions to the "during speech" state.
[0157] During the calculation process of the during speech state transition:
[0158] It is judged that the "during speech" state transitions to the "too long speech" state through the following conditions: the current state is in the "during speech" state, and the cumulative time of the "during speech" state satisfies and it still has not transitioned to the "end of speech" state, then it directly transitions to the "too long speech" state. The voice segment from the "before speech" state to the time threshold can be intercepted and output as a complete voice segment of the child's human voice. The state is reset to the "before speech" state.
[0159] It is judged that the "during speech" state transitions to the "end of speech" state through the following conditions: the current state is in the "during speech" state, and the voice transition segment classification neural network detects an inflection point where the "during speech" state transitions to the "end of speech" state, and it transitions to the "end of speech" state.
[0160] It is judged that the "end of speech" state transitions to the "too short speech" state through the following conditions: the current state is in the "end of speech" state, and the cumulative time of the "during speech" state before the "end of speech" state satisfies , and it directly transitions to the "too short speech" state. The voice segment from the "before speech" state to Output of the voice segment. Because too short a voice may not be a complete voice segment, or it may be misjudged by the voice category classification neural network as a children's voice, or it may not be suitable for subsequent operations such as recognizing the voice content and pronunciation scoring. Reset to the "before voice" state.
[0161] Judge the transition from the "during voice" state to the "during voice" state through the following conditions: The current state is in the "during voice" state, the voice transition segment classification neural network does not detect the inflection point of the transition from the "during voice" state to the "end of voice" state, and the cumulative time of the "during voice" state does not meet , and transfer back to the "during voice" state.
[0162] After judging all voice states and state transitions for multiple consecutive audio frames to be recognized in time, extract a voice segment marked as from the "before voice" state to the "during voice" state to the "end of voice" state in the training audio frames as a continuous voice segment of children's voices.
[0163] Note that when judging the "before voice" state, "during voice" state, and "end of voice" state, relaxation processing can be performed on the before-voice inflection point and the end-of-voice inflection point, that is, control the judgment accuracy of the before-voice inflection point and the end-of-voice inflection point, and judge the "before voice" state, "during voice" state, and "end of voice" state according to the judgment accuracy.
[0164] The relaxation factors of the before-voice inflection point and the end-of-voice inflection point are 、 , and the relaxation factor of the before-voice inflection point means that after judging the before-voice inflection point, extend forward frames to avoid the position of the before-voice inflection point being too tight (too high in accuracy), resulting in too much voice being cut; conversely, the relaxation factor of the end-of-voice inflection point means that after judging the end-of-voice inflection point, extend backward frames to avoid the position of the end-of-voice inflection point being too tight (too high in accuracy), resulting in the voice being cut off too early.
[0165] The relaxation factors of the before-voice end point and the end-of-voice inflection point can be calculated through the following formula:
[0166]
[0167] respectively represent the relaxed before-voice inflection point and the before-voice inflection point judged by the voice transition segment classification neural network. respectively represent the relaxed end-of-voice inflection point and the end-of-voice inflection point judged by the voice transition segment classification neural network.
[0168] In this way, it is possible to determine the states from "before speech" to "during speech" to "end of speech" through Markov state transition relationships and the discrimination of inflection points in speech states.
[0169] Then, based on the determined states from "before speech" to "during speech" to "end of speech", the transition segments of the speech states of multiple audio frames to be recognized are labeled. Then, a speech segment labeled with the states from "before speech" to "during speech" to "end of speech" among the multiple audio frames to be recognized is extracted as a continuous speech segment of the children's human voice.
[0170] In this way, the present solution also creatively uses the Markov state transition relationship of speech states to determine a continuous speech segment of the children's human voice from before speech to during speech to end of speech, which also helps to further verify whether the previously recognized children's human voice is correct or to eliminate overly short continuous speech segments.
[0171] In this way, according to the various embodiments of the present application, the detection ability of continuous speech segments of human voices in complex noise backgrounds is improved, ensuring the stability of the recognition performance and evaluation performance of speech recognition and speech evaluation systems in, for example, educational scenarios.
[0172] Figure 9 A block diagram of a device 900 for recognizing continuous speech segments of human voices according to at least one embodiment of the present disclosure is shown.
[0173] As Figure 9 shown, the device 900 for recognizing continuous speech segments of human voices includes: an extraction device 910, a determination device 920, and an identification device 930.
[0174] The extraction device 910 may be configured to extract multiple features to be recognized regarding the simulated basilar membrane induction information of the human voice from multiple temporally continuous audio frames to be recognized through variational mode decomposition and gamma-tone frequency cepstral coefficient filters.
[0175] The determination device 920 may be configured to input the extracted multiple features to be recognized into a speech category classification neural network to determine multiple posterior probabilities of having human voices in the multiple audio frames to be recognized from the multiple audio frames to be recognized.
[0176] The identification device 930 may be configured to identify one or more continuous speech segments of the human voice based on the multiple posterior probabilities of having human voices determined in the multiple audio frames to be recognized.
[0177] In one embodiment, a speech category classification neural network can be trained through the following steps: generating speech with noise and reverberation based on speech sample data; extracting sample features about human voice that simulate human ear basilar membrane sensing information from the speech sample data through variational mode decomposition and gammatone frequency cepstral coefficient filter; generating time information of speech state granularity related to the speech sample data through the sample features; based on the time correspondence between the speech sample data and the speech with noise and reverberation and the time information of the speech state granularity related to the speech sample data, cutting the speech with noise and reverberation into multiple training audio frames and obtaining the speech category and speech category corresponding to each training audio frame; using the multiple training audio frames and speech categories as training sets to train the speech category classification neural network.
[0178] In one embodiment, the extraction device 910 can be configured to: decompose the speech sample data into multiple intrinsic mode function components through variational mode decomposition; and use a gamma-tone frequency cepstral coefficient filter to extract characteristics of simulated human ear basilar membrane sensing information from each of the multiple intrinsic mode function components.
[0179] In one embodiment, a training audio frame is decomposed into a plurality of intrinsic mode function components by variational mode decomposition, including: framing and preprocessing speech with noise and reverberation; calculating power spectral density of the framing and preprocessed speech with noise and reverberation; and decomposing the framing and preprocessed speech with noise and reverberation into a plurality of intrinsic mode function components based on the power spectral density.
[0180] In one embodiment, the speech with added noise and reverberation is generated by adding at least one of one or more human voices, noise, and reverberation to speech sample data, wherein the speech categories include at least predetermined human voices, human voices other than predetermined human voices, noise other than human voices, and silence.
[0181] In one embodiment, multiple features to be identified about human voice that simulate the sensing information of the basilar membrane of the human ear are extracted from multiple temporally consecutive audio frames to be identified by variational modal decomposition, including: decomposing each of the multiple audio frames to be identified into multiple intrinsic mode function components by variational modal decomposition; extracting the features to be identified of the simulated sensing information of the basilar membrane of the human ear from each intrinsic mode function component of the multiple intrinsic mode function components by a gamma-tone frequency cepstral coefficient filter.
[0182] In one embodiment, the recognition device 930 can be configured to: determine the posterior probability of the speech transition segments of the multiple audio frames to be recognized by using the speech transition segment classification neural network based on the posterior probabilities of the multiple human voices determined in the multiple audio frames to be recognized.
[0183] In one embodiment, the voice transition segment classification neural network is trained as follows: Generate noisy and reverberant speech based on voice sample data and cut the noisy and reverberant speech into multiple training audio frames; Label the noisy and reverberant speech with voice transition segment labels, where the voice transition segment labels include "before voice" to "before voice", "before voice" to "during voice", "during voice" to "during voice", "during voice" to "end of voice"; Input multiple training audio frames and voice transition segment labels to train the voice transition segment classification neural network.
[0184] In one embodiment, the recognition device 930 may be configured to: Determine one or more continuous voice segments of the human voice by using the Markov state transition relationship according to the posterior probabilities of the voice transition segments of the determined multiple audio frames to be recognized.
[0185] In one embodiment, the Markov state transition relationship is determined as follows:
[0186] The states set for the Markov state transition relationship include the "before voice" state, the "too long voice silence before voice" state, the "during voice" state, the "end of voice" state, the "too long voice" state, and the "too short voice" state;
[0187] Set the transition relationships between the states of the Markov state transition relationship, where the transition relationships include: the "before voice" state transitions to the "before voice" state, the "before voice" state transitions to the "during voice" state, the "before voice" state transitions to the "too long voice silence before voice" state, the "during voice" state transitions to the "during voice" state, the "during voice" state transitions to the "end of voice" state, the "too long voice silence before voice" state transitions to the "before voice" state, the "too long voice" state transitions to the "before voice" state, the "too short voice" state transitions to the "before voice" state, the "end of voice" state transitions to the "before voice" state, and the "end of voice" state transitions to the "too short voice" state.
[0188] In one embodiment, the recognition device 930 may be configured to: starting from the "before speech" state, determine whether the "before speech" state transitions to the "too long silence before speech" state through the following conditions: the current state is in the "before speech" state, the speech transition segment classification neural network does not detect a transition from the "before speech" state to the "during speech" state, and the cumulative time of the "before speech" state satisfies being greater than or equal to a first threshold, then the state directly transitions to the "too long silence before speech" state, and the state is reset to the "before speech" state; determine whether the "before speech" state transitions to the "before speech" state through the following conditions: the current state is in the "before speech" state, and the speech transition segment classification neural network detects whether there is an inflection point of the transition from the "before speech" state to the "during speech" state. If the cumulative time of the "before speech" state does not satisfy being greater than or equal to the first threshold, and the speech transition segment classification neural network does not detect an inflection point of the transition from the "before speech" state to the "during speech" state, then the state transitions to the "before speech" state; determine whether the "before speech" state transitions to the "during speech" state through the following conditions: the current state is in the "before speech" state, and the speech transition segment classification neural network detects an inflection point of the transition from the "before speech" state to the "during speech" state, then the state transitions to the "during speech" state; determine whether the "during speech" state transitions to the "too long speech" state through the following conditions: the current state is in the "during speech" state, and the cumulative time of the "during speech" state satisfies being greater than or equal to a second threshold and has not yet transitioned to the "end of speech" state, then the state transitions to the "too long speech" state, and the speech segment from the "before speech" state to the second threshold is intercepted and output as a continuous speech segment of the human voice, and the state is reset to the "before speech" state; determine whether the "during speech" state transitions to the "end of speech" state through the following conditions: the current state is in the "during speech" state, and the speech transition segment classification neural network detects an inflection point of the transition from the "during speech" state to the "end of speech" state, then the state transitions to the "end of speech" state; determine whether the "end of speech" state transitions to the "too short speech" state through the following conditions: the current state is in the "end of speech" state, and the cumulative time of the "during speech" state before the "end of speech" state satisfies being less than or equal to a third threshold, then the state transitions to the "too short speech" state, and the speech segment with a duration of the third threshold starting from the "before speech" state is discarded and output, and the state is reset back to the "before speech" state; determine whether the "during speech" state transitions to the "during speech" state through the following conditions: the current state is in the "during speech" state, the speech transition segment classification neural network does not detect an inflection point of the transition from the "during speech" state to the "end of speech" state, and the cumulative time of the "during speech" state does not satisfy being less than or equal to the third threshold, then the state transitions back to the "during speech" state; after determining the states and the transitions between the states for a plurality of consecutive audio frames to be recognized in terms of time, a speech segment labeled from the "before speech" state to the "during speech" state to the "end of speech" state in the training audio frames is extracted as a continuous speech segment of the human voice.
[0189] Thus, according to various embodiments of the present application, the detection ability of continuous speech segments of human voices in a complex noise background is improved, ensuring the stability of the recognition performance and evaluation performance of speech recognition and speech evaluation systems in educational scenarios, for example.
[0190] Figure 10 A block diagram of an exemplary device for identifying continuous speech segments of human voices suitable for implementing at least one embodiment of the present disclosure is shown.
[0191] The device for identifying continuous speech segments of human voices may include a processor 1010 and a memory 1020. The memory 1020 is coupled to the processor 1010 and stores computer instructions therein for performing the steps of the various methods of at least one embodiment of the present disclosure when executed by the processor 1010. The device for identifying continuous speech segments of human voices may be manufactured in the form of a smart phone, a tablet computer, a wearable device, etc.
[0192] The processor 1010 may include, but is not limited to, for example, one or more processors or microprocessors, etc.
[0193] The memory 1020 may include, but is not limited to, for example, a random access memory (RAM), a read-only memory (ROM), a flash memory, an EPROM memory, an EEPROM memory, a register, a computer storage medium (such as a hard disk, a floppy disk, a solid state drive, a removable disk, a CD-ROM, a DVD-ROM, a Blu-ray disc, etc.).
[0194] In addition, the device for identifying continuous speech segments of human voices may further include (but is not limited to) a data bus 1030, an input / output (I / O) bus 1040, a display 1050, and an input / output device 1060 (such as a keyboard, a mouse, a speaker, etc.).
[0195] The processor 1010 may communicate with an external display 1050 and an input / output device 1060, etc. through the I / O bus 1040.
[0196] In one embodiment, the at least one computer instruction may also be compiled into or form a computer program product or software product, and when one or more computer instructions are executed by a processor, the steps of the various functions and / or methods described in the embodiments of the present technology are performed.
[0197] At least one embodiment according to the present disclosure also provides a non-transitory computer-readable storage medium. Instructions are stored on the non-transitory computer-readable storage medium, and the instructions are, for example, computer instructions. When the computer instructions are executed by a processor, the various methods described above can be executed. The non-transitory computer-readable storage medium includes, but is not limited to, for example, random access memory (RAM), read-only memory (ROM), flash memory, EPROM memory, EEPROM memory, registers, computer storage media (such as hard disks, floppy disks, solid-state drives, removable disks, CD-ROMs, DVD-ROMs, Blu-ray discs, etc.). For example, the non-transitory computer-readable storage medium can be connected to a computing device such as a computer. Then, in the case where the computing device executes the computer instructions stored on the non-transitory computer-readable storage medium, the various methods described above can be performed.
[0198] The present disclosure may also include a computer program product, wherein the computer program product can perform the methods, steps, and operations given herein. For example, such a computer program product can be a computer software package, computer code instructions, a computer-readable tangible medium having computer instructions tangibly stored (and / or encoded) thereon, and the instructions can be executed by a processor to perform the operations described herein. The computer program product may include packaging materials.
[0199] The block diagrams of the devices, apparatuses, equipment, and systems involved in the present disclosure are only illustrative examples and are not intended to require or imply that they must be connected, arranged, and configured in the manner shown in the block diagrams. As those skilled in the art will recognize, these devices, apparatuses, equipment, and systems can be connected, arranged, and configured in any manner. Words such as "including", "comprising", "having", etc. are open-ended words, meaning "including but not limited to", and can be used interchangeably with each other. The phrase "such as / for example" used herein means "such as / for example but not limited to", and can be used interchangeably with it.
[0200] The step flowcharts in the present disclosure and the above method descriptions are only illustrative examples and are not intended to require or imply that the steps of the various embodiments must be performed in the order given. As those skilled in the art will recognize, the steps in the above embodiments can be performed in any order. Words such as "subsequently", "then", "next", etc. are not intended to limit the order of the steps; these words are only used to guide the reader through the description of these methods. In addition, for example, any reference to a singular element using the articles "a", "an", or "the" is not to be construed as limiting that element to the singular.
[0201] In addition, the steps and devices in each of the embodiments in at least one embodiment of the present disclosure are not limited to being implemented in a certain embodiment. In fact, according to the concept of the present disclosure, relevant partial steps and partial devices in at least one embodiment of the present disclosure can be combined to conceive new embodiments, and these new embodiments are also included within the scope of the present disclosure.
[0202] The above-described method can be implemented in hardware, software, firmware, or any combination thereof. Additionally, modules and / or other suitable means for performing the methods and techniques described herein can be downloaded wirelessly from a server as appropriate. Alternatively, the various methods described herein can be provided via a storage component so as to obtain the various methods when coupled to the storage component. Further, any other suitable technique for providing the methods and techniques described herein to a device can be utilized.
[0203] The foregoing description has been presented for purposes of illustration and description. Furthermore, this description is not intended to limit at least one embodiment of the present disclosure to the form disclosed herein. Although several example aspects and embodiments have been discussed above, those skilled in the art will recognize some of their variations, modifications, alterations, additions, and subcombinations.
Claims
1. A method for identifying continuous speech segments of human voices, comprising extracting, from a plurality of consecutively time - continuous audio frames to be recognized, a plurality of features to be recognized regarding the simulated human cochlear basilar membrane induction information of human voices through variational mode decomposition and gamma - tone frequency cepstral coefficient filters; inputting the extracted plurality of features to be recognized into a speech category classification neural network to determine, from the plurality of audio frames to be recognized, a plurality of posterior probabilities of the plurality of audio frames to be recognized having the human voices; identifying one or more continuous speech segments of the human voices according to the plurality of posterior probabilities of the plurality of audio frames to be recognized having the human voices determined; 2. The method according to claim 1, wherein The speech category classification neural network is trained through the following steps: generating speech with added noise and reverberation based on speech sample data; extracting, from the speech sample data, sample features regarding the simulated human cochlear basilar membrane induction information of the human voices through the variational mode decomposition and the gamma - tone frequency cepstral coefficient filters; generating time information of speech state granularity related to the speech sample data through the sample features; based on the time correspondence relationship between the speech sample data and the speech with added noise and reverberation and the time information of the speech state granularity related to the speech sample data, cutting the speech with added noise and reverberation into a plurality of training audio frames and obtaining the speech category corresponding to each training audio frame; using the plurality of training audio frames and the speech categories as a training set to train the speech category classification neural network.
3. The method according to claim 2, wherein the extracting, from the speech sample data, sample features regarding the simulated human cochlear basilar membrane induction information of human voices through variational mode decomposition and gamma - tone frequency cepstral coefficient filters comprises: decomposing the speech sample data into a plurality of intrinsic mode function components through the variational mode decomposition; extracting, for each of the plurality of intrinsic mode function components, sample features regarding the simulated human cochlear basilar membrane induction information by using the gamma - tone frequency cepstral coefficient filters.
4. The method according to claim 3, wherein, The decomposing the training audio frames into a plurality of intrinsic mode function components through the variational mode decomposition comprises: framing and pre - processing the speech with added noise and reverberation; calculating the power spectral density of the framed and pre - processed speech with added noise and reverberation; decomposing the framed and pre - processed speech with added noise and reverberation into the plurality of intrinsic mode function components based on the power spectral density.
5. The method according to claim 2, wherein, The speech with added noise and reverberation is generated by adding at least one of one or more human voices, noises, and reverberations to the speech sample data, wherein the speech categories at least include a predetermined human voice, a non - predetermined human voice, a non - human voice noise, and silence.
6. The method according to claim 1, wherein The extracting, from a plurality of consecutively time - continuous audio frames to be recognized, a plurality of features to be recognized regarding the simulated human cochlear basilar membrane induction information of human voices through variational mode decomposition comprises: decomposing each of the plurality of audio frames to be recognized into a plurality of intrinsic mode function components through the variational mode decomposition; Extracting the to-be-identified features of the simulated human cochlear basilar membrane induction information from each of the plurality of intrinsic mode function components through the gamma-common frequency cepstral coefficient filter.
7. The method according to claim 1, wherein The identifying one or more continuous speech segments of the human voice according to the plurality of posterior probabilities of the human voice determined from the plurality of to-be-identified audio frames includes: According to the posterior probabilities of multiple human voices determined from the plurality of to-be-identified audio frames, using a speech transition segment classification neural network to determine the posterior probabilities of the speech transition segments of the plurality of to-be-identified audio frames.
8. The method according to claim 7, wherein the speech transition segment classification neural network is trained in the following manner: Generating noisy and reverberated speech based on speech sample data and cutting the noisy and reverberated speech into a plurality of training audio frames; Label the speech transition segments for the speech with added noise and reverberation, where, The speech transition segment labels include "before speech" to "before speech", "before speech" to "during speech", "during speech" to "during speech", "during speech" to "end of speech"; Inputting the plurality of training audio frames and the speech transition segment labels to train the speech transition segment classification neural network.
9. The method according to claim 7, wherein, The identifying one or more continuous speech segments of the human voice according to the plurality of posterior probabilities of the human voice determined from the plurality of to-be-identified audio frames includes: According to the posterior probabilities of the speech transition segments of the determined plurality of to-be-identified audio frames, using a Markov state transition relationship to determine one or more continuous speech segments of the human voice.
10. The method according to claim 9, wherein, The Markov state transition relationship is determined in the following manner: Setting the states of the Markov state transition relationship to include a "before speech" state, a "before speech, speech silence too long" state, a "during speech" state, an "end of speech" state, a "speech too long" state, and a "speech too short" state; Setting the transition relationships between the states of the Markov state transition relationship, wherein the transition relationships include: the "before speech" state transitions to the "before speech" state, the "before speech" state transitions to the "during speech" state, the "before speech" state transitions to the "before speech, speech silence too long" state, the "during speech" state transitions to the "during speech" state, the "during speech" state transitions to the "end of speech" state, the "before speech, speech silence too long" state transitions to the "before speech" state, the "speech too long" state transitions to the "before speech" state, the "speech too short" state transitions to the "before speech" state, the "end of speech" state transitions to the "before speech" state, and the "end of speech" state transitions to the "speech too short" state.
11. The method according to claim 10, wherein, The determining one or more continuous speech segments of the human voice according to the posterior probabilities of the speech transition segments of the determined plurality of to-be-identified audio frames and using the Markov state transition relationship includes: Starting from the "before speech" state, the "before speech" state is determined to transition to the "before speech silent too long" state through the following conditions: The current state is in the "before speech" state, the speech transition segment classification neural network does not detect the "before speech" state transitioning to the "during speech" state, and the cumulative time of the "before speech" state meets or exceeds the first threshold. The state directly transitions to the "before speech silent too long" state, where the state is reset to the "before speech" state; The "before speech" state is determined to transition to the "before speech" state through the following conditions: The current state is in the "before speech" state, and the speech transition segment classification neural network detects whether there is an inflection point where the "before speech" state transitions to the "during speech" state. If the cumulative time of the "before speech" state does not meet or exceed the first threshold and the speech transition segment classification neural network does not detect an inflection point where the "before speech" state transitions to the "during speech" state, then the state transitions to the "before speech" state; The "before speech" state is determined to transition to the "during speech" state through the following conditions: The current state is in the "before speech" state, and the speech transition segment classification neural network detects an inflection point where the "before speech" state transitions to the "during speech" state. The state transitions to the "during speech" state; The "during speech" state is determined to transition to the "speech too long" state through the following conditions: The current state is in the "during speech" state, and the cumulative time of the "during speech" state meets or exceeds the second threshold and has not yet transitioned to the "end of speech" state. Then the state transitions to the "speech too long" state, and the speech segment from the "before speech" state to the second threshold is intercepted and output as a continuous speech segment of the human voice, where the state is reset to the "before speech" state; The "during speech" state is determined to transition to the "end of speech" state through the following conditions: The current state is in the "during speech" state, and the speech transition segment classification neural network detects an inflection point where the "during speech" state transitions to the "end of speech" state. The state transitions to the "end of speech" state; The "end of speech" state is determined to transition to the "speech too short" state through the following conditions: The current state is in the "end of speech" state, and the cumulative time of the "during speech" state before the "end of speech" state meets or is less than the third threshold. The state transitions to the "speech too short" state, where the speech segment with a duration of the third threshold starting from the "before speech" state is discarded and output, and the state is reset back to the "before speech" state; The "during speech" state is determined to transition to the "during speech" state through the following conditions: The current state is in the "during speech" state, the speech transition segment classification neural network does not detect an inflection point where the "during speech" state transitions to the "end of speech" state, and the cumulative time of the "during speech" state does not meet or is less than the third threshold. The state transitions back to the "during speech" state; After judging the states of a plurality of consecutively-time audio frames to be recognized and the transitions between the states, a continuous speech segment of the human voice is extracted from a speech segment in the training audio frames labeled with the states from "before speech" to "during speech" to "end of speech".
12. A device for recognizing a continuous speech segment of a human voice, comprising: An extraction device configured to extract a plurality of features to be recognized regarding the simulated basilar membrane induction information of the human voice from a plurality of consecutively-time audio frames to be recognized through variational mode decomposition and gammatone frequency cepstral coefficient filters; A determination device configured to input the extracted plurality of features to be recognized into a speech category classification neural network so as to determine a plurality of posterior probabilities of the plurality of audio frames to be recognized having the human voice from the plurality of audio frames to be recognized; A recognition device configured to recognize one or more continuous speech segments of the human voice according to the plurality of posterior probabilities of the plurality of audio frames to be recognized having the human voice determined; 13. A device for recognizing a continuous speech segment of a human voice, comprising: A memory for storing instructions; A processor for executing the instructions in the memory and executing the method according to any one of claims 1-11.
14. A non-transitory storage medium having instructions stored thereon, Among them, wherein when the instructions are executed by a processor, the processor is caused to execute the method according to any one of claims 1-11.
15. A computer program product having instructions stored thereon, Among them, wherein when the instructions are executed by a processor, the processor is caused to execute the method according to any one of claims 1-11.
Citation Information
Patent Citations
Voice keyword identification method and apparatus based on deep neural network
CN105679316A
Continuous speech recognition method and continuous speech recognition system
CN106157953A
Voice recognition method, device and equipment and computer readable medium
CN111798853A
Many-to-one voice conversion method based on Gammatone frequency cepstrum coefficient
CN114283822A
Audio recognition method and device, computing equipment and storage medium
CN114360513A