Speech Enhancement Method, Device and Robot for Robot
By performing multi-step enhanced processing of the voice information collected by robot pets, including downsampling, excitation spectrum generation, noise reduction and spectrum broadening, the poor voice quality caused by robot pets due to noise and distance in voice interaction is solved, and the accuracy of speech recognition is improved.
Patent Information
- Application Number
- CN202110192265.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-02-18
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2041-02-18
AI Technical Summary
When robot pets interact with human voice, their speech recognition accuracy is low due to noise interference caused by constant movement and poor speech quality in long-distance interaction.
By acquiring the voice information collected by the robot, the first enhancement is performed to generate initial enhanced voice information, and the channel parameters are generated based on the voice information, and the second enhancement is performed to generate enhanced voice information. The method includes steps such as downsampling, excitation spectrum generation, noise reduction, pitch estimation and spectrum broadening, with the aim of restoring low-frequency information polluted by noise and high-frequency information complementary attenuation.
By enhancing the voice information collected by the robot, the quality of voice information is improved, and the accuracy of voice recognition is improved, solving the problem of poor voice information quality.
Smart Images

Figure CN114974275B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of speech processing technology, and in particular, to a speech enhancement method, device, and robot for a robot. Background Art
[0002] With the continuous development of robots, robot pets are becoming more and more popular. However, robot pets, such as legged robots, are constantly moving during the voice interaction with people. Different from traditional fixed intelligent devices (such as smart speakers), due to the continuous movement of the robot pet, it will generate a lot of noise by itself, such as the noise of the drive motor, the mechanical transmission noise of the joint part during the movement, etc. These noises will cause great interference to the speech recognition.
[0003] In addition, since the robot is always in a moving state, it may be very far away from the user. At this time, due to the influence of its own noise and environmental noise, the voice quality of the user will be poor, resulting in the robot being unable to accurately respond to the user's instructions. Summary of the Invention
[0004] This application provides a speech enhancement method, device, and robot for a robot to at least solve the problem of poor quality of speech information in the related art. The technical solution of this application is as follows:
[0005] According to the first aspect of the embodiments of this application, a speech enhancement method for a robot is provided, including:
[0006] Obtain the speech information collected by the robot;
[0007] Perform a first enhancement on the speech information to generate initial enhanced speech information, and generate the vocal tract parameters of the speech information according to the speech information; and
[0008] Perform a second enhancement on the basis of the vocal tract parameters and the initial enhanced speech information to generate enhanced speech information.
[0009] Optionally, the performing a first enhancement on the speech information to generate initial enhanced speech information includes:
[0010] Downsample the speech information to generate first speech information;
[0011] Generate the excitation spectrum corresponding to the first speech information according to the first speech information;
[0012] Perform noise reduction on the first speech information to generate the first speech information after noise reduction;
[0013] Generate the initial enhanced speech information according to the excitation spectrum and the first speech information after noise reduction.
[0014] Optionally, generating an excitation spectrum corresponding to the first voice information according to the first voice information includes:
[0015] Performing non-negative matrix factorization (NMF) on the first voice information to generate a voice frame probability of the first voice information;
[0016] Performing a first pitch estimation on the first voice information to generate an initial pitch estimation value of the first voice information;
[0017] Performing a second pitch estimation on the initial pitch estimation value according to the voice frame probability to generate a target pitch estimation value;
[0018] Generating the excitation spectrum according to the target pitch estimation value.
[0019] Optionally, performing non-negative matrix factorization (NMF) on the first voice information to generate a voice frame probability of the first voice information includes:
[0020] Performing a Fourier transform on each frame of voice signal in the first voice information to generate a spectral matrix of each frame of voice signal;
[0021] Based on a pre-acquired voice basis matrix, performing a decomposition operation on the spectral matrix of each frame of voice signal to obtain a voice activation matrix corresponding to the spectral matrix of each frame of voice signal and an updated voice basis matrix;
[0022] Based on a pre-acquired interference basis matrix, performing a decomposition operation on the spectral matrix of each frame of voice signal to obtain an interference activation matrix corresponding to the spectral matrix of each frame of voice signal and an updated interference basis matrix;
[0023] Based on the updated voice basis matrix and the updated interference basis matrix, performing a repeated decomposition operation on the spectral matrix of each frame of voice signal to obtain a target interference activation matrix and a target voice activation matrix corresponding to the spectral matrix of each frame of voice signal;
[0024] Determining the voice frame probability of each frame of voice signal in the first voice information according to the target interference activation matrix and the target voice activation matrix.
[0025] Optionally, the pre-acquired voice basis matrix includes the following steps:
[0026] Obtaining a sample voice signal;
[0027] Framing the sample voice signal to obtain N sub-signals, where N is a positive integer;
[0028] Perform Fourier transform on each frame of the N-frame sub-signals to determine the spectral values of F frequency points included in each frame of signal, where F is a positive integer;
[0029] Generate a first sample spectral matrix corresponding to the sample speech signal, which includes F rows and N columns, based on the spectral values of F frequency points included in each frame of sub-signal;
[0030] Perform clustering on the first sample spectral matrix to generate the pre-acquired speech basis matrix including K column vectors, where K is a positive integer less than or equal to N.
[0031] Optionally, the steps for pre-acquiring the interference basis matrix include:
[0032] Acquire a sample interference signal, which is generated when the robot moves;
[0033] Frame the sample interference signal to obtain M-frame sub-signals, where M is a positive integer;
[0034] Perform Fourier transform on each frame of the M-frame sub-signals to determine the spectral values of G frequency points included in each frame of signal, where G is a positive integer;
[0035] Generate a second sample spectral matrix corresponding to the sample speech signal, which includes G rows and M columns, based on the spectral values of G frequency points included in each frame of sub-signal;
[0036] Perform clustering on the second sample spectral matrix to generate the pre-acquired interference basis matrix including L column vectors, where L is a positive integer less than or equal to M.
[0037] Optionally, the vocal tract parameter is the linear prediction coding residual, and the steps for generating the vocal tract parameter of the speech information based on the speech information include:
[0038] Perform linear prediction coding LPC envelope estimation on the speech information to generate the linear prediction coding residual.
[0039] Optionally, the steps for performing second enhancement on the initial enhanced speech information based on the vocal tract parameter to generate the enhanced speech information include:
[0040] Generate the unvoiced information of the speech information based on the initial enhanced speech information, the target pitch estimation value, and the speech frame probability;
[0041] Generate the whispered information of the speech information based on the speech frame probability, the linear prediction coding residual, and the random noise;
[0042] Perform first voice synthesis based on the voiced information and unvoiced information of the voice information to generate first synthesized voice information;
[0043] Perform spectral broadening on the first synthesized voice information to generate voice information after spectral broadening;
[0044] Perform second voice synthesis on the first synthesized voice information and the voice information after spectral broadening to generate second synthesized voice information;
[0045] Perform upsampling on the second synthesized voice information to generate the enhanced voice information.
[0046] Optionally, the method further includes:
[0047] Determine a control instruction for the robot in response to an identification result of the enhanced voice information;
[0048] Control the robot according to the control instruction.
[0049] According to a second aspect of the embodiments of the present application, there is provided a voice enhancement device for a robot, including:
[0050] An acquisition module, configured to acquire voice information collected by the robot;
[0051] A first processing module, configured to perform first enhancement on the voice information to generate initial enhanced voice information, and generate vocal tract parameters of the voice information; and
[0052] A second processing module, configured to perform second enhancement on the basis of the vocal tract parameters and the initial enhanced voice information to generate enhanced voice information.
[0053] Optionally, the first processing module includes:
[0054] A downsampling unit, configured to perform downsampling on the voice information to generate first voice information;
[0055] A first generation unit, configured to generate an excitation spectrum corresponding to the first voice information according to the first voice information;
[0056] A noise reduction unit, configured to perform noise reduction on the first voice information to generate first voice information after noise reduction;
[0057] A second generation unit, configured to generate the initial enhanced voice information according to the excitation spectrum and the first voice information after noise reduction.
[0058] Optionally, the first generation unit is specifically configured to:
[0059] Perform non - negative matrix factorization (NMF) on the first voice information to generate the voice frame probability of the first voice information;
[0060] Perform a first pitch estimation on the first voice information to generate an initial pitch estimation value of the first voice information;
[0061] According to the voice frame probability, perform a second pitch estimation on the initial pitch estimation value to generate a target pitch estimation value;
[0062] Generate the excitation spectrum according to the target pitch estimation value.
[0063] Optionally, the first generation unit is specifically configured to:
[0064] Perform a Fourier transform on each frame of voice signal in the first voice information to generate a spectral matrix of each frame of voice signal;
[0065] Based on a pre - acquired voice basis matrix, perform a decomposition operation on the spectral matrix of each frame of voice signal to obtain a voice activation matrix corresponding to the spectral matrix of each frame of voice signal and an updated voice basis matrix;
[0066] Based on a pre - acquired interference basis matrix, perform a decomposition operation on the spectral matrix of each frame of voice signal to obtain an interference activation matrix corresponding to the spectral matrix of each frame of voice signal and an updated interference basis matrix;
[0067] Based on the updated voice basis matrix and the updated interference basis matrix, perform a repeated decomposition operation on the spectral matrix of each frame of voice signal to obtain a target interference activation matrix and a target voice activation matrix corresponding to the spectral matrix of each frame of voice signal;
[0068] Determine the voice frame probability of each frame of voice signal in the first voice information according to the target interference activation matrix and the target voice activation matrix.
[0069] Optionally, the device further includes:
[0070] The above - mentioned acquisition module is further configured to acquire a sample voice signal;
[0071] A framing module for framing the sample voice signal to obtain N sub - signals, where N is a positive integer;
[0072] A determination module for performing a short - time Fourier transform on each of the N sub - signals to determine the spectral values of F frequency points included in each frame of signal, where F is a positive integer;
[0073] A generation module, configured to generate a first sample spectrum matrix corresponding to the sample speech signal, including F rows and N columns, according to spectral values of F frequency points included in each frame of sub-signals; perform clustering on the first sample spectrum matrix to generate the pre-acquired speech basis matrix including K column vectors, where K is a positive integer less than or equal to N.
[0074] Optionally, the above-mentioned acquisition module is configured to acquire a sample interference signal generated when the target device moves;
[0075] The above-mentioned frame division module is configured to perform frame division on the sample interference signal to obtain M frames of sub-signals, where M is a positive integer;
[0076] The above-mentioned determination module is configured to perform short-time Fourier transform on each frame of sub-signals among the M frames of sub-signals to determine spectral values of G frequency points included in each frame of signal, where G is a positive integer;
[0077] The above-mentioned generation module is configured to generate a second sample spectrum matrix corresponding to the sample speech signal, including G rows and M columns, according to spectral values of G frequency points included in each frame of sub-signals; perform clustering on the second sample spectrum matrix to generate the pre-acquired interference basis matrix including L column vectors, where L is a positive integer less than or equal to M.
[0078] Optionally, the vocal tract parameter is linear prediction coding residual, and the generating the vocal tract parameter of the speech information according to the speech information includes:
[0079] Performing linear prediction coding LPC envelope estimation on the speech information to generate linear prediction coding residual.
[0080] Optionally, the second processing module is specifically configured to:
[0081] Generate the voiced sound information of the speech information according to the initial enhanced speech information, the target pitch estimate value, and the speech frame probability;
[0082] Generate the unvoiced sound information of the speech information according to the speech frame probability, the linear prediction coding residual, and random noise;
[0083] Perform first speech synthesis according to the voiced sound information and the unvoiced sound information of the speech information to generate first synthesized speech information;
[0084] Perform spectral broadening on the first synthesized speech information to generate speech information after spectral broadening;
[0085] Perform second speech synthesis on the first synthesized speech information and the speech information after spectral broadening to generate second synthesized speech information;
[0086] Upsample the second synthesized speech information to generate the enhanced speech information.
[0087] Optionally, the apparatus further includes:
[0088] A control module, configured to determine a control instruction for the robot in response to an identification result of the enhanced speech information; and control the robot according to the control instruction.
[0089] According to a third aspect of the embodiments of the present application, there is provided a robot, including: at least one processor; and
[0090] a memory communicatively connected to the at least one processor; wherein,
[0091] the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method described in the first aspect.
[0092] According to a fourth aspect of the embodiments of the present application, there is provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to execute the method described in the first aspect.
[0093] According to a fifth aspect of the embodiments of the present application, there is provided a computer program product, including a computer program which, when executed by a processor, implements the method described in the first aspect.
[0094] The technical solutions provided by the embodiments of the present application at least bring the following beneficial effects:
[0095] Obtain the speech information collected by the robot, generate the vocal tract parameters of the speech information according to the speech information, perform first enhancement on the speech information to generate the initial enhanced speech information, and further perform second enhancement according to the vocal tract parameters and the initial enhanced speech information to generate the enhanced speech information. By enhancing the speech information collected by the robot, the quality of the speech information is improved, and thus the accuracy of speech recognition is improved.
[0096] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0097] The accompanying drawings herein are incorporated into the specification and constitute a part of the specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application, and do not constitute an improper limitation of the present application.
[0098] Figure 1It is a schematic flowchart of a voice enhancement method for a robot provided by an embodiment of the present application;
[0099] Figure 2 It is a schematic flowchart of another voice enhancement method for a robot provided by an embodiment of the present application;
[0100] Figure 3 It is a flowchart block diagram for generating initial enhanced voice information provided by an embodiment of the present application;
[0101] Figure 4 It is a schematic flowchart of another voice enhancement method for a robot provided by an embodiment of the present application;
[0102] Figure 5 It is a block diagram of bandwidth expansion provided by an embodiment of the present application;
[0103] Figure 6 It is a schematic structural diagram of a voice enhancement device for a robot provided by an embodiment of the present application;
[0104] Figure 7 It is a schematic block diagram of a robot 800 provided by an embodiment of the present application. Detailed implementation manners
[0105] In order to enable those of ordinary skill in the art to better understand the technical solutions of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings.
[0106] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present application described herein can be implemented in an order different from those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present application. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.
[0107] In the scenario where a robot interacts with a human, the robot is constantly moving and generates a lot of noise on its own. For example, the sound of motor drive, and the noise generated by the motor is usually at 300HZ and below. In many cases, the low-frequency information of the collected voice information is covered by the noise. At the same time, during the interaction with the robot, the distance is usually far, such as 3-5 meters, that is, far-field voice interaction. During far-field voice interaction, the high-frequency components attenuate severely, resulting in a low accuracy of speech recognition. To solve this problem, the present application proposes a voice enhancement method for a robot, which obtains the voice information collected by the robot, performs a first enhancement on the voice information to generate initial enhanced voice information, and generates the vocal tract parameters of the voice information according to the voice information. Furthermore, a second enhancement is performed according to the vocal tract parameters and the initial enhanced voice information to generate enhanced voice information. Through the enhancement process, the low-frequency voice signal contaminated by noise and the high-frequency voice signal with severe attenuation are restored, the quality of the voice signal is improved, and thus the accuracy of speech recognition is improved.
[0108] Figure 1 is a schematic flowchart of a voice enhancement method for a robot provided by an embodiment of the present application, as Figure 1 shown, the method includes the following steps.
[0109] Step 101, obtain the voice information collected by the robot, and generate the vocal tract parameters of the voice information according to the voice information.
[0110] In an embodiment of the present application, the robot can be a legged robot, such as a quadruped robot or a biped robot, etc. Among them, a microphone can be set in the robot, and the collected voice information is obtained according to the microphone. This voice information is the originally collected voice information, and there is interference in the low-frequency part and attenuation in the high-frequency part of this voice information, and the voice information is distorted and of poor quality.
[0111] In an embodiment of the present application, the vocal tract parameters include linear predictive coding residuals, and linear predictive coding (LPC) envelope estimation is performed on the first voice information to generate linear predictive coding residuals.
[0112] Step 102, perform a first enhancement on the voice information to generate initial enhanced voice information.
[0113] In one embodiment of the present application, the collected voice information is downsampled to generate first voice information. Among them, since the high-frequency part of the voice is far away, the attenuation of the collected voice information is relatively large and can be ignored. For example, the sampling rate of the voice information is reduced from 16KHz to 8KHz, that is to say, subsequent enhancement processing of the voice information is based on an 8K sampling rate. And an excitation spectrum corresponding to the first voice information is generated according to the first voice information. Further, the first voice information is denoised to generate the first voice information after denoising. Furthermore, an initial enhanced voice information is generated according to the excitation spectrum and the first voice information after denoising.
[0114] Step 103, perform second enhancement according to the vocal tract parameters and the initial enhanced voice information to generate enhanced voice information.
[0115] In one embodiment of the present application, according to the initial enhanced voice information, the target pitch estimate value, and the voice frame probability, the voiced feature of the voice information is determined, and according to the voice frame probability, the linear prediction coding residual, and the random noise, the unvoiced feature of the voice information is determined. Furthermore, according to the voiced feature and the unvoiced feature of the voice information, first voice synthesis is performed to generate first synthesized voice information. Furthermore, the first synthesized voice information is spectrally broadened to obtain the broadened voice information. Through spectral broadening, the generation of the high-frequency part of the voice with relatively large attenuation is realized.
[0116] Further, second voice synthesis is performed on the first synthesized voice information and the broadened voice information to generate second synthesized voice information. Among them, the second synthesized voice information combines the first synthesized voice information and the second synthesized voice information to obtain enhanced voice information containing noise-free low-frequency information and spectrally broadened frequency band information;
[0117] Further, the second synthesized voice information is upsampled to generate enhanced voice information. The enhanced voice information contains noise-free low-frequency voice information and non-attenuated high-frequency voice information. That is to say, the enhanced voice information restores the contaminated low-frequency information in the voice information collected by the robot and complements the attenuated high-frequency information to generate voice information with enhanced voice containing the entire frequency band. At the same time, the quality of the voice information is relatively high, improving the voice recognition effect.
[0118] In the voice enhancement method for a robot according to an embodiment of the present application, voice information collected by the robot is obtained, the voice information is first enhanced to generate initial enhanced voice information, voice channel parameters of the voice information are generated, and second enhancement is performed based on the channel parameters and the initial enhanced voice information to generate enhanced voice information. By enhancing the voice information collected by the robot, contaminated low-frequency information can be restored, and attenuated high-frequency information can be complemented, improving the quality of the voice information and thus providing the accuracy of voice recognition.
[0119] Based on the above embodiment, this embodiment further provides an implementation method. After generating the enhanced language information, the following steps may further be included:
[0120] In response to the recognition result of the enhanced voice information, determine the control instruction of the robot;
[0121] According to the control instruction, control the robot.
[0122] In this embodiment, according to the generated enhanced voice information, semantic recognition is performed on the enhanced voice information, and corresponding control instructions are determined according to the recognition result to control the robot to perform corresponding actions. For example, if the semantic recognition result of the generated enhanced voice information is "the floor is too dirty", a control instruction for controlling the robot to perform cleaning or mopping is generated, so that the robot executes the cleaning or mopping instruction to achieve human-machine interaction.
[0123] Based on the above embodiment, this embodiment provides another voice enhancement method for a robot. Figure 2 As a schematic flowchart of another voice enhancement method for a robot according to an embodiment of the present application, it specifically illustrates how to generate initial enhanced voice information. As Figure 2 shown, step 102 includes the following steps:
[0124] Step 201, perform downsampling on the voice information to generate first voice information.
[0125] Among them, Figure 3 is a flowchart for generating initial enhanced voice information according to an embodiment of the present application, which will be described in combination with Figure 3 for illustration.
[0126] In this embodiment, due to the relatively long interaction distance, the high-frequency components of the voice information collected by the robot are severely attenuated, and the high-frequency part of the voice can be ignored. To reduce the amount of data processing, downsampling is performed on the voice information collected by the robot. For example, the sampling rate is reduced from 16KHz to 8KHz, and the corresponding frequency spectrum is 0 - 4Khz, that is, the frequency spectrum corresponding to the first voice information is 0 - 4KHz.
[0127] Step 202: Generate an excitation spectrum corresponding to the first voice message according to the first voice message.
[0128] As Figure 3 shown, perform a first pitch estimation on the first voice message to generate an initial pitch estimation value of the first voice message. Among them, the initial gene estimation value is the estimated gene period. As a possible implementation, frame the first voice message to obtain multiple frames of sub-messages. For each frame of sub-message, use the autocorrelation function to determine the possible initial pitch estimation value of the corresponding frame. Among them, the initial pitch estimation value corresponding to each frame of sub-message is multiple possible values, and the pitch estimation value is the pitch period pitch.
[0129] Furthermore, in an implementation manner of the embodiment of the present application, perform non-negative matrix factorization (NMF) on the first voice message to generate the voice frame probability of the first voice message. The voice frame probability indicates the probability that each frame in the first voice message is a voice frame. According to the value of this probability, it can be determined whether the corresponding voice frame is a voice frame or a non-voice frame, that is, a noise frame.
[0130] Further, according to the voice frame probability, perform a second pitch estimation on the initial pitch estimation value to generate a target pitch estimation value. As a possible implementation, for each frame of sub-message determined to be a voice frame, according to the multiple initial pitch estimation values of each frame of sub-message determined to be a voice frame, since the pitch period is continuous between adjacent frames of sub-messages and there will be no jumps, the discontinuous initial pitch estimation values can be deleted. According to the remaining continuous initial pitch estimation values, determine the target pitch estimation value. As an implementation manner, if the remaining initial pitch estimation value is one, then use this initial pitch estimation value as the target pitch estimation value; as another implementation manner, if the remaining initial pitch estimation values are multiple, then use the average value of the multiple pitch estimation values as the target pitch estimation value, that is, the target gene period is determined.
[0131] Further, generate an excitation spectrum according to the target pitch estimation value. Among them, the excitation spectrum is a pulse train with a pulse period of the target pitch estimation value, which is used to simulate the excitation signal in the voice generation process. The frequency of the pulse train is the reciprocal of the target pitch estimation value, that is, the reciprocal of the target gene period.
[0132] Step 203: Denoise the first voice message to generate the denoised first voice message.
[0133] Step 204: Generate an initial enhanced voice message according to the excitation spectrum and the denoised first voice message.
[0134] In this embodiment, according to the generated excitation spectrum and the denoised first voice message, the low-frequency signal contaminated by noise can be restored, that is, the initially enhanced voice message.
[0135] In this embodiment, according to the sound source characteristics of the extracted voice information, a first pitch estimation is performed to generate an initial pitch estimation value of the first voice information. Furthermore, according to the determined probability of the voice frame, a second pitch estimation is performed on the initial pitch estimation value to generate a target pitch estimation value. Furthermore, according to the generated excitation spectrum and the first voice information after noise reduction, the low-frequency signal contaminated by noise can be restored, that is, the initially enhanced voice information.
[0136] Based on the previous embodiment, this embodiment provides a method for determining the probability of a voice frame, which includes the following steps:
[0137] Perform a short-time Fourier transform on each frame of voice signal in the first voice information to generate a spectral matrix for each frame of voice signal.
[0138] In this embodiment, for each frame of voice signal obtained by frame-dividing the first voice information, a short-time Fourier transform is performed to generate a spectral matrix for each frame of voice signal, where the spectral matrix includes an amplitude spectral matrix or a power spectral matrix. For example, based on the amplitude spectrum of the sound signal to be detected in the frequency domain, an amplitude spectral matrix can be determined, or based on the power spectrum of the sound signal to be detected in the frequency domain, a power spectral matrix can be determined.
[0139] Based on the pre-obtained voice basis matrix, perform a decomposition operation on the spectral matrix of each frame of voice signal to obtain the voice activation matrix corresponding to the spectral matrix of each frame of voice signal and the updated voice basis matrix.
[0140] Among them, the pre-obtained voice basis matrix is generated through the following method:
[0141] Obtain a sample voice signal, where the sample voice signal is a voice signal containing typical interference obtained in advance. Perform frame division on the sample voice signal to obtain N sub-signals, where N is a positive integer. Furthermore, perform a short-time Fourier transform on each of the N sub-signals to determine the spectral values of F frequency points included in each frame of signal, where F is a positive integer. According to the spectral values of F frequency points included in each frame of sub-signal, generate a first sample spectral matrix corresponding to the sample voice signal, which includes F rows and N columns. Perform clustering on the first sample spectral matrix to generate the pre-obtained voice basis matrix, where the pre-obtained voice basis matrix includes K column vectors, where K is a positive integer less than or equal to N.
[0142] Based on the pre-obtained interference basis matrix, perform a decomposition operation on the spectral matrix of each frame of voice signal to obtain the interference activation matrix corresponding to the spectral matrix of each frame of voice signal and the updated interference basis matrix.
[0143] In this embodiment, the pre-obtained interference basis matrix can be obtained through the following method:
[0144] Obtain a sample interference signal, which is generated when the target device is in motion. For example, during the interaction process, a robot is in a moving state. Thus, during the movement of the robot, obtain the sample interference signal. Frame the sample interference signal to obtain M sub-signals, where M is a positive integer. Furthermore, perform a short-time Fourier transform on each of the M sub-signals to determine the spectral values of G frequency points included in each frame of the signal, where G is a positive integer. Further, based on the spectral values of G frequency points included in each frame of the sub-signal, generate a second sample spectral matrix corresponding to the sample speech signal, which includes G rows and M columns. Cluster the second sample spectral matrix to generate a pre-obtained interference basis matrix, where the pre-obtained interference basis matrix includes L column vectors, and L is a positive integer less than or equal to M.
[0145] Based on the updated speech basis matrix and the updated interference basis matrix, perform repeated decomposition operations on the spectral matrix of each frame of the speech signal to obtain the target interference activation matrix and the target speech activation matrix corresponding to the spectral matrix of each frame of the speech signal.
[0146] Among them, the decomposition operation is the decomposition operation of non-negative matrix factorization (NMF). Through continuous iterative operations, the target interference activation matrix and the target speech activation matrix are solved.
[0147] In this embodiment, after obtaining the updated speech basis matrix, perform a decomposition operation on the spectral matrix of each frame of the speech signal based on the updated speech basis matrix to obtain the speech activation matrix corresponding to the corresponding spectral matrix and the updated speech basis matrix. And, after obtaining the updated interference basis matrix, perform a decomposition operation on the corresponding spectral matrix based on the updated interference basis matrix to obtain the interference activation matrix corresponding to the spectral matrix and the updated interference basis matrix.
[0148] Taking the spectral matrix V of a frame of speech information as an example for illustration:
[0149] As an example, the spectral matrix is V, and the pre-obtained speech basis matrix is W1. When performing the decomposition operation for the first time, perform the decomposition operation on the spectral matrix V based on W1, and the updated speech basis matrix is W2. Furthermore, when performing the decomposition operation for the second time, perform the decomposition operation on the spectral matrix V based on W2, and the updated speech basis matrix is W3, and so on.
[0150] As another example, the spectral matrix is V, and the pre-obtained interference basis matrix is W1'. When performing the decomposition operation for the first time, perform the decomposition operation on the spectral matrix V based on W1', and the updated interference basis matrix is W2'. Furthermore, when performing the decomposition operation for the second time, perform the decomposition operation on the spectral matrix V based on W2', and the updated interference basis matrix is W3', and so on.
[0151] In this embodiment, the iteration conditions can be preset. When the iteration termination conditions are met, the currently decomposed voice activation matrix is used as the target voice activation matrix, and the currently decomposed interference activation matrix is used as the target interference activation matrix. When the iteration termination conditions are not met, an operation of repeatedly decomposing the spectral matrix based on the updated voice basis matrix and the updated interference basis matrix is performed.
[0152] Among them, the iteration termination conditions may include the number of iterations, cost functions, etc. The cost function includes, for example, the Euclidean distance, etc. As an example, when the number of iterations meets the preset number, it is determined that the iteration termination conditions are met. By repeatedly decomposing the spectral matrix, the target voice basis matrix, the target voice activation matrix, the target interference basis matrix, and the target interference activation matrix corresponding to the spectral matrix are generated.
[0153] Similarly, the target voice activation matrix and the target voice activation matrix corresponding to the spectral matrix of each frame of voice signal can be obtained. According to the target interference activation matrix and the target voice activation matrix, the voice frame probability of each frame of voice signal in the first voice information is determined.
[0154] In one implementation manner of the embodiment of the present application, the target voice activation matrix and the target interference activation matrix of each frame of voice signal are obtained. According to the target voice activation matrix and the target interference activation matrix of each frame of voice signal, the voice interference ratio of each frame of voice signal is determined. According to the voice interference ratio, the voice frame probability of each frame of voice signal is determined, improving the accuracy of voice frame determination. Among them, the voice frame probability can indicate whether each frame of voice signal is a voice frame or a non-voice frame.
[0155] Based on the above embodiment, this embodiment provides another method for enhancing the voice of a robot. Figure 4 It is a schematic flowchart of another method for enhancing the voice of a robot provided by the embodiment of the present application. As Figure 4 shown, the above step 103 includes the following steps:
[0156] Step 401, generate the voiced sound information of the voice information according to the initial enhanced voice information, the target pitch estimate value, and the voice frame probability.
[0157] In one implementation manner of this embodiment, the initial enhanced voice information, the target pitch estimate value, and the voice frame probability are input into the inverse filter of linear prediction, so that the inverse filter generates a pulse train according to the initial enhanced voice information, the target pitch estimate value, and the voice frame probability to generate the voiced sound information in the voice information.
[0158] Step 402, generate the unvoiced sound information of the voice information according to the voice frame probability, the linear prediction coding residual, and the random noise.
[0159] Among them, the random noise is the set comfortable Gaussian white noise, which is used to simulate the weak sound in vocalization.
[0160] In an example of this embodiment, when it is determined to be a speech frame, random noise is used as the unvoiced excitation, and according to the linear prediction coding residual corresponding to the vocal tract parameters, the weak sound information of the speech information is generated.
[0161] In another example of this embodiment, the speech frame probability, the linear prediction coding residual, and the random noise are input into the inverse filter of linear prediction, so that the inverse filter determines the random noise corresponding white noise as unvoiced sound to generate the corresponding unvoiced sound information.
[0162] Step 403: Perform first speech synthesis according to the voiced sound information and the weak sound information of the speech information to generate first synthesized speech information.
[0163] In this embodiment, according to the voiced sound information and the weak sound information of the generated speech information, a speech synthesizer is called to synthesize and obtain the first synthesized speech information.
[0164] Step 404: Perform spectral broadening on the first synthesized speech information to generate broadened speech information.
[0165] For example, the spectrum corresponding to the first synthesized speech information obtained by synthesis is 0 - 4KHZ, and the spectrum corresponding to the broadened speech information generated by spectral broadening is 4 - 8KHZ.
[0166] As a possible implementation manner, as Figure 5 shown, bandwidth detection is performed on the first synthesized speech information. Among them, the bandwidth detection is composed of 8 band - pass filters with a bandwidth of 1KHz. By comparing the energy magnitudes of the 8 band - pass filters, the bandwidth of the first synthesized speech information is determined. Then, the first synthesized speech information with the determined bandwidth is input into the sub - band filter for pre - processing, so that the bandwidth of the broadened speech information obtained after subsequent non - linear processing is constant. The non - linear processing is, for example, the conventional non - linear processing of full - wave rectification. Thus, the first synthesized speech information with a spectrum of 0 - 4KHz is spectrally broadened to generate broadened speech information with a spectrum of 4 - 8KHz. Further, through the gain adjustment of post - processing, the gain of the broadened speech information is increased, improving the effect of the speech information.
[0167] It should be noted that, in order to avoid distortion of the generated broadened speech information, delay alignment is performed to align the first synthesized speech information and the broadened speech information to avoid distortion of the broadened speech information.
[0168] Step 405: Perform second voice synthesis on the first synthesized voice information and the broadened voice information to generate second synthesized voice information.
[0169] In the embodiment of the present application, the first synthesized voice information and the broadened voice information are combined to obtain the second synthesized voice information. For example, the spectrum corresponding to the first synthesized voice information is 0 - 4KHz, the spectrum corresponding to the broadened voice information is 4 - 8KHz, and the spectrum corresponding to the second synthesized voice information obtained after combination is 0 - 8KHz. The second synthesized voice information contains the restored high-frequency information, that is, in the voice information number corresponding to the collected voice information, the high-frequency information with a large attenuation due to distance.
[0170] Step 406: Upsample the second synthesized voice information to generate enhanced voice information.
[0171] In this embodiment, since the spectrum of the voice signal collected by the robot corresponds to 0 - 16Khz, while the spectrum corresponding to the second synthesized voice information is 0 - 8KHz, the enhanced voice information obtained by upsampling, that is, the interpolation method, has a spectrum corresponding to 0 - 16KHz, realizing the generation of the voice information in the full frequency band, restoring the contaminated low-frequency information and the high-frequency information with a large attenuation, and improving the generation quality of the voice information.
[0172] In the voice enhancement method for a robot according to the embodiment of the present application, the voice information collected by the robot is obtained, the voice information is first enhanced to generate initial enhanced voice information, the vocal tract parameters of the voice information are generated according to the voice information, and the second enhancement is performed according to the vocal tract parameters and the initial enhanced voice information to generate enhanced voice information. By enhancing the voice information collected by the robot, the contaminated low-frequency information is restored, and the attenuated high-frequency information is complemented, improving the quality of the voice information, and further providing the accuracy of voice recognition.
[0173] To implement the above embodiment, the embodiment of the present application provides a voice enhancement device for a robot.
[0174] Figure 6 Shown in the following is a structural schematic diagram of a voice enhancement device for a robot provided by the embodiment of the present application. Figure 6 As shown, the device includes the following steps:
[0175] An acquisition module 61, configured to acquire the voice information collected by the robot and generate the vocal tract parameters of the voice information according to the voice information;
[0176] A first processing module 62, configured to perform first enhancement on the voice information to generate initial enhanced voice information.
[0177] The second processing module 63 is used to perform a second enhancement according to the vocal tract parameters and the initial enhanced speech information to generate enhanced speech information.
[0178] Further, in an implementation of the embodiment of the present application, the first processing module 62 includes:
[0179] A downsampling unit, configured to downsample the voice information to generate first voice information;
[0180] A first generating unit, configured to generate an excitation spectrum corresponding to the first voice information according to the first voice information;
[0181] A noise reduction unit, configured to perform noise reduction on the first voice information to generate first voice information after noise reduction;
[0182] The second generating unit is used to generate the initial enhanced speech information according to the excitation spectrum and the first speech information after noise reduction.
[0183] In one implementation of the embodiment of the present application, the first generating unit is specifically configured to:
[0184] Performing non-negative matrix factorization (NMF) on the first speech information to generate a speech frame probability of the first speech information;
[0185] Performing a first pitch estimation on the first speech information to generate an initial pitch estimation value of the first speech information;
[0186] Performing a second pitch estimation on the initial pitch estimation value according to the speech frame probability to generate a target pitch estimation value;
[0187] The excitation spectrum is generated according to the target pitch estimation value.
[0188] In one implementation of the embodiment of the present application, the first generating unit is specifically configured to:
[0189] Performing short-time Fourier transform on each frame of speech signal in the first speech information to generate a spectrum matrix of each frame of speech signal;
[0190] Based on the pre-acquired speech basis matrix, a decomposition operation is performed on the spectrum matrix of each frame of speech signal to obtain a speech activation matrix and an updated speech basis matrix corresponding to the spectrum matrix of each frame of speech signal;
[0191] Based on the pre-acquired interference basis matrix, a decomposition operation is performed on the spectrum matrix of each frame of speech signal to obtain an interference activation matrix and an updated interference basis matrix corresponding to the spectrum matrix of each frame of speech signal;
[0192] Based on the updated speech basis matrix and the updated interference basis matrix, perform repeated decomposition operations on the spectral matrix of each frame of speech signal to obtain a target interference activation matrix and a target speech activation matrix corresponding to the spectral matrix of each frame of speech signal;
[0193] Determine the speech frame probability of each frame of speech signal in the first speech information according to the target interference activation matrix and the target speech activation matrix.
[0194] In an implementation manner of the embodiment of the present application, the device further includes:
[0195] The above-mentioned acquisition module 61 is further configured to acquire a sample speech signal;
[0196] The framing module is configured to frame the sample speech signal to obtain N sub-signals, where N is a positive integer;
[0197] The determination module is configured to perform short-time Fourier transform on each frame of the N sub-signals to determine the spectral values of F frequency points included in each frame of signal, where F is a positive integer;
[0198] The generation module is configured to generate a first sample spectral matrix corresponding to the sample speech signal including F rows and N columns according to the spectral values of F frequency points included in each frame of sub-signal; perform clustering on the first sample spectral matrix to generate the pre-acquired speech basis matrix including K column vectors, where K is a positive integer less than or equal to N.
[0199] In an implementation manner of the embodiment of the present application,
[0200] The above-mentioned acquisition module 61 is configured to acquire a sample interference signal, which is generated when the target device moves;
[0201] The above-mentioned framing module is configured to frame the sample interference signal to obtain M sub-signals, where M is a positive integer;
[0202] The above-mentioned determination module is configured to perform short-time Fourier transform on each frame of the M sub-signals to determine the spectral values of G frequency points included in each frame of signal, where G is a positive integer;
[0203] The above-mentioned generation module is configured to generate a second sample spectral matrix corresponding to the sample speech signal including G rows and M columns according to the spectral values of G frequency points included in each frame of sub-signal; perform clustering on the second sample spectral matrix to generate the pre-acquired interference basis matrix including L column vectors, where L is a positive integer less than or equal to M.
[0204] In an implementation of the embodiment of the present application, the vocal tract parameter is a linear prediction coding residual, and the first processing module 62 is specifically used to:
[0205] Performing linear predictive coding (LPC) envelope estimation on the speech information to generate a linear predictive coding residual.
[0206] In one implementation of the embodiment of the present application, the second processing module 63 is specifically used to:
[0207] Generate voiced sound information of the speech information according to the initial enhanced speech information, the target pitch estimation value and the speech frame probability; generate unstressed sound information of the speech information according to the speech frame probability, the linear prediction coding residual and random noise; perform first speech synthesis according to the voiced sound information of the speech information and the unstressed sound information of the speech information to generate first synthesized speech information; perform spectrum widening on the first synthesized speech information to generate widened speech information; perform second speech synthesis on the first synthesized speech information and the widened speech information to generate second synthesized speech information; upsample the second synthesized speech information to generate the enhanced speech information.
[0208] As a possible implementation, the device further includes:
[0209] The control module is used to determine the control instructions of the robot in response to the recognition result of the enhanced voice information; and control the robot according to the control instructions.
[0210] It should be noted that the above explanation of the method embodiment is also applicable to the device of this embodiment, and the principles are the same, which will not be repeated here.
[0211] In the speech enhancement device for a robot of the present embodiment, speech information collected by the robot is acquired, a first enhancement is performed on the speech information to generate initial enhanced speech information, vocal tract parameters of the speech information are generated based on the speech information, and a second enhancement is performed based on the vocal tract parameters and the initial enhanced speech information to generate enhanced speech information. By enhancing the speech information collected by the robot, the contaminated low-frequency information is restored and the attenuated high-frequency information is supplemented, thereby improving the quality of the speech information and further improving the accuracy of speech recognition.
[0212] In order to implement the above embodiment, this embodiment provides a robot, including: at least one processor; and
[0213] a memory communicatively connected to the at least one processor; wherein,
[0214] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method described in the foregoing method embodiments.
[0215] To implement the above embodiments, this embodiment provides a non-transitory computer-readable storage medium storing computer instructions, where the computer instructions are used to cause the computer to execute the method described in the foregoing method embodiments.
[0216] To implement the above embodiments, this embodiment provides a computer program product including a computer program that, when executed by a processor, implements the method described in the foregoing method embodiments.
[0217] Figure 7 FIG. is a schematic block diagram of a robot 800 provided in an embodiment of the present application.
[0218] As Figure 7 shown, the robot 800 includes a computing unit 801, which can execute various appropriate actions and processes according to a computer program stored in a ROM (Read-Only Memory) 802 or a computer program loaded from a storage unit 808 into a RAM (Random Access Memory) 803. In the RAM 803, various programs and data required for the operation of the robot 800 can also be stored. The computing unit 801, the ROM 802, and the RAM 803 are connected to each other through a bus 804. An I / O (Input / Output) interface 805 is also connected to the bus 804.
[0219] Multiple components in the robot 800 are connected to the I / O interface 805, including: an input unit 806; an output unit 807, such as various types of displays, speakers, etc.; a storage unit 808, such as a magnetic disk, an optical disc, etc.; and a communication unit 809, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 809 allows the robot 800 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.
[0220] The computing unit 801 can be various general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a CPU (Central Processing Unit), a GPU (Graphic Processing Units), various dedicated AI (Artificial Intelligence) computing chips, various computing units running machine learning model algorithms, a DSP (Digital Signal Processor), and any suitable processor, controller, microcontroller, etc. The computing unit 801 executes the various methods and processes described above, such as the voice enhancement method for a robot. For example, in some embodiments, the voice enhancement method for a robot can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 808. In some embodiments, part or all of the computer program can be loaded and / or installed onto the robot 800 via the ROM 802 and / or the communication unit 809. When the computer program is loaded into the RAM 803 and executed by the computing unit 801, one or more steps of the voice enhancement method for a robot described above can be executed. Alternatively, in other embodiments, the computing unit 801 can be configured to execute the foregoing method in any other suitable manner (e.g., by means of firmware).
[0221] The various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuitry, integrated circuit systems, FPGAs (Field Programmable Gate Arrays), ASICs (Application-Specific Integrated Circuits), ASSPs (Application Specific Standard Products), SOCs (System On Chip), CPLDs (Complex Programmable Logic Devices), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a dedicated or general-purpose programmable processor that receives data and instructions from a storage system, at least one input device, and at least one output device, and transmits the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0222] The program code for implementing the method of the present application can be written in any combination of one or more programming languages. These program codes can be provided to a processor or a controller of a general-purpose computer, a special-purpose computer, or other programmable data processing devices, such that when the program codes are executed by the processor or the controller, the functions / operations specified in the flowchart and / or block diagram are implemented. The program codes can be executed entirely on the machine, partially on the machine, executed partially on the machine as an independent software package and partially on a remote machine, or executed entirely on a remote machine or a server.
[0223] In the context of the present application, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media would include electrical connections based on one or more wires, portable computer disks, hard disks, RAM, ROM, EPROM (Electrically Programmable Read-Only-Memory), or flash memory, optical fibers, CD-ROM (Compact Disc Read-Only Memory), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0224] In order to provide interaction with a user, the systems and techniques described herein can be implemented on a computer having a display device (e.g., a CRT (Cathode-Ray Tube) or an LCD (Liquid Crystal Display) monitor) for displaying information to the user. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0225] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: LAN (Local Area Network), WAN (Wide Area Network), the Internet, and blockchain networks.
[0226] It should be noted that artificial intelligence is a discipline that studies enabling a computer to simulate certain thinking processes and intelligent behaviors of humans (such as learning, reasoning, thinking, planning, etc.), and it has both hardware-level technologies and software-level technologies. Artificial intelligence hardware technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, and big data processing; artificial intelligence software technologies mainly include several major directions such as computer vision technology, speech recognition technology, natural language processing technology, and machine learning / deep learning, big data processing technology, and knowledge graph technology.
[0227] Those skilled in the art will readily conceive of other embodiments of the present application after considering the specification and practicing the invention disclosed herein. The present application is intended to cover any variations, uses, or adaptations of the present application that follow the general principles of the present application and include known common knowledge or conventional technical means in the technical field not disclosed in the present application. The specification and examples are only regarded as exemplary, and the true scope and spirit of the present application are pointed out by the following claims.
[0228] It should be understood that the present application is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present application is only limited by the appended claims.
[0229] Those skilled in the art will readily conceive of other embodiments of the present application after considering the specification and practicing the invention disclosed herein. The present application is intended to cover any variations, uses, or adaptations of the present application that follow the general principles of the present application and include known common knowledge or conventional technical means in the technical field not disclosed in the present application. The specification and examples are only regarded as exemplary, and the true scope and spirit of the present application are pointed out by the following claims.
[0230] It should be understood that the present application is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present application is only limited by the appended claims.
Claims
1. A voice enhancement method for a robot, characterized in that, it includes: Obtain the voice information collected by the robot, and generate the vocal tract parameters of the voice information according to the voice information; Perform a first enhancement on the voice information to generate initial enhanced voice information; Perform a second enhancement according to the vocal tract parameters and the initial enhanced voice information to generate enhanced voice information; The performing a first enhancement on the voice information to generate initial enhanced voice information includes: Perform downsampling on the voice information to generate first voice information; Generate an excitation spectrum corresponding to the first voice information according to the first voice information; Perform noise reduction on the first voice information to generate the first voice information after noise reduction; Generate the initial enhanced voice information according to the excitation spectrum and the first voice information after noise reduction; The generating an excitation spectrum corresponding to the first voice information according to the first voice information includes: Perform non-negative matrix factorization (NMF) on the first voice information to generate the voice frame probability of the first voice information; Perform a first pitch estimation on the first voice information to generate an initial pitch estimation value of the first voice information; Perform a second pitch estimation on the initial pitch estimation value according to the voice frame probability to generate a target pitch estimation value; Generate the excitation spectrum according to the target pitch estimation value.
2. The method according to claim 1, characterized in that, The performing non-negative matrix factorization (NMF) on the first voice information to generate the voice frame probability of the first voice information includes: Perform Fourier transform on each frame of voice signal in the first voice information to generate a spectral matrix of each frame of voice signal; Based on a pre-obtained voice basis matrix, perform a decomposition operation on the spectral matrix of each frame of voice signal to obtain a voice activation matrix corresponding to the spectral matrix of each frame of voice signal and an updated voice basis matrix; Based on a pre-obtained interference basis matrix, perform a decomposition operation on the spectral matrix of each frame of voice signal to obtain an interference activation matrix corresponding to the spectral matrix of each frame of voice signal and an updated interference basis matrix; Based on the updated voice basis matrix and the updated interference basis matrix, perform repeated decomposition operations on the spectral matrix of each frame of voice signal to obtain a target interference activation matrix and a target voice activation matrix corresponding to the spectral matrix of each frame of voice signal; Determine the voice frame probability of each frame of voice signal in the first voice information according to the target interference activation matrix and the target voice activation matrix.
3. The method according to claim 2, characterized in that, The pre-obtained voice basis matrix includes the following steps: Obtain a sample voice signal; Perform frame division on the sample voice signal to obtain N sub-signals, where N is a positive integer; Perform Fourier transform on each frame of the N sub-signals to determine the spectral values of F frequency points included in each frame of signal, where F is a positive integer; Generate a first sample spectral matrix corresponding to the sample voice signal, including F rows and N columns, according to the spectral values of F frequency points included in each frame of sub-signal. Cluster the first sample spectral matrix to generate the pre-acquired speech basis matrix including K column vectors, where K is a positive integer less than or equal to N.
4. The method according to claim 3, wherein, the steps for the pre-acquired interference basis matrix include: Obtain sample interference signals, which are generated when the robot moves; Frame the sample interference signals to obtain M frame sub-signals, where M is a positive integer; Perform Fourier transform on each of the M frame sub-signals to determine the spectral values of G frequency points included in each frame signal, where G is a positive integer; Generate a second sample spectral matrix corresponding to the sample speech signal including G rows and M columns according to the spectral values of G frequency points included in each frame sub-signal; Cluster the second sample spectral matrix to generate the pre-acquired interference basis matrix including L column vectors, where L is a positive integer less than or equal to M.
5. The method according to claim 1, wherein, the vocal tract parameter is the linear prediction coding residual, and the steps for generating the vocal tract parameter of the speech information according to the speech information include: Perform linear prediction coding LPC envelope estimation on the speech information to generate a linear prediction coding residual.
6. The method according to claim 5, wherein, the steps for performing second enhancement on the initial enhanced speech information according to the vocal tract parameter and the initial enhanced speech information to generate enhanced speech information include: Generate the voiced information of the speech information according to the initial enhanced speech information, the target pitch estimate value, and the speech frame probability; Generate the unvoiced information of the speech information according to the speech frame probability, the linear prediction coding residual, and random noise; Perform first speech synthesis according to the voiced information and the unvoiced information of the speech information to generate first synthesized speech information; Perform spectral broadening on the first synthesized speech information to generate speech information after spectral broadening; Perform second speech synthesis on the first synthesized speech information and the speech information after spectral broadening to generate second synthesized speech information; Perform upsampling on the second synthesized speech information to generate the enhanced speech information.
7. The method according to any one of claims 1-6, wherein, the method further includes: Determine the control instruction of the robot in response to the recognition result of the enhanced speech information; Control the robot according to the control instruction.
8. A speech enhancement device for a robot, wherein, it includes: An acquisition module for acquiring speech information collected by the robot; A first processing module for performing first enhancement on the speech information to generate initial enhanced speech information, and generating the vocal tract parameter of the speech information according to the speech information; and A second processing module for performing second enhancement on the initial enhanced speech information according to the vocal tract parameter and the initial enhanced speech information to generate enhanced speech information; The first processing module includes: A downsampling unit for downsampling the speech information to generate first speech information; A first generation unit, configured to generate an excitation spectrum corresponding to the first voice information according to the first voice information; A noise reduction unit, configured to perform noise reduction on the first voice information to generate the first voice information after noise reduction; A second generation unit, configured to generate the initial enhanced voice information according to the excitation spectrum and the first voice information after noise reduction; The first generation unit is specifically configured to: Perform non-negative matrix factorization (NMF) on the first voice information to generate a voice frame probability of the first voice information; Perform a first pitch estimation on the first voice information to generate an initial pitch estimation value of the first voice information; Perform a second pitch estimation on the initial pitch estimation value according to the voice frame probability to generate a target pitch estimation value; Generate the excitation spectrum according to the target pitch estimation value.
9. A robot, Characterized in that, It includes: At least one processor; And A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method according to any one of claims 1-7.
10. A non-transitory computer-readable storage medium storing computer instructions, Characterized in that, Wherein, The computer instructions are used to cause the computer to execute the method according to any one of claims 1-7.
11. A computer program product, Characterized in that, It includes a computer program, and the computer program realizes the method according to any one of claims 1-7 when executed by a processor.
Citation Information
Patent Citations
Speech enhancement method and device applicable to strong noise environments
CN103208291A
Speech enhancement method for speech recognition in noise environment
CN108831495A