A voice enhancement method, device, computer device, and storage medium
Through the voice enhancement method combined with audio and video, the noise gain factor is adjusted using facial information, which solves the problem of non-steady state noise suppression in the prior art, and achieves higher quality voice signal recognition and suppression effects.
Patent Information
- Application Number
- CN202211458680.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-21
- Publication Date
- 2025-08-05
- Estimated Expiration
- 2042-11-21
AI Technical Summary
Existing speech enhancement methods are difficult to effectively suppress non-steady state noise at low signal-to-noise ratio, resulting in unclear noise removal and damage to vocal quality. Deep learning methods have poor denoising effects on unseen noises.
By combining audio and video information, facial information is used to adjust the noise gain factor, especially the lip action information and the fundamental tone formant frequency, accurately identify the voice signal and suppress noise, including neural network models and noise estimation technology.
Better suppress non-steady state noise, improve speech quality and robustness, accurately identify speech signals, reduce speech distortion, and enhance the clarity of speech signals.
Smart Images

Figure CN115910095B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer application technology, and in particular to a speech enhancement method, apparatus, computer equipment, and computer-readable storage medium. Background Art
[0002] In many video conversation scenarios, the microphone collects background noise while capturing the human voice, which greatly reduces the user experience and makes it more difficult for the person on the other end of the video to understand the spoken content. Therefore, it is necessary to perform voice enhancement processing on the sound signal, including noise removal and improving the quality of the human voice.
[0003] Existing speech enhancement methods can be categorized into traditional methods and deep learning. Traditional methods involve two steps: noise estimation and noise suppression. They determine the presence of noise based on the input speech signal. If speech is absent, the noise estimate is updated. Then, noise suppression is performed on the noisy signal using statistical methods, Wiener filtering, or spectral subtraction. However, traditional methods cannot suppress non-stationary noise. At low signal-to-noise ratios, the accuracy of noise estimation decreases, and weak vocal components may be misinterpreted as noise. This results in incomplete noise removal and impairs vocal quality. Furthermore, at low signal-to-noise ratios, the accuracy of pitch estimation and formant estimation also decreases, making it impossible to protect the pitch and its multiples, and unable to use formants to reduce speech distortion. Another deep learning approach involves building a deep learning model to learn the mapping from the noisy speech spectrum to the clean speech spectrum. This method can remove non-stationary noise, but the denoising effect is dataset-dependent and may not be effective for noise not present in the dataset. Summary of the Invention
[0004] The purpose of the present invention is to provide a speech enhancement method, apparatus, computer device and computer-readable storage medium. Compared with existing speech enhancement methods, the present invention realizes speech enhancement by combining audio and video information, avoids the influence of environmental noise, better suppresses non-steady-state noise, can more accurately recognize speech signals, improves speech quality and has higher robustness.
[0005] According to one aspect of the present invention, the present invention provides a method for speech enhancement, comprising:
[0006] Acquiring audio and video data, wherein the audio and video data includes image information and voice signals;
[0007] Determining whether there is a human voice in the speech signal;
[0008] If the human voice is present, determining whether corresponding facial information exists in the image information;
[0009] If the facial information exists, adjusting the noise gain factor according to the facial information;
[0010] The noise gain factor is used to suppress the noise to obtain the enhanced speech signal.
[0011] Optionally, adjusting the noise gain factor according to the facial information includes:
[0012] Extracting lip movement information from the facial information, and using a movement recognition module to identify the lip movement information to obtain phonemes for pronunciation;
[0013] According to the phoneme, extracting the fundamental pitch and formant frequency of normal pronunciation from a database;
[0014] The noise gain factor is adjusted according to the fundamental pitch and the formant frequency.
[0015] Optionally, extracting lip movement information from the facial information includes:
[0016] The facial information is extracted using a neural network model to obtain the lip movement information.
[0017] Optionally, after obtaining the audio and video data, the method further includes:
[0018] Extracting the speech signal to obtain audio features;
[0019] Extracting the image information to obtain lip information;
[0020] splicing the audio features and the lip information using time synchronization to obtain audio and video fusion information;
[0021] Accordingly, determining whether corresponding facial information exists in the image information includes:
[0022] Determine whether the lip information corresponding to the audio feature exists in the audio and video fusion information.
[0023] Optionally, extracting the image information to obtain lip information includes:
[0024] performing lip positioning on the image information;
[0025] According to the lip positioning, the lip information corresponding to the lip positioning is extracted.
[0026] Optionally, determining whether a human voice is present in the speech signal includes:
[0027] A human voice detection module is used to determine whether the human voice exists in the speech signal.
[0028] Optionally, the method further includes:
[0029] If the human voice is not present, obtaining a noise estimate based on the speech signal;
[0030] Accordingly, the speech signal enhanced by suppressing noise using the noise gain factor includes:
[0031] The speech signal enhanced by suppressing noise is obtained by utilizing the noise estimate and the noise gain factor.
[0032] The present invention provides a speech enhancement device, comprising:
[0033] A receiving module, configured to obtain audio and video data, wherein the audio and video data includes image information and voice signals;
[0034] A first judgment module is used to determine whether there is a human voice in the speech signal;
[0035] a second determination module, configured to determine whether corresponding facial information exists in the image information if the human voice exists;
[0036] an adjustment module, configured to adjust a noise gain factor according to the facial information if the facial information exists;
[0037] The speech enhancement module is configured to suppress the noise by using the noise gain factor to obtain the enhanced speech signal.
[0038] The present invention provides a computer device, characterized by comprising:
[0039] Memory for storing computer programs;
[0040] A processor is configured to implement the above-mentioned speech enhancement method when executing the computer program.
[0041] The present invention provides a computer-readable storage medium, characterized in that the computer-readable storage medium stores computer-executable instructions, and when the computer-executable instructions are loaded and executed by a processor, the steps of the speech enhancement method as described above are implemented.
[0042] As can be seen, compared to existing speech enhancement methods, the present invention achieves speech enhancement by combining audio and video information, avoiding the influence of environmental noise, better suppressing non-stationary noise, and more accurately recognizing speech signals, thereby improving speech quality and providing higher robustness. The present application also provides a speech enhancement device, computer equipment, and computer-readable storage medium, all of which have the aforementioned beneficial effects. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are merely embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying any creative work.
[0044] Figure 1 A flow chart of a speech enhancement method provided by an embodiment of the present invention;
[0045] Figure 2 A flowchart of another speech enhancement method provided by an embodiment of the present invention;
[0046] Figure 3 A flowchart of a non-human voice enhancement method provided by an embodiment of the present invention;
[0047] Figure 4 A structural block diagram of a speech enhancement device provided by an embodiment of the present invention;
[0048] Figure 5 This is a structural block diagram of a computer device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0049] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.
[0050] Based on the problems existing in the prior art, the present invention provides a speech enhancement method. Compared with the existing speech enhancement method, the present invention realizes speech enhancement by combining audio and video information, avoids the influence of environmental noise, better suppresses non-steady-state noise, can more accurately recognize speech signals, improves the quality of speech and has higher robustness.
[0051] The following is a detailed description, please refer to Figure 1 , Figure 1 This is a flow chart of a speech enhancement method provided by an embodiment of the present invention. The speech enhancement method according to an embodiment of the present invention may include:
[0052] Step S101: Acquire audio and video data, wherein the audio and video data includes image information and voice signals.
[0053] In the embodiments of the present invention, audio and video data may be data combining voice signals and image information. The voice signals may include voice data containing both human and non-human voices, and the image information may include a large amount of image data obtained through filming, including facial information, environmental information, and the like. In the embodiments of the present invention, there is no limitation on the method for obtaining the audio and video data, and the data may be obtained via a mobile phone device or other audio and video recording device.
[0054] Step S102: Determine whether there is a human voice in the speech signal. Then, step S103 is performed: if there is a human voice, determine whether there is corresponding facial information in the image information.
[0055] In the embodiment of the present invention, it is first determined whether a human voice is present in the speech signal. If a human voice is present, step S103 is executed: if a human voice is present, corresponding facial information is determined in the image information. Facial information may include face information, lip movement information, and eye information. It should be noted that in the embodiment of the present invention, the human voice detection module may determine whether a human voice is present based on spectral changes in the speech signal, or based on other audio features of the speech signal, and this is not limited in the embodiment of the present invention.
[0056] Step S104: If there is facial information, the noise gain factor is adjusted according to the facial information. Then, according to the noise gain factor, step S105 is executed: the noise gain factor is used to suppress the noise to obtain an enhanced speech signal.
[0057] In an embodiment of the present invention, lip movement information can be extracted from facial information, and the lip movement information can be identified using a motion recognition module to obtain the phonemes of the pronunciation. Then, based on the phonemes, the fundamental pitch and formant frequency of normal pronunciation can be extracted from a database. Finally, the noise gain factor can be adjusted based on the fundamental pitch and formant frequency. It should be noted that motion recognition can utilize a hidden Markov method, a neural network method, or a combination of the two, and this is not limited in the embodiment of the present invention. In particular, in an embodiment of the present invention, a neural network model can be used to extract facial information to obtain lip movement information, and then, based on the lip movement information, motion recognition can be used to obtain the phonemes of the current pronunciation. For example, based on the lip movement in the facial information, a convolutional neural network and a recurrent neural network can be used to identify the phonemes corresponding to the current action. For example, if the lip movement is identified as "hello", the three phonemes "ni3", "h", and "ao3" can be obtained, where "3" represents the third tone. It should be noted that a convolutional neural network is a type of feedforward neural network that includes convolution calculations and has a deep structure. It is one of the representative algorithms of deep learning. A convolutional neural network has representational learning capabilities and can perform translation-invariant classification of input information according to its hierarchical structure. A recurrent neural network is a type of recursive neural network that takes sequence data as input, recursively evolves in the direction of the sequence, and all nodes are connected in a chain. It has applications in natural language processing such as speech recognition, language modeling, machine translation, and other fields, and is also used for various time series forecasts.
[0058] Specifically, in the embodiment of the present invention, the noise gain factor is adjusted according to the fundamental pitch and the formant frequency, and the adjustment calculation can be performed using the following formula:
[0059]
[0060] in, is the adjusted noise gain factor. If k is the formant frequency or the multiple of the fundamental frequency, then Keep unchanged; if k is not a multiple of the formant frequency and the fundamental frequency, then = In practical applications The value range is 0-1, for example A value of 0.3 is acceptable, meaning that less suppression is applied at formant and octave frequencies, which can reduce speech distortion and improve intelligibility, while more suppression is applied at non-formant frequencies and non-octave frequencies.
[0061] It should be noted that, in the embodiment of the present invention, the noise gain factor can be calculated by using methods such as Wiener filtering, statistical methods, and spectral subtraction.
[0062] In a specific embodiment, the noise gain factor is obtained using the formula in the Wiener filtering method, and its calculation formula is as follows:
[0063]
[0064] in, is the Wiener filter gain factor, is the prior signal-to-noise ratio of frequency k, where The decision-guided method can be used to estimate the following formula:
[0065]
[0066] in, is a smoothing constant, represents the frequency k-enhanced signal obtained in the m-1th frame, and represent noisy speech and noise spectrum respectively.
[0067] Accordingly, according to the obtained gain factor, step S105 is performed: the noise gain factor is used to suppress the noise to obtain an enhanced speech signal. In the embodiment of the present invention, the enhanced speech signal can be obtained by multiplying the noise gain factor with the frequency domain. For example, in the Wiener filtering method, the enhanced speech signal is obtained according to the formula, which is as follows:
[0068]
[0069] in, represents the estimated denoised speech signal, is the representation of the noisy signal in the frequency domain.
[0070] It should be noted that, in the embodiment of the present invention, the spectral minimum mean square error method MMSE can also be used to obtain an enhanced speech signal. For example, in the spectral minimum mean square error method, the enhanced speech signal is obtained according to the formula, which is as follows:
[0071]
[0072] in, denote the zero-order and first-order modified Bessel functions, respectively, represents the estimated denoised speech signal, is the posterior signal-to-noise ratio, which can be calculated according to the formula:
[0073]
[0074] Where, represents the frequency domain represents the noise estimate.
[0075] It can be seen that in the embodiment of the present invention, if the facial information exists, the noise gain factor is adjusted according to the facial information, and then the noise gain factor is used to suppress the noise to obtain an enhanced voice signal, thereby achieving voice enhancement. In the embodiment of the present invention, through the combination of audio and video, the pronunciation phonemes in the voice signal can be more accurately identified, the quality of the voice is improved, and the normal pronunciation is obtained by adjusting the noise gain factor through phonemes, which can better suppress non-steady-state noise, avoid the influence of environmental noise, and have higher robustness.
[0076] Please refer to Figure 2 , which is another speech enhancement method provided by an embodiment of the present invention.
[0077] Step S201: Acquire audio and video data, wherein the audio and video data includes voice signals and image information.
[0078] Step S202: extracting the speech signal to obtain audio features.
[0079] In the embodiment of the present invention, speech features can be extracted from speech signals in audio and video data to obtain audio features, such as extracting human voices. Extracting speech features can improve the efficiency of speech recognition and ensure the quality of speech recognition.
[0080] Step S203: extracting image information to obtain lip information.
[0081] In the embodiment of the present invention, the lip information includes lip movement features, etc. It should be noted that, in the embodiment of the present invention, it is necessary to first perform face detection on the image information, extract the face image, and then extract the lip information from the face image. In some embodiments, the extracted face image can be compressed first, and the compressed image data can be processed accordingly to further reduce the complexity. It should be noted that, in the embodiment of the present invention, lip positioning is performed on the compressed image data, and then, based on the lip positioning, the lip information corresponding to the lip positioning is extracted. There is no limitation on the method of compressing the image, and the image information can be compressed using the principal component analysis method, the discrete cosine transform method, or the wavelet transform method.
[0082] Step S204: splicing the audio features and lip information using time synchronization to obtain audio and video fusion information.
[0083] In the embodiment of the present invention, the audio and video fusion information includes audio features and visual features. The time-synchronized audio features and visual features can be spliced together according to time information, and then the dimensionality of the spliced fusion features can be reduced to obtain the audio and video fusion information. In the embodiment of the present invention, LDA (Linear Discriminant Analysis) can be used, followed by an MLLT (Maximum Likelihood Data Rotation) to transform the fusion feature data to obtain the audio and video fusion information, which can improve the efficiency of speech recognition.
[0084] Step S205: Determine whether there is lip information corresponding to the audio feature in the audio and video fusion information.
[0085] Step S206: If lip information exists, adjust the noise gain factor according to the lip information.
[0086] Step S207: using the noise gain factor to suppress the noise to obtain an enhanced speech signal.
[0087] Based on the above embodiments, an embodiment of the present invention provides a speech enhancement method. Compared with the existing method for enhancing speech, the present invention adjusts the noise gain factor according to the lips, and then uses the noise gain factor to suppress the noise to obtain an enhanced speech signal to achieve speech enhancement, thereby avoiding the influence of environmental noise, better suppressing non-steady-state noise, and being able to more accurately recognize speech signals, thereby improving the quality of speech and having higher robustness.
[0088] Please refer to Figure 3 , a flowchart of a non-human voice enhancement method provided by an embodiment of the present invention, a non-human voice enhancement method according to an embodiment of the present invention may include:
[0089] Step S301: Acquire audio and video data, where the audio and video data includes image information and voice signals.
[0090] Step S302: Determine whether there is a human voice in the speech signal.
[0091] Step S303: If there is no human voice, obtain a noise estimate based on the speech signal.
[0092] Noise estimation in the embodiment of the present invention is to use an algorithm to numerically estimate the size of the noise. Common noise estimation algorithms include recursive averaging, minimum tracking, and histogram statistics. In the embodiment of the present invention, the recursive averaging method can be used to obtain noise estimation.
[0093] In the embodiment of the present invention, when there is no human voice, the noise estimate can be updated based on the speech signal; when there is a human voice, the existing noise estimate is not updated. It should be noted that the noise estimate can be obtained by performing a first-order recursion using a recursive averaging method, where the formula for the first-order recursion is as follows:
[0094]
[0095] in, is the noise estimate, For the Frame speech signal, It can be regarded as the probability of speech existence. If it is 1, it means that the frequency band k speech exists, that is, the speech when there is human voice, use As the noise estimate of the current frame l; when 0 means there is no human voice and only speech signal. It is equivalent to In practical applications, if it is non-human voice, then = , and complete the updated noise estimation; if it is human voice, then the fundamental frequency and the fundamental frequency corresponding to It should be set to a value close to 1, such as 0.98, to reduce noise estimation and protect the human voice. The fundamental pitch of a male is generally 0 to 200Hz. For example, if the fundamental pitch is 100Hz, then the multiplication frequency is 200 / 300 / 400 / 500 / etc. The multiplication frequency is used to reduce noise while protecting the human voice information. It should be noted that It can be obtained according to the formula, which is as follows:
[0096]
[0097] in, is the lth frame speech signal, is the noisy speech of frequency k, For clean speech and noise. When the non-human voice microphone collects It is equivalent to .
[0098] In the embodiment of the present invention, by performing noise estimation based on the speech signal, non-stationary noise can be better suppressed, the speech signal can be recognized more accurately, the speech quality is improved, and the robustness is higher.
[0099] Step S304: using the noise estimation and the noise gain factor, suppressing the noise to obtain an enhanced speech signal.
[0100] In the embodiment of the present invention, the noise can be suppressed based on the noise estimation and the noise gain factor to obtain an enhanced speech signal. The noise estimation can be used as a parameter in the Wiener filtering method to obtain an enhanced speech signal by calculating the noise estimation and the noise gain factor. The noise estimation can also be used as a parameter in the spectral minimum mean square error method to obtain an enhanced speech signal by calculating the noise estimation and the noise gain factor. The embodiment of the present invention does not limit the method for suppressing noise.
[0101] Based on the above embodiments, an embodiment of the present invention provides a speech enhancement method. Compared with the existing method for enhancing speech, the present invention obtains a noise estimate based on the speech signal, and then suppresses the noise based on the noise estimate and the noise gain factor to obtain an enhanced speech signal to achieve speech enhancement. This avoids the influence of environmental noise, better suppresses non-steady-state noise, can more accurately recognize speech signals, improves the quality of speech, and has higher robustness.
[0102] A speech enhancement device and a computer device provided by an embodiment of the present invention are introduced below. The speech enhancement device and the computer device described below can be referenced to the speech enhancement method described above.
[0103] Please refer to Figure 4 , Figure 4 This is a structural block diagram of a speech enhancement device provided by an embodiment of the present invention. The device may include:
[0104] A receiving module 10 is configured to obtain audio and video data, wherein the audio and video data includes image information and voice signals;
[0105] A first judgment module 20 is used to determine whether there is a human voice in the speech signal;
[0106] A second determination module 30 is configured to determine whether corresponding facial information exists in the image information if the human voice exists;
[0107] an adjustment module 40 for adjusting a noise gain factor according to the facial information if the facial information exists;
[0108] The speech enhancement module 50 is configured to suppress noise using the noise gain factor to obtain a clean speech signal.
[0109] Based on the above embodiment, the adjustment module 40 may include:
[0110] a recognition unit, configured to extract lip movement information from the facial information, and identify the lip movement information using movement recognition to obtain phonemes for pronunciation;
[0111] An extraction unit, configured to extract the fundamental pitch and formant frequency of normal pronunciation from a database according to the phoneme;
[0112] An adjustment unit is configured to adjust the noise gain factor according to the fundamental tone and the formant frequency.
[0113] Based on any of the above embodiments, the identification unit may include:
[0114] an extraction subunit, configured to extract the facial information using a neural network model to obtain the lip movement information;
[0115] The recognition subunit is used to obtain the phoneme of the current pronunciation by using the lip movement information and the movement recognition.
[0116] Based on any of the above embodiments, the receiving module 10 may include
[0117] An audio extraction module, configured to extract the speech signal to obtain audio features;
[0118] A visual extraction module, configured to extract the image information to obtain lip information;
[0119] A fusion module is used to use time synchronization to splice the audio features and the lip information to obtain audio and video fusion information.
[0120] In the embodiment of the present invention, after obtaining the audio and video fusion information, it can be determined whether the lip information corresponding to the audio feature exists in the audio and video fusion information.
[0121] Based on any of the above embodiments, the visual extraction module may include:
[0122] a positioning unit, configured to perform lip positioning on the image information;
[0123] An extraction unit is configured to extract, based on the lip positioning, the lip information corresponding to the lip positioning.
[0124] Based on any of the above embodiments, the first determination module 20 may include:
[0125] The judgment unit is used to use the human voice detection module to determine whether the human voice exists in the speech signal.
[0126] Based on any of the above embodiments, after the first determining module 20, the following steps may be further included:
[0127] a noise estimation module, configured to obtain a noise estimate based on the speech signal if the human voice is not present;
[0128] In the embodiment of the present invention, the noise estimation and the noise gain factor can be used to suppress the speech signal with enhanced noise.
[0129] In the embodiment of the present invention, a second judgment module 30 is used to determine whether corresponding facial information exists in the image information if the human voice exists, and an adjustment module 40 is used to adjust the noise gain factor according to the facial information if the facial information exists. The method of combining audio and video information to achieve speech enhancement avoids the influence of environmental noise, better suppresses non-steady-state noise, can more accurately recognize speech signals, improve the quality of speech, and has higher robustness.
[0130] Please refer to Figure 5 , Figure 5 This is a structural block diagram of a computer device provided in an embodiment of the present invention, the computer device comprising:
[0131] Memory 10, for storing computer programs;
[0132] The processor 20 is configured to implement the above-mentioned speech enhancement method when executing the computer program.
[0133] like Figure 5 , which is a schematic diagram of the structure of a computer device, may include: a memory 10 , a processor 20 , a communication interface 31 , an input / output interface 32 , and a communication bus 33 .
[0134] In an embodiment of the present invention, the memory 10 is used to store one or more programs. The program may include program code, and the program code includes computer operating instructions. In an embodiment of the present application, the memory 10 may store programs for implementing the following functions:
[0135] Acquiring audio and video data, wherein the audio and video data includes image information and voice signals;
[0136] Determining whether there is a human voice in the speech signal;
[0137] If the human voice is present, determining whether corresponding facial information exists in the image information;
[0138] If the facial information exists, adjusting the noise gain factor according to the facial information;
[0139] The noise gain factor is used to suppress the noise to obtain the enhanced speech signal.
[0140] In one possible implementation, the memory 10 may include a program storage area and a data storage area, wherein the program storage area may store an operating system and applications required for at least one function, etc.; the data storage area may store data created during use.
[0141] In addition, the memory 10 may include a read-only memory and a random access memory, and provides instructions and data to the processor. A portion of the memory may also include NVRAM. The memory stores an operating system and operating instructions, executable modules or data structures, or a subset or an extended set thereof. The operating instructions may include various operating instructions for implementing various operations. The operating system may include various system programs for implementing various basic tasks and processing hardware-based tasks.
[0142] The processor 20 may be a central processing unit (CPU), an application-specific integrated circuit, a digital signal processor, a field programmable gate array, or other programmable logic device. The processor 20 may be a microprocessor or any conventional processor. The processor 20 may call a program stored in the memory 10 .
[0143] The communication interface 31 may be an interface for connecting to other devices or systems.
[0144] The input / output interface 32 may be an interface for obtaining external input data or outputting data to the outside world.
[0145] Of course, it needs to be explained that Figure 5 The structure shown does not constitute a limitation on the computer device in the embodiment of the present application. In actual applications, the computer device may include Figure 5 More or fewer components than shown, or combinations of certain components.
[0146] The method for achieving speech enhancement by combining audio and video information in the embodiment of the present invention avoids the influence of environmental noise, better suppresses non-stationary noise, can more accurately recognize speech signals, improves the quality of speech and has higher robustness.
[0147] An embodiment of the present invention also provides a computer-readable storage medium storing computer-executable instructions. When loaded and executed by a processor, the computer-executable instructions acquire audio and video data, wherein the audio and video data includes image information and a voice signal; determine whether a human voice exists in the voice signal; if the human voice exists, determine whether corresponding facial information exists in the image information; if the facial information exists, adjust a noise gain factor based on the facial information; and utilize the noise gain factor to suppress noise and thereby enhance the voice signal. Compared to existing methods for enhancing voice, the method for enhancing voice through the combination of audio and video information in the embodiments of the present invention avoids the influence of ambient noise, better suppresses non-stationary noise, more accurately recognizes voice signals, improves voice quality, and exhibits higher robustness.
[0148] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. Reference can be made to the common and similar parts between the various embodiments. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the method description.
[0149] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present invention.
[0150] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein may be implemented directly using hardware, a software module executed by a processor, or a combination of the two. The software module may be placed in random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art.
[0151] The above is a detailed introduction to the speech enhancement method, device, computer equipment, and storage medium provided by the present invention. Specific examples are used herein to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only intended to help understand the method and core ideas of the present invention. It should be pointed out that, for those skilled in the art, without departing from the principles of the present invention, several improvements and modifications can be made to the present invention, and these improvements and modifications also fall within the scope of protection of the claims of the present invention.
Claims
1. A speech enhancement method, characterized in that: include: Acquiring audio and video data, wherein the audio and video data includes image information and voice signals; Determining whether there is a human voice in the speech signal; If the human voice is present, determining whether corresponding facial information exists in the image information; If the facial information exists, adjusting the noise gain factor according to the facial information; Suppressing noise using the noise gain factor to enhance the speech signal; The adjusting the noise gain factor according to the facial information includes: Extracting lip movement information from the facial information, and using a movement recognition module to identify the lip movement information to obtain phonemes for pronunciation; According to the phoneme, extracting the fundamental pitch and formant frequency of normal pronunciation from a database; adjusting the noise gain factor according to the fundamental pitch and the formant frequency; The method of suppressing noise by using the noise gain factor to enhance the speech signal includes: Multiplying the noise gain factor with the frequency domain yields the enhanced speech signal.
2. The speech enhancement method according to claim 1, wherein: The extracting lip movement information from the facial information includes: The facial information is extracted using a neural network model to obtain the lip movement information.
3. The speech enhancement method according to claim 1, wherein: After obtaining the audio and video data, the method further includes: Extracting the speech signal to obtain audio features; Extracting the image information to obtain lip information; splicing the audio features and the lip information using time synchronization to obtain audio and video fusion information; Accordingly, determining whether corresponding facial information exists in the image information includes: Determine whether the lip information corresponding to the audio feature exists in the audio and video fusion information.
4. The speech enhancement method according to claim 3, wherein: The extracting the image information to obtain lip information includes: performing lip positioning on the image information; According to the lip positioning, the lip information corresponding to the lip positioning is extracted.
5. The speech enhancement method according to claim 1, wherein: Determining whether a human voice is present in the speech signal includes: A human voice detection module is used to determine whether the human voice exists in the speech signal.
6. The speech enhancement method according to claim 1, wherein: The method further comprises: If the human voice is not present, obtaining a noise estimate based on the speech signal; Accordingly, the speech signal enhanced by suppressing noise using the noise gain factor includes: The speech signal enhanced by suppressing noise is obtained by utilizing the noise estimate and the noise gain factor.
7. A speech enhancement device, characterized in that: include: A receiving module, configured to obtain audio and video data, wherein the audio and video data includes image information and voice signals; A first judgment module is used to determine whether there is a human voice in the speech signal; a second determination module, configured to determine whether corresponding facial information exists in the image information if the human voice exists; an adjustment module, configured to adjust a noise gain factor according to the facial information if the facial information exists; A speech enhancement module, configured to suppress the noise by using the noise gain factor to obtain the enhanced speech signal; The adjustment module is specifically configured to extract lip movement information from the facial information, identify the lip movement information using the movement recognition module to obtain phonemes of pronunciation; extract the fundamental pitch and formant frequency of normal pronunciation from a database based on the phonemes; and adjust the noise gain factor based on the fundamental pitch and the formant frequency. The speech enhancement module is specifically configured to multiply the noise gain factor by the frequency domain to obtain an enhanced speech signal.
8. A computer device, characterized in that: include: Memory for storing computer programs; A processor, configured to implement the speech enhancement method according to any one of claims 1 to 6 when executing the computer program.
9. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer-executable instructions, and when the computer-executable instructions are loaded and executed by the processor, the steps of the speech enhancement method according to any one of claims 1 to 6 are implemented.
Citation Information
Patent Citations
Formant dependent speech signal enhancement
CN104704560A
Audio signal processing equipment and method as well as electronic equipment
CN106653041A