Speech enhancement device and method
By detecting and enhancing formants, combining adjacent audio frames to form audio segments and performing gain processing, the problem of insufficient speech clarity is solved, and speech recognition ability is improved.
Patent Information
- Application Number
- CN202410598668.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-15
- Publication Date
- 2025-11-18
AI Technical Summary
Existing technologies are insufficient to effectively enhance speech clarity, especially for people with hearing loss, particularly those who have lost sensitivity to certain frequency bands, leading to difficulties in speech recognition.
By detecting and enhancing formants, combining adjacent audio frames to form audio segments, and then performing gain processing on these segments, a clearer speech data is finally generated.
It improves the clarity and uniqueness of speech, making speech recognition easier, especially for people with hearing loss, enhancing their speech recognition ability.
Smart Images

Figure CN120977320A_ABST
Abstract
Description
Technical Field
[0001] This application relates to apparatus and methods for speech enhancement, and more particularly to apparatus and methods for enhancing speech by detecting formants. Background Technology
[0002] Today, electronic products typically use technologies such as increased volume, noise reduction, or echo cancellation to enhance speech clarity. These technologies primarily utilize the difference in energy characteristics between noise and speech to separate the two. However, in some cases of hearing loss, the loss occurs in sensitivity to certain frequency bands. Some people only lose hearing in high-frequency sounds, while others experience hearing loss in only a few frequencies. Still others not only experience decreased sensitivity to certain frequencies but also require a greater difference in volume between two adjacent sounds for their hearing to distinguish them.
[0003] Therefore, there is a need for a speech enhancement device and method that can improve speech clarity. Summary of the Invention
[0004] This application provides a speech enhancement device. The speech enhancement device includes an audio input circuit and a processor. The audio input circuit is configured to convert an audio input signal into first audio data. The processor is configured to perform: generating a plurality of audio frames based on the first audio data; performing formant analysis on the audio frames to determine whether to combine adjacent audio frames of the audio frames into an audio segment; performing gain processing on the audio segment including the combined audio frames; and combining the audio segment with one or more uncombined audio frames of the audio frames to form second audio data.
[0005] Furthermore, this application provides a speech enhancement method. The speech enhancement method includes: converting an audio input signal into first audio data; generating a plurality of audio frames based on the first audio data; performing formant analysis on the audio frames to determine whether to combine adjacent audio frames into an audio segment; performing gain processing on the audio segment including the combined audio frames; and combining the audio segment with one or more uncombined audio frames of the audio frames into second audio data.
[0006] In summary, the speech enhancement device and method provide speech with enhanced resonant waves, making the provided speech more unique. Therefore, speech clarity can be improved, thereby enhancing speech recognition. Attached Figure Description
[0007] The various embodiments disclosed herein can be best understood by reading the following description and the accompanying drawings. It should be noted that, in accordance with standard practice in the art, the various features in the figures are not drawn to scale. In fact, the dimensions of certain features may be intentionally enlarged or reduced for clarity of description.
[0008] Figure 1 This disclosure describes an embodiment of an electronic device.
[0009] Figure 2 This is a schematic diagram of formants in speech as described in one embodiment of this disclosure.
[0010] Figure 3 This disclosure presents an embodiment of a speech enhancement method.
[0011] Figure 4A and Figure 4B This is a schematic diagram showing the sampling of audio data and the generation of multiple audio frames.
[0012] Figure 5 This is a schematic diagram showing the window frame processing, linear predictive coding, and formant analysis performed on the audio frame.
[0013] Figure 6 It is a table that displays the formants of an audio segment formed by multiple audio frames.
[0014] Figure 7A As described in one embodiment of this disclosure Figure 6 The flowchart for gain processing of audio segments.
[0015] Figure 7B It is a display Figure 7A A schematic diagram of the signals at each stage of the process.
[0016] Figure 8 This is a schematic diagram showing how audio segments are combined with uncombined audio frames to form audio data. Detailed Implementation
[0017] Hearing loss takes many forms. This disclosure presents a speech enhancement device and method to enhance the formant frequency bands that are acoustically associated with the human speech process. Since each person's physiological structure is different, the formant frequency bands vary from person to person; therefore, enhancing specific frequency bands can result in clearer speech intelligibility.
[0018] Figure 1This disclosure describes an embodiment of an electronic device 100. The electronic device 100 includes a voice enhancement device 10, a microphone 20, and a playback device 30. The microphone 20 receives ambient speech and provides an audio input signal Sin to the voice enhancement device 10. The voice enhancement device 10 enhances the speech portion of the audio input signal Sin and provides an enhanced audio output signal Sout to the playback device 30 for playback. In some embodiments, the microphone 20 may be a microphone, and the playback device 30 may be a speaker. In some embodiments, the voice enhancement device 10 is implemented in an integrated circuit (IC). In some embodiments, the electronic device 100 can be applied to devices for receiving and playing audio, such as hearing aids, portable electronic devices, and home appliances.
[0019] The voice enhancement device 10 includes an audio input circuit 12, a storage device 13, a processor 14, and an audio output circuit 16. The audio input circuit 12 is configured to generate audio data D1 based on an audio input signal Sin. The audio input signal Sin is an analog signal, while the audio data D1 is a digital signal. In some embodiments, the audio input circuit 12 includes an analog-to-digital converter (ADC), a noise suppressor, etc. After receiving the audio data D1, the processor 14 enhances the speech characteristics (e.g., formants) in the audio data D1 to generate audio data D2. The audio output circuit 16 is configured to generate an audio output signal Sout based on the audio data D2. In some embodiments, the audio output signal Sout is an analog signal, while the audio data D2 is a digital signal. In some embodiments, the audio output circuit 16 includes a digital-to-analog converter (DAC), an amplifier, etc.
[0020] Generally, speech is composed of a series of sound waves, mostly below 5000Hz. The sound source begins with the vibration of the vocal cords, and the shape of the vocal cords and the muscles controlling their vibration affect the frequency of the vibration, i.e., pitch, also known as F0. When the sound source passes through the vocal tract, the airflow and the structure of the vocal tract resonate, giving certain frequencies a distinct intensity. These frequencies are called formants, and the formants determine the timbre. The lowest frequency formant is called F1, the second lowest is F2, and so on. Because everyone's physiological structure is different, the measured formants for the same pronunciation will vary from person to person. Speech typically contains 4 to 5 stable and relatively strong formants. Furthermore, connecting the formant frequencies at each time point forms multiple formant curves, i.e., voiceprints, such as... Figure 2 As shown.
[0021] exist Figure 1In the speech enhancement device 10, the processor 14 can execute an audio processing program to detect and enhance formants in the speech from the audio data D1 to provide audio data D2. Furthermore, the processor 14 can execute the audio processing program in conjunction with a storage device 13. In some embodiments, the processor 14 is a digital signal processor (DSP) or a central processing unit (CPU). In some embodiments, the storage device 13 may include memory or a buffer.
[0022] In some embodiments, the voice enhancement device 10 further includes a communication module 17. The communication module 17 is configured to connect to other electronic devices (e.g., televisions, Bluetooth speakers, etc.) via wired or wireless means. In some embodiments, the processor 14 may encode audio data D2 according to a known audio encoding format (e.g., Pulse Code Modulation (PCM) format) and provide audio data D3 to the communication module 17 so that the audio data D3 can be transmitted to other electronic devices for playback. In some embodiments, the processor 14 may receive and decode audio data D3 from other electronic devices via the communication module 17. The processor 14 may then detect formants from the decoded audio data and enhance the voice based on the detected formants to provide audio data D2 to the audio output circuit 16.
[0023] Figure 3 This disclosure presents an embodiment of a speech enhancement method 200. The speech enhancement method 200 can be derived from... Figure 1 The voice enhancement device 10 performs this function. In some embodiments, Figure 1 The processor 14 may be implemented using any suitable form, including hardware circuitry, software, firmware, or any combination thereof. In some embodiments, at least a portion of the processor 14 may be selectively implemented as computer software running on one or more image processors, data processors, and / or digital signal processors or configurable modular elements (e.g., FPGAs).
[0024] Please include Figure 1 refer to Figure 3 During operation S202, the audio input circuit 12 converts the audio input signal Sin into audio data D1 and provides it to the processor 14.
[0025] In operation S204, the processor 14 samples the audio data D1 to generate consecutive audio frames Fr and performs windowing on the audio frames Fr, i.e., executes a window function on the audio frames Fr. In some embodiments, two adjacent audio frames Fr may partially overlap. In some embodiments, the sampling rate of the audio data D1 is determined by the sampling frequency of the audio input circuit 12. Furthermore, the sampling rate and the window function are stored in the storage device 13. In some embodiments, the processor 14 decodes, samples, and performs windowing on the audio data D3 from the communication module 17 to generate audio frames Fr.
[0026] Please include Figure 1 refer to Figure 4A , Figure 4A This is a schematic diagram illustrating the sampling of audio data D1 and the generation of audio frames Fr_0 to Fr_n according to an embodiment of this disclosure. Processor 14 first samples the audio data D1 to generate raw audio frames F_0 to F_n. Then, processor 14 overlaps adjacent raw audio frames to obtain audio frames Fr_0 to Fr_n. In the embodiment of FIG. 4, the raw audio frames F_0-F_n and the audio frames Fr_0-Fr_n have the same frame size T_frame, while there is a frame overlap T_overlap between the audio frames Fr_0-Fr_n. For example, audio frame Fr_2 partially overlaps with the adjacent preceding audio frame Fr_1 and following audio frame Fr_3. In some embodiments, the audio frame count T_frame of audio frames Fr_0-Fr_n includes 1024 sampling points, while the audio frame overlap T_overlap is 256 sampling points. Furthermore, the audio frame hop size T_hop of audio frames Fr_0-Fr_n represents the starting distance between two adjacent audio frames Fr, where the audio frame hop size T_hop is equal to the audio frame count T_frame minus the audio frame overlap T_overlap. The audio frame count T_frame, the audio frame hop size T_hop, and the audio frame overlap T_overlap can be stored in storage device 13. In other embodiments, adjacent audio frames Fr do not overlap, that is, the audio frame overlap T_overlap is 0 sampling points, making the audio frame hop size T_hop the same as the audio frame count T_frame, such as... Figure 4B As shown.
[0027] Back Figure 3In operation S206, processor 14 performs linear prediction coding (LPC) on each audio frame Fr generated in operation S204 to obtain formants. In LPC, formants are high-energy frequency peaks in the spectrum. Furthermore, the number of formants is determined by the order of the LPC equations. For simplicity, the LPC calculation process will be omitted.
[0028] During operation S208, processor 14 analyzes the formants of each audio frame Fr to obtain information such as the number, frequency, and bandwidth of the formants.
[0029] refer to Figure 5 , Figure 5 This diagram illustrates the windowing, LPC, and formant analysis performed on the audio frame Fr_m. The window function can be a known function such as a Gaussian window, Hann window, or Hamming window. In formant analysis, the start time of the audio frame Fr_m is tm. Furthermore, after LPC, multiple formants (e.g., six) are obtained from the audio frame Fr_m. Based on the frequency and bandwidth of the formants, the processor 14 can determine which formants correspond to speech. For example, when the frequency of a formant is greater than 250Hz and less than 3000Hz, and the bandwidth is less than 600Hz, the processor 14 determines that the formant corresponds to speech, i.e., it is a valid formant.
[0030] Back Figure 3 In operation S210, based on the analysis results of the formants of each audio frame Fr in operation S208, the processor 14 determines whether to combine adjacent audio frames Fr. When one or more audio frames Fr do not have formants corresponding to speech (effective formants), the processor 14 decides that the one or more audio frames do not need to be combined, and the process proceeds to operation S216. Conversely, in operation S212, when multiple consecutive audio frames Fr have effective formants, the processor 14 decides to combine these audio frames Fr into an audio segment SEG. In addition, the processor 14 groups the effective formants in the audio segment SEG to obtain the average frequency, average bandwidth, etc. of the effective formants in each group.
[0031] refer to Figure 6 , Figure 6The table displays formant information for the audio segment SEG_k formed by audio frames Fr_(m-2) to Fr_(m+2). In some embodiments, when the number of valid formants in N consecutive audio frames Fr is greater than 0, the first audio frame in the N consecutive audio frames Fr is assigned by processor 14 as the starting audio frame of the audio segment SEG. Furthermore, when the number of valid formants in M consecutive audio frames Fr following this starting audio frame is equal to 0, the preceding audio frame in the M consecutive audio frames Fr is assigned by processor 14 as the ending audio frame of the audio segment SEG. In some embodiments, N can be equal to M. In some embodiments, N is different from M. Figure 6 In this embodiment, M and N are set to 3, and the audio segment SEG_k comprises five audio frames Fr_(m-2) to Fr_(m+2). Furthermore, audio frame Fr_(m-2) is the starting audio frame, and audio frame Fr_(m+2) is the ending audio frame. In other words, the effective formants of the three consecutive audio frames Fr following audio frame Fr_(m+2) are 0.
[0032] like Figure 6 As shown, after obtaining all audio frames of audio segment SEG_k, processor 14 divides the effective formants of audio frames Fr_(m-2) to Fr_(m+2) into groups C1 and C2 based on the maximum number of effective formants, 2. Group C1 includes the first effective formants of audio frames Fr_(m-2) to Fr_(m+2), while group C2 includes the second effective formants of audio frames Fr_(m-2), Fr_(m-1), and Fr_(m+2). Next, processor 14 averages the frequency and bandwidth of each group to obtain the average frequencies Avg_C1 and Avg_C2, and the average bandwidths Avg_C1_bw and Avg_C2_bw for groups C1 and C2, respectively.
[0033] Please refer to this again. Figure 3 In operation S214, the processor 14 performs gain processing on the audio segment SEG based on the average effective formant value of each group obtained in operation S212.
[0034] Also refer to Figure 7A and Figure 7B , Figure 7A As described in one embodiment of this disclosure Figure 6 The flowchart for gain processing of the audio segment SEG_k, and Figure 7B It is a display Figure 7A A schematic diagram of the signals at each stage of the process. First, in operation S310, Figure 1Processor 14 converts the audio segment SEG_k into a spectrum (e.g., performs a Fourier transform), that is, converts it into a frequency domain signal. In operation S320, processor 14 generates corresponding gain values Gain_C1 and Gain_C2 based on the average frequencies Avg_C1 and Avg_C2 of groups C1 and C2, and their average bandwidths Avg_C1_bw and Avg_C2_bw, and applies these gain values to the spectrum of the audio segment SEG_k. In operation S330, processor 14 converts the audio segment SEG_k into a time domain signal. Therefore, in the audio segment SEG_k, processor 14 only performs gain processing on the effective formant average values of groups C1 and C2.
[0035] Back Figure 3 In operation S216, processor 14 combines the gain-processed audio segment SEG with the uncombined audio frame Fr to form audio data D2 (e.g., an audio stream). For example, as Figure 8 As shown, audio segment SEG_k is combined with audio frames Fr_(m-3) and Fr_(m+3) to form audio data D2. As previously described, uncombined audio frames (e.g., audio frames Fr_(m-3) and Fr_(m+3)) do not have formants corresponding to speech. In some embodiments, processor 14 provides audio data D3 to communication module 17 based on audio data D2 so that audio data D3 can be transmitted to other electronic devices for playback.
[0036] During operation S218, the audio output circuit 16 converts the audio data D2 into an audio output signal Sout and provides it to the broadcasting device 30.
[0037] According to speech enhancement method 200, electronic device 100 can detect and enhance formants in speech to strengthen vowel characteristics, making the enhanced speech more distinctive. Therefore, speech clarity can be improved, thereby enhancing speech recognition. For example, when electronic device 100 is a hearing aid, the speech of others around the user (e.g., an elderly person or a hearing-impaired person) becomes clearer through the hearing aid, making it easier for the wearer to recognize the content of the speech. Furthermore, speech enhancement method 200 can be applied to electronic products capable of performing speech recognition, enabling the electronic product to more accurately and quickly recognize the input speech content and perform corresponding operations.
[0038] Although the present invention has been described above with reference to preferred embodiments, it is not intended to limit the invention. Anyone skilled in the art, including those skilled in the art, may make some modifications and refinements without departing from the spirit and scope of the invention. Therefore, the scope of protection of the present invention shall be determined by the appended claims.
[0039] Symbol Explanation
[0040] 10: Voice enhancement device
[0041] 12: Audio Input Circuit
[0042] 13: Storage device
[0043] 14: Processor
[0044] 16: Audio output circuit
[0045] 17: Communication Module
[0046] 20: Radio device
[0047] 30: Broadcasting equipment
[0048] 100: Electronic devices
[0049] 200: Speech Enhancement Methods
[0050] D1-D3: Audio Data
[0051] F_0-F_n: Original audio frames
[0052] Fr_0-Fr_n, Fr_(m-3)-Fr_(m+3): audio frame
[0053] Gain_C1, Gain_C2: Gain values
[0054] LPC: Linear Predictive Coding
[0055] S202-S218, S310-S330: Operation
[0056] SEG_k: Audio segment
[0057] Sin: Audio input signal
[0058] Sout: Audio output signal
[0059] T_frame: Number of audio frames
[0060] T_hop: Audio frame jump
[0061] T_overlap: Audio frame overlap
Claims
1. A speech enhancement apparatus, comprising: an audio input circuit configured to convert an audio input signal into first audio data; and a processor configured to perform: generating a plurality of audio frames from the first audio data; performing formant analysis on the audio frames to determine whether to combine adjacent ones of the audio frames into an audio segment; performing gain processing on the audio segment including the combined audio frames; and combining the audio segment with one or more of the audio frames that are not combined into second audio data.
2. The speech enhancement apparatus of claim 1, wherein the processor performs formant analysis on each of the audio frames to obtain a number of formants for each of the audio frames.
3. The speech enhancement apparatus of claim 2, wherein a first one of the audio frames of a consecutive N number of the audio frames is a starting audio frame of the audio segment when the number of formants of the consecutive N number of the audio frames are all greater than zero.
4. The speech enhancement apparatus of claim 3, wherein a preceding one of the audio frames of a consecutive M number of the audio frames after the starting audio frame is an ending audio frame of the audio segment when the number of formants of the consecutive M number of the audio frames are all equal to zero.
5. The speech enhancement apparatus of claim 1, wherein the processor divides the formants of the audio frames of the audio segment into a plurality of groups according to a maximum number of the formants of the audio frames of the audio segment, and obtains an average frequency and an average bandwidth of the formants of each of the groups.
6. The speech enhancement apparatus of claim 5, wherein the processor performs gain on the audio segment according to the average frequency and the average bandwidth of the groups.
7. The speech enhancement apparatus of claim 1, wherein the processor performs sampling and windowing on the first audio data to generate the audio frames, wherein each of the audio frames is partially overlapped with a preceding one of the audio frames and a succeeding one of the audio frames.
8. The speech enhancement apparatus of claim 1, further comprising an audio output circuit configured to convert the second audio data into an audio output signal.
9. The speech enhancement apparatus of claim 1, further comprising a communication module configured to transmit the second audio data to an electronic device in a wired or wireless manner.
10. A speech enhancement method, comprising: converting an audio input signal into first audio data; generating a plurality of audio frames from the first audio data; performing formant analysis on the audio frames to determine whether to combine adjacent ones of the audio frames into an audio segment; performing gain processing on the audio segment including the combined audio frames; and combining the audio segment with one or more of the audio frames that are not combined into second audio data.