A method and apparatus for processing own voice of a hearing aid and a hearing aid device
Patent Information
- Application Number
- CN202610962677.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-30
- Publication Date
- 2026-08-18
AI Technical Summary
[0021]本发明提供的助听器自身语音的处理方法中,对助听器的输入语音进行低频校准后,检测符合自身语音特征的条件的声音作为自身语音,对自身语音的低频输出部分进行抑制,并将抑制的增益限制到环境背景声水平,从而有效地抑制通过空气传导至助听器的自身语音的增益,解决用户听感不适问题。
Smart Images

Figure CN122602047A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of sound acquisition and processing technology, and in particular to a method for processing the speech of a hearing aid, a speech processing device, and a hearing aid device including the speech processing device. Background Technology
[0002] When a user wears a hearing aid, their own voice is transmitted to the ear through the following two pathways: (1) air conduction: sound waves travel from the mouth through the air to the hearing aid and are amplified to reach the ear; (2) bone conduction: sound is transmitted directly to the cochlea through the vibration of the skull.
[0003] The amplified sound from a hearing aid may be much louder than the ambient sound, making the user feel that their voice is too loud, unnatural, or exaggerated, reducing their ability to understand language and causing a muffled, buzzing, or clogging sensation in their ears. These phenomena are particularly noticeable in users with normal low-frequency hearing but high-frequency hearing loss.
[0004] Therefore, there is an urgent need for a technical solution that can alleviate the problem caused by the amplification of one's own speech when it is transmitted through the air to the hearing aid. Summary of the Invention
[0005] In view of the above problems, the present invention proposes a method for processing the speech of a hearing aid itself, a speech processing device, and a hearing aid device including the speech processing device to overcome or at least partially solve the above problems.
[0006] One objective of this invention is to suppress the gain of the user's own speech transmitted through the air to the hearing aid, thereby solving the problem of hearing discomfort for the user.
[0007] A further objective of this invention is to improve the accuracy of its own speech detection.
[0008] Another further objective of the present invention is to improve the auditory experience after self-speech suppression.
[0009] In particular, according to one aspect of the present invention, a method for processing the speech of a hearing aid itself is provided, comprising: The input speech from the hearing aid is calibrated at low frequencies, and the calibrated speech signal is acquired. Sounds that meet the conditions of their own speech feature parameters in the collected speech signal are taken as their own speech, and the suppression gain of their own speech is calculated. Based on the ambient background noise in the input speech, the suppression gain of the speaker's own speech is limited to ensure that the level of the speaker's own speech does not exceed the level of the ambient background noise. The input speech is amplified and output by combining the suppression gain of the restricted speech and the set gain of the hearing aid.
[0010] Optionally, the steps of performing low-frequency calibration on the input speech of the hearing aid and acquiring the calibrated speech signal include: Calculate the calibration filter coefficients; The adjustment value for low-frequency calibration of the time-domain data sequence of the input speech is calculated based on the calibration filter coefficients; The input speech is calibrated using the adjustment value of low-frequency calibration to obtain the calibrated speech signal; The acquisition dataset consists of frame data of the calibrated speech signal. B .
[0011] Optionally, calculating the calibration filter coefficients includes: The calibration filter coefficients are calculated according to the following formula (1) based on the cutoff frequency of the calibration filter and the system sampling rate: (1) in, e It is a natural constant. π Pi fc The cutoff frequency, fs The system sampling rate, α To calibrate the filter coefficients; The adjustment values for low-frequency calibration of the time-domain data sequence of the input speech, calculated based on the calibration filter coefficients, include: The adjustment value for low-frequency calibration of the time-domain data sequence of the input speech is calculated according to the following formula (2): (2) in, This is the adjustment value for the current low-frequency calibration. The input speech is a time-domain data sequence. This is the adjustment value for the low-frequency calibration at the previous moment, and in the initial... At any time ; The input speech is calibrated using the adjustment value of low-frequency calibration, resulting in a calibrated speech signal including: The calibrated speech signal is obtained by calibrating the input speech using the adjustment value of low-frequency calibration according to the following formula (3): (3) in, c The set calibration adjustment constant, For the calibrated speech signal; The acquisition dataset consists of frame data of the calibrated speech signal. B include: Collect frame data of the calibrated speech signal and add each frame data to the acquired dataset. B Among these, when adding the latest frame data, the dataset is collected.B If the total number of elements exceeds the set cache limit, discard the collected dataset. B The oldest frame data is used to obtain the collected dataset represented by the following equation (4). B : (4) in, To collect datasets B Set the cache size.
[0012] Optionally, the step of detecting sounds in the acquired speech signal that meet the conditions of their own speech feature parameters as their own speech, and calculating the suppression gain of their own speech, includes: Statistical Data Collection B The zero-crossing count of mid-frame data per unit time is calculated, and the zero-crossing count is converted into an equivalent zero-crossing frequency. Based on the equivalent zero-crossing frequency, calculate the percentage of the high part of the speech signal that is higher than the preset target equivalent zero-crossing frequency and the percentage of the low part that is lower than the preset target equivalent zero-crossing frequency. Then, calculate the equivalent frequency suppression percentage based on the percentage of the high part and the percentage of the low part. The preset target equivalent zero-crossing frequency is the center value of the equivalent zero-crossing frequency of the speech signal itself. Calculate the collected dataset B The current average energy of the mid-frame data is used to calculate the energy suppression percentage of the speech signal based on the current average energy. The energy suppression percentage represents the percentage of speech in the speech signal that meets its own energy characteristics. The self-speech confidence of the speech signal is calculated based on the equivalent frequency suppression percentage and the energy suppression percentage. The suppression gain of one's own speech is calculated based on its own speech confidence and a preset expected suppression gain table.
[0013] Optionally, statistically collect datasets B The zero-crossing count of mid-frame data per unit time, and the conversion of the zero-crossing count into an equivalent zero-crossing frequency, includes: Collect the dataset according to the following formula (5). B Zero-crossing count of mid-frame data per unit time: (5) in, sign ( ) indicates taking the corresponding sign bit. To collect datasets B Set the cache size. This indicates a zero-crossing count, and n =1 The initial value is 0; and The zero-crossing count is converted into an equivalent zero-crossing frequency by normalization according to the following formula (6): (6) in, To convert the zero-crossing count to the equivalent zero-crossing frequency after normalization, fs The system sampling rate; Based on the equivalent zero-crossing frequency, calculate the percentage of the high portion of the speech signal that is higher than the preset target equivalent zero-crossing frequency and the percentage of the low portion that is lower than the preset target equivalent zero-crossing frequency, and calculate the equivalent frequency suppression percentage based on the high and low percentages, including: The percentage of the high portion of the speech signal that is higher than the preset target equivalent zero-crossing frequency is calculated according to the following formula (7). : (7) in, The preset target equivalent zero-crossing frequency, The preset equivalent zero-crossing frequency high value is the upper limit of the equivalent zero-crossing frequency of its own speech. The percentage of the lower part of the speech signal below the preset target equivalent zero-crossing frequency is calculated according to the following formula (8). : (8) in, The preset equivalent zero-crossing frequency low value is defined as the lower limit of the equivalent zero-crossing frequency of the speech itself; and The smoothed equivalent frequency suppression percentage is calculated according to the following formula (9) based on the high and low percentages: (9) in, This represents the current percentage of the smoothed equivalent frequency suppression. The percentage of equivalent frequency suppression after smoothing from the previous sampling point, when hour The value is 0; Calculate the collected dataset B The current average energy of the mid-frame data, and the percentage of energy suppression of the speech signal calculated based on the current average energy, include: The collected dataset is calculated according to the following formula (10). B Current average energy of Chinese speech signal : (10) The percentage of energy suppression after smoothing of the speech signal is calculated according to the following formula (11): (11) in, This represents the current smoothed percentage of energy suppression. The percentage of energy suppression after smoothing from the previous sampling point and when Its value is 0 at that time. To preset a low value for the logarithm of energy, The preset high value of the logarithm of energy, the preset low value of the logarithm of energy, and the preset high value of the logarithm of energy are respectively the lower limit and upper limit of the temporal energy of the speech itself; The self-speech confidence of a speech signal is calculated based on the equivalent frequency suppression percentage and the energy suppression percentage, including: The self-speech confidence of the speech signal is obtained by mixing and smoothing according to the following formulas (12) and (13): (12) (13) in, The confidence level of the mixture for the current sampling point. The confidence level of the speech after mixing and smoothing at the previous sampling point and when Its value is 0 at that time. To preset the smooth ascent parameter, To preset the smooth descent parameters, The confidence level of the current sampling point's own speech after mixing and smoothing; The suppression gain of one's own speech is calculated based on its own speech confidence and a preset expected suppression gain table, including: The suppression gain of the speech itself is calculated using the following formula (14): (14) in, To preset the trigger sensitivity parameters, To preset the desired suppression gain table, This is a suppression gain table of the speaker's own speech obtained through confidence transformation.
[0014] Optionally, the step of limiting the suppression gain of the input speech to ensure that the level of the input speech does not exceed the level of the ambient background noise includes: Frequency domain analysis is performed on the input speech to obtain the time-frequency domain signal of the input speech; Calculate the short-time sound level and long-time background noise level of the input speech based on the time-frequency domain signal; The level of sound change that can be suppressed relative to the long-term background noise is calculated based on the short-term sound level and the long-term background noise level. The suppression gain of one's own speech is limited to the level of sound change.
[0015] Optionally, frequency domain analysis of the input speech to obtain the time-frequency domain signal of the input speech includes: Frequency domain analysis is performed on the input speech, and the time-frequency domain signal of the input speech is obtained by the following equation (15): (15) in, For the length of the window, For the time-domain data sequence of the input speech in Data within the window length, This is the time-frequency domain signal whose logarithmic energy is obtained after frequency domain transformation; The calculation of short-time sound level and long-time background noise level of the input speech based on the time-frequency domain signal includes: The short-time sound level of the input speech is calculated using the following formula (16): (16) in, To preset the parameters for tracking short-term sound levels, The short-term sound level at the previous sampling point. The short-time sound level at the current sampling point, and when hour ;as well as The long-term background noise level of the input speech is calculated using the following formula (17): (17) in, To preset parameters for tracking long-term background noise levels, The long-term background sound level of the previous sampling point, The long-term background noise level at the current sampling point, and when hour ; The level of sound change that can be suppressed relative to long-term background noise is calculated based on short-term sound level and long-term background noise level, including: The level of sound change that can be suppressed by short-term sound relative to long-term background sound is calculated using the following formula (18). : (18) Limiting the suppression gain of one's own speech to below the level of sound variation includes: The suppression gain of the restricted self-speech is calculated using the following formula (19), ensuring that the level of the self-speech does not exceed the level of the long-term background noise: (19) in, The coefficients for tracking changes in speech are set. The suppression gain of the speech after the constraint of the previous sampling point and when Its value is 0 at that time. This is the suppression gain of the speech itself after the current sampling point is restricted.
[0016] Optionally, the steps of amplifying and outputting the input speech by combining the suppression gain of the restricted speech and the set gain of the hearing aid include: The real-time gain of the hearing aid is obtained by summing the set gain of the hearing aid with the suppressed gain of the speaker's own speech after limiting it. The input speech is amplified using real-time gain and then output.
[0017] Optionally, before identifying sounds in the acquired speech signal that meet the conditions of their own speech feature parameters as their own speech, the processing method further includes: The self-speech feature parameters are obtained, including the center value, upper limit value, and lower limit value of the equivalent zero-crossing frequency of the self-speech, as well as the upper limit value and lower limit value of the time-domain energy.
[0018] Optionally, the steps for obtaining one's own speech feature parameters include: Collect the audio signal of the user reading a given text aloud; From the collected speech signals, select sounds whose equivalent zero-crossing frequency is within a preset range and whose time-domain energy is higher than a preset threshold. The center, upper, and lower limits of the equivalent zero-crossing frequency of the selected sounds were statistically analyzed, as well as the upper and lower limits of the time-domain energy. If the statistical feature parameters meet the hearing aid usage effect, the center value, upper limit value and lower limit value of the equivalent zero-crossing frequency and the upper limit value and lower limit value of the time domain energy are saved as its own speech feature parameters to a fixed memory. Read its own speech feature parameters from a fixed memory.
[0019] According to another aspect of the present invention, a speech processing apparatus is also provided. The speech processing apparatus includes a memory, a processor, and a machine-executable program stored in the memory and running on the processor, wherein the processor, when executing the machine-executable program, implements any of the aforementioned methods for processing the speech of a hearing aid itself.
[0020] According to another aspect of the present invention, a hearing aid device is also provided, comprising a voice acquisition device, the aforementioned voice processing device, and a speaker connected in sequence.
[0021] In the hearing aid self-speech processing method provided by the present invention, after low-frequency calibration of the input speech of the hearing aid, the sound that meets the conditions of self-speech characteristics is detected as self-speech, the low-frequency output part of self-speech is suppressed, and the suppression gain is limited to the level of ambient background sound, thereby effectively suppressing the gain of self-speech transmitted to the hearing aid through the air and solving the problem of user hearing discomfort.
[0022] Furthermore, in the hearing aid speech processing method provided by the present invention, the equivalent frequency suppression percentage is calculated by calculating the percentage of the high portion of the collected speech data that is higher than the preset target equivalent zero-crossing frequency and the percentage of the low portion that is lower than the preset target equivalent zero-crossing frequency. At the same time, the energy suppression percentage of the collected speech data is calculated. Then, the self-speech confidence of the speech signal is calculated based on the equivalent frequency suppression percentage and the energy suppression percentage. And the self-speech suppression gain is calculated based on the self-speech confidence of the speech signal, thereby improving the detection accuracy and suppression accuracy of the self-speech.
[0023] Furthermore, in the hearing aid's own speech processing method provided by the present invention, by calculating the short-term sound level and long-term background sound level of the input speech, the theoretically suppressable sound change level of the short-term sound relative to the long-term background sound is obtained, thereby limiting the suppression gain of the own speech to the range of sound change, so that the hearing of the own speech is not abrupt and the hearing of the own speech after suppression is improved.
[0024] Furthermore, in the hearing aid's own speech processing method provided by the present invention, parameters are configured reasonably to minimize damage to other people's speech, thereby preserving better speech recognition after processing.
[0025] The technical solution of this invention realizes a lightweight, real-time, controllable, and low-power self-speech processing program, which is particularly suitable for hearing aids and other hearing aid devices.
[0026] The above and other objects, advantages and features of the present invention will become more apparent to those skilled in the art from the following detailed description of specific embodiments of the invention in conjunction with the accompanying drawings. Attached Figure Description
[0027] The following sections will describe some specific embodiments of the invention in detail by way of example and not limitation, with reference to the accompanying drawings. The same reference numerals in the drawings denote the same or similar parts or portions. Those skilled in the art should understand that these drawings are not necessarily drawn to scale. In the drawings: Figure 1 This is a flowchart illustrating a method for processing the speech of a hearing aid according to an embodiment of the present invention. Figure 2 This is a flowchart illustrating the input calibration and data acquisition steps according to an embodiment of the present invention; Figure 3 This is a flowchart illustrating the steps of calculating the suppression gain of one's own speech according to an embodiment of the present invention; Figure 4 This is a flowchart illustrating the steps of limiting the suppression gain of one's own speech according to an embodiment of the present invention; Figure 5This is a flowchart illustrating the steps of applying limited gain according to an embodiment of the present invention. Figure 6 This is a flowchart illustrating the steps of obtaining one's own speech feature parameters according to an embodiment of the present invention; Figure 7 This is a schematic diagram of the structure of a voice processing device according to an embodiment of the present invention; Figure 8 This is a schematic diagram of the structure of a hearing aid device according to an embodiment of the present invention. Detailed Implementation
[0028] Exemplary embodiments of the invention will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the invention are shown in the drawings, it should be understood that the invention may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the invention to those skilled in the art.
[0029] To address the numerous problems caused by hearing aids amplifying the user's own speech as it travels through the air, resulting in auditory discomfort, this invention proposes a method for processing the user's own speech in hearing aids.
[0030] Figure 1 This is a schematic flowchart illustrating a method for processing the speech of a hearing aid according to an embodiment of the present invention. See also... Figure 1 As shown, the hearing aid's own speech processing method in this embodiment may include at least the following steps S102 to S108.
[0031] Step S102: Perform low-frequency calibration on the input speech of the hearing aid and collect the calibrated speech signal.
[0032] Input speech can be captured by a speech acquisition device (such as a microphone) and fed into the hearing aid.
[0033] Step S104: Detect sounds in the collected speech signal that meet the conditions of their own speech feature parameters as their own speech, and calculate the suppression gain of their own speech.
[0034] Step S106: Based on the ambient background noise in the input speech, limit the suppression gain of the speech itself to ensure that the level of the speech does not exceed the level of the ambient background noise.
[0035] Step S109: Combine the suppression gain of the restricted speech with the set gain of the hearing aid to amplify the input speech before outputting it.
[0036] In the hearing aid self-speech processing method provided in this embodiment, after low-frequency calibration of the input speech of the hearing aid, the sound that meets the conditions of its own speech characteristics is detected as its own speech, the low-frequency output part of the self-speech is suppressed, and the suppression gain is limited to the level of the ambient background sound, thereby effectively suppressing the gain of the self-speech transmitted to the hearing aid through the air and solving the problem of user hearing discomfort.
[0037] Figure 2 This is a flowchart illustrating the input calibration and data acquisition steps according to an embodiment of the present invention. See also... Figure 2 As shown, in one embodiment, step S102 (input calibration and data acquisition step) may include steps S1021 to S1024.
[0038] Step S1021: Calculate the calibration filter coefficients.
[0039] In a specific embodiment, the calibration filter coefficients are calculated according to the following formula (1) based on the cutoff frequency of the calibration filter and the system sampling rate: (1) in, e It is a natural constant. π Pi fc The cutoff frequency, fs The system sampling rate, α To calibrate the filter coefficients.
[0040] Step S1022: Calculate the adjustment value of low-frequency calibration of the time-domain data sequence of the input speech based on the calibration filter coefficients.
[0041] In one specific embodiment, the adjustment value for low-frequency calibration of the time-domain data sequence of the input speech is calculated according to the following formula (2): (2) in, This is the adjustment value for the current low-frequency calibration. The input speech is a time-domain data sequence. This is the adjustment value for the low-frequency calibration at the previous moment, and in the initial... time, The time-domain data sequence of the input speech can be acquired by a speech acquisition device (such as a microphone).
[0042] Step S1023: Use the adjustment value of low-frequency calibration to perform low-frequency calibration on the input speech to obtain the calibrated speech signal.
[0043] In one specific embodiment, the calibrated speech signal is obtained by calibrating the input speech using the adjustment value of low-frequency calibration according to the following formula (3): (3) in, c The set calibration adjustment constant, This is the calibrated speech signal, used for subsequent calculations of its own speech suppression.
[0044] Set calibration adjustment constant c The value can be determined based on the type of voice acquisition device (such as a microphone) used, in order to flatten the response curve of the voice acquisition device.
[0045] Step S1024: Collect frame data of the calibrated speech signal to form a collection dataset. B .
[0046] In one specific embodiment, frame data of the calibrated speech signal is acquired, and each frame data is added to the acquired dataset. B As a dataset to be collected B Element.
[0047] In a further embodiment, a dataset is collected. B It has a set number of cache entries (cache size). That is, collecting datasets B Maximum capacity Each element. When adding the latest frame data (i.e., the latest element), the dataset is collected. B If the total number of elements exceeds the set cache limit, discard the collected dataset. B The oldest frame data (i.e., the oldest element) is used to obtain the collected dataset represented by the following equation (4). B : (4) Equation (4) represents discarding the collected dataset. B The oldest element in Central China B1 Simultaneously, the current frame data of the calibrated voice signal collected will be... Added to the dataset as the latest element B In this way, an updated collection dataset is obtained. .
[0048] Figure 3 This is a flowchart illustrating the steps of calculating the suppression gain of one's own speech according to an embodiment of the present invention. See also... Figure 3 As shown, in one embodiment, step S104 (the step of calculating the suppression gain of its own speech) may include steps S1041 to S1045.
[0049] Step S1041: Collect statistical datasets BThe zero-crossing count of the mid-frame data per unit time is calculated, and the zero-crossing count is converted into an equivalent zero-crossing frequency.
[0050] In a specific embodiment, firstly, the dataset is statistically collected according to the following formula (5). B Zero-crossing count of mid-frame data per unit time: (5) in, sign ( ) indicates taking the corresponding sign bit. To collect datasets B Set the cache size (i.e. the number of elements). This indicates a zero-crossing count, and n =1 The initial value is 0. Statistics are calculated through a loop. m -1 times the zero count is obtained.
[0051] Next, the zero-crossing count is converted into an equivalent zero-crossing frequency by normalization according to the following equation (6): (6) in, To convert the zero-crossing count to the equivalent zero-crossing frequency after normalization, fs This represents the system sampling rate.
[0052] Step S1042: Based on the equivalent zero-crossing frequency, calculate the percentage of the high portion of the speech signal that is higher than the preset target equivalent zero-crossing frequency and the percentage of the low portion that is lower than the preset target equivalent zero-crossing frequency, and calculate the equivalent frequency suppression percentage based on the high and low percentages. The preset target equivalent zero-crossing frequency is the center value of the equivalent zero-crossing frequency of the speech signal itself.
[0053] In one specific embodiment, the percentage of the high portion of the speech signal above the preset target equivalent zero-crossing frequency is calculated according to the following formula (7). : (7) in, The preset target equivalent zero-crossing frequency, The preset equivalent zero-crossing frequency high value is the upper limit of the equivalent zero-crossing frequency of its own speech.
[0054] The percentage of the lower part of the speech signal below the preset target equivalent zero-crossing frequency is calculated according to the following formula (8). : (8) in, The preset equivalent zero-crossing frequency is the lower limit of the equivalent zero-crossing frequency of its own speech.
[0055] In this embodiment, the equivalent zero-crossing frequency of the collected speech signal is close to a preset target equivalent zero-crossing frequency. The degree to which the effective zero-crossing frequency of the acquired speech signal is divided into a high portion above the preset target effective zero-crossing frequency and a low portion below the preset target effective zero-crossing frequency. The closer the effective zero-crossing frequency of the acquired speech signal is to the target effective zero-crossing frequency, the better. The higher the probability, the greater the likelihood that the speech is the expected speech itself. Meanwhile, speech with an equivalent zero-crossing frequency less than a preset low equivalent zero-crossing frequency or greater than a preset high equivalent zero-crossing frequency is considered not to be the expected speech itself.
[0056] Finally, the smoothed equivalent frequency suppression percentage is calculated according to the following formula (9) based on the high and low percentages: (9) in, This represents the current percentage of the smoothed equivalent frequency suppression. This represents the percentage of equivalent frequency suppression after smoothing from the previous sampling point, and when... hour The value is 0 (that is, when time The initial value is preset to 0.
[0057] Step S1043, calculate the collected dataset B The current average energy of the mid-frame data is used to calculate the energy suppression percentage of the speech signal. The energy suppression percentage represents the percentage of speech in the speech signal that meets its own energy characteristics.
[0058] In one specific embodiment, the collected dataset is calculated according to the following formula (10). B Current average energy of Chinese speech signal : (10) in To collect datasets B Set the cache size (i.e. the number of elements).
[0059] Next, the percentage of energy suppression after smoothing of the speech signal is calculated according to the following formula (11): (11) in, This represents the current smoothed percentage of energy suppression. The percentage of energy suppression after smoothing from the previous sampling point and when Its value is 0 (that is, when time The initial value is set to 0. To preset a low value for the logarithm of energy, The preset high value is the logarithmic energy value. The preset low value and preset high value are the lower and upper limits of the temporal energy of the speech itself, respectively.
[0060] Step S1044: Calculate the self-speech confidence of the speech signal based on the equivalent frequency suppression percentage and the energy suppression percentage.
[0061] In a specific embodiment, the self-speech confidence of the speech signal is calculated by mixing and smoothing according to the following equations (12) and (13): (12) (13) in, The confidence level of the mixture for the current sampling point. The confidence level of the speech after mixing and smoothing at the previous sampling point and when Its value is 0 (that is, when time The initial value is preset to 0). To preset the smooth ascent parameter, To preset the smooth descent parameters, This represents the confidence level of the speaker's own speech after mixing and smoothing at the current sampling point. The higher the confidence level of the speaker's own speech, the greater the probability that the speech is the expected speaker's own speech.
[0062] The preset smooth rise and preset smooth fall parameters can be set according to the user's listening preferences to adjust the response and release speed of speech suppression.
[0063] Step S1045: Calculate the suppression gain of your own speech based on your own speech confidence and the preset expected suppression gain table.
[0064] In a specific embodiment, the suppression gain of the speech itself is calculated using the following formula (14): (14) in, To preset the trigger sensitivity parameters, To preset the desired suppression gain table, This is a suppression gain table of the speaker's own speech obtained through confidence transformation.
[0065] Preset Desired Suppression Gain Table Specifically, it could be a low-frequency gain table with preset desired suppression, which can be set based on empirical values and adjusted during the fitting of hearing aids to make the output speech more suitable for the user's listening experience needs.
[0066] The preset trigger sensitivity parameter can be set according to actual application requirements. Generally, it can be set to an integer in the range of greater than or equal to 0 and less than 5.
[0067] In the hearing aid speech processing method provided in this embodiment of the invention, the equivalent frequency suppression percentage is calculated by calculating the percentage of the high portion of the collected speech data that is higher than the preset target equivalent zero-crossing frequency and the percentage of the low portion that is lower than the preset target equivalent zero-crossing frequency. At the same time, the energy suppression percentage of the collected speech data is calculated. Then, the self-speech confidence of the speech signal is calculated based on the equivalent frequency suppression percentage and the energy suppression percentage. And the self-speech suppression gain is calculated based on the self-speech confidence of the speech signal, thereby improving the detection accuracy and suppression accuracy of the self-speech.
[0068] Figure 4 This is a flowchart illustrating the steps of limiting the suppression gain of one's own speech according to an embodiment of the present invention. See also... Figure 4 As shown, in one embodiment, step S106 (the step of limiting the suppression gain of its own speech) may include steps S1061 to S1064.
[0069] Step S1061: Perform frequency domain analysis on the input speech to obtain the time-frequency domain signal of the input speech.
[0070] In a specific embodiment, frequency domain analysis is performed on the input speech, and the time-frequency domain signal of the input speech is obtained by the following equation (15): (15) in, For the length of the window, For the time-domain data sequence of the input speech in Data within the window length, This is the time-frequency domain signal whose logarithmic energy is obtained after frequency domain transformation.
[0071] Step S1062: Calculate the short-time sound level and long-time background noise level of the input speech based on the time-frequency domain signal.
[0072] In one specific embodiment, the short-time sound level of the input speech is calculated using the following formula (16): (16) in, To preset the parameters for tracking short-term sound levels, The short-term sound level at the previous sampling point. The short-time sound level at the current sampling point, and when The time is initialized as follows: .
[0073] Preset parameters for tracking short-term sound levels It can be set according to the actual application requirements, and can generally be set to a value in the range of 0 to 1.
[0074] The long-term background noise level of the input speech is calculated using the following formula (17): (17) in, To preset parameters for tracking long-term background noise levels, The long-term background sound level of the previous sampling point, The long-term background noise level at the current sampling point, and when The time is initialized as follows: .
[0075] Preset parameters for tracking long-term background noise levels It can be set according to the actual application requirements, and can generally be set to a value in the range of 0 to 1.
[0076] Step S1063: Calculate the level of sound change that the short-time sound can suppress relative to the long-time background sound based on the short-time sound level and the long-time background sound level.
[0077] In one specific embodiment, the level of suppressible sound change of short-time sound relative to long-time background sound is calculated by the following formula (18). : (18).
[0078] Step S1064: Limit the suppression gain of your own speech to below the level of sound change.
[0079] In one specific embodiment, the suppression gain of the restricted self-speech is calculated using the following formula (19) to ensure that the level of the self-speech does not exceed the level of the long-term background noise: (19) in, The coefficients for tracking changes in speech are set. The suppression gain of the speech after the constraint of the previous sampling point and when Its value is 0 (that is, when time The initial value is preset to 0). This is the suppression gain of the speech itself after the current sampling point is restricted.
[0080] The set coefficients for speech sound tracking changes The value can be determined based on the adjustment of the hearing aid, and it can generally be set to a value in the range of 0 to 1.
[0081] In this applicationn and ω These represent the sampled frame sequence values and the angular frequency sequence values at the frequency domain resolution, respectively.
[0082] In the hearing aid speech processing method provided in this embodiment of the invention, the short-time sound level and long-time background sound level of the input speech are calculated to obtain the theoretically suppressable sound change level of the short-time sound relative to the long-time background sound. In this way, the suppression gain of the speech is limited to the range of sound change, so that the speech does not sound abrupt and the listening experience after the speech is suppressed is improved.
[0083] Figure 5 This is a flowchart illustrating the steps of applying limited gain according to an embodiment of the present invention. See also... Figure 5 As shown, in one embodiment, step S108 (the step of applying the limited gain) may include steps S1081 to S1082.
[0084] Step S1081: The real-time gain of the hearing aid is obtained by summing the set gain of the hearing aid with the suppressed gain of the speech after limitation.
[0085] Specifically, the real-time gain of the hearing aid is calculated according to the following formula (20): (20) in, It is the set gain of the hearing aid, which is the gain value configured for the user under normal conditions; It is the real-time gain value of the hearing aid after limiting its own speech gain.
[0086] Step S1082: Amplify the input speech using real-time gain and then output it.
[0087] Specifically, the time-frequency domain signal after the hearing aid restricts its own speech output is obtained according to the following formula (21). : (twenty one) in, The input speech signal is in the time-frequency domain.
[0088] Then, frequency domain synthesis tools can be applied to... Restored to time domain signal For example, the output can be played through a speaker.
[0089] This completes the entire process of processing the sound signal from the hearing aid's input limiting amplification of its own speech to the output of the speaker.
[0090] In some embodiments, prior to step S104, the method for processing the hearing aid's own speech may further include: acquiring its own speech feature parameters. Specifically, the own speech feature parameters may include the center value, upper limit value, and lower limit value of the equivalent zero-crossing frequency of the own speech, as well as the upper limit value and lower limit value of the time-domain energy.
[0091] Optionally, the aforementioned self-speech feature parameters can be obtained by learning the self-speech feature parameters, or the learned self-speech feature parameters can be used as a reference for adjustment and configuration.
[0092] Figure 6 This is a flowchart illustrating the steps of obtaining one's own speech feature parameters according to an embodiment of the present invention. See also... Figure 6 As shown, in a specific embodiment, the step of obtaining one's own speech feature parameters may include the following steps S1101 to S1105.
[0093] Step S1101: Collect the voice signal of the user reading the given text.
[0094] Before collecting speech signals, a process to learn one's own speech characteristics can be triggered, thus initiating the ear-blocking test. During the test, the user's speech signal is collected as they read a given text. The given text can be pre-prepared audio text that is likely to induce ear blockage.
[0095] Step S1102: Select sounds from the collected speech signals whose equivalent zero-crossing frequency is within a preset range and whose time-domain energy is higher than a preset threshold.
[0096] The preset range is a preset initial screening range, typically set to, for example, between 0 and 4000 Hz.
[0097] The preset threshold is a pre-defined initial screening threshold. Generally, it can be set to, for example, 40 dBSPL, so that speech with a speech energy of 40 dBSPL or higher can be screened out.
[0098] Optionally, after signal filtering, it can be determined whether valid speech that meets the requirements (i.e., the equivalent zero-crossing frequency is within the preset range and the time domain energy is higher than the preset threshold) has been successfully filtered out; if not, return to step S1101, re-acquire the speech signal of the user reading the given text and filter it; if yes, execute step S1103.
[0099] Step S1103: Statistically calculate the center value (i.e., average value), upper limit value (i.e., minimum value), and lower limit value (i.e., maximum value) of the equivalent zero-crossing frequency of the selected sound, as well as the upper limit value (i.e., maximum value) and lower limit value (i.e., minimum value) of the time-domain energy.
[0100] Optionally, after performing the statistics, it can be determined whether all parameter statistics have been completed; if not, return to step S1103 to continue performing the statistics; if yes, proceed to the next step.
[0101] After obtaining the statistical characteristic parameters, these newly obtained statistical characteristic parameters can be used to conduct user voice tests to verify the effectiveness of the hearing aid. If the user feedback is satisfactory (e.g., no occlusion, no jarring linearity, natural hearing, etc.), then it is determined that the statistical characteristic parameters meet the requirements for the effectiveness of the hearing aid.
[0102] Step S1104: If the statistical feature parameters meet the requirements for hearing aid performance, save the center value, upper limit value, and lower limit value of the equivalent zero-crossing frequency, as well as the upper limit value and lower limit value of the time-domain energy, as its own speech feature parameters to a fixed memory.
[0103] Step S1105: Read the speech feature parameters of the device itself from the fixed memory.
[0104] If the statistically analyzed feature parameters do not meet the requirements for hearing aid performance, these parameters can be discarded and not stored. Then, after re-initialization, the testing and learning of the user's own speech feature parameters can be performed again.
[0105] Alternatively, if the statistical characteristic parameters do not meet the requirements for hearing aid performance, the statistical characteristic parameters can be adjusted (e.g., manually) before conducting a performance test.
[0106] The hearing aid's own speech processing method of the present invention realizes a lightweight, real-time, controllable, and low-power self-speech processing program, which alleviates the problem caused by the self-speech being conducted through the air to the hearing aid for amplification. At the same time, by reasonably configuring parameters, it minimizes damage to other people's speech, thereby preserving better speech recognition after processing.
[0107] Based on the same technical concept, the present invention also provides a voice processing device 700. Figure 7 This is a schematic diagram of the structure of a voice processing device 700 according to an embodiment of the present invention.
[0108] like Figure 7 As shown, the voice processing device 700 may include a memory 720, a processor 710, and a machine-executable program 721 stored in the memory 720 and running on the processor 710. When the processor 710 executes the machine-executable program 721, it implements the hearing aid's own voice processing method of any of the above embodiments.
[0109] Based on the same technical concept, the present invention also provides a hearing aid device 800. Figure 8 This is a schematic diagram of the structure of a hearing aid device 800 according to an embodiment of the present invention.
[0110] like Figure 8 As shown, the hearing aid device 800 may include a voice acquisition device 810, a voice processing device 700, and a speaker 820 connected in sequence.
[0111] The speech processing device 700 may be any of the speech processing devices 700 of the foregoing embodiments, which executes the speech processing method steps of the hearing aid itself in any of the foregoing embodiments.
[0112] In some embodiments, the hearing aid device 800 can be any device with sound processing capabilities, including but not limited to hearing aids, assistive hearing devices, smart hearing glasses, etc.
[0113] It should be noted that the logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be specifically implemented in any computer program product for use by, or in conjunction with, instruction execution systems, apparatus or devices (such as computer-based systems, processor-included systems or other systems that can fetch and execute instructions from, an instruction execution system, apparatus or device).
[0114] It should be understood that various parts of the present invention can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented using software or firmware stored in memory and executed by a suitable instruction execution system.
[0115] The processor 710 can be a single-core processor, a multi-core processor, a computing cluster, or any other configuration. The memory 720 can include random access memory (RAM), read-only memory, flash memory, or any other suitable storage system.
[0116] The flowchart provided in this embodiment is not intended to indicate that the operations of the method will be performed in any particular order, or that all operations of the method are included in every case. Furthermore, the method may include additional operations. Within the scope of the technical concept provided by the method in this embodiment, additional variations can be made to the above method.
[0117] Therefore, those skilled in the art should recognize that although numerous exemplary embodiments of the present invention have been shown and described in detail herein, many other variations or modifications conforming to the principles of the present invention can be directly determined or derived from the disclosure of the present invention without departing from the spirit and scope of the invention. Thus, the scope of the present invention should be understood and construed as covering all such other variations or modifications.
Claims
1. A method for processing the speech of a hearing aid itself, comprising: The input speech from the hearing aid is calibrated at low frequencies, and the calibrated speech signal is acquired. Sounds that meet the conditions of their own speech feature parameters in the collected speech signal are taken as their own speech, and the suppression gain of their own speech is calculated. Based on the ambient background noise in the input speech, the suppression gain of the self-speech is limited to ensure that the level of the self-speech does not exceed the level of the ambient background noise; The input speech is amplified and output by combining the suppression gain of the restricted speech and the set gain of the hearing aid.
2. The method for processing the hearing aid's own speech according to claim 1, wherein, The steps for performing low-frequency calibration on the input speech of the hearing aid and acquiring the calibrated speech signal include: Calculate the calibration filter coefficients; The adjustment value for low-frequency calibration of the time-domain data sequence of the input speech is calculated based on the calibration filter coefficients. The input speech is calibrated using the adjustment value of the low-frequency calibration to obtain a calibrated speech signal; The collected data set consists of frame data of the calibrated speech signal. B .
3. The method for processing the hearing aid's own speech according to claim 2, wherein, The calculation of the calibration filter coefficients includes: The calibration filter coefficients are calculated according to the following formula (1) based on the cutoff frequency of the calibration filter and the system sampling rate: (1) in, e It is a natural constant. π Pi fc The cutoff frequency, fs The system sampling rate, α To calibrate the filter coefficients; The step of calculating the low-frequency calibration adjustment value of the time-domain data sequence of the input speech based on the calibration filter coefficients includes: The adjustment value for low-frequency calibration of the time-domain data sequence of the input speech is calculated according to the following formula (2): (2) in, This is the adjustment value for the current low-frequency calibration. The input speech is a time-domain data sequence. This is the adjustment value for the low-frequency calibration at the previous moment, and in the initial... At any time ; The step of performing low-frequency calibration on the input speech using the adjustment value of the low-frequency calibration to obtain the calibrated speech signal includes: The calibrated speech signal is obtained by calibrating the input speech using the adjustment value of the low-frequency calibration according to the following formula (3): (3) in, c The set calibration adjustment constant, For the calibrated speech signal; The frame data of the calibrated speech signal acquired constitute the acquisition dataset. B include: Collect frame data of the calibrated speech signal and add each frame data to the collected dataset. B Among them, when adding the latest frame data, in the acquired dataset B If the total number of elements exceeds the set cache limit, the collected dataset will be discarded. B The oldest frame data is used to obtain the collected dataset represented by the following equation (4). B : (4) in, For the collected dataset B Set the cache size.
4. The method for processing the hearing aid's own speech according to claim 2, wherein, The steps of detecting sounds in the acquired speech signal that meet the conditions of their own speech feature parameters as their own speech, and calculating the suppression gain of their own speech, include: Statistics on the collected dataset B The zero-crossing count of the mid-frame data per unit time is calculated, and the zero-crossing count is converted into an equivalent zero-crossing frequency. Based on the equivalent zero-crossing frequency, calculate the percentage of the high portion of the speech signal that is higher than the preset target equivalent zero-crossing frequency and the percentage of the low portion that is lower than the preset target equivalent zero-crossing frequency, and calculate the equivalent frequency suppression percentage based on the percentage of the high portion and the percentage of the low portion. The preset target equivalent zero-crossing frequency is the center value of the equivalent zero-crossing frequency of its own speech. Calculate the collected dataset B The current average energy of the mid-frame data is used to calculate the energy suppression percentage of the speech signal based on the current average energy. The energy suppression percentage represents the percentage of speech in the speech signal that satisfies the energy characteristics of its own speech. The self-speech confidence of the speech signal is calculated based on the equivalent frequency suppression percentage and the energy suppression percentage; The suppression gain of the speech is calculated based on the self-speech confidence level and the preset expected suppression gain table.
5. The method for processing the hearing aid's own speech according to claim 4, wherein, The statistics mentioned above are collected datasets. B The zero-crossing count of mid-frame data per unit time, and the conversion of the zero-crossing count into an equivalent zero-crossing frequency, includes: The collected dataset is statistically analyzed according to the following formula (5). B Zero-crossing count of mid-frame data per unit time: (5) in, sign ( ) indicates taking the corresponding sign bit. For the collected dataset B Set the cache size. This indicates a zero-crossing count, and n =1 The initial value is 0; and The zero-crossing count is converted into an equivalent zero-crossing frequency by normalization according to the following formula (6): (6) in, To convert the zero-crossing count to the equivalent zero-crossing frequency after normalization, fs The system sampling rate; The step of calculating the percentage of the high portion of the speech signal that is higher than the preset target equivalent zero-crossing frequency and the percentage of the low portion that is lower than the preset target equivalent zero-crossing frequency based on the equivalent zero-crossing frequency, and calculating the equivalent frequency suppression percentage based on the high portion percentage and the low portion percentage includes: The percentage of the high portion of the speech signal that is higher than the preset target equivalent zero-crossing frequency is calculated according to the following formula (7). : (7) in, The preset target is the equivalent zero-crossing frequency. The preset equivalent zero-crossing frequency high value is the upper limit of the equivalent zero-crossing frequency of its own speech. The percentage of the lower portion of the speech signal below the preset target equivalent zero-crossing frequency is calculated according to the following formula (8). : (8) in, The preset equivalent zero-crossing frequency low value is defined as the lower limit of the equivalent zero-crossing frequency of the speech itself; and The smoothed equivalent frequency suppression percentage is calculated according to the following formula (9) based on the high percentage and the low percentage: (9) in, This represents the current percentage of the smoothed equivalent frequency suppression. The percentage of equivalent frequency suppression after smoothing from the previous sampling point, when hour The value is 0; The calculation of the collected dataset B The current average energy of the mid-frame data, and the calculation of the energy suppression percentage of the speech signal based on the current average energy, include: The collected dataset is calculated according to the following formula (10). B Current average energy of Chinese speech signal : (10) The percentage of energy suppression after smoothing of the speech signal is calculated according to the following formula (11): (11) in, This represents the current smoothed percentage of energy suppression. The percentage of energy suppression after smoothing from the previous sampling point and when Its value is 0 at that time. To preset a low value for the logarithm of energy, The preset high value of the logarithm of energy is defined as the preset low value of the logarithm of energy, and the preset high value of the logarithm of energy is defined as the lower limit and upper limit of the temporal energy of the speech itself, respectively. The calculation of the self-speech confidence of the speech signal based on the equivalent frequency suppression percentage and the energy suppression percentage includes: The self-speech confidence of the speech signal is obtained by mixing and smoothing according to the following formulas (12) and (13): (12) (13) in, The confidence level of the mixture for the current sampling point. The confidence level of the speech after mixing and smoothing at the previous sampling point and when Its value is 0 at that time. To preset the smooth ascent parameter, To preset the smooth descent parameters, The confidence level of the current sampling point's own speech after mixing and smoothing; The step of calculating the suppression gain of one's own speech based on the self-speech confidence level and a preset expected suppression gain table includes: The suppression gain of the self-speech is calculated using the following formula (14): (14) in, To preset the trigger sensitivity parameters, To preset the desired suppression gain table, This is a suppression gain table of the self-speech obtained through confidence conversion.
6. The method for processing the speech of a hearing aid itself according to claim 1, wherein, The step of limiting the suppression gain of the speaker's own speech to such that the level of the speaker's own speech does not exceed the level of the ambient background noise in the input speech includes: The input speech is analyzed in the frequency domain to obtain the time-frequency domain signal of the input speech; Calculate the short-time sound level and long-time background noise level of the input speech based on the time-frequency domain signal; Calculate the level of sound change that can be suppressed relative to the long-term background noise based on the short-term sound level and the long-term background noise level; The suppression gain of the self-speech is limited below the level of the sound change.
7. The method for processing the hearing aid's own speech according to claim 6, wherein, The step of performing frequency domain analysis on the input speech to obtain the time-frequency domain signal of the input speech includes: The input speech is analyzed in the frequency domain, and the time-frequency domain signal of the input speech is obtained by the following formula (15): (15) in, For the length of the window, The time-domain data sequence of the input speech in Data within the window length, This is the time-frequency domain signal whose logarithmic energy is obtained after frequency domain transformation; The calculation of the short-time sound level and long-time background noise level of the input speech based on the time-frequency domain signal includes: The short-time sound level of the input speech is calculated using the following formula (16): (16) in, To preset the parameters for tracking short-term sound levels, The short-term sound level at the previous sampling point. The short-time sound level at the current sampling point, and when hour ;as well as The long-term background noise level of the input speech is calculated using the following formula (17): (17) in, To preset parameters for tracking long-term background noise levels, The long-term background sound level of the previous sampling point, The long-term background noise level at the current sampling point, and when hour ; The step of calculating the suppressible sound change level of short-time sound relative to long-time background sound based on the short-time sound level and the long-time background sound level includes: The level of sound change that can be suppressed by short-term sound relative to long-term background sound is calculated using the following formula (18). : (18) The step of limiting the suppression gain of the speech itself below the level of sound change includes: The suppression gain of the restricted self-speech is calculated using the following formula (19), such that the level of the self-speech does not exceed the level of the long-term background noise: (19) in, The coefficients for tracking changes in speech are set. The suppression gain of the speech after the constraint of the previous sampling point and when Its value is 0 at that time. This is the suppression gain of the speech itself after the current sampling point is restricted.
8. The method for processing the speech of a hearing aid itself according to claim 1, wherein, The steps of amplifying and outputting the input speech by combining the suppression gain of the restricted speech and the set gain of the hearing aid include: The real-time gain of the hearing aid is obtained by summing the set gain of the hearing aid with the suppressed gain of the speech after limiting it. The input speech is amplified using the real-time gain and then output.
9. The method for processing the speech of a hearing aid itself according to claim 5, wherein, Before identifying sounds in the collected speech signal that meet the conditions of their own speech feature parameters as their own speech, the processing method further includes: The self-speech feature parameters are obtained, including the center value, upper limit value and lower limit value of the equivalent zero-crossing frequency of the self-speech, as well as the upper limit value and lower limit value of the time-domain energy.
10. The method for processing the speech of a hearing aid itself according to claim 9, wherein, The steps to obtain one's own speech feature parameters include: Collect the audio signal of the user reading a given text aloud; From the collected speech signals, select sounds whose equivalent zero-crossing frequency is within a preset range and whose time-domain energy is higher than a preset threshold. The center, upper, and lower limits of the equivalent zero-crossing frequency of the selected sounds were statistically analyzed, as well as the upper and lower limits of the time-domain energy. If the statistical feature parameters meet the hearing aid usage effect, the center value, upper limit value and lower limit value of the equivalent zero-crossing frequency and the upper limit value and lower limit value of the time domain energy are saved as their own speech feature parameters to a fixed memory. The self-voice feature parameters are read from the fixed memory.
11. A speech processing apparatus, comprising a memory, a processor, and a machine-executable program stored in the memory and running on the processor, wherein the processor, when executing the machine-executable program, implements a method for processing the speech of a hearing aid itself according to any one of claims 1 to 10.
12. A hearing aid device, comprising a speech acquisition device, a speech processing device, and a speaker connected in sequence, wherein, The speech processing device is the speech processing device according to claim 11.