Voice data processing method and binaural hearing aid

By collecting speech data through the microphone array of binaural hearing aids, extracting and comparing voiceprint features, the problem of misjudgment in traditional self-speaking voice recognition is solved, achieving more accurate self-speaking voice recognition and enhancement processing, and improving the user experience.

CN121924430APending Publication Date: 2026-04-24ANKER INNOVATIONS TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411497312.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-10-24
Publication Date
2026-04-24

AI Technical Summary

Technical Problem

Traditional self-speaking voice recognition methods, which rely on spatial cues of the self-speaking voice, are prone to misjudgment, resulting in inaccurate self-speaking judgments and a poor user experience.

Method used

The speech data processing method is adopted to acquire speech data by the microphone array of the binaural hearing aid, extract voiceprint features, and compare them with the pre-stored user standard voiceprint features to determine whether it is self-spoken speech, and perform enhancement processing to reduce the feeling of ear blockage and sound distortion.

Benefits of technology

It improves the accuracy of voice recognition, reduces misrecognition caused by environmental noise or background sounds, enhances the user's auditory experience, and maintains a high recognition rate even in noisy environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121924430A_ABST
    Figure CN121924430A_ABST
Patent Text Reader

Abstract

The invention relates to a voice data processing method and a binaural hearing aid. The method is applied to the binaural hearing aid, the binaural hearing aid comprises a first hearing aid and a second hearing aid, and the method comprises the following steps: respectively acquiring voice data acquired by the first hearing aid and the second hearing aid, extracting voiceprint features in the voice data under the condition of judging that potential self-talking voice exists in the voice data, the voiceprint features are compared with stored standard voiceprint features of the user, a comparison result is obtained, the standard voiceprint features are obtained by processing voice sample data of the user, and the voice sample data comprise near-talking voice data collected by a first hearing aid close to the mouth of the user; and if the comparison result shows that the self-talking voice exists in the voice data, performing enhancement processing on the voice data. By adopting the method, more accurate self-talking voice detection can be realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of audio processing technology, and in particular to a speech data processing method and a binaural hearing aid. Background Technology

[0002] Binaural hearing aids are hearing-assistive devices designed to improve hearing by using two hearing aids simultaneously, with both aids sharing microphone information. Because hearing aid users often complain about issues such as excessive volume when hearing their own voice (self-speaking), a feeling of occlusion, and severe distortion of their own voice, self-speaking sound optimization solutions have emerged to reduce this occlusion and improve the self-speaking experience. Self-speaking sound optimization primarily involves detecting the user's own air conduction sound and processing it through algorithms to identify whether the received sound is the wearer's own voice or an external sound. This results in "reduced self-speaking sound," eliminating the common occlusion sensation associated with hearing aids, providing a clearer and more comfortable self-speaking experience.

[0003] However, traditional self-speaking voice recognition is usually based on spatial cues of the self-speaking voice, which is prone to errors in self-speaking judgment. Summary of the Invention

[0004] Therefore, it is necessary to provide a speech data processing method and a binaural hearing aid with higher accuracy in self-talk judgment to address the above-mentioned technical problems.

[0005] In a first aspect, this application provides a speech data processing method applied to a binaural hearing aid, the binaural hearing aid comprising a first hearing aid and a second hearing aid, including:

[0006] Acquire speech data collected by the first hearing aid and the second hearing aid respectively;

[0007] If it is determined that there is potential self-speaking speech in the speech data, the voiceprint features in the speech data are extracted;

[0008] The voiceprint features are compared with the stored standard voiceprint features of the user to obtain the comparison result. The standard voiceprint features are obtained by processing the user's speech sample data, which includes near speech data collected by a first hearing aid close to the user's mouth.

[0009] If the comparison result indicates that the speech data contains self-speaking speech, then the speech data is enhanced.

[0010] Secondly, this application provides a binaural hearing aid, including a first hearing aid, a second hearing aid, and a processor, wherein the processor is connected to both the first hearing aid and the second hearing aid, wherein:

[0011] The first microphone array in the first hearing aid and the second microphone array in the second hearing aid collect speech data and send the collected speech data to the processor, which then executes the steps of the speech data processing method described in any one of the above descriptions.

[0012] The aforementioned speech data processing method and binaural hearing aid differ from traditional self-speaking detection methods based on spatial cues of self-speaking sound. This method introduces near-speaking speech data and voiceprint verification. First, a first hearing aid placed close to the user's mouth collects clearer near-speaking speech data, extracting and storing the user's standard voiceprint features. In subsequent applications, the system continuously checks whether potential self-speaking sounds exist in the speech data collected by the first and second hearing aids. If potential self-speaking sounds are detected, a voiceprint verification process is performed, extracting voiceprint features from the speech data and comparing them with the stored standard voiceprint features to further accurately determine if self-speaking sounds exist. If self-speaking sounds are detected, the speech data is enhanced to reduce ear blockage and sound distortion, improving the user's auditory experience. The entire solution does not rely on specific spatial orientation information for self-talk detection. It only needs to compare the voiceprint features in the real-time collected voice data with the user's standard voiceprint features. Since the user's standard voiceprint features are extracted from near-speech data containing clear and distinct user voice characteristics, the presence of self-talk can be accurately detected by comparing the voiceprint features with the user's standard voiceprint features. This reduces misidentification caused by environmental noise or background noise, maintains a high recognition rate even in noisy environments, and improves the robustness of the system. Attached Figure Description

[0013] To more clearly illustrate the technical solutions in the embodiments of this application or related technologies, the drawings used in the description of the embodiments of this application or related technologies will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0014] Figure 1 This is an application environment diagram of a voice data processing method in one embodiment;

[0015] Figure 2 This is a flowchart illustrating a voice data processing method in one embodiment;

[0016] Figure 3 This is a flowchart illustrating the self-voiceprint registration step in one embodiment;

[0017] Figure 4 This is a flowchart illustrating the step of determining whether there is potential self-talk in the voice data in one embodiment;

[0018] Figure 5 This is a flowchart illustrating the step of determining whether there is potential self-talk in the voice data in another embodiment;

[0019] Figure 6 This is a flowchart illustrating the voice data processing method in another embodiment;

[0020] Figure 7 This is a structural block diagram of a voice data processing device in one embodiment;

[0021] Figure 8 This is a structural block diagram of the voice data processing device in another embodiment;

[0022] Figure 9 This is an internal structural diagram of a computer device in one embodiment. Detailed Implementation

[0023] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0024] The voice data processing method provided in this application embodiment can be applied to, for example... Figure 1 In the application environment shown, the hearing aid 102, such as a binaural hearing aid including a first hearing aid and a second hearing aid, communicates with the server 104 via a network. The data storage system can store the data that the server 104 needs to process, including the user's standard voiceprint features. These standard voiceprint features are obtained by processing near-speech speech data collected by a microphone array located near the user's mouth. The data storage system can be integrated onto the server 104 or placed in the cloud or on another network server. Specifically, the binaural hearing aids can acquire speech data collected by the first and second hearing aids respectively, and send the acquired speech data to the server 104 in real time. The server 104 determines in real time whether there might be potential self-speech in the speech data. If it determines that there is potential self-speech, a voiceprint verification process is performed to extract the voiceprint features from the speech data. Then, the extracted voiceprint features are compared with the stored standard voiceprint features of the user to obtain a comparison result. If the comparison result indicates that there is self-speech in the speech data, the speech data is enhanced so that the user can clearly hear their own voice.

[0025] Understandably, this method can also be applied to hearing aids such as binaural hearing aids or other intelligent devices with computing capabilities, as well as to systems that include binaural hearing aids and servers, and can be implemented through the interaction between binaural hearing aids and servers.

[0026] Hearing assistive devices include, but are not limited to, headphones and hearing aids. Smart devices can be, but are not limited to, various personal computers, laptops, smartphones, tablets, IoT devices, and portable wearable devices. IoT devices can include smart speakers, smart TVs, smart air conditioners, smart in-vehicle systems, and projection devices. Portable wearable devices can include smartwatches, smart bracelets, and head-mounted displays. Head-mounted displays can be virtual reality (VR) devices, augmented reality (AR) devices, and smart glasses. Server 104 can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services.

[0027] In one exemplary embodiment, such as Figure 2 As shown, a voice data processing method is provided, which can be applied to... Figure 1 Taking the binaural hearing aid 102 as an example, in this embodiment, the method includes steps S200 to S800, wherein:

[0028] S200 acquires speech data collected by the first hearing aid and the second hearing aid, respectively.

[0029] In this embodiment, the first hearing aid and the second hearing aid are hearing aids worn on or near the user's ears, respectively, and can also be referred to as the left ear hearing aid and the first hearing aid. In practical applications, it can be a binaural hearing aid where the microphone arrays in the left ear hearing aid and the first hearing aid each include two or more microphones. The microphone arrays in the left ear hearing aid and the first hearing aid continuously collect speech data, including environmental noise and human voice signals, and send the collected speech data to the processor or control module of the binaural hearing aid. It is understood that the speech data sent by the hearing aid to the server will carry the hearing aid's device identification information so that the server can identify which hearing aid sent the data.

[0030] S400 extracts voiceprint features from speech data when it determines that there is potential self-speaking speech in the speech data.

[0031] Self-spoken speech refers to the sound signal emitted by the hearing aid wearer themselves, typically the speech signal captured by a microphone or other audio acquisition equipment and converted into an electrical or digital signal. Potential self-spoken speech data refers to unrecognized user self-spoken speech, that is, speech data emitted by the hearing aid wearer that has not been analyzed and identified. Voiceprint features refer to the unique acoustic characteristics of an individual's voice; these features can be used to identify and distinguish different speakers. Voiceprint formation is influenced by various factors, including the shape and size of the vocal cords, the structure of the oral and nasal cavities, and pronunciation habits.

[0032] After receiving voice data from the first and second hearing aids, the processor determines in real time whether there is potential self-speaking speech (i.e., whether self-speaking speech is possible). If potential self-speaking speech is found, the voiceprint verification process begins. Specifically, this may involve preprocessing the voice data to separate the voice signal and extract the voiceprint features.

[0033] S600 compares the voiceprint features with the stored standard voiceprint features of the user to obtain the comparison results. The standard voiceprint features are obtained by processing the user's speech sample data, which includes near speech data collected by the first hearing aid close to the user's mouth.

[0034] In this embodiment, the voice sample data includes voice data recorded by the user through the microphone array of a first hearing aid placed near their mouth, containing a specified text. Specifically, during the self-speaking voiceprint registration process, the user removes one of their hearing aids (e.g., the first hearing aid in the right ear) and holds it to their mouth. The microphone array in the removed hearing aid (i.e., the first microphone array) can be understood as a proximity microphone array. The user records a specified text using this proximity microphone array and the second hearing aid still worn in the left ear, thus obtaining the user's voice sample data. Subsequently, the voice sample data is preprocessed to separate the user's pure self-speaking voice. Next, voiceprint features are extracted from the self-speaking voice, and these extracted voiceprint features are identified as the user's standard voiceprint features, which are then stored in a database. Specifically, the standard voiceprint features can be associated with the device identification information of the hearing aid worn by the user or the user's identification information and stored in the database.

[0035] Following the previous step, after extracting the voiceprint features from the speech data, the user's standard voiceprint features can be obtained based on the hearing aid's device identifier or user identifier. The user's standard voiceprint features are then compared with the extracted voiceprint features to determine whether they match, thus confirming whether the speech data contains self-spoken voices from the user.

[0036] S800: If the comparison result indicates that the speech data contains self-speaking speech, then the speech data is enhanced.

[0037] If the comparison results indicate the presence of self-speaking voice in the voice data, the voice data is enhanced to allow the user to hear their own voice more clearly. Specifically, the voice data can be enhanced according to preset enhancement parameters and the user's personal preferences. For example, the user can adjust the volume or frequency response through the self-speaking control in the mobile application interface. If, during the fitting process, the wearer reports that the self-speaking voice sounds too low-frequency, the audiologist can use fitting tools to reduce the gain of the low frequencies below 1000Hz (Hertz) of the self-speaking voice.

[0038] The aforementioned voice data processing method first collects clearer near-speech voice data recorded by the user through a first hearing aid placed close to the user's mouth, and extracts and stores the user's standard voiceprint features from this near-speech voice data. In subsequent practical applications, it determines in real time whether there is potential self-speech voice in the voice data collected by the first and second hearing aids. If potential self-speech voice is determined to exist, a voiceprint verification step is performed to extract the voiceprint features from the voice data and compare them with the stored standard voiceprint features of the user to further accurately determine whether there is self-speech voice in the voice data. If self-speech voice is determined to exist, the voice data is enhanced to reduce the problem of ear blockage and sound distortion, thereby improving the user's auditory experience. The entire solution does not rely on specific spatial orientation information for self-talk detection. It only needs to compare the voiceprint features in the real-time collected voice data with the user's standard voiceprint features. Since the user's standard voiceprint features are extracted from near-speech data containing clear and distinct user voice characteristics, the presence of self-talk can be accurately detected by comparing the voiceprint features with the user's standard voiceprint features. This reduces misidentification caused by environmental noise or background noise, maintains a high recognition rate even in noisy environments, and improves the robustness of the system.

[0039] In practical applications, before detecting self-speaking voices, it is necessary to register the user's self-speaking voiceprint for subsequent voiceprint verification to further determine whether self-speaking voices exist. The self-speaking voiceprint registration stage of this application differs from the traditional method that relies on professional audiologists recording user voice samples in a listening room; this application provides a more applicable self-speaking voiceprint registration method. For example... Figure 3 As shown, in an exemplary embodiment, before comparing the voiceprint features with the stored standard voiceprint features of the user, the method further includes:

[0040] S102, Obtain user's voice sample data.

[0041] S104, preprocesses the speech sample data.

[0042] S106, extract audio feature data from the preprocessed speech sample data.

[0043] S108 trains a pre-built deep neural network based on audio feature data.

[0044] S110 extracts voiceprint features from the trained deep neural network and identifies the extracted voiceprint features as the user's standard voiceprint features.

[0045] In this embodiment, the pre-built deep neural network can be a TDNN (Time Delay Neural Network). TDNNs can process information with time delays and are well-suited for processing sequential data such as speech signals, where dependencies exist between data points at different times. It is understood that in other embodiments, recurrent neural networks, long short-term memory networks, and other deep learning networks may also be used.

[0046] In practice, the user can remove the first hearing aid worn in the right ear and hold it to their mouth. The microphone array in the first hearing aid near the mouth (i.e., the first microphone array) can be understood as a proximity microphone array, and the microphone in the second hearing aid worn in the left ear is the second microphone array. The user records speech data of a specified text using both the first and second microphone arrays, obtaining the user's speech sample data. After obtaining this speech sample data, the processor preprocesses the data to separate the user's pure self-spoken speech.

[0047] Specifically, preprocessing can include, but is not limited to, frame segmentation, windowing, and adaptive filtering. After obtaining clean self-spoken speech through preprocessing, MFCC features can be extracted from the clean self-spoken speech using Mel Frequency Cepstral Coefficients (MFCC) or other suitable feature extraction methods, and then normalized. Next, the normalized MFCC features are input into a pre-constructed TDNN network, which includes an input layer, several hidden layers, and an output layer. The neural network is trained using this voiceprint feature sample data to obtain a trained TDNN network. Then, the trained TDNN network can be used to perform forward propagation on the input MFCC features, extracting voiceprint features such as X-vectors from the network layers. The X-vectors represent the speaker's voiceprint features and are used to distinguish different speakers. Finally, the extracted X-vectors are determined as the user's standard voiceprint features. Furthermore, the user's standard voiceprint features can be associated with a user identifier or hearing aid device identifier and stored in a database.

[0048] In this embodiment, by extracting audio features from speech samples to train a deep neural network model, the user's voiceprint features can be extracted efficiently and accurately, and it is applicable to self-speaking voice verification in different environments.

[0049] There are various ways to preprocess speech sample data. In one exemplary embodiment, preprocessing speech sample data includes at least one of the following methods:

[0050] The first step is to perform frame-by-frame processing on the speech sample data.

[0051] The second step is to perform windowing processing on the speech sample data.

[0052] The third step is to extract the differential signal from the output signal of the microphone array in the first and second hearing aids.

[0053] The fourth step involves constructing an adaptive blocking matrix based on speech sample data, and then using this matrix to filter the input signals from the microphone arrays in the first and second hearing aids.

[0054] The fifth step involves using a preset adaptive interference canceller to remove interfering speech and noise signals from the microphone array's output signal.

[0055] In this embodiment, in order to obtain a purer self-speaking voice, at least one of the following processing methods can be used to process the voice sample data: framing, windowing, extracting the Differential Microphone Array (DMA), constructing an Adaptive Blocking Matrix (ABM), and eliminating interference and noise signals through an Adaptive Interference Canceller (AIC).

[0056] Framing involves dividing a continuous audio signal into small, overlapping frames for subsequent processing. Specifically, the continuous audio signal can be divided into small frames of a fixed length (e.g., 256 or 512 samples), with each frame typically having 50% overlap to ensure signal continuity and reduce boundary effects.

[0057] Windowing involves applying a window function to each frame to reduce spectral leakage and boundary effects. Specifically, this can be done by selecting an appropriate window function (such as the Hanning window, Hamming window, etc.) and multiplying the samples of each frame by the value of the window function.

[0058] Differential microphone arrays enhance desired signals and suppress interference signals by calculating the differential signal between microphone pairs. Specifically, this can be achieved through differential signal processing techniques to extract the differential signal from the output signals of the two microphones, thereby enhancing the directionality of the desired signal.

[0059] An adaptive blocking matrix is ​​used to generate a blocking signal that primarily contains interfering speech and noise while minimizing the presence of self-speaking components. Specifically, an adaptive blocking matrix can be constructed based on speech sample data. Subsequently, the coefficients of the adaptive blocking matrix are adjusted using an adaptive algorithm such as LMS (Least Mean Square) to filter the input signal of the microphone array and minimize leakage of the target speech signal.

[0060] An adaptive interference canceller is used to remove interfering speech and noise signals from the output signal of a microphone array using adaptive filtering techniques. Specifically, it can use a blocking signal generated by a constructed blocking matrix as a reference signal, adjust the filter coefficients using an adaptive filter (such as an LMS filter) to make the filtered reference signal match the interference components in the microphone array's output signal as closely as possible, and subtract the filtered reference signal from the microphone array's output signal to remove interfering speech and noise signals, thus obtaining an output that mainly contains the desired signal.

[0061] In this embodiment, by performing the above-mentioned data preprocessing on the voice sample data, a purer user's self-spoken voice can be obtained, providing a data foundation for extracting more accurate voiceprint features in the future.

[0062] In an exemplary embodiment, the method further includes: acquiring speech activity detection information of the first microphone array of the first hearing aid, determining the signal-to-noise ratio of the speech data collected by the second microphone array of the second hearing aid, and updating the parameters of the adaptive blocking matrix and the adaptive filter if the speech activity detection information indicates that speech activity has been detected and the signal-to-noise ratio is higher than a preset signal-to-noise ratio threshold.

[0063] Speech activity detection information is a binary signal, where 1 indicates that the current frame contains speech activity, and 0 indicates that the current frame does not contain speech activity. Signal-to-noise ratio (SNR) represents the ratio of signal strength to noise strength. A higher SNR indicates higher signal quality.

[0064] Following the previous embodiment, to better separate the speaker's voice from other interference and noise, the speech activity of the first microphone array can be detected using a VAD (Voice Activity Detection) algorithm to obtain the speech activity detection information of the first microphone array (hereinafter referred to as VAD detection result). Furthermore, the signal-to-noise ratio (SNR) of the speech signal is calculated based on the speech signal collected by the second microphone array worn by the user's ear. Specifically, the SNR can be determined based on the following method:

[0065] [\text{SNR} = 10 \log_{10} \left( \frac{P_{\text{signal}}}{P_{\text{noise}}} \right)]

[0066] Where (P_{\text{signal}}) is the signal power, (P_{\text{noise}}) is the noise power, and SNR is the signal-to-noise ratio, usually expressed in decibels (dB).

[0067] After obtaining the VAD detection result and signal-to-noise ratio (SNR), the parameters of the adaptive blocking matrix and adaptive interference canceller can be updated based on these results. Specifically, if the VAD detection result is 1 and the SNR of the second microphone array is higher than a preset SNR threshold, such as 10 dB, then the parameters of the adaptive blocking matrix and adaptive interference canceller are updated. Conversely, if the VAD detection result is 0, or the SNR of the second microphone array is lower than the preset SNR threshold of 10 dB, then the parameters of the adaptive blocking matrix and adaptive interference canceller are not updated. It is understood that the SNR threshold can be adjusted according to the actual application scenario to balance update frequency and performance.

[0068] In this embodiment, the update timing of the adaptive blocking matrix and the adaptive interference canceller is dynamically adjusted by using the VAD detection result of the microphone of the first hearing aid held by the user near the mouth and the signal-to-noise ratio of the speech data collected by the microphone array of the second hearing aid worn by the user. This enables better separation of the user's own voice from other interference and noise.

[0069] In practical applications, there are several methods for self-talk detection, including detection based on spatial cues of the self-talk sound or through self-talk scanning. Since microphones at different locations receive signals from the same sound source at different times, this time difference is called relative delay. Therefore, by measuring this relative delay, the location of the sound source can be estimated, and thus, it can be determined whether the speech originates from a specific user.

[0070] like Figure 4 As shown, in an exemplary embodiment, based on the collected voice data, determining whether the voice data contains potential self-talk includes:

[0071] S420, acquires voice data collected by the first microphone array and the second microphone array respectively.

[0072] S440, based on the voice data collected by the first microphone array and the second microphone array, determine the relative delay between two microphones in the first microphone array and the second microphone array.

[0073] S460 determines whether there is potential self-talk in the voice data based on the relative delay between two microphones in the first microphone array and the second microphone array.

[0074] The relative delay between microphones refers to the time difference between microphones at different locations receiving signals from the same sound source. This time difference is mainly due to the different path lengths sound waves travel to reach microphones at different locations. When a user speaks, the sound reaches both microphones almost simultaneously, so the relative delay is very small or close to zero. Therefore, in a multi-microphone system, by measuring the relative delay between different microphones, the direction and location of the sound source can be determined, the angle of the user's mouth relative to the microphone array can be estimated, and thus, it can be determined whether the signal received by the microphone is the user's spoken voice.

[0075] In specific implementation, the voice data collected by the first microphone array in the first hearing aid and the second microphone array in the second hearing aid can be acquired. The presence of potential self-talking sounds in the voice data can be detected by calculating the relative delay between two microphones in the first microphone array of the first hearing aid and the second microphone array of the second hearing aid. Specifically, a time difference-based estimation algorithm can be used to measure the time difference (relative delay) between different microphones receiving signals from the same sound source, and the location or direction of the sound source can be inferred from this. Alternatively, a frequency-based estimation method can be used to determine the relative delay, such as calculating the phase difference between the signals from two microphones in the first and second microphone arrays, converting the phase difference into a time difference, and thus obtaining the relative delay between the two microphones in the first and second microphone arrays. In other embodiments, the relative delay between the two microphones in the first and second microphone arrays can be directly estimated using a cross-correlation function.

[0076] In this embodiment, the relative delay between two microphones in the first microphone array and two microphones in the second microphone array of the binaural hearing aid can be determined using a cross-correlation function. Subsequently, based on the relative delay between the two microphones in the first and second microphone arrays, the angle of the user's mouth relative to the microphone array is estimated to determine whether the sound source originates from the user, and thus to determine whether there is potential self-speaking speech in the speech data.

[0077] In this embodiment, by determining whether the voice data contains the voice signal of a specific user based on the relative delay between two microphones in the first microphone array and the second microphone array, the accuracy and robustness of voice recognition can be improved in complex environments.

[0078] In an exemplary embodiment, the relative delay between two microphones in the first microphone array and the second microphone array includes a first relative delay between two microphones in the same microphone array and a second relative delay between two microphones in different microphone arrays. Figure 5 As shown, S460 includes:

[0079] S462, determine the absolute value of the delay difference between the first relative delay duration and the preset relative delay duration reference value, wherein the relative delay duration reference value is the relative delay duration between two microphones in the second microphone array, and the second hearing aid is a hearing aid in the wearing state.

[0080] S464, if the absolute value of the delay difference is less than the preset relative delay duration error and the second relative delay durations all meet the preset relative delay duration error range, then it is determined that the voice data contains potential self-talking voice.

[0081] In this embodiment, the first relative delay between two microphones in the same microphone array includes the relative delay between any two microphones in the same microphone array of the first hearing aid and the relative delay between any two microphones in the same microphone array of the second hearing aid. The second relative delay between two microphone arrays in different microphone arrays includes the relative delay between microphones in the first microphone array and microphones in the second microphone array. It can be understood that in this embodiment, the terms "first relative delay" and "second relative delay" are used to distinguish the relative delay between different microphones in the left and right ear microphone arrays; essentially, both refer to the relative delay between microphones.

[0082] Taking a microphone array containing two microphones as an example, where the microphones in the microphone array of the hearing aid worn in the left ear are micL1 and micL2, and the microphones in the microphone array of the hearing aid worn in the right ear are micR1 and micR2, the relative delay duration between micL1 and micL2 (i.e., the first relative delay duration) can be determined using the GCC method, denoted as delay L; the relative delay duration between micR1 and micR2 (i.e., the first relative delay duration) can be determined, denoted as delay R; the relative delay duration between micL1 and micR1 (i.e., the second relative delay duration) can be determined, denoted as delay1; and the relative delay duration between micL2 and micR2 (i.e., the second relative delay duration) can be determined, denoted as delay2.

[0083] In practical applications, when the detected relative delay is close to zero or within a very small range, it indicates the possible presence of user-generated speech. Conversely, a larger relative delay indicates that the sound source is located to one side or behind the microphone array, which is usually someone else's voice. Therefore, an allowable relative delay error, denoted as the relative delay error D, can be pre-set based on relevant experience and experimental results. Furthermore, a reference value for the relative delay between microphones in the second microphone array of the hearing aid worn on the ear is determined to estimate the angle of the mouth relative to the second microphone array. It is understood that if one of the binaural hearing aids includes two or more microphone arrays, the relative delay between each pair of microphones in the microphone arrays of the first and second hearing aids can be calculated. Based on the relative delay between each pair of microphones, it can be determined whether the speech data potentially contains user-generated speech.

[0084] In one example embodiment, the relative delay duration reference value can be obtained based on the following:

[0085] Determine the relative impulse response of each microphone in the second hearing aid to the user's mouth.

[0086] A reference value for the relative delay duration is obtained based on the relative delay duration between each relative impulse response.

[0087] The relative impulse response (RIR) is the impulse response of a system relative to a reference point or reference condition. It is usually used to compare impulse responses obtained at different locations, at different times, or under different conditions.

[0088] Specifically, since the microphone array captures both direct sound and ambient reflected sound from the speaker's mouth, the relative transfer function of the direct sound path can be estimated using the signal processed by AIC. Specifically, this can be achieved by using the NLMS (Normalized Least Mean Square) coefficients in the adaptive interference cancellation method to predict the relative impulse responses RIR1 and RIR2 from the user's mouth to the two microphones in the second hearing aid. Then, the maximum cross-correlation point (peak value) between RIR1 and RIR2 is calculated using the cross-correlation function. The delay corresponding to this point is the relative delay duration between RIR1 and RIR2 (i.e., the relative delay duration reference value), denoted as delayRIR. Furthermore, the angular position of the user's mouth relative to the microphone array can be estimated using triangulation techniques based on the known microphone array geometry and relative delay duration.

[0089] Furthermore, after acquiring the user's voice sample data during the user registration process, the voice sample data can be preprocessed to separate out the pure self-spoken speech. Based on the self-spoken speech, the correlation transfer functions RIR1 and RIR2 of the two microphones in the user's left-ear hearing aid to the mouth are estimated using the NLMS coefficients in the adaptive interference cancellation method. The relative delay duration of the two transfer functions (i.e., the relative delay duration reference value) is calculated and denoted as delayRIR. Then, based on the aforementioned relative delay duration error D and the relative delay duration reference value delayRIR, corresponding self-spoken speech judgment conditions are set. For example, the absolute difference between the relative delay duration and the relative delay duration reference value between two microphones in the same microphone array must be less than a preset relative delay duration error, and the relative delay duration between two microphones in different microphone arrays must be within a preset relative delay duration error range.

[0090] In the subsequent self-talk detection process, the presence of self-talk in the voice data is determined based on the self-talk judgment criteria. Specifically, this can be achieved by first determining the absolute value of the delay difference between delayL and delayRIR, and the absolute value of the delay difference between delayR and delayRIR. Then, a judgment is made based on preset self-talk judgment criteria. If the relative delay durations between microphones simultaneously meet the following self-talk judgment criteria, then self-talk is considered to exist. The self-talk judgment criteria are as follows:

[0091] 1) The delay L of micL1 and micL2, |delayL - delayRIR| <D

[0092] 2) The delays of micR1 and micR2 are: |delayR - delayRIR| <D

[0093] 3) Delay 1 for mic L1 and mic R1, -D <delay1<D

[0094] 4) Delay 2 for micL2 and micR2, -D <delay2<D

[0095] If the relative delay between any two microphones in different microphone arrays meets the above-mentioned conditions for determining self-speaking speech, then the speech data is determined to contain potential self-speaking speech, and the voiceprint verification process begins. It is understandable that if a microphone array contains more than two microphones, the relative delay between any two microphones in a single microphone array, as well as the relative delay between any two microphones in different microphone arrays, are calculated. For example, if the left and right ear microphone arrays each contain three microphones, with the left ear array containing micL1, micL2, and micL3, and the right ear array containing micR1, micR2, and micR3, the relative delay between micL1 and micL2, micL1 and micL3, and micL2 and micL3 (i.e., the first relative delay) can be determined using the GCC method, as well as the relative delay between micR1 and micR2, micR1 and micR3, and micR2 and micR3 (i.e., the first relative delay). In addition, the relative delay durations (i.e., the second relative delay durations) between micL1 and micR1, micL2 and micR2, and micL3 and micR3 were determined.

[0096] In this embodiment, by determining the relative delay duration error, the relative delay duration reference value, and the self-speaking voice judgment conditions, it is possible to quickly and accurately judge the existence of potential self-speaking voice in practical applications.

[0097] Voiceprint verification typically compares the degree of match between the voiceprint features to be verified and standard voiceprint features to determine whether they belong to the same speaker. In an exemplary embodiment, such as... Figure 6 As shown, S600 includes: S620, determining the similarity between the voiceprint features and the standard voiceprint features of the user that have been stored.

[0098] S800 includes: S820, which determines that the voice data contains self-speaking speech when the similarity is higher than a preset similarity threshold, and performs enhancement processing on the voice data.

[0099] In practical applications, a preset similarity threshold can be set based on application requirements and experience. This threshold is used to determine whether the similarity between two voiceprint features is high enough to ensure accurate recognition. In this embodiment, the similarity threshold can be 90%. It is understood that the similarity threshold can be adjusted according to actual circumstances.

[0100] In this embodiment, the similarity between the extracted X-vector and the user's X-vector can be determined by calculating the Euclidean distance. If the similarity is higher than 90%, it is determined that self-speaking speech exists in the voice data; if the similarity is lower than 90%, it is determined that self-speaking speech does not exist in the voice data. In other embodiments, the similarity between the extracted X-vector and the user's X-vector can also be determined by calculating the Mahalanobis distance or the cosine of the angle between them, thereby verifying whether self-speaking speech exists in the voice data.

[0101] In this embodiment, by determining the similarity between the extracted voiceprint features and the user's standard voiceprint features, it is possible to accurately determine whether the voice data collected by the microphone originates from the user, and thus determine whether there is self-speaking voice.

[0102] To provide a clearer explanation of the voice data processing method provided in this application, a specific embodiment is described below, which includes the following:

[0103] I. Self-speaking voiceprint registration

[0104] Step 1: Obtain user voice sample data.

[0105] Step 2: Preprocess the speech sample data.

[0106] Step 3: Extract audio feature data from the preprocessed speech sample data.

[0107] Step 4: Train the pre-built deep neural network based on the audio feature data.

[0108] Step 5: Extract voiceprint features from the trained deep neural network, identify the extracted voiceprint features as the user's standard voiceprint features, and store the user's standard voiceprint features in the database.

[0109] Specifically, a user wearing a hearing aid in their left ear can remove the first hearing aid and hold it to their mouth. The microphone array in the first hearing aid (i.e., the first microphone array) placed near the mouth can be understood as a proximity microphone array. The user records speech data of a specified text (i.e., self-speaking speech) using this proximity microphone array and the second hearing aid still worn in the left ear, thus obtaining the user's speech sample data. To obtain cleaner self-speaking speech, the received speech sample data can be framed, windowed, the differential microphone array extracted, an adaptive blocking matrix constructed, and interference and noise signals eliminated using an adaptive interference canceller. The specific processing steps are described in the above embodiments and will not be repeated here.

[0110] After obtaining clean self-spoken speech through preprocessing, MFCC features can be extracted from the clean self-spoken speech using Mel Frequency Cepstral Coefficients (MFCC) or other suitable feature extraction methods. The extracted MFCC features are then normalized. Next, the normalized MFCC features are input into a pre-constructed TDNN network, which includes an input layer, several hidden layers, and an output layer. The neural network is trained using this voiceprint feature sample data to obtain a trained TDNN network. Then, the trained TDNN network is used to perform forward propagation on the input MFCC features, extracting X-vectors from the network layers. These extracted X-vectors are identified as the user's standard voiceprint features and stored in a database.

[0111] II. Self-speaking voiceprint detection and verification

[0112] During daily use of hearing aids, the first and second hearing aids in a binaural hearing aid continuously collect speech data through their built-in microphones. This speech data includes environmental noise and human voice signals, and is then sent to the processor of the binaural hearing aid. Upon receiving the speech data, the processor determines the first relative delay between two microphones in the same microphone array and the second relative delay between two microphones in different microphone arrays. Then, based on preset self-speaking sound detection conditions, it judges the first and second relative delays to determine whether there is potential self-speaking speech in the speech data.

[0113] If the voice data is determined to contain potential self-speaking speech, the self-speaking voiceprint verification process begins. Specifically, the voice data is first preprocessed, following the same preprocessing procedure as the self-speaking voiceprint registration process. Using the same feature extraction method as for self-speaking voiceprint registration, an X-vector is extracted from the preprocessed voice data. Then, the Euclidean distance between the extracted X-vector and the user's X-vector is calculated to determine their similarity. If the similarity is higher than 90%, it is determined that self-speaking speech exists in the voice data; if the similarity is lower than 90%, it is determined that self-speaking speech does not exist in the voice data.

[0114] III. Self-Speaking Voice Optimization

[0115] If self-speaking is detected in the voice data, the voice data is enhanced by methods such as reducing low-frequency gain and adjusting volume to allow the user to hear clearer self-speaking voice.

[0116] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.

[0117] Based on the same inventive concept, this application also provides a voice data processing apparatus for implementing the voice data processing method described above. The solution provided by this apparatus is similar to the implementation described in the above method; therefore, the specific limitations in one or more voice data processing apparatus embodiments provided below can be found in the limitations of the voice data processing method described above, and will not be repeated here.

[0118] In one exemplary embodiment, such as Figure 7 As shown, a voice data processing device 700 is provided for use in a binaural hearing aid, which includes a first hearing aid and a second hearing aid. The device includes: a data acquisition module 710, a feature extraction module 720, a feature comparison module 730, and an enhancement processing module 740, wherein:

[0119] The data acquisition module 710 is used to acquire the speech data collected by the first hearing aid and the second hearing aid, respectively.

[0120] The feature extraction module 720 is used to extract the voiceprint features from the voice data when it is determined that there is potential self-speaking speech in the voice data.

[0121] The feature comparison module 730 is used to compare the voiceprint features with the stored standard voiceprint features of the user to obtain the comparison results. The standard voiceprint features are obtained by processing the user's speech sample data, which includes near speech data collected by the first hearing aid close to the user's mouth.

[0122] The enhancement processing module 740 is used to enhance the speech data if the comparison result indicates that the speech data contains self-speaking speech.

[0123] like Figure 8 As shown, in an exemplary embodiment, the device further includes a data judgment module 702, which is used to acquire voice data collected by the first microphone array and the second microphone array respectively, determine the relative delay time between two microphones in the first microphone array and the second microphone array based on the voice data collected by the first microphone array and the second microphone array, and determine whether there is potential self-talk in the voice data based on the relative delay time between two microphones in the first microphone array and the second microphone array.

[0124] In an exemplary embodiment, the relative delay between two microphones in the first microphone array and the second microphone array includes a first relative delay between two microphones in the same microphone array and a second relative delay between two microphones in different microphone arrays. Figure 7 As shown, the data judgment module 702 is further used to determine the absolute value of the delay difference between the first relative delay duration and the preset relative delay duration reference value. The relative delay duration reference value is the relative delay duration between the second microphone array in the second hearing aid worn by the user. If the absolute value of the delay difference is less than the preset relative delay duration error and the second relative delay durations all meet the preset relative delay duration error range, then it is determined that the voice data contains potential self-talking voice.

[0125] In an exemplary embodiment, the device further includes a voiceprint registration module 701, which is used to acquire the user's voice sample data, preprocess the voice sample data, extract audio feature data from the preprocessed voice sample data, train a pre-constructed deep neural network based on the audio feature data, extract voiceprint features from the trained deep neural network, and determine the extracted voiceprint features as the user's standard voiceprint features.

[0126] In an exemplary embodiment, the voiceprint registration module 701 is further configured to acquire the user's self-spoken speech data collected by the first hearing aid and the second hearing aid, and to determine the self-spoken speech data as speech sample data; wherein, the first hearing aid is a hearing aid held close to the user's mouth in a handheld state, and the second hearing aid is a hearing aid worn in a wearing state.

[0127] In an exemplary embodiment, the voiceprint registration module 701 is further configured to construct an adaptive blocking matrix based on voice sample data, filter the input signal of the microphone array through the adaptive blocking matrix, and remove interfering voice and noise signals from the output signal of the microphone array through a preset adaptive interference canceller.

[0128] In an exemplary embodiment, the voiceprint registration module 701 is further configured to acquire voice activity detection information of the first microphone array, determine the signal-to-noise ratio of the voice data collected by the second microphone array of the second hearing aid, and update the parameters of the adaptive blocking matrix and the adaptive interference canceller if the voice activity detection information indicates that voice activity has been detected and the signal-to-noise ratio is higher than a preset signal-to-noise ratio threshold.

[0129] like Figure 8 As shown, in an exemplary embodiment, the device further includes a relative delay duration determination module 704, which is used to determine the relative impulse response of each microphone in the second microphone array to the user's mouth, and obtain a relative delay duration reference value based on the relative delay duration between the relative impulse responses.

[0130] In an exemplary embodiment, the feature comparison module 730 is further configured to determine the similarity between the voiceprint feature and the stored standard voiceprint feature of the user, and if the similarity is higher than a preset similarity threshold, determine that the voice data contains self-speaking voice.

[0131] Each module in the aforementioned voice data processing device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the operations corresponding to each module.

[0132] In one exemplary embodiment, a binaural hearing aid is provided, including a first hearing aid, a second hearing aid, and a processor, wherein the processor is connected to the first and second hearing aids. The first microphone array in the first hearing aid and the second microphone array in the second hearing aid collect speech data and send the collected speech data to the processor. The processor executes the steps of any of the above-described speech data processing methods to optimize the self-spoken sound in the speech data, thereby improving the user's hearing experience.

[0133] It is understood that the devices listed above for binaural hearing aids are only those related to the present application and do not constitute a limitation on the binaural hearing aids to which the present application is applied. In addition to the components listed above, power modules, speakers, and other components may also be included.

[0134] In one exemplary embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 9 As shown, this computer device includes a processor, memory, input / output (I / O) interfaces, and a communication interface. The processor, memory, and I / O interfaces are connected via a system bus, and the communication interface is also connected to the system bus via the I / O interfaces. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides the environment for the operating system and computer programs stored in the non-volatile storage media. The database stores user's standard voiceprint characteristics and relative latency errors, among other data. The I / O interfaces are used for exchanging information between the processor and external devices. The communication interface is used for communication with external terminals via a network connection. When executed by the processor, the computer program implements a voice data processing method.

[0135] Those skilled in the art will understand that Figure 9 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0136] In one exemplary embodiment, a computer device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in any of the above-described embodiments of the voice data processing method.

[0137] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the steps in any of the above-described voice data processing method embodiments.

[0138] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in any of the speech data processing method embodiments.

[0139] It should be noted that the user information (including but not limited to user device information, user voice sample data, user personal information, etc.) and data (including but not limited to data used for analysis such as voice data, stored data such as standard voiceprint features, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data must comply with relevant regulations.

[0140] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, artificial intelligence (AI) processors, etc., and are not limited to these.

[0141] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.

[0142] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.

Claims

1. A speech data processing method applied to a binaural hearing aid, wherein the binaural hearing aid comprises a first hearing aid and a second hearing aid, characterized in that, The method includes: Acquire speech data collected by the first hearing aid and the second hearing aid respectively; If it is determined that there is potential self-speaking speech in the speech data, the voiceprint features in the speech data are extracted; The voiceprint features are compared with the stored standard voiceprint features of the user to obtain the comparison result. The standard voiceprint features are obtained by processing the user's speech sample data, which includes near speech data collected by a first hearing aid close to the user's mouth. If the comparison result indicates that the speech data contains self-speaking speech, then the speech data is enhanced.

2. The method according to claim 1, characterized in that, The first hearing aid includes a first microphone array, the second hearing aid includes a second microphone array, and the method further includes: Acquire the voice data collected by the first microphone array and the second microphone array respectively; Based on the voice data collected by the first microphone array and the second microphone array, the relative delay time between two microphones in the first microphone array and the second microphone array is determined; Based on the relative delay between two microphones in the first microphone array and the second microphone array, it is determined whether the voice data contains potential self-talk.

3. The method according to claim 2, characterized in that, The relative delay between two microphones in the first and second microphone arrays includes a first relative delay between two microphones in the same microphone array and a second relative delay between two microphones in different microphone arrays; determining whether the voice data contains potential self-talk based on the relative delay between two microphones in the first and second microphone arrays includes: The absolute value of the delay difference between the first relative delay duration and a preset relative delay duration reference value is determined, wherein the relative delay duration reference value is the relative delay duration between two microphones in the second microphone array, and the second hearing aid is a hearing aid in the wearing state; If the absolute value of the delay difference is less than a preset relative delay duration error, and the second relative delay durations all meet the preset relative delay duration error range, then it is determined that the voice data contains potential self-speaking voice.

4. The method according to claim 3, characterized in that, Before determining the absolute value of the delay difference between the first relative delay duration and a preset relative delay duration reference value, the method further includes: Determine the relative impulse response of each microphone in the second hearing aid to the user's mouth; A reference value for the relative delay duration is obtained based on the relative delay duration between each relative impulse response.

5. The method according to claim 2 or 3, characterized in that, Before comparing the voiceprint feature with the stored standard voiceprint feature of the user, the method further includes: Obtain the user's voice sample data; The speech sample data is preprocessed; Audio feature data is extracted from the preprocessed speech sample data; Based on the audio feature data, a pre-constructed deep neural network is trained; Voiceprint features are extracted from the trained deep neural network, and the extracted voiceprint features are determined as the user's standard voiceprint features.

6. The method according to claim 5, characterized in that, The acquisition of the user's voice sample data includes: Acquire the user's self-spoken speech data collected by the first hearing aid and the second hearing aid; The self-spoken voice data is identified as voice sample data; The first hearing aid is a hearing aid held close to the user's mouth in a handheld state, while the second hearing aid is a hearing aid worn in a wearing state.

7. The method according to claim 6, characterized in that, The preprocessing of the speech sample data includes: Based on the speech sample data, an adaptive blocking matrix is ​​constructed, and the input signals of the microphone arrays in the first and second hearing aids are filtered by the adaptive blocking matrix. Interfering speech and noise signals in the output signal of the microphone array are eliminated by a preset adaptive interference canceller.

8. The method according to claim 7, characterized in that, The method further includes: Obtain speech activity detection information of the first microphone array of the first hearing aid; Determine the signal-to-noise ratio of the speech data acquired by the second microphone array of the second hearing aid; If the speech activity detection information indicates that speech activity has been detected and the signal-to-noise ratio is higher than a preset signal-to-noise ratio threshold, then the parameters of the adaptive blocking matrix and the adaptive interference canceller are updated.

9. The method according to any one of claims 1 to 4, characterized in that, The step of comparing the voiceprint features with the stored standard voiceprint features of the user includes: Determine the similarity between the voiceprint feature and the standard voiceprint feature of the user that has been stored; If the similarity is higher than a preset similarity threshold, the voice data is determined to contain the user's self-spoken voice.

10. A binaural hearing aid, comprising a first hearing aid, a second hearing aid, and a processor, wherein the processor is connected to the first hearing aid and the second hearing aid respectively, characterized in that: The first microphone array in the first hearing aid and the second microphone array in the second hearing aid collect speech data and send the collected speech data to the processor, wherein the processor executes the steps of the speech data processing method according to any one of claims 1 to 9.