Hearing loss and compensation-based hearing aid and speech coaching integrated method, system
By working together with Bluetooth headsets and mobile terminals, a personalized hearing loss curve is established, providing intelligent hearing aids and oral practice modes. This solves the problems of auditory input distortion and oral language improvement for hearing-impaired users, achieving a dual improvement in auditory assistance and language learning.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SHENZHEN INNOTRIK TECH
- Filing Date
- 2026-03-02
- Publication Date
- 2026-04-28
AI Technical Summary
Existing hearing aids and language learning tools are independent and cannot be interconnected, failing to meet the needs of hearing-impaired users for auditory input distortion and oral language improvement, and lacking personalized training and auditory assistive parameter optimization mechanisms.
By working together with Bluetooth headsets and mobile terminals, a personalized hearing loss curve is established, providing intelligent hearing aid and oral practice modes. User voice data is collected for voice enhancement processing, and audio adaptation and feedback analysis are performed based on the hearing loss curve to build a closed-loop feedback mechanism of hearing and pronunciation.
It achieves a seamless integration of auditory input and oral training for hearing-impaired users, provides highly personalized training content, dynamically optimizes auditory assistive parameters, and improves communication and learning outcomes.
Smart Images

Figure CN121771613B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the fields of hearing-assisted technology, artificial intelligence and language learning, and in particular relates to an integrated method and system for hearing aids and oral practice based on hearing loss and compensation. Background Technology
[0002] The demand for hearing aids is growing. Traditional hearing aids mainly focus on amplifying sound and reducing noise to compensate for users' hearing loss. However, their functions are limited and cannot meet the needs of users, especially language learners with hearing impairments, who have an urgent need to improve their oral communication skills.
[0003] On the other hand, existing AI-powered spoken language tutoring applications or devices are typically designed for users with normal hearing, and their audio output uses standard sound effects without considering the specific hearing conditions of hearing-impaired users. This can lead to hearing-impaired users being unable to clearly perceive key phonemes (such as high-frequency consonants) in the teaching audio, making it difficult for them to effectively imitate and self-correct, thus significantly reducing their learning effectiveness.
[0004] Currently, hearing aids and language learning tools operate independently, with their functions and data not being shared. Users frequently need to switch between devices in daily life and learning scenarios, resulting in a fragmented experience. More importantly, there is a lack of a mechanism that can link a user's "auditory perception ability" with "pronunciation performance" and use this link to dynamically optimize auditory assistive parameters, thereby creating positive feedback for learning.
[0005] Therefore, there is an urgent need for an integrated solution that can deeply integrate intelligent hearing aids and personalized oral training, optimize auditory compensation through user pronunciation feedback, and ultimately improve overall communication and learning outcomes. Summary of the Invention
[0006] Purpose of the Invention: To address the shortcomings of the existing technologies, the purpose of this invention is to provide an intelligent method, system, and device that integrates hearing aids and oral language practice. It aims to solve the core problems faced by hearing-impaired users or those with hearing loss in oral language learning, such as distorted auditory input, lack of personalized training, and a disconnect between the improvement of listening and speaking abilities. Through technological integration and closed-loop feedback, it achieves a dual improvement in the effectiveness of both hearing assistance and language learning.
[0007] Technical Solution: The present invention provides an integrated method for hearing aids and oral practice based on hearing loss and compensation, which is executed collaboratively by a Bluetooth headset and an application running on a mobile terminal, and includes the following steps:
[0008] S1. Establish a personalized hearing loss curve for each user;
[0009] S2. The Bluetooth headset includes two working modes: intelligent hearing aid mode and oral practice mode;
[0010] In intelligent hearing aid mode, ambient sound signals are collected, and real-time hearing compensation and enhancement processing is performed on the ambient sound signals based on a personalized hearing loss curve to generate and play the compensated audio signal.
[0011] In the oral practice mode, a practice process is provided, including real-time practice mode, preset scenario practice mode, and vocabulary scenario practice mode.
[0012] S3. During the training process, the user's voice signal is collected, the collected voice is enhanced, and the processed user voice data is sent to the mobile terminal.
[0013] S4. On the mobile terminal, the application receives voice data and generates corresponding target language audio data for coaching based on the coaching mode selected by the user.
[0014] S5. Based on the personalized hearing loss curve, the training audio data is adapted to generate an adapted audio signal and sent to the Bluetooth headset for playback.
[0015] S6. In the oral practice mode, record and analyze the user's interaction data and pronunciation performance, and generate learning suggestions based on the analysis results;
[0016] S7. Based on interactive data and pronunciation performance, analyze the user's perception of the audio they hear, and dynamically adjust the compensation parameters in the personalized hearing loss curve accordingly.
[0017] Further, step S1 specifically involves: obtaining the user's pure-tone hearing thresholds at 7 frequency points and generating personalized hearing loss curves, where the 7 frequency points are 125Hz, 250Hz, 500Hz, 1000Hz, 2000Hz, 4000Hz and 8000Hz.
[0018] Furthermore, in step S2, the training process is specifically as follows:
[0019] Real-time coaching mode: Based on the semantic content of the received user voice data, generate corresponding coaching audio data in the target language;
[0020] Preset scenario practice mode: Based on the target scenario selected by the user from multiple preset typical spoken language scenarios, a series of dialogue audio data matching the target scenario is generated;
[0021] Vocabulary Scenario Tutoring Mode: Obtain the target vocabulary selected by the user; based on the semantics and usage of the target vocabulary, automatically generate related oral tutoring scenarios and corresponding tutoring audio data.
[0022] Furthermore, step S5 specifically involves: performing adaptation processing on the training audio data using a multi-channel loudness compensation algorithm, the specific process of which includes:
[0023] S5.1 Divide the training audio signal into K independent compensation channels in the frequency domain, with each channel corresponding to a preset frequency band. ,in ;f k,low and f k,high Indicates the lower and upper limits of the frequency band;
[0024] S5.2. Based on the personalized hearing loss curve, determine the user's hearing loss level in each frequency band. Hearing loss value ;
[0025] S5.3, for each frequency band Based on its corresponding hearing loss value The target compensation gain of the channel is calculated using a preset gain mapping function. :
[0026]
[0027] in, It is the gain mapping function. These are adjustable coefficients related to frequency band characteristics and audio content;
[0028] S5.4, All frequency bands The original audio signal components within Multiply by the corresponding target compensation gain The compensated signal components are obtained. :
[0029]
[0030] S5.5, Compensate all K channel signal components Frequency domain synthesis is performed, and the final, adapted time-domain audio signal is generated through inverse transformation.
[0031] Furthermore, step S7 specifically includes the following steps:
[0032] S7.1 The system generates the training text, synthesizes it into audio, and then plays it to the user;
[0033] S7.2 The user repeats the text aloud, and the user's audio is recorded;
[0034] S7.3 The system converts the user's audio into a text sequence and aligns the converted text sequence with standard text; let the standard text sequence be... xM This represents the Mth word in the standard text sequence, and the user text sequence is... y L Represent the Lth word in the user's text sequence; construct the alignment cost matrix. ,in Indicates will The former Each element and The former The minimum cost of aligning elements;
[0035] Construct the cumulative distance matrix:
[0036]
[0037] in The distance is Euclidean; by backtracking the optimal path, a frame-level correspondence between the two audio segments is established.
[0038] Based on the alignment results, calculate the similarity score for each word region:
[0039]
[0040] in, For words Corresponding all frame alignment pairs A set; For set The number of elements in the middle; The maximum value among all corresponding distances across all frames;
[0041] when hour, This represents the pronunciation similarity threshold, marking the word as mispronounced.
[0042] S7.4 Frequency Compensation Based on FFT Spectral Analysis: For each identified mispronounced word, the system analyzes its acoustic features and designs a compensation strategy; based on the text alignment results, it locates the temporal position of the mispronounced word in the user's audio and extracts the corresponding audio segment. Simultaneously, extract the audio segments corresponding to the words from the standard audio. ;
[0043] S7.5 Iterative training until mastery: The system plays the compensation audio, and the user repeats it again, forming a training loop. In continuous iteration, when the same erroneous word is no longer detected, the compensation strategy is considered correct.
[0044] Furthermore, step S7.4 specifically includes the following steps:
[0045] S7.4.1 Perform FFT analysis on the two audio segments respectively to obtain the spectral representation:
[0046]
[0047]
[0048] in For window functions, For FFT points, For frequency index; x ref (n) represents the reference audio played by the device, x user (n) represents the user repeating audio, and j is the imaginary number;
[0049] S7.4.2 Calculate the degree of missing data in the user's spectrum relative to the standard spectrum; define the frequency band. ,in Covering the entire audio segment, f k,low and f k,high Indicates the lower and upper limits of the frequency band, which is divided according to the Bark scale;
[0050] S7.4.3. For each frequency band, calculate the energy ratio:
[0051]
[0052] in To prevent division by zero for small constants, S user (f) represents the FFT spectrum of the user's audio, S ref (f) represents the FFT spectrum of the reference audio;
[0053] When the energy ratio in the high-frequency region When the frequency range is below the mid-low frequency range, the frequency domain correction function is as follows: ;
[0054] Among them, compensation filter The design is as follows:
[0055]
[0056] in, For threshold frequency; Maximum frequency; To compensate for the intensity coefficients, the compensated audio is transformed back to the time domain via inverse FFT and then smoothed to avoid artificial artifacts.
[0057] The present invention also discloses an integrated system for hearing aids and oral practice based on hearing loss and compensation, including Bluetooth headsets and an application installed on a mobile terminal;
[0058] The Bluetooth headset includes:
[0059] The voice acquisition unit is used to collect user voice and ambient sound.
[0060] The audio processing unit is used to perform speech enhancement processing and audio playback;
[0061] The first communication unit is used to perform wireless data interaction with the mobile terminal;
[0062] The application includes:
[0063] The second communication unit is used for data interaction with the Bluetooth headset;
[0064] The hearing management module is used to create, store, and adjust users' personalized hearing loss curves.
[0065] An audio compensation engine is used to perform real-time hearing compensation processing on the audio signal sent to the headphones based on the hearing loss curve.
[0066] The mode management module provides a user interface to select either the intelligent hearing aid mode or the oral practice mode, and within the oral practice mode, to select either the real-time practice mode, the preset scenario practice mode, or the vocabulary scenario practice mode.
[0067] The tutoring service module is used to generate tutoring content based on the selected mode in the oral tutoring mode.
[0068] The feedback analysis module is used to analyze the correlation between the user's pronunciation performance and auditory perception, and generate adjustment suggestions for the personalized hearing loss curve.
[0069] The present invention also discloses a Bluetooth headset, which is an ear-hook or ear clip type, including a housing and a voice acquisition unit, an audio processing unit and a first communication unit integrated within the housing; the voice acquisition unit includes an omnidirectional microphone for acquiring ambient sound and a directional microphone for focusing user voice; the housing is provided with physical buttons or touch areas for switching working modes.
[0070] Beneficial effects: Compared with the prior art, the present invention has the following significant advantages:
[0071] 1. Deep integration of functions: This invention is the first to deeply integrate professional-grade intelligent hearing aids and multimodal AI oral practice into a single Bluetooth headset, eliminating the hassle of switching devices and providing hearing-impaired users with a seamless life assistance and learning tool.
[0072] 2. Highly Personalized Training: This invention innovatively applies hearing loss curves to the preprocessing of spoken language teaching audio, ensuring that the teaching content heard by each user is optimized by their personal listening model, greatly improving the comprehensibility and effectiveness of the training materials.
[0073] 3. Pioneering a New Biofeedback Mechanism: This invention constructs a unique "hearing-pronunciation" biofeedback closed loop. The system not only "teaches" and "assists hearing," but also obtains feedback from the user's "speaking" to optimize the parameters of "hearing," realizing dynamic synergy and self-evolution in the hearing aid and language rehabilitation process.
[0074] 4. Scene Adaptation and Intelligent Mode Collaboration: The system can intelligently recommend or switch working modes based on the environment and user behavior, and maintain awareness of key environmental sounds during training, balancing learning focus and life safety.
[0075] 5. Multi-mode oral practice closed loop: It retains three major practice modes: real-time dialogue, scenario practice, and customized vocabulary, and can intelligently switch between them based on user data, forming a complete personalized language learning ecosystem. Attached Figure Description
[0076] Figure 1 This is a flowchart of the intelligent method provided in the embodiments of the present invention;
[0077] Figure 2 This is a schematic diagram of the overall system architecture provided in an embodiment of the present invention;
[0078] Figure 3 A schematic diagram illustrating the dual-mode (hearing aid / tutoring) collaborative operation provided in an embodiment of the present invention;
[0079] Figure 4 A flowchart for dynamically adjusting hearing loss compensation parameters provided in an embodiment of the present invention;
[0080] Figure 5 This is a schematic diagram of the functional modules of a mobile terminal application provided in an embodiment of the present invention. Detailed Implementation
[0081] The technical solution of the present invention will be further described below with reference to the accompanying drawings.
[0082] Example 1: System Overall Architecture
[0083] like Figure 2 As shown, the system of this invention adopts an "end-cloud-end" collaborative architecture. The Bluetooth headset, as the user-side terminal device, is responsible for high-fidelity sound acquisition and playback. The built-in voice acquisition unit clearly captures the user's voice at close range; the voice processing unit performs echo cancellation, automatic gain control, background noise suppression, and hearing compensation to ensure the quality of uploaded and played voice; the voice playback unit plays training audio from the app, providing hearing compensation for hearing-impaired patients; and the communication unit is responsible for data transmission and reception with the mobile app.
[0084] A dedicated application (App) running on the smartphone (mobile terminal) serves as the local processing hub, integrating core modules such as pattern management, hearing management, practice session generation, and feedback analysis. Complex AI models for speech recognition, natural language processing, and dialogue generation can be deployed on cloud servers and invoked through the App to balance device computing power and processing effectiveness. Bluetooth headsets and mobile phones communicate stably and with low latency via Bluetooth Low Energy (BLE) or Bluetooth Classic.
[0085] Example 2: Overall Flow of the Intelligent Method
[0086] like Figure 1 As shown, this demonstrates the entire process from user initialization to the formation of a learning loop. Upon first use, users are guided by the app to complete an online listening test (or import an audiogram), creating and storing a personalized hearing loss curve. During daily use, users can select modes. In intelligent hearing aid mode, the system cyclically performs environmental sound acquisition, hearing compensation, and playback. In oral practice mode, users further select sub-modes (real-time scenario, preset scenario, vocabulary scenario), and the system executes the corresponding practice process, recording all interaction data. Each time the user repeats aloud, the feedback analysis module is triggered, analyzing the practice data to determine if and how to optimize the user's hearing compensation parameters, completing one loop, as shown below. Figure 4 As shown.
[0087] Example 3: Dual-mode collaboration and audio processing
[0088] like Figure 3 The diagram illustrates the data flow in both modes. After mode determination, if the user enters hearing aid mode, the captured sound is processed by the hearing compensation module. If the user enters coaching mode, the enhanced voice is uploaded to the app or cloud to generate coaching audio. This audio must then undergo personalized adaptation processing by the same hearing compensation module before being sent back to the headphones for playback. This ensures that regardless of the mode, the sound heard by the user is optimized based on their individual hearing model, which is the core of "personalized training based on hearing compensation."
[0089] Example 4: The process of dynamically adjusting hearing loss compensation parameters
[0090] Step 1: Text generation and audio synthesis playback for coaching:
[0091] The system generates a text for practice sessions, synthesizes it into audio, and then plays it to the user.
[0092] Step 2: User repetition and audio recording:
[0093] The user repeats the text aloud, and the user's audio is recorded.
[0094] Step 3: Error identification based on text alignment:
[0095] The system converts the user's audio into a text sequence and then aligns the transcribed text with standard text. Let the standard text sequence be... The user text sequence is Construct the alignment cost matrix. ,in Indicates will The former Each element and The former The minimum cost of aligning elements.
[0096] Construct the cumulative distance matrix:
[0097]
[0098] in The distance is Euclidean. By backtracking the optimal path, a frame-level correspondence between the two audio segments is established.
[0099] Based on the alignment results, calculate the similarity score for each word region:
[0100]
[0101] in, For words Corresponding all frame alignment pairs A set; For set The number of elements in the middle (number of aligned frame pairs); : The maximum value of the distance among all possible frame pairs (normalization factor).
[0102] when When the value is typically 0.6-0.8, the word is marked as mispronounced.
[0103] Step 4: Frequency compensation based on FFT spectrum analysis:
[0104] For each identified erroneous word, the system analyzes its acoustic features, particularly its frequency distribution, to design a precise compensation strategy. Based on the text alignment results, the system locates the temporal position of the erroneous word in the user's audio and extracts the corresponding audio segment. Similarly, extract the audio segments corresponding to the words from the standard audio. .
[0105] The compensation steps are as follows:
[0106] 1) Perform FFT analysis on the two audio segments respectively to obtain the spectral representation:
[0107]
[0108]
[0109] in For window functions (such as Hamming windows). For FFT points, For frequency indexing.
[0110] 2) Calculate the degree of missing data in the user's spectrum relative to the standard spectrum. Define frequency bands. ,in It covers the entire speech frequency band. The frequency band is divided according to the Bark scale.
[0111] 3) For each frequency band, calculate the energy ratio:
[0112]
[0113] in To prevent division by zero for small constants.
[0114] When in the high-frequency region (e.g., >3kHz) The frequency response is significantly lower than the mid-low frequency range, so a compensation audio spectral characteristic is designed. The frequency domain correction function is then: .
[0115] Among them, compensation filter The design is as follows:
[0116]
[0117] in, This is the threshold frequency, typically set to 2000-3000Hz; This is the maximum frequency, typically 8000Hz; The compensation intensity coefficient is typically set to 1.5-2.0. The compensated audio is then transformed back to the time domain via inverse FFT and appropriately smoothed to avoid artifacts.
[0118] Step 5: Iterative training until mastery is achieved:
[0119] The system plays the compensated audio, and the user repeats it again, forming a training loop. The compensation strategy is considered correct when the same erroneous word is no longer detected after three consecutive iterations.
[0120] Example 5: Mobile Terminal Application Module
[0121] like Figure 5As shown, the App serves as the control and processing center, and its software modules include: a communication module responsible for interaction with the headphones and the cloud; a UI and mode management module providing the user interface; a hearing management module maintaining the hearing curve; a training service module comprising three sub-modules; and a feedback analysis module that implements closed-loop intelligence and is used to update the hearing compensation parameters in the device's voice processing unit. All modules work together to drive the entire system's operation.
[0122] Example 6: Experimental Procedure Example: Taking the Chinese word "test" as an example
[0123] 1. Experimental prerequisites and parameter settings:
[0124] Target word: "test" (standard pinyin: cè shì);
[0125] User hearing loss curve (simplified example): Assume that the user has approximately 30 dB of hearing loss in the high-frequency range (>2 kHz) and normal hearing in the mid- and low-frequency range.
[0126] Pronunciation similarity threshold: ;
[0127] Spectrum compensation intensity coefficient: ;
[0128] Frequency allocation: Divided into 3 key frequency bands according to the Bark scale (simplified explanation):
[0129] 0 - 1 kHz (low frequency);
[0130] 1 - 3 kHz (intermediate frequency);
[0131] 3 - 8 kHz (high frequency);
[0132] 2. Step Execution and Data Input:
[0133] Steps 1 and 2: Audio playback and user repetition;
[0134] The system plays a standard audio "test" after compensation based on the current hearing curve.
[0135] The user repeats the pronunciation. Due to the user's high-frequency hearing loss, their perception of the fricative / c / in "cè" (with its main energy concentrated in 3-5 kHz) is blurred, resulting in the absence or distortion of this phoneme when imitating the pronunciation.
[0136] Step 3: Error identification based on DTW alignment;
[0137] The system converts standard audio and user audio into text sequences.
[0138] Data substitution: Assume that the standard text subsequence corresponding to the character "测" is , and the text subsequence pronounced by the user is .
[0139] Calculate the alignment cost: Calculate the cumulative distance matrix through the DTW algorithm and backtrack to obtain the optimal alignment path .
[0140] Calculate the similarity: Assume that the alignment path contains 5 frame pairs, calculate the Euclidean distance of each frame pair , and obtain the maximum distance . The total distance sum of the path is .
[0141]
[0142] Judgment: , so the system marks "测" as a mispronounced word.
[0143] Step 4: Frequency compensation based on FFT spectrum analysis;
[0144] Location and extraction: According to the alignment result, extract the audio segment of "测" in the standard audio and the corresponding segment in the user audio .
[0145] FFT analysis and energy ratio calculation:
[0146] Perform FFT on the two audio segments to obtain the spectra and .
[0147] Calculate the energy ratios of three frequency bands which are 0.97, 0.83 and 0.3 respectively.
[0148] Analysis result: The energy ratio of the user in the high-frequency band (0.30) is significantly lower than that in the middle and low-frequency bands, which is consistent with the user's hearing loss curve (30 dB high-frequency loss), confirming that the reason for "mispronunciation" is "inability to hear clearly".
[0149] Design and apply a compensation filter:
[0150] Based on the above analysis, the system designs a compensation filter for high-frequency loss . Substitute into the formula:
[0151]
[0152] Substitute the data ( , the intermediate-frequency reference energy ratio is taken as 0.83, ), in Local calculation:
[0153]
[0154] The system generates compensated audio , that is, selective enhancement of about 38% is performed in the high-frequency region where the user has serious losses (especially the / c / phoneme frequency band of 3 - 5 kHz).
[0155] Step 5: Iterative training until mastered;
[0156] The system plays the "test" audio processed by the new compensation strategy to the user.
[0157] The user repeats the reading. Due to the high-frequency phonemes being clearer and more audible now, the accuracy of their imitated pronunciation improves.
[0158] Experimental result: After three iterations of training, the pronunciation similarity of the user for the word "test" has increased from 0.28 to 0.82, exceeding the threshold of 0.7. The system determines that this compensation strategy (boosting by about 38% in the high-frequency band for the individual user) is effective for teaching this word, and integrates this fine-tuning experience into the personalized hearing loss compensation model of this user.
[0159] It should be noted that in the embodiments of the present invention, functions such as dialogue generation, scenario construction, and pronunciation evaluation can partially or entirely rely on the local models running on the mobile terminal, or some functions can be implemented by the App calling the cloud API service, which does not deviate from the core idea of the system architecture of the present invention.
[0160] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method integrating hearing aids and oral practice based on hearing loss and compensation, characterized in that, This is performed collaboratively by an application running on a Bluetooth headset and a mobile device, and includes the following steps: S1. Establish a personalized hearing loss curve for each user; S2. The Bluetooth headset includes two working modes: intelligent hearing aid mode and oral practice mode; In intelligent hearing aid mode, ambient sound signals are collected, and real-time hearing compensation and enhancement processing is performed on the ambient sound signals based on a personalized hearing loss curve to generate and play the compensated audio signal. In the oral practice mode, a practice process is provided, including real-time practice mode, preset scenario practice mode, and vocabulary scenario practice mode. S3. During the training process, the user's voice signal is collected, the collected voice is enhanced, and the processed user voice data is sent to the mobile terminal. S4. On the mobile terminal, the application receives voice data and generates corresponding target language audio data for coaching based on the coaching mode selected by the user. S5. Based on the personalized hearing loss curve, the training audio data is adapted to generate an adapted audio signal and sent to the Bluetooth headset for playback. S6. In the oral practice mode, record and analyze the user's interaction data and pronunciation performance, and generate learning suggestions based on the analysis results; S7. Based on interactive data and pronunciation performance, analyze the user's perception of the audio they hear, and dynamically adjust the compensation parameters in the personalized hearing loss curve accordingly.
2. The integrated method for hearing aids and oral practice based on hearing loss and compensation as described in claim 1, characterized in that, Step S1 specifically involves obtaining the user's pure-tone hearing thresholds at 7 frequency points and generating a personalized hearing loss curve. The 7 frequency points are 125Hz, 250Hz, 500Hz, 1000Hz, 2000Hz, 4000Hz and 8000Hz.
3. The integrated method for hearing aids and oral practice based on hearing loss and compensation according to claim 1, characterized in that, In step S2, the training process is specifically as follows: Real-time coaching mode: Based on the semantic content of the received user voice data, generate corresponding coaching audio data in the target language; Preset scenario practice mode: Based on the target scenario selected by the user from multiple preset typical spoken language scenarios, a series of dialogue audio data matching the target scenario is generated; Vocabulary Practice Mode: Acquires target vocabulary words selected by the user; Based on the semantics and usage of the target vocabulary, related oral practice scenarios and corresponding audio data are automatically generated.
4. The integrated method for hearing aids and oral practice based on hearing loss and compensation as described in claim 1, characterized in that, Step S5 specifically involves: using a multi-channel loudness compensation algorithm to adapt the training audio data, the specific process of which includes: S5.1 Divide the training audio signal into K independent compensation channels in the frequency domain, with each channel corresponding to a preset frequency band. ,in ;f k,low and f k,high Indicates the lower and upper limits of the frequency band; S5.
2. Based on the personalized hearing loss curve, determine the user's hearing loss level in each frequency band. Hearing loss value ; S5.3, for each frequency band Based on its corresponding hearing loss value The target compensation gain of the channel is calculated using a preset gain mapping function. : ; in, It is the gain mapping function. These are adjustable coefficients related to frequency band characteristics and audio content; S5.4, All frequency bands The original audio signal components within Multiply by the corresponding target compensation gain The compensated signal components are obtained. : ; S5.5, Compensate all K channel signal components Frequency domain synthesis is performed, and the final, adapted time-domain audio signal is generated through inverse transformation.
5. The integrated method for hearing aids and oral practice based on hearing loss and compensation according to claim 1, characterized in that, Step S7 specifically includes the following steps: S7.1 The system generates the training text, synthesizes it into audio, and then plays it to the user; S7.2 The user repeats the text aloud, and the user's audio is recorded; S7.3 The system converts the user's audio into a text sequence and aligns the converted text sequence with standard text; let the standard text sequence be... x M This represents the Mth word in the standard text sequence, and the user text sequence is... y L Represent the Lth word in the user's text sequence; construct the alignment cost matrix. ,in Indicates will The former Each element and The former The minimum cost of aligning elements; Construct the cumulative distance matrix: ; in The distance is Euclidean; by backtracking the optimal path, a frame-level correspondence between the two audio segments is established. Based on the alignment results, calculate the similarity score for each word region: ; in, For words Corresponding all frame alignment pairs A set; For set The number of elements in the middle; The maximum value among all corresponding distances across all frames; when hour, This represents the pronunciation similarity threshold, marking the word as mispronounced. S7.4 Frequency Compensation Based on FFT Spectral Analysis: For each identified mispronounced word, the system analyzes its acoustic features and designs a compensation strategy; based on the text alignment results, it locates the temporal position of the mispronounced word in the user's audio and extracts the corresponding audio segment. Simultaneously, extract the audio segments corresponding to the words from the standard audio. ; S7.5 Iterative training until mastery: The system plays the compensation audio, and the user repeats it again, forming a training loop. In continuous iteration, when the same erroneous word is no longer detected, the compensation strategy is considered correct.
6. The integrated method for hearing aids and oral practice based on hearing loss and compensation according to claim 5, characterized in that, Step S7.4 specifically includes the following steps: S7.4.1 Perform FFT analysis on the two audio segments respectively to obtain the spectral representation: ; ; in For window functions, For FFT points, For frequency index; x ref (n) represents the reference audio played by the device, x user (n) represents the user repeating audio, and j is the imaginary number; S7.4.2 Calculate the degree of missing data in the user's spectrum relative to the standard spectrum; define the frequency band. ,in Covering the entire audio segment, f k,low and f k,high Indicates the lower and upper limits of the frequency band, which is divided according to the Bark scale; S7.4.
3. For each frequency band, calculate the energy ratio: ; in To prevent division by zero for small constants, S user (f) represents the FFT spectrum of the user's audio, S ref (f) represents the FFT spectrum of the reference audio; When the energy ratio in the high-frequency region When the frequency range is below the mid-low frequency range, the frequency domain correction function is as follows: ; Among them, compensation filter The design is as follows: ; in, For threshold frequency; Maximum frequency; To compensate for the intensity coefficients, the compensated audio is transformed back to the time domain via inverse FFT and then smoothed to avoid artificial artifacts.
7. An integrated system for hearing aids and oral practice based on hearing loss and compensation, used to implement the method as described in claim 1, characterized in that, This includes Bluetooth headsets and applications installed on mobile devices; The Bluetooth headset includes: The voice acquisition unit is used to collect user voice and ambient sound. The audio processing unit is used to perform speech enhancement processing and audio playback; The first communication unit is used to perform wireless data interaction with the mobile terminal; The application includes: The second communication unit is used for data interaction with the Bluetooth headset; The hearing management module is used to create, store, and adjust users' personalized hearing loss curves. An audio compensation engine is used to perform real-time hearing compensation processing on the audio signal sent to the headphones based on the hearing loss curve. The mode management module provides a user interface to select either the intelligent hearing aid mode or the oral practice mode, and within the oral practice mode, to select either the real-time practice mode, the preset scenario practice mode, or the vocabulary scenario practice mode. The tutoring service module is used to generate tutoring content based on the selected mode in the oral tutoring mode. The feedback analysis module is used to analyze the correlation between the user's pronunciation performance and auditory perception, and generate adjustment suggestions for the personalized hearing loss curve.
8. A Bluetooth headset, integrating the system of claim 7, characterized in that, The Bluetooth headset is an ear-hook or ear clip type, including a shell and a voice acquisition unit, an audio processing unit and a first communication unit integrated within the shell; the voice acquisition unit includes an omnidirectional microphone for acquiring ambient sound and a directional microphone for focusing user voice; the shell is provided with physical buttons or touch areas for switching working modes.
Citation Information
Patent Citations
Speech enhancement hearing aid method
CN109147808A
Hearing and monitoring system
US20200268260A1