Intelligent fitting method and system based on fitting database
By building an intelligent fitting method based on a fitting database and utilizing multi-task learning and deep learning algorithms, the problem that smart hearing aids cannot be accurately adjusted in real time in complex environments is solved, personalized parameter adjustment of hearing aids is achieved, and the user experience is improved.
Patent Information
- Application Number
- CN202511109511.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-08
- Publication Date
- 2025-09-05
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing smart hearing aids cannot achieve real-time and accurate personalized parameter adjustment in complex environments, and cannot fully consider the user's hearing characteristics and environmental changes, resulting in a poor user experience.
By building an intelligent fitting method based on a fitting database, collecting user hearing data and environmental sound data, and using multi-task learning and deep learning algorithms to extract features, personalized hearing aid adjustment parameters are generated, and the hearing aid parameters are dynamically adjusted according to real-time feedback and environmental changes.
It enables precise adjustment of hearing aids in complex environments, improves users' hearing experience and satisfaction, and enhances the intelligence level of hearing aids.
Smart Images

Figure CN120602878A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of intelligent hearing aids, and in particular to an intelligent fitting method and system based on a fitting database. Background Art
[0002] Existing hearing aids face numerous challenges during use, particularly regarding the technical bottleneck of precisely adjusting device parameters to suit the individual needs of different users and changing environmental conditions. Traditional hearing aid adjustment methods rely on manual fitting, typically using audiometric tests to obtain the user's hearing curve and then manually adjusting the hearing aid's gain and frequency response based on this information. However, this manual fitting method is not only time-consuming and labor-intensive, but also struggles to meet the diverse needs of users in real-world environments.
[0003] With the development of intelligent technology, smart hearing aids based on digital signal processing (DSP) are becoming mainstream. Smart hearing aids use built-in processing chips to analyze ambient sound in real time and automatically adjust parameters, enabling users to achieve better hearing in different environments. For example, hearing aids can enhance speech signals in noisy environments and reduce background noise in quiet ones. However, despite advances in automatic adjustment, existing technologies still cannot achieve truly personalized adjustments, especially when fully considering the user's hearing characteristics, environmental changes, and real-time feedback.
[0004] In recent years, with the development of big data technology, fitting databases have become an important tool for improving the adjustment of smart hearing aids. By collecting a large amount of user hearing data and environmental sound data, combined with machine learning and deep learning algorithms, potential patterns can be extracted from them to help the hearing aid system perform more intelligent personalized adjustments. For example, based on the analysis of historical fitting data, smart hearing aids can more accurately predict the user's hearing needs and automatically adjust the gain, frequency response and noise reduction strategy according to the current environmental sound conditions. However, the existing intelligent adjustment methods based on fitting databases still have shortcomings, especially in how to combine multi-task learning and deep neural networks for accurate personalized parameter generation, environmental change prediction and real-time feedback optimization. A complete and systematic solution has not yet been formed.
[0005] Therefore, there is an urgent need for an intelligent fitting method based on a fitting database, which can automatically extract features from the user's hearing data and environmental sound data through deep learning algorithms, generate personalized hearing aid adjustment parameters, and dynamically optimize the performance of the hearing aid based on real-time feedback and environmental changes, thereby improving the user's hearing experience and usage satisfaction. Summary of the Invention
[0006] In view of the above-mentioned problems, the present invention is proposed.
[0007] Therefore, the technical problem solved by the present invention is: the technical problem of intelligent hearing aids adjusting parameters in real time and accurately in complex environments to adapt to changes in different noises and speech signals.
[0008] To solve the above technical problems, the present invention provides the following technical solutions: an intelligent fitting method based on a fitting database, comprising: Collect user hearing data and environmental sound data from the fitting database and perform preprocessing; Build an adjustment model and use it to automatically extract features from pre-processed user hearing data and environmental sound data to generate personalized hearing aid adjustment parameters; Dynamically adjust hearing aid parameters based on real-time user feedback and environmental changes, and achieve real-time synchronization and optimization through the APP.
[0009] As a preferred solution of the intelligent fitting method based on the fitting database described in the present invention, the preprocessing includes cleaning the collected user hearing data and environmental sound data, standardizing data of different scales, and filling in missing data.
[0010] As a preferred solution of the intelligent fitting method based on the fitting database described in the present invention, the construction of the adjustment model includes using a 1-second sliding window to segment the user hearing data and the environmental sound data, and extracting the audio features within each window.
[0011] As a preferred solution of the intelligent fitting method based on the fitting database of the present invention, wherein: the construction of the adjustment model further includes constructing a multi-task learning feature; Constructing multi-task learning features includes noise prediction tasks and speech recognition tasks.
[0012] As a preferred solution of the intelligent fitting method based on the fitting database described in the present invention, the noise prediction task extracts the frequency components of the noise, the time series of the noise intensity and the signal-to-noise ratio through spectrum analysis to predict the noise type in the current environment; The speech recognition task extracts the spectral features of speech through short-time Fourier transform and Mel-frequency cepstral coefficients to determine whether there is a speech signal in the current environment; The spectral features include the frequency, pitch, speaking speed and volume of the speech signal.
[0013] As a preferred embodiment of the intelligent fitting method based on the fitting database of the present invention, the construction of the adjustment model further comprises: using four GRU layers to construct a shared network layer including pre-processed user hearing data and environmental sound data as input; The temporal features extracted by the shared network layer are used as common feature representations for all tasks; Design independent output layers for noise prediction and speech recognition tasks; The output layer of the noise prediction task includes the output features of the shared network layer as input; a fully connected layer as the network structure to predict the noise type; the output is the noise type; Noise types include background noise, mechanical noise, human voice and no noise; The noise prediction task is formulated as follows: ; ; in, Indicates the The prediction results of each category; Softmax(·) means converting the input into a probability distribution; Indicates The weight matrix associated with each output is the model's prediction value for the noise type; represents the time series feature matrix, represents element-wise multiplication, Represents the input data Perform natural logarithm operation on each element in; represents the bias term; represents the increment function; Indicates A task-related weight matrix; H represents the product of the weight matrix and the input data; represents the scaling factor; Indicates the total number of tasks.
[0014] As a preferred embodiment of the intelligent fitting method based on the fitting database of the present invention, the construction and adjustment model further comprises: the output layer of the speech recognition task includes the output features of the shared network layer as input, a fully connected layer as the network structure, and determines whether there is a speech signal in the current environment; the output is whether there is a speech signal, and the formula is expressed as: ; in, Represents the output of the speech recognition task, and the output value is between [0,1]; exp represents the exponential function, represents the activation function, represents the regulating factor; when When , it indicates that there is a voice signal; when When , it means there is no speech signal; Represents the Sigmoid activation function; Represents the mapping weight matrix, which transforms the time series feature matrix Mapping to the output space of the speech recognition task; Indicates a nonlinear transformation of the time series feature matrix; Indicates the end point of the time interval. Represents the time series feature matrix at time characteristics.
[0015] As a preferred solution of the intelligent fitting method based on the fitting database described in the present invention, wherein: the construction of the adjustment model also includes using the cross entropy loss function as the loss function of the noise prediction task, represents the loss of the noise prediction task, and the formula is expressed as: ; in, represents the number of samples; Indicates the The true labels of samples; Indicates the The predicted probability of a sample; Function as the loss function for speech recognition tasks, Represents the loss of the speech recognition task, and the formula is expressed as: ; in, Indicates the The true labels of samples, It is The predicted probability of a sample; Indicates that the loss is weighted according to the importance or complexity of the sample; The overall loss function is expressed as the weighted sum of all task loss functions, and the formula is expressed as: ; in, represents the contribution of the loss of the noise prediction task to the total loss, represents the contribution of the loss of the control speech recognition task to the total loss, represents the loss of the noise prediction task, represents the loss of the speech recognition task; The pre-processed user hearing data and environmental sound data are divided into a 70% training set and a 30% validation set; The pre-processed user hearing data and environmental sound data are processed through a shared network layer to extract common time series features; The task-specific output layer further processes the features output by the shared network layer to generate prediction results for each task; Calculate the loss function of each task and calculate the gradient through backpropagation; The gradient is propagated back to each layer of the adjustment model, adjusting the weights of the adjustment model; Adam is used to update the parameters in the network. After each round of training, the loss function value decreases.
[0016] As a preferred embodiment of the intelligent fitting method based on the fitting database of the present invention, the dynamic adjustment of hearing aid parameters includes, when the output of the noise prediction task is background noise, +5dB, 250Hz to 1000Hz frequency band enhancement, and moderate noise reduction; When the output of the noise prediction task is mechanical noise, +3dB, suppression of the frequency band greater than 2000Hz, strong noise reduction; When the output of the noise prediction task is human voice, +10dB, 250Hz to 4kHz frequency band enhancement, mild noise reduction; When the output of the noise prediction task is noise-free, no gain adjustment is performed and noise reduction is turned off; When the output of the speech recognition task is a speech signal, +10dB, 300Hz to 3kHz frequency band enhancement, and mild noise reduction; When the output of the speech recognition task is no speech signal, no gain adjustment is performed and strong noise reduction is performed.
[0017] As a preferred solution of the intelligent fitting method based on the fitting database of the present invention, wherein: the data module collects user hearing data and environmental sound data from the fitting database and performs preprocessing; The adjustment model module builds an adjustment model and uses the adjustment model to automatically extract features from the pre-processed user hearing data and environmental sound data to generate personalized hearing aid adjustment parameters; The adjustment module dynamically adjusts hearing aid parameters based on real-time user feedback and environmental changes, and achieves real-time synchronization and optimization through the APP.
[0018] Beneficial effects of the present invention: The intelligent fitting method based on the fitting database provided by the present invention collects and preprocesses user hearing data and environmental sound data to provide high-quality input for the model; uses a deep learning model to extract time series features to achieve accurate recognition and prediction of noise and speech signals; dynamically adjusts the hearing aid parameters in accordance with different environmental changes, thereby improving the adaptability of hearing devices in complex environments; and optimizes the performance of hearing aids and enhances the user's auditory experience through the synergy of a real-time feedback mechanism and an APP. The overall intelligence level of the hearing aid is improved, enabling it to automatically adapt to various noise and speech changes, significantly improving the user's usage experience and satisfaction. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0020] Figure 1 This is an overall flow chart of an intelligent fitting method based on a fitting database provided in the first embodiment of the present invention. DETAILED DESCRIPTION
[0021] To make the above-mentioned objects, features, and advantages of the present invention more clearly understood, the following detailed description of the specific embodiments of the present invention is given in conjunction with the accompanying drawings. It is obvious that the described embodiments are only part of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary persons in this field without creative work should fall within the scope of protection of the present invention.
[0022] Example 1, reference Figure 1 , as an embodiment of the present invention, provides an intelligent fitting method based on a fitting database, comprising: S1: Collect user hearing data and environmental sound data from the fitting database and perform preprocessing.
[0023] User hearing data and environmental sound data include test results related to the user's personal hearing condition, hearing curves, hearing aid debugging parameters, hearing change records, user preferences, and types of hearing impairment; environmental sound data covers noise type, noise intensity, signal-to-noise ratio, speech signal characteristics, and judgment of whether speech exists or not.
[0024] The collected user hearing data and environmental sound data are cleaned, data of different scales are standardized, and missing data are filled.
[0025] Furthermore, by collecting and preprocessing user hearing data and environmental sound data, we can fully understand the user's personal hearing condition, environmental noise characteristics, and user preferences, and provide accurate input for the model. These data include hearing test results, hearing aid debugging parameters, environmental noise information, etc., which help to build personalized adjustment strategies. Data cleaning and standardization are the key to solving data quality problems. They can eliminate noise interference, fill missing values, and ensure that all data are compared and calculated on the same scale, thereby avoiding model errors and improving the training effect of subsequent algorithms. Overall, this step ensures the accuracy and completeness of the input data, laying a solid foundation for the realization of intelligent hearing aid adjustment.
[0026] S2: Build an adjustment model and use the adjustment model to automatically extract features from the preprocessed user hearing data and environmental sound data to generate personalized hearing aid adjustment parameters.
[0027] A 1-second sliding window is used to segment user hearing data and environmental sound data, and the audio features within each window are extracted.
[0028] Constructing multi-task learning features includes noise prediction tasks and speech recognition tasks.
[0029] The noise prediction task extracts the frequency components of the noise, the time series of the noise intensity, and the signal-to-noise ratio through spectrum analysis, and predicts the noise type in the current environment.
[0030] The speech recognition task extracts the spectral features of speech through short-time Fourier transform and Mel-frequency cepstral coefficients to determine whether there is a speech signal in the current environment.
[0031] The Short-Time Fourier Transform (STFT) divides the audio signal into multiple one-second sliding windows and calculates the signal's spectrum within each window. After obtaining the spectral features from the STFT, Mel-Frequency Cepstral Coefficients (MFCCs) are used to further extract features from the speech signal. MFCCs are a commonly used feature in speech processing that simulate the human ear's perception of sound and emphasize important frequency components in speech signals. MFCCs compress the spectral features into a set of coefficients that contain key information about the speech signal.
[0032] The extracted MFCC features typically require normalization to eliminate amplitude differences between different speech signals. The processed speech features (including MFCC features) are fed into the speech recognition network within the deep learning model. This network analyzes the input's temporal characteristics to determine whether a speech signal is present in the current environment.
[0033] The network uses the patterns learned during training to analyze the input features and make a classification judgment. If the system recognizes that the spectral characteristics of the input signal match those of a speech signal, the output result is "speech signal present." If no signal matching the speech characteristics is detected, the judgment is "no speech signal."
[0034] Ultimately, the system outputs a binary value indicating whether a speech signal is present in the current environment. This result is used as a basis for subsequent hearing aid adjustments, helping the system decide whether to strengthen speech signal processing or adjust other audio characteristics.
[0035] The spectral features include the frequency, pitch, speaking speed and volume of the speech signal.
[0036] The shared network layer constructed using 4 GRU layers includes preprocessed user hearing data and environmental sound data as input.
[0037] The temporal features extracted by the shared network layer are used as common feature representations for all tasks.
[0038] Design independent output layers for noise prediction and speech recognition tasks.
[0039] The output layer of the noise prediction task includes the output features of the shared network layer as input; a fully connected layer as the network structure to predict the noise type; and the output is the noise type.
[0040] Noise types include background noise, mechanical noise, human voice and no noise.
[0041] The noise prediction task is formulated as follows: ; ; in, Indicates the The prediction results of each category; Softmax(·) means converting the input into a probability distribution; Indicates The weight matrix associated with each output is the model's prediction value for the noise type; represents the time series feature matrix, represents element-wise multiplication, Represents the input data Perform natural logarithm operation on each element in; represents the bias term; represents the increment function; Indicates A task-related weight matrix; H represents the product of the weight matrix and the input data; represents the scaling factor; Indicates the total number of tasks.
[0042] The output layer of the speech recognition task includes the output features of the shared network layer as input and a fully connected layer as the network structure to determine whether there is a speech signal in the current environment; the output is whether there is a speech signal, which is expressed as follows: ; in, Represents the output of the speech recognition task, and the output value is between [0,1]; exp represents the exponential function, represents the activation function, represents the regulating factor; when When , it indicates that there is a voice signal; when When , it means there is no speech signal; Represents the Sigmoid activation function; Represents the mapping weight matrix, which transforms the time series feature matrix Mapping to the output space of the speech recognition task; Indicates a nonlinear transformation of the time series feature matrix; Indicates the end point of the time interval. Represents the time series feature matrix at time characteristics.
[0043] As a preferred solution of the intelligent fitting method based on the fitting database described in the present invention, wherein: the construction of the adjustment model also includes using the cross entropy loss function as the loss function of the noise prediction task, represents the loss of the noise prediction task, and the formula is expressed as: ; in, represents the number of samples; Indicates the The true labels of samples; Indicates the The predicted probability of a sample; Function as the loss function for speech recognition tasks, Represents the loss of the speech recognition task, and the formula is expressed as: ; in, Indicates the The true labels of samples, It is The predicted probability of a sample; Indicates that the loss is weighted according to the importance or complexity of the sample; The overall loss function is expressed as the weighted sum of all task loss functions, and the formula is expressed as: ; in, represents the contribution of the loss of the noise prediction task to the total loss, represents the contribution of the loss of the control speech recognition task to the total loss, represents the loss of the noise prediction task, represents the loss of the speech recognition task;
[0044] The preprocessed user hearing data and environmental sound data are divided into a ratio of 70% training set and 30% validation set.
[0045] The pre-processed user hearing data and environmental sound data are input and processed through a shared network layer to extract common time series features.
[0046] The task-specific output layer further processes the features output by the shared network layer to generate prediction results for each task.
[0047] Calculate the loss function for each task and calculate the gradient through backpropagation.
[0048] The gradients are propagated back to each layer of the network, adjusting the network weights.
[0049] Adam is used to update the parameters in the network. After each round of training, the loss function value decreases.
[0050] Furthermore, through deep learning and multi-task learning methods, dynamic adjustment and personalized optimization of hearing aid parameters can be achieved. First, sliding window and audio feature extraction technology are used to segment and process user hearing data and environmental sound data to ensure the time series and real-time nature of the data, which provides efficient input for subsequent model training. Then, through the multi-task learning framework, noise prediction tasks and speech recognition tasks are designed separately. These two tasks can independently handle the changes in different environmental sounds while sharing the common time series features extracted by the network layer, reducing model redundancy and improving learning efficiency. The noise prediction task enables the hearing aid to recognize and distinguish multiple types of noise, while the speech recognition task determines whether there is a speech signal in the environment, providing intelligent environmental adaptability. By adopting the cross-entropy loss function, the model can effectively optimize the accuracy of noise and speech recognition tasks, further improving the accuracy and real-time response capability of the hearing aid. The design of the overall loss function allows the optimization objectives of the two tasks to be balanced in the same training process, ensuring the simultaneous accuracy of noise and speech recognition. Through the Adam optimization algorithm and back-propagation mechanism, the network weights are continuously updated and optimized, ultimately achieving personalized hearing aid parameter adjustment, allowing the hearing aid to dynamically adjust according to the real-time environment and enhance the user's listening experience.
[0051] It should be noted that the four GRU layers are stacked sequentially, with the output hidden state of each GRU layer serving as the input to the next, forming a layer-by-layer network structure. The first GRU layer receives preprocessed user hearing data and environmental sound data as input, and its output hidden state is passed to the second GRU layer, and so on, until the last GRU layer. The output of each layer is processed using a nonlinear activation function (such as tanh) to ensure that the network can capture complex temporal features.
[0052] S3: Dynamically adjust hearing aid parameters based on real-time user feedback and environmental changes, and achieve real-time synchronization and optimization through the APP.
[0053] When the output of the noise prediction task is background noise, +5dB, 250Hz to 1000Hz frequency band enhancement, and moderate noise reduction.
[0054] When the output of the noise prediction task is mechanical noise, +3dB is applied to suppress frequencies greater than 2000Hz, resulting in strong noise reduction.
[0055] When the output of the noise prediction task is human voice, +10dB is used to enhance the 250Hz to 4kHz frequency band and provide mild noise reduction.
[0056] When the output of the noise prediction task is noise-free, no gain adjustment is performed and noise reduction is turned off.
[0057] When the output of the speech recognition task is a speech signal, +10dB, 300Hz to 3kHz frequency band enhancement, and mild noise reduction.
[0058] When the output of the speech recognition task is no speech signal, no gain adjustment is performed and strong noise reduction is performed.
[0059] Furthermore, through precise noise identification and speech signal detection, the hearing aid's gain and noise reduction strategies are intelligently adjusted, significantly improving the user's hearing experience in various environments. Specifically, the output of the noise prediction task distinguishes different types of noise (such as background noise, mechanical noise, human voice, or no noise) and adjusts the frequency band gain and noise reduction strength accordingly, enabling the hearing aid to adapt to complex auditory environments. For example, when background noise is strong, the gain of the low-frequency band is increased and moderate noise reduction is applied to effectively improve hearing clarity. When mechanical noise is high, strong noise reduction is applied to reduce interference. When a human voice is detected, the voice band is enhanced and background noise is reduced to optimize speech perception. In the speech recognition task, if a speech signal is detected, light noise reduction is applied and the voice band is enhanced to ensure clear hearing. Conversely, when no speech signal is present, strong noise reduction is applied to reduce background interference. This precise dynamic adjustment not only enables the hearing aid to provide optimal audio quality in various environments, but also responds to user needs in real time, improving wearer comfort and user experience.
[0060] Embodiment 2, an embodiment of the present invention, provides an intelligent fitting system based on a fitting database, including: The data module collects user hearing data and environmental sound data from the fitting database and performs preprocessing; The adjustment model module builds an adjustment model and uses the adjustment model to automatically extract features from the pre-processed user hearing data and environmental sound data to generate personalized hearing aid adjustment parameters; The adjustment module dynamically adjusts hearing aid parameters based on real-time user feedback and environmental changes, and achieves real-time synchronization and optimization through the APP.
[0061] Example 3, an embodiment of the present invention, is different from the previous two embodiments in that: If the functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, or the portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage media include various media that can store program code, such as USB flash drives, mobile hard drives, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical disks.
[0062] The logic and / or steps represented in the flowcharts or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing the logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (e.g., a computer-based system, a system including a processor, or other system that can fetch and execute instructions from an instruction execution system, apparatus, or device). For purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by, or in conjunction with, an instruction execution system, apparatus, or device.
[0063] More specific examples (a non-exhaustive list) of computer-readable media include the following: an electrical connection with one or more wires (electronic devices), a portable computer disk cartridge (magnetic devices), a random access memory (RAM), a read-only memory (ROM), an erasable and programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). In addition, the computer-readable medium may even be paper or other suitable medium on which the program is printed, since the program may be obtained electronically, for example, by optically scanning the paper or other medium, and then editing, interpreting, or processing in another suitable manner as necessary, and then storing it in a computer memory.
[0064] Example 4 is an embodiment of the present invention, which provides an intelligent fitting method and system based on a fitting database. In order to verify the beneficial effects of the present invention, scientific demonstration is carried out through simulation experiments.
[0065] In this example, hearing data and environmental sound data from 100 users were first collected from a smart hearing aid fitting database. This data includes user feedback in a variety of environments, such as background noise, mechanical noise, human voices, and a noiseless environment. Data collection lasted for six months, with each user recording at least 15 minutes of environmental data and personal feedback daily. All data included hearing loss (HL) and hearing aid adjustment requirements. The environmental sound data covered the frequency range from 250 Hz to 4 kHz.
[0066] Data preprocessing steps include cleaning invalid data (such as incorrect records caused by equipment failure or user error) and normalizing data of different scales so that all data values are between 0 and 1. In addition, missing values are filled through interpolation to ensure data integrity.
[0067] Next, a 1-second sliding window is used to segment the user's hearing data and ambient sound data. Audio features within each window are extracted, primarily using short-time Fourier transforms (STFTs) and Mel-frequency cepstral coefficients (Mel-frequency cepstral coefficients) for feature extraction in the speech recognition task. In the noise prediction task, spectral analysis is used to extract the frequency components, intensity time series, and signal-to-noise ratio of the noise, thereby determining the type of noise in the current environment.
[0068] To build the adjustment model, a shared network structure consisting of four GRU layers was used. Preprocessed data was used as input, and common time series features were extracted through the shared network layer. These time series features were used as common feature representations for all tasks. Separate output layers were then designed for the noise prediction and speech recognition tasks. The output layer of the noise prediction task used a Softmax activation function to analyze the spectral data and predict the noise type (background noise, mechanical noise, human voice, or no noise). The output layer of the speech recognition task used a Sigmoid activation function to determine whether a speech signal was present in the environment.
[0069] During the dynamic adjustment of hearing aid parameters, the system adjusts the hearing aid's gain and noise reduction intensity in real time based on the output of the noise prediction and speech recognition tasks. For example, when background noise is detected, the system boosts the gain of the 250 Hz to 1000 Hz frequency band and enables moderate noise reduction. When mechanical noise is detected, the system suppresses frequencies above 2000 Hz and enables strong noise reduction. If a speech signal is recognized, the 250 Hz to 4 kHz frequency band is boosted and light noise reduction is applied. All parameter adjustments are synchronized and optimized in real time through the app to ensure a personalized and accurate user experience.
[0070] In this example, the environmental and hearing data of 100 users were used. The following are the parameters and results of hearing aid adjustment in different environments: Background Noise Environment: When the noise prediction task output was background noise, the system enhanced the 250 Hz to 1000 Hz frequency band and enabled medium-intensity noise reduction. The user's final hearing score was 12.15 dB.
[0071] Mechanical noise environment: When the noise prediction task output was mechanical noise, the system suppressed frequencies above 2000 Hz and enabled strong noise reduction. The user's final hearing score was 9.45 dB.
[0072] Human Voice Environment: When the noise prediction task outputs a human voice, the system enhances the 250 Hz to 4 kHz frequency band and enables mild noise reduction. The user's final hearing rating is 15.85 dB.
[0073] Noiseless environment: When the noise prediction task outputs "no noise," the system does not perform gain adjustment and disables noise reduction. The user's final hearing score is 14.25 dB.
[0074] Speech Signal Environment: When the speech recognition task outputs a speech signal, the system enhances the 300 Hz to 3 kHz frequency band and enables mild noise reduction. The user's final hearing score was 16.20 dB.
[0075] It can be seen from the experimental data that the intelligent fitting method of the present invention provides flexible hearing adjustment in different environments, ensuring that users can obtain the best hearing experience in various environments.
[0076] First, in a noisy environment, the system enhances the low-frequency band (250 Hz to 1000 Hz) and moderately reduces noise, resulting in a user's hearing score of 12.15 dB, significantly better than a traditional hearing aid without intelligent adjustments. Compared to the fixed frequency enhancement of traditional hearing aids, this invention provides personalized frequency band selection and gain adjustment based on the noise type, significantly improving hearing clarity in noisy environments.
[0077] In mechanically noisy environments, this invention can identify mechanical noise and automatically perform strong noise reduction, suppressing high-frequency noise, thereby significantly reducing the impact of mechanical noise on the user's hearing. The user's final hearing score was 9.45 dB, far exceeding the performance of traditional hearing aids in such noisy environments. Compared to traditional technologies, this method can ensure a good listening experience even in complex noisy environments through precise noise prediction and intelligent noise reduction.
[0078] In a human voice environment, this invention provides users with clearer speech information through precise speech signal recognition and frequency band enhancement (250 Hz to 4 kHz). Compared with traditional hearing aids, the system adjusts to the real-time environment, providing a better speech recognition experience. The final hearing score was 15.85 dB, which is relatively ideal.
[0079] In a noise-free environment, with the system's noise reduction function disabled, the user's final hearing score was 14.25 dB, demonstrating that the user's hearing remained at a high level without external interference. Compared to traditional technologies, this method is more flexible in adapting to different environments, avoiding unnecessary gain adjustments and ensuring natural and comfortable hearing in quiet environments.
[0080] Overall, the method of the present invention, through multi-task learning and noise prediction, can accurately identify and adapt to various environmental changes, adjust hearing aid parameters according to different noise types and speech signals, and optimize the user experience. Compared with traditional technologies, the present invention provides a more personalized and intelligent solution that can better cope with the complex and changing noise environments in real life. This method not only effectively improves hearing performance but also enhances the actual user experience, showing significant innovation and technical advantages.
[0081] It should be understood that various components of the present invention may be implemented using hardware, software, firmware, or a combination thereof. In the above-described embodiments, multiple steps or methods may be implemented using software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented using hardware, as in another embodiment, any one of the following technologies known in the art or a combination thereof may be used: a discrete logic circuit having logic gate circuits for implementing logic functions on data signals, an application-specific integrated circuit having suitable combinational logic gate circuits, a programmable gate array (PGA), a field-programmable gate array (FPGA), etc.
[0082] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention, which should all be included in the scope of the claims of the present invention.
Claims
1. An intelligent fitting method based on a fitting database, characterized in that: include: Collect user hearing data and environmental sound data from the fitting database and perform preprocessing; Build an adjustment model and use it to automatically extract features from pre-processed user hearing data and environmental sound data to generate personalized hearing aid adjustment parameters; Dynamically adjust hearing aid parameters based on real-time user feedback and environmental changes, and achieve real-time synchronization and optimization through the APP; The constructing and adjusting model further includes constructing a shared network layer using four GRU layers, with pre-processed user hearing data and environmental sound data as input; The temporal features extracted by the shared network layer are used as common feature representations for all tasks; Design independent output layers for noise prediction and speech recognition tasks; The output layer of the noise prediction task includes,the output features of the shared network layer as input; 1 fully connected layer as the network structure to predict the noise type; The output is noise type; Noise types include background noise, mechanical noise, human voice and no noise; The noise prediction task is formulated as follows: ; ; in, Indicates the The prediction results of each category; Softmax(·) means converting the input into a probability distribution; Indicates The weight matrix associated with each output is the model's prediction value for the noise type; represents the time series feature matrix, represents element-wise multiplication, Represents the input data Perform natural logarithm operation on each element in; represents the bias term; represents the increment function; Indicates A task-related weight matrix; H represents the product of the weight matrix and the input data; represents the scaling factor; Indicates the total number of tasks.
2. The intelligent fitting method based on the fitting database according to claim 1, characterized in that: The preprocessing includes cleaning the collected user hearing data and environmental sound data, standardizing data of different scales, and filling in missing data.
3. The intelligent fitting method based on the fitting database according to claim 2, characterized in that: The construction of the adjustment model includes using a 1-second sliding window to segment the user hearing data and the environmental sound data, and extracting audio features within each window.
4. The intelligent fitting method based on the fitting database according to claim 3, characterized in that: The constructing and adjusting the model further includes constructing a multi-task learning feature; Constructing multi-task learning features includes noise prediction tasks and speech recognition tasks.
5. The intelligent fitting method based on the fitting database according to claim 4, characterized in that: The noise prediction task extracts the frequency components of the noise, the time series of the noise intensity and the signal-to-noise ratio through spectrum analysis, and predicts the noise type in the current environment; The speech recognition task extracts the spectral features of speech through short-time Fourier transform and Mel-frequency cepstral coefficients to determine whether there is a speech signal in the current environment; The spectral features include the frequency, pitch, speaking speed and volume of the speech signal.
6. The intelligent fitting method based on the fitting database according to claim 5, characterized in that: The construction and adjustment model also includes an output layer of the speech recognition task, which includes the output features of the shared network layer as input and a fully connected layer as a network structure to determine whether there is a speech signal in the current environment; the output is whether there is a speech signal, and the formula is expressed as: ; in, Represents the output of the speech recognition task, and the output value is between [0,1]; exp represents the exponential function, represents the activation function, represents the regulating factor; when When , it indicates that there is a voice signal; when When , it means there is no speech signal; Represents the Sigmoid activation function; Represents the mapping weight matrix, which transforms the time series feature matrix Mapping to the output space of the speech recognition task; Indicates a nonlinear transformation of the time series feature matrix; Indicates the end point of the time interval. Represents the time series feature matrix at time characteristics.
7. The intelligent fitting method based on the fitting database according to claim 6, characterized in that: The constructing and adjusting the model further includes using a cross entropy loss function as a loss function for the noise prediction task, represents the loss of the noise prediction task, and the formula is expressed as: ; in, represents the number of samples; Indicates the The true labels of samples; Indicates the The predicted probability of a sample; Function as the loss function for speech recognition tasks, Represents the loss of the speech recognition task, and the formula is expressed as: ; in, Indicates the The true labels of samples, It is The predicted probability of a sample; Indicates that the loss is weighted according to the importance or complexity of the sample; The overall loss function is expressed as the weighted sum of all task loss functions, and the formula is expressed as: ; in, represents the contribution of the loss of the noise prediction task to the total loss, represents the contribution of the loss of the control speech recognition task to the total loss, represents the loss of the noise prediction task, represents the loss of the speech recognition task; The pre-processed user hearing data and environmental sound data are divided into a 70% training set and a 30% validation set; The pre-processed user hearing data and environmental sound data are processed through a shared network layer to extract common time series features; The task-specific output layer further processes the features output by the shared network layer to generate prediction results for each task; Calculate the loss function of each task and calculate the gradient through backpropagation; The gradient is propagated back to each layer of the adjustment model, adjusting the weights of the adjustment model; Adam is used to update the parameters in the network. After each round of training, the loss function value decreases.
8. The intelligent fitting method based on the fitting database according to claim 7, characterized in that: The dynamic adjustment of hearing aid parameters includes, when the output of the noise prediction task is background noise, +5dB, 250Hz to 1000Hz frequency band enhancement, and moderate noise reduction; When the output of the noise prediction task is mechanical noise, +3dB, suppression of the frequency band greater than 2000Hz, strong noise reduction; When the output of the noise prediction task is human voice, +10dB, 250Hz to 4kHz frequency band enhancement, mild noise reduction; When the output of the noise prediction task is noise-free, no gain adjustment is performed and noise reduction is turned off; When the output of the speech recognition task is a speech signal, +10dB, 300Hz to 3kHz frequency band enhancement, and mild noise reduction; When the output of the speech recognition task is no speech signal, no gain adjustment is performed and strong noise reduction is performed.
9. An intelligent fitting system based on a fitting database using the method according to any one of claims 1 to 8, characterized in that: The data module collects user hearing data and environmental sound data from the fitting database and performs preprocessing; The adjustment model module builds an adjustment model and uses the adjustment model to automatically extract features from the pre-processed user hearing data and environmental sound data to generate personalized hearing aid adjustment parameters; The adjustment module dynamically adjusts hearing aid parameters based on real-time user feedback and environmental changes, and achieves real-time synchronization and optimization through the APP.