Hearing aid speech quality self-evaluation method based on hearing loss classification

By constructing a hearing aid speech quality self-evaluation network based on hearing loss classification, and utilizing a recurrent neural network with convolutional neural networks and attention mechanisms, the problem of low accuracy in no-reference speech quality evaluation is solved, thus simplifying and improving the accuracy of hearing aid speech quality self-evaluation.

CN116453547BActive Publication Date: 2026-01-27NANJING INST OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202210620231.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-02
Publication Date
2026-01-27
Estimated Expiration
2042-06-02

AI Technical Summary

Technical Problem

Existing methods for objectively evaluating speech quality in hearing aids without reference are not very accurate and are difficult to apply effectively in the hearing aid fitting process.

Method used

A multi-task training approach is adopted, combining convolutional neural networks, recurrent neural networks with attention mechanisms, and softmax classification networks. By adjusting the importance of tasks through weight factors, a self-evaluation network for hearing aid speech quality is constructed, including frame-level feature extraction, hearing loss classification, and quality prediction subnetworks, simplifying the processing.

Benefits of technology

It improves the accuracy of no-reference speech quality assessment, simplifies processing steps, enriches application scenarios, and enhances the accuracy of hearing aid speech quality self-evaluation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116453547B_ABST
    Figure CN116453547B_ABST
Patent Text Reader

Abstract

The application discloses a hearing aid speech quality self-evaluation method based on hearing loss classification, comprising the following steps: constructing a speech quality self-evaluation network composed of a frame-level feature extraction network, a hearing loss classification subnetwork and a quality prediction subnetwork; calculating shallow features based on a hearing aid processed signal, learning deep representation of a distorted signal by using the frame-level feature extraction network, so as to obtain frame-level features; and obtaining the classification of hearing loss degree before distortion speech compensation and the predicted value of quality score by the hearing loss classification subnetwork and the quality prediction subnetwork respectively after the shape resetting of the frame-level features. According to the multi-task training strategy, the application takes the quality score of the predicted distorted signal as a main task, takes the quality classification of the predicted distorted signal as an auxiliary task, adjusts the importance of the main task and the auxiliary task in the network by the weight factor of the loss function during the training, and improves the accuracy of the no-reference hearing aid speech quality evaluation method and simplifies the processing process.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of hearing aid speech quality assessment technology, and in particular to a self-assessment method for hearing aid speech quality based on hearing loss classification. Background Technology

[0002] Traditional hearing aids primarily compensate for the missing sound energy and frequency components by amplifying the sound signal. They rely on the experience and skills of the audiologist to adjust the algorithm parameters to achieve optimal performance. However, this method of fitting hearing aids is inefficient and difficult to pass on effectively, presenting significant limitations. Fitting-free hearing aids represent a major future trend. These devices allow for an initial fitting based on the patient's hearing loss, followed by updating the algorithm parameters through a speech quality self-evaluation method until the speech quality meets the standard or the patient is satisfied.

[0003] Speech quality assessment methods can be categorized into subjective and objective methods based on the evaluator. Subjective assessment, under certain conditions, involves a human subject classifying distorted speech according to standard pronunciation. Common subjective assessment methods include Mean Opinion Score (MOS), Diagnostic Rhyme Test (DRT), and Diagnostic Approval Metric (DAM). Considering that humans are the ultimate recipients of speech quality assessment, subjective assessment is the most direct and accurate method, often referred to as the "gold standard." However, subjective assessment requires strict control of the testing environment and the hiring of evaluators, placing significant demands on time, money, and manpower, making it difficult to conduct in daily life.

[0004] Objective evaluation methods simulate the human auditory process using computers to provide a quality rating that is highly correlated with subjective ratings. Depending on whether an original reference signal is required, objective speech quality evaluation models can be divided into "full-reference" and "no-reference" models. Full-reference speech quality evaluation algorithms require both the original clean signal as a reference and the distorted signal to be evaluated. The Perceptual Objective Speech Quality Assessment (POLOA) standardized by the International Telecommunication Union (ITU) is a widely used full-reference speech quality algorithm in the telecommunications field. In the field of hearing aid speech quality, the Hearing Aid Speech Quality Index (HASQI) and Perceptual Model-Hearing Impairment Speech Quality (PEMO-Q-HI) are two typical full-reference speech quality evaluation models established by considering the cochlear damage of hearing-impaired patients. Although full-reference speech quality evaluation algorithms have a high correlation with subjective ratings, the clean reference signal is often difficult to obtain, greatly limiting its application. No-reference speech quality evaluation algorithms do not require an original signal as a reference; they directly extract feature parameters from the distorted signal and map them into a quality score using prior knowledge or a trained model. In the telecommunications field, there has been considerable research on no-reference speech quality assessment models, such as ITU standard P.563; Low Complexity Quality Assessment (LCQA); and with the development of deep learning, several deep learning-based no-reference speech quality assessment methods have been proposed in recent years, such as QualityNet, NISQA, and MOSNet. In the field of hearing aid quality, existing research is primarily an extension of no-reference models from the telecommunications domain, such as LCQA-HA and SRMR-HA. No-reference speech quality indices specifically proposed for hearing aids include PLP-HL and FBE-HL. No-reference quality assessment methods offer greater flexibility, but due to the lack of references, their accuracy is relatively low and requires further improvement. Summary of the Invention

[0005] To address the shortcomings of existing methods for objectively evaluating the speech quality of hearing aids without reference, which suffer from low accuracy, this invention discloses a self-evaluation method for hearing aid speech quality based on hearing loss classification. This method employs a multi-task training approach, with quality prediction as the primary task and hearing loss classification as the secondary task. Weighting factors are used to adjust the importance of the primary and secondary tasks within the network. This fully leverages the feature extraction capabilities of convolutional neural networks, combined with the temporal modeling capabilities of recurrent neural networks using attention mechanisms, and the classification capabilities of the Softmax function. By utilizing the advantages of different network models, this method improves the accuracy of objective evaluation methods for speech quality without reference and simplifies the self-evaluation process for hearing aid speech quality.

[0006] To address the aforementioned technical problems, this invention provides a self-evaluation method for hearing aid speech quality based on hearing loss classification, comprising the following steps:

[0007] S1: Construct a hearing aid speech quality self-evaluation network that includes a frame-level feature extraction network, a hearing loss classification sub-network, and a quality prediction sub-network;

[0008] S2: Input the shallow features of the speech to be tested into the frame-level feature extraction network to obtain frame-level features;

[0009] S3: Input the obtained frame-level features into the hearing loss classification sub-network to obtain the classification of the degree of hearing loss before distorted speech compensation;

[0010] S4: Input the obtained frame-level features into the quality prediction sub-network to obtain the predicted quality score;

[0011] S5: The hearing aid speech quality self-evaluation network is trained using training data labeled with hearing aid speech quality indicators. The loss function is a weighted combination of the loss functions of the quality prediction subnetwork and the hearing loss classification subnetwork.

[0012] Preferably, in S2, the shallow features of the speech to be tested are input into the frame-level feature extraction network to obtain frame-level features. The specific process is as follows:

[0013] Shallow features are calculated based on the processed signal from the hearing aid. A frame-level feature extraction network is used to learn the deep representation of the distorted signal, thereby obtaining frame-level features. The frame-level feature extraction network is composed of a convolutional neural network.

[0014] Preferably, in the S2 process, the shallow features calculated based on the processed signal from the hearing aid are as follows:

[0015]

[0016] This feature represents the average filter bank energy in each channel of the Gammatone filter for each frame, where S represents the short-time logarithmic amplitude spectrum of the distorted signal on the auditory frequency scale after framing and windowing, and c, t, and n are the number of channels C, the number of frames T, and the frame length N of the Gammatone filter bank, respectively. The shape of the final shallow feature of the distorted signal is T×32.

[0017] Preferably, the convolutional neural network used for frame-level feature extraction consists of four stacked convolutional networks, each of which contains a two-dimensional convolutional layer, a batch normalization layer, and a PReLU activation function layer. The number of output features of each two-dimensional convolutional layer is 8, 8, 16, and 16, respectively; the kernel size is [5, 5], [5, 5], [3, 5], and [3, 5], respectively; the stride is [1, 1], [1, 2], [1, 2], and [1, 2], respectively; and the padding width is [2, 2], [2, 2], [1, 2], and [1, 2], respectively. The frame-level feature shape extracted by the convolutional neural network from the shallow features of the distorted signal is represented as 16×T×4, where T is the number of frames T.

[0018] Preferably, in S3, the obtained frame-level features are input into the hearing loss classification sub-network to obtain the classification of the degree of hearing loss before distorted speech compensation. The process is as follows:

[0019] The frame-level features after shape resetting are first globally averaged to obtain segment-level features. Then, through a set of batch-normalized fully connected layers and Softmax layers, the classification of hearing loss degree before distortion speech compensation is obtained. The hearing loss classification subnetwork consists of a set of batch-normalized fully connected layers and Softmax layers.

[0020] Preferably, in S3, the obtained frame-level features are input into the hearing loss classification sub-network to obtain the classification of the degree of hearing loss before distorted speech compensation. The specific process is as follows:

[0021] S31: First, reshape the extracted frame-level features to generate T×64 features, where T is the number of frames. Then, perform global averaging on the reshaped frame-level features to achieve segment-level features.

[0022] S32: The segment-level features are input into two fully connected layers of the hearing loss classification sub-network, which consists of two fully connected layers and a Softmax layer stacked together. The output of each fully connected layer passes through a batch normalization layer. The output of the first batch normalization layer is activated by the ReLU function and then fed into the second fully connected layer. The output of the second batch normalization layer is then used as the final output of the two fully connected layers and fed into the Softmax layer. The final output of the segment-level features after passing through the two fully connected layers is of length N. l The vector, N l The total number of categories representing the degree of hearing loss;

[0023] S33: Feed the final outputs of the two fully connected layers into the Softmax layer: The outputs of the two fully connected layers are of length N. l The vector, after passing through the Softmax layer, provides a classification of the degree of hearing loss before distorted speech compensation, specifically represented as follows:

[0024]

[0025] In the formula This represents the speech features fed into the Softmax layer, where the subscript i indicates the level of hearing loss classification, i∈{1,2,...,N}. l}, o i (z) represents the predicted probability of the Softmax layer for each hearing loss level classification based on the input speech features.

[0026] Preferably, in S4, the obtained frame-level features are simultaneously input into the quality prediction sub-network to obtain the predicted quality score. The process is as follows:

[0027] The frame-level features after shape resetting are first passed through a recurrent neural network with a self-attention mechanism to obtain segment-level features. These segment-level features are then mapped through a fully connected layer to obtain the predicted value of the quality score. The quality prediction sub-network consists of a recurrent neural network, an attention mechanism layer, and an activated fully connected layer.

[0028] Preferably, in S4, the obtained frame-level features are simultaneously input into the quality prediction sub-network to obtain the predicted quality score. The specific process is as follows:

[0029] S41. Reshape the frame-level features to generate T×64 frame-level features, where T is the number of frames;

[0030] S42. The frame-level features after shape resetting are fed into a recurrent neural network with an attention mechanism to obtain segment-level features: BiLSTM is used to learn the bidirectional dependency of time series data, and an attention mechanism is used to compensate for the possible information loss in the hidden state output by BiLSTM.

[0031] S43. The segment-level features output by the recurrent neural network with attention mechanism are fed into a fully connected layer to obtain the predicted quality score. After passing through the fully connected layer, the segment-level features are activated by the sigmoid function to give the final predicted quality score.

[0032] Preferably, the input feature dimension and hidden state dimension of the BiLSTM are both set to 64, then the output feature size is T×64. The attention mechanism calculates the hidden state output after weighting by the attention weight vector at each time step of the bidirectional output of the BiLSTM. At the t-th time step of forward propagation, the output of the attention mechanism is:

[0033]

[0034] In the formula, m represents the current time step. Let f be the hidden output column vector at time step i, where the superscript f indicates forward computation. Let represent the attention weight vector from time step i to the current time step m, calculated as follows:

[0035]

[0036] In the formula W a M is the learnable weight matrix, M is the number of time steps, and the superscript T indicates matrix transpose. Finally, the segment-level features output by the recurrent neural network combined with the attention mechanism are a vector of length 128 points.

[0037] Preferably, the specific process of training using training data labeled with hearing aid speech quality indicators in S5 is as follows:

[0038] During training, the loss function of the speech quality self-evaluation network is calculated until the loss function value is less than a threshold, thus completing the training. The loss function of the speech quality self-evaluation network is a weighted combination of the loss functions of the quality prediction sub-network and the hearing loss classification sub-network. The weight factor β adjusts the importance of the hearing loss classification sub-network in the speech quality self-evaluation network.

[0039] Loss=(1-β)Loss score +βLoss level

[0040] Among them, Loss score The loss function of the quality prediction subnetwork is expressed as follows:

[0041]

[0042] Loss level The loss function for the hearing loss classification subnetwork is expressed as follows:

[0043]

[0044] Among them, o i N represents the probability predicted by the Softmax classifier for the i-th degree of hearing loss. l The total number of categories representing the degree of hearing loss, level i The true probability of the i-th degree of hearing loss is derived from the hearing loss degree classification level:

[0045]

[0046] In the formula, level represents the classification level of hearing loss, which is composed of the average hearing threshold loss avg at 500Hz, 1kHz, 2kHz, and 4kHz. HL The calculation yields the following result:

[0047]

[0048] In the formula, ceil represents rounding down, and level is 1 to N. l Integers between [a certain range].

[0049] Compared with the prior art, the beneficial effects of the present invention are:

[0050] 1. This invention is a hearing aid speech quality self-evaluation method based on hearing loss classification. Compared with the full-reference speech quality evaluation method, it does not require the collection of a clean speech signal as a reference, which simplifies the processing steps and enriches the application scenarios.

[0051] 2. This invention organically combines convolutional neural networks, recurrent neural networks with attention mechanisms, fully connected layers, and Softmax classification networks into a whole, making full use of the feature mining capabilities of convolutional neural networks, the temporal modeling capabilities of recurrent neural networks with attention mechanisms, and the classification capabilities of Softmax networks, thereby improving the evaluation accuracy of no-reference quality assessment networks.

[0052] 3. This invention incorporates an attention mechanism into the traditional LSTM model, enabling recurrent units to filter out rich and useful information from the hidden outputs;

[0053] 4. This invention employs a multi-task training strategy, with quality prediction as the primary task and hearing loss classification as the secondary task. The importance of the primary and secondary tasks in the quality assessment network is adjusted by weighting factors. The selected secondary task has a high correlation with the primary task in hearing aid quality assessment, thereby improving the accuracy of hearing aid speech quality self-evaluation. Attached Figure Description

[0054] Figure 1 This is a flowchart of the hearing aid speech quality self-evaluation method based on hearing loss classification provided by the present invention;

[0055] Figure 2 This is a flowchart of the training set training for the hearing aid speech quality self-evaluation method based on hearing loss classification provided by the present invention;

[0056] Figure 3 This is a flowchart of the frame-level feature extraction network used in the embodiments of the present invention;

[0057] Figure 4 This is a flowchart of the hearing loss classification subnetwork used in the embodiments of the present invention;

[0058] Figure 5 This is a waveform diagram of speech processed by noise addition, enhancement, and hearing loss conditions in an embodiment of the present invention;

[0059] Table 1 is an evaluation table of the prediction results of the network trained on the training set and the comparison network on the test set in the embodiments of the invention. Detailed Implementation

[0060] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. The advantages and features of the present invention will become clearer from the following description and claims. It should be noted that the drawings are all in a very simplified form and use non-precise proportions, and are only used to facilitate and clarify the illustration of the embodiments of the present invention.

[0061] Example

[0062] This invention provides a method for self-evaluation of hearing aid speech quality based on hearing loss classification. Please refer to [link / reference]. Figure 1 and Figure 2 It includes the following steps:

[0063] S1: Construct a hearing aid speech quality self-evaluation network including a frame-level feature extraction network, a hearing loss classification sub-network, and a quality prediction sub-network; wherein, the frame-level feature extraction network is composed of four stacked two-dimensional convolutional networks; the hearing loss classification sub-network consists of a set of batch-normalized fully connected layers and a Softmax layer; the quality prediction sub-network consists of a recurrent neural network, an attention mechanism layer, and an activated fully connected layer.

[0064] S2: Input the shallow features of the speech to be tested into the frame-level feature extraction network, calculate the shallow features based on the signal after hearing aid processing, and use the frame-level feature extraction network to learn the deep representation of the distorted signal, thereby obtaining the frame-level features;

[0065] First, the audio to be tested is framed. In this embodiment, the sampling rate is set to 16kHz, the frame length is 320 points, and the frame shift is half the frame length. Assuming the number of points in the distorted signal is L, then the number of frames... in This indicates rounding down to the nearest integer, ultimately resulting in a speech data matrix of the form T×320;

[0066] After the speech to be tested is framed, its shallow features are calculated. The shallow features calculated based on the signal after hearing aid processing are the FBEs features of the signal, specifically expressed as follows:

[0067]

[0068] In the formula, S represents the short-time logarithmic amplitude spectrum of the distorted signal on the auditory frequency scale after framing and windowing, and c, t, and n are the number of channels C, the number of frames T, and the frame length N of the Gammantone filter bank, respectively. Here, t represents the label of the speech frame, and T represents the total number of frames after framing the speech. The relationship between the two is t = 0, 1, 2, ... T. Since the Gammantone filter bank used in this embodiment is 32-channel, the shape of the shallow features of the speech to be tested is T×32.

[0069] The structure of the frame-level feature extraction network is shown in the attached figure. Figure 3 As shown, this is a convolutional neural network composed of four convolutional networks. Each convolutional network contains a two-dimensional convolutional layer, a batch normalization layer, and a PReLU activation function layer. The number of output features of each two-dimensional convolutional layer is 8, 8, 16, and 16, respectively. The kernel sizes are [5, 5], [5, 5], [3, 5], and [3, 5], respectively. The stride is [1, 1], [1, 2], [1, 2], and [1, 2], respectively. The padding width is [2, 2], [2, 2], [1, 2], and [1, 2], respectively. After processing by the convolutional neural network, the shallow features of the distorted signal are extracted to form frame-level features of the form 16×T×4.

[0070] S3: Input the obtained frame-level features into the hearing loss classification sub-network, first perform global averaging on the frame-level features after shape reset to obtain segment-level features, and then pass through a set of batch-standardized fully connected layers and Softmax layers to obtain the classification of the degree of hearing loss before distortion speech compensation;

[0071] Specifically, it includes the following steps:

[0072] S31: First, reshape the frame-level features extracted by the convolutional neural network to generate features of the form T×64, where T is the number of frames T. Then, perform global averaging on the reshaped frame-level features to obtain segment-level features of the form 1×64.

[0073] S32: Input segment-level features into two fully connected layers of the hearing loss classification subnetwork: The structure of the hearing loss subnetwork is shown in the appendix. Figure 4 As shown, the structure consists of two fully connected layers and a softmax layer stacked together. The output of each fully connected layer passes through a batch normalization layer. The output of the first batch normalization layer is activated by the ReLU function and then fed into the second fully connected layer. The output of the second batch normalization layer is then used as the final output of the two fully connected layers and fed into the softmax layer. The final output of the segment-level features after passing through the two fully connected layers is of length N. l The vector, N l The total number of categories representing the degree of hearing loss is N in this embodiment. l =16;

[0074] S33: Feed the final outputs of the two fully connected layers into the Softmax layer: The outputs of the two fully connected layers are of length N. l The vector, after passing through the Softmax layer, provides a classification of the degree of hearing loss before distorted speech compensation, specifically represented as follows:

[0075]

[0076] In the formula This represents the speech features fed into the Softmax layer, where the subscript i indicates the level of hearing loss classification, i∈{1,2,...,N}. l}, o i (z) represents the predicted probability of the Softmax layer for each hearing loss level classification based on the input speech features.

[0077] S34: The frame-level features of the speech to be tested are simultaneously input into the quality prediction sub-network: the frame-level features after shape resetting are first passed through a recurrent neural network with a self-attention mechanism to obtain segment-level features, and these segment-level features are then mapped through a fully connected layer to obtain the predicted value of the quality score.

[0078] S4: The obtained frame-level features are simultaneously input into the quality prediction sub-network to obtain the predicted quality score. The process is as follows:

[0079] The frame-level features after shape resetting are first passed through a recurrent neural network with a self-attention mechanism to obtain segment-level features. These segment-level features are then mapped through a fully connected layer to obtain the predicted value of the quality score. The quality prediction sub-network consists of a recurrent neural network, an attention mechanism layer, and an activated fully connected layer.

[0080] Specifically, the steps include the following:

[0081] S41: Reshape the frame-level features output by the convolutional neural network to generate frame-level features in the form of T×64;

[0082] S42: The frame-level features after shape resetting are fed into a recurrent neural network combined with an attention mechanism to obtain segment-level features: BiLSTM is used to learn the bidirectional dependency of temporal data, and an attention mechanism is used to compensate for possible information loss in the hidden state of the BiLSTM output. The input feature dimension and hidden state dimension of BiLSTM are both set to 64, so the output feature has the form T×64, where T is the frame number T. The attention mechanism calculates the hidden state output after weighting by the attention weight vector for each time step of the bidirectional output of BiLSTM. Taking the t-th time step of forward propagation as an example, the output of the attention mechanism is:

[0083]

[0084] In the formula, m represents the current time step. Let f be the hidden output column vector at time step i, where the superscript f indicates forward computation. Let represent the attention weight vector from time step i to the current time step m, calculated as follows:

[0085]

[0086] In the formula W aThe weight matrix is ​​a learnable matrix, M is the number of time steps, and the superscript T indicates matrix transpose; the segment-level features output by the recurrent neural network combined with the attention mechanism are a vector of length 128 points.

[0087] S43: The segment-level features output by the recurrent neural network with attention mechanism are fed into the fully connected layer for mapping to obtain the predicted quality score: The fully connected layer used for quality score mapping consists of 128 input nodes and 1 output node. After the segment-level features pass through this fully connected layer, they are activated by the sigmoid function to give the final predicted quality score.

[0088] S5: The hearing aid speech quality self-evaluation network is trained using training data labeled with hearing aid speech quality indicators. The loss function is a weighted combination of the loss functions of the quality prediction subnetwork and the hearing loss classification subnetwork.

[0089] The training data used in this embodiment consisted of 11,572 sentences spoken by 28 speakers (14 males and 14 females) from the Voice Bank Corpus speech database. These sentences were randomly superimposed with one of 15 noise types from the NoiseX-92 noise set at any signal-to-noise ratio between -5 and 15 dB. Each noisy speech sentence was then denoised using one of three enhancement algorithms: traditional Wiener filtering, Wiener filtering based on prior signal-to-noise ratio, and multi-band spectral subtraction. The unprocessed noisy speech and the denoised speech processed by the three enhancement algorithms were used as distorted audio samples for the training set. The training set contained 11,572 × 4 samples, which were then used as training labels by calculating the Hearing Aid Speech Quality Index (HASQI) under random hearing loss conditions. The hearing loss data constituting the condition for hearing loss came from audiograms of 338 ears of 169 hearing-impaired patients. According to the 1997 WHO hearing loss classification table, there were 8 ears with normal hearing loss, 18 ears with mild hearing loss, 103 ears with moderate hearing loss, 128 ears with severe hearing loss, and 81 ears with profound hearing loss.

[0090] During training, the loss function of the speech quality self-evaluation network is calculated. This loss function is a weighted combination of the loss functions of the quality prediction sub-network and the hearing loss classification sub-network. A weighting factor β adjusts the importance of the hearing loss classification sub-network in the speech quality self-evaluation network; in this embodiment, β = 0.2.

[0091] Loss=(1-β)Loss score +βLoss level

[0092] Among them, Loss score The loss function of the quality prediction subnetwork is specifically expressed as:

[0093]

[0094] Loss level The loss function of the hearing loss classification subnetwork is specifically expressed as:

[0095]

[0096] Among them, o i N represents the probability predicted by the Softmax classifier for the i-th degree of hearing loss. l The total number of categories representing the degree of hearing loss, level i The true probability of the i-th degree of hearing loss is derived from the hearing loss level classification.

[0097]

[0098] In the formula, level represents the classification level of hearing loss, which is composed of the average hearing threshold loss avg at 500Hz, 1kHz, 2kHz, and 4kHz. HL The calculation yields the following result, specifically expressed as:

[0099]

[0100] In the formula, ceil represents rounding down, and level is 1 to N. l Integers between [a certain range].

[0101] The network described in this embodiment is trained according to the above five steps, using SQINet and SQINet-Class as comparison networks, and trained with the same training data. The SQINet comparison network is a network structure based on this invention with the hearing loss classification sub-network removed; that is, it consists only of a frame-level feature extraction network and a quality prediction network. SQINet-Class is a quality classification-based SQINet, also using a multi-task training strategy, except that the training objective of the auxiliary task is the quality classification label.

[0102] The test data used in this embodiment also comes from the Voice Bank Corpus speech database, constructed using 824 sentences spoken by the remaining two speakers (one male and one female) in the database. Each sentence was processed using one type of noise from the NoiseX-92 noise set, randomly superimposed with a signal-to-noise ratio between 5 and 15 dB. Then, one of the three speech enhancement algorithms used in constructing the training set, or the option to not enhance the speech, was applied to the noisy speech. The test set contains a total of 824 samples, and the Hearing Aid Speech Quality Index (HASQI) was calculated as the true value for these samples under random hearing loss conditions. Here, the hearing loss data constituting the hearing loss condition comes from audiograms of 14 ears from 7 hearing-impaired patients, different from the training set. According to the 1997 WHO hearing loss classification table, there are 2 ears with moderate hearing loss, 10 ears with severe hearing loss, and 2 ears with profound hearing loss.

[0103] To verify the accuracy of the hearing aid speech quality self-evaluation of this invention, the aforementioned test set was used to predict the results using the described method and two contrast networks, and the difference between the predicted values ​​and the true values ​​of each network was calculated. The results were measured using three evaluation metrics: Pearson correlation coefficient (PCC), root mean square error (RMSE), and mean absolute error (MAE). PCC describes the degree of linear correlation between two variables, and its calculation formula is as follows:

[0104]

[0105] In the formula, N is the number of samples, and MOS is... o For objective scoring, MOS s Subjective rating and These are the average of the objective scores and the average of the subjective scores, respectively; RMSE and MAE both describe the error between the two variables, and are calculated as follows:

[0106]

[0107]

[0108] The prediction accuracy of the network described in this embodiment and the comparison network on the test set is shown in Appendix Table 1.

[0109]

[0110] Table 1

[0111] As can be seen from the table, the method described in this invention significantly outperforms the comparison network in all indicators. Specifically, the method described in this invention exhibits the highest linear correlation with the Hearing Aid Speech Quality Index (HASQI) and the smallest difference. Furthermore, the waveforms of the speech processed by this invention after noise addition, enhancement, and hearing loss condition treatment are shown in the table below. Figure 5 As shown, all of these demonstrate that the hearing aid of the present invention has a higher accuracy in self-assessing the voice quality.

[0112] The above description is merely a description of preferred embodiments of the present invention and is not intended to limit the scope of the present invention in any way. Any changes or modifications made by those skilled in the art based on the above disclosure shall fall within the protection scope of the claims.

Claims

1. A self-evaluation method for hearing aid speech quality based on hearing loss classification, characterized in that, Includes the following steps: S1: Construct a hearing aid speech quality self-evaluation network that includes a frame-level feature extraction network, a hearing loss classification sub-network, and a quality prediction sub-network; S2: Input the shallow features of the speech to be tested into the frame-level feature extraction network to obtain frame-level features; S3: Input the obtained frame-level features into the hearing loss classification sub-network to obtain the classification of the degree of hearing loss before distorted speech compensation; S4: Input the obtained frame-level features into the quality prediction sub-network to obtain the predicted quality score; S5: The hearing aid speech quality self-evaluation network is trained using training data labeled with hearing aid speech quality indicators. The loss function is a weighted combination of the loss functions of the quality prediction subnetwork and the hearing loss classification subnetwork. The specific process of training using training data labeled with hearing aid speech quality indicators is as follows: During training, the loss function of the speech quality self-evaluation network is calculated until the loss function value is less than a threshold, thus completing the training. The loss function of the speech quality self-evaluation network is a weighted combination of the loss functions of the quality prediction sub-network and the hearing loss classification sub-network. The weight factor β adjusts the importance of the hearing loss classification sub-network in the speech quality self-evaluation network. Loss=(1-β)Loss score +βLoss level Among them, Loss score The loss function of the quality prediction subnetwork is expressed as follows: Loss level The loss function for the hearing loss classification subnetwork is expressed as follows: Among them, o i N represents the probability predicted by the Softmax classifier for the i-th degree of hearing loss. l The total number of categories representing the degree of hearing loss, level i The true probability of the i-th degree of hearing loss is derived from the hearing loss degree classification level: In the formula, level represents the classification level of hearing loss, which is composed of the average hearing threshold loss avg at 500Hz, 1kHz, 2kHz, and 4kHz. HL The calculation yields the following result: In the formula, ceil represents rounding down, and level is... Integers between [a certain range].

2. The hearing aid speech quality self-evaluation method based on hearing loss classification as described in claim 1, characterized in that, In S2, the shallow features of the speech to be tested are input into the frame-level feature extraction network to obtain frame-level features. The specific process is as follows: Shallow features are calculated based on the processed signal from the hearing aid. A frame-level feature extraction network is used to learn the deep representation of the distorted signal, thereby obtaining frame-level features. The frame-level feature extraction network is composed of a convolutional neural network.

3. The hearing aid speech quality self-evaluation method based on hearing loss classification as described in claim 2, characterized in that, In the S2 process, shallow features are calculated based on the signal processed by the hearing aid: This feature represents the average filter bank energy in each channel of the Gammatone filter for each frame, where S represents the short-time logarithmic amplitude spectrum of the distorted signal on the auditory frequency scale after framing and windowing, and c, t, and n are the number of channels C, the number of frames T, and the frame length N of the Gammatone filter bank, respectively. The shape of the final shallow feature of the distorted signal is T×32.

4. The hearing aid speech quality self-evaluation method based on hearing loss classification as described in claim 3, characterized in that, The convolutional neural network used for frame-level feature extraction consists of four stacked convolutional networks. Each convolutional network contains a two-dimensional convolutional layer, a batch normalization layer, and a PReLU activation function layer. The number of output features of each two-dimensional convolutional layer is 8, 8, 16, and 16, respectively. The kernel sizes are [5, 5], [5, 5], [3, 5], and [3, 5], respectively. The stride sizes are [1, 1], [1, 2], [1, 2], and [1, 2], respectively. The padding widths are [2, 2], [2, 2], [1, 2], and [1, 2], respectively. The frame-level feature shape extracted by the convolutional neural network from the shallow features of the distorted signal is represented as 16×T×4, where T is the number of frames.

5. The hearing aid speech quality self-evaluation method based on hearing loss classification as described in claim 3, characterized in that, In S3, the obtained frame-level features are input into the hearing loss classification subnetwork to obtain the classification of the degree of hearing loss before distorted speech compensation. The process is as follows: The frame-level features after shape resetting are first globally averaged to obtain segment-level features. Then, through a set of batch-normalized fully connected layers and Softmax layers, the classification of hearing loss degree before distortion speech compensation is obtained. The hearing loss classification subnetwork consists of a set of batch-normalized fully connected layers and Softmax layers.

6. The self-evaluation method for hearing aid speech quality based on hearing loss classification as described in claim 5, characterized in that, In S3, the obtained frame-level features are input into the hearing loss classification subnetwork to obtain the classification of the degree of hearing loss before distorted speech compensation. The specific process is as follows: S31: First, reshape the extracted frame-level features to generate T×64 features, where T is the number of frames. Then, perform global averaging on the reshaped frame-level features to achieve segment-level features. S32: The segment-level features are input into two fully connected layers of the hearing loss classification sub-network, which consists of two fully connected layers and a Softmax layer stacked together. The output of each fully connected layer passes through a batch normalization layer. The output of the first batch normalization layer is activated by the ReLU function and then fed into the second fully connected layer. The output of the second batch normalization layer is then used as the final output of the two fully connected layers and fed into the Softmax layer. The final output of the segment-level features after passing through the two fully connected layers is of length N. l The vector, N l The total number of categories representing the degree of hearing loss; S33: Feed the final outputs of the two fully connected layers into the Softmax layer: The outputs of the two fully connected layers are of length N. l The vector, after passing through the Softmax layer, provides a classification of the degree of hearing loss before distorted speech compensation, specifically represented as follows: In the formula, z = [z1, z2, ..., z Nl The image represents the speech features fed into the Softmax layer, and the subscript i indicates the level of hearing loss classification, i∈{1,2,...,N}. l }, o i (z) represents the predicted probability of the Softmax layer for each hearing loss level classification based on the input speech features.

7. The hearing aid speech quality self-evaluation method based on hearing loss classification as described in claim 3, characterized in that, In S4, the obtained frame-level features are simultaneously input into the quality prediction sub-network to obtain the predicted quality score. The process is as follows: The frame-level features after shape resetting are first passed through a recurrent neural network with a self-attention mechanism to obtain segment-level features. These segment-level features are then mapped through a fully connected layer to obtain the predicted value of the quality score. The quality prediction sub-network consists of a recurrent neural network, an attention mechanism layer, and an activated fully connected layer.

8. The self-evaluation method for hearing aid speech quality based on hearing loss classification as described in claim 7, characterized in that, In S4, the obtained frame-level features are simultaneously input into the quality prediction sub-network to obtain the predicted quality score. The specific process is as follows: S41. Reshape the frame-level features to generate T×64 frame-level features, where T is the number of frames; S42. The frame-level features after shape reshaping are fed into a recurrent neural network with an attention mechanism to obtain segment-level features: BiLSTM is used to learn the bidirectional dependency of time series data, and an attention mechanism is used to compensate for the possible information loss in the hidden state output by BiLSTM. S43. The segment-level features output by the recurrent neural network with attention mechanism are fed into a fully connected layer to obtain the predicted quality score. After passing through the fully connected layer, the segment-level features are activated by the sigmoid function to give the final predicted quality score.

9. The self-evaluation method for hearing aid speech quality based on hearing loss classification as described in claim 8, characterized in that, The input feature dimension and hidden state dimension of BiLSTM are both set to 64, so the output feature size is T×64, where T is the number of frames. The attention mechanism calculates the hidden state output after weighting by the attention weight vector at each time step of the bidirectional output of BiLSTM. At the m-th time step of forward propagation, the output of the attention mechanism is: In the formula, m represents the current time step. Let f be the hidden output column vector at time step i, where the superscript f indicates forward computation. Let represent the attention weight vector from time step i to the current time step m, calculated as follows: In the formula W a M is the learnable weight matrix, M is the number of time steps, and the superscript T indicates matrix transpose. Finally, the segment-level features output by the recurrent neural network combined with the attention mechanism are a vector of length 128 points.