Non-intrusive binaural speech evaluation method based on multi-source sonogram fusion

By constructing a binaural speech evaluation method based on multi-source sound-image fusion, and utilizing a self-supervised learning model and LoRA technology, the shortcomings of non-invasive speech quality evaluation for hearing aids are addressed, achieving high-precision and efficient speech quality evaluation, and improving the model's generalization ability and computational efficiency.

CN122002204AInactive Publication Date: 2026-05-08NANJING INST OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
NANJING INST OF TECH
Filing Date
2026-02-05
Publication Date
2026-05-08
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing non-invasive speech quality assessment methods for hearing aids lack a holistic binaural assessment model. Traditional subjective assessments are cumbersome and time-consuming, while existing objective assessment methods are ineffective under individualized hearing loss and environmental interference, and consume high computational resources.

Method used

A non-invasive binaural speech evaluation method based on multi-source acoustic-image fusion is constructed. Through a binaural speech quality evaluation framework, a self-supervised learning model is used to optimize feature fusion and prediction, and LoRA technology is combined for fine-tuning to achieve high-precision evaluation.

Benefits of technology

It achieves high-precision evaluation of the output speech quality of hearing aids, reduces computational resource consumption, improves the generalization ability and prediction accuracy of the model, and meets the evaluation needs of different application scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122002204A_ABST
    Figure CN122002204A_ABST
Patent Text Reader

Abstract

The invention discloses a non-intrusive binaural speech evaluation method based on multi-source sound image fusion, which belongs to the technical field of audio signal processing and deep learning, and comprises the following steps of: extracting high-dimensional weighted representation of speech to be detected by using a leading-edge pre-training self-supervised learning speech basic model, extracting a frequency spectrum amplitude, and extracting a high-dimensional weighted representation of the speech to be detected; and the personalized audiogram data and the personalized audiogram data expanded along the frequency dimension are jointly used as the input of a feature fusion module in the binaural speech evaluation network so as to deeply mine the frame-level potential features of the speech to be detected. A dual-branch evaluation network is constructed: one combines the fused features with SSL representation to predict a speech quality score, and the other combines the fused features with audiogram data to output a quality level. According to the method, synchronous evaluation of the hearing aid voice quality score and grade is effectively realized, the Pearson linear correlation and the Spearman monotonous correlation between the predicted value and the real evaluation index are obviously improved, and a high-precision solution is provided for personalized hearing evaluation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of audio signal processing and deep learning technology, specifically to a non-invasive binaural speech evaluation method based on multi-source acoustic-graphic fusion. Background Technology

[0002] Hearing loss remains a global public health challenge. Hearing aids, as the best current option for improving hearing in individuals with hearing impairments, directly impact the wearing rate and satisfaction of patients. The effectiveness of hearing aid compensation is often assessed by the quality of the output speech. However, traditional subjective hearing aid speech quality assessment relies on audiologists and a large number of test subjects scoring speech quality; this process is cumbersome, time-consuming, and prone to individual variability. Existing objective hearing aid speech quality assessment also faces numerous challenges and room for improvement.

[0003] Early objective speech quality assessment methods, such as the speech quality assessment metrics PESQ and POLQA proposed by the International Telecommunication Union (ITU), and Quality-Net and MOS-Net proposed by researchers in related fields using deep learning technology, have proven effective. However, these metrics are applicable to people with normal hearing and are used to predict subjective average opinion scores. The objective evaluation methods for hearing aid speech quality are more complex and challenging due to the influence of various factors, including individualized hearing loss and the speech environment. Based on the presence or absence of a reference speech, hearing aid speech quality assessment metrics can be divided into invasive and non-invasive methods. Invasive speech quality assessment metrics, such as HASQI and PEMO-Q-HI, offer high accuracy, but clean signals are often difficult to obtain in real-world environments. Non-invasive hearing aid speech quality assessment, which does not require a reference signal, is more favored by researchers in related fields and is gradually becoming a key research direction. In recent years, deep learning technology and speech fundamental models have developed rapidly, and their powerful nonlinear modeling capabilities have shown great potential in non-invasive hearing aid speech assessment methods. However, there is currently a lack of models for comprehensive binaural assessment. Furthermore, given the current state of advanced self-supervised learning speech models, it is worthwhile to further explore their potential role in hearing aid speech quality assessment.

[0004] Therefore, it is evident that existing non-invasive speech quality assessment methods for hearing aids still have significant room for improvement. Given the current challenges, developing a non-invasive binaural speech assessment method based on multi-source acoustic image fusion is of great research value and significance. Summary of the Invention

[0005] Objective: This invention aims to eliminate reliance on subjective speech quality assessment for hearing aids and to further optimize and overcome the shortcomings of existing non-invasive objective speech quality assessment methods for hearing aids. By constructing a comprehensive assessment framework applicable to binaural speech quality and utilizing cutting-edge speech models for deep optimization and adaptation, high-precision assessment of hearing aid output speech quality can be achieved. While improving model prediction accuracy, computational resource consumption is effectively reduced, achieving a dual optimization of assessment performance and efficiency.

[0006] Technical Solution: To solve the above-mentioned technical problems and achieve the purpose of this invention, the technical solution adopted by this invention is: a non-invasive binaural speech evaluation method based on multi-source acoustic-image fusion, comprising the following steps:

[0007] Step 1: Obtain the speech dataset for model training, evaluation, and testing; simulate the acquisition using binaural clean speech through three stages: adding noise, enhancing denoising, and individual compensation;

[0008] Step 2: Based on the speech dataset, calculate the speech quality scores and speech quality levels for the left and right audio channels in binaural speech;

[0009] Step 3: Extract the spectral amplitude of the speech from the left and right ears, and extract the weighted speech representation output by the pre-trained self-supervised learning speech base model. First, the spectral amplitude is obtained by framing and windowing the speech to be tested. Then, STFT is performed on each frame of data and the weighted representation is obtained by weighted calculation of the output of each hidden layer of the self-supervised learning model. Finally, the audiogram is expanded along the frequency and embedded into the speech features.

[0010] Step 4: Construct a binaural speech evaluation model, including a feature fusion network, a speech scoring prediction network, and a speech quality grading network; design a loss function and use backpropagation to train the binaural speech evaluation model;

[0011] Step 5: Evaluate binaural speech based on the trained binaural speech evaluation model.

[0012] Furthermore, step 2 specifically involves: both left and right ear audiograms are 1*8 dimensions, denoted as... and The audiogram dimensions correspond to frequencies [250Hz, 500Hz, 1kHz, 2kHz, 3kHz, 4kHz, 6kHz, 8kHz]; during each speech processing step, the following are recorded: the true value of the 3D speech quality evaluation, for the left ear, right ear, and the overall score, denoted as […]. , ,and ; 2D speech level label values, denoted as follows for the left and right ears respectively: and .

[0013] Furthermore, in step 3, the spectral amplitudes of the speech signals from the left and right ears are denoted as follows: and The speech weighted representations are denoted as follows: and Taking the left ear as an example, the expression is as follows:

[0014]

[0015] in, , Indicates frame length, and These represent the audio frames before and after windowing, respectively. The Hamming window function is specifically shown in equation (2):

[0016]

[0017] in, It is a constant value, set to 0.46;

[0018] Secondly, a 512-point STFT is performed on each frame of data, with the frame shift set to half the frame length. The spectral amplitude of the speech frame is calculated, and the STFT is performed on the windowed speech. The specific expression is as follows:

[0019]

[0020] in, express The corresponding complex spectrum, Here, f is the frequency index, and f is the frequency corresponding to a single feature frequency point of the speech.

[0021]

[0022] in, express Spectral amplitude, Indicates the real part, The imaginary part is represented; the spectral amplitude characteristics of the entire left ear speech are denoted as: The same applies to the right ear, denoted as... ;

[0023] The speech weighted representation is obtained by weighting the output features of each hidden layer of a self-supervised speech base model. Taking the left ear as an example, the specific expression is as follows:

[0024]

[0025] in, Indicates the first The layer outputs the weights corresponding to the speech representation. The sum of the weights of each layer is 1; Indicates the first The right ear outputs a speech representation, and the same logic applies to the other ear. .

[0026] Furthermore, in step 3, the audiogram is expanded along the frequency and mapped according to the following formula to obtain a set of 1*256 audiogram loss features;

[0027]

[0028] in, The frequency corresponding to a single feature frequency point of speech; HL data after expanding the left ear audiogram i L The left ear audiogram is number 1 Each hearing valve, i=0,1…7, corresponds to the speech spectrum characteristics. and speech weighted representation After being combined, its data dimension size is [3, number of frames, number of features], where 3 is the number of channels, the first channel is the speech weighted representation, the second channel is the speech spectrum feature, and the third channel is the extended audiogram.

[0029] Further, step 4 specifically involves: constructing a feature fusion network model, a speech scoring prediction network model, and a speech quality grading network model; combining the spectral features of the left and right ears respectively, weighting the representation, and generating an audiogram, denoted as […]. , , ]and[ , , The joint features serve as the input to the feature fusion network model. The outputs of the left and right ear feature fusion network models are denoted as follows: and The features output by the fusion network are combined with the weighted representation and the audiogram, respectively. , ]and[ , These are used as inputs to the speech quality score prediction network model and the speech quality classification network model, respectively. Similarly, the input features for the right ear are denoted as […]. , ]and[ , The predicted values ​​output by the left and right ear fraction prediction networks are denoted as follows: and The overall voice quality score obtained by the decision-making level is recorded as follows: The output labels of the speech quality grading network are respectively denoted as... and The overall loss of the binaural speech evaluation model consists of the joint mean absolute error loss of the left and right ear speech quality prediction scores, the overall speech quality prediction score, and the left and right ear speech quality rating labels. The overall network model is trained by backpropagation using the loss function. Based on this, the pre-trained speech base model is fine-tuned using low-rank adaptive (LoRA) technology to make it better suited for downstream speech quality scoring and grading tasks, ultimately obtaining a high-performance hearing aid speech evaluation model.

[0030] Furthermore, in step 4, the activation layer of the feature fusion network adopts... The activation function, its mathematical expression is shown below:

[0031]

[0032] in, The slope of the negative region;

[0033] The pooling layer uses L2 norm pooling, as shown in the following formula:

[0034]

[0035] in, This represents the input data for the pooling layer. Represents the output of the pooling layer; coefficients Determine the pooling norm type;

[0036] The output of the left ear feature fusion network is denoted as: The same applies to the right ear, denoted as... .

[0037] Furthermore, in step 4, the BiLSTM layer and multi-head attention mechanism layer in the speech scoring prediction network are shown in expressions (9)-(12);

[0038] The BiLSTM layer consists of a forward LSTM and a backward LSTM, learning bidirectional dependencies in time series data and utilizing contextual information. Its formula is expressed as follows:

[0039]

[0040] in, Indicates the index of the current frame. Indicates activation function function, These are the activation values ​​for the input gate, forget gate, and output gate, respectively. and These represent the current cell state and the cell state at the previous time step, respectively. and These represent the current hidden state and the hidden state at the previous time step, respectively. Indicates the current input frame. These are the weight matrices for the input gate, forget gate, cell state, and output gate, respectively. Then these are the bias vectors for each gate; This indicates that each element is multiplied one by one. The hyperbolic tangent activation function is defined as follows:

[0041]

[0042] in, This represents the natural exponential function. This indicates the input to the function;

[0043] The mathematical expression for the multi-head attention layer is as follows:

[0044]

[0045]

[0046] in, They are the first Size Query, key and value matrix, This indicates the dimensions of the query and the key, while This represents the projection matrix that maps the cascaded outputs of all heads to the final output, where T represents the matrix transpose.

[0047] The outputs of the speech scoring prediction networks for the left and right ears are denoted as follows: and .

[0048] Furthermore, in step 4, the activation layer and Softmax classification layer in the speech quality classification network are as shown in expressions (13)-(14);

[0049] Activation layer selection The activation function, mathematically expressed as follows:

[0050]

[0051] in, The slope of the negative region is a learnable parameter that is updated during training; the output of the Softmax layer represents the probability distribution of each speech quality level; argmax is used to extract the index corresponding to the maximum probability in the Softmax output, which serves as the classification label; its mathematical expression is as follows:

[0052]

[0053] in, It is the normalized output of the fully connected layer, which is then activated by the activation function. The output, The dimension is the number of category labels; This represents the output of the speech quality grading network for the left ear; the same applies to the right ear, denoted as: .

[0054] Furthermore, in step 4, the loss function is as shown in expression (15).

[0055]

[0056] in, The mean absolute error between the overall speech quality score output by the decision fusion layer and the actual overall speech quality score; and These are the mean absolute errors between the predicted speech quality scores for the left and right ears and the actual scores, respectively. and These represent the average absolute error between the predicted labels and the actual labels for the speech quality grading of the left and right ears, respectively.

[0057] The present invention also discloses a computer device, including a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the steps of the method of the present invention.

[0058] The present invention also discloses a computer-readable storage medium having a computer program / instructions stored thereon, which, when executed by a processor, implements the steps of the method of the present invention.

[0059] Beneficial effects: Compared with the prior art, the present invention has the following significant advantages:

[0060] 1. The network model proposed in this invention overcomes the shortcomings of existing monoaural assessments by integrating binaural signal features to achieve comprehensive quality assessment. Furthermore, it can simultaneously predict the overall quality and quality level of the hearing aid's binaural signals, meeting the assessment needs of different application scenarios.

[0061] 2. This invention utilizes LoRA technology to fine-tune a pre-trained self-supervised speech foundation model. Compared to traditional methods, this model can more accurately capture subtle speech quality degradation, thus achieving a significant breakthrough in objective prediction accuracy.

[0062] 3. This invention designs a multi-task learning framework based on feature sharing. By utilizing quality grading tasks to effectively assist the learning of regression prediction tasks, and leveraging the co-correlation between tasks, it significantly improves the model's generalization ability and various prediction indicators, outperforming independent single-task models. Attached Figure Description

[0063] Figure 1 This is a model structure diagram of a non-invasive binaural speech evaluation method based on multi-source acoustic-graphic fusion according to the present invention;

[0064] Figure 2 This is a diagram of the feature fusion network structure of the present invention.

[0065] Figure 3 This is a network structure diagram of the quality scoring and grading system of the present invention.

[0066] Figure 4 This is a schematic diagram comparing the scatter plot mapping between the model prediction and the actual value in an embodiment of the present invention.

[0067] Figure 5 This is a performance comparison chart of the method proposed in this invention and related cutting-edge indicators. Detailed Implementation

[0068] The invention will be further described below with reference to the accompanying drawings. This invention discloses a non-invasive binaural speech evaluation method based on multi-source audiogram fusion, aiming to eliminate reliance on subjective speech evaluation and overcome and optimize some of the challenges of current non-invasive hearing aid speech evaluation methods. First, this invention acquires left and right dual-channel speech datasets. -5dB to -5dB white noise is randomly added to the clean dual-channel speech, and Wiener filtering is used for noise reduction. Subsequently, WDRC combined with the audiogram is used to perform personalized compensation processing on the speech, simulating the output speech of a hearing aid. Next, the spectral amplitude of the speech is extracted, along with the weighted representation of the speech output by a pre-trained self-supervised learning speech base model, while the audiogram is expanded along the frequency. These three are combined as input to a feature fusion network, whose output is concatenated with the original weighted representation and the original audiogram, respectively, and used as input to a scoring prediction network and a quality grading network. The predicted speech quality scores for both ears are jointly input to a decision fusion layer, whose output serves as the overall predicted score for binaural speech quality. By calculating the sum of the average absolute error loss between the left and right ear speech quality rating labels, the left and right ear speech quality prediction scores, and the overall speech quality prediction score and their respective true scores, and using this sum as the loss function for backpropagation optimization, a high-performance binaural speech evaluation model for hearing aids is finally obtained.

[0069] like Figure 1 As shown, this invention discloses a non-invasive binaural speech evaluation method based on multi-source acoustic-graphic fusion, comprising the following steps:

[0070] Step 1: Obtain the dataset for model training and testing; randomly select 6000 audio samples from the clean dataset in the DNS2020 database, expanding from mono audio to dual-channel audio; of these, 4800 samples are used for training, 800 for evaluation during training, and 400 for testing after training; the audio sampling rate is 16000Hz, and the duration is 10s; randomly add -5 to the audio samples. After adding 5dB of white noise, Wiener filtering is used for noise reduction. After noise reduction, the speech is personalized and compensated using multi-channel WDRC combined with the prescription formula FIG6.

[0071] Step 2: Based on the speech obtained in Step 1, calculate the speech quality score and quality level label required for training and testing; the original audiograms for both ears are 1*8 in dimension, and each value is the minimum threshold for a hearing-impaired patient to hear a pure tone at the corresponding frequency, denoted as follows: and The audiogram dimensions correspond to frequencies of [250Hz, 500Hz, 1kHz, 2kHz, 3kHz, 4kHz, 6kHz, 8kHz]; During each speech processing step, the following are recorded: the true value of the 3D speech quality assessment, and the true values ​​of the speech quality assessment for both the left and right ears, calculated using HASQI, and denoted as follows: and The true value of the overall binaural speech quality assessment is taken as the maximum value of the speech quality scores of the left and right ears, denoted as: [denoted as...]. The rating values ​​all range from 0 to 1; the 2D speech level label value is denoted as... and The labels are based on the subjective speech quality level perception classification of ITU-T P.800 P.800, and are divided into 5 levels: excellent, good, average, poor, and very poor, corresponding to label values ​​of 0, 1, 2, 3, and 4 respectively.

[0072] Step 3: Based on the binaural speech obtained in Step 1, obtain the input features of the binaural evaluation framework; first, extract the spectral amplitude of the left and right channels in the binaural signal. Taking the left ear as an example, the speech is framed and windowed.

[0073]

[0074] in, , Indicates frame length, and These are the audio frames before and after the windowing function, respectively. Let be the Hamming window function, and its mathematical expression is shown in equation (2):

[0075]

[0076] in, It is a constant value, set to 0.46;

[0077] Secondly, a 512-point STFT is performed on each frame of data, with the frame shift set to half the frame length, to calculate the spectral amplitude of the speech frame. The specific expression for STFT of windowed speech is as follows:

[0078]

[0079] in, express The corresponding complex spectrum, Frequency index;

[0080]

[0081] in, express Spectral amplitude, Indicates the real part, The imaginary part is represented; the spectral amplitude characteristics of the entire left ear speech are denoted as: The same applies to the right ear, denoted as... ;

[0082] The speech weighted representation is obtained by weighting the output features of each hidden layer in a self-supervised learning speech base model. The speech base model used is WavLM; taking the left ear as an example, the specific expression is as follows:

[0083]

[0084] in, Indicates the first The layer outputs the weights corresponding to the speech representation. The sum of the weights of each layer is 1. Indicates the first The layer outputs the speech representation. The same logic applies to the right ear. .

[0085] The original audiogram data has a dimension of 1*8 and is expanded along the frequency direction of the speech spectrum amplitude features; according to the mapping rule of Equation (6), a set of 1*256 audiogram loss features is obtained;

[0086]

[0087] in, This refers to the frequency corresponding to a single feature point in speech; taking the left ear as an example. Personalized information features after expanding the left ear audiogram, and speech spectrum features and speech weighted representation After being combined, its data dimension size is [3, number of frames, number of features], where 3 is the number of channels, the first channel is the speech weighted representation, the second channel is the speech spectrum feature, and the third channel is the extended audiogram;

[0088] Step 4: Construct a binaural speech evaluation framework and train the model; such as Figure 2 As shown in the feature fusion network structure diagram, the construction of the feature fusion network model aims to further integrate and explore the intrinsic connections and deep representations between multi-source acoustic image features. The input features of the feature fusion network models for the left and right ears are respectively, [ , , ]and[ , , The input features are processed through 5 layers of 2D convolutional layers and 3 layers of L2-norm pooling. Each convolutional layer has a 3×3 kernel, a stride of (1, 1), and edge padding of (1, 1). Each convolutional layer is paired with a batch normalization layer and a Leaky ReLU activation layer. The p-value of the pooling layers is set to 2, the pooling window is 1×4, and the pooling stride is 1×4. The feature extraction network employs a multi-layer CNN structure to capture deep representations in the input features. The Leaky ReLU activation layer effectively alleviates the "neuron death" problem caused by sparse gradients by introducing a small slope on the negative half-axis, ensuring stable backpropagation of gradients in deep networks while introducing nonlinearity. L2-norm pooling more comprehensively preserves feature information in local regions and significantly enhances the model's robustness to phase shifts and small feature deformations.

[0089] in, The activation layer, mathematically expressed as follows:

[0090]

[0091] in, The slope of the negative region, which is set to 0.01 by default;

[0092] The L2 norm pooling layer is shown in the following equation:

[0093]

[0094] in, This represents the input data for the pooling layer. Represents the output of the pooling layer; coefficients The pooling norm type is determined and set to 2 in this invention.

[0095] The output of the left ear feature fusion network is denoted as: The same applies to the right ear, denoted as... ;

[0096] like Figure 3 The network structure diagram for quality scoring and grading is presented, which constructs predictions for speech quality scores and grade labels. The quality scoring and grading employs a multi-task learning strategy, based on feature sharing from the output of the feature fusion model.

[0097] The aim is to leverage the synergistic relationships between tasks to improve the model's generalization ability and predictive performance; taking the left ear as an example, the input for its speech quality score is […], fusing the features and weighted representations output by the network, […]. , This task consists of a bidirectional LSTM, multiple attention mechanism layers, an adaptive global max pooling layer, and a fully connected layer. The output of the fully connected layer is activated by a sigmoid function, causing its prediction score to be between 0 and 1.

[0098] The bidirectional LSTM layer consists of a forward LSTM and a backward LSTM, which can learn bidirectional dependencies in time series data and utilize contextual information. Its formula is expressed as follows:

[0099]

[0100] in, Indicates the index of the current frame. Indicates activation function function, These are the activation values ​​for the input gate, forget gate, and output gate, respectively. For the unit state, and These represent the current hidden state and the hidden state at the previous time step, respectively. Indicates the current input frame. These are the weight matrices for the input gate, forget gate, cell state, and output gate, respectively. Then these are the bias vectors for each gate; This indicates that each element is multiplied one by one. The hyperbolic tangent activation function is defined as follows:

[0101]

[0102] Furthermore, the multi-head attention layer can learn the impact of each frame of speech on the task prediction result, assigning different attention weights to different frames, which is beneficial to improving the accuracy of the prediction task. Its mathematical expression is as follows:

[0103]

[0104]

[0105] in, They are the first Size Query, key and value matrix, This indicates the dimensions of the query and the key, while This represents the projection matrix that maps the cascaded outputs of all heads to the final output.

[0106] The outputs of the speech scoring prediction networks for the left and right ears are denoted as follows: and ;

[0107] Taking the left ear as an example, the input for speech quality grading is the features output by the fusion network and the original audiogram. , This task consists of an adaptive global average pooling layer, a fully connected layer, and a Softmax layer.

[0108] Among them, the activation layer is selected The activation function, mathematically expressed as follows:

[0109]

[0110] in, The slope, which is the negative part, is a learnable parameter that is updated during training.

[0111] The Softmax layer is used to obtain the probability distribution of each speech quality level. argmax is then used to extract the index corresponding to the maximum probability value in the Softmax output, which serves as the level label. Its mathematical expression is as follows:

[0112]

[0113] in, It is the normalized output of the fully connected layer, which is then activated by the activation function. The output, The dimension representing the number of category labels. This represents the output of the speech quality grading network for the left ear; the same applies to the right ear, denoted as: ;

[0114] Finally, the predicted speech quality values ​​for the left and right ears. and The input-decision fusion layer consists of fully connected layers, with the last layer outputting a Sigmoid activation function whose value ranges from 0 to 1. The mathematical expression of the Sigmoid activation function is as follows:

[0115]

[0116] The overall speech quality score for both ears is denoted as: ;

[0117] The overall loss of the hearing aid's binaural speech assessment framework is calculated as the joint mean absolute error loss of the left and right ear speech quality prediction scores, the overall speech quality prediction score, and the left and right ear speech quality grading labels. The overall network model is trained via backpropagation using the loss function. Based on this, the pre-trained speech base model is fine-tuned using Low-Rank Adaptive Regression (LoRA) to better suit downstream speech quality scoring and grading tasks. During training, the batch size is set to 16, the initial learning rate is 0.0001, and the adaptive Adam method is used for optimization. Training is performed for 300 epochs, with cosine annealing dynamically adjusting the learning rate for each epoch. Finally, a high-performance compensated speech evaluation model is obtained. The loss function formula is:

[0118]

[0119] in, The mean absolute error between the overall speech quality score output by the decision fusion layer and the actual overall speech quality score. , Number of voices; and Let be the mean absolute error between the predicted speech quality scores for the left and right ears and the actual scores, respectively. Its mathematical expression is denoted as: and ; and Let be the mean absolute error between the predicted labels and the true labels for the left and right ear speech quality grading, respectively. and ;

[0120] Figure 4Scatter plots showing the model's predicted values ​​versus the actual values ​​of speech quality and speech intelligibility are presented. The results indicate that the binaural overall speech quality prediction results of this invention exhibit a high degree of linear and monotonic correlation with the actual scores, with the LCC and SRCC of the overall speech quality score reaching 0.991 and 0.989, respectively. Figure 5 The fitting performance of this invention, a pre-trained self-supervised learning speech base model, a LoRA-free fine-tuning algorithm, and other cutting-edge speech quality assessment algorithms was compared on the three metrics of LCC, SRCC, and RMSE. Compared with the pre-trained model and the LoRA-free version, this invention improves the RMSE fitting accuracy by approximately 14.29% and 5.26%, respectively, and outperforms other methods on all metrics, fully validating the effectiveness and superiority of the proposed method.

[0121] The above description is only a preferred embodiment of the present invention. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.

Claims

1. A non-invasive binaural speech assessment method based on multi-source acoustic-graphic fusion, characterized in that, Includes the following steps: Step 1: Obtain the speech dataset for model training, evaluation, and testing; simulate the acquisition using binaural clean speech through three stages: noise addition, denoising enhancement, and individual compensation. Step 2: Based on the speech dataset, calculate the speech quality scores and speech quality levels for the left and right audio channels in binaural speech; Step 3: Extract the spectral amplitude of the speech from the left and right ears, and extract the weighted speech representation output by the pre-trained self-supervised learning speech base model. First, the spectral amplitude is obtained by framing and windowing the speech to be tested. Then, STFT is performed on each frame of data and the weighted representation is obtained by weighted calculation of the output of each hidden layer of the self-supervised learning model. Finally, the audiogram is expanded along the frequency and embedded into the speech features. Step 4: Construct a binaural speech evaluation model, including a feature fusion network, a speech scoring prediction network, and a speech quality grading network; design a loss function and use backpropagation to train the binaural speech evaluation model; Step 5: Evaluate binaural speech based on the trained binaural speech evaluation model.

2. The non-invasive binaural speech assessment method based on multi-source acoustic-graphic fusion according to claim 1, characterized in that, Step 2 specifically involves: The audiograms for both ears are 1*8 dimensions, denoted as... and The audiogram dimensions correspond to frequencies [250Hz, 500Hz, 1kHz, 2kHz, 3kHz, 4kHz, 6kHz, 8kHz]; during each speech processing step, the following are recorded: the true value of the 3D speech quality evaluation, for the left ear, right ear, and the overall score, denoted as […]. , ,and ; 2D speech level label values, denoted for the left and right ears respectively: and .

3. The non-invasive binaural speech assessment method based on multi-source acoustic-graphic fusion according to claim 1, characterized in that, In step 3, the spectral amplitudes of the speech in the left and right ears are recorded as follows: and The speech weighted representations are denoted as follows: and Taking the left ear as an example, the expression is as follows: ; in, , Indicates frame length, and These represent the audio frames before and after windowing, respectively. The Hamming window function is specifically shown in equation (2): ; in, It is a constant value, set to 0.46; Secondly, a 512-point STFT is performed on each frame of data, with the frame shift set to half the frame length. The spectral amplitude of the speech frame is calculated, and the STFT is performed on the windowed speech. The specific expression is as follows: ; in, express The corresponding complex spectrum, Here, f is the frequency index, and f is the frequency corresponding to a single feature frequency point of the speech. ; in, express Spectral amplitude, Indicates the real part, The imaginary part is represented; the spectral amplitude characteristics of the entire left ear speech are denoted as: The same applies to the right ear, denoted as... ; The speech weighted representation is obtained by weighting the output features of each hidden layer of a self-supervised speech base model. Taking the left ear as an example, the specific expression is as follows: ; in, Indicates the first The layer outputs the weights corresponding to the speech representation. The sum of the weights of each layer is 1; Indicates the first The right ear outputs a speech representation, and the same logic applies to the other ear. .

4. The non-invasive binaural speech assessment method based on multi-source acoustic-graphic fusion according to claim 1, characterized in that, In step 3, the audiogram is expanded along the frequency and mapped according to the following formula to obtain a set of 1*256 audiogram loss features; ; in, The frequency corresponding to a single feature frequency point of speech; HL data after expanding the left ear audiogram i L The left ear audiogram is number 1 Each hearing valve, i=0,1…7, corresponds to the speech spectrum characteristics. and speech weighted representation After being combined, its data dimension size is [3, number of frames, number of features], where 3 is the number of channels, the first channel is the speech weighted representation, the second channel is the speech spectrum feature, and the third channel is the extended audiogram.

5. A non-invasive binaural speech assessment method based on multi-source acoustic-graphic fusion according to claim 1, characterized in that, In step 4, the activation layer of the feature fusion network adopts... The activation function, its mathematical expression is shown below: ; in, The slope of the negative region; The pooling layer uses L2 norm pooling, as shown in the following formula: ; in, This represents the input data for the pooling layer. Represents the output of the pooling layer; coefficients Determine the pooling norm type; The output of the left ear feature fusion network is denoted as: The same applies to the right ear, denoted as... .

6. The non-invasive binaural speech assessment method based on multi-source acoustic-graphic fusion according to claim 1, characterized in that, In step 4, the BiLSTM layer and multi-head attention mechanism layer in the speech rating prediction network are shown in expressions (9)-(12); The BiLSTM layer consists of a forward LSTM and a backward LSTM, learning bidirectional dependencies in time series data and utilizing contextual information. Its formula is expressed as follows: ; in, Indicates the index of the current frame. Indicates activation function function, These are the activation values ​​for the input gate, forget gate, and output gate, respectively. and These represent the current cell state and the cell state at the previous time step, respectively. and These represent the current hidden state and the hidden state at the previous time step, respectively. Indicates the current input frame. These are the weight matrices for the input gate, forget gate, cell state, and output gate, respectively. Then these are the bias vectors for each gate; This indicates that each element is multiplied one by one. The hyperbolic tangent activation function is defined as follows: ; in, This represents the natural exponential function. This indicates the input to the function; The mathematical expression for the multi-head attention layer is as follows: ; ; in, They are the first Size Query, key and value matrix, This indicates the dimensions of the query and the key, while This represents the projection matrix that maps the cascaded outputs of all heads to the final output, where T represents the matrix transpose. The outputs of the speech scoring prediction networks for the left and right ears are denoted as follows: and .

7. The non-invasive binaural speech assessment method based on multi-source acoustic-graphic fusion according to claim 1, characterized in that, In step 4, the activation layer and the Softmax classification layer in the speech quality classification network are as shown in expressions (13)-(14); Activation layer selection The activation function, mathematically expressed as follows: ; in, The slope of the negative region is a learnable parameter that is updated during training; the output of the Softmax layer represents the probability distribution of each speech quality level; argmax is used to extract the index corresponding to the maximum probability in the Softmax output, which serves as the classification label; its mathematical expression is as follows: ; in, It is the normalized output of the fully connected layer, which is then activated by the activation function. The output, The dimension is the number of category labels; This represents the output of the speech quality grading network for the left ear; the same applies to the right ear, denoted as: .

8. A non-invasive binaural speech assessment method based on multi-source acoustic-graphic fusion according to claim 1, characterized in that, In step 4, the loss function is as shown in expression (15); ; in, The mean absolute error between the overall speech quality score output by the decision fusion layer and the actual overall speech quality score; and These are the mean absolute errors between the predicted speech quality scores for the left and right ears and the actual scores, respectively. and These represent the average absolute error between the predicted labels and the actual labels for the speech quality grading of the left and right ears, respectively.

9. A computer device comprising a memory, a processor, and a computer program stored in the memory, characterized in that, The processor executes the computer program to implement the steps of the method of claim 1.

10. A computer-readable storage medium having a computer program / instructions stored thereon, characterized in that, When the computer program / instructions are executed by the processor, they implement the steps of the method of claim 1.