Audio data evaluation method, apparatus, device, storage medium, and program product
By performing human ear perception processing and evaluation network evaluation on audio data, the problem of low correlation between audio data evaluation results and human subjective perception in existing technologies is solved, and more accurate audio quality evaluation is achieved.
Patent Information
- Application Number
- CN202311128401.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-09-01
- Publication Date
- 2026-01-23
- Estimated Expiration
- 2043-09-01
AI Technical Summary
Existing audio data evaluation methods ignore human subjective perception when measuring the degree of distortion and spatial quality of spatial audio, resulting in a low correlation between the evaluation results and human subjective evaluation.
By performing human ear perception processing on the reference audio and the target audio, reference human ear perception domain audio and target human ear perception domain audio are generated. An evaluation network is then used for evaluation, taking into account the characteristics of the human ear perception domain, thereby improving the accuracy of the evaluation results and their relevance to subjective human ear evaluation.
It improves the accuracy of audio data evaluation, makes the evaluation results more consistent with human subjective perception, and enhances the correlation between the evaluation results and human subjective evaluation.
Smart Images

Figure CN119559965B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of audio processing technology, and in particular to an audio data evaluation method, apparatus, device, storage medium, and program product. Background Technology
[0002] In the process of evaluating audio data, related technologies consider two quality attributes: objective attributes that measure the degree of distortion in spatial audio and spatial quality. Through multi-task learning and triplet training, neural networks can simultaneously learn the basic audio quality and spatial quality attributes of spatial audio, thereby making a more complete evaluation of the overall quality of spatial audio. In this way, the evaluation process only considers the objective quality of distortion and spatial quality, resulting in a low correlation between the evaluation results and human subjective evaluation. Summary of the Invention
[0003] Therefore, it is necessary to provide an audio data evaluation method, apparatus, electronic device, storage medium, and program product that can improve the accuracy of audio data evaluation in response to the above-mentioned technical problems.
[0004] Firstly, this application provides an audio data evaluation method. The method includes:
[0005] Acquire the raw audio, which includes a reference audio and a target audio to be evaluated;
[0006] The original audio is processed by human ear perception to obtain human ear perception domain audio, which includes reference human ear perception domain audio corresponding to the reference audio and target human ear perception domain audio corresponding to the target audio.
[0007] An evaluation network is used to evaluate and process the original audio and the audio in the human ear's perception domain;
[0008] Based on the evaluation results, the difference information between the reference audio and the target audio is determined, and the evaluation result of the target audio is determined based on the difference information.
[0009] Secondly, this application also provides an audio data evaluation device. The device includes:
[0010] The first acquisition module is used to acquire the original audio, which includes reference audio and target audio to be evaluated.
[0011] The first processing module is used to perform human ear perception processing on the original audio to obtain human ear perception domain audio, wherein the human ear perception domain audio includes reference human ear perception domain audio corresponding to the reference audio and target human ear perception domain audio corresponding to the target audio.
[0012] The first evaluation module is used to evaluate the original audio and the audio in the human ear perception domain using an evaluation network;
[0013] The first determining module is used to determine the difference information between the reference audio and the target audio based on the evaluation processing results, and to determine the evaluation result of the target audio based on the difference information.
[0014] Thirdly, this application also provides an electronic device. The electronic device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the steps of the first aspect described above.
[0015] Fourthly, this application also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program thereon, which, when executed by a processor, implements the steps of the first aspect described above.
[0016] Fifthly, this application also provides a computer program product. The computer program product includes a computer program that, when executed by a processor, implements the steps of the first aspect described above.
[0017] The aforementioned audio data evaluation method, apparatus, electronic device, storage medium, and program product, after acquiring the target audio and reference audio, obtains reference and target human hearing domain audio and target human hearing domain audio by performing human hearing perception processing on the reference and target audio respectively. An evaluation network is then used to evaluate the original audio and the human hearing domain audio. Because the characteristics of the human hearing domain are comprehensively considered in both the reference and target human hearing domain audio, the difference information between the reference and target audio can be more accurately determined based on the evaluation processing results. This allows for a more accurate evaluation of the target audio based on this difference information, and also results in a higher correlation between the evaluation results and subjective human hearing evaluation. Attached Figure Description
[0018] Figure 1 This is a diagram illustrating the application environment of an audio data evaluation method in one embodiment.
[0019] Figure 2 This is a flowchart illustrating an audio data evaluation method in one embodiment;
[0020] Figure 3 This is another flowchart illustrating the audio data evaluation method in one embodiment;
[0021] Figure 4 This is another flowchart illustrating the audio data evaluation method in one embodiment;
[0022] Figure 5 This is another flowchart illustrating the audio data evaluation method in one embodiment;
[0023] Figure 6 This is another flowchart illustrating the audio data evaluation method in one embodiment;
[0024] Figure 7 This is a schematic diagram illustrating the implementation process of human ear perception processing in one embodiment;
[0025] Figure 8 This is a schematic diagram illustrating the implementation process of time-frequency transformation of audio data in one embodiment;
[0026] Figure 9 This is a schematic diagram of the implementation framework of an audio data evaluation method in one embodiment;
[0027] Figure 10 This is a schematic diagram of the composition structure of the evaluation network in one embodiment;
[0028] Figure 11 This is a schematic diagram of the implementation architecture of the evaluation network in one embodiment;
[0029] Figure 12 This is a schematic diagram of another implementation framework of the audio data evaluation method in one embodiment;
[0030] Figure 13 This is a schematic diagram of a framework for determining the evaluation results of audio data in one embodiment;
[0031] Figure 14 This is a structural block diagram of an audio data evaluation device in one embodiment;
[0032] Figure 15 This is a diagram of the internal structure of an electronic device in one embodiment. Detailed Implementation
[0033] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0034] In one embodiment, spatial audio objective evaluation metrics are determined through a variety of methods:
[0035] Method 1: Assessment and Prediction of Binaural Aspects of Audio Quality (BAMQ) Based on Binaural Cueing. BAMQ takes a pair of matched spatial audio samples as input, processes them in the front end, and then inputs them into the back end for objective quality assessment. The BAMQ front end includes two stages: outer ear sound processing and binaural feature extraction. In the outer ear sound processing part, the BAMQ algorithm uses four sets of filters to simulate the sound processing process of the human outer ear, basilar membrane, cochlea, and cochlear hairs, mapping spatial audio to the human ear's perceptual domain. In the binaural feature extraction stage, BAMQ performs fine-structure filtering and modulation filtering on the audio mapped to the human ear's perceptual domain. Then, it calculates the interaural transfer function (ITF) and extracts four binaural features: interaural phase difference (IPD), interaural time difference (ITD), interaural vector strength (IVS), and interaural level difference (ILD). Finally, the extracted binaural features are input to the backend. In the backend, BAMQ first calculates corresponding sub-indices based on the binaural features extracted from the frontend. These sub-indices are then input into a series of weights calculated using the Multivariate Adaptive Regression Splines (MARS) method to calculate the final spatial audio objective quality score.
[0036] Method 2: A Deep Perceptual Spatial Audio Localization Metric (DPLM): DPLM first trains a deep neural network to predict the direction of arrival (DOA) of both ears at the frame level. Then, it uses the depth feature distance of this neural network to evaluate the spatial localization of spatial audio. Specifically, the input to DPLM is two binaural spatial audio files, reference audio x1 and target audio x2. After performing a Short Time Fourier Transform (STFT), the amplitude and phase spectra of both ears are extracted, stacked, and then input into the neural network. The neural network in DPLM consists of three layers: a feature extraction layer based on an Inception network, a temporal aggregation layer based on a Long Short Time Memory (LSTM) network, and a fully connected fifty-class localization head. After the amplitude / phase spectra of the stacked spatial audio files are input into the network, the network calculates the corresponding embedding layers. Finally, the DPLM algorithm uses the depth feature distance between the embedding layers of reference audio x1 and target audio x2 as the spatial audio localization score.
[0037] Method 3: Spatial Audio Quality Assessment Metric (SAQAM) Based on Deep Learning: SAQAM is based on a multi-task learning framework, simultaneously considering two quality attributes of spatial audio: listening quality (LQ) and spatial quality (SQ). A deep neural network model is trained to perform objective quality assessment of spatial audio. For spatial quality, SAQAM uses a similar approach to DPLM, implicitly learning the spatial quality of spatial audio in the embedding layer by predicting frame-level binaural arrival directions. For listening quality, SAQAM utilizes contrastive learning and triplet training methods. By constructing triplet training samples of <anchor sample, positive sample, negative sample>, and ensuring that |anchor sample - positive sample| <|anchor sample - negative sample|, the network implicitly learns the listening quality of spatial audio through its embedded layer. The SAQAM network structure consists of a feature extraction layer, a temporal aggregation layer, and two multi-task heads (auditory quality and spatial quality). The feature extraction layer employs an Inception structure similar to DPLM, while the temporal aggregation block uses a temporal convolutional network based on causal convolution and dilated convolution. The two multi-task heads are implemented using one-dimensional convolution, weight normalization, and fully connected layers, respectively. For the input reference audio x1 and target audio x2, SAQAM first performs a short-time Fourier transform on the input audio, then stacks the extracted left and right ear amplitude and phase spectra before inputting them into the network for calculation, outputting embedding layers for the feature extraction layer, temporal aggregation layer, auditory quality task head, and spatial quality task head. Finally, the network calculates the Overall Quality (OVRL) score based on the shared depth feature distance of the feature extraction layer and temporal aggregation layer, calculates the auditory quality score based on the feature extraction layer, temporal aggregation layer, and auditory quality task head, and calculates the spatial quality score based on the same features. In SAQAM, the auditory quality of audio data is an indicator that measures the degree of audio distortion, an objective attribute, rather than a subjective attribute such as the timbre of the audio. Auditory quality is unrelated to human subjective perception.
[0038] Method 1 considers human perception of spatial audio, mapping the audio to the human auditory domain to better align with subjective human perception. It then uses binaural cues and multiple linear regression to predict spatial audio quality. However, the BAMQ algorithm has stringent input requirements, demanding temporal alignment and equal length between the reference and target audio, making it difficult to apply in real-world scenarios in certain situations.
[0039] Method 2 introduces the concept of depth feature distance. By predicting the direction of arrival of spatial audio and utilizing the information implicit in the network embedding layer, it completes the prediction of spatial audio localization deviation, which can reflect the spatial quality of spatial audio. However, in addition to spatial quality, the auditory quality of spatial audio (such as changes in fidelity) also affects the overall quality of spatial audio. Therefore, DPLM's objective evaluation of spatial audio is flawed and incomplete.
[0040] Method 3 addresses the shortcomings of DPLM by considering two quality attributes of spatial audio: auditory quality and spatial quality. Through multi-task learning and triplet training, the algorithm's neural network can simultaneously learn both auditory and spatial quality attributes, enabling a more objective assessment of the overall quality of spatial audio. However, SAQAM trains its neural network solely on a series of objective quality attributes, neglecting the subjective characteristics of the human ear, thus failing to guarantee the correlation between the assessment results and subjective human evaluations.
[0041] Of the three implementations described above, using only the raw spectrogram ignores the complex psychoacoustic phonemes that affect human auditory perception. While capturing the objective level of distortion, it fails to consider how the auditory system responds to different frequencies and masks certain sounds. This can lead to an incomplete understanding of the audio, resulting in poor subjective evaluations of sound distortion. Conversely, downsampled spectrograms that rely solely on the physiological limitations of human hearing ignore the objective characteristics of the original audio signal. While capturing psychoacoustic properties, they do not cover the entire frequency range and details of the source signal. This may hinder the model's ability to accurately capture various types of audio distortion.
[0042] Based on this, this application provides an audio data evaluation method. To obtain a comprehensive and accurate audio representation, it combines human-perceived audio data and raw audio data, enabling the evaluation network to fully utilize the raw spectral content while also considering the complexity of human auditory perception, thus improving the performance of subjective evaluation. This audio data evaluation method performs human-perceived processing on the acquired reference audio and target audio separately, obtaining reference human-perceived domain audio and target human-perceived domain audio. An evaluation network is then used to evaluate the raw audio and the human-perceived domain audio. Thus, because the characteristics of the human-perceived domain are comprehensively considered in the reference and target human-perceived domain audio, the difference information between the reference audio and the target audio can be more accurately determined based on the evaluation results. This allows for a more accurate evaluation of the target audio based on this difference information, and also results in a higher correlation between the evaluation results and subjective human auditory evaluation.
[0043] The audio data evaluation method provided in this application embodiment can be applied to, for example... Figure 1In the application environment shown, after acquiring the reference audio 102 and the target audio 103 to be evaluated, the audio data evaluation terminal 101 performs human ear perception processing on the reference audio and the target audio respectively to obtain the reference human ear perception domain audio and the target human ear perception domain audio. An evaluation network is then used to evaluate the original audio and the human ear perception domain audio. Because the characteristics of the human ear perception domain are comprehensively considered in the reference and target human ear perception domain audio, the difference information between the reference audio and the target audio can be more accurately determined based on the evaluation processing results. This allows for a more accurate evaluation of the target audio based on this difference information, and also results in a higher correlation between the evaluation results and subjective human ear evaluation. The audio data evaluation terminal 101 can be, but is not limited to, various personal computers, laptops, smartphones, tablets, IoT devices, and portable wearable devices. IoT devices can include smart speakers, smart TVs, smart air conditioners, smart in-vehicle devices, etc. Portable wearable devices can include smartwatches, smart bracelets, head-mounted devices, etc.
[0044] In one embodiment, such as Figure 2 As shown, an audio data evaluation method is provided, which is then applied to... Figure 1 Taking the audio data evaluation terminal in the example, the following steps are included:
[0045] Step 201: Obtain the original audio.
[0046] The original audio includes a reference audio and a target audio to be evaluated. The reference and target audio can be time-aligned and of equal length, or they can be time-unaligned and of unequal length. The audio content of the reference and target audio can be the same or different.
[0047] For example, the reference audio can be audio data of good quality that can be used as a benchmark, such as undamaged audio data. The target audio can be any type of audio data, such as audio data including noise, compressed audio data, or damaged audio data. The target audio can be acquired by an audio data evaluation terminal or by receiving target audio sent from other terminals.
[0048] Step 202: Perform human ear perception processing on the original audio to obtain human ear perception domain audio.
[0049] The audio in the human ear perception domain includes a reference human ear perception domain audio corresponding to the reference audio and a target human ear perception domain audio corresponding to the target audio. Human ear perception processing involves mapping the reference audio and target audio from the original time domain to the human ear perception domain to perform human ear simulation on the reference audio and the target audio, thereby obtaining the reference human ear perception domain audio and the target human ear perception domain audio.
[0050] In some possible implementations, multiple filters are used to filter both the reference and target audio to simulate the human ear's sound perception and processing, thereby transferring the original time-domain audio data to the human ear's perception domain. Taking the processing of the reference audio as an example, first, a bandpass filter is used to filter the reference audio to simulate the frequency response and sound perception of the human outer and middle ear. Then, a higher-order filter is used to filter the filtered reference audio to simulate the sound perception of the basilar membrane. Next, the higher-order filtered reference audio is instantaneously compressed to simulate the cochlear compression process of sound in the human cochlea. Finally, a low-pass filter is used to filter the instantaneously compressed audio data to simulate the mechatronic transmission process of the cochlear hair cells, obtaining the audio data of the reference audio in the human ear's perception domain, i.e., the reference human ear perception domain audio. Similarly, the same steps can be used to convert the target audio to the human ear's perception domain to obtain the target human ear perception domain audio.
[0051] Step 203: Use an evaluation network to evaluate the original audio and the audio in the human ear perception domain.
[0052] The evaluation network is a trained network used to evaluate the audio quality of the target audio. The evaluation network is used to process both the original audio and the audio in the human auditory perception domain; that is, the original audio is processed by the evaluation network to obtain the evaluation result of the original audio, and the audio in the human auditory perception domain is processed by the same evaluation network to obtain the evaluation result of the audio in the human auditory perception domain.
[0053] For example, the evaluation network may include an objective distortion quality evaluation subnetwork and a spatial quality evaluation subnetwork. The objective distortion quality evaluation subnetwork evaluates the original audio, while the spatial quality evaluation network evaluates the audio in the human auditory perception domain, thus enabling the evaluation network to process both the original audio and the audio in the human auditory perception domain. In this embodiment, considering the loss caused by dimensionality reduction in perceptual audio processing, to enhance the correlation between the evaluation results output by the evaluation network and human subjective perception, the evaluation network is trained by using human auditory perceptual audio data to assist the original audio, thereby achieving audio data evaluation. This approach not only obtains more accurate evaluation results but also addresses the problem of audio loss. Although the auditory quality of audio data in SAQAM is an indicator of the degree of objective audio distortion, the degree of objective audio distortion is not directly related to human auditory perception. The way the human ear perceives objective audio distortion is affected by the complexity of auditory cognition and the physiological limitations of human hearing. Therefore, in this embodiment, considering the overall interaction between human psychoacoustic factors and objective distortion is crucial when evaluating the impact of objective audio distortion on human auditory perception.
[0054] Step 204: Determine the difference information between the reference audio and the target audio based on the evaluation processing results, and determine the evaluation result of the target audio based on the difference information.
[0055] The evaluation results can include the evaluation results of the original audio and the evaluation results of the audio in the human hearing perception domain. The evaluation results of the original audio represent the depth features of the original audio in objective distortion quality evaluation, while the evaluation results of the audio in the human hearing perception domain represent the depth features of the audio in the human hearing perception domain and the original audio in spatial quality. This difference information is obtained by determining the distance between the depth features corresponding to different audio data.
[0056] The difference information is used to reflect the differences between the target audio and a higher-quality reference audio in terms of spatial characteristics, audio fidelity, and overall quality. For example, the difference information represents the spatial feature distance, objective feature distance, and overall quality feature distance between the reference audio and the target audio.
[0057] In some feasible implementations, the original audio and the audio in the human ear's perception domain are converted through time-frequency transformation to obtain first and second spectral information. The first and second spectral information are then separately input into an evaluation network, which predicts the feature set of the reference audio and the feature set of the target audio. For the target audio, since the second spectral information includes both the original time-domain and human ear-perception domain spectral information, the evaluation network first performs feature extraction and time-domain aggregation on both spectral information. Then, using the spatial quality head and objective distortion quality evaluation head within the evaluation network, it determines the spatial quality recognition features and objective distortion quality evaluation recognition features.
[0058] Similarly, the evaluation network processes the reference audio, performing feature extraction and temporal aggregation. Then, using the spatial quality head and objective distortion quality evaluation head within the network, spatial quality recognition features and objective distortion quality evaluation recognition features are determined. Finally, by using the features of the reference audio output by the evaluation network and the features of the target audio output by the same network, the depth feature distance between the reference audio and the target audio is determined, thus obtaining the difference information.
[0059] The evaluation results of the target audio include: spatial quality, objective distortion quality assessment, and overall quality. Here, objective distortion quality assessment is an objective attribute that measures audio fidelity and is independent of subjective human perception. Because the description of audio quality in SAQAM can be interpreted as a measure of subjective attributes such as timbre, this embodiment uses objective distortion quality assessment to describe this objective attribute of audio fidelity, making the expression of the evaluation results more direct and accurate.
[0060] In some possible implementations, the evaluation result of the target audio can be obtained by converting the difference information into a quality score. For example, converting the feature distance in the objective distortion quality assessment of the difference information into a quality score yields the objective distortion quality assessment score of the target audio; converting the feature distance in the spatial quality assessment of the difference information into a quality score yields the spatial quality score of the target audio; and converting the comprehensive feature distance in the difference information into a quality score yields the overall quality score of the target audio. Then, combining the objective distortion quality assessment score, the spatial quality score, and the overall quality score allows for a more comprehensive and accurate evaluation of the target audio.
[0061] In this embodiment, after acquiring the target audio and reference audio, human hearing perception processing is applied to both the reference and target audio to obtain reference and target human hearing perception domain audio, respectively. Since the characteristics of the human hearing perception domain are comprehensively considered in both the reference and target audio, using an evaluation network to evaluate the original audio and the audio within the human hearing perception domain makes the evaluation results more consistent with human subjective perception. Subsequently, based on the evaluation results, the difference information between the reference and target audio can be determined more accurately, allowing for a more accurate evaluation of the target audio based on this difference information, and resulting in a higher correlation between the evaluation results and subjective human hearing assessments.
[0062] In one embodiment, the evaluation network includes an objective distortion quality evaluation subnetwork and a spatial quality evaluation subnetwork, and step 203 above can be achieved through... Figure 3 The steps shown are to be implemented as follows:
[0063] Step 301: The original audio is objectively evaluated at least by the objective distortion quality evaluation sub-network to obtain multiple reference objective evaluation features corresponding to the reference audio and multiple target objective evaluation features corresponding to the target audio.
[0064] Specifically, after performing time-domain transformation on the reference and target audio in the original audio, the amplitude and phase spectra obtained from the time-domain transformation are input into the objective distortion quality evaluation subnetwork to obtain multiple reference objective evaluation features and multiple target objective evaluation features. Alternatively, the amplitude and phase spectra corresponding to the original audio, as well as the amplitude and phase spectra of the audio in the human auditory perception domain, are combined and input into the objective distortion quality evaluation subnetwork to obtain multiple reference objective evaluation features and multiple target objective evaluation features.
[0065] In some possible implementations, the objective distortion quality assessment subnetwork includes a cascaded first feature extraction layer, a first temporal aggregation layer, and an objective distortion quality assessment head. Step 301 described above can be implemented through the following steps 311 to 315 (not shown in the figure):
[0066] Step 311: The original audio is processed by the first feature extraction layer to obtain a first spectral feature. The first spectral feature includes a first reference spectral feature corresponding to the reference audio and a first target spectral feature corresponding to the target audio.
[0067] The process involves performing time-frequency conversion on the original audio to obtain spectral information, which is then input into the objective distortion quality evaluation subnetwork. A first feature extraction layer is used to extract features from the spectral information of the input original audio to obtain the first spectral feature. For example, this feature extraction layer includes six stacked feature extraction blocks. After the sixth feature extraction block, a 1*1 convolution is used to compress the channel dimension and the frequency dimension. The input dimension of the feature extraction block is [B, 4, f, 257], and the dimension of the output first spectral feature is [B, f, 64]. Thus, by extracting features from the input spectral information, the original time-domain spectral features can be extracted more accurately and fully.
[0068] Step 312: The first time-domain aggregation layer is used to perform time-domain aggregation processing on the first spectral features to obtain the first aggregated features.
[0069] The first aggregated feature includes a first reference aggregated feature corresponding to the reference audio and a first target aggregated feature corresponding to the target audio. The first spectral feature output from the first feature extraction layer is input into the first temporal aggregation layer to obtain the aggregated feature output by the temporal aggregation layer.
[0070] In some possible implementations, a first time-domain aggregation layer is used to aggregate the first reference spectral features and the first target spectral features in the original time domain to obtain a first aggregated feature. In the first time-domain aggregation layer, the first spectral features are first subjected to an ascending-channel convolution, and then a descending-channel convolution to obtain the first aggregated feature.
[0071] Step 313: The objective distortion quality assessment head is used to perform objective distortion quality assessment and recognition processing on the first aggregated feature to obtain objective distortion quality assessment and recognition features. The objective distortion quality assessment and recognition features include reference objective distortion quality assessment and recognition features corresponding to the reference audio and target objective distortion quality assessment and recognition features corresponding to the target audio.
[0072] Specifically, the first aggregated feature is input into the objective distortion quality assessment head to predict the objective distortion quality assessment of the target audio from the original first aggregated feature, thus obtaining the objective distortion quality assessment recognition feature.
[0073] The objective distortion quality assessment head can include two stacked 1*3 convolutional kernels and weight normalization layers, as well as two fully connected layers. After the first aggregated feature is input into the objective distortion quality assessment head, it is processed sequentially through the two stacked 1*3 convolutional kernels and weight normalization layers, as well as the two fully connected layers, to obtain the objective distortion quality assessment recognition feature. In this way, the objective distortion quality assessment recognition feature is independent of human auditory perception, so it can purely represent the objective distortion quality assessment of the original temporal domain target audio.
[0074] Step 314: The first reference spectral feature, the first reference aggregation feature, and the reference objective distortion quality assessment identification feature are used as the plurality of reference objective assessment features.
[0075] Step 315: Use the first target spectral feature, the first target aggregation feature, and the target objective distortion quality assessment and identification feature as the multiple target objective assessment features.
[0076] In this embodiment of the application, the original audio is evaluated by the first feature extraction layer, the first temporal aggregation layer, and the objective distortion quality evaluation head cascaded in the objective distortion quality evaluation subnetwork. Since the input features are the features of the original temporal reference audio and the target audio, which are unrelated to human hearing perception, the output of the objective distortion quality evaluation subnetwork can intuitively represent the objective distortion quality evaluation of the target audio in the original temporal domain.
[0077] Step 302: The spatial quality evaluation sub-network is used to perform spatial evaluation processing on the human ear perception domain audio, and the evaluation features obtained from the objective evaluation processing are fused during the spatial evaluation processing to obtain multiple reference spatial evaluation features corresponding to the reference audio and multiple target spatial evaluation features corresponding to the target audio.
[0078] Specifically, after performing time-domain transformation on the reference audio and target audio in the human ear perception domain, the amplitude spectrum and phase spectrum obtained after time-domain transformation are input into the objective distortion quality evaluation sub-network to obtain multiple reference objective evaluation features and multiple target objective evaluation features.
[0079] In some possible implementations, the spatial quality evaluation subnetwork includes a cascaded second feature extraction layer, a second temporal aggregation layer, and a spatial quality head. Step 302 described above can be implemented through the following steps 321 to 326 (not shown in the figure):
[0080] Step 321: The second feature extraction layer is used to perform feature extraction processing on the audio in the human ear perception domain to obtain a second spectral feature. The second spectral feature includes a second reference spectral feature corresponding to the reference audio and a second target spectral feature corresponding to the target audio.
[0081] The second feature extraction layer has the same structure and shares parameters as the first feature extraction layer. After time-frequency conversion of the audio in the human ear's perceptual domain, the spectral information is obtained and input into the spatial quality evaluation subnetwork. The second feature extraction layer is used to extract features from the spectral information of the input audio in the human ear's perceptual domain to obtain the second spectral features.
[0082] Step 322: The second time-domain aggregation layer is used to perform time-domain aggregation processing on the second spectral features to obtain the second aggregated features.
[0083] The second aggregated feature includes a second reference aggregated feature corresponding to the reference audio and a second target aggregated feature corresponding to the target audio. The second spectral feature output from the second feature extraction layer is input into the second temporal aggregation layer to obtain the second aggregated feature output by the temporal aggregation layer.
[0084] In some possible implementations, the second aggregation layer has the same structure and shares parameters with the first aggregation layer. The second temporal aggregation layer aggregates the second reference spectral features and the second target spectral features to obtain the second aggregated features. For example, in the second temporal aggregation layer, the second spectral features are first subjected to an ascending-channel convolution, and then a descending-channel convolution to obtain the second aggregated features. This can be achieved through the following process:
[0085] In the temporal aggregation layer of the evaluation network, the original temporal features are first subjected to an ascending-channel convolution, and then a descending-channel convolution to obtain aggregated features. That is, step 322 above can be achieved through the following process:
[0086] First, in the second temporal aggregation layer, the second spectral feature is subjected to at least one level of convolution processing based on multiple preset channel numbers that increase progressively to obtain the first convolution feature.
[0087] The second temporal aggregation layer includes two weight normalizations, two 1*3 convolutional kernels, and a residual module. After inputting the original temporal spectral features and the spectral features of the human ear perception domain into the temporal aggregation layer, the number of channels increases progressively in the four-layer temporal aggregation layer, with numbers of 64, 128, and 256. Thus, in the second temporal aggregation layer, the second spectral features are convolved according to the channel number increasing from 64 to 128 and then to 256 to obtain the first convolutional features.
[0088] Secondly, based on the progressively decreasing number of preset channels, the first convolutional feature is subjected to at least one level of convolutional processing to obtain the second convolutional feature;
[0089] The number of channels decreases progressively: 256, 128, and 64. The first convolutional feature is convolved with the number of channels decreasing from 256 to 128 and 64, and the second convolutional feature is output. In this way, the receptive field of the second convolutional feature can be expanded through dilated convolution.
[0090] Finally, residual processing is performed on the second convolutional feature to obtain the second aggregated feature.
[0091] Specifically, the second convolutional features are processed using a reference module in the temporal aggregation layer. For example, a residual function or residual network is used to process the second convolutional features to obtain the output aggregated features. Thus, by first performing convolutions on the spectral features with progressively increasing preset channel numbers, and then performing convolutions with progressively decreasing preset channel numbers, the receptive field of the second convolutional features is increased through dilated convolution, thereby increasing the receptive field of the aggregated features and making the aggregated features richer.
[0092] Step 323: Perform a weighted fusion process on the second aggregated feature and the first aggregated feature to obtain a fused feature.
[0093] The fusion features include reference fusion features corresponding to the reference audio and target fusion features corresponding to the target audio. The second aggregated feature and the first aggregated feature are weighted using preset weights and then summed to obtain the fusion features.
[0094] The preset weight includes two different weight values, and the sum of these two different weight values is 1.
[0095] Step 324: The spatial quality head is used to perform spatial quality recognition processing on the fused features to obtain spatial quality recognition features. The spatial quality recognition features include reference spatial quality recognition features corresponding to the reference audio and target spatial quality recognition features corresponding to the target audio.
[0096] The spatial quality identification features are multi-dimensional matrices. The fused features are input into the spatial quality head, which performs convolution and weight normalization operations on the input fused features to output the spatial quality identification features.
[0097] In some possible implementations, in the spatial quality head, spatial quality recognition features for evaluating the spatial quality of the target audio are obtained by performing azimuth classification on the input fused features. That is, step 324 above can be implemented through the following process:
[0098] First, the fused features are preprocessed to obtain preprocessed features.
[0099] The spatial quality head comprises two stacked 1x3 convolutional layers with weight normalization, two fully connected layers, and a soft maximum activation function. The input features are preprocessed by performing convolution and weight normalization on the two stacked 1x3 convolutional layers, resulting in preprocessed features. This convolution and weight normalization of the input features improves the network's training speed and thus enhances prediction accuracy.
[0100] Then, the azimuth angles corresponding to the preprocessed features are classified by angle to obtain the probability distribution of the azimuth angles in the preset angle interval set.
[0101] In this spatial quality head, the interval [0°, 360°] is first divided into multiple equidistant sectors, thus obtaining a set of preset angle intervals. The azimuth angle corresponding to the preprocessed feature is the azimuth angle of the audio frame to which the preprocessed feature belongs, including: the horizontal angle and elevation angle of the audio frame in the polar coordinate system. This polar coordinate system has the human ear as the origin and faces the sound source. The polar coordinates of the vector represented by the preprocessed feature in this polar coordinate system, namely the horizontal angle and elevation angle, are the azimuth angle corresponding to the preprocessed feature.
[0102] For example, the azimuth angles corresponding to the preprocessed features are classified by two fully connected layers and a soft maximum activation function in the spatial quality head to obtain the probability distribution of the azimuth angle in a preset angle range set.
[0103] Finally, the probability distributions corresponding to the preset angle interval set are fused to obtain the spatial quality identification features.
[0104] In this process, after obtaining the probability distribution of the azimuth angle of the preprocessed features in each preset angle interval, these probability distributions are concatenated to obtain a spatial quality recognition feature that includes multiple probability distributions. Thus, by predicting the azimuth angle of the audio frame containing the obtained preprocessed features, the azimuth angle is classified to obtain a probability that reflects the azimuth angle in each preset angle interval. By combining the probability distributions corresponding to each preset angle interval, a spatial quality recognition feature is formed, ensuring that the value of each latitude in this feature reflects the probability distribution of the azimuth angle of the target audio within the preset angle interval.
[0105] Step 325: Use the second reference spectral feature, the second reference aggregation feature, and the reference space quality identification feature as the plurality of reference space evaluation features.
[0106] Step 326: Use the second target spectral feature, the second target aggregation feature, and the target spatial quality identification feature as the plurality of target spatial evaluation features.
[0107] In this embodiment of the application, the spatial quality head is used to learn the weighted and combined fusion features in order to evaluate the spatial quality of the target audio. In this way, the human ear perception domain features and the original temporal domain features are incorporated into the process of determining the spatial quality recognition features, so that the spatial quality recognition features can better match the characteristics of the human ear perception domain.
[0108] In one embodiment, the difference information includes at least one of: overall feature distance, spatial feature distance, and objective feature distance, and step 204 above can be achieved through... Figure 4 The steps shown are to be implemented as follows:
[0109] Step 401: Determine the overall feature distance between the reference audio and the target audio based on the first spectral feature, the second spectral feature, the first aggregated feature, and the second aggregated feature.
[0110] Specifically, the distances between the first reference spectral feature of the reference audio and the first target spectral feature of the target audio, the distances between the second reference spectral feature of the reference human ear perception domain audio and the second target spectral feature of the target human ear perception domain audio, the distances between the first reference aggregation feature of the reference audio and the first target aggregation feature of the target audio, and the distances between the second reference aggregation feature of the reference human ear perception domain audio and the second target aggregation feature of the target human ear perception domain audio are determined. These distances are then summed, and the summation results are averaged to obtain the overall feature distance.
[0111] Step 402: Determine the objective feature distance between the reference audio and the target audio based on the first spectral feature, the first aggregation feature, and the objective distortion quality assessment and identification feature.
[0112] In the same way as step 401, the distances between the first spectral feature, the first aggregate feature, and the objective distortion quality assessment and identification feature are summed and then averaged to obtain the objective feature distance.
[0113] Step 403: Determine the spatial feature distance between the reference audio and the target audio based on the second spectral feature, the second aggregation feature, and the spatial quality identification feature.
[0114] In the same manner as step 401, the distances between the second spectral feature, the second aggregated feature, and the spatial quality identification feature are summed and then averaged to obtain the objective feature distance. Thus, by determining the overall feature distance between the reference audio and the target audio, the objective feature distance in objective distortion quality evaluation, and the spatial feature distance in spatial quality, these feature distances are combined as difference information between the reference audio and the target audio. Furthermore, the spatial feature distance considers the characteristics of the human auditory perception domain. Therefore, this difference information can more fully and accurately reflect the difference between the reference audio and the target audio, and also ensure that the final evaluation result is correlated with human auditory perception features.
[0115] After determining the difference information, the audio quality of the target audio relative to the reference audio is obtained by using multiple feature distances within the difference information, thus yielding the evaluation result of the target audio. This can be achieved through the following process:
[0116] First, based on the overall feature distance in the difference information, the overall quality of the target audio relative to the reference audio is determined.
[0117] In some possible implementations, a mapping table between feature distance and quality score is first obtained, and then the quality score corresponding to the comprehensive feature distance is determined according to the mapping table, thus obtaining the overall quality.
[0118] Secondly, based on the spatial feature distance, the spatial quality of the target audio relative to the reference audio is determined.
[0119] In some possible implementations, the spatial quality is obtained by determining the corresponding quality score in a mapping table based on the spatial feature distance. Since audio features from the human ear's perceptual domain are incorporated when determining the spatial feature distance, the evaluation of the spatial quality of the target audio can better align with the subjective perception characteristics of the human ear.
[0120] Next, based on the objective feature distance, an objective distortion quality assessment of the target audio relative to the reference audio is determined.
[0121] In some possible implementations, the corresponding quality score is determined in the mapping table according to the distance of the objective feature, thus obtaining the objective distortion quality assessment.
[0122] Finally, the overall quality, spatial quality, and objective distortion quality assessments are combined to obtain the assessment results.
[0123] In some possible implementations, the overall quality, spatial quality, and objective distortion quality assessments are concatenated to form an array, thus obtaining the assessment result for the target audio. This assessment result can reflect the quality of the target audio from multiple aspects, including overall quality, spatial quality, and objective distortion quality. Therefore, by combining the overall quality, spatial quality, and objective distortion quality assessments of the target audio relative to the reference audio, the assessment result can accurately reflect the quality of the target audio and also improve its correlation with the human auditory perception domain.
[0124] In one embodiment, a time-frequency transformation is performed on the reference audio and the reference human ear perception domain audio to obtain spectral information that can be input into the evaluation network. This process is as follows: Figure 5 As shown, this can be achieved through the following steps:
[0125] Step 501: Perform time-frequency transformation on the reference audio and the reference human ear perception domain audio to obtain the first spectrum information.
[0126] Specifically, the reference audio and the reference human ear perception domain audio are converted from the time domain to the frequency domain to obtain the first spectrum information, which includes the amplitude spectrum and phase spectrum of the reference audio, as well as the amplitude spectrum and phase spectrum of the reference human ear perception domain audio.
[0127] In some possible implementations, a short-time Fourier transform is performed on the reference audio and the reference human ear perception domain audio to obtain a time spectrum. The amplitude spectrum and phase spectrum are then extracted from the obtained time spectrum to obtain the first spectral information. In this way, during the time-frequency transformation process, the original time-domain audio data and the human ear perception domain audio data are combined, ensuring that the obtained spectral information fully considers the human ear perception domain audio data. This facilitates subsequent evaluation of the target audio by combining the human ear perception domain audio data with the original time-domain audio data.
[0128] For example, firstly, the reference audio and the reference human hearing perception domain audio are respectively transformed into the complex domain to obtain the time spectrum of the reference audio in the complex domain and the time spectrum of the reference human hearing perception domain audio in the complex domain; then, the amplitude spectrum and phase spectrum of the time spectrum are extracted; finally, the amplitude spectrum and phase spectrum corresponding to the reference audio, and the amplitude spectrum and phase spectrum corresponding to the reference human hearing perception domain audio are stacked to obtain the first spectrum information. This ensures that the first spectrum information fully incorporates the phase spectrum and amplitude spectrum of the human hearing perception domain, which is beneficial for subsequent audio evaluation processes.
[0129] Step 502: Perform time-frequency transformation on the target audio and the target human ear perception domain audio respectively to obtain the second spectrum information.
[0130] Specifically, the target audio and the target human ear perception domain audio are transformed from the time domain to the frequency domain to obtain the second spectrum information, which includes the amplitude spectrum and phase spectrum of the target audio and the amplitude spectrum and phase spectrum of the target human ear perception domain audio.
[0131] In some possible implementations, a short-time Fourier transform is performed on the target audio and the target human ear perception domain audio to obtain the time spectrum, and then the amplitude spectrum and phase spectrum are extracted from the obtained time spectrum to obtain the second spectral information.
[0132] By performing a complex domain transformation on the target audio and the target human ear perception domain audio, and stacking the extracted amplitude spectrum and phase spectrum, a second spectral information containing the human ear perception domain audio can be obtained; that is, step 502 above can be implemented through steps 521 to 523 (not shown in the figure):
[0133] Step 521: Perform complex domain transformation on the target audio and the target human ear perception domain audio to obtain the time spectrum of the target audio in the complex domain and the time spectrum of the target human ear perception domain audio in the complex domain.
[0134] Both the target audio and the target human ear perception domain audio are stored with left and right channel audio. The time-frequency spectrum is the spectrum of the audio data in the complex domain. The time-frequency spectrum of the target audio in the complex domain includes: the time-frequency spectrum of the left channel and the time-frequency spectrum of the right channel in the complex domain. The time-frequency spectrum of the target human ear perception domain audio in the complex domain includes: the time-frequency spectrum of the left channel and the time-frequency spectrum of the right channel in the complex domain.
[0135] In some possible implementations, a complex domain conversion between the target audio and the target human auditory perception domain audio is achieved by performing a short-time Fourier transform on the left and right channels of the target audio and the left and right channels of the target human auditory perception domain audio. For example, a short-time Fourier transform of the audio data can be performed using a length of 512 points and a 50% (256 points) jump length, and using a Hamming window of length 512 points.
[0136] Step 522: Extract the amplitude spectrum and phase spectrum of the time spectrum.
[0137] Specifically, the amplitude spectrum and phase spectrum are extracted from the time-frequency spectra of the complex domain of the left and right channels, respectively, to obtain the amplitude spectrum of the left channel of the target audio, the amplitude spectrum of the right channel of the target audio, the phase spectrum of the left channel of the target audio, the phase spectrum of the right channel of the target audio, the amplitude spectrum of the left channel of the audio to be evaluated in the human ear perception domain, the amplitude spectrum of the right channel of the audio in the target human ear perception domain, the phase spectrum of the left channel of the audio in the target human ear perception domain, and the phase spectrum of the right channel of the audio in the target human ear perception domain.
[0138] Step 523: Stack the amplitude spectrum and phase spectrum corresponding to the target audio, and the amplitude spectrum and phase spectrum corresponding to the target human ear perception domain audio to obtain the second spectrum information.
[0139] Specifically, the amplitude spectrum of the target audio's left channel, amplitude spectrum of the target audio's right channel, phase spectrum of the target audio's left channel, and phase spectrum of the target audio's right channel, along with the amplitude spectrum of the target audio's left channel, amplitude spectrum of the target audio's right channel, phase spectrum of the target audio's left channel, and phase spectrum of the target audio's right channel, are stacked along the channel dimension to obtain a matrix composed of amplitude and phase spectra, which is the second spectral information. Thus, by performing complex-domain transformations on the original time-domain target audio and the target audio's human-perceived domain audio respectively, their respective time spectra can be obtained. The phase and amplitude spectra extracted from these time spectra are then stacked to obtain the second spectral information, which reflects both the phase and amplitude spectra of the original time domain and the phase and amplitude spectra of the human-perceived domain. This second spectral information fully incorporates the phase and amplitude spectra of the human-perceived domain, which is beneficial for incorporating the human-perceived domain into subsequent audio evaluations, thereby improving the correlation between the evaluation results and the human ear's primary evaluation criteria.
[0140] Step 503: Input the first spectrum information and the second spectrum information into the evaluation network to evaluate the original audio and the human ear perception domain audio using the evaluation network.
[0141] After obtaining the first and second spectral information, the evaluation network processes the first and second spectral information to determine the difference information between the reference audio and the target audio.
[0142] The evaluation network splits the first and second spectral information by channel to obtain the original time-domain spectral information (i.e., the phase spectrum and amplitude spectrum of the original time domain) and the spectral information of the human ear perception domain (the phase spectrum and amplitude spectrum of the human ear perception domain). Then, the first feature extraction layer extracts features from the phase spectrum and amplitude spectrum of the original time domain to obtain the first spectral features. The second feature extraction layer extracts features from the phase spectrum and amplitude spectrum of the human ear perception domain to obtain the second spectral features.
[0143] Since the stacking order of amplitude spectrum and phase spectrum in the first spectral information is: reference audio left channel amplitude spectrum, reference audio right channel amplitude spectrum, reference audio left channel phase spectrum, reference audio right channel phase spectrum, reference human ear perception domain audio left channel amplitude spectrum, reference human ear perception domain audio right channel amplitude spectrum, reference human ear perception domain audio left channel phase spectrum, and reference human ear perception domain audio right channel phase spectrum; and the stacking order of amplitude spectrum and phase spectrum in the second spectral information is: target audio left channel amplitude spectrum, target audio right channel amplitude spectrum, target audio left channel phase spectrum, target audio right channel phase spectrum, target human ear perception domain audio left channel amplitude spectrum, target human ear perception domain audio right channel amplitude spectrum, target human ear perception domain audio left channel phase spectrum, and target human ear perception domain audio right channel phase spectrum; therefore, in the evaluation network, the first 4 channels of the first spectral information and the second spectral information (latitude: [8,f,257]) are used as the original time domain spectral information, and the last 4 channels are used as the human ear perception domain spectral information. In this way, after splitting the first and second spectral information, the input to the corresponding feature extraction layer can make the features output by the feature extraction layer more accurate.
[0144] In one embodiment, the evaluation process of the network training is as follows: Figure 6 As shown, this can be achieved through steps 601 to 604:
[0145] Step 601: Obtain sample audio.
[0146] The sample audio includes at least three sample numbers, such as anchor samples, positive samples, and negative samples.
[0147] In some possible implementations, the sample audio is constructed using triples to enrich the sample audio; that is, step 601 above can be achieved through the following process:
[0148] First, retrieve at least three preset audio tracks from the preset audio database.
[0149] The preset audio database includes multiple audio data sets labeled with audio quality tags. At least three preset audio tracks are randomly selected from this preset audio database and their energy is normalized to the same signal-to-noise ratio. For example, three audio data sets are randomly sampled from the preset audio database, and their energy is normalized to -25 dB.
[0150] Secondly, according to different preset signal-to-noise ratios, preset noise is added to the at least three preset audio tracks to obtain at least three processed audio tracks.
[0151] The preset noise can be a randomly sampled noise data point from a noise library. Three signal-to-noise ratios (SNRs) are sampled within a preset SNR range, for example, three different SNRs within the range of [-20dB, +30dB], to obtain three different preset SNRs. Then, according to the magnitude of the preset SNR, the preset noise is added to the audio data of each day, so that the SNR of the processed audio data after adding noise is the preset SNR.
[0152] Next, from the at least three processed audio samples, identify the positive sample, the anchor sample with the highest signal-to-noise ratio, and the negative sample with the lowest signal-to-noise ratio.
[0153] Taking three processed audio data as an example, the three processed audio data are arranged in descending order of signal-to-noise ratio (SNR). The processed audio data with the highest SNR is used as the anchor sample, the processed audio data with the lowest SNR is used as the negative sample, and the remaining processed audio data is used as the positive sample.
[0154] Finally, the anchor sample, the positive sample, and the negative sample are determined as the sample audio.
[0155] In this way, after obtaining anchor samples, positive samples, and negative samples according to different signal-to-noise ratios, the anchor samples, positive samples, and negative samples are formed into triples as sample audio. This enables the evaluation network to be trained to learn feature representations with better discriminative power, making samples of the same category closer together, while samples of different categories are more dispersed.
[0156] Step 602: Perform human ear perception processing on the sample audio to obtain human ear perception domain sample audio.
[0157] In this process, the anchor samples, positive samples, and negative samples in the sample audio are transformed from the original time domain to the human ear perception domain to simulate the human ear's perception and processing of these sample data, thus obtaining the human ear perception domain sample audio, namely the anchor samples, positive samples, and negative samples of the human ear perception domain.
[0158] Step 603: The training evaluation network is used to evaluate the sample audio and the human ear perception domain sample audio to obtain the objective distortion quality evaluation loss and spatial quality loss of the sample audio.
[0159] Specifically, after performing time-frequency transformation on the sample audio and the sample audio in the human ear's perceptual domain to obtain sample spectral information, it is input into the evaluation network to be trained. The objective distortion quality assessment loss of the sample audio is used to represent the difference between the predicted objective distortion quality assessment of the sample audio and the true value of the sample audio. The spatial quality loss is used to represent the difference between the predicted spatial quality of the sample audio and the true value of the sample audio.
[0160] For example, by performing a short-time Fourier transform on the sample audio and the sample audio in the human ear's perceptual domain, and extracting the amplitude and phase spectra in the resulting complex domain, sample spectral information is obtained. This sample spectral information includes: the amplitude and phase spectra of the original time-domain sample audio, and the amplitude and phase spectra of the sample audio in the human ear's perceptual domain. The amplitude and phase spectra of the original time-domain sample audio, and the amplitude and phase spectra of the sample audio in the human ear's perceptual domain, are then input into the evaluation network to be trained.
[0161] Step 604: Adjust the network parameters of the evaluation network to be trained based on the objective distortion quality assessment loss and spatial quality loss to obtain the evaluation network.
[0162] The network parameters for the evaluation network to be trained include the weights and learning rate. The network parameters are iteratively combined with objective distortion quality assessment loss and spatial quality loss until the function values corresponding to the output objective distortion quality assessment loss and spatial quality loss converge, thus obtaining the evaluation network.
[0163] In this embodiment, during the training process of the evaluation network to be trained, by converting the sample audio to the human ear perception domain, the sample spectral information can be fused with the audio data of the human ear perception domain. Finally, using the sample spectral information, the objective distortion quality assessment loss and spatial quality loss of the sample audio are determined, and based on the objective distortion quality assessment loss and spatial quality loss, the network parameters of the evaluation network to be trained are adjusted to obtain the evaluation network. Thus, since the sample spectral information comprehensively considers the sample audio from the human ear perception domain and the original temporal domain sample audio, the evaluation network to be trained can learn the characteristics of the human ear perception domain, thereby improving the performance of the evaluation network.
[0164] In some possible implementations, the network parameters of the evaluation network to be trained are adjusted by combining the objective distortion quality assessment loss and the spatial quality loss as the overall loss. That is, step 604 above can be achieved through the following steps 641 and 642 (not shown in the figure):
[0165] Step 641: Determine the overall loss based on the objective distortion quality assessment loss and the spatial quality loss.
[0166] The total loss is obtained by summing the objective distortion quality assessment loss and the spatial quality loss and then averaging them.
[0167] Step 642: Using the overall loss, adjust the network parameters of the evaluation network to be trained to obtain the evaluation network.
[0168] Specifically, the network parameters of the evaluation network to be trained are iteratively applied using this overall loss to ensure that the total loss output by the evaluation network converges. Thus, since the spatial loss learns from the audio samples of the human auditory perception domain, using the average of the objective distortion quality evaluation loss and the spatial quality loss as the overall loss to train the evaluation network allows it to fully learn the perceptual characteristics of the human auditory perception domain for audio data. This results in the evaluation network's predictions of the target audio being more closely aligned with human auditory perception characteristics.
[0169] In one embodiment, the evaluation network to be trained includes an objective distortion quality evaluation subnetwork and a spatial quality evaluation subnetwork. Evaluation processes such as feature extraction and temporal aggregation are performed on the sample spectral information of anchor samples, positive samples, and negative samples to determine the objective distortion quality evaluation loss and spatial quality loss. That is, step 603 above can be implemented through the following steps 631 to 633:
[0170] Step 631: The objective distortion quality evaluation subnetwork in the evaluation network to be trained is used to evaluate the anchor sample, the positive sample and the negative sample to obtain the objective distortion quality evaluation loss of the sample audio.
[0171] Specifically, the first sample spectral information of the anchor sample, the first sample spectral information of the positive sample, and the first sample spectral information of the negative sample are input into the objective distortion quality assessment subnetwork. The input spectral information is processed by the first feature extraction layer, the first temporal aggregation layer, and the objective distortion quality assessment head within the objective distortion quality assessment subnetwork to obtain the objective distortion quality assessment loss of the sample audio. In this way, the objective distortion quality assessment loss is independent of the human auditory perception domain, enabling accurate learning of the original temporal domain sample features.
[0172] Step 632: Use the spatial quality evaluation subnetwork in the evaluation network to be trained to evaluate the anchor samples of the human ear perception domain.
[0173] Specifically, the second sample spectral information of the anchor sample in the human ear perception domain is input into the spatial quality evaluation subnetwork. The input spectral information is processed through the second feature extraction layer and the second time-domain aggregation layer to obtain the output of the second time-domain aggregation layer.
[0174] Step 633: Based on the evaluation and processing results of the anchor sample in the human ear perception domain and the evaluation and processing results of the anchor sample, determine the spatial quality loss.
[0175] Specifically, the output of the second temporal aggregation layer is weighted and summed with the evaluation results of the anchor samples output by the objective distortion quality evaluation subnetwork, and then input into the spatial quality head to obtain the spatial quality loss. Thus, by inputting only the anchor samples from the human ear perception domain into the spatial quality head of the evaluation network to be trained, this spatial quality loss is obtained. In this way, using the sample aggregation features of anchor samples with a high signal-to-noise ratio to determine the spatial quality loss can improve the accuracy of the evaluation network's spatial quality prediction.
[0176] In some possible implementations, the objective distortion quality assessment loss is determined by analyzing the depth feature distance between the anchor sample and the positive sample, as well as the depth feature distance between the positive sample and the negative sample. That is, step 631 above can be achieved through the following process:
[0177] First, determine the first sample depth feature distance between the original time domain sample features corresponding to the anchor sample and the original time domain sample features corresponding to the positive sample.
[0178] The first sample depth feature distance can be obtained by determining the 1-norm of the difference between the original time domain sample features and the original time domain sample features corresponding to the positive sample.
[0179] Secondly, the second sample depth feature distance is determined between the original time domain sample features corresponding to the anchor sample and the original time domain sample features corresponding to the negative sample.
[0180] The second sample depth feature distance can be obtained by determining the 1-norm of the difference between the original time domain sample features corresponding to the anchor sample and the original time domain sample features corresponding to the negative sample.
[0181] Finally, the objective distortion quality assessment loss is determined based on the difference between the first sample depth feature distance and the first sample depth feature distance.
[0182] In this process, after determining the difference between the first sample depth feature distance and the second sample depth feature distance, the difference is summed with a preset distance parameter. The maximum value between this summation and 0 is used as the objective distortion quality assessment loss. To prevent the network from outputting all zeros, the preset distance parameter is initially set to 0.5 and gradually increased to 1.5 as training progresses. This setting of the preset distance parameter ensures that the differences between different data points increase during the learning process, resulting in better clustering performance. Thus, by using the maximum value between the difference between the first sample depth feature distance and the second sample depth feature distance and the constant value 0 as the objective distortion quality assessment loss, the objective distortion quality assessment loss can more fully learn the differences between different samples, thereby improving the clustering performance of the network being trained.
[0183] In some possible implementations, spatial quality loss is obtained by classifying and predicting the azimuth of anchor samples, which can be achieved through the following process:
[0184] First, based on the evaluation and processing results of the human ear perception domain anchor samples, the prediction probability distribution of the predicted azimuth angle of the human ear perception domain anchor samples within a preset angle interval set is predicted.
[0185] Specifically, after obtaining the features of the human ear's perceptual domain corresponding to the anchor sample, the evaluation network to be trained predicts the horizontal and vertical angles of the features of the human ear's perceptual domain of the anchor sample in polar coordinates, thereby obtaining the predicted azimuth angle of the anchor sample. The predicted azimuth angles falling within a preset set of angle intervals are then classified to obtain the predicted probability distribution of the predicted azimuth angle in each preset angle interval.
[0186] Secondly, based on the true azimuth angle and the predicted azimuth angle of the evaluation processing results of the anchor sample, the included angle difference information is determined.
[0187] The included angle difference information is the difference between the polar coordinates representing the true azimuth angle and the polar coordinates representing the predicted azimuth angle, reflecting the angle difference between the vectors corresponding to the true azimuth angle and the vectors corresponding to the predicted azimuth angle. Since the azimuth angle includes elevation and horizontal angles, this included angle difference information can be determined by the true horizontal angle of the original time-domain sample features and the predicted horizontal angle in the predicted azimuth angle, as well as the true elevation angle of the original time-domain sample features and the predicted elevation angle in the predicted azimuth angle. For example, it can be determined as the large spherical distance between the true value of the sample features in the original time-domain and the predicted value of the features in the human ear perception domain.
[0188] Next, based on the true value distribution of the true azimuth angle in the preset angle interval set and the predicted probability distribution, the distribution difference information is determined.
[0189] In this process, after obtaining the predicted probability distribution of the predicted azimuth angle in each preset angle interval, the Earth's moving average distance between the true value distribution and the predicted probability distribution is determined to represent the difference between the two probability densities, thus obtaining the distribution difference information.
[0190] Finally, the spatial mass loss is determined based on the included angle difference information and the distribution difference information.
[0191] Specifically, the angular difference information and the distribution difference information are summed to obtain the spatial quality loss. Thus, by determining the predicted probability distribution of the anchor sample's predicted azimuth angle within a preset angle interval set, the distribution difference information between the true value distribution and this predicted probability distribution can be analyzed. Combining the angular difference information between the true value azimuth angle and the predicted azimuth angle with this distribution difference information allows the features of the human ear's perception domain to be introduced into the spatial quality loss. This enables the spatial quality loss to learn the features of the human ear's perception domain, and consequently, the evaluation network to learn the features of the human ear's perception domain, improving the evaluation results output by the evaluation network to better reflect the characteristics of human ear perception.
[0192] In one embodiment, the audio data evaluation method provided in this application performs human ear perception processing on the original spatial audio to obtain spatial audio in the human ear perception domain as auxiliary information to assist in spatial quality prediction, thereby improving the correlation between the model evaluation results and human subjective perception. The original spatial audio and the spatial audio in the human ear perception domain are first subjected to short-time Fourier transforms to extract amplitude and phase spectra, and the stacked amplitude / phase spectrum matrix is used as input to the neural network. Then, the amplitude / phase spectrum matrices of the original and the spatial audio in the human ear perception domain are simultaneously input into a deep neural network for feature extraction and temporal aggregation to calculate the corresponding embedding layers. The embedding layer of the original spatial audio is directly input into the objective distortion quality assessment task head for objective distortion quality prediction, because the objective distortion quality assessment task is an objective task unrelated to human ear perception characteristics. Correspondingly, for the spatial quality task head related to human ear perception characteristics, we use the embedding layer of the spatial audio in the human ear perception domain, weighted by learnable weights, to assist the embedding layer of the original spatial audio in spatial quality prediction, thereby obtaining prediction results that better conform to human subjective perception characteristics. Finally, we obtain the spatial audio quality score by calculating the deep feature distance between the embedding layers of the reference audio and the target audio, including the objective distortion quality assessment score, the spatial quality score, and the overall quality score. Detailed explanations of each part are given below:
[0193] Multiple filtering operations are used to simulate the human ear's perception and processing of sound, mapping the original spatial audio to the human ear's perceptual domain, such as... Figure 7 As shown, this can be achieved through the following steps:
[0194] Step 701: Input reference audio 71 and target audio 72.
[0195] Step 702: Filter the input spatial audio to obtain filtered audio.
[0196] Here, a second-order Butterworth bandpass filter with a bandpass frequency range of [500Hz, 2000Hz] is used to filter the input spatial audio to simulate the frequency response of the human outer and middle ear and the perception and processing of sound.
[0197] Step 703: Filter the filtered audio to simulate the human basilar membrane's perception and processing of sound.
[0198] Specifically, a fourth-order Gammatone filter bank is used to filter the filtered audio output from step 702, simulating the human basilar membrane's perception and processing of sound. The sampling rate of this Gammatone filter bank is equivalent to the sampling rate of spatial audio, with a minimum cutoff frequency of 50Hz, a maximum cutoff frequency of 14000Hz, a fundamental frequency of 700Hz, and one Gammatone filter per equivalent rectangular bandwidth (order 4, bandwidth factor 1).
[0199] Step 704: Process the output of step 703 to simulate the cochlear compression process of sound in the human cochlea.
[0200] Here, the output of step 703 is processed using instantaneous compression with an exponent of 0.4 to simulate the cochlear compression process of sound in the human cochlea.
[0201] Step 705: Low-pass filter is applied to the output of step 704 to simulate the mechatronic transmission process of human cochlear hair cells.
[0202] In this process, a half-wave correction and a fifth-order Butterworth low-pass filter with a cutoff frequency of 770Hz are used to filter the output of step 704 to simulate the mechatronic transmission process of the cochlear hair cells in the human ear. The output of step 705 is the spatial audio mapped to the human ear's perceptual domain, thus obtaining the reference audio 73 and the target audio 74 in the human ear's perceptual domain.
[0203] After obtaining the reference audio 73 and the target audio 74 in the human ear's perceptual domain, short-time Fourier transforms are performed on the pairs <reference audio, reference audio in the human ear's perceptual domain> and <target audio, target audio in the human ear's perceptual domain> to extract the amplitude spectrum and phase spectrum. These are then stacked to form an amplitude / phase spectrum matrix, which serves as the input to the deep neural network, as follows: Figure 8 As shown:
[0204] Step 801: Perform a short-time Fourier transform on the input raw audio to extract the time spectrum, and perform a short-time Fourier transform on the audio in the human ear perception domain to extract the time spectrum.
[0205] Specifically, a short-time Fourier transform (SFT) is performed on the input spatial audio (i.e., reference audio 71, reference audio 73 in the human ear's perceptual domain, target audio 72, and target audio 74 in the human ear's perceptual domain) using a 512-point length, a 50% (256-point) skip length, and a 512-point Hamming window. This yields the complex time-frequency spectrum 811 and the complex time-frequency spectrum 812 in the human ear's perceptual domain, respectively. The SFT reduces spectral leakage by employing a 512-point length, a 50% (256-point) skip length, and a 512-point Hamming window.
[0206] Step 802: Extract the amplitude spectrum and phase spectrum from the complex time spectrum 811; and extract the amplitude spectrum and phase spectrum from the complex time spectrum 812 in the human ear perception domain.
[0207] Step 803: Stack the amplitude spectrum and phase spectrum in the channel dimension to obtain the input matrix.
[0208] For the <reference audio, the reference audio in the human ear's perceptual domain>, we stack the amplitude and phase spectra in the channel dimension according to the order of [reference audio left ear amplitude spectrum, reference audio right ear amplitude spectrum, reference audio left ear phase spectrum, reference audio right ear phase spectrum, reference human ear perceptual domain audio left ear amplitude spectrum, reference human ear perceptual domain audio right ear amplitude spectrum, reference human ear perceptual domain audio left ear phase spectrum, reference human ear perceptual domain audio right ear phase spectrum], obtaining an input matrix of dimension [c=8, f, T=257], i.e., the reference audio amplitude spectrum / phase spectrum 813. Here, c represents 8 different channels, f represents a frame (proportional to the duration of the input audio), and T represents the time resolution (frequency), which is half the number of points in the short-time Fourier transform. For the <target audio, the target audio in the human ear's perceptual domain>, the amplitude and phase spectra are stacked in the same stacking order as the reference audio, obtaining an input matrix of dimension [8, f, 257], i.e., the target audio amplitude spectrum / phase spectrum 814.
[0209] Step 804: Input the input matrix into the neural network to determine the quality of the target audio relative to the reference audio.
[0210] In this embodiment, the calculated amplitude / phase spectrum matrices of the reference audio and the target audio are input into the deep neural network to calculate the embedding layer. The training process, loss function, and prediction process of the deep neural network are described below:
[0211] For the training process of the neural network, this application embodiment is based on a multi-task learning framework and uses a triplet approach to train the neural network, such as... Figure 9 As shown:
[0212] First, input the stacked amplitude and phase spectrum matrices.
[0213] In this process, three spatial audio samples are randomly sampled from the spatial audio dataset and their energy is normalized to -25dB. One noise data point is randomly sampled from the noise dataset, and three signal-to-noise ratios (SNRs) are uniformly sampled across the range [-20dB, +30dB], ensuring that SNR A > P > N to satisfy the condition |AP| < |AN| in triplet training. Then, the three different spatial audio samples, but with the same energy distribution normalized, are combined with the same noise sample and noise is added at three different SNR levels A, P, and N to construct a <anchor sample, positive sample, negative sample> triplet. Next, the anchor sample a, positive sample p, and negative sample n undergo human auditory processing and short-time Fourier transform, respectively, and the amplitude / phase spectrum matrices are extracted and input into the neural network to calculate the embedding layer. Specifically, the amplitude / phase spectrum matrix 91 of anchor sample a, the amplitude / phase spectrum matrix 93 of negative sample n, and the amplitude / phase spectrum matrix 92 of positive sample p are input into the neural network through input module 94 to calculate the embedding layer. The first four channels (0 to 3) input the amplitude and phase spectrum matrix 95 of the original audio, and the last four channels (4 to 7) input the amplitude and phase spectrum 96 of the human ear's perceptual domain.
[0214] Secondly, the amplitude / phase spectrum matrices of the anchor samples, negative samples, and positive samples are processed through the feature extraction layer 901, the temporal aggregation layer 902, the objective distortion quality assessment head 903, and the spatial quality head 904, respectively, to calculate the embedding layer and prediction results. Furthermore, based on the embedding layer and prediction results, the objective distortion quality assessment loss 905 and the spatial quality loss 906 are determined, and the overall quality loss 907 is calculated based on the objective distortion quality assessment loss 905 and the spatial quality loss 906 for backpropagation.
[0215] In deep neural networks, feature extraction layers are used to learn signal features while preserving the signal's temporal structure. The feature extraction layer consists of six stacked Inception-style feature extraction blocks, with the number of channels increasing progressively. After the sixth feature extraction block, 1x1 convolutions are used to compress the channel dimension and frequency dimension. The input dimension of the feature extraction block is [B, 4, f, 257], and the output dimension is [B, f, 64]. Feature extraction is performed on both the original spatial audio and the spatial audio in the human auditory perception domain, using feature extraction layers with shared parameters. The output of the feature extraction block serves as the input to the temporal aggregation layer. In the feature extraction layer, the number of inputs, outputs, and convolutional kernels used are shown in Table 1. The "parameter column" (such as [[1],[2,4],[1,2],[3,1]) represents [[the number of output channels of the first 1*1 convolution], [the number of output channels of the second 1*1 convolution and the number of channels of the 3*3 convolution], [the number of output channels of the third 1*1 convolution and the number of channels of the 5*5 convolution], [the size of the maximum pooling kernel and the number of output channels of the maximum pooling kernel], respectively.
[0216] Table 1 Parameter settings in the feature extraction layer structure
[0217] Layer type kernel size Input dimensions Output size Parameter settings Input layer - B*4*f*257 B*4*f*257 - Feature extraction layer 1 - B*4*f*257 B*8*f*257 [[1],[2,4],[1,2],[3,1]] Feature extraction layer 2 - B*8*f*257 B*16*f*257 [[2],[4,8],[2,4],[3,2]] Feature extraction layer 3 - B*16*f*257 B*32*f*257 [4],[8,16],[4,8],[3,4] Feature extraction layer 4 - B*32*f*257 B*32*f*257 [4],[8,16],[4,8],[3,4] Feature extraction layer 5 - B*32*f*257 B*64*f*257 [8],[16,32],[8,16],[3,8] Feature extraction layer 6 - B*64*f*257 B*64*f*257 [8],[16,32],[8,16],[3,8] Frequency Max Pooling 1*4 B*64*f*257 B*64*f*64 - 1x1 convolution 1*1 B*64*f*257 B*1*f*64 - Batch Normalization (BN) - B*1*f*257 B*1*f*64 - Extruded layer - B*1*f*257 B*f*64 -
[0218] The structure of the feature extraction layer is as follows: Figure 10 The feature extraction layer 1000 in the diagram includes: a pooling layer 1001 with a stride of 4, an aggregation layer 1002, and four batch normalization layers and multiple convolutional layers.
[0219] In deep neural networks, a temporal aggregation layer is used to learn long-term dependencies from the frame-level representations generated by the feature extraction layer. The temporal aggregation block consists of four stacked temporal convolutional networks, each composed of two 1*3 convolutional kernels, weight normalization, and a residual module. Causal convolution is used to extract temporal information, and dilated convolution is used to continuously expand the receptive field. In the four-layer temporal convolutional network, the number of input channels first increases from 64 to 128 and 256, then decreases from 256 to 128 and 64 before output. Therefore, the input and output dimensions of the temporal aggregation layer are both [B, f, 64]. Temporal aggregation is performed on both the original spatial audio and the spatial audio in the human ear's perceptual domain, using a temporal aggregation layer with shared parameters. The output of the temporal aggregation layer serves as the input to the objective distortion quality assessment head and the spatial quality head, as shown in Table 2.
[0220] Table 2 Temporal Aggregation Layer Structure
[0221] Layer type kernel size Expansion parameters Number of channels Input dimensions Output size Input layer - - 64 B*f*64 B*f*64 Temporal convolution block 1 1*3 1 128 B*f*64 B*f*128 Temporal convolution block 2 1*3 2 256 B*f*128 B*f*256 Temporal convolution block 3 1*3 4 128 B*f*258 B*f*128 Temporal convolution block 4 1*3 8 64 B*f*128 B*f*64
[0222] Here, the structure of the temporal aggregation layer is as follows: Figure 10The temporal aggregation layer 1010 in the figure includes two 1*3 convolutional kernels and two weight normalizations 1011 and 1012.
[0223] In deep neural networks, an objective distortion quality assessment task head is used to learn an objective distortion quality assessment representation of spatial audio from the embedding layer of the temporal aggregation layer. For example... Figure 10 As shown in the Objective Distortion Quality Evaluation Head 1020, it consists of two stacked 1*3 convolutional layers and weight normalization layers 1021 and 1022, with a stride of 1 and padding of 1 to ensure that the input and output dimensions remain unchanged. Following this are two fully connected layers that do not alter the channel dimensions. Since the objective distortion quality evaluation task is independent of human perception characteristics, the objective distortion quality evaluation head directly uses the temporal aggregation layer embedding layer of the original spatial audio as input. The input and output dimensions are both [B, f, 64], and the structure is shown in Table 3.
[0224] Table 3 Objective Distortion Quality Assessment Task Header Structure
[0225] Layer type kernel size Step length Number of channels Input dimensions Output size Input layer - - 64 B*f*64 B*f*64 Convolutional layer 1 1*3 1 64 B*f*64 B*f*64 Weight normalization - - 64 B*f*64 B*f*64 Convolutional layer 2 1*3 1 64 B*f*64 B*f*64 Weight normalization - - 64 B*f*64 B*f*64 Fully connected layer 1 - - 64 B*f*64 B*f*64 Fully connected layer 2 - - 64 B*f*64 B*f*64
[0226] In the neural network, a spatial quality task head is used to learn the spatial quality representation of spatial audio from the embedding layer of the temporal aggregation layer. For the spatial quality prediction task, i.e., the azimuth prediction task, the interval [0°, 360°] is divided into 50 equidistant sectors and treated as a 50-class classification task. The spatial quality task head is as follows: Figure 10 The spatial quality head (1030) consists of two stacked 1x3 convolutional layers and weight normalization layers (1031 and 1032), two fully connected layers, and a softmax activation function. Its input dimension is [B, f, 64], and its output dimension is [B, f, 50], representing the probability distribution of the spatial audio's azimuth corner in a specific sector for each frame. Since the spatial quality task head is strongly correlated with human subjective perception, a temporal aggregation layer embedding the spatial audio from the human ear's perception domain is used to assist in the spatial quality task head's prediction. Specifically, a learnable weight is used to weight and combine the temporal aggregation layer embedding the original spatial audio with the temporal aggregation layer embedding the spatial audio from the human ear's perception domain. This weighted combination is then input into the spatial quality task head for prediction, utilizing the characteristics of human hearing and perception to assist in the objective distortion quality assessment prediction of spatial audio, as shown in Table 4.
[0227] Table 4 Space Mass Mission Header Structure
[0228] Layer type kernel size Step length Number of channels Input dimensions Output size Input layer - - 64 B*f*64 B*f*64 Convolutional layer 1 1*3 1 64 B*f*64 B*f*64 Weight normalization - - 64 B*f*64 B*f*64 Convolutional layer 2 1*3 1 64 B*f*64 B*f*64 Weight normalization - - 64 B*f*64 B*f*64 Fully connected layer 1 - - 64 B*f*64 B*f*64 Fully connected layer 2 - - 50 B*f*64 B*f*50 Soft Maximum B*f*50 B*f*50
[0229] Next, calculate the objective distortion quality assessment loss, i.e., the triplet loss function, as shown in formula (1):
[0230] (1);
[0231] in, The depth feature distance of the full feature activation stack between positive samples and anchor samples. It is the depth feature distance of the full feature activation stack between positive and negative samples, as shown in formula (2):
[0232] (2);
[0233] here, These are temporal resolution, frequency band, and number of channels, respectively. In model training, the deep feature stack involved in the objective distortion quality assessment loss consists of a feature extraction layer, a temporal aggregation layer, and an embedding layer of the objective distortion quality assessment task head output. To prevent the network from outputting all zeros, the spacing parameter is initially set to 0.5 and gradually increased to 1.5 as the number of training epochs increases.
[0234] Here, the objective distortion quality assessment task head is as follows: Figure 10 The objective distortion quality assessment head 1020 in the example includes two 1*3 convolutional kernels and two weight normalization layers 1021 and 1022. The objective distortion quality assessment task head is as follows: Figure 10 The spatial quality head 1030 in the diagram includes two 1*3 convolutional kernels and two weight normalization layers 1031 and 1032.
[0235] Next, spatial quality loss is calculated. Here, the spatial localization task is treated as a classification task, and the angle [0°, 360°] is mapped to 50 equidistant sectors such as [0, 49]. The washersine distance, reflecting continuous angle changes, and the Earth Movement Distance (EMD distance), which emphasizes inter-class relationships, are used as loss functions. Since excessive noise may affect the prediction of the objective distortion quality assessment head, the anchor spatial audio with the highest signal-to-noise ratio (lowest noise) is used to calculate the spatial quality loss. The anchor sample is labeled 'a', and the output of the objective distortion quality assessment head in the human auditory perception domain is... The loss function is then shown in formula (3):
[0236] (3);
[0237] Where loss_haversine represents the large sphere distance, or the semi-versus distance, as shown in formula (4):
[0238] (4);
[0239] in, , The actual and predicted horizontal angles are represented by a and , respectively, based on a and . It is calculated from the median of the sector in which it is located; , These represent the actual and predicted elevation angles, respectively. Since this embodiment does not perform elevation angle prediction, it is assumed that... = =True elevation angle.
[0240] Here, loss_emd represents the average distance traveled on Earth, as shown in formula (5):
[0241] (5);
[0242] in yes The cumulative probability density distribution function, and It is a probability distribution of a normal distribution calculated based on the correct sector index v in a, as shown in formula (6). This is the probability distribution of network prediction. The probability density distribution function:
[0243] (6);
[0244] Finally, the total loss is calculated as shown in formula (7):
[0245] (7);
[0246] Backpropagation is performed based on the loss to train the network, enabling it to predict objective quality attributes of spatial audio, such as spatial quality and objective distortion quality assessment, with the assistance of human auditory perception.
[0247] In this embodiment of the application, for the prediction process of the neural network, the amplitude / phase spectrum matrix of the input reference audio and target audio is calculated to obtain the corresponding calculation results, that is, the corresponding embedding layer, which is used for subsequent calculation of deep feature distance and objective score. Figure 11The network structure of the spatial audio objective evaluation model for human ear perception characteristics, as shown, processes the amplitude / phase spectrum matrices of the input reference audio and target audio to obtain the embedding layers of the reference audio and target audio. The amplitude / phase spectrum 1101 of the reference audio and the amplitude / phase spectrum 1102 of the target audio are respectively input to the network architecture through the input module 1103. Here, the dimension of the input amplitude / phase spectrum matrix is [8, f, 257], the first 4 channels [4, f, 257] are the amplitude / phase spectrum matrix 1104 of the original spatial audio, and the last 4 channels [4, f, 257] are the amplitude / phase spectrum matrix 1105 of the spatial audio in the human ear perception domain. The feature extraction layer 1106 extracts features from the amplitude / phase spectrum matrix 1104 of the original spatial audio in the first 4 channels, and uses the extracted features as the embedding layer of the feature extraction layer and inputs them into the temporal aggregation layer 1109. The feature extraction layer 1106 extracts features from the amplitude / phase spectrum matrix 1105 of the spatial audio in the human ear perception domain on the last four channels, and inputs the extracted features into the temporal aggregation layer 1109. The temporal aggregation layer 1109 performs convolution, weight normalization, and residual processing on the amplitude / phase spectrum matrix 1104 of the original spatial audio and the amplitude / phase spectrum matrix 1105 of the spatial audio in the human ear perception domain, respectively. The two output results are used as inputs to the objective distortion quality evaluation head 1107 to obtain the output embedding layer 1110. By weighting and combining the output of the amplitude / phase spectrum matrix 1104 of the original spatial audio in the temporal aggregation layer 1109, and the output of the amplitude / phase spectrum matrix 1105 of the spatial audio in the human ear perception domain in the temporal aggregation layer 1109, the weighted output is used as the input of the spatial quality head 1108 to obtain the output embedding layer 1110; alternatively, the output of the amplitude / phase spectrum matrix 1104 of the original spatial audio in the temporal aggregation layer 1109, and the output of the amplitude / phase spectrum matrix 1105 of the spatial audio in the human ear perception domain in the temporal aggregation layer 1109 can be concatenated, and the concatenated output is used as the input of the spatial quality head 1108 to obtain the output embedding layer 1110; thus, the reference audio embedding layer 1111 and the target audio embedding layer 1112 are finally obtained.
[0248] In other embodiments, the input to the objective distortion quality assessment head 1107 may be the output of the amplitude / phase spectrum matrix 1104 of the original spatial audio in the temporal aggregation layer 1109, or it may be the result of fusing the output of the amplitude / phase spectrum matrix 1104 of the original spatial audio in the temporal aggregation layer 1109 with the output of the amplitude / phase spectrum matrix 1105 of the spatial audio in the human ear perception domain in the temporal aggregation layer 1109. For example, the output results of the amplitude / phase spectrum matrix 1104 of the original spatial audio in the temporal aggregation layer 1109, and the output results of the amplitude / phase spectrum matrix 1105 of the spatial audio in the human ear perception domain in the temporal aggregation layer 1109, are weighted and combined before being input into the objective distortion quality evaluation head 1107; or, the output results of the amplitude / phase spectrum matrix 1104 of the original spatial audio in the temporal aggregation layer 1109, and the output results of the amplitude / phase spectrum matrix 1105 of the spatial audio in the human ear perception domain in the temporal aggregation layer 1109 are spliced together before being input into the objective distortion quality evaluation head 1107, so as to further improve the accuracy of the output results of the objective distortion quality evaluation head 1107.
[0249] In this process, the same feature extraction layer, temporal aggregation layer, and multi-task head processing are used for both reference / target audio and original spatial audio / spatial audio in the human ear's perceptual domain. Figure 12 The process shown is implemented as follows:
[0250] First, the input amplitude / phase spectrum is split according to the channel, into the original audio amplitude / phase spectrum and the amplitude / phase spectrum of the human ear perception domain.
[0251] Here, the human hearing perception processing module 1201 processes the input reference audio 1202 and target audio 1203 to obtain audio 1204 in the human hearing perception domain. Then, the reference audio 1202, target audio 1203, reference audio 1205 in the human hearing perception domain, and target audio 1206 in the human hearing perception domain are input to the time-frequency transformation module 1207. The time-frequency transformation module 1207 performs time-frequency transformation on the input audio to obtain a complex time spectrum 1220 and a complex time spectrum 1221 in the human hearing perception domain, and extracts the amplitude spectrum and phase spectrum 1222 from these time spectra. Then, the extracted amplitude and phase spectra are stacked in the channel dimension to obtain the input 1223 to the neural network, namely, the reference audio amplitude and phase spectrum matrix 1208 and the target audio amplitude and phase spectrum matrix 1209.
[0252] Next, the reference audio amplitude and phase spectrum matrix 1208 and the target audio amplitude and phase spectrum matrix 1209 are input into the deep neural network module 1210. The input phase spectrum and amplitude spectrum are divided into the original audio amplitude and phase spectrum 1224 and the human ear perception domain audio amplitude and phase spectrum 1225, and input into the same feature extraction layer 1230 to calculate the feature extraction embedding layer.
[0253] Furthermore, in the deep neural network module 1210, the original audio feature extraction embedding layer and the human ear perception domain audio feature extraction embedding layer are input into the same temporal aggregation layer 1231 to calculate the temporal aggregation embedding layer.
[0254] Next, the original audio temporal aggregation embedding layer is directly input into the objective distortion quality evaluation head 1232 to determine the objective distortion quality evaluation task head embedding layer (objective distortion quality evaluation is unrelated to human auditory perception).
[0255] Next, the original audio temporal aggregation embedding layer and the human ear perception domain audio temporal aggregation embedding layer are weighted and combined with a learnable weight α, and the spatial quality head 1233 is input to calculate the spatial quality task head embedding layer (spatial quality is a task related to human ear perception, and is predicted using an audio-assisted network in the human ear perception domain).
[0256] Finally: The original audio feature extraction embedding layer ([B,f,64]), the weighted combined temporal aggregation embedding layer ([B,f,64]), the objective distortion quality assessment embedding layer ([B,f,64]), and the spatial quality embedding layer output ([B,f,50]) are returned to obtain the final output embedding layer 1234, which includes a reference audio embedding layer and a target audio embedding layer. The deep feature distance is input as the scoring module 1211 to calculate the subsequent deep feature distance as the objective score of the spatial audio, thereby obtaining the objective distortion quality assessment 1212, the overall quality 1213, and the spatial quality 1214. Two different sets of embedding layers are calculated using the same network for both the reference audio and the target audio.
[0257] After obtaining the embedding layers of the target audio and the reference audio, the overall quality, objective distortion quality assessment, and spatial quality score of the target audio relative to the reference audio are obtained by calculating their depth feature distance. Figure 13 As shown:
[0258] First, the depth feature distance 1303 between the shared feature extraction layer and temporal aggregation layer embedding layer between the reference audio embedding layer 1301 and the target audio embedding layer 1302 is calculated as the overall quality score 1304 of the spatial audio.
[0259] Secondly, the depth feature distance 1305 between the feature extraction layer, temporal aggregation layer, and objective distortion quality assessment task head embedding layer of the reference audio and target audio is calculated as the objective distortion quality assessment score 1306 for spatial audio.
[0260] Next, the spatial quality score 1308 of the spatial audio is obtained by calculating the deep feature distance 1307 between the feature extraction layer, temporal aggregation layer, and spatial quality task head embedding layer of the reference audio and the target audio.
[0261] The depth feature distance is shown in formula (8):
[0262] (8);
[0263] in These are the temporal resolution, bandwidth, and number of channels of the l-th embedding layer, respectively.
[0264] Finally, the overall quality score, objective distortion quality assessment score, and spatial quality score are output as the evaluation results of the "Spatial Audio Objective Evaluation Method Based on Human Ear Perception Characteristics".
[0265] The audio data evaluation method proposed in this application can effectively perform objective evaluation of spatial audio, has better objective attributes, and its correlation with human subjective ratings in the human subjective evaluation dataset exceeds the baseline indices DPLM and SAQAM.
[0266] For objective testing, public area testing, clustering testing, and monotonicity testing were conducted respectively:
[0267] Common Area Test: In the common area test, two sets of data are constructed. The first set contains spatial audio pairs with different content but the same quality attribute (LQ or SQ). The second set contains spatial audio pairs with different content but different quality attributes. Quality attribute scores (LQ or SQ) are assigned to each audio pair, and the standard normalized normal distribution parameters (mean and variance) are calculated using maximum likelihood estimation. The standard normalized normal distribution curves for the two sets of data are then plotted. The area of the common region (overlapping region) of the two curves is the desired result. A smaller common region demonstrates stronger robustness to content changes and greater content invariance of the model.
[0268] Clustering Test: The clustering test verifies the model's retrieval performance for quality attributes (LQ or SQ). Specifically, 10 different quality level groups are generated according to different quality attributes. Each quality level group consists of 100 spatial audio tracks with the same quality attribute but mismatched content. The quality attributes of all 1000 spatial audio tracks are calculated, and the k (=10 / 25) audio tracks with the best scores are selected. The probability that these tracks belong to the quality level group with the highest quality attribute is calculated. The higher this probability, the stronger the network's clustering and retrieval characteristics, indicating that the embedding layer in the model truly captures quality attributes such as objective distortion quality assessment and spatial localization attributes.
[0269] Monotonicity Test: In the monotonicity test, we verify the monotonic behavior of the model as the perturbation level (LQ or SQ) increases. The monotonicity test generates a series of spatial audio samples with progressively increasing perturbation levels and calculates the corresponding quality attribute scores. Finally, we calculate the Spearman correlation coefficient between the quality attribute score (LQ or SQ) and the increasing perturbation levels (signal-to-noise ratio or deflection angle). A higher correlation coefficient demonstrates stronger monotonicity of the neural network to perturbations.
[0270] Meanwhile, in order to verify the correlation between the "spatial quality objective evaluation method based on human ear perception characteristics" proposed in this application embodiment and human subjective perception, the spatial audio human subjective evaluation dataset "Bitrate Compression Ambisonic" was used for subjective testing.
[0271] The purpose of this subjective test dataset is to study the impact of bitrate compression on Ambisonic audio. It includes 88 spatial audio tracks in 8 groups, covering six audio types: pink noise, human voice, castanets, bells, electronic music (2 groups), and acoustic music (2 groups). The subjective rating for this dataset is the MUSHAR rating, which assesses the degree of audio timbre distortion.
[0272] In this embodiment, objective tests are performed on the subjective test set, and the Spearman correlation coefficient between the objective test scores and the human subjective scores in the subjective test set is calculated. The larger the correlation coefficient, the higher the correlation between the model provided in this embodiment and human subjective perception. Moreover, experimental results demonstrate that the model provided in this embodiment has better subjective intuition and a stronger correlation with human subjective perception, thus proving the benefit and superiority of using human auditory perception characteristics to assist in the objective evaluation of spatial audio in this embodiment.
[0273] In this embodiment, the input spatial audio is first processed by human ear perception to obtain spatial audio mapped to the human ear perception domain. This spatial audio is used to assist in the objective evaluation of spatial audio, solving the problem that objective evaluation algorithms cannot reasonably consider both the subjective and objective attributes of spatial audio, and increasing the correlation between the evaluation results and human subjective perception. By introducing spatial audio from the human ear perception domain as auxiliary information, and using learnable weights to guide the embedding layer in the neural network to focus on the human ear perception attributes of spatial audio, the correlation between the model's subjective evaluation results and human subjective ratings can be enhanced. Furthermore, a series of subjective and objective evaluations were conducted on the proposed model. Experimental results show that, in both subjective and objective tests, the model proposed in this embodiment achieves better overall performance, possesses better objective characteristics, and has a stronger correlation with human subjective perception. This also proves the effectiveness of introducing human ear perception processing in this embodiment.
[0274] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages in other steps. Based on the same inventive concept, this application also provides an audio data evaluation apparatus for implementing the audio data evaluation method described above. The solution provided by this apparatus is similar to the implementation scheme described in the above method. Therefore, the specific limitations in one or more audio data evaluation apparatus embodiments provided below can be found in the limitations of the audio data evaluation method above, and will not be repeated here.
[0275] In one embodiment, such as Figure 14As shown, an audio data evaluation device 1400 is provided, including: an acquisition module 1401, a processing module 1402, an evaluation module 1403, and a determination module 1404, wherein: the acquisition module 1401 is used to acquire raw audio, the raw audio including a reference audio and a target audio to be evaluated; the processing module 1402 is used to perform human ear perception processing on the raw audio to obtain human ear perception domain audio, the human ear perception domain audio including a reference human ear perception domain audio corresponding to the reference audio and a target human ear perception domain audio corresponding to the target audio; the evaluation module 1403 is used to perform evaluation processing on the raw audio and the human ear perception domain audio using an evaluation network; the determination module 1404 is used to determine the difference information between the reference audio and the target audio based on the evaluation processing result, and determine the evaluation result of the target audio based on the difference information.
[0276] In one embodiment, the evaluation module 1403 is further configured to: perform objective evaluation processing on the original audio using the objective distortion quality evaluation subnetwork to obtain multiple reference objective evaluation features corresponding to the reference audio and multiple target objective evaluation features corresponding to the target audio; perform spatial evaluation processing on the human ear perception domain audio using the spatial quality evaluation subnetwork, and fuse the evaluation features obtained from the objective evaluation processing during the spatial evaluation processing to obtain multiple reference spatial evaluation features corresponding to the reference audio and multiple target spatial evaluation features corresponding to the target audio.
[0277] In one embodiment, the objective distortion quality evaluation subnetwork includes a cascaded first feature extraction layer, a first temporal aggregation layer, and an objective distortion quality evaluation head. The evaluation module 1403 is further configured to: perform feature extraction processing on the original audio using the first feature extraction layer to obtain a first spectral feature, the first spectral feature including a first reference spectral feature corresponding to the reference audio and a first target spectral feature corresponding to the target audio; and perform temporal aggregation processing on the first spectral feature using the first temporal aggregation layer to obtain a first aggregated feature, the first aggregated feature including a first reference aggregated feature corresponding to the reference audio and a first target spectral feature corresponding to the target audio. A target aggregation feature is obtained; the objective distortion quality assessment head is used to perform objective distortion quality assessment and identification processing on the first aggregation feature to obtain an objective distortion quality assessment and identification feature, which includes a reference objective distortion quality assessment and identification feature corresponding to the reference audio and a target objective distortion quality assessment and identification feature corresponding to the target audio; the first reference spectrum feature, the first reference aggregation feature, and the reference objective distortion quality assessment and identification feature are used as the plurality of reference objective assessment features; the first target spectrum feature, the first target aggregation feature, and the target objective distortion quality assessment and identification feature are used as the plurality of target objective assessment features.
[0278] In one embodiment, the spatial quality evaluation subnetwork includes a cascaded second feature extraction layer, a second temporal aggregation layer, and a spatial quality head. The evaluation module 1403 is further configured to: perform feature extraction processing on the human ear perception domain audio using the second feature extraction layer to obtain a second spectral feature, the second spectral feature including a second reference spectral feature corresponding to the reference audio and a second target spectral feature corresponding to the target audio; perform temporal aggregation processing on the second spectral feature using the second temporal aggregation layer to obtain a second aggregated feature, the second aggregated feature including a second reference aggregated feature corresponding to the reference audio and a second target aggregated feature corresponding to the target audio; and combine the second aggregated feature with the first... The aggregated features are weighted and fused to obtain fused features, which include reference fused features corresponding to the reference audio and target fused features corresponding to the target audio. The spatial quality head is used to perform spatial quality recognition processing on the fused features to obtain spatial quality recognition features, which include reference spatial quality recognition features corresponding to the reference audio and target spatial quality recognition features corresponding to the target audio. The second reference spectral feature, the second reference aggregated feature, and the reference spatial quality recognition feature are used as the plurality of reference spatial evaluation features. The second target spectral feature, the second target aggregated feature, and the target spatial quality recognition feature are used as the plurality of target spatial evaluation features.
[0279] In one embodiment, the evaluation module 1403 is further configured to perform at least one level of convolution processing on the second spectral feature according to a plurality of preset channel numbers that increase progressively in the second temporal aggregation layer to obtain a first convolution feature; perform the at least one level of convolution processing on the first convolution feature according to the plurality of preset channel numbers that decrease progressively in the second convolution feature to obtain a second convolution feature; and perform residual processing on the second convolution feature to obtain the second aggregated feature.
[0280] In one embodiment, the determining module 1404 is further configured to determine the overall feature distance between the reference audio and the target audio based on the first spectral feature, the second spectral feature, the first aggregated feature, and the second aggregated feature; determine the objective feature distance between the reference audio and the target audio based on the first spectral feature, the first aggregated feature, and the objective distortion quality assessment identification feature; and determine the spatial feature distance between the reference audio and the target audio based on the second spectral feature, the second aggregated feature, and the spatial quality identification feature.
[0281] In one embodiment, the evaluation module 1403 is further configured to perform time-frequency transformation on the reference audio and the reference human ear perception domain audio to obtain first spectrum information; perform time-frequency transformation on the target audio and the target human ear perception domain audio respectively to obtain second spectrum information; and input the first spectrum information and the second spectrum information into the evaluation network to perform evaluation processing on the original audio and the human ear perception domain audio using the evaluation network.
[0282] In one embodiment, the evaluation module 1403 is further configured to perform a complex domain transformation on the target audio and the target human ear perception domain audio to obtain the time spectrum of the target audio in the complex domain and the time spectrum of the target human ear perception domain audio in the complex domain; extract the amplitude spectrum and phase spectrum of the time spectrum; and stack the amplitude spectrum and phase spectrum corresponding to the target audio and the amplitude spectrum and phase spectrum corresponding to the target human ear perception domain audio to obtain the second spectrum information.
[0283] In one embodiment, the apparatus further includes: a sample acquisition module for acquiring sample audio; a perception processing module for performing human ear perception processing on the sample audio to obtain human ear perception domain sample audio; an evaluation processing module for using a training evaluation network to evaluate the sample audio and the human ear perception domain sample audio to obtain the objective distortion quality evaluation loss and spatial quality loss of the sample audio; and an adjustment module for adjusting the network parameters of the training evaluation network according to the objective distortion quality evaluation loss and the spatial quality loss to obtain the evaluation network. In one embodiment, the sample acquisition module is further configured to acquire at least three preset audio tracks from a preset audio database; add preset noise to the at least three preset audio tracks according to different preset signal-to-noise ratios to obtain at least three processed audio tracks; determine a positive sample, an anchor sample with the highest signal-to-noise ratio, and a negative sample with the lowest signal-to-noise ratio from the at least three processed audio tracks; and determine the anchor sample, the positive sample, and the negative sample as the sample audio.
[0284] In one embodiment, the human ear perception domain sample audio includes: human ear perception domain anchor samples, human ear perception domain positive samples, and human ear perception domain negative samples. The evaluation processing module is further configured to use an objective distortion quality evaluation subnetwork in the evaluation network to be trained to evaluate the anchor samples, the positive samples, and the negative samples to obtain the objective distortion quality evaluation loss of the sample audio; use a spatial quality evaluation subnetwork in the evaluation network to be trained to evaluate the human ear perception domain anchor samples; and determine the spatial quality loss based on the evaluation processing results of the human ear perception domain anchor samples and the evaluation processing results of the anchor samples.
[0285] In one embodiment, the evaluation processing module is used to predict the prediction probability distribution of the predicted azimuth angle of the anchor sample in the human ear perception domain within a preset angle interval set based on the evaluation processing result of the anchor sample; determine the included angle difference information based on the true azimuth angle of the evaluation processing result of the anchor sample and the predicted azimuth angle; determine the distribution difference information based on the true distribution of the true azimuth angle within the preset angle interval set and the predicted probability distribution; and determine the spatial quality loss based on the included angle difference information and the distribution difference information. In one embodiment, the adjustment module is further used to determine the overall loss based on the objective distortion quality evaluation loss and the spatial quality loss; and use the overall loss to adjust the network parameters of the evaluation network to be trained to obtain the evaluation network.
[0286] Each module in the aforementioned audio data evaluation device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in the processor of the electronic device in hardware form or independent of it, or stored in the memory of the electronic device in software form, so that the processor can call and execute the corresponding operations of each module.
[0287] In one embodiment, an electronic device is provided, which may be a terminal, and its internal structure diagram may be as follows: Figure 15 As shown, the electronic device includes a data acquisition device, a processor, a memory, an input / output interface, a communication interface, a display unit, and an input device. The data acquisition device, processor, memory, and input / output interface are connected via a system bus, and the communication interface, display unit, and input device are also connected to the system bus via the input / output interface. The processor provides computing and control capabilities. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The input / output interface is used for exchanging information between the processor and external devices. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, mobile cellular networks, NFC (Near Field Communication), or other technologies. When the computer program is executed by the processor, it implements an audio data evaluation method. The display unit is used to form a visually visible image and can be a display screen, a projection device, or a virtual reality imaging device. The display screen can be an LCD screen or an e-ink screen. The input device of the electronic device can be a touch layer covering the display screen, or buttons, trackballs, or touchpads set on the casing of the electronic device, or external keyboards, touchpads, or mice, etc.
[0288] Those skilled in the art will understand that Figure 15 The structures shown are merely block diagrams of some structures related to the present application and do not constitute a limitation on the electronic devices to which the present application is applied. Specific electronic devices may include more or fewer components than shown in the figures, or combine certain components, or have different component arrangements. In one embodiment, an electronic device is provided, including a memory and a processor. The memory stores a computer program, and the processor executes the computer program to perform the steps in the above-described method embodiments. In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored, and when executed by a processor, the computer program implements the steps in the above-described method embodiments. In one embodiment, a computer program product is provided, including a computer program, which, when executed by a processor, implements the steps in the above-described method embodiments.
[0289] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of related data must comply with the relevant laws, regulations and standards of relevant countries and regions. Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, database or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory may include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic resistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory may include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the various embodiments provided in this application may include at least one of relational databases and non-relational databases. Non-relational databases may include blockchain-based distributed databases, etc., and are not limited thereto. The processors involved in the various embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited thereto. The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as the combination of these technical features does not contradict each other, it should be considered within the scope of this specification. The above embodiments only illustrate several implementation methods of this application, and their descriptions are relatively specific and detailed, but they should not be construed as limiting the scope of this patent application.It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the appended claims.
Claims
1. An audio data evaluation method, characterized in that, The method includes: Acquire the raw audio, which includes a reference audio and a target audio to be evaluated; The original audio is processed by human ear perception to obtain human ear perception domain audio, which includes reference human ear perception domain audio corresponding to the reference audio and target human ear perception domain audio corresponding to the target audio. An objective distortion quality assessment subnetwork is used to perform objective assessment processing on the original audio at least to obtain multiple reference objective assessment features corresponding to the reference audio and multiple target objective assessment features corresponding to the target audio; a spatial quality assessment subnetwork is used to perform spatial assessment processing on the audio in the human ear perception domain, and the assessment features obtained from the objective assessment processing are fused during the spatial assessment processing to obtain multiple reference spatial assessment features corresponding to the reference audio and multiple target spatial assessment features corresponding to the target audio; Based on the evaluation results, the difference information between the reference audio and the target audio is determined, and the evaluation result of the target audio is determined based on the difference information.
2. The method according to claim 1, characterized in that, The objective distortion quality assessment subnetwork includes a cascaded first feature extraction layer, a first temporal aggregation layer, and an objective distortion quality assessment head. The objective distortion quality assessment subnetwork performs objective assessment processing on the original audio at least to obtain multiple reference objective assessment features corresponding to the reference audio and multiple target objective assessment features corresponding to the target audio, including: The original audio is processed by the first feature extraction layer to obtain a first spectral feature, which includes a first reference spectral feature corresponding to the reference audio and a first target spectral feature corresponding to the target audio. The first temporal aggregation layer is used to perform temporal aggregation processing on the first spectral features to obtain a first aggregated feature. The first aggregated feature includes a first reference aggregated feature corresponding to the reference audio and a first target aggregated feature corresponding to the target audio. The objective distortion quality assessment head is used to perform objective distortion quality assessment and identification processing on the first aggregated feature to obtain objective distortion quality assessment and identification features. The objective distortion quality assessment and identification features include reference objective distortion quality assessment and identification features corresponding to the reference audio and target objective distortion quality assessment and identification features corresponding to the target audio. The first reference spectral feature, the first reference aggregation feature, and the reference objective distortion quality assessment and identification feature are used as the plurality of reference objective assessment features; The first target spectral feature, the first target aggregation feature, and the target objective distortion quality assessment and identification feature are used as the objective assessment features of the plurality of targets.
3. The method according to claim 2, characterized in that, The spatial quality assessment subnetwork includes a cascaded second feature extraction layer, a second temporal aggregation layer, and a spatial quality head. The spatial quality assessment subnetwork is used to perform spatial assessment processing on the audio in the human ear's perceptual domain, and the assessment features obtained from objective assessment processing are fused during the spatial assessment process, including: The second feature extraction layer is used to perform feature extraction processing on the audio in the human ear perception domain to obtain a second spectral feature. The second spectral feature includes a second reference spectral feature corresponding to the reference audio and a second target spectral feature corresponding to the target audio. The second temporal aggregation layer is used to perform temporal aggregation processing on the second spectral features to obtain the second aggregated features. The second aggregated features include a second reference aggregated feature corresponding to the reference audio and a second target aggregated feature corresponding to the target audio. The second aggregated feature and the first aggregated feature are weighted and fused to obtain a fused feature, which includes a reference fused feature corresponding to the reference audio and a target fused feature corresponding to the target audio. The spatial quality head is used to perform spatial quality recognition processing on the fused features to obtain spatial quality recognition features, which include reference spatial quality recognition features corresponding to the reference audio and target spatial quality recognition features corresponding to the target audio. The second reference spectral feature, the second reference aggregation feature, and the reference space quality identification feature are used as the plurality of reference space evaluation features; The second target spectral feature, the second target aggregation feature, and the target spatial quality identification feature are used as the plurality of target spatial evaluation features.
4. The method according to claim 3, characterized in that, The step of using the second time-domain aggregation layer to perform time-domain aggregation processing on the second spectral features to obtain the second aggregated features includes: In the second temporal aggregation layer, the second spectral feature is subjected to at least one level of convolution processing based on multiple preset channel numbers that increase progressively to obtain the first convolution feature; Based on the progressively decreasing number of preset channels, the first convolutional feature is subjected to at least one level of convolution processing to obtain the second convolutional feature; The second convolutional feature is processed by residual processing to obtain the second aggregated feature.
5. The method according to claim 3, characterized in that, The difference information includes at least one of overall feature distance, objective feature distance, and spatial feature distance. Determining the difference information between the reference audio and the target audio based on the evaluation process includes: Based on the first spectral feature, the second spectral feature, the first aggregated feature, and the second aggregated feature, the overall feature distance between the reference audio and the target audio is determined; Based on the first spectral feature, the first aggregation feature, and the objective distortion quality assessment and identification feature, the objective feature distance between the reference audio and the target audio is determined; The spatial feature distance between the reference audio and the target audio is determined based on the second spectral feature, the second aggregation feature, and the spatial quality identification feature.
6. The method according to claim 1, characterized in that, The method further includes: Perform time-frequency transformation on the reference audio and the reference human ear perception domain audio to obtain the first spectrum information; The target audio and the target human ear perception domain audio are respectively subjected to time-frequency transformation to obtain the second spectrum information; The first and second spectral information are input into the evaluation network, which then splits the first and second spectral information by channel to obtain the original time-domain spectral information and the spectral information in the human ear perception domain. The objective distortion quality evaluation subnetwork performs objective evaluation processing on the original time-domain spectral information. The spatial quality evaluation subnetwork performs spatial evaluation processing on the spectral information in the human ear perception domain.
7. The method according to claim 6, characterized in that, The step of performing time-frequency transformation on the target audio and the target human ear perception domain audio respectively to obtain the second spectrum information includes: The target audio and the target human ear perception domain audio are transformed into a complex domain to obtain the time spectrum of the target audio in the complex domain and the time spectrum of the target human ear perception domain audio in the complex domain; Extract the amplitude spectrum and phase spectrum of the time spectrum; The amplitude spectrum and phase spectrum corresponding to the target audio, as well as the amplitude spectrum and phase spectrum corresponding to the target human ear perception domain audio, are stacked to obtain the second spectrum information.
8. The method according to claim 1, characterized in that, The method further includes: Obtain sample audio; The sample audio is processed by human hearing perception to obtain sample audio in the human hearing perception domain; The sample audio and the human ear perception domain sample audio are evaluated using an evaluation network to obtain the objective distortion quality evaluation loss and spatial quality loss of the sample audio. Based on the objective distortion quality assessment loss and the spatial quality loss, the network parameters of the evaluation network to be trained are adjusted to obtain the evaluation network; the evaluation network includes the objective distortion quality assessment subnetwork and the spatial quality assessment subnetwork.
9. The method according to claim 8, characterized in that, The acquisition of sample audio includes: Retrieve at least three preset audio tracks from the preset audio database; According to different preset signal-to-noise ratios, preset noise is added to the at least three preset audios to obtain at least three processed audios; From the at least three processed audio samples, identify the positive sample, the anchor sample with the highest signal-to-noise ratio, and the negative sample with the lowest signal-to-noise ratio; The anchor sample, the positive sample, and the negative sample are determined as the sample audio.
10. The method according to claim 9, characterized in that, The human ear perception domain sample audio includes: human ear perception domain anchor samples, human ear perception domain positive samples, and human ear perception domain negative samples. The evaluation network to be trained evaluates the sample audio and the human ear perception domain sample audio to obtain the objective distortion quality evaluation loss and spatial quality loss of the sample audio, including: The objective distortion quality evaluation subnetwork in the evaluation network to be trained is used to evaluate the anchor sample, the positive sample, and the negative sample to obtain the objective distortion quality evaluation loss of the sample audio. The spatial quality evaluation subnetwork in the evaluation network to be trained is used to evaluate the anchor samples of the human ear perception domain. Based on the evaluation and processing results of the anchor samples in the human ear perception domain and the evaluation and processing results of the anchor samples, the spatial quality loss is determined.
11. The method according to claim 10, characterized in that, The determination of the spatial quality loss based on the evaluation results of the anchor samples in the human ear perception domain and the evaluation results of the anchor samples includes: Based on the evaluation and processing results of the human ear perception domain anchor sample, predict the prediction probability distribution of the predicted azimuth angle of the human ear perception domain anchor sample within a preset angle interval set. Based on the true azimuth angle and the predicted azimuth angle of the evaluation processing results of the anchor sample, the included angle difference information is determined; Based on the true value distribution of the true azimuth angle in the preset angle interval set and the predicted probability distribution, the distribution difference information is determined; The spatial mass loss is determined based on the included angle difference information and the distribution difference information.
12. The method according to claim 8, characterized in that, The step of adjusting the network parameters of the evaluation network to be trained based on the objective distortion quality assessment loss and the spatial quality loss to obtain the evaluation network includes: The overall loss is determined based on the objective distortion quality assessment loss and the spatial quality loss. The network parameters of the evaluation network to be trained are adjusted using the overall loss to obtain the evaluation network.
13. An audio data evaluation device, characterized in that, The device includes: The acquisition module is used to acquire raw audio, which includes reference audio and target audio to be evaluated; The processing module is used to perform human ear perception processing on the original audio to obtain human ear perception domain audio, wherein the human ear perception domain audio includes reference human ear perception domain audio corresponding to the reference audio and target human ear perception domain audio corresponding to the target audio. The evaluation module is used to perform objective evaluation processing on the original audio at least using an objective distortion quality evaluation subnetwork to obtain multiple reference objective evaluation features corresponding to the reference audio and multiple target objective evaluation features corresponding to the target audio; and to perform spatial evaluation processing on the audio in the human ear perception domain using a spatial quality evaluation subnetwork, and to fuse the evaluation features obtained from the objective evaluation processing during the spatial evaluation processing to obtain multiple reference spatial evaluation features corresponding to the reference audio and multiple target spatial evaluation features corresponding to the target audio. The determination module is used to determine the difference information between the reference audio and the target audio based on the evaluation processing results, and to determine the evaluation result of the target audio based on the difference information.
14. An electronic device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 12.
15. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 12.
16. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 12.
Citation Information
Patent Citations
Three-dimensional (3D) audio quality objective evaluation method
CN102664017A