Audio processing method and apparatus
The audio processing method enhances model performance by optimizing training data selection using objective and subjective metrics, addressing issues of noise and label inaccuracies in real-world data to improve model efficiency.
Patent Information
- Application Number
- PCT/CN2024/107216
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-24
- Publication Date
- 2026-01-29
AI Technical Summary
Current audio processing technologies face suboptimal model performance due to the presence of undesirable factors in real-world speech data, such as excessive background noise, data missing, and inaccurate data labels, which are not effectively addressed by simple data pre-processing.
An audio processing method that optimizes training data selection by using a combination of objective evaluation metrics and a correlation vector associated with subjective metrics to enhance the quality of training data, ensuring only high-quality data is used for model training.
The method improves the quality of training data, leading to enhanced performance of audio processing models by accurately filtering and organizing training data, thereby improving model efficiency and effectiveness.
Smart Images

Figure CN2024107216_29012026_PF_FP_ABST
Abstract
Description
AUDIO PROCESSING METHOD AND APPARATUSTECHNICAL FIELD
[0001] The present disclosure relates to a field of audio processing, and in particular, to an audio processing method, an audio processing apparatus and device, and a computer-readable storage medium.BACKGROUND
[0002] Audio processing, as an important information processing technology, involves collection, processing, analysis, and utilization of audio signals. As Artificial Intelligence (AI) advances, neural networks such as deep learning networks have been widely applied in audio processing, resulting in a tremendous amount of audio processing models such as speech recognition models, speech synthesis models, and the like. High-quality training data is crucial for training audio processing models and improving performance of these models. However, current training audio data such as real-world speech data often contains many undesirable factors, leading to suboptimal model performance.
[0003] SUMMARY OF THE DISCLOSURE
[0004] The present disclosure proposes an audio processing method, an audio processing apparatus and device, and a computer-readable storage medium, which optimize training data using a plurality of objective evaluation metrics in combination with a correlation vector associated with a plurality of subjective metrics, thereby enhancing the quality of training data and improving model performance.
[0005] According to an aspect of the present disclosure, there is provided an audio processing method comprising: acquiring a set of audio data; for each audio data of the set of audio data, evaluating the audio data by using a plurality of objective metrics to generate an objective metric vector including a plurality of objective metric values; acquiring a correlation vector representing correlation between the plurality of objective metrics and a plurality of subjective metrics, and determining a weighting coefficient vector for each audio data based at least on the correlation vector; performing weighted summation on the plurality of objective metric values in the objective metric vector of each audio data by using the weighting coefficient vector to obtain a weighted sum as an evaluation score for the audio data; and screening, from the set of audio data, audio data with the evaluation score greater than a first predetermined threshold into a set of target audio data.
[0006] According to another aspect of the present disclosure, there is provided an audio processing method, comprising: receiving an audio signal; processing the audio signal using an audio processing model to generate a processed audio signal; and outputting the processed audio signal, wherein audio data for training the audio processing model is determined based on a set of target audio data set obtained by: acquiring a set of audio data; for each audio data of the set of audio data, evaluating the audio data by using a plurality of objective metrics to generate an objective metric vector including a plurality of objective metric values; acquiring a correlation vector representing correlation between the plurality of objective metrics and a plurality of subjective metrics, and determining a weighting coefficient vector for each audio data based at least on the correlation vector; performing weighted summation on the objective metric vector of each audio data by using the weighting coefficient vector to obtain a weighted sum as an evaluation score for the audio data; and screening, from the set of audio data, audio data with the evaluation score greater than a first predetermined threshold into a set of target audio data.
[0007] According to another aspect of the present disclosure, there is provided an audio processing apparatus, comprising: an acquisition unit configured to acquire a set of audio data; an evaluation unit configured to evaluate, for each audio data of the set of audio data, the audio data by using a plurality of objective metrics to generate an objective metric vector including a plurality of objective metric values; a weighted summation unit configured to acquire a correlation vector representing correlation between the plurality of objective metrics and a plurality of subjective metrics, determine a weighting coefficient vector for each audio data based at least on the correlation vector, and perform weighted summation on the objective metric vector of each audio data by using the weighting coefficient vector to obtain a weighted sum as an evaluation score for the audio data; and a screening unit configured to screen, from the set of audio data, audio data with the evaluation score greater than a first predetermined threshold into a set of target audio data.
[0008] According to another aspect of the present disclosure, there is provided a data loading system comprising: a multi-evaluation module configured to generate a set of target audio data by using the audio processing method as described above; a noise data acquisition module configured to acquire a set of noise data; a data selector configured to select a subset of target audio data from the set of target audio data and a subset of noise data from the set of noise data according to a predetermined rule; a simulation module configured to multiply the subset of target audio data and the subset of noise data by an impulse response function, respectively, to obtain a set of simulated audio data and a set of simulated noise data; and a mixer configured to mix the set of simulated audio data and the set of simulated noise data to obtain a set of training audio data.
[0009] According to another aspect of the present disclosure, there is provided an audio processing device, comprising: one or more processors; and one or more memories, wherein the one or more memories have stored therein computer-readable instructions that, when executed by the one or more processors, cause the one or more processors to execute the audio processing methods as described above.
[0010] According to another aspect of the present disclosure, there is provided a computer-readable storage medium having stored thereon computer-readable instructions which, when executed by a processor, cause the processor to execute the audio processing methods as described above.
[0011] According to another aspect of the present disclosure, there is provided a computer program product or computer program including computer-readable instructions, which, when executed by a processor, cause the processor to execute the audio processing methods as described above.BRIEF DESCRIPTION OF THE DRAWINGS
[0012] The above and other objects, features and advantages of embodiments of the present disclosure will become obvious from the following detailed description of embodiments of the present disclosure taken in conjunction with accompanying drawings. The accompanying drawings are used to provide further understanding of the embodiments of the present disclosure, constitute a part of the specification, explain the present disclosure together with the embodiments of the present disclosure, and do not constitute a limitation of the present disclosure. In the drawings, like reference numerals generally represent like components or steps.
[0013] FIG. 1 illustrates a framework of an audio processing method according to one or more embodiments of the present disclosure;
[0014] FIG. 2 illustrates a flow diagram of an audio processing method according to one or more embodiments of the present disclosure;
[0015] FIG. 3 illustrates a flow diagram of a method for determining a correlation vector according to one or more embodiments of the present disclosure;
[0016] FIG. 4 illustrates an example of an audio test room according to one or more embodiments of the present disclosure;
[0017] FIG. 5 illustrates a framework of a data loading system according to one or more embodiments of the present disclosure;
[0018] FIG. 6 illustrates a flow diagram of an audio processing method according to one or more embodiments of the present disclosure; and
[0019] FIG. 7 illustrates a schematic structural diagram of an audio processing apparatus according to one or more embodiments of the disclosure.
[0020] DESCRIPTION OF THE EMBODIMENTS
[0021] In order to make objectives, technical solutions and advantages of the embodiments of the present disclosure clearer, the technical solutions in the embodiments of the present disclosure will be described clearly and thoroughly with reference to the accompanying drawings. Obviously, these described embodiments are only a part of the present disclosure, not all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present disclosure without paying creative efforts fall into the protection scope of the present disclosure.
[0022] As used herein and in the claims, the words “a, ” “an, ” “an, ” and / or “the” do not refer to the singular, but may include the plural unless the context clearly dictates otherwise. In general, the terms “comprise” and “comprising” only imply the inclusion of steps and elements specifically identified, these steps and elements do not constitute an exclusive list and a method or apparatus may also contain other steps or elements.
[0023] As used herein and in the claims, audio data or audio signals may refer to any sound signals in any form collected by any sound acquisition device and may include speech signals, music signals, natural sound signals, machine-generated sound signals, and the like, which are not particularly limited in the present disclosure unless explicitly stated otherwise. As used herein and in the claims, a headphone may refer to an electro-acoustic transducer for converting electric signals into sounds, and may be or include a headset, an earphone, an earbud, or any other transducer devices, which is not particularly limited in the present disclosure unless explicitly stated otherwise.
[0024] Flowcharts are used herein to illustrate steps of a method according to one or more embodiments of the present disclosure. It should be understood that preceding or subsequent steps do not have to be performed exactly in order. Rather, various steps may be processed in reverse order or simultaneously, as desired. Meanwhile, other steps may also be added to the method, or certain step or steps may be removed from the method.
[0025] Audio processing technologies such as Automatic Speech Recognition (ASR) and Text-To-Speech (TTS) have developed rapidly and heavily rely on neural network models. Audio data used for training audio processing models may be, for example, real-world speech data, which usually contains undesirable factors like excessive background noise, data missing, distortion, inaccurate data labels, and the like, leading to suboptimal model performance. Generally, data pre-processing may be performed by using a suite of established quality metrics to filter training data, but simple data pre-processing has little effect on improving the quality of the training data.
[0026] To solve the above technical problems, the present disclosure provides an audio processing method that can optimize training data selection by using a plurality of objective evaluation metrics in combination with a correlation vector associated with a plurality of subjective metrics, thereby enhancing the quality of the training data and improving model performance. One or more embodiments of the present disclosure also provide a training data loading system that enhances the effectiveness of model training by ensuring that only the most relevant and high-quality training data is used.
[0027] FIG. 1 illustrates a framework 100 of an audio processing method according to one or more embodiments of the present disclosure. The framework 100 schematically shows an overall evaluation flow of the audio processing method that aims to obtain optimal target audio data used for training audio processing models. Specifically, candidate audio data 102 may be acquired, for example, from public audio databases or collected in strict experimental environments, and may be pre-processed, as described below. The candidate audio data 102 may be subjected to a series of processing 104 to obtain the target audio data 114. The processing 104 may include, for example, objective metric calculation 104_1 to obtain an objective metric vector of the candidate audio data, outlier processing 104_2 to label outliers of objective metric values of the objective metric vector, weighted evaluation 104_3 to obtain an evaluation score for the candidate audio data by jointly considering the outliers and a correlation vector 106 characterizing the correlation between objective metrics and subjective metrics, and training data selection 104_4 to select target audio data 114 based on the evaluation score. For example, candidate audio data with an evaluation score greater than a first predetermined threshold may be selected as target audio data 114, while candidate audio data with an evaluation score smaller than the first predetermined threshold may be collected as pseudo audio data 116 for further processing, such as data enhancement. In addition, to obtain the correlation vector 106, test audio data 108 may be collected in strict experimental environments, and undergo subject metric evaluation 110 and correlation analysis 112. Specific operations of the processing 104, 110, and 112 will be described in detail below.
[0028] FIG. 2 illustrates a flow diagram of an audio processing method 200 according to one or more embodiments of the present disclosure. The method 200 may be implemented by a computer or a server, which is not particularly limited in the present disclosure. As shown in FIG. 2, in step S202, a set of audio data is acquired. For example, the set of audio data may be acquired from public audio databases or collected in strict experimental environments, and in particular, may be speech data from human beings.
[0029] In one or more embodiments of the present disclosure, the audio data may be pre-processed because not all raw audio data is of sufficient quality. Some data at the beginning or end may have no signal or be in a convergence phase, some data may suffer from excessive background noise, and the accuracy of data labels is also a concern. The pre-processing of the raw audio data may involve format normalization and managing missing values to filter the data, and then the filtered data may be initially checked using a suite of established quality metrics, such as frequently used Signal-to-Noise Ration (SNR) , Peak SNR (PSNR) and the like. Audio data that conforms to the established quality metrics may be incorporated into a candidate audio database, from which the set of audio data may be acquired in step S202. On the other hand, audio data that fails the initial quality check may be refined, for example, by using a pre-trained, advanced audio data enhancement system and re-checked after the refinement.
[0030] In step S204, for each audio data of the set of audio data, the audio data may be evaluated by using a plurality of objective metrics to generate an objective metric vector including a plurality of objective metric values. Step S204 may correspond to the objective metric calculation 104_1 shown in FIG. 1. The plurality of objective metrics may be used to assess clarity, intelligibility, and the like of the audio data. In one or more embodiments of the present disclosure, the plurality of objective metrics may include, for example, at least one of Signal-to-Noise Ratio (SNR) , Peak Signal-to-Noise Ratio (PSNR) , Segmented Signal-to-Noise Ratio (SSNR) , Linear Prediction Coefficient (LPC) , Spectral Distance (SD) , Perceptual Evaluation of Speech Quality (PESQ) , Mean Opinion Score (MOS) based on neural networks (e.g., AutoMOS, QualityNet, MOSNet, etc. ) , Short-Time Objective Intelligibility (STOI) , Normalized Covariance Measure (NCM) or any other metrics that can evaluate audio data objectively. For each audio data x of the set of audio data, the objective metric vector M generated by using the plurality of objective metrics may be expressed as: M= [M1 (x) , M2 (x) …MN (x) ] (1)
[0031] where N is the number of the plurality of objective metrics.
[0032] In step S206, a correlation vector representing the correlation between the plurality of objective metrics and a plurality of subjective metrics may be obtained and then a weighting coefficient vector for each audio data may be determined based at least on the correlation vector. The plurality of subjective metrics may include, for example, at least one of Mean Opinion Score (MOS) , CrowdMOS (CMOS) , clarity of audio, naturalness of audio, a retention degree of background sound, Absolute Category Rating (ACR) , Degradation Category Rating (DCR) , Comparative Category Rating (CCR) , ABX Test or any other metrics that can evaluate audio data subjectively.
[0033] As noted above, the correlation vector β may be determined and stored in advance by subject metric evaluation and correlation analysis based on test audio data collected in strict experimental environments, such as an audio test room established according to ETSI (European Telecommunications Standards Institute) standards. Herein, the correlation vector β may be expressed as: β= [β1, β2…βN] (2)
[0034] where a correlation value βi (1 ≤ i ≤ N) of the correlation vector β may characterize the correlation of the ith objective metric with the plurality of subjective metrics. A specific method for determining the correlation vector will be described in further detail below.
[0035] Before determining the weighting coefficient vector, outlier processing of the plurality of objective metric values in the objective metric vector may be performed to eliminate significant outliers. Specifically, an average value μ and a standard deviation σ of the plurality of objective metric values in the objective metric vector may be calculated, and then an outlier degree for each objective metric value may be calculated as:
[0036] where Z is an outlier degree vector representing the outlier degree for each objective metric value, μ and σ are the average value and the standard deviation of the plurality of objective metric values, respectively.
[0037] Then, a validity vector α associated with the objective metric vector may be determined based on the outlier degree vector Z, and each element of the validity vector characterizes the validity of a corresponding objective metric value of the objective metric vector. For example, the validity vector α may be expressed as: α=[α1, α2, …αN] (4)
[0038] and
[0039] where αi is the ith element of the validity vector corresponding to the ith objective metric value of the objective metric vector, and 1 ≤ i ≤ N; T is a predetermined threshold to decide the validity of the objective metric value, which may be referred to as a second predetermined threshold to differentiate it from the first predetermined threshold used to screen target audio data; v1 is a first value, v2 is a second value and v1 < v2.
[0040] That is, if the ith outlier degree value is greater than or equal to the second predetermined threshold, the validity value corresponding to the ith objective metric value may be determined as the first value v1; otherwise if the ith outlier degree value is less than the second predetermined threshold, the validity value corresponding to the ith objective metric value may be determined as the second value v2. In one or more embodiments of the present disclosure, the second predetermined threshold T may be 3, or a larger or smaller value or any other value set per practical requirements, which is not particularly limited in the present disclosure. In one or more embodiments of the present disclosure, the first value v1 may be 0 and the second value v2 may be 1, or the first and second values may be any other values as long as the first value is less than the second value, which is not particularly limited in the present disclosure.
[0041] Aweighting coefficient vector W for each audio data may be determined based on the validity vector and the correlation vector. Basically, the higher the correlation value βi, the closer its corresponding ith objective metric value is to subjective metric values, and the more it conforms to actual auditory experience, and thus a greater weight should be allocated to the ith objective metric value. On the other hand, the validity value αi of the validity vector equal to the first value (e.g., 0) indicates that the ith objective metric value significantly deviates from other objective metric values, and thus should be allocated with small weight or even eliminated (if the first value is 0) . Therefore, the weighting coefficient vector W for each audio data may be expressed as: W= [α1β1, α2β2, …αNβN] (6)
[0042] Then, in step S208, weighted summation may be performed on the objective metric vector of each audio data by using the weighting coefficient vector to obtain a weighted sum as an evaluation score for the audio data. Specifically, the evaluation score S (x) for each audio data may be expressed as: S (x) =W·M=α1β1M1 (x) +α2β2M2 (x) +…+αNβNMN (x) (7)
[0043] The evaluation score for each of the set of audio data may be used as a criterion to decide whether the audio data is qualified target audio data. Specifically, in step S210, audio data with an evaluation score greater than the first predetermined threshold may be screened into a set of target audio data. The first predetermined threshold may be determined according to empirical parameters or distribution of the evaluation scores of the objective metric values, for example. Audio data with an evaluation score greater than or equal to the first predetermined threshold means that it is pure enough to serve as clean, target audio data for training audio processing models, such as AI noise reduction models. On the other hand, audio data with an evaluation score less than the first predetermined threshold means that it may still contain non-ignorable noise artifacts or other undesirable factors which may lead to suboptimal training results.
[0044] The specific method for determining the correlation vector will be described below in connection with FIG. 3. FIG. 3 illustrates a flow diagram of a method 300 for determining a correlation vector according to one or more embodiments of the present disclosure. In step S302, a set of test audio data including a plurality of test audio data is acquired. For example, the set of test audio data may be collected in an audio test room, such as an ETSI room established according to ETSI standards. ETSI room is commonly used in industries where noise suppression performance is measured and optimized. In general, the most important characteristic of the ETSI room is strong sound isolation from external sound sources to ensure reproducible results in the same environment, given identical test parameters.
[0045] In one or more embodiments of the present disclosure, the ETSI room may be established as shown in FIG. 4, which includes four speakers 402 and a head and torso simulator (HATS) 404 located in the middle of the speakers. In the ETSI room, reverberations may be maintained, for example, in a range of 300 ms to 700 ms, and a noise floor should meet ETSI standards, for example, being less than 30 dBA, which are not particularly limited in the present disclosure. In FIG. 4, by way of example and not limitation, the four speakers may be evenly distributed around the HATS. For example, a distance between two adjacent speakers may be 2.828 m and a distance between the HATS and each speaker may be 2 m, but the distances may be any other appropriate values. The four speakers may generate environment noise to create diffuse noise conditions in the room.
[0046] A set of test audio data may be recorded under various noise conditions, including a quiet scenario and scenarios with different levels of noise. In one or more embodiments of the present disclosure, the set of test audio data may be a result of collecting audio played in the audio test room by using a first sound collection device located in the audio test room. The first sound collection device may be, for example, a microphone, a headphone, a smartphone, a computer, or any other device integrated with a sound collection component. During the process of audio collection, both audio signals and noise signals (no noise signal in the quiet scenario) may be played in the audio test room, and the first sound collection device may record the played audio and noise signals as the test audio data.
[0047] Alternatively, in one or more embodiments of the present disclosure, the set of test audio data may be a result of collecting audio from the first sound collection device by using a second sound collection device located outside the audio test room and in communication with the first sound collection device located in the audio test room. The second sound collection device may be a communication device integrated with a sound collection component, for example, a smartphone or a computer. For example, the test audio data collected in this way may be suitable for determining a correlation vector for training data applied in an AI noise reduction model of a headphone, a smartphone, or a computer.
[0048] By way of example rather than limitation and with reference to the ETSI room shown in FIG. 4, the process of collecting test audio data used for an AI noise reduction model of a headphone or a smartphone may be: establishing a connection between a headphone (e.g., a Bluetooth headphone) and a first smartphone located in the ETSI room, and putting the headphone in an ear of the HATS; calling a second smartphone located outside the ETSI room; playing speech and noise files at the same time; opening the recording function of the second smartphone and setting it in a mute mode; recording and saving the audio data; performing post-recording processing such as format translation and audio trimming to obtain the test audio data.
[0049] To obtain the correlation vector, both objective evaluation and subjective evaluation will be performed for the test audio data, and their evaluation results will be combined and analyzed to derive a correlation between objective metrics and subjective perceptions, which may allow for more precise filtering and optimization of training data. The correlation analysis may be performed by using statistical methods such as correlation coefficient calculation, regression analysis, and the like.
[0050] Returning to FIG. 3, in Step 304, for each evaluator of a plurality of evaluators, the plurality of test audio data in the set of test audio data may be evaluated by the evaluator using the plurality of subjective metrics to determine a subjective metric matrix of the evaluator, and an overall three-dimensional subjective metric matrix will be obtained for the plurality of evaluators. In one or more embodiments of the present disclosure, assuming that there are K test audio data in the set of test audio data, a three-dimensional subjective metric matrix X for the kth test audio data (1 ≤ k ≤ K) obtained by G evaluators using U subjective metrics may be expressed as:
[0051] where K is the number of test audio data in the set of test audio data, 1 ≤ k ≤ K; G is the number of evaluators, 1 ≤ g ≤ G; and U is the number of the subjective metrics, 1 ≤ u ≤ U.
[0052] In order to balance the impact of different evaluators and different scoring systems, the three-dimensional subjective metric matrix X may be normalized and averaged first to obtain an averaged subjective metric matrix For ease of explanation, assuming that a uniform scoring system is used by the G evaluators for the U subjective metrics, meaning that the maximum values for the U subjective metrics are the same, the averaged subjective metric matrix may be expressed as:
[0053] where Smax represents the maximum value, which may be 5, 10, 100, or any other value.
[0054] It should be appreciated that if different scoring systems are used for the U subjective metrics, that is, the maximum values for the U subjective metrics are different, respective subjective metric values of the three-dimensional subjective metric matrix X may be normalized relative to corresponding maximum values, for example, being divided by 5, 10, 100, or any other value.
[0055] On the other hand, in step S306, the plurality of test audio data in the set of test audio data may be evaluated by using the plurality of objective metrics to determine an objective metric matrix. For K test audio data, the objective metric matrix Y obtained by using N objective metrics may be expressed as:
[0056] Then in step S308, the correlation vector β may be determined based on the three-dimensional subjective metric matrix X and the objective metric matrix Y. Specifically, for each of the set of test audio data, a correlation matrix characterizing the correlation between each objective metric value of the objective metric matrix for the test audio data and each averaged subjective metric value of the averaged subjective metric matrix for the test audio data may be determined. In other words, an overall three-dimensional correlation matrix which includes K two-dimensional correlation matrices each for one test audio data, may be obtained as :
[0057] Taking correlation coefficient calculation as an example of determining the correlation vector β, may be calculated as:
[0058] where represents a correlation efficient between the nth objective metric value and the uth subjective metric value for the kth test audio data.
[0059] To determine the coefficient vector β, first, correlation matrices for the plurality of test audio data in the set of test audio data may be averaged to obtain an averaged correlation matrix. In other words, the three-dimensional correlation matrix may be averaged relative to the K test audio data to obtain the averaged correlation matrix as:
[0060] Then, a plurality of vectors in the averaged correlation matrix respectively corresponding to the plurality of subjective metrics may be averaged to determine the correlation vector for each test audio data. In other words, the averaged correlation matrix may be further averaged relative to the U subjective metrics to determine the correlation vector β, with the ith correlation value βi equal to:
[0061] where 1 ≤ i ≤ N, and the function max () is used to eliminate undesirable negative correlation values.
[0062] The determined correlation vector β may be stored, for example, in a memory, and used to weight the plurality of objective metric values of the objective metric vector in conjunction with the validity vector to obtain a final evaluation score of the audio data, thereby realizing screening of the set of audio data based on the evaluation score. In one or more embodiments of the present disclosure, the correlation vector may be updated if the test audio data is updated or optimized, or based on practical training requirements.
[0063] Since the correlation vector may reflect the correlation between objective metrics and subjective metrics, the data screening may be more accurate and effective compared to purely objective or purely subjective evaluation-based data screening. The set of target audio data obtained by the data processing method 200 of the present disclosure may be clean enough to serve as target data for the training process of audio processing models such as an AI noise reduction model.
[0064] One or more embodiments of the present disclosure also provide a data loading system to facilitate data information organization, data management and retrieval, data evaluation and screening, training data construction and loading, and the like. FIG. 5 illustrates a framework of a data loading system 500 according to one or more embodiments of the present disclosure.
[0065] As shown in FIG. 5, a set of audio data 502 may be retrieved from a memory of the system 500 or a server, or received via an input interface. As previously discussed, the audio data 502 may undergo preliminary processing to ensure it meets basic quality standards before it is subjected to further analysis by the multi-evaluation module 504. The multi-evaluation module 504, in turn, may employ the audio processing method 200 as detailed with reference to FIG. 2 to screen out a set of target audio data, or to put it differently, qualified audio data 506. In one or more embodiments of the present disclosure, the data loading system may further include a data enhancement module (not illustrated in FIG. 5) to enhance quality of data that fails the screening of the multi-evaluation module 504.
[0066] On the other hand, a set of noise data 508 may be retrieved from a memory of the system or a server, or received via an input interface. As previously discussed, the noise data 508 may be pre-processed by a filter module 510. The filter module 510 may filter out bad noise data such as those being overly brief, having improper clipping, or suffering from significant packet loss during recording. After that, the noise data 508 may undergo classification by a classification module 512, which assigns each noise data with a specific noise type label. The noise type label may indicate, for example, an energy amplitude, an MOS score, a surrounding environment (e.g., outdoor or indoor) of the noise data, or the like, which may aid in subsequent training data construction and training processes.
[0067] An audio selector 516 may be configured to select a subset of target audio data from the set of target audio data (i.e., the qualified audio data 506) according to a predetermined rule. Likewise, a noise selector 518 may be configured to select a subset of noise data from the qualified, labeled noise data 514, or directly from the set of noise data 508 in one or more embodiments, according to the predetermined rule. The predetermined rule may be designed to ensure a balanced data selection, and may indicate, for example, a length of audio data, a number of audio data, characteristics of sound sources of audio data such as an age distribution, a gender distribution, language species, and the like. Besides, for the noise data selection, the predetermined rule may also indicate noise type labels and a number of noise data for each noise type label.
[0068] The subset of target audio data and the subset of noise data may be used to determine a set of training audio data used for training audio processing models such as an AI noise reduction model applied in a headphone or a smartphone. A simulation module 520 may be configured to apply an impulse response to the subset of target audio data and the subset of noise data to simulate real-world responses of a device such as a headphone or a smartphone. For example, if the training audio data is to be used in an AI noise reduction model for a headphone, the simulation module 520 may apply a headphone impulse response function in the subset of target audio data and the subset of noise data, respectively, to simulate headphone-related data. This process may be expressed as: Audio (t) simu=R*Audio (t) Noise (t) simu=R*Noise (t) (15)
[0069] where Audio (t) represents target audio data at time t, Noise (t) represents noise data at time t, R represents the headphone impulse response function, Audio (t) simu represents simulated audio data output by a headphone, and Noise (t) simu represents simulated noise data output by the headphone.
[0070] A mixer 522 may be configured to mix the set of simulated audio data and the set of simulated noise data to obtain a set of mixed audio data. In one or more embodiments of the present disclosure, the set of mixed audio data may be further enhanced to better simulate real data output by a device like a headphone. To achieve this, one or more of signal-to-noise ratio (SNR) scaling, random distortion processing, random filtering processing, and the like may be performed on the set of mixed audio data to determine a set of training audio data.
[0071] In one or more embodiments of the present disclosure, the system 500 may include an SNR scaling module (not illustrated) configured to adjust the energy of the audio data and the noise data to simulate various SNR scenarios, enhancing the realism of data simulation by controlling SNR across different SNR ranges. This may be essential for processing large-scale data and simulating real-world use cases.
[0072] In one or more embodiments of the present disclosure, the system 500 may further include a random distortion module (not illustrated) configured to introduce distortions such as reverberation distortion and random clipping distortion to the mixed audio data. For example, the reverberation distortion may be generated by a room-pulse response generator that convolves pulse signals with the mixed audio data and tailors the result signals with a predefined room size; the random clipping distortion may be generated by adding a small amount of clipping to the mixed audio data to enhance the robustness of models to be trained.
[0073] In one or more embodiments of the present disclosure, the system 500 may further include a random filter module (not illustrated) configured to further augment the mixed audio data. For example, the random filter module may use a variety of random filters applied in a controlled ratio, such as a random low-pass filter configured to simulate the absence of high-frequency components, and a random spectrum frequency mask filter used to obscure certain parts of audio spectrums, which may further enhance the robustness of models to be trained.
[0074] Finally, a set of training data pair 524 may be obtained, each training data pair including mixed audio data that serves as a training input of an audio processing model and corresponding target audio data that serves as a training target of the model.
[0075] The data loading system 500 according to one or more embodiments of the present disclosure may record detailed data information and its corresponding training details. The data information may include but is not limited to, data sources, processing steps, evaluation metric values, training parameters, and training outcomes, for example. With the data loading system 500, it can efficiently manage and retrieve different training data combinations, and thus enhance the effectiveness of model training. In one or more embodiments of the present disclosure, the data loading system 500 may be further optimized based on specific device channel information.
[0076] One or more embodiments of the present disclosure also provide an audio processing method that may process audio signals by using a trained audio processing model. FIG. 6 illustrates a flow diagram of an audio processing method 600 according to one or more embodiments of the present disclosure. The method 600 may be performed by an audio device such as a headphone, a smartphone, a computer, a smart wearable device, or any other device integrated with audio processing functions, which is not particularly limited in the present disclosure. In addition, since some steps of the method 600 are similar to the steps of the method 200, repetitive content will be omitted herein for the sake of brevity.
[0077] As shown in FIG. 6, in step S602, an audio signal may be received. The audio signal may be any sound signal to be processed, such as a speech signal generated during a conversation over a smartphone and captured by microphones within a headphone connected to the smartphone (e.g., via Bluetooth) .
[0078] In step S604, the audio signal may be processed using an audio processing model to generate a processed audio signal. In one or more embodiments of the present disclosure, the audio processing model may be any pre-trained model such as an AI noise reduction model, a speech enhancement model, and the like, which is not particularly limited in the present disclosure. For example, the audio processing model may be an AI noise reduction model applied in an audio device such as a headphone, a smartphone, a computer, and the like, and used to reduce noise existing in the received audio signal. Then, in step S606, the processed audio signal may be output, for example, to ears of a user or for further processing.
[0079] The audio processing model used in the method 600 may be trained by using audio data determined based on a set of target audio data. The set of target audio data may be obtained by steps S604_1, S604_2, S604_3, S604_4, and S604_5, which may correspond to steps S202 to S210 of the method 200 described above, respectively, and thus will not be repeatedly described herein.
[0080] An audio processing apparatus according to one or more embodiments of the present disclosure will be described below with reference to FIG. 7. FIG. 7 illustrates a schematic structural diagram of an audio processing apparatus 700 in accordance with one or more embodiments of the disclosure. As shown in FIG. 7, the audio processing apparatus 700 may include an acquisition unit 702, an evaluation unit 704, a weighted summation unit 706, and a screening unit 708. In addition to these four units, the audio processing apparatus 700 may further include other related components, but since these components are not relevant to the present disclosure, a detailed description thereof is omitted herein. In addition, since details of part of the functions of the audio processing apparatus 700 are similar to details of the steps of the audio processing method 200 as described with reference to FIG. 2, repeated descriptions of some content are omitted herein for brevity. The audio processing apparatus 700 may be a computer or a server, which is not particularly limited in the present disclosure.
[0081] The acquisition unit 702 may be configured to acquire a set of audio data. For example, the acquisition unit 702 may acquire the set of audio data from public audio databases or collected in strict experimental environments, and in particular, the set of audio data may be speech data from human beings.
[0082] In one or more embodiments of the present disclosure, the audio data may be pre-processed because not all raw audio data is of sufficient quality. Some data at the beginning or end may have no signal or be in a convergence phase, some data may suffer from excessive background noise, and the accuracy of data labels is also a concern. The pre-processing of the raw audio data may involve format normalization and managing missing values to filter the data, and then the filtered data may be initially checked using a suite of established quality metrics, such as frequently used Signal-to-Noise Ration (SNR) , Peak SNR (PSNR) and the like. Audio data that conforms to the established quality metrics may be incorporated into a candidate audio database, from which the set of audio data may be acquired. On the other hand, audio data that fails the initial quality check may be refined, for example, by using a pre-trained, advanced audio data enhancement system and re-checked after the refinement.
[0083] The evaluation unit 704 may be configured to evaluate, for each audio data of the set of audio data, the audio data by using a plurality of objective metrics to generate an objective metric vector including a plurality of objective metric values. The plurality of objective metrics may be used to assess clarity, intelligibility, and the like of the audio data. In one or more embodiments of the present disclosure, the plurality of objective metrics may include, for example, at least one of Signal-to-Noise Ratio (SNR) , Peak Signal-to-Noise Ratio (PSNR) , Segmented Signal-to-Noise Ratio (SSNR) , Linear Prediction Coefficient (LPC) , Spectral Distance (SD) , Perceptual Evaluation of Speech Quality (PESQ) , Mean Opinion Score (MOS) based on neural networks (e.g., AutoMOS, QualityNet, MOSNet, etc. ) , Short-Time Objective Intelligibility (STOI) , Normalized Covariance Measure (NCM) or any other metrics that can evaluate audio data objectively.
[0084] The weighted summation unit 706 may be configured to obtain a correlation vector representing the correlation between the plurality of objective metrics and a plurality of subjective metrics and then determine a weighting coefficient vector for each audio data based at least on the correlation vector. The plurality of subjective metrics may include, for example, at least one of Mean Opinion Score (MOS) , CrowdMOS (CMOS) , clarity of audio, naturalness of audio, a retention degree of background sound, Absolute Category Rating (ACR) , Degradation Category Rating (DCR) , Comparative Category Rating (CCR) , ABX Test or any other metrics that can evaluate audio data subjectively.
[0085] As described above, the correlation vector may be determined and stored in advance by subject metric evaluation and correlation analysis based on test audio data collected in strict experimental environments, such as an audio test room established according to ETSI (European Telecommunications Standards Institute) standards. In one or more embodiments of the present disclosure, the correlation vector may be determined according to Equations (8) -(14) as described above.
[0086] Before determining the weighting coefficient vector, the weighted summation unit 706 may perform outlier processing of the plurality of objective metric values in the objective metric vector to eliminate significant outliers. Specifically, the weighted summation unit 706 may calculate an average value and a standard deviation of the plurality of objective metric values in the objective metric vector, and then calculate an outlier degree for each objective metric value as, for example, Equation (3) . The weighted summation unit 706 may determine a validity vector associated with the objective metric vector based on the outlier degree vector, for example, according to Equations (4) - (5) , and each element of the validity vector characterizes validity of a corresponding objective metric value of the objective metric vector.
[0087] Then, the weighted summation unit 706 may determine a weighting coefficient vector for each audio data based on the validity vector and the correlation vector, for example, according to Equation (6) , and perform weighted summation on the objective metric vector of each audio data by using the weighting coefficient vector to obtain a weighted sum as an evaluation score for the audio data, for example, according to Equation (7) .
[0088] The evaluation score for each of the set of audio data may be used as a criterion to decide whether the audio data is qualified target audio data. Specifically, the screening unit 708 may be configured to screen audio data with the evaluation score greater than a first predetermined threshold into a set of target audio data. The first predetermined threshold may be determined according to empirical parameters or distribution of the evaluation scores of the objective metric values, for example. Audio data with an evaluation score greater than or equal to the first predetermined threshold means that it is pure enough to serve as clean, target audio data for training audio processing models, such as AI noise reduction models. On the other hand, audio data with an evaluation score less than the first predetermined threshold means that it may still contain non-ignorable noise artifacts or other undesirable factors which may lead to suboptimal training results.
[0089] The set of target audio data may be used to train audio processing models such as AI noise reduction models, speech enhancement models, and the like. Because the set of target audio data is screened using a plurality of objective metrics in combination with a correlation vector that considers the correlation between the plurality of objective metrics and a plurality of subjective metrics, training data combinations with the best quality may be accurately filtered and organized, thereby improving training efficiency and enhancing the performance of the trained models.
[0090] In one or more embodiments of the present disclosure, there is further provided an audio processing device comprising one or more processors and one or more memories, where the one or more memories have stored therein computer-readable instructions which, when executed by the one or more processors, cause the one or more processors to execute the audio processing methods as described in the various embodiments.
[0091] One or more embodiments of the present disclosure also provide a computer-readable storage medium having stored thereon computer-readable instructions which, when executed by a processor, cause the processor to execute the audio processing methods as described in the various embodiments. The computer-readable storage medium may include, but is not limited to, volatile memory and / or nonvolatile memory, for example. The volatile memory may include, for example, a Random Access Memory (RAM) , a cache, and the like. The nonvolatile memory may include, for example, a Read-Only Memory (ROM) , a hard disk, a flash memory, and the like.
[0092] One or more embodiments of the present disclosure also provide a computer program product or computer program including computer-readable instructions stored in a computer-readable storage medium. A processor of a computer may read the computer-readable instructions from the computer-readable storage medium, and the processor executes the computer-readable instructions so that the computer performs the audio processing methods as described in the various embodiments.
[0093] The following is a non-limiting list of examples that are in accordance with one or more techniques of this disclosure.
[0094] Example 1. An audio processing method comprising: acquiring a set of audio data; for each audio data of the set of audio data, evaluating the audio data by using a plurality of objective metrics to generate an objective metric vector including a plurality of objective metric values; acquiring a correlation vector representing correlation between the plurality of objective metrics and a plurality of subjective metrics, and determining a weighting coefficient vector for each audio data based at least on the correlation vector; performing weighted summation on the plurality of objective metric values in the objective metric vector of each audio data by using the weighting coefficient vector to obtain a weighted sum as an evaluation score for the audio data; and screening, from the set of audio data, audio data with the evaluation score greater than a first predetermined threshold into a set of target audio data.
[0095] Example 2. The method of Example 1, wherein determining the weighting coefficient vector for each audio data based at least on the correlation vector comprises: calculating an average value and a standard deviation of the plurality of objective metric values in the objective metric vector of the audio data; determining an outlier degree for each objective metric value of the objective metric vector based on the average value and the standard deviation; determining a validity vector associated with the objective metric vector based on the outlier degree for each objective metric value of the objective metric vector, wherein the validity vector characterizes validity of the plurality of objective metric values in the objective metric vector; and determining the weighting coefficient vector for the audio data based on the validity vector and the correlation vector.
[0096] Example 3. The method of any one of Examples 1-2, wherein determining the validity vector associated with the objective metric vector based on the outlier degree for each objective metric value of the objective metric vector comprises: in the validity vector, determining a validity value corresponding to an objective metric value with an outlier degree greater than or equal to a second predetermined threshold as a first value, and determining a validity value corresponding to an objective metric value with an outlier degree less than the second predetermined threshold as a second value, wherein the first value is less than the second value.
[0097] Example 4. The method of any one of Examples 1-3, wherein determining the weighting coefficient vector for the audio data based on the validity vector and the correlation vector comprises: multiplying validity values of the validity vector with corresponding correlation values of the correlation vector to determine the weighting coefficient vector for the audio data.
[0098] Example 5. The method of any one of Examples 1-4, wherein the correlation vector is determined by: acquiring a set of test audio data including a plurality of test audio data; for each evaluator of a plurality of evaluators, evaluating the plurality of test audio data in the set of test audio data by the evaluator using the plurality of subjective metrics to determine a subjective metric matrix of the evaluator; evaluating the plurality of test audio data in the set of test audio data by using the plurality of objective metrics to determine an objective metric matrix; and determining the correlation vector based on subjective metric matrices of the plurality of evaluators and the objective metric matrix.
[0099] Example 6. The method of any one of Examples 1-5, wherein determining the correlation vector based on the subjective metric matrices of the plurality of evaluators and the objective metric matrix comprises: normalizing and averaging the subjective metric matrices of the plurality of evaluators to determine an averaged subjective metric matrix; for each of the set of test audio data, calculating a correlation matrix characterizing correlation between each objective metric value of the objective metric matrix for the test audio data and each averaged subjective metric value of the averaged subjective metric matrix for the test audio data; and determining the correlation vector based on the correlation matrix for each of the set of test audio data.
[0100] Example 7. The method of any one of Examples 1-6, wherein determining the correlation vector based on the correlation matrix for each of the set of test audio data comprises: averaging correlation matrices for the plurality of test audio data in the set of test audio data to obtain an averaged correlation matrix; and averaging a plurality of vectors in the averaged correlation matrix respectively corresponding to the plurality of subjective metrics to determine the correlation vector.
[0101] Example 8. The method of any one of Examples 1-7, wherein determining the correlation vector based on the subjective metric matrices of the plurality of evaluators and the objective metric matrix comprises: performing regression analysis on the subjective metric matrices of the plurality of evaluators and the objective metric matrix to determine the correlation vector.
[0102] Example 9. The method of any one of Examples 1-8, wherein the set of test audio data is collected in an audio test room, the audio test room being a space isolated from external sound sources, and wherein the set of test audio data is a result of collecting audio played in the audio test room by using a first sound collection device located in the audio test room, or the set of test audio data is a result of collecting audio from the first sound collection device by using a second sound collection device located outside the audio test room and in communication with the first sound collection device located in the audio test room.
[0103] Example 10. The method of any one of Examples 1-9, wherein the first sound collection device is one of a microphone, a headphone, a smartphone, a computer, or other devices integrated with a sound collection component, and the second sound collection device is a communication device integrated with a sound collection component.
[0104] Example 11. The method of any one of Examples 1-10, further comprising: training an audio processing model by using the set of target audio data.
[0105] Example 12. The method of any one of Examples 1-11, wherein training the audio processing model by using the set of target audio data comprises: acquiring a set of noise data; selecting a subset of target audio data from the set of target audio data and a subset of noise data from the set of noise data according to a predetermined rule; determining a set of training audio data based at least on the subset of target audio data and the subset of noise data; and training the audio processing model by using the set of training audio data and the set of target audio data.
[0106] Example 13. The method of any one of Examples 1-12, wherein determining the set of training audio data based at least on the subset of target audio data and the subset of noise data comprises: multiplying the subset of target audio data and the subset of noise data by an impulse response function, respectively, to obtain a set of simulated audio data and a set of simulated noise data; mixing the set of simulated audio data and the set of simulated noise data to obtain a set of mixed audio data; and determining the set of training audio data based at least on the set of mixed audio data.
[0107] Example 14. The method of any one of Examples 1-13, wherein determining the set of training audio data based at least on the set of mixed audio data comprises: performing one or more of signal-to-noise ratio (SNR) scaling, random distortion processing and random filtering processing on the set of mixed audio data to determine the set of training audio data.
[0108] Example 15. The method of any one of Examples 1-14, wherein the plurality of objective metric values include at least one of Signal-to-Noise Ratio (SNR) , Peak Signal-to-Noise Ratio (PSNR) , Segmented Signal-to-Noise Ratio (SSNR) , Linear Prediction Coefficient (LPC) , Spectral Distance (SD) , Perceptual Evaluation of Speech Quality (PESQ) , Mean Opinion Score (MOS) based on neural networks, Short-Time Objective Intelligibility (STOI) and Normalized Covariance Measure (NCM) .
[0109] Example 16. The method of any one of Examples 1-15, wherein the plurality of subjective metrics include at least one of Mean Opinion Score (MOS) , CrowdMOS (CMOS) , clarity of audio, naturalness of audio, a retention degree of background sound, Absolute Category Rating (ACR) , Degradation Category Rating (DCR) , Comparative Category Rating (CCR) , and ABX Test.
[0110] Example 17. An audio processing method, comprising: receiving an audio signal; processing the audio signal using an audio processing model to generate a processed audio signal; and outputting the processed audio signal, wherein audio data for training the audio processing model is determined based on a set of target audio data set obtained by: acquiring a set of audio data; for each audio data of the set of audio data, evaluating the audio data by using a plurality of objective metrics to generate an objective metric vector including a plurality of objective metric values; acquiring a correlation vector representing correlation between the plurality of objective metrics and a plurality of subjective metrics, and determining a weighting coefficient vector for each audio data based at least on the correlation vector; performing weighted summation on the objective metric vector of each audio data by using the weighting coefficient vector to obtain a weighted sum as an evaluation score for the audio data; and screening, from the set of audio data, audio data with the evaluation score greater than a first predetermined threshold into a set of target audio data.
[0111] Example 18. An audio processing apparatus, comprising: an acquisition unit configured to acquire a set of audio data; an evaluation unit configured to evaluate, for each audio data of the set of audio data, the audio data by using a plurality of objective metrics to generate an objective metric vector including a plurality of objective metric values; a weighted summation unit configured to acquire a correlation vector representing correlation between the plurality of objective metrics and a plurality of subjective metrics, determine a weighting coefficient vector for each audio data based at least on the correlation vector, and perform weighted summation on the objective metric vector of each audio data by using the weighting coefficient vector to obtain a weighted sum as an evaluation score for the audio data; and a screening unit configured to screen, from the set of audio data, audio data with the evaluation score greater than a first predetermined threshold into a set of target audio data.
[0112] Example 19. A data loading system, comprising: a multi-evaluation module configured to generate a set of target audio data by using the audio processing method of any one of Examples 1-16; a noise data acquisition module configured to acquire a set of noise data; a data selector configured to select a subset of target audio data from the set of target audio data and a subset of noise data from the set of noise data according to a predetermined rule; a simulation module configured to multiply the subset of target audio data and the subset of noise data by an impulse response function, respectively, to obtain a set of simulated audio data and a set of simulated noise data; and a mixer configured to mix the set of simulated audio data and the set of simulated noise data to obtain a set of training audio data.
[0113] Example 20. An audio processing device, comprising: one or more processors; and one or more memories having stored therein computer-readable instructions which, when executed by the one or more processors, cause the one or more processors to execute the method of any one of Examples 1-17.
[0114] Example 21. A computer-readable storage medium having stored thereon computer-readable instructions which, when executed by a processor, cause the processor to execute the method of any one of Examples 1-17.
[0115] Example 22. A computer program product or computer program including computer-readable instructions, which, when executed by a processor, cause the processor to execute the method of any one of Examples 1-17.
[0116] It is to be recognized that depending on the examples, certain acts or events of any of the techniques described herein may be performed in a different sequence, may be added, merged, or left out altogether (e.g., not all described acts or events are necessary for the practice of the techniques) . Moreover, in certain examples, acts or events may be performed concurrently, e.g., through multi-threaded processing, interrupt processing, or multiple processors, rather than sequentially.
[0117] Program portions of the technology may be considered to be “product” or “article” that exists in the form of executable codes and / or related data, which are embodied or implemented by a computer-readable medium. A tangible, permanent storage medium may include an internal memory, or a storage used by computers, processors, or similar devices or associated modules. For example, various semiconductor memories, tape drivers, disk drivers, or any similar devices capable of providing storage functionality for software.
[0118] All software or parts of it may sometimes communicate over a network, such as the Internet or other communication networks. Such communication can load software from one computer device or processor to another. For example, loading from one server or host computer to a hardware environment of one computer environment, or other computer environment implementing the system, or a system having a similar function associated with providing information needed for the communication method. Therefore, another medium capable of transmitting software elements can also be used as a physical connection between local devices, such as light waves, electric waves, electromagnetic waves, etc., to be propagated through cables, optical cables, or air. A physical medium used for carrying the waves such as cables, wireless connections, or fiber optic cables may also be considered as a medium for carrying the software. In usage herein, unless a tangible “storage” medium is defined, other terms referring to a computer or machine “readable medium” mean a medium that participates in the execution of any instruction by the processor.
[0119] The present application uses specific words to describe embodiments of the present disclosure. Reference to “an embodiment, ” “one or more embodiments, ” and / or “some embodiments” means a feature, structure, or characteristic in connection with at least one embodiment of the present disclosure. Therefore, it should be emphasized and noted that two or more references to “an embodiment, ” “one embodiment, ” or “an alternative embodiment” in various places throughout this specification do not necessarily refer to the same embodiment. Furthermore, certain features, structures, or characteristics may be combined as suitable in one or more embodiments of the application.
[0120] Moreover, one skilled in the art will appreciate that aspects of the present disclosure may be illustrated and described in terms of a number of patentable categories or instances, including any new and useful process, machine, manufacture, or combination of matter, or any new and useful improvement thereof. Accordingly, aspects of the present disclosure may be performed entirely by hardware, entirely by software (including firmware, resident software, micro-code, etc. ) , or by a combination of hardware and software. The above hardware or software may each be referred to as a “data block, ” “module, ” “engine, ” “unit, ” “component, ” or “system. ” Furthermore, aspects of the present disclosure may be embodied as a computer product embodied in one or more computer-readable media including computer-readable program code.
[0121] Unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs. It will be further understood that terms, such as those defined in commonly used dictionaries, should be interpreted as having a meaning that is consistent with their meaning in the context of the relevant art and will not be interpreted in an idealized or extremely formal sense unless expressly so defined herein.
[0122] While various embodiments of the disclosure have been described, it will be apparent to those of ordinary skill in the art that many more embodiments and implementations are possible that are within the scope of the disclosure. Accordingly, the disclosure is not to be restricted except in light of the attached claims and their equivalents.
Claims
1.An audio processing method, comprising:acquiring a set of audio data;for each audio data of the set of audio data, evaluating the audio data by using a plurality of objective metrics to generate an objective metric vector including a plurality of objective metric values;acquiring a correlation vector representing correlation between the plurality of objective metrics and a plurality of subjective metrics, and determining a weighting coefficient vector for each audio data based at least on the correlation vector;performing weighted summation on the plurality of objective metric values in the objective metric vector of each audio data by using the weighting coefficient vector to obtain a weighted sum as an evaluation score for the audio data; andscreening, from the set of audio data, audio data with the evaluation score greater than a first predetermined threshold into a set of target audio data.2.The method of claim 1, wherein determining the weighting coefficient vector for each audio data based at least on the correlation vector comprises:calculating an average value and a standard deviation of the plurality of objective metric values in the objective metric vector of the audio data;determining an outlier degree for each objective metric value of the objective metric vector based on the average value and the standard deviation;determining a validity vector associated with the objective metric vector based on the outlier degree for each objective metric value of the objective metric vector, wherein the validity vector characterizes validity of the plurality of objective metric values in the objective metric vector; anddetermining the weighting coefficient vector for the audio data based on the validity vector and the correlation vector.3.The method of claim 2, wherein determining the validity vector associated with the objective metric vector based on the outlier degree for each objective metric value of the objective metric vector comprises:in the validity vector, determining a validity value corresponding to an objective metric value with an outlier degree greater than or equal to a second predetermined threshold as a first value, and determining a validity value corresponding to an objective metric value with an outlier degree less than the second predetermined threshold as a second value, wherein the first value is less than the second value.4.The method of claim 2, wherein determining the weighting coefficient vector for the audio data based on the validity vector and the correlation vector comprises:multiplying validity values of the validity vector with corresponding correlation values of the correlation vector to determine the weighting coefficient vector for the audio data.5.The method of claim 1, wherein the correlation vector is determined by:acquiring a set of test audio data including a plurality of test audio data;for each evaluator of a plurality of evaluators, evaluating the plurality of test audio data in the set of test audio data by the evaluator using the plurality of subjective metrics to determine a subjective metric matrix of the evaluator;evaluating the plurality of test audio data in the set of test audio data by using the plurality of objective metrics to determine an objective metric matrix; anddetermining the correlation vector based on subjective metric matrices of the plurality of evaluators and the objective metric matrix.6.The method of claim 5, wherein determining the correlation vector based on the subjective metric matrices of the plurality of evaluators and the objective metric matrix comprises:normalizing and averaging the subjective metric matrices of the plurality of evaluators to determine an averaged subjective metric matrix;for each of the set of test audio data, calculating a correlation matrix characterizing correlation between each objective metric value of the objective metric matrix for the test audio data and each averaged subjective metric value of the averaged subjective metric matrix for the test audio data; anddetermining the correlation vector based on the correlation matrix for each of the set of test audio data.7.The method of claim 6, wherein determining the correlation vector based on the correlation matrix for each of the set of test audio data comprises:averaging correlation matrices for the plurality of test audio data in the set of test audio data to obtain an averaged correlation matrix; andaveraging a plurality of vectors in the averaged correlation matrix respectively corresponding to the plurality of subjective metrics to determine the correlation vector.8.The method of claim 5, wherein determining the correlation vector based on the subjective metric matrices of the plurality of evaluators and the objective metric matrix comprises:performing regression analysis on the subjective metric matrices of the plurality of evaluators and the objective metric matrix to determine the correlation vector.9.The method of claim 5, wherein the set of test audio data is collected in an audio test room, the audio test room being a space isolated from external sound sources, and whereinthe set of test audio data is a result of collecting audio played in the audio test room by using a first sound collection device located in the audio test room, orthe set of test audio data is a result of collecting audio from the first sound collection device by using a second sound collection device located outside the audio test room and in communication with the first sound collection device located in the audio test room.10.The method of claim 9, wherein the first sound collection device is one of a microphone, a headphone, a smartphone, a computer, or other devices integrated with a sound collection component, and the second sound collection device is a communication device integrated with a sound collection component.11.The method of claim 1, further comprising:training an audio processing model by using the set of target audio data.12.The method of claim 11, wherein training the audio processing model by using the set of target audio data comprises:acquiring a set of noise data;selecting a subset of target audio data from the set of target audio data and a subset of noise data from the set of noise data according to a predetermined rule;determining a set of training audio data based at least on the subset of target audio data and the subset of noise data; andtraining the audio processing model by using the set of training audio data and the set of target audio data.13.The method of claim 12, wherein determining the set of training audio data based at least on the subset of target audio data and the subset of noise data comprises:multiplying the subset of target audio data and the subset of noise data by an impulse response function, respectively, to obtain a set of simulated audio data and a set of simulated noise data;mixing the set of simulated audio data and the set of simulated noise data to obtain a set of mixed audio data; anddetermining the set of training audio data based at least on the set of mixed audio data.14.The method of claim 13, wherein determining the set of training audio data based at least on the set of mixed audio data comprises:performing one or more of signal-to-noise ratio (SNR) scaling, random distortion processing, and random filtering processing on the set of mixed audio data to determine the set of training audio data.15.The method of claim 2, wherein the plurality of objective metric values include at least one of Signal-to-Noise Ratio (SNR) , Peak Signal-to-Noise Ratio (PSNR) , Segmented Signal-to-Noise Ratio (SSNR) , Linear Prediction Coefficient (LPC) , Spectral Distance (SD) , Perceptual Evaluation of Speech Quality (PESQ) , Mean Opinion Score (MOS) based on neural networks, Short-Time Objective Intelligibility (STOI) and Normalized Covariance Measure (NCM) .16.The method of claim 2, wherein the plurality of subjective metrics include at least one of Mean Opinion Score (MOS) , CrowdMOS (CMOS) , clarity of audio, naturalness of audio, a retention degree of background sound, Absolute Category Rating (ACR) , Degradation Category Rating (DCR) , Comparative Category Rating (CCR) , and ABX Test.17.An audio processing method, comprising:receiving an audio signal;processing the audio signal using an audio processing model to generate a processed audio signal; andoutputting the processed audio signal,wherein audio data for training the audio processing model is determined based on a set of target audio data set obtained by:acquiring a set of audio data;for each audio data of the set of audio data, evaluating the audio data by using a plurality of objective metrics to generate an objective metric vector including a plurality of objective metric values;acquiring a correlation vector representing correlation between the plurality of objective metrics and a plurality of subjective metrics, and determining a weighting coefficient vector for each audio data based at least on the correlation vector;performing weighted summation on the objective metric vector of each audio data by using the weighting coefficient vector to obtain a weighted sum as an evaluation score for the audio data; andscreening, from the set of audio data, audio data with the evaluation score greater than a first predetermined threshold into the set of target audio data.18.An audio processing apparatus, comprising:an acquisition unit configured to acquire a set of audio data;an evaluation unit configured to evaluate, for each audio data of the set of audio data, the audio data by using a plurality of objective metrics to generate an objective metric vector including a plurality of objective metric values;a weighted summation unit configured to acquire a correlation vector representing correlation between the plurality of objective metrics and a plurality of subjective metrics, determine a weighting coefficient vector for each audio data based at least on the correlation vector, and perform weighted summation on the objective metric vector of each audio data by using the weighting coefficient vector to obtain a weighted sum as an evaluation score for the audio data; anda screening unit configured to screen, from the set of audio data, audio data with the evaluation score greater than a first predetermined threshold into a set of target audio data.19.An audio processing device, comprising:one or more processors; andone or more memories having stored therein computer-readable instructions which, when executed by the one or more processors, cause the one or more processors to execute the method of any of claims 1-17.20.A computer-readable storage medium having stored thereon computer-readable instructions which, when executed by a processor, cause the processor to execute the method of any of claims 1-17.
Citation Information
Patent Citations
Apparatus and method for audio data analysis
US20220111294A1
Method for learning an audio quality metric combining labeled and unlabeled data
US20230245674A1