Audio quality evaluation method and device

By combining the input audio with N reference audio and their quality scores, the problem of retraining the model when evaluating the criteria in the prior art is solved, and the high universality and adaptability of the audio quality evaluation model is achieved.

CN119993210APending Publication Date: 2025-05-13WEBANK (CHINA)
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510131926.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-06
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

Existing audio quality evaluation methods require retraining the model when evaluating criteria changes, resulting in limited model universality.

Method used

By obtaining the input audio and N reference audio, the quality score for each reference audio is determined and these features are fused to generate fusion features, thereby evaluating the quality of the input audio. This method does not require retraining the model when evaluating criterion changes.

Benefits of technology

This improves the universality of the audio quality evaluation model, so that it does not need to retrain the model when evaluating criteria, reduces training costs and improves the adaptability of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119993210A_ABST
    Figure CN119993210A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides an audio quality evaluation method and device, which are applied to the technical field of machine learning, and the method comprises the steps: obtaining an input audio and N reference audios, the reference audios being determined based on audio samples in an audio quality data set and quality scores of the audio samples; determining a quality score corresponding to each reference audio according to the N reference audios; fusing the reference audio features of the N reference audios, the quality score features of the N quality scores and the input audio features of the input audio to obtain fused features; the fusion feature represents the incidence relation between the input audio and the reference audio and the quality score; and determining the audio quality of the input audio features according to the fusion features. According to the method and the device, when the evaluation standard changes, the audio quality evaluation model does not need to be retrained, so that the universality of the audio quality evaluation model is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present invention relate to the technical field of machine learning, and in particular to an audio quality assessment method and device. Background Art

[0002] In the wave of artificial intelligence generated content (AIGC), the generation of images, audio, and video through AI has become the norm in mainstream media. The audio and video uploaded by users to various platforms may be generated by AI, such as the audio and video uploaded by users of short video platforms. However, due to the uneven quality of AI technologies used in the current technological development process, there are certain differences in the quality of the generated audio and video. In the process of generating audio, it is necessary to perform automated testing on the generated audio to evaluate whether the quality of the generated audio is normal.

[0003] Existing audio quality assessment methods are divided into two categories, one is reference-based method and the other is reference-free method. Reference-based methods often require the provision of clean original audio as a reference benchmark to evaluate the audio to be tested. However, the same clean original sound as the evaluation sample is often difficult to obtain in actual scenarios; while reference-free methods require a large number of data sets to be used to learn the evaluation criteria of a large number of data sets through models in order to evaluate new samples. However, when the evaluation criteria change, it is necessary to retrain on the corresponding data set with the same evaluation criteria so that the model can adapt to the new evaluation criteria. At this time, the cost of model training is huge and the versatility of the model is limited. Summary of the invention

[0004] The embodiments of the present application provide an audio quality assessment method and device for determining the audio quality of input audio. When the assessment criteria change, there is no need to retrain the audio quality assessment model, thereby improving the versatility of the audio quality assessment model.

[0005] In a first aspect, the present application provides an audio quality assessment method, comprising:

[0006] Obtain input audio and N reference audios, where the reference audios are determined based on audio samples in an audio quality dataset and quality scores of the audio samples;

[0007] Determine, according to the N reference audios, a quality score corresponding to each reference audio;

[0008] The reference audio features of the N reference audios, the quality score features of the N quality scores, and the input audio features of the input audio are fused to obtain fused features, wherein the fused features represent the correlation relationship between the input audio, the reference audio, and the quality score;

[0009] The audio quality of the input audio feature is determined according to the fused feature.

[0010] For the quality assessment of the input audio in the target scenario, N reference audios and the corresponding N quality scores constitute a reference sequence of the input audio, and the quality score of the input audio is estimated. Since the reference sequence contains the relative relationship between the reference audio and the quality score in the reference sequence, combined with the relative relationship between the samples and the quality score in the input audio and reference sequence, the audio quality score estimate of the input audio under the quality assessment standard can be obtained. If the quality assessment standard needs to be changed, the reference sequence can be redefined without replacing the model.

[0011] Optionally, fusing the reference audio features of the N reference audios, the quality score features of the N quality scores, and the input audio features of the input audio to obtain the fused features includes:

[0012] The reference audio features of the N reference audios and the quality score features of the N quality scores are merged to obtain reference sequence related features; the reference sequence related features represent the correlation relationship between the reference audios and the quality scores;

[0013] The reference sequence related features and the input audio features of the input audio are fused to obtain the fused features.

[0014] Optionally, the fusing the reference audio features of the N reference audios and the quality score features of the N quality scores to obtain reference sequence related features includes:

[0015] Using a cross attention mechanism to fuse the reference audio features of the N reference audios and the quality score features of the N quality scores to obtain a first fused feature;

[0016] The first fusion feature is extracted through a self-attention mechanism to obtain N reference sequence related features.

[0017] Optionally, the audio quality dataset is composed of the quality scores of the audio samples in different dimensions;

[0018] The reference audio is determined based on audio samples in the audio quality dataset and quality scores of the audio samples, including:

[0019] For any audio sample, other audio samples in any dimension and corresponding quality scores are selected as reference audio of the audio sample.

[0020] Optionally, a normalization operation is performed on each quality score in the reference audio under any audio sample to obtain an optimized quality score under each audio sample;

[0021] The mass score is replaced by the optimized mass score.

[0022] Optionally, a gradient optimization algorithm is used to train the audio quality assessment model, specifically the following formula:

[0023]

[0024] Where L is the loss function, is the predicted audio quality score, and y is the actual audio quality score.

[0025] In a second aspect, an embodiment of the present application provides an audio quality assessment device, comprising:

[0026] An acquisition module, configured to acquire input audio and N reference audios, wherein the reference audios are determined based on audio samples in an audio quality dataset and quality scores of the audio samples;

[0027] A processing module, configured to determine a quality score corresponding to each reference audio according to the N reference audios;

[0028] The processing module is further configured to fuse the reference audio features of the N reference audios, the quality score features of the N quality scores, and the input audio features of the input audio to obtain a fused feature; the fused feature represents the correlation between the input audio, the reference audio, and the quality score;

[0029] The processing module is further used to determine the audio quality of the input audio feature according to the fusion feature.

[0030] For the quality assessment of the input audio in the target scenario, N reference audios and the corresponding N quality scores constitute a reference sequence of the input audio, and the quality score of the input audio is estimated. Since the reference sequence contains the relative relationship between the reference audio and the quality score in the reference sequence, combined with the relative relationship between the samples and the quality score in the input audio and reference sequence, the audio quality score estimate of the input audio under the quality assessment standard can be obtained. If the quality assessment standard needs to be changed, the reference sequence can be redefined without replacing the model.

[0031] Optionally, the processing module is specifically used for:

[0032] The fusing the reference audio features of the N reference audios, the quality score features of the N quality scores, and the input audio features of the input audio to obtain the fused features includes:

[0033] The reference audio features of the N reference audios and the quality score features of the N quality scores are merged to obtain reference sequence related features; the reference sequence related features represent the correlation relationship between the reference audios and the quality scores;

[0034] The reference sequence related features and the input audio features of the input audio are fused to obtain the fused features.

[0035] Optionally, the processing module is specifically used for:

[0036] The fusing the reference audio features of the N reference audios and the quality score features of the N quality scores to obtain reference sequence related features includes:

[0037] Using a cross attention mechanism to fuse the reference audio features of the N reference audios and the quality score features of the N quality scores to obtain a first fused feature;

[0038] The first fusion feature is extracted through a self-attention mechanism to obtain N reference sequence related features.

[0039] Optionally, the processing module is specifically used for:

[0040] The audio quality dataset is composed of the quality scores of the audio samples in different dimensions;

[0041] The reference audio is determined based on audio samples in the audio quality dataset and quality scores of the audio samples, including:

[0042] For any audio sample, other audio samples in any dimension and corresponding quality scores are selected as reference audio of the audio sample.

[0043] Optionally, the processing module is specifically used for:

[0044] Processing each quality score in the reference audio under any audio sample by a normalization operation to obtain an optimized quality score under each audio sample;

[0045] The mass score is replaced by the optimized mass score.

[0046] Optionally, the processing module is specifically used for:

[0047] The audio quality assessment model is trained using a gradient optimization algorithm, specifically as follows:

[0048]

[0049] Where L is the loss function, is the predicted audio quality score, and y is the actual audio quality score.

[0050] In a third aspect, an embodiment of the present application provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes any of the audio quality assessment methods described in the first aspect.

[0051] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium storing a computer program executable by a computer device. When the program is run on the computer device, the computer device executes any of the audio quality assessment methods described in the first aspect.

[0052] In a fifth aspect, an embodiment of the present application provides a computer program product, which includes a computer program / instruction, and when the computer program / instruction is executed by a processor, it implements any audio quality assessment method described in the first aspect above. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.

[0054] Figure 1 A flowchart of an audio quality assessment method provided in an embodiment of the present application;

[0055] Figure 2 A schematic diagram of the structure of an audio quality assessment model provided in an embodiment of the present application;

[0056] Figure 3 A schematic diagram of a feature fusion module structure provided in an embodiment of the present application;

[0057] Figure 4 A schematic diagram of an audio quality assessment device provided in an embodiment of the present application;

[0058] Figure 5 A schematic diagram of the structure of a computer device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0059] In order to make the purpose, technical solution and beneficial effects of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0060] Text to Speech (TTS): also known as speech synthesis, uses artificial intelligence to convert text into speech and written content into natural and realistic speech.

[0061] In the prior art, users convert a text into corresponding audio by using text-to-speech technology, and input the generated audio into a trained quality score estimation module to determine the quality of the audio. If the quality score estimation module is trained from a large number of scene data sets collected under three evaluation indicators: audio noise, tone change, and clear pronunciation. Once the evaluation indicator is changed, it is necessary to re-collect the data set under the new evaluation indicator and retrain the quality score estimation module. Therefore, it will cost a lot to train the quality score estimation module.

[0062] Therefore, if Figure 1 As shown, this is an audio quality assessment method proposed in this application, which adaptively assesses audio quality based on a small number of samples in the current scene without retraining the model. Specifically, it includes the following steps:

[0063] Step S101: obtaining input audio and N reference audios, where the reference audios are determined based on audio samples in an audio quality dataset and quality scores of the audio samples.

[0064] Step S102: determining a quality score corresponding to each reference audio according to N reference audios.

[0065] Step S103, the reference audio features of the N reference audios, the quality score features of the N quality scores, and the input audio features of the input audio are fused to obtain fused features; the fused features represent the correlation between the input audio, the reference audio, and the quality score.

[0066] Step S104: determining the audio quality of the input audio feature according to the fused feature.

[0067] For the quality assessment of the input audio in the target scenario, N reference audios and the corresponding N quality scores constitute a reference sequence of the input audio, and the quality score of the input audio is estimated. Since the reference sequence contains the relative relationship between the reference audio and the quality score in the reference sequence, combined with the relative relationship between the samples and the quality score in the input audio and reference sequence, the audio quality score estimate of the input audio under the quality assessment standard can be obtained. If the quality assessment standard needs to be changed, the reference sequence can be redefined without replacing the model.

[0068] Among them, the audio quality data set is composed of the quality scores of audio samples in different dimensions; the reference audio is determined based on the audio samples and the quality scores of the audio samples in the audio quality data set, including: for any audio sample, other audio samples and corresponding quality scores in any dimension are selected as the reference audio of the audio sample.

[0069] Specifically, for an audio, it can be evaluated from multiple dimensions, which may include noise intensity, clarity, fullness, stereo, timbre expression, and emotional expression ability. The quality scores of the same audio in different dimensions may be the same or different. For example, there is an audio sample 1, and the audio sample 1 is scored from the noise intensity of the audio, and the quality score is 0.2; the audio sample 1 is scored from the clarity of the audio pronunciation, and the quality score is 0.8; the audio sample 1 is scored from the strength of the audio tone, and the quality score is 0.4. Therefore, the present application constructs an audio quality data set according to different dimensions. For example, suppose there are 3 audio quality data sets, namely, an audio noise data set (the score label is the noise intensity of the audio), an audio clarity data set (the score label is the clarity of audio pronunciation), and an audio tone data set (the score label is the strength of the tone). The three data sets score n audios from the dimensions of noise, clarity of pronunciation, and tone. The audio noise dataset consists of n audio samples and the quality scores obtained by scoring the n audio samples based on the noise intensity of the audio. The audio noise dataset = {(audio 1, 0.2), (audio 2, 0.7), (audio 3, 0.5), (audio 4, 0.1) ... (audio n, 0.8)}. Similarly, the audio clarity dataset = {(audio 1, 0.8), (audio 2, 0.3), (audio 3, 0.5), (audio 4, 0.9) ... (audio n, 0.2)}; the audio tone dataset = {(audio 1, 0.3), (audio 2, 0.4), (audio 3, 0.5), (audio 4, 0.1) ... (audio n, 0.4)}.

[0070] In the present application, based on the above-mentioned audio quality data set, it is also necessary to select other audio samples and corresponding quality scores under any dimension for any audio sample in the audio quality data set as the reference audio of the audio sample, so as to jointly construct an audio sample. For example, there are n audio samples: audio sample 1, audio sample 2, audio sample 3...audio sample n, and each audio sample has a corresponding quality score. Take audio sample 1 as an example for illustration. If audio sample 1 comes from an audio noise data set, the quality score of audio sample 1 in the audio noise data set is 0.2. From the audio noise data set = {(audio 1, 0.2), (audio 2, 0.7), (audio 3, 0.5), (audio 4, 0.1)...(audio n, 0.8)}, randomly select n audios and their corresponding quality scores to form a reference sequence = [(audio 1, 0.2), (audio 2, 0.7), (audio 3, 0.5), (audio 4, 0.1)...(audio n, 0.8)]. In summary, audio sample 1, the quality score of audio sample 1 of 0.2, and the reference sequence of audio sample 1 together constitute a data sample. The same method is used to construct corresponding reference sequences for audio sample 2 with a quality score of 0.5 from the audio clarity dataset and audio sample 3 with a quality score of 0.8 from the audio pitch dataset, and finally constitute the data samples shown in Table 1 below.

[0071] Table 1

[0072] Serial number Audio Quality score Reference sequence 1 Audio Sample 1 0.2 [(audio 1, 0.2), (audio 3, 0.7), …, (audio n, 0.8)] 2 Audio Sample 2 0.5 [(audio 1, 0.8), (audio 2, 0.3), …, (audio n, 0.2)] 3 Audio Sample 3 0.8 [(audio 1, 0.3), (audio 2, 0.4), ..., (audio n, 0.4)]

[0073] Since audio and quality scores correspond one-to-one, and the quality score expresses the relative relationship under a certain evaluation standard, if the audio sample and the reference sequence are input into the audio quality assessment model, the audio quality assessment model cannot effectively obtain the impact of the reference sequence on the quality score. The audio quality assessment model only captures the relationship between the input audio and the reference sequence.

[0074] Therefore, the present application proposes a sample construction strategy to establish the relationship between the audio in the input audio quality assessment model and the reference sequence and the quality score output by the audio quality assessment model, so that the output quality score expresses the relative relationship under a certain evaluation standard. Therefore, when the evaluation standard changes, the quality score in the reference sequence will also change accordingly. Then, the new reference sequence is input into the audio quality assessment model to obtain the value of the quality score of the input audio under the new evaluation standard, so that it can be associated with the reference sequence and can correctly express the relative relationship between the quality scores.

[0075] In another possible embodiment, in order to improve the generalization ability of the audio quality assessment model, a normalization operation is used to process each quality score in the reference audio under any audio sample to obtain an optimized quality score under each audio sample; and the quality score is replaced with the optimized quality score. For example, for an audio sample A, if its sample quality score in the audio noise dataset is g, the quality score of the reference sequence selected for the audio sample A is [s1, s2, s3…s n ], then the sample score S of the currently selected audio sample A is recorded as S = [g, s1, s2, s3…s n ], the relationship between the sample quality score g and the reference sequence quality score S can be established through normalization operation, and the following two groups of reference sequence optimization quality scores S1 and S2 can be obtained:

[0076] S1=S / max(S)=[g / max(S),s1 / max(S),s2 / max(S),…s n / max(S)];

[0077] You can continue to invert the reference sequence scores to obtain another set of reference sequence optimization quality score values:

[0078] S2=1-S1=1-S / max(S)=[1-g / max(S),1-s1 / max(S),1-s2 / max(S),…1-s n / max(S)];

[0079] The optimized quality score values ​​of the two groups of reference sequences S1 and S2 represent the quality score results under different evaluation criteria. For example, S1 represents the evaluation criterion that the greater the noise, the greater the quality score value, and S2 represents the evaluation criterion that the smaller the noise, the greater the quality score value. On the basis of Table 1, all the quality score values ​​in Table 1 are replaced by the above transformations of S→S1 and S→S2, so as to obtain the optimized quality score values ​​of the reference sequence. New sample data can be generated without changing the relative relationship of the quality scores, so that the quality score values ​​of the samples change with the changes of the reference sequence. Taking audio sample 1 as an example, through the above transformations of S→S1 and S→S2, audio sample 1 can obtain 2 new quality scores and reference sequences, and audio sample 2 and audio sample 3 in the above Table 1 can obtain 2 new quality scores and reference sequences respectively. Through the sample construction strategy proposed in this embodiment, 6 audio sample data can be obtained from the 3 original audio sample data in Table 1, which expands the training data for training the audio quality assessment model and improves the generalization ability of the audio quality assessment model.

[0080] After constructing the audio samples for training the audio quality assessment model, an audio quality assessment model provided by the embodiment of the present application is as follows: Figure 2As shown, it specifically includes an audio feature encoding module 201, a reference sequence learning module 202, a feature fusion module 203 and a quality score estimation module 204.

[0081] exist Figure 2 Where x is the input audio, [r1, r2…r n ] is the audio sample of the reference sequence, [s1, s2…s n ] is the quality score corresponding to the audio sample of the reference sequence. Input audio x and audio samples of the reference sequence [r1, r2…r n ] The audio feature encoding module 201 obtains the input audio feature F0 and the audio sample features of the reference sequence [F1, F2…F n ]. The audio feature encoding module 201 may use an autoregressive module, such as a wav2vec network. The quality scores [s1, s2, ..., s n ] is expanded into the audio sample features [F1, F2…F n ]Quality score features of the same dimension [V1, V2…V n ]. The audio sample features of the reference sequence [F1, F2…F n ] and quality score features [V1, V2…V n ] together constitute the reference sequence features. The reference sequence features are input into the reference sequence learning module 202 to extract the quality assessment features F r , the quality assessment feature F r The input audio feature F0 is input into the feature fusion module 203 for feature fusion to obtain the fusion feature F fusion . The fusion feature F fusion The input quality score estimation module 204 obtains a predicted quality score of the input audio, wherein the quality score estimation module 204 is composed of a regression model. The audio quality assessment model takes the input audio and the reference sequence as inputs and jointly predicts the audio quality score of the input audio.

[0082] Since the input audio has different quality scores under different quality assessment standards, this embodiment is designed to take the reference sequence as input and extract quality assessment features according to the reference sequence. The audio quality assessment model needs to fully understand the assessment standard represented by the reference sequence, thereby affecting the subsequent quality score estimation of the input audio.

[0083] In this application, the feature fusion module is a hierarchical feature fusion module based on the attention mechanism, which can fully extract the information of the reference sequence based on the hierarchical mechanism of cross attention, self attention and cross attention. Figure 3The feature fusion module structure diagram shown in FIG. 1 is a fusion module that combines the extracted reference sequence features with the input audio feature F0 and processes the feature fusion module to obtain a fusion feature F containing the input audio feature and the reference sequence information. fusion .

[0084] The cross-attention mechanism is used to fuse the reference audio features of N reference audios and the quality score features of N quality scores to obtain the first fused feature; the self-attention mechanism is used to extract the first fused feature to obtain N reference sequence related features. The reference audio features of N reference audios and the quality score features of N quality scores are fused to obtain the reference sequence related features; the reference sequence related features represent the correlation between the reference audio and the quality score; the reference sequence related features and the input audio features of the input audio are fused to obtain the fused features.

[0085] Specifically, by Figure 2 It can be seen that the reference sequence features include the audio sample features of the reference sequence [F1, F2…F n ] and quality score features [V1, V2…V n ], since the two belong to different domains, the cross attention module can be used to fuse F and V to obtain the first fusion feature A, F n , V n →A n ; A=[A1、A2…A n ], the correlation between multiple samples and quality scores in the reference sequence can be extracted by autonomously extracting the first fusion feature A, and after extraction, F p . F p The correlation of the reference sequence is obtained, which characterizes the relationship between samples and quality scores in the reference sequence, namely the reference sequence feature. Then, the relationship between the input audio and the related features of the reference sequence is further explored through cross attention to obtain the fusion feature F fusion . The fusion feature F fusion The input quality score estimation module 204 obtains a predicted quality score for the input audio.

[0086] The following is a detailed description of the process of training the audio quality assessment model in this application. For each input audio sample x, a reference sequence [(r1, s1), (r2, s2), …, (r3, s3)] in the same data set is selected, and the audio sample x and the reference sequence are input into Figure 2 In the model, we get the prediction score The actual quality score is y. This application defines the loss function as L, and uses a gradient optimization algorithm to train the audio quality assessment model, as shown in the following formula:

[0087]

[0088] In the above formula, L is the loss function, is the predicted audio quality score, y is the real audio quality score. The gradient optimization algorithm can use adam (adaptive moment estimation) to optimize the weights of the audio quality assessment model.

[0089] For the quality assessment of the input audio in the target scenario, N reference audios and the corresponding N quality scores constitute a reference sequence of the input audio, and the quality score of the input audio is estimated. Since the reference sequence contains the relative relationship between the reference audio and the quality score in the reference sequence, combined with the relative relationship between the samples and the quality score in the input audio and reference sequence, the audio quality score estimate of the input audio under the quality assessment standard can be obtained. If the quality assessment standard needs to be changed, the reference sequence can be redefined without replacing the model.

[0090] like Figure 4 As shown, an audio quality assessment device 400 is provided for an embodiment of the present application, including:

[0091] An acquisition module 401 is used to acquire input audio and N reference audios, where the reference audios are determined based on audio samples in an audio quality dataset and quality scores of the audio samples;

[0092] The processing module 402 is used to determine a quality score corresponding to each reference audio according to the N reference audios;

[0093] The processing module 402 is further configured to fuse the reference audio features of the N reference audios, the quality score features of the N quality scores, and the input audio features of the input audio to obtain a fused feature; the fused feature represents the correlation between the input audio, the reference audio, and the quality score;

[0094] The processing module 402 is further configured to determine the audio quality of the input audio feature according to the fusion feature.

[0095] For the quality assessment of the input audio in the target scenario, N reference audios and the corresponding N quality scores constitute a reference sequence of the input audio, and the quality score of the input audio is estimated. Since the reference sequence contains the relative relationship between the reference audio and the quality score in the reference sequence, combined with the relative relationship between the samples and the quality score in the input audio and reference sequence, the audio quality score estimate of the input audio under the quality assessment standard can be obtained. If the quality assessment standard needs to be changed, the reference sequence can be redefined without replacing the model.

[0096] Optionally, the processing module 402 is specifically configured to:

[0097] The fusing the reference audio features of the N reference audios, the quality score features of the N quality scores, and the input audio features of the input audio to obtain the fused features includes:

[0098] The reference audio features of the N reference audios and the quality score features of the N quality scores are merged to obtain reference sequence related features; the reference sequence related features represent the correlation relationship between the reference audios and the quality scores;

[0099] The reference sequence related features and the input audio features of the input audio are fused to obtain the fused features.

[0100] Optionally, the processing module 402 is specifically configured to:

[0101] The fusing the reference audio features of the N reference audios and the quality score features of the N quality scores to obtain reference sequence related features includes:

[0102] Using a cross attention mechanism to fuse the reference audio features of the N reference audios and the quality score features of the N quality scores to obtain a first fused feature;

[0103] The first fusion feature is extracted through a self-attention mechanism to obtain N reference sequence related features.

[0104] Optionally, the processing module 402 is specifically configured to:

[0105] The audio quality dataset is composed of the quality scores of the audio samples in different dimensions;

[0106] The reference audio is determined based on audio samples in the audio quality dataset and quality scores of the audio samples, including:

[0107] For any audio sample, other audio samples in any dimension and corresponding quality scores are selected as reference audio of the audio sample.

[0108] Optionally, the processing module 402 is specifically configured to:

[0109] Processing each quality score in the reference audio under any audio sample by a normalization operation to obtain an optimized quality score under each audio sample;

[0110] The mass score is replaced by the optimized mass score.

[0111] Optionally, the processing module 402 is specifically configured to:

[0112] The audio quality assessment model is trained using a gradient optimization algorithm, specifically as follows:

[0113]

[0114] Where L is the loss function, is the predicted audio quality score, and y is the actual audio quality score.

[0115] Based on the same technical concept, the embodiment of the present application provides a computer device, which can be Figure 5 The terminal device 501 and / or the server 502 shown in FIG. Figure 5 As shown, it includes at least one processor 501 and a memory 502 connected to the at least one processor. The specific connection medium between the processor 501 and the memory 502 is not limited in the embodiment of the present application. Figure 5 For example, the processor 501 and the memory 502 are connected via a bus. The bus can be divided into an address bus, a data bus, a control bus, etc.

[0116] In the embodiment of the present application, the memory 502 stores instructions that can be executed by at least one processor 501, and the at least one processor 501 can perform the steps of the above-mentioned audio quality assessment method by executing the instructions stored in the memory 502.

[0117] Among them, the processor 501 is the control center of the computer device, and various interfaces and lines can be used to connect various parts of the computer device. By running or executing instructions stored in the memory 502 and calling data stored in the memory 502, the audio quality of the input audio is determined, and when the evaluation criteria change, there is no need to retrain the audio quality evaluation model, thereby improving the versatility of the audio quality evaluation model. Optionally, the processor 501 may include one or more processing units, and the processor 501 may integrate an application processor and a modem processor, wherein the application processor mainly processes an operating system, a user interface, and an application program, and the modem processor mainly processes wireless communications. It is understandable that the above-mentioned modem processor may not be integrated into the processor 501. In some embodiments, the processor 501 and the memory 502 may be implemented on the same chip, and in some embodiments, they may also be implemented separately on independent chips.

[0118] The processor 501 may be a general-purpose processor, such as a central processing unit (CPU), a digital signal processor, an application-specific integrated circuit (ASIC), a field programmable gate array or other programmable logic device, a discrete gate or transistor logic device, or a discrete hardware component, and may implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of the present application. A general-purpose processor may be a microprocessor or any conventional processor, etc. The steps of the method disclosed in the embodiments of the present application may be directly embodied as being executed by a hardware processor, or may be executed by a combination of hardware and software modules in the processor.

[0119] The memory 502 is a non-volatile computer-readable storage medium that can be used to store non-volatile software programs, non-volatile computer executable programs and modules. The memory 502 may include at least one type of storage medium, such as a flash memory, a hard disk, a multimedia card, a card-type memory, a random access memory (Random Access Memory, RAM), a static random access memory (Static Random Access Memory, SRAM), a programmable read-only memory (Programmable Read Only Memory, PROM), a read-only memory (Read Only Memory, ROM), an electrically erasable programmable read-only memory (Electrically Erasable Programmable Read-Only Memory, EEPROM), a magnetic memory, a disk, an optical disk, etc. The memory 502 is any other medium that can be used to carry or store a desired program code in the form of an instruction or data structure and can be accessed by a computer device, but is not limited thereto. The memory 502 in the embodiment of the present application can also be a circuit or any other device that can realize a storage function, for storing program instructions and / or data.

[0120] Based on the same inventive concept, an embodiment of the present application provides a computer-readable storage medium storing a computer program executable by a computer device. When the program runs on the computer device, the computer device executes the steps of the above-mentioned audio quality assessment method.

[0121] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application may adopt the form of a computer program product implemented in one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that include computer-usable program code.

[0122] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0123] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0124] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0125] Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalents, the present application is also intended to include these modifications and variations.

Claims

1. An audio quality assessment method, applied to an audio quality assessment model, characterized in that: include: Obtain input audio and N reference audios, where the reference audios are determined based on audio samples in an audio quality dataset and quality scores of the audio samples; Determine, according to the N reference audios, a quality score corresponding to each reference audio; fusing the reference audio features of the N reference audios, the quality score features of the N quality scores, and the input audio features of the input audio to obtain a fused feature; The fusion feature represents the correlation relationship between the input audio, the reference audio and the quality score; The audio quality of the input audio feature is determined according to the fused feature.

2. The method according to claim 1, characterized in that The fusing the reference audio features of the N reference audios, the quality score features of the N quality scores, and the input audio features of the input audio to obtain the fused features includes: The reference audio features of the N reference audios and the quality score features of the N quality scores are merged to obtain reference sequence related features; the reference sequence related features represent the correlation relationship between the reference audios and the quality scores; The reference sequence related features and the input audio features of the input audio are fused to obtain the fused features.

3. The method according to claim 2, characterized in that The fusing the reference audio features of the N reference audios and the quality score features of the N quality scores to obtain reference sequence related features includes: Using a cross attention mechanism to fuse the reference audio features of the N reference audios and the quality score features of the N quality scores to obtain a first fused feature; The first fusion feature is extracted through a self-attention mechanism to obtain N reference sequence related features.

4. The method according to any one of claims 1 to 3, characterized in that: The audio quality dataset is composed of the quality scores of the audio samples in different dimensions; The reference audio is determined based on audio samples in the audio quality dataset and quality scores of the audio samples, including: For any audio sample, other audio samples in any dimension and corresponding quality scores are selected as reference audio of the audio sample.

5. The method according to claim 4, characterized in that Also includes: Processing each quality score in the reference audio under any audio sample by a normalization operation to obtain an optimized quality score under each audio sample; The mass score is replaced by the optimized mass score.

6. The method according to claim 5, characterized in that Also includes: The audio quality assessment model is trained using a gradient optimization algorithm, specifically as follows: Where L is the loss function, is the predicted audio quality score, and y is the actual audio quality score.

7. An audio quality assessment device, characterized in that: include: An acquisition module, configured to acquire input audio and N reference audios, wherein the reference audios are determined based on audio samples in an audio quality dataset and quality scores of the audio samples; A processing module, configured to determine a quality score corresponding to each reference audio according to the N reference audios; The processing module is further configured to fuse the reference audio features of the N reference audios, the quality score features of the N quality scores, and the input audio features of the input audio to obtain a fused feature; the fused feature represents the correlation between the input audio, the reference audio, and the quality score; The processing module is further used to determine the audio quality of the input audio feature according to the fusion feature.

8. A computer device comprising a memory, a processor and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the steps of any one of the methods of claims 1 to 6 are implemented.

9. A computer-readable storage medium, characterized in that: It stores a computer program executable by a computer device. When the program is run on the computer device, the computer device executes the steps of any method described in claims 1 to 6.

10. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instruction is executed by a processor, the steps of any method described in claims 1 to 6 are implemented.