Audio Quality Analysis Method, Device, Electronic Device and Computer Readable Storage Medium

The audio quality analysis method automates the scoring process by extracting multiple audio features and using a trained model, addressing the inefficiencies and inaccuracies of manual scoring, thereby enhancing the speed and precision of audio quality evaluation.

CN115240711BActive Publication Date: 2025-07-15FACE CUTE CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202110449490.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-04-25
Publication Date
2025-07-15
Estimated Expiration
2041-04-25

AI Technical Summary

Technical Problem

In the prior art, MOS scoring for audio quality mainly relies on manual scoring, which is time-consuming and labor-intensive and has low accuracy.

Method used

By obtaining multiple audio characteristics of the audio to be analyzed, including the Meer frequency cepspectral coefficient, fundamental frequency, gradient, Meer spectrum and spectral flux, the trained audio quality analysis model is used for automatic analysis, combining multiple sub-quality analysis models and frame-level feature extraction, the average result is calculated to obtain the audio quality analysis results.

Benefits of technology

Automated and fast audio quality analysis is achieved, reducing the time cost of manual ratings and improving analysis accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115240711B_ABST
    Figure CN115240711B_ABST
Patent Text Reader

Abstract

An embodiment of the present disclosure discloses an audio quality analysis method, apparatus, electronic device, and computer-readable storage medium. The audio quality analysis method includes: obtaining an audio to be analyzed; extracting a plurality of audio features of the audio to be analyzed; where at least two types of audio features are included in the plurality of audio features; obtaining a first quantity of audio features from the plurality of audio features of the audio to be analyzed and inputting them into an audio quality analysis model; and obtaining a quality analysis result of the audio to be analyzed according to an output result of the audio quality analysis model. The above method solves the technical problems of time consumption and low analysis accuracy caused by manual analysis through various audio features and an analysis model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of audio quality analysis, and in particular, to an audio quality analysis method, apparatus, electronic device, and computer-readable storage medium. Background Art

[0002] In recent years, with the rapid development of the mobile Internet, the multimedia industry has risen rapidly. Audio content has been favored by a large number of users and creators due to its fast dissemination speed, low production threshold, and strong social attributes. In order to provide users with audio of various qualities, it is necessary to label each audio category, for example, to score the quality of the audio with MOS (Mean Opinion Score).

[0003] Currently, the MOS scoring for audio sound quality mainly relies on manual scoring, which is time-consuming and laborious, and the scoring accuracy is low. Summary of the Invention

[0004] This Summary of the Invention section is provided to introduce concepts in a brief form. These concepts will be described in detail in the following Detailed Description section. This Summary of the Invention section is not intended to identify key features or essential features of the claimed technical solution, nor is it intended to be used to limit the scope of the claimed technical solution.

[0005] To solve the above technical problems, the embodiments of the present disclosure propose the following technical solutions.

[0006] In a first aspect, an embodiment of the present disclosure provides an audio quality analysis method, including:

[0007] Obtain an audio to be analyzed;

[0008] Extract a plurality of audio features of the audio to be analyzed; wherein, at least two types of audio features are included in the plurality of audio features;

[0009] Obtain a first quantity of audio features from the plurality of audio features of the audio to be analyzed and input them into an audio quality analysis model;

[0010] Obtain a quality analysis result of the audio to be analyzed according to an output result of the audio quality analysis model.

[0011] Further, the plurality of audio features include at least two of Mel Frequency Cepstral Coefficients, fundamental frequency, gradient, Mel spectrum, and spectral flux.

[0012] Further, the audio quality analysis model is trained through the following steps:

[0013] Obtain a training data set, where the training data set includes sample audios and sample quality analysis results labeled for the sample audios;

[0014] Extract multiple audio features of the sample audio; wherein, at least two types of audio features are included in the multiple audio features;

[0015] Randomly obtain a first quantity of audio features from the multiple audio features of the sample audio and input them into the audio quality analysis model to be trained;

[0016] Update the parameters of the audio quality analysis model to be trained according to the quality analysis result output by the audio quality analysis model to be trained and the quality scoring annotation data of the sample audio;

[0017] Use other sample audios in the training dataset to iterate the above steps of feature extraction, audio feature input, and updating the parameters of the audio quality analysis model to be trained until the convergence condition of the model is reached to obtain the audio quality analysis model.

[0018] Further, the audio quality analysis model includes multiple sub-quality analysis models. The step of obtaining a first quantity of audio features from the multiple audio features of the audio to be analyzed and inputting them into the audio quality analysis model, and obtaining the quality of the audio to be analyzed according to the output result of the audio quality analysis model includes:

[0019] Each of the sub-quality analysis models of the audio quality analysis model randomly obtains a preset quantity of audio features from the multiple audio features of the audio to be analyzed; wherein the sum of the quantities of the audio features obtained by the sub-quality analysis models is equal to the first quantity;

[0020] Obtain the output result of each sub-quality analysis model;

[0021] Use the output result with the largest quantity among the same output results as the quality analysis result of the audio to be analyzed.

[0022] Further, the step of extracting multiple audio features of the audio to be analyzed and obtaining a first quantity of audio features from the multiple audio features of the audio to be analyzed and inputting them into the audio quality analysis model includes:

[0023] Divide the audio to be analyzed into multiple audio frames;

[0024] Extract multiple audio features for each audio frame;

[0025] Randomly obtain a first quantity of audio features from the multiple audio features of each audio frame among the multiple audio frames and input them into the audio quality analysis model respectively.

[0026] Further, the step of obtaining the analysis result of the audio to be analyzed according to the output result of the audio quality analysis model includes:

[0027] The audio quality analysis model outputs a plurality of output results according to the plurality of audio frames; wherein, each audio frame corresponds to one output result;

[0028] Calculate the average result of the plurality of output results as the quality analysis result of the audio to be analyzed.

[0029] Further, the obtaining of the audio to be analyzed includes:

[0030] Collecting an audio signal through an audio acquisition interface to obtain the audio to be analyzed; or

[0031] Receiving the audio to be analyzed through a data transmission interface.

[0032] Further, after obtaining the quality analysis result of the audio to be analyzed according to the output result of the audio quality analysis model, it further includes:

[0033] Displaying a first page, where the quality analysis result of the audio to be analyzed is included in the first page.

[0034] Further, the training data set is obtained through the following steps:

[0035] Determining the number of sample audios in the sample acquisition terminal;

[0036] Collecting the number of sample audios through the sample acquisition terminal;

[0037] Sending the sample audio to a first platform through the sample acquisition terminal;

[0038] Receiving the grade annotation data of the sample audio through the first platform to obtain the training data set.

[0039] In a second aspect, an embodiment of the present disclosure provides an audio quality analysis device, including:

[0040] An audio acquisition module, configured to acquire an audio to be analyzed;

[0041] A feature extraction module, configured to extract a plurality of audio features of the audio to be analyzed; wherein, at least two types of audio features are included in the plurality of audio features;

[0042] An input module, configured to obtain a first number of audio features from the plurality of audio features of the audio to be analyzed and input them into an audio quality analysis model;

[0043] An analysis result acquisition module, configured to obtain the quality analysis result of the audio to be analyzed according to the output result of the audio quality analysis model.

[0044] Further, the multiple audio features include at least two of Mel Frequency Cepstral Coefficients (MFCCs), fundamental frequency, gradient, Mel spectrogram, and spectral flux.

[0045] Further, the audio quality analysis model is trained through the following steps:

[0046] Obtain a training data set, which includes sample audio and sample quality analysis results annotated for the sample audio;

[0047] Extract multiple audio features of the sample audio; among them, the multiple audio features include at least two types of audio features;

[0048] Randomly obtain a first number of audio features from the multiple audio features of the sample audio and input them into the audio quality analysis model to be trained;

[0049] Update the parameters of the audio quality analysis model to be trained according to the quality analysis result output by the audio quality analysis model to be trained and the quality scoring annotation data of the sample audio;

[0050] Use other sample audio in the training data set to iterate the above steps of feature extraction, audio feature input, and updating the parameters of the audio quality analysis model to be trained until the convergence condition of the model is reached to obtain the audio quality analysis model.

[0051] Further, the audio quality analysis model includes multiple sub-quality analysis models, and the input module and the analysis result acquisition module are also used for:

[0052] Each sub-quality analysis model of the audio quality analysis model randomly obtains a preset number of audio features from the multiple audio features of the audio to be analyzed; the sum of the numbers of audio features obtained by the sub-quality analysis models is equal to the first number;

[0053] Obtain the output results of each sub-quality analysis model;

[0054] Take the output result with the largest number in the same output results as the quality analysis result of the audio to be analyzed.

[0055] Further, the feature extraction module and the input module are also used for:

[0056] Divide the audio to be analyzed into multiple audio frames;

[0057] Extract multiple audio features for each audio frame;

[0058] Randomly obtain a first number of audio features from the multiple audio features of each audio frame among the multiple audio frames and input them into the audio quality analysis model respectively.

[0059] Further, the analysis result obtaining module is further configured to:

[0060] The audio quality analysis model outputs a plurality of output results according to the plurality of audio frames; wherein, each audio frame corresponds to an output result;

[0061] Calculate the average result of the plurality of output results as the quality analysis result of the audio to be analyzed.

[0062] Further, the audio obtaining module is further configured to:

[0063] Collect an audio signal through an audio acquisition interface to obtain the audio to be analyzed; or

[0064] Receive the audio to be analyzed through a data transmission interface.

[0065] Further, the audio quality analysis device further includes a display module, configured to:

[0066] Display a first page, where the first page includes the quality analysis result of the audio to be analyzed.

[0067] Further, the training data set is obtained through the following steps:

[0068] Determine the number of sample audios in the sample collection terminal;

[0069] Collect the number of sample audios through the sample collection terminal;

[0070] Send the sample audio to a first platform through the sample collection terminal;

[0071] Receive the grade annotation data of the sample audio through the first platform to obtain the training data set.

[0072] In a third aspect, an embodiment of the present disclosure provides an electronic device, including: at least one processor; and,

[0073] A memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute any one of the methods in the foregoing first aspect.

[0074] In a fourth aspect, an embodiment of the present disclosure provides a non-transitory computer-readable storage medium, characterized in that the non-transitory computer-readable storage medium stores computer instructions for causing a computer to execute any one of the methods in the foregoing first aspect.

[0075] Embodiments of the present disclosure disclose an audio quality analysis method, apparatus, electronic device, and computer-readable storage medium. The audio quality analysis method includes: obtaining an audio to be analyzed; extracting a plurality of audio features of the audio to be analyzed; where the plurality of audio features include at least two types of audio features; obtaining a first quantity of audio features from the plurality of audio features of the audio to be analyzed and inputting them into an audio quality analysis model; and obtaining a quality analysis result of the audio to be analyzed according to an output result of the audio quality analysis model. The above method solves the technical problems of time consumption and low analysis accuracy caused by manual analysis through various audio features and analysis models.

[0076] The above description is only an overview of the technical solution of the present disclosure. In order to understand the technical means of the present disclosure more clearly, it can be implemented according to the content of the specification. In order to make the above and other purposes, features, and advantages of the present disclosure more obvious and understandable, the following preferred embodiments are specifically given, and in conjunction with the accompanying drawings, the details are described as follows. BRIEF DESCRIPTION OF THE DRAWINGS

[0077] In combination with the accompanying drawings and referring to the following specific embodiments, the above and other features, advantages, and aspects of the embodiments of the present disclosure will become more obvious. Throughout the drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic, and the original elements and elements are not necessarily drawn to scale.

[0078] Figure 1 It is a flowchart of the audio quality analysis method provided by the embodiments of the present disclosure;

[0079] Figure 2 It is a schematic diagram of audio frame segmentation in the audio quality analysis method provided by the embodiments of the present disclosure;

[0080] Figure 3 It is a schematic diagram of an application scenario of the audio quality analysis method provided by the embodiments of the present disclosure;

[0081] Figure 4 It is a schematic diagram of the structure of an embodiment of the audio quality analysis apparatus provided by the embodiments of the present disclosure;

[0082] Figure 5 It is a schematic diagram of the structure of an electronic device provided according to the embodiments of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0083] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Instead, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not used to limit the protection scope of the present disclosure.

[0084] It should be understood that the various steps recited in the method embodiments of the present disclosure can be executed in a different order and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this regard.

[0085] As used herein, the term "including" and its variants are open-ended, i.e., "including but not limited to". The term "based on" is "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". The relevant definitions of other terms will be given in the following description.

[0086] It should be noted that the concepts such as "first", "second", etc. mentioned in the present disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence relationship of the functions executed by these devices, modules or units.

[0087] It should be noted that the modifications of "one" and "multiple" mentioned in the present disclosure are illustrative rather than restrictive. Those skilled in the art should understand that, unless otherwise clearly stated in the context, it should be understood as "one or more".

[0088] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only for illustrative purposes and are not used to limit the scope of these messages or information.

[0089] Figure 1 For the flowchart of the method embodiment of the audio quality analysis provided for the embodiments of the present disclosure, the audio quality analysis method provided in this embodiment can be executed by an audio quality analysis device, which can be implemented as software, or implemented as a combination of software and hardware. The audio quality analysis device can be integrally arranged in a certain device in the audio quality analysis system, such as an audio quality analysis server or an audio quality analysis terminal device. As Figure 1 shown, the method includes the following steps:

[0090] Step S101, obtain the audio to be analyzed;

[0091] Optionally, the level of the audio to be analyzed represents the quality of the audio to be analyzed, such as sound quality, etc. The audio to be analyzed may include various types of audio content such as speech, music, songs, etc.

[0092] Optionally, step S101 includes:

[0093] Collecting an audio signal through an audio acquisition interface to obtain the audio to be analyzed; or

[0094] Receiving the audio to be analyzed through a data transmission interface.

[0095] Among them, the audio acquisition interface is an interface built into other audio systems for real-time acquisition of audio data. Exemplarily, by receiving the audio recorded by the user in real time through a server, the server is connected to the user's terminal device through an http interface. The user terminal device collects the sound in the environment through a microphone and uploads it to the server through the http. Thus, remote real-time audio acquisition can be achieved.

[0096] Among them, the data transmission interface is a web page interface. For example, an upload entry is set on the page, and the user uploads the audio to be analyzed through this upload entry.

[0097] Return attachment Figure 1 The audio quality analysis method further includes step S102 of extracting multiple audio features of the audio to be analyzed; among them, at least two types of audio features are included in the multiple audio features;

[0098] Among them, any audio feature extraction method can be used to extract the multiple audio features of the audio to be analyzed, and the type of the extracted audio features can also be any type. Exemplarily, the audio features may include at least two of Mel Frequency Cepstral Coefficients, fundamental frequency, gradient, Mel spectrogram, and spectral flux.

[0099] Exemplarily, the above five types of audio features of the audio to be analyzed are extracted. Among them, the Mel Frequency Cepstral Coefficient type includes 40 audio features, the fundamental frequency type includes 3 features, the gradient type includes 2 features, the Mel spectrogram type includes 256 features, and the spectral flux type includes 3 features. Each type of feature will include the standard deviation and average value of the feature, and the maximum value information is included in the gradient type and spectral flux type. Taking the above five types of audio features as an example, 304 audio features of the above five types are obtained by extracting the audio features of the audio to be analyzed, and these 304 audio features are respectively stored in 304 arrays.

[0100] Optionally, step S102 includes: dividing the audio to be analyzed into multiple audio frames; extracting multiple audio features for each audio frame. Among them, dividing the audio to be analyzed into multiple audio frames includes: dividing the audio to be analyzed into multiple audio frames according to a preset frame length and a preset step size. Exemplarily, as Figure 2 shown, the duration of an audio is 5 seconds, the preset frame length is 1 second, and the preset step size is 0.5 second. Then the above 5-second audio is divided into 9 audio frames, and there is an overlap of 0.5 second between two adjacent audio frames. After obtaining the multiple audio frames, multiple audio features are extracted for each audio frame. The process of extracting audio features from the multiple audio frames is the same as the process of extracting audio features of the audio to be analyzed, which will not be elaborated here.

[0101] Return to Appendix Figure 1 , the audio quality analysis method further includes step S103 of obtaining a first number of audio features from the multiple audio features of the audio to be analyzed and inputting them into an audio quality analysis model.

[0102] In this step, the first number is related to the audio quality analysis model used. Obtaining a first number of audio features from the multiple audio features of the audio to be analyzed includes: randomly obtaining a first number of audio features from the at least two types of audio features. Exemplarily, in the above example, the extracted audio features include 304 features of five types, then a first number of audio features are randomly obtained from the 304 features as the input of the audio quality analysis model. Among them, the first number can be any value, and can be greater than, equal to, or less than the total number of audio features.

[0103] Optionally, when randomly obtaining a first number of audio features from the multiple audio features of the audio to be analyzed, the audio features are numbered, and then the first number of audio features are obtained by executing the preset random function the first number of times.

[0104] Optionally, the audio quality analysis model includes multiple sub-quality analysis models, and step S103 includes:

[0105] Each of the sub-quality analysis models of the audio quality analysis model randomly obtains a preset number of audio features from the multiple audio features of the audio to be analyzed; where the sum of the numbers of audio features obtained by the sub-quality analysis models is equal to the first number.

[0106] Among them, each of the multiple sub-quality analysis models is used to analyze the audio to be analyzed, and each sub-quality analysis model randomly obtains a preset number of audio features from the multiple audio features; wherein, the preset number can be the same or different for each sub-quality analysis model. For example, the audio quality analysis model includes 4 sub-quality analysis models. The first sub-quality analysis model randomly obtains 10 audio features, the second sub-quality analysis model randomly obtains 100 audio features, the third sub-quality analysis model randomly obtains 30 audio features, and the fourth sub-quality analysis model randomly obtains 90 audio features. The sum of the numbers of audio features obtained by all sub-quality analysis models is equal to the first number.

[0107] When the audio to be analyzed is divided into multiple audio frames, step S103 includes: randomly obtaining a first number of audio features from the multiple audio features of each audio frame among the multiple audio frames and inputting them into the audio quality analysis model respectively.

[0108] Return attachment Figure 1 The audio quality analysis method further includes step S104 of obtaining the quality analysis result of the audio to be analyzed according to the output result of the audio quality analysis model.

[0109] In this step, the audio quality analysis model outputs the analysis result of the audio to be analyzed.

[0110] Optionally, when the audio quality analysis model includes multiple sub-quality analysis models, step S104 includes:

[0111] Obtaining the output result of each sub-quality analysis model;

[0112] Taking the output result with the largest number in the same output results as the analysis result of the audio to be analyzed.

[0113] In the above step, each sub-quality analysis model will output an output result, that is, the analysis result of the audio to be analyzed, and then take the output result with the largest number in the same output results as the final analysis result. For example, the number of sub-quality analysis models is 50, among which the analysis results output by 15 sub-quality analysis models are the first result, the analysis results of 25 sub-quality analysis models are the second result, and the analysis results output by 10 sub-quality analysis models are the third result. Then the output result of the audio quality analysis model is the second result. If there are multiple output results with the largest number in the same output results, randomly select one of them as the analysis result.

[0114] Optionally, when the audio to be analyzed is divided into multiple audios, step S104 includes:

[0115] The audio quality analysis model outputs multiple output results according to the multiple audio frames;

[0116] Calculate the average result of the multiple output results as the analysis result of the audio to be analyzed.

[0117] In the above steps, for each audio frame, the audio quality analysis model outputs one output result to obtain multiple output results, and finally takes the average result of the multiple output results as the analysis result of the audio to be analyzed. For example, the analysis result represents the sound quality level of the audio, which can be represented by a numerical value. For example, the 5 levels are represented as 1-5 respectively, then each audio frame corresponds to one level. At this time, calculate the average level of the levels corresponding to multiple audio frames as the analysis result.

[0118] In the above steps S101 - S104, by extracting at least two types of audio features of the audio to be analyzed and inputting the obtained first quantity of audio features into the trained audio quality analysis model, the analysis result of the audio to be analyzed is obtained. Due to the use of the audio quality analysis model, the problem of time-consuming and laborious manual audio quality analysis is avoided; due to the use of at least two types of audio features, the problem of low analysis accuracy is solved.

[0119] Further, the audio quality analysis model used in the above steps is a pre-trained model, which is trained through the following steps:

[0120] Obtain a training data set, where the training data set includes sample audio and sample quality analysis results annotated for the sample audio;

[0121] Extract multiple audio features of the sample audio; among them, the multiple audio features include at least two types of audio features;

[0122] Randomly obtain a first quantity of audio features from the multiple audio features of the sample audio and input them into the audio quality analysis model to be trained;

[0123] Update the parameters of the audio quality analysis model to be trained according to the quality analysis result output by the audio quality analysis model to be trained and the quality scoring annotation data of the sample audio;

[0124] Use other sample audio in the training data set to iterate the above steps of feature extraction, audio feature input, and updating the parameters of the audio quality analysis model to be trained until the convergence condition of the model is reached to obtain the audio quality analysis model..

[0125] It is understandable that the above steps may further include an initialization step for initializing the parameters of the audio quality analysis model to be trained. The initialization step may include randomly initializing the parameters of the audio quality analysis model to be trained or initializing the parameters of the audio quality analysis model to be trained to preset values. When the audio quality analysis model to be trained includes multiple sub-quality analysis models, initializing the parameters of the audio quality analysis model to be trained includes initializing the framework parameters of the audio quality analysis model and the parameters of the sub-quality analysis models. Exemplarily, if the audio quality analysis model is a random forest model, the parameters of the audio quality analysis model include framework parameters, such as the number of decision trees in the random forest, and also include the parameters of the sub-quality analysis models, such as the parameters of the decision trees in the random forest, such as the maximum number of features of the decision tree, that is, the preset number of the above randomly obtained features, the depth of the decision tree, the maximum number of leaf nodes, and other parameters.

[0126] Obtain a training sample set, which includes sample audio and grade annotation data of the sample audio. The grade annotation data may represent the sound quality of the sample audio in the present disclosure. For example, the sound quality is divided into 5 grades, and the grade annotation data is one of the 5 grades. The grade annotation data can be represented by a vector. For example, the 5 grades can be represented by a 5-bit one-hot encoding.

[0127] Extract multiple audio features of the sample audio. The execution process of this step is the same as that of the above step S102 and will not be elaborated here.

[0128] Obtain a first number of audio features from the multiple audio features of the sample audio and input them into the audio quality analysis model to be trained. Among them, the process of obtaining the first number of audio features is the same as that in step S103 and will not be elaborated here.

[0129] Update the parameters of the audio quality analysis model to be trained according to the analysis result output by the audio quality analysis model to be trained and the grade annotation data of the sample audio. Exemplarily, calculate the loss of the audio quality analysis model to be trained through a pre-set loss function, and the loss function is a function related to the analysis result and the grade annotation data. After obtaining the loss of the audio quality analysis model to be trained, calculate the gradient of the loss function according to the backpropagation algorithm, and update the parameters of the audio quality analysis model to be trained according to the gradient.

[0130] After that, use other sample audio in the training set to iterate the above steps of feature extraction, input of audio features, and updating the parameters of the audio quality analysis model to be trained until the convergence condition of the model is reached to obtain the audio quality analysis model. Exemplarily, the convergence condition includes that the value of the loss function is less than a preset value.

[0131] Further, the training data set is obtained through the following steps:

[0132] Determine the number of sample audios in the sample collection terminal;

[0133] Collect the determined number of sample audios through the sample collection terminal;

[0134] Send the sample audios to the first platform through the sample collection terminal;

[0135] Receive the grading annotation data of the sample audios through the first platform to obtain the training data set.

[0136] Among them, the first platform includes a crowdsourcing testing platform. After determining the number of samples in the training data set, the sample collection terminal collects a sufficient number of sample audios and sends them to the first platform, which is used to display the sample audios to users and accept the annotation of the sample audios by users. Exemplarily, the first platform shows the annotation criteria to users, such as the division criteria of sound quality, and plays the sample audios, and then receives the selected grading annotation data by users to obtain the training data set. Exemplarily, each sample audio may be annotated multiple times. If multiple users annotate the same sample audio, then at this time, select the grading annotation data with the most annotation times as the grading annotation data of the sample audio or use the average value of the grading annotation data of multiple users as the grading annotation data of the sample audio.

[0137] Further, after the step S104, it further includes: displaying a first page, where the first page includes the analysis result of the audio to be analyzed.

[0138] In the above steps, after obtaining the analysis result, display the first page showing the analysis result on the terminal. Optionally, the first page may only show the analysis result of the current audio, or may show the analysis result of the current audio and the analysis results of other historical audios.

[0139] Figure 3 It is an application scenario of the audio quality analysis method in the above embodiment. Such as Figure 2As shown, in the described application scenario, a sound quality score is given to the audio to be analyzed, where the scores range from 1 to 5 points, corresponding to 5 levels respectively. After obtaining the audio to be analyzed, through frame segmentation operation, it is divided into multiple audio frames with a frame length of 1 second and a frame interval of 0.5 second; then, feature extraction is performed on each audio frame, such as extracting five audio features including Mel-frequency cepstral coefficients, fundamental frequency, gradient, Mel-spectrum, and spectral flux of each audio frame. Among them, the Mel-frequency cepstral coefficient type includes 40 audio features, the fundamental frequency type includes 3 features, the gradient type includes 2 features, the Mel-spectrum type includes 256 features, and the spectral flux type includes 3 features, totaling 304 features. Then, the audio quality analysis model randomly selects all or part of the features from these 304 features with repetition as the input of the model, including each sub-model randomly selecting all or part of the features from the 304 features as the input of the sub-model. Each sub-model outputs a score, and then the score that appears the most times is selected as the analysis result of the audio quality analysis model. The above operations are performed on each audio frame of the same audio to be analyzed to obtain multiple analysis results of the same audio to be analyzed, and the average analysis result of the multiple analysis results is calculated as the analysis result of the audio to be analyzed.

[0140] An embodiment of the present disclosure discloses an audio quality analysis method. The audio quality analysis method includes: obtaining an audio to be analyzed; extracting multiple audio features of the audio to be analyzed; where at least two types of audio features are included in the multiple audio features; obtaining a first quantity of audio features from the multiple audio features of the audio to be analyzed and inputting them into an audio quality analysis model; obtaining a quality analysis result of the audio to be analyzed according to the output result of the audio quality analysis model. The above method solves the technical problems of time-consuming and low analysis accuracy caused by manual analysis through multiple audio features and an analysis model.

[0141] In the above text, although the steps in the above method embodiments are described in the above order, those skilled in the art should understand that the steps in the embodiments of the present disclosure do not necessarily need to be executed in the above order, and they can also be executed in reverse order, in parallel, in cross, or other orders. Moreover, based on the above steps, those skilled in the art can also add other steps, and these obvious variations or equivalent replacement methods should also be included in the protection scope of the present disclosure, which will not be elaborated here.

[0142] Figure 4 It is a schematic structural diagram of an embodiment of an audio quality analysis device provided by an embodiment of the present disclosure. As Figure 4 shown, the device 400 includes: an audio acquisition module 401, a feature extraction module 402, an input module 403, and an analysis result acquisition module 404. Among them,

[0143] The audio acquisition module 401 is used to acquire the audio to be analyzed;

[0144] The feature extraction module 402 is used to extract multiple audio features of the audio to be analyzed; wherein, at least two types of audio features are included in the multiple audio features;

[0145] The input module 403 is used to obtain a first number of audio features from the multiple audio features of the audio to be analyzed and input them into the audio quality analysis model;

[0146] The analysis result acquisition module 404 is used to obtain the quality analysis result of the audio to be analyzed according to the output result of the audio quality analysis model.

[0147] Further, the multiple audio features include at least two of mel-frequency cepstral coefficients, fundamental frequency, gradient, mel-spectrum, and spectral flux.

[0148] Further, the audio quality analysis model is trained through the following steps:

[0149] Obtain a training data set, where the training data set includes sample audio and sample quality analysis results annotated for the sample audio;

[0150] Extract multiple audio features of the sample audio; wherein, at least two types of audio features are included in the multiple audio features;

[0151] Randomly obtain a first number of audio features from the multiple audio features of the sample audio and input them into the audio quality analysis model to be trained;

[0152] Update the parameters of the audio quality analysis model to be trained according to the quality analysis result output by the audio quality analysis model to be trained and the quality scoring annotation data of the sample audio;

[0153] Use other sample audio in the training data set to iterate the above steps of feature extraction, audio feature input, and updating the parameters of the audio quality analysis model to be trained until the convergence condition of the model is reached to obtain the audio quality analysis model.

[0154] Further, the audio quality analysis model includes multiple sub-quality analysis models, and the input module 403 and the analysis result acquisition module 404 are also used for:

[0155] Each of the sub-quality analysis models of the audio quality analysis model randomly obtains a preset number of audio features from the multiple audio features of the audio to be analyzed; wherein the sum of the numbers of audio features obtained by the sub-quality analysis models is equal to the first number;

[0156] Obtain the output results of each of the sub-quality analysis models;

[0157] Take the output result with the largest quantity among the same output results as the quality analysis result of the audio to be analyzed.

[0158] Furthermore, the feature extraction module 402 and the input module 403 are also used for:

[0159] Divide the audio to be analyzed into multiple audio frames;

[0160] Extract multiple audio features for each audio frame;

[0161] Randomly obtain a first quantity of audio features from the multiple audio features of each audio frame among the multiple audio frames and input them into the audio quality analysis model respectively.

[0162] Furthermore, the analysis result acquisition module 404 is also used for:

[0163] The audio quality analysis model outputs multiple output results according to the multiple audio frames; wherein, each audio frame corresponds to one output result;

[0164] Calculate the average result of the multiple output results as the quality analysis result of the audio to be analyzed.

[0165] Furthermore, the audio acquisition module 401 is also used for:

[0166] Collect audio signals through an audio acquisition interface to obtain the audio to be analyzed; or

[0167] Receive the audio to be analyzed through a data transmission interface.

[0168] Furthermore, the audio quality analysis device further includes a display module, which is used for:

[0169] Display a first page, wherein the first page includes the quality analysis result of the audio to be analyzed.

[0170] Furthermore, the training data set is obtained through the following steps:

[0171] Determine the quantity of sample audio in the sample acquisition terminal;

[0172] Collect the quantity of sample audio through the sample acquisition terminal;

[0173] Send the sample audio to a first platform through the sample acquisition terminal;

[0174] Receive the grade annotation data of the sample audio through the first platform to obtain the training data set.

[0175] Figure 4 The device shown can execute Figures 1-2 the method of the embodiment shown. For parts not described in detail in this embodiment, reference can be made to the relevant descriptions of Figures 1-2 the embodiment shown. For the execution process and technical effects of this technical solution, refer to the description in Figures 1-2 the embodiment shown, which will not be elaborated herein.

[0176] Next, refer to Figure 5 , which shows a schematic structural diagram of an electronic device 500 suitable for implementing the embodiments of the present disclosure. The terminal devices in the embodiments of the present disclosure may include, but are not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Tablet Computers), PMPs (Portable Multimedia Players), in-vehicle terminals (such as in-vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 5 The electronic device shown is only an example and should not impose any limitations on the functions and usage scope of the embodiments of the present disclosure.

[0177] As Figure 5 shown, the electronic device 500 may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 501, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 502 or the program loaded from the storage device 508 into the random access memory (RAM) 503. In the RAM 503, various programs and data required for the operation of the electronic device 500 are also stored. The processing device 501, the ROM 502, and the RAM 503 are connected to each other through a bus 504. The input / output (I / O) interface 505 is also connected to the bus 504.

[0178] Generally, the following devices may be connected to the I / O interface 505: an input device 506 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 507 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 508 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 509. The communication device 509 can allow the electronic device 500 to communicate with other devices wirelessly or wiredly to exchange data. Although Figure 5 the electronic device 500 with various devices is shown, it should be understood that it is not required to implement or have all the shown devices. Instead, more or fewer devices can be implemented or had.

[0179] In particular, according to embodiments of the present disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of the present disclosure include a computer program product that includes a computer program carried on a non-transitory computer-readable medium, and the computer program includes program code for performing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from the network through the communication device 509, or installed from the storage device 508, or installed from the ROM 502. When the computer program is executed by the processing device 501, the above-mentioned functions defined in the methods of the embodiments of the present disclosure are performed.

[0180] It should be noted that the above-mentioned computer-readable medium in the present disclosure can be a computer-readable signal medium or a computer-readable storage medium or any combination of the two. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, the computer-readable storage medium can be any tangible medium that contains or stores a program, and the program can be used by or in combination with an instruction execution system, apparatus, or device. In the present disclosure, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, in which computer-readable program code is carried. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, and the computer-readable signal medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted by any suitable medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.

[0181] In some embodiments, the client and the server can communicate using any currently known or future-developed network protocol such as HTTP (HyperText Transfer Protocol), and can be interconnected with digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include local area networks ("LANs"), wide area networks ("WANs"), the Internet (e.g., the Internet), and end-to-end networks (e.g., ad hoc end-to-end networks), as well as any currently known or future-developed networks.

[0182] The above computer-readable medium can be included in the above electronic device; or can exist separately without being assembled into the electronic device.

[0183] The above computer-readable medium carries one or more programs, which, when executed by the electronic device, cause the electronic device to perform the above audio quality analysis method.

[0184] Computer program code for performing the operations of the present disclosure can be written in one or more programming languages or combinations thereof. The programming languages include, but are not limited to, object-oriented programming languages such as Java, Smalltalk, and C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (e.g., by using an Internet service provider to connect through the Internet).

[0185] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a portion of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that, in some alternative implementations, the functions noted in the blocks may occur in a different order than noted in the drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or by a combination of dedicated hardware and computer instructions.

[0186] The units involved in the embodiments described in the present disclosure can be implemented in software or in hardware. In some cases, the name of the unit does not constitute a limitation on the unit itself.

[0187] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. By way of example, and without limitation, exemplary types of hardware logic components that can be used include: Field Programmable Gate Arrays (FPGA), Application Specific Integrated Circuits (ASIC), Application Specific Standard Products (ASSP), System on a Chip (SOC), Complex Programmable Logic Devices (CPLD), and the like.

[0188] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a Random Access Memory (RAM), a Read-Only Memory (ROM), an Erasable Programmable Read-Only Memory (EPROM or Flash memory), an optical fiber, a portable Compact Disc Read-Only Memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0189] According to one or more embodiments of the present disclosure, an electronic device is provided, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute any of the audio quality analysis methods in the foregoing first aspect.

[0190] According to one or more embodiments of the present disclosure, a non-transitory computer-readable storage medium is provided, characterized in that the non-transitory computer-readable storage medium stores computer instructions for causing a computer to execute any of the audio quality analysis methods in the foregoing first aspect.

[0191] The above description is only a preferred embodiment of the present disclosure and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of disclosure involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosure concept. For example, the technical solutions formed by mutually replacing the above features with the (but not limited to) technical features having similar functions disclosed in the present disclosure.

Claims

1. An audio quality analysis method, characterized in that, Including: Obtain the audio to be analyzed; Extract multiple audio features of the audio to be analyzed; wherein, at least two types of audio features are included in the multiple audio features; Obtain a first quantity of audio features from the multiple audio features of the audio to be analyzed and input them into an audio quality analysis model; the audio quality analysis model includes multiple sub-quality analysis models. The process of obtaining a first quantity of audio features from the multiple audio features of the audio to be analyzed and inputting them into the audio quality analysis model, and obtaining the quality of the audio to be analyzed according to the output result of the audio quality analysis model includes: each of the sub-quality analysis models of the audio quality analysis model randomly obtains a preset quantity of audio features from the multiple audio features of the audio to be analyzed; wherein the sum of the quantities of the audio features obtained by the sub-quality analysis model is equal to the first quantity; obtain the output result of each sub-quality analysis model; use the output result with the largest quantity among the same output results as the quality analysis result of the audio to be analyzed; Obtain the quality analysis result of the audio to be analyzed according to the output result of the audio quality analysis model.

2. The audio quality analysis method according to claim 1, wherein The multiple audio features include at least two of Mel Frequency Cepstral Coefficients, fundamental frequency, gradient, Mel spectrogram, and spectral flux.

3. The audio quality analysis method according to claim 1, wherein, The audio quality analysis model is trained through the following steps: Obtain a training data set, which includes sample audio and the sample quality analysis result annotated for the sample audio; Extract multiple audio features of the sample audio; wherein, at least two types of audio features are included in the multiple audio features; Randomly obtain a first quantity of audio features from the multiple audio features of the sample audio and input them into the audio quality analysis model to be trained; Update the parameters of the audio quality analysis model to be trained according to the quality analysis result output by the audio quality analysis model to be trained and the quality scoring annotation data of the sample audio; Use other sample audio in the training data set to iterate the above steps of feature extraction, audio feature input, and updating the parameters of the audio quality analysis model to be trained until the convergence condition of the model is reached to obtain the audio quality analysis model.

4. The audio quality analysis method according to claim 1, wherein The process of extracting multiple audio features of the audio to be analyzed and obtaining a first quantity of audio features from the multiple audio features of the audio to be analyzed and inputting them into the audio quality analysis model includes: Divide the audio to be analyzed into multiple audio frames; Extract multiple audio features for each audio frame; Randomly obtain a first quantity of audio features from the multiple audio features of each audio frame among the multiple audio frames and input them into the audio quality analysis model respectively.

5. The audio quality analysis method according to claim 4, wherein The process of obtaining the analysis result of the audio to be analyzed according to the output result of the audio quality analysis model includes: The audio quality analysis model outputs multiple output results according to the multiple audio frames; wherein, each audio frame corresponds to an output result; Calculate the average result of the multiple output results as the quality analysis result of the audio to be analyzed.

6. The audio quality analysis method according to claim 1, wherein The process of obtaining the audio to be analyzed includes: Collect an audio signal through an audio acquisition interface to obtain the audio to be analyzed; or Receive the audio to be analyzed through a data transmission interface.

7. The audio quality analysis method according to claim 1, wherein After obtaining the quality analysis result of the audio to be analyzed according to the output result of the audio quality analysis model, the following steps are further included: Display a first page, where the quality analysis result of the audio to be analyzed is included in the first page.

8. The audio quality analysis method according to claim 3, wherein The training data set is obtained through the following steps: Determine the number of sample audios in the sample collection terminal; Collect the number of sample audios through the sample collection terminal; Send the sample audios to a first platform through the sample collection terminal; Receive the grade annotation data of the sample audios through the first platform to obtain the training data set.

9. An audio quality analysis device, characterized in that, Include: An audio acquisition module for acquiring the audio to be analyzed; A feature extraction module for extracting multiple audio features of the audio to be analyzed; wherein, at least two types of audio features are included in the multiple audio features; An input module for randomly acquiring a first number of audio features from the multiple audio features of the audio to be analyzed and inputting them into an audio quality analysis model; the audio quality analysis model includes multiple sub-quality analysis models. Obtaining the quality of the audio to be analyzed by inputting a first number of audio features from the multiple audio features of the audio to be analyzed according to the output result of the audio quality analysis model includes: each of the sub-quality analysis models of the audio quality analysis model randomly acquires a preset number of audio features from the multiple audio features of the audio to be analyzed; wherein the sum of the numbers of audio features acquired by the sub-quality analysis models is equal to the first number; obtaining the output result of each sub-quality analysis model; and taking the output result with the largest number among the same output results as the quality analysis result of the audio to be analyzed. An analysis result acquisition module for obtaining the analysis result of the audio to be analyzed according to the output result of the audio quality analysis model.

10. An electronic device, including: A memory for storing computer-readable instructions; And A processor for running the computer-readable instructions, so that when the processor runs, it implements the method according to any one of claims 1-8.

11. A non-transitory computer-readable storage medium for storing computer-readable instructions, when the computer-readable instructions are executed by a computer, enabling the computer to execute the method according to any one of claims 1-8.

Citation Information

Patent Citations

  • Audio processing method, audio processing device, storage medium and electronic equipment

    CN110739006A