Audio quality determination method and device, electronic equipment and storage medium

By acquiring the recording signal and reference signal of the audio playback device, and evaluating the audio quality using the perceptual model, the problem of low accuracy caused by subjective listening is solved, and a more accurate and efficient audio quality assessment is achieved.

CN120581039APending Publication Date: 2025-09-02XIAOMI EV TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410797571.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-06-19
Publication Date
2025-09-02

AI Technical Summary

Technical Problem

In the prior art, the quality evaluation of audio playback devices mainly relies on subjective listening, resulting in low evaluation accuracy and time-consuming and labor-intensive.

Method used

By acquiring the recording signal, the tone reference signal and the spatial reference signal of the audio playback device, the tone and spatial characteristics are determined using the perception model, and the audio quality is then evaluated.

Benefits of technology

It improves the accuracy and efficiency of the quality assessment of audio playback devices, realizes multi-dimensional quality assessment, and reduces subjective influence.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120581039A_ABST
    Figure CN120581039A_ABST
Patent Text Reader

Abstract

The invention provides an audio quality determination method and device, electronic equipment and a storage medium, and the method comprises the steps: obtaining a first sound recording signal obtained by recording preset first audio data played by a to-be-tested audio playing device, and obtaining a first sound recording signal compared with a mode of obtaining the first sound recording signal through synthesis; the accuracy of obtaining the first recording signal is improved, then a tone reference signal and a spatial reference signal corresponding to the to-be-tested audio playing device are obtained, and a first tone feature and a first spatial feature corresponding to the first recording signal are determined according to the tone reference signal, the spatial reference signal and the first recording signal; the first timbre feature and the first spatial feature are input into the perception model obtained through training, the first timbre quality score and the first spatial quality score of the to-be-tested audio playing device are obtained, the reference signal is not fixed any more, and the accuracy of subsequent multi-dimensional quality scoring is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of audio data processing, and in particular to a method, device, electronic device, and storage medium for determining audio quality. Background Art

[0002] At present, science and technology are developing rapidly, and various audio playback devices have also made great progress.

[0003] In the prior art, the quality of audio played by audio playback devices is primarily evaluated through subjective listening. While subjective methods can, to a certain extent, evaluate and quantify the quality of audio playback, they are highly subjective, time-consuming, and labor-intensive, and their accuracy is low. Summary of the Invention

[0004] The present application aims to solve one of the technical problems in the related art at least to a certain extent.

[0005] To this end, the present application proposes a method, device, electronic device and storage medium for determining audio quality to improve the accuracy of determining a quality score corresponding to audio data played by an audio playback device.

[0006] In one aspect, an embodiment of the present application provides a method for determining audio quality, including:

[0007] Acquire a first recording signal obtained by recording a preset first audio data played by the audio playback device to be tested;

[0008] Acquire a timbre reference signal and a spatial reference signal corresponding to the audio playback device to be tested;

[0009] determining a first timbre feature corresponding to the first recording signal based on the timbre reference signal and the first recording signal, and determining a first spatial feature corresponding to the first recording signal based on the spatial reference signal and the first recording signal;

[0010] The first timbre feature and the first spatial feature are input into the trained perception model to obtain a first timbre quality score and a first spatial quality score of the audio playback device to be tested.

[0011] Another embodiment of the present application provides a device for determining audio quality, including:

[0012] A first acquisition module is used to acquire a first recording signal obtained by recording the first audio data played by the audio playback device to be tested;

[0013] A second acquisition module, configured to acquire a timbre reference signal and a spatial reference signal corresponding to the audio playback device to be tested;

[0014] a determination module, configured to determine a first timbre feature corresponding to the first recording signal based on the timbre reference signal and the first recording signal, and to determine a first spatial feature corresponding to the first recording signal based on the spatial reference signal and the first recording signal;

[0015] The processing module is configured to input the first timbre feature and the first spatial feature into the trained perception model to obtain a first timbre quality score and a first spatial quality score of the audio playback device to be tested.

[0016] Another embodiment of the present application provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, the method described in the above aspect is implemented.

[0017] Another aspect of the present application provides a non-transitory computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the method described in the aforementioned aspect is implemented.

[0018] Another embodiment of the present application provides a computer program product having a computer program stored thereon, which implements the method described in the above aspect when the program is executed by a processor.

[0019] The audio quality determination method, device, electronic device, and storage medium proposed in this application obtain a first recording signal obtained by recording a preset first audio data played by an audio playback device to be tested. The method of obtaining the first recording signal through actual measurement improves the accuracy of obtaining the first recording signal compared to the method of obtaining the first recording signal through synthesis. Furthermore, a timbre reference signal and a spatial reference signal corresponding to the audio playback device to be tested are obtained. A first timbre feature corresponding to the first recording signal is determined based on the timbre reference signal and the first recording signal, and a first spatial feature corresponding to the first recording signal is determined based on the spatial reference signal and the first recording signal. The first timbre feature and the first spatial feature are input into a trained perceptual model to obtain a first timbre quality score and a first spatial quality score of the audio playback device to be tested. By using a timbre reference signal and a spatial reference signal that match the audio playback device to be tested, the reference signal is no longer fixed, thereby improving the accuracy of subsequent quality scores. At the same time, a multi-dimensional quality assessment of the spatial and timbre aspects of the audio playback device to be tested is performed, thereby improving the accuracy of the quality assessment.

[0020] Additional aspects and advantages of the present application will be given in part in the description below, and in part will become apparent from the description below, or will be learned through practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] The above and / or additional aspects and advantages of the present application will become apparent and easily understood from the following description of the embodiments in conjunction with the accompanying drawings, in which:

[0022] Figure 1 A flowchart of a method for determining audio quality provided in an embodiment of the present application;

[0023] Figure 2 A flowchart of another method for determining audio quality provided in an embodiment of the present application;

[0024] Figure 3 A flowchart of a method for training a perception model provided in an embodiment of the present application;

[0025] Figure 4 A schematic diagram of the structure of an audio quality determination device provided in an embodiment of the present application;

[0026] Figure 5 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0027] The following describes in detail embodiments of the present application, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present application, and should not be construed as limiting the present application.

[0028] The following describes the audio quality determination method, device, electronic device, and storage medium according to embodiments of the present application with reference to the accompanying drawings.

[0029] Figure 1 A flowchart of a method for determining audio quality provided in an embodiment of the present application.

[0030] In the embodiment of the present application, the audio quality determination method is configured in an audio quality determination device as an example. The audio quality determination device can be applied to any electronic device so that the electronic device can perform the audio quality determination function.

[0031] Among them, the electronic device can be any device with computing capabilities, such as a personal computer, mobile terminal, server (or cloud), etc. The mobile terminal can be, for example, a vehicle-mounted device, a mobile phone, a tablet computer, a personal digital assistant, a wearable device, and other hardware devices with various operating systems, touch screens and / or display screens.

[0032] like Figure 1 As shown, the method may include the following steps:

[0033] Step 101 : obtaining first recording data obtained by recording preset first audio data played by an audio playback device to be tested.

[0034] The audio playback device to be tested may be a car stereo or multi-channel audio system, a home theater, or an audio / video conferencing system, etc., which is not limited in this embodiment. The first audio data is preset audio data, which may be a special test signal (noise, sweep signal), or a representative music clip, etc.

[0035] As an implementation method, a sound collection device is placed within a target area covered by the sound of first audio data played by the audio playback device to be tested. The audio data is recorded using the sound collection settings to obtain a first recording signal. The first recording signal can be a binaural recording signal, which includes two-channel recording signals, namely, a left channel recording signal and a right channel recording signal. In this application, the actual binaural recording signal is obtained through recording. Compared with the related art method of synthesizing the first recording signal through the head-related transfer function (HRTF), this method is not limited by the limited sound source distance and spatial sampling resolution of the HRTF database, thereby improving the accuracy of obtaining the first recording signal. In the scenario of testing the in-vehicle audio system in a vehicle, the target area can be a car seat, such as the driver's seat; in the scenario of testing a home theater, the target area can be the center of the sofa in the home theater, etc. The sound collection device is an artificial head (HATS) with a built-in microphone. The collected first recording signal is subjected to some preprocessing, including alignment with an internal reference signal, segment truncation, and level amplitude normalization.

[0036] Step 102: Acquire a timbre reference signal and a spatial reference signal corresponding to the audio playback device to be tested.

[0037] In the embodiment of the present application, the timbre reference signal and the spatial reference signal are both timbre speech signals and spatial speech signals in a relatively ideal state or state. Among them, the spatial reference signal is a recording signal related to the spatial attributes of audio quality. The spatial attributes of audio quality refer to those characteristics that can affect the listener's perception of the spatial positioning, spatial sense and overall immersion of the sound, such as stereo positioning, surround sound, spatial reverberation, sound field width and depth, etc. The timbre reference signal is a recording signal related to timbre attributes, and has timbre characteristics, which are characteristics used to describe the unique quality of sound in music and acoustics. It allows different instruments or voices to be distinguished even when playing notes of the same pitch and loudness. Timbre is determined by factors such as the frequency composition, harmonic structure, dynamic range, and envelope (attack, decay, sustain and release) of the sound. Among them, small distortion means that the distortion is less than the set threshold.

[0038] In the embodiment of the present application, the timbre reference signal and the spatial reference signal are not fixed, but are regularly updated as the audio playback device progresses, so as to obtain the timbre reference signal and the spatial reference signal corresponding to the audio playback device to be tested, thereby improving the accuracy of the reference recording signal and making the predicted score more in line with the new application scenario, thereby avoiding the problem of reduced accuracy of subsequent quality assessment due to the large difference between the audio playback device to be tested and the fixed reference recording signal used.

[0039] Step 103 : determining a first timbre feature corresponding to the first recording signal based on the timbre reference signal and the first recording signal, and determining a first spatial feature corresponding to the first recording signal based on the spatial reference signal and the first recording signal.

[0040] The timbre reference signal and the spatial reference signal refer to ideal or distortion-free recording signals, and are used to determine the quality difference in timbre and spatiality between the recording signal and the first recording signal to be evaluated.

[0041] In one implementation of the embodiment of the present application, a timbre reference signal and a first recording signal are input into a model of the international standard ITU-R BS.1387 Perceptual Evaluation of Audio Quality (PEAQ). The model is used to calculate the difference in timbre between the first recording signal and the timbre reference signal to obtain a first timbre feature corresponding to the first recording signal.

[0042] In one implementation of an embodiment of the present application, the spatial reference signal and the first recorded signal are input into a pre-trained recognition model, which is a multilayer perceptron (MLP) neural network. The spatial quality difference between the first recorded signal and the spatial reference signal is determined through the recognition model, and the quality difference is used as the first spatial feature corresponding to the first recorded signal.

[0043] Step 104 : Input the first timbre feature and the first spatial feature into the trained perception model to obtain a first timbre quality score and a first spatial quality score of the audio playback device to be tested.

[0044] Among them, the perception model adopts a linear regression model or a neural network model.

[0045] In an embodiment of the present application, the first timbre feature and the first spatial feature are input into a trained perceptual model, i.e., the first timbre feature and the first spatial feature are analyzed separately by the perceptual model. Since the first timbre feature indicates timbre difference information between the first recording signal and the timbre reference signal, a first timbre quality score of the audio playback device to be tested is determined based on this difference information. Since the first spatial feature indicates spatial difference information between the first recording signal and the spatial reference signal, a first spatial quality score of the audio playback device to be tested is determined based on this difference information.

[0046] It should be understood that since the timbre reference signal and the spatial reference signal serve as a unified reference standard, the first timbre quality score in terms of timbre and the first spatial quality in terms of space of the audio playback device to be tested are determined. The first timbre quality score and the first spatial quality score are both differences relative to the unified evaluation standard. The use of a unified evaluation standard ensures that there is a consistent benchmark for comparison between different tests, thereby improving the reliability and comparability of the evaluation results.

[0047] In the method for determining audio quality of the embodiment of the present application, a first recording signal is obtained by recording a preset first audio data played by the audio playback device to be tested. The method of obtaining the first recording signal by actual measurement improves the accuracy of obtaining the first recording signal compared to the method of obtaining the first recording signal by synthesis. Then, a timbre reference signal and a spatial reference signal corresponding to the audio playback device to be tested are obtained. A first timbre feature corresponding to the first recording signal is determined based on the timbre reference signal and the first recording signal, and a first spatial feature corresponding to the first recording signal is determined based on the spatial reference signal and the first recording signal. The first timbre feature and the first spatial feature are input into the trained perceptual model to obtain a first timbre quality score and a first spatial quality score of the audio playback device to be tested. By using a timbre reference signal and a spatial reference signal that match the audio playback device to be tested, the reference signal is no longer fixed, thereby improving the accuracy of subsequent quality scores. At the same time, a multi-dimensional quality assessment of the spatial and timbre aspects of the audio playback device to be tested is performed, thereby improving the accuracy of the quality assessment.

[0048] Based on the above embodiments, Figure 2 A flowchart of another method for determining audio quality provided in an embodiment of the present application is shown as follows: Figure 2 As shown, the method comprises the following steps:

[0049] Step 201 : Acquire a first recording signal obtained by recording a preset first audio data played by an audio playback device to be tested.

[0050] Among them, step 201 can refer to the explanation in the above embodiment, the principle is the same, and it will not be repeated here.

[0051] In step 202 , the second audio data played by the first reference audio playback device is recorded to obtain a second recording signal, and the third audio data played by the second reference audio playback device is recorded to obtain a third recording signal.

[0052] The fact that the timbre quality of the first reference audio playback device is greater than a set timbre quality threshold means that it has been previously determined through testing that the timbre quality of the first reference audio playback device is greater than the set timbre quality threshold. The first reference audio playback device is one that, when playing a sound source, causes less linear or nonlinear distortion to the sound source and outputs a higher timbre quality voice signal. The set timbre quality threshold is typically a relatively high value to identify whether the timbre quality of the audio played by the audio playback device meets requirements, thereby determining whether the audio playback device can serve as a reference audio playback device.

[0053] Among them, the spatial quality of the second reference audio playback device is greater than the set spatial quality threshold, which means that it has been determined in advance through testing that the spatial quality of the second reference audio playback device is greater than the set spatial quality threshold. The second reference audio playback device refers to a device that can better present the expected sound field, immersive feeling, etc. when playing the sound source. The set spatial quality threshold is usually a higher value, which is used to identify whether the spatial quality of the audio playback device meets the requirements, so as to determine whether the audio playback device can be used as a reference audio playback device.

[0054] It should be noted that the first reference audio playback device may be the same audio playback device, which needs to simultaneously meet the requirements that the spatial quality is greater than a set spatial quality threshold and the timbre quality is greater than a set timbre quality threshold.

[0055] The second audio data and the third audio data correspond to different audio formats of the same audio segment. For example, the second audio data is a stereo version of audio segment A, and the third audio data is a Dolby version of audio segment A. The first reference audio playback device is controlled to play the second audio data, and a second recording signal is obtained through artificial head recording. The second reference audio playback device is controlled to play the third audio data, and a third recording signal is obtained through artificial head recording. The second and third recording signals have different data, with the second recording signal carrying more of the first timbre characteristics, and the third recording signal carrying more of the first spatial characteristics.

[0056] Step 203: Use the left channel data or the right channel data in the second recording signal as a timbre reference signal.

[0057] In an embodiment of the present application, the second recording signal is a dual-channel voice signal. By performing signal separation on the second recording signal, left channel data and right channel data can be separated. Since the first timbre characteristics of the left and right channels are not much different, the left channel data or the right channel data in the second recording signal can be used as a timbre reference signal.

[0058] Step 204: Use the third recorded signal as a spatial reference signal.

[0059] In an embodiment of the present application, the third recording signal is a dual-channel speech signal. Since there are differences in the first spatial features of the left and right channels, the third recording signal is not separated, and the dual-channel speech signal is used as the spatial reference signal, thereby improving the accuracy of the spatial reference signal.

[0060] It should be understood that with the development of electronic technology, audio coding and decoding technology is constantly changing, the number of audio channels is also increasing, etc. Therefore, it is necessary to update the spatial reference signal and sound quality reference signal used as a reference in a timely manner to avoid the situation where the configuration gap between the audio playback device to be tested and the reference audio playback device becomes larger and larger. The quality score is evaluated based on the timbre reference signal or spatial reference signal generated by the reference audio playback device, which will make the score output by the subsequent perception model increasingly unreliable. The spatial reference signal and sound quality reference signal of this application are constantly updated, which improves the accuracy and reliability of the quality score.

[0061] Step 205 : determining a first timbre feature corresponding to the first recording signal based on the timbre reference signal and the first recording signal, and determining a first spatial feature corresponding to the first recording signal based on the spatial reference signal and the first recording signal.

[0062] Among them, step 205 can refer to the relevant explanations in the above-mentioned embodiment, and the principle is the same, so it will not be repeated here.

[0063] Step 206 : Input the first timbre feature and the first spatial feature into the trained perception model to obtain a first timbre quality score and a first spatial quality score of the audio playback device to be tested.

[0064] In one implementation of the present embodiment, the perception model includes a timbre perception network and a spatial perception network. A first timbre feature is input into the trained timbre perception network to obtain a first timbre quality score for the audio playback device under test. A first spatial feature is input into the trained spatial perception network to obtain a first spatial quality score for the audio playback device under test. The explanations regarding the first timbre quality score and the first spatial quality score in the previous embodiment also apply to this embodiment, and the principles are the same, so they are not further elaborated here.

[0065] The first timbre quality score and first spatial quality score of the tested audio playback device can be used to generate an evaluation report for use by the technical department or quality control department in quality adjustments and optimizations. This can be applied in scenarios such as: evaluation organizations or vehicle quality departments evaluating the sound quality of in-car audio systems across multiple models; or audio algorithm development teams and tuners evaluating whether audio quality has improved after algorithm optimization or tuning, among other scenarios, which are not limited in this embodiment.

[0066] The method for determining audio quality in an embodiment of the present application obtains a first recording signal obtained by recording a preset first audio data played by an audio playback device to be tested. The method of obtaining the first recording signal by actual measurement improves the accuracy of obtaining the first recording signal compared to the method of obtaining the first recording signal by synthesis. Then, a timbre reference signal and a spatial reference signal corresponding to the audio playback device to be tested are obtained. A first timbre feature corresponding to the first recording signal is determined based on the timbre reference signal and the first recording signal, and a first spatial feature corresponding to the first recording signal is determined based on the spatial reference signal and the first recording signal. The first timbre feature and the first spatial feature are input into a trained perceptual model to obtain a first timbre quality score and a first spatial quality score of the audio playback device to be tested. By using a timbre reference signal and a spatial reference signal that match the audio playback device to be tested, the reference signal is no longer fixed, thereby improving the accuracy of subsequent quality scores. At the same time, a multi-dimensional quality assessment of the spatial and timbre aspects of the audio playback device to be tested is performed, thereby improving the accuracy of the quality assessment.

[0067] Based on the above embodiments, the present invention provides a method for training a perception model. Figure 3 A flowchart of a method for perceptual model training provided in an embodiment of the present application specifically illustrates a method for perceptual model training for determining audio quality, such as Figure 3 As shown, the method comprises the following steps:

[0068] Step 301: Obtain sample data.

[0069] The sample data includes sample timbre features and sample space features.

[0070] As an implementation method, the sample timbre features and the sample space features are stored in a database. The sample timbre features can be generated by obtaining a timbre reference signal as a reference and a sample recording signal as an evaluation, and determining the sample timbre features based on the timbre reference signal and the sample recording signal. The sample space features can be generated by obtaining a space reference signal as a reference and a sample recording signal as an evaluation, and determining the sample space features based on the space reference signal and the sample recording signal. The timbre reference signal and the space reference signal can be referred to the relevant explanations in the aforementioned embodiments, and the principles are the same, so they will not be repeated here.

[0071] The sample data carries label data. The label data is obtained by equalizing the timbre reference signal and the sample recording signal and using it as listening material for the listening experiment. As an implementation method, the listening experiment follows mainstream standard processes, such as the Multiple Stimuli with Hidden Reference and Anchor (MUSHRA) method and Category Comparison Rating (CCR), allowing listeners with listening ability to give subjective scores. The subjective scores indicate the quality difference between the audio playback device corresponding to the sample recording signal and the reference audio playback device.

[0072] Furthermore, the sample timbre features, sample spatial features and corresponding subjective scores are stored in a database for training the perception model.

[0073] In step 302 , the sample timbre features are input into a timbre perception network to predict a second timbre quality score, and the sample spatial features are input into a spatial perception model to predict a second spatial quality score.

[0074] Among them, the second timbre quality score and the second space quality score can refer to the relevant explanations in the above embodiments, and the principles are the same, so they will not be repeated here.

[0075] Step 303: Adjust the parameters of the perception model based on the difference between the second timbre quality score and the subjective timbre quality score in the label information carried by the sample data, and the difference between the second spatial quality score and the subjective spatial quality score in the label information to obtain a trained perception model.

[0076] In one implementation of the embodiment of the present application, a timbre loss function is determined based on the difference between the second timbre quality score and the subjective timbre quality score in the label information, a spatial loss function is determined based on the difference between the second spatial quality score and the subjective spatial quality score in the label information, and a target loss function is determined based on the timbre loss function and the spatial loss function. The target loss function is used to adjust the parameters of the perception model to obtain a trained perception model.

[0077] The perceptual model parameter adjustment process can be repeated multiple times, each time using different training samples. Training is terminated when the loss function is less than a threshold, or when the number of repetitions exceeds a threshold. The perceptual model obtained after the final model parameter adjustment is used as the trained perceptual model to improve the effectiveness of model training.

[0078] It should be noted that the data in the database will be continuously streamlined with the development of audio technology, that is, outdated data that is far away from the current acoustic configuration will be deleted, and data that is close to the current acoustic configuration will be expanded to achieve continuous updating of the database. The trained perception model will then be trained or fine-tuned based on the updated database, making the quality results of the perception model recognition more accurate.

[0079] In the training method of the perceptual model of the embodiment of the present application, by continuously updating the database, the trained perceptual model is trained or fine-tuned based on the updated database, so that the quality results recognized by the perceptual model are more accurate, and then the quality of the voice signal output by the audio playback device to be tested is evaluated by the perceptual model. The output quality score result is an objective evaluation result, which is not affected by human subjective factors. The subjective score results of human beings are fully considered in the training process of the perceptual model, so that the quality score results output by the trained perceptual model are correlated with the subjective evaluation results, which can more fully reflect the listener's perception of sound quality and improve the accuracy and reliability of quality evaluation. In order to realize the above embodiment, the embodiment of the present application also proposes an audio quality determination device.

[0080] Figure 4 A schematic diagram of the structure of an audio quality determination device provided in an embodiment of the present application.

[0081] like Figure 4 As shown, the device may include:

[0082] The first acquisition module 41 is configured to acquire a first recording signal obtained by recording the audio playback device to be tested playing preset first audio data.

[0083] The second acquisition module 42 is configured to acquire a timbre reference signal and a spatial reference signal corresponding to the audio playback device to be tested.

[0084] The determination module 43 is used to determine the first timbre feature corresponding to the first recording signal based on the timbre reference signal and the first recording signal, and to determine the first spatial feature corresponding to the first recording signal based on the spatial reference signal and the first recording signal.

[0085] The processing module 44 is configured to input the first timbre feature and the first spatial feature into the trained perception model to obtain a first timbre quality score and a first spatial quality score of the audio playback device to be tested.

[0086] Furthermore, in one implementation of the embodiment of the present application, the perception model includes a timbre perception network and a spatial perception network, and the processing module 44 is configured to:

[0087] Inputting the first timbre feature into the timbre perception network to obtain a first timbre quality score of the audio playback device to be tested;

[0088] The first spatial feature is input into the spatial perception network to obtain a first spatial quality score of the audio playback device to be tested.

[0089] In one implementation of the embodiment of the present application, the second obtaining module 42 is further configured to:

[0090] Recording second audio data played by the first reference audio playback device to obtain a second recording signal, and recording third audio data played by the second reference audio playback device to obtain a third recording signal; wherein the timbre quality of the first reference audio playback device is greater than a set timbre quality threshold; and the spatial quality of the second reference audio playback device is greater than a set spatial quality threshold; and wherein the second audio data and the third audio data correspond to different audio formats of the same audio segment;

[0091] using the left channel data or the right channel data in the second recording signal as the timbre reference signal;

[0092] The third recorded signal is used as the spatial reference signal.

[0093] In one implementation of the embodiment of the present application, the apparatus further includes a training module for training the perception model, the training module being configured to:

[0094] Acquire sample data; wherein the sample data includes sample timbre features and sample space features;

[0095] Inputting the sample timbre feature into the timbre perception network to predict a second timbre quality score, and inputting the sample space feature into the space perception model to predict a second space quality score;

[0096] Based on the difference between the second timbre quality score and the subjective timbre quality score in the label information carried by the sample data, and the difference between the second spatial quality score and the subjective spatial quality score in the label information, the parameters of the perceptual model are adjusted to obtain the trained perceptual model.

[0097] In one implementation of the embodiment of the present application, the training module is further configured to:

[0098] determining a timbre loss function according to a difference between the second timbre quality score and the subjective timbre quality score in the label information;

[0099] determining a spatial loss function according to a difference between the second spatial quality score and the subjective spatial quality score in the label information;

[0100] Determining a target loss function according to the timbre loss function and the spatial loss function;

[0101] The target loss function is used to adjust the parameters of the perception model to obtain the trained perception model.

[0102] It should be noted that the above explanation of the method embodiment is also applicable to the device of this embodiment and will not be repeated here.

[0103] The audio quality determination device of the embodiment of the present application obtains a first recording signal obtained by recording a preset first audio data played by an audio playback device to be tested. The method of obtaining the first recording signal by actual measurement improves the accuracy of obtaining the first recording signal compared to the method of obtaining the first recording signal by synthesis. Then, a timbre reference signal and a spatial reference signal corresponding to the audio playback device to be tested are obtained. A first timbre feature corresponding to the first recording signal is determined based on the timbre reference signal and the first recording signal, and a first spatial feature corresponding to the first recording signal is determined based on the spatial reference signal and the first recording signal. The first timbre feature and the first spatial feature are input into the trained perceptual model to obtain a first timbre quality score and a first spatial quality score of the audio playback device to be tested. By using a timbre reference signal and a spatial reference signal that match the audio playback device to be tested, the reference signal is no longer fixed, thereby improving the accuracy of subsequent quality scores. At the same time, a multi-dimensional quality assessment of the spatial and timbre aspects of the audio playback device to be tested is performed, thereby improving the accuracy of the quality assessment.

[0104] In order to implement the above embodiments, the present application also proposes an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, the method described in the above method embodiments is implemented.

[0105] In order to implement the above embodiments, the present application also proposes a non-transitory computer-readable storage medium on which a computer program is stored. When the program is executed by a processor, the method described in the above method embodiments is implemented.

[0106] In order to implement the above embodiments, the present application further proposes a computer program product on which a computer program is stored. When the computer program is executed by a processor, the method described in the above method embodiments is implemented.

[0107] Figure 5This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. For example, the electronic device 800 can be a mobile phone, a computer, a digital broadcast terminal, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc.

[0108] Reference Figure 5 , the electronic device 800 may include one or more of the following components: a processing component 802 , a memory 804 , a power component 806 , a multimedia component 808 , an audio component 810 , an input / output (I / O) interface 812 , a sensor component 814 , and a communication component 816 .

[0109] The processing component 802 generally controls the overall operation of the electronic device 800, such as operations associated with display, phone calls, data communications, camera operation, and recording operations. The processing component 802 may include one or more processors 820 to execute instructions to perform all or part of the steps of the above-described method. In addition, the processing component 802 may include one or more modules to facilitate interaction between the processing component 802 and other components. For example, the processing component 802 may include a multimedia module to facilitate interaction between the multimedia component 808 and the processing component 802.

[0110] The memory 804 is configured to store various types of data to support operations on the electronic device 800. Examples of such data include instructions for any application or method operating on the electronic device 800, contact data, phone book data, messages, pictures, videos, etc. The memory 804 can be implemented by any type of volatile or non-volatile storage device, or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk, or optical disk.

[0111] The power component 806 provides power to the various components of the electronic device 800. The power component 806 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to the electronic device 800.

[0112] The multimedia component 808 includes a screen that provides an output interface between the electronic device 800 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen can be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touch, slide, and gestures on the touch panel. The touch sensor can not only sense the boundaries of the touch or slide action, but also detect the duration and pressure associated with the touch or slide operation. In some embodiments, the multimedia component 808 includes a front camera and / or a rear camera. When the electronic device 800 is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera can receive external multimedia data. Each front camera and rear camera can be a fixed optical lens system or have a focal length and optical zoom capability.

[0113] The audio component 810 is configured to output and / or input audio signals. For example, the audio component 810 includes a microphone (MIC), which is configured to receive external audio signals when the electronic device 800 is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signal can be further stored in the memory 804 or transmitted via the communication component 816. In some embodiments, the audio component 810 also includes a speaker for outputting audio signals.

[0114] I / O interface 812 provides an interface between processing component 802 and peripheral interface modules, such as a keyboard, click wheel, buttons, etc. These buttons may include but are not limited to: a home button, volume buttons, a start button, and a lock button.

[0115] The sensor assembly 814 includes one or more sensors for providing various aspects of status assessment for the electronic device 800. For example, the sensor assembly 814 can detect the open / closed state of the electronic device 800, the relative positioning of components, such as the display and keypad of the electronic device 800. The sensor assembly 814 can also detect changes in the position of the electronic device 800 or a component of the electronic device 800, the presence or absence of user contact with the electronic device 800, the orientation or acceleration / deceleration of the electronic device 800, and temperature changes of the electronic device 800. The sensor assembly 814 may include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor assembly 814 may also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor assembly 814 may also include an accelerometer, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.

[0116] The communication component 816 is configured to facilitate wired or wireless communication between the electronic device 800 and other devices. The electronic device 800 can access a wireless network based on a communication standard, such as WiFi, 4G or 5G, or a combination thereof. In an exemplary embodiment, the communication component 816 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 816 also includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology and other technologies.

[0117] In an exemplary embodiment, the electronic device 800 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the above methods.

[0118] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 804 including instructions, and the instructions can be executed by the processor 820 of the electronic device 800 to perform the above method. For example, the non-transitory computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc.

[0119] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art can combine and combine different embodiments or examples described in this specification and features of different embodiments or examples without contradiction.

[0120] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of the technical features being referred to. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of such features. Throughout the description of this application, "plurality" means at least two, for example, two, three, etc., unless otherwise specifically defined.

[0121] Any process or method description in a flowchart or otherwise described herein may be understood to represent a module, segment or portion of code comprising one or more executable instructions for implementing the steps of a custom logical function or process, and the scope of the preferred embodiments of the present application includes alternative implementations in which functions may be performed out of the order shown or discussed, including performing functions in a substantially simultaneous manner or in the reverse order depending on the functions involved, which should be understood by those skilled in the art to which the embodiments of the present application belong.

[0122] The logic and / or steps represented in the flowcharts or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing the logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (e.g., a computer-based system, a system including a processor, or other system that can fetch and execute instructions from an instruction execution system, apparatus, or device). For purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by, or in conjunction with, an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of computer-readable media include the following: an electrical connection with one or more wires (electronic devices), a portable computer disk cartridge (magnetic device), random access memory (RAM), read-only memory (ROM), erasable and programmable read-only memory (EPROM or flash memory), fiber optic devices, and a portable compact disc read-only memory (CDROM). Furthermore, the computer-readable medium may even be paper or other suitable medium on which the program is printed, since the program may be obtained electronically, for example, by optically scanning the paper or other medium and then editing, interpreting or processing it in another suitable manner if necessary, and then storing it in a computer memory.

[0123] It should be understood that various parts of the present application can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented using software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented using hardware, as in another embodiment, any one of the following technologies known in the art or a combination thereof can be used to implement: a discrete logic circuit having a logic gate circuit for implementing a logic function on a data signal, an application-specific integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.

[0124] Those skilled in the art will understand that all or part of the steps in the method of the above embodiment can be completed by instructing related hardware through a program, and the program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiment.

[0125] In addition, the functional units in the various embodiments of the present application may be integrated into a processing module, or each unit may exist physically separately, or two or more units may be integrated into a module. The above-mentioned integrated module may be implemented in the form of hardware or in the form of a software functional module. If the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium.

[0126] The storage medium mentioned above may be a read-only memory, a magnetic disk, or an optical disk, etc. Although the embodiments of the present application have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting the present application. Persons skilled in the art may make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present application.

Claims

1. A method for determining audio quality, characterized in that: include: Acquire a first recording signal obtained by recording a preset first audio data played by the audio playback device to be tested; Acquire a timbre reference signal and a spatial reference signal corresponding to the audio playback device to be tested; determining a first timbre feature corresponding to the first recording signal based on the timbre reference signal and the first recording signal, and determining a first spatial feature corresponding to the first recording signal based on the spatial reference signal and the first recording signal; The first timbre feature and the first spatial feature are input into the trained perception model to obtain a first timbre quality score and a first spatial quality score of the audio playback device to be tested.

2. The method according to claim 1, wherein The perception model includes a timbre perception network and a spatial perception network, and inputting the first timbre feature and the first spatial feature into the trained perception model to obtain a first timbre quality score and a first spatial quality score of the audio playback device to be tested includes: Inputting the first timbre feature into the timbre perception network to obtain a first timbre quality score of the audio playback device to be tested; The first spatial feature is input into the spatial perception network to obtain a first spatial quality score of the audio playback device to be tested.

3. The method according to claim 1 or 2, wherein: The acquiring of a timbre reference signal and a spatial reference signal corresponding to the audio playback device to be tested includes: Recording second audio data played by the first reference audio playback device to obtain a second recording signal, and recording third audio data played by the second reference audio playback device to obtain a third recording signal; wherein the timbre quality of the first reference audio playback device is greater than a set timbre quality threshold; and the spatial quality of the second reference audio playback device is greater than a set spatial quality threshold; and wherein the second audio data and the third audio data correspond to different audio formats of the same audio segment; using the left channel data or the right channel data in the second recording signal as the timbre reference signal; The third recorded signal is used as the spatial reference signal.

4. The method according to claim 2, wherein The training method of the perception model includes: Acquire sample data; wherein the sample data includes sample timbre features and sample space features; Inputting the sample timbre feature into the timbre perception network to predict a second timbre quality score, and inputting the sample space feature into the space perception model to predict a second space quality score; Based on the difference between the second timbre quality score and the subjective timbre quality score in the label information carried by the sample data, and the difference between the second spatial quality score and the subjective spatial quality score in the label information, the parameters of the perceptual model are adjusted to obtain the trained perceptual model.

5. The method according to claim 4, wherein The step of adjusting the parameters of the perceptual model based on a difference between the second timbre quality score and the subjective timbre quality score in the label information carried by the sample data, and a difference between the second spatial quality score and the subjective spatial quality score in the label information, to obtain the trained perceptual model, includes: determining a timbre loss function according to a difference between the second timbre quality score and the subjective timbre quality score in the label information; determining a spatial loss function according to a difference between the second spatial quality score and the subjective spatial quality score in the label information; Determining a target loss function according to the timbre loss function and the spatial loss function; The target loss function is used to adjust the parameters of the perception model to obtain the trained perception model.

6. A device for determining audio quality, characterized in that: include: A first acquisition module is used to acquire a first recording signal obtained by recording the first audio data played by the audio playback device to be tested; A second acquisition module, configured to acquire a timbre reference signal and a spatial reference signal corresponding to the audio playback device to be tested; a determination module, configured to determine a first timbre feature corresponding to the first recording signal based on the timbre reference signal and the first recording signal, and to determine a first spatial feature corresponding to the first recording signal based on the spatial reference signal and the first recording signal; The processing module is configured to input the first timbre feature and the first spatial feature into the trained perception model to obtain a first timbre quality score and a first spatial quality score of the audio playback device to be tested.

7. The device according to claim 6, characterized in that The perception model includes a timbre perception network and a space perception network, and the processing module is used to: Inputting the first timbre feature into the timbre perception network to obtain a first timbre quality score of the audio playback device to be tested; The first spatial feature is input into the spatial perception network to obtain a first spatial quality score of the audio playback device to be tested.

8. An electronic device, characterized in that: The method comprises a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the method according to any one of claims 1 to 5 is implemented.

9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 5 is implemented.

10. A computer program product, characterized in that The invention comprises a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 5.