Method for training voice quality detection model, voice quality detection method, electronic device and medium
The audio quality detection model is trained using voice activity detection and feature extraction with MOS labels to assess audio quality accurately without clean reference audio, addressing limitations in existing methods and enhancing detection efficiency.
Patent Information
- Application Number
- CN202210333127.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-31
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2042-03-31
AI Technical Summary
In the prior art, audio quality detection is difficult to accurately evaluate in the absence of pure clean audio, especially the problems of audio time misalignment and spectrum distortion in VOIP network communication.
By obtaining the initial training audio and average opinion score tags, voice endpoint detection is performed to remove non-vocals, feature extraction is performed, and the sound quality detection model is trained using convolutional neural network, long-term memory network and full-connection layer, and the loss value is calculated to adjust the model parameters until the training completion conditions are met.
Accurate evaluation of audio quality without pure clean audio is achieved, expanding the detection range and improving detection efficiency.
Smart Images

Figure CN114694678B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of audio processing, and particularly to a method for training a sound quality detection model, a sound quality detection method, an electronic device, and a computer-readable storage medium. Background Art
[0002] Currently, the PESQ (Perceptual Evaluation of Speech Quality) method is usually used to detect the audio sound quality and obtain a detection result representing the quality of the sound. The PESQ method is usually for audio in VOIP (Voice over Internet Protocol) network communication and can evaluate problems such as audio time misalignment and spectrum distortion caused by frame loss, jitter, etc. during the network transmission of the audio signal. Calculating the PESQ score requires preparing the pure clean audio corresponding to the noisy audio, and this sound quality evaluation method is called reference-based sound quality evaluation. However, in applications, it is difficult to obtain pure clean audio, making it difficult to perform the sound quality detection of most audio. Summary of the Invention
[0003] In view of this, the purpose of the present application is to provide a method for training a sound quality detection model, a sound quality detection method, an electronic device, and a computer-readable storage medium to accurately evaluate the quality of the audio to be tested.
[0004] To solve the above technical problems, in a first aspect, the present application provides a method for training a sound quality detection model, including:
[0005] Obtaining an initial training audio and a corresponding mean opinion score label; the mean opinion score label is used to represent the average sound quality evaluation parameter obtained after multiple evaluation objects evaluate the sound quality of the initial training audio;
[0006] Performing non-human voice filtering processing on the initial training audio based on voice activity detection to obtain a training audio;
[0007] Performing feature extraction processing on the training audio to obtain training features;
[0008] Inputting the training features into an initial model to obtain a corresponding training sound quality detection result;
[0009] Calculating a loss value using the training sound quality detection result and the mean opinion score label, and adjusting the model parameters of the initial model using the loss value;
[0010] When it is detected that the training completion condition is met, the adjusted initial model is determined as the sound quality detection model.
[0011] Optionally, the process of obtaining the mean opinion score label includes:
[0012] Playing the initial training audio to each of the evaluation objects;
[0013] Receiving the initial audio quality data obtained by each of the evaluation objects after evaluating the audio quality of the initial training audio;
[0014] Generating the mean opinion score label by using each of the initial audio quality data.
[0015] Optionally, the generating the mean opinion score label by using each of the initial audio quality data includes:
[0016] Performing an averaging process on each of the initial audio quality data to obtain a first score label;
[0017] Inputting the initial training audio into an audio defect detection model to obtain a defect detection result; wherein, the audio defect detection model is used to detect audio defects that can affect the auditory perception;
[0018] Generating a second score label based on the defect detection result;
[0019] Generating the mean opinion score label by using the first score label and the second score label.
[0020] Optionally, the performing feature extraction processing on the training audio to obtain training features includes:
[0021] Resampling the training audio based on the maximum sampling rate perceptible by the human ear to obtain intermediate data;
[0022] Performing sliding window framing on the intermediate data based on a preset window length to obtain a plurality of audio frames;
[0023] Performing feature extraction processing on each of the audio frames to obtain the training features.
[0024] Optionally, the performing non-human voice filtering processing on the initial training audio based on voice activity detection to obtain the training audio includes:
[0025] Performing voice activity detection on the initial training audio to obtain voice endpoint moments;
[0026] Segmenting the initial training audio according to the voice endpoint moments to obtain a plurality of audio segments, and removing the non-human voice frequency bands in the audio segments to obtain human voice frequency bands;
[0027] Concatenating the human voice frequency bands to obtain the training audio.
[0028] Optionally, the initial model includes a convolutional neural network, a long short-term memory network, a fully connected layer, and an average pooling layer;
[0029] The step of inputting the training features into the initial model to obtain corresponding training speech quality detection results includes:
[0030] Inputting the training features into the convolutional neural network to obtain training intermediate features;
[0031] Inputting the training intermediate features into the long short-term memory network to obtain training initial detection results;
[0032] Inputting the training initial detection results into the fully connected layer to obtain training intermediate detection results;
[0033] Inputting the training intermediate detection results into the average pooling layer to obtain the training speech quality detection results.
[0034] Optionally, the step of obtaining the initial training audio and the corresponding mean opinion score label includes:
[0035] Obtaining a batch of multiple initial training audios and the mean opinion score labels from the training dataset according to a preset batch size;
[0036] Correspondingly, the step of calculating the loss value by using the training speech quality detection results and the mean opinion score labels includes:
[0037] When obtaining the training speech quality detection results corresponding to all the initial training audios in a batch, calculating the loss value by using the training speech quality detection results, the mean opinion score labels, and the training intermediate detection results in this batch.
[0038] Optionally, the step of calculating the loss value by using the training speech quality detection results, the mean opinion score labels, and the training intermediate detection results in this batch includes:
[0039] According to
[0040]
[0041] obtaining the loss value;
[0042] where the loss is the loss value, S is the preset batch size, T S is the number of frames corresponding to the training features, M S is the mean opinion score label, is the training speech quality detection result, is the value corresponding to the t-th frame in the training features in the training intermediate detection results, and α is a preset weight.
[0043] Optionally, the detection of meeting the training completion condition includes:
[0044] Determine whether the loss value is less than a preset threshold;
[0045] If so, it is determined that the training completion condition is met.
[0046] In a second aspect, the present application further provides a sound quality detection method, including:
[0047] Obtain an initial audio to be measured;
[0048] Perform non-human voice filtering processing on the initial audio to be measured based on voice activity detection to obtain an audio to be measured;
[0049] Perform feature extraction processing on the audio to be measured to obtain a feature to be measured;
[0050] Input the feature to be measured into a sound quality detection model to obtain a sound quality detection result corresponding to the initial audio to be measured.
[0051] Optionally, the sound quality detection model includes a convolutional neural network, a long short-term memory network, a fully connected layer, and an average pooling layer;
[0052] The inputting the feature to be measured into the sound quality detection model to obtain a sound quality detection result corresponding to the initial audio to be measured includes:
[0053] Input the feature to be measured into the convolutional neural network to obtain an intermediate feature to be measured;
[0054] Input the intermediate feature to be measured into the long short-term memory network to obtain an initial detection result;
[0055] Input the initial detection result into the fully connected layer to obtain an intermediate detection result;
[0056] Input the intermediate detection result into the average pooling layer to obtain the sound quality detection result.
[0057] In a third aspect, the present application further provides an electronic device, including a memory and a processor, wherein:
[0058] The memory is used to store a computer program;
[0059] The processor is used to execute the computer program to implement the above-mentioned sound quality detection model training method and / or the above-mentioned sound quality detection method.
[0060] In a fourth aspect, the present application also provides a computer-readable storage medium for storing a computer program, where the computer program, when executed by a processor, implements the above-mentioned method for training a sound quality detection model, and / or the above-mentioned sound quality detection method.
[0061] The method for training a sound quality detection model provided by the present application includes: obtaining an initial training audio and a corresponding mean opinion score label; performing interference filtering processing based on voice activity detection on the initial training audio to obtain a training audio; performing feature extraction processing on the training audio to obtain training features; inputting the training features into an initial model to obtain a corresponding training sound quality detection result; calculating a loss value using the training sound quality detection result and the mean opinion score label, and adjusting the model parameters of the initial model using the loss value; when it is detected that the training completion condition is met, determining the adjusted initial model as the sound quality detection model.
[0062] It can be seen that this method uses the mean opinion score (MOS) as the label for the initial training audio and the training audio. The mean opinion score is evaluated by a large number of listeners for the quality of audio of sentences read aloud by male or female speakers transmitted through a communication circuit. Listeners score each sentence according to the following criteria: (1) very poor (2) poor (3) fair (4) good (5) very good. MOS is an arithmetic method of all listeners' individual scores, ranging from 1 (worst) to 5 (best). The obtained MOS label can accurately represent the quality of the sentence audio in the initial training audio. Through voice activity detection, the non-human voice part of the audio can be regarded as interference and filtered out, and only the human voice part is retained as the training audio. After obtaining the training features through feature extraction, the initial model is used for quality detection, and the loss value is calculated using the obtained training sound quality detection result and the mean opinion score label. The initial model is adjusted using the loss value so that the initial model learns the correct way to evaluate the audio quality, and the obtained training sound quality detection result is as close as possible to the MOS label. After training is completed, the adjusted initial model can be determined as the sound quality detection model. The obtained sound quality detection model can accurately evaluate the quality of the audio to be measured without having pure clean audio.
[0063] In addition, the present application also provides a sound quality detection method, an electronic device, and a computer-readable storage medium, which also have the above-mentioned beneficial effects. BRIEF DESCRIPTION OF THE DRAWINGS
[0064] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained according to the provided drawings without creative efforts.
[0065] Figure 1 Schematic diagram of the hardware composition framework applicable to a method for training a sound quality detection model provided by an embodiment of the present application;
[0066] Figure 2 Schematic diagram of the hardware composition framework applicable to another method for training a sound quality detection model provided by an embodiment of the present application;
[0067] Figure 3 Flowchart of a method for training a sound quality detection model provided by an embodiment of the present application;
[0068] Figure 4 Flowchart of a method for detecting sound quality provided by an embodiment of the present application;
[0069] Figure 5 Flowchart of another method for detecting sound quality provided by an embodiment of the present application. Detailed implementation manners
[0070] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Apparently, the described embodiments are only a part rather than all of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.
[0071] For ease of understanding, first, an introduction will be given to the hardware composition framework used in the method for training a sound quality detection model provided by the embodiments of the present application and / or the corresponding solution of the sound quality detection method. Please refer to Figure 1 , Figure 1 Schematic diagram of the hardware composition framework applicable to a method for training a sound quality detection model provided by an embodiment of the present application. The electronic device 100 may include a processor 101 and a memory 102, and may further include one or more of a multimedia component 103, an information input / information output (I / O) interface 104, and a communication component 105.
[0072] Among them, the processor 101 is used to control the overall operation of the electronic device 100 to complete all or part of the steps in the method for training a sound quality detection model and / or the sound quality detection method; the memory 102 is used to store various types of data to support the operation of the electronic device 100. These data may include, for example, instructions for any application or method operating on the electronic device 100, as well as application-related data. The memory 102 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read-Only Memory (EPROM), Programmable Read-Only Memory (PROM), Read-Only Memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disc, or one or more of them. In this embodiment, the memory 102 stores at least programs and / or data for implementing the following functions:
[0073] Obtain the initial training audio and the corresponding mean opinion score label;
[0074] Perform non-human voice filtering processing on the initial training audio based on voice activity detection to obtain the training audio;
[0075] Perform feature extraction processing on the training audio to obtain training features;
[0076] Input the training features into the initial model to obtain the corresponding training sound quality detection result;
[0077] Calculate the loss value using the training sound quality detection result and the mean opinion score label, and adjust the model parameters of the initial model using the loss value;
[0078] When it is detected that the training completion condition is satisfied, determine the adjusted initial model as the sound quality detection model.
[0079] And / or,
[0080] Obtain the initial audio to be measured;
[0081] Perform non-human voice filtering processing on the initial audio to be measured based on voice activity detection to obtain the audio to be measured;
[0082] Perform feature extraction processing on the audio to be measured to obtain the features to be measured;
[0083] Input the feature to be measured into the sound quality detection model to obtain the sound quality detection result corresponding to the initial audio to be measured.
[0084] The multimedia component 103 may include a screen and an audio component. The screen may be, for example, a touch screen, and the audio component is used to output and / or input audio signals. For example, the audio component may include a microphone for receiving external audio signals. The received audio signal may be further stored in the memory 102 or transmitted through the communication component 105. The audio component further includes at least one speaker for outputting audio signals. The I / O interface 104 provides an interface between the processor 101 and other interface modules, and the other interface modules may be a keyboard, a mouse, buttons, etc. These buttons may be virtual buttons or physical buttons. The communication component 105 is used for wired or wireless communication between the electronic device 100 and other devices. Wireless communication, such as Wi-Fi, Bluetooth, Near Field Communication (NFC), 2G, 3G, or 4G, or a combination of one or more of them. Accordingly, the communication component 105 may include: a Wi-Fi component, a Bluetooth component, an NFC component.
[0085] The electronic device 100 may be implemented by one or more Application Specific Integrated Circuits (ASICs), Digital Signal Processors (DSPs), Digital Signal Processing Devices (DSPDs), Programmable Logic Devices (PLDs), Field Programmable Gate Arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components, and is used to execute the sound quality detection model training method.
[0086] Of course, Figure 1 The structure of the illustrated electronic device 100 does not limit the electronic device in the embodiments of the present application. In practical applications, the electronic device 100 may include more or fewer components than Figure 1 shown, or combine certain components.
[0087] It can be understood that the number of electronic devices is not limited in the embodiments of the present application. It may be multiple electronic devices that cooperate to complete the sound quality detection model training method and / or the sound quality detection method. In a possible implementation manner, please refer to Figure 2 ,Figure 2 This is a schematic diagram of the hardware composition framework applicable to another method for training a sound quality detection model provided by an embodiment of the present application. As can be seen from Figure 2 this, the hardware composition framework may include: a first electronic device 11 and a second electronic device 12, which are connected through a network 13.
[0088] In the embodiment of the present application, the hardware structures of the first electronic device 11 and the second electronic device 12 may refer to Figure 1 the electronic device 100 in. That is, it can be understood that there are two electronic devices 100 in this embodiment, and the two perform data interaction. Further, in the embodiment of the present application, the form of the network 13 is not limited, that is, the network 13 may be a wireless network (such as WIFI, Bluetooth, etc.) or a wired network.
[0089] Among them, the first electronic device 11 and the second electronic device 12 may be the same type of electronic device. For example, both the first electronic device 11 and the second electronic device 12 are servers; they may also be different types of electronic devices. For example, the first electronic device 11 may be a smart phone or other smart terminal, and the second electronic device 12 may be a server. In a possible implementation manner, a server with strong computing power may be used as the second electronic device 12 to improve data processing efficiency and reliability, and thus improve the processing efficiency of training the sound quality detection model. At the same time, a smart phone with low cost and wide application range is used as the first electronic device 11 to implement the interaction between the second electronic device 12 and the user. It can be understood that the interaction process may be: the user obtains an initial training audio on the smart phone and gives a corresponding mean opinion score label, and the smart phone sends the initial training audio and the mean opinion score label to the server, and the server uses the initial training audio and the MOS label to train to obtain a sound quality detection model. The server sends the sound quality detection model to the smart phone for sound quality detection on the smart phone.
[0090] Or, the sound quality detection model is deployed on the server, and the smart phone can interact with the user, obtain an initial audio to be measured, and send it to the server. The server uses the sound quality detection model to detect the initial audio to be measured, obtains a corresponding sound quality detection result, and sends the sound quality detection result to the smart phone for output to the user.
[0091] Please refer to Figure 3 , Figure 3 which is a schematic flowchart of a method for training a sound quality detection model provided by an embodiment of the present application. The method in this embodiment includes:
[0092] S101: Obtain an initial training audio and a corresponding mean opinion score label.
[0093] The Mean Opinion Score (MOS) is evaluated by a large number of listeners to assess the quality of audio of sentences read aloud by male or female speakers transmitted through a communication circuit. Listeners rate each sentence according to the following criteria: (1) very poor, (2) poor, (3) fair, (4) good, (5) very good. MOS is the arithmetic mean of all individual listener ratings, ranging from 1 (worst) to 5 (best). This evaluation method is widely used in the subjective evaluation of audio quality. However, the subjective evaluation process is time-consuming and laborious. Therefore, the PESQ (Perceptual Evaluation of Speech Quality) method is widely used in the automatic detection of audio quality (or objective evaluation) to improve the efficiency of audio quality detection. However, PESQ is a reference-based audio quality assessment and can only detect the audio to be tested when there is a pure clean audio of the audio to be tested (which can be regarded as lossless audio), resulting in a limited detection range.
[0094] In this application, the initial training audio and the corresponding MOS labels are used to form training data, and a voice quality detection model is trained to enable automatic voice quality detection even without pure clean audio, improve the detection efficiency, and expand the detection range. Specifically, the initial training audio refers to the directly obtained training data, which may contain one or more sentences of speech. The Mean Opinion Score label refers to the MOS score label obtained by manually performing MOS scoring on the speech part of the initial training audio, and is an average voice quality evaluation parameter used to characterize the voice quality evaluation obtained by multiple evaluation objects for the initial training audio. The initial training audio and the corresponding MOS labels can be prepared in advance, or when it is necessary to train the voice quality detection model, the initial training audio can be temporarily selected and MOS scoring can be performed on it by the user to obtain the corresponding MOS labels.
[0095] In one implementation, the Mean Opinion Score label can be generated when it is obtained. Specifically, the initial training audio can be played to each evaluation object. For example, the initial training audio can be sent to the electronic devices used by each evaluator and a play control instruction can be sent. The initial voice quality data obtained after each evaluation object evaluates the voice quality of the initial training audio is received. The initial voice quality data can be generated and sent by the electronic devices used by each evaluator. The Mean Opinion Score label is generated using each initial voice quality data.
[0096] This embodiment does not limit the specific manner of generating the mean opinion score. For example, the average calculation can be performed on each initial audio quality data. In another embodiment, in order to improve the reliability of the mean opinion score label, the average processing can be first performed on each initial audio quality data to obtain the first score label. In addition, the initial training audio can be input into the audio defect detection model to obtain the defect detection result; the audio defect detection model is used to detect the audio defects that can affect the auditory perception, and the types of audio defects can be set according to needs. For example, it can include plosives (generated due to the mouth being too close to the microphone, manifested as po sounds, hu sounds, snoring sounds, etc.), hissing sounds / sibilants (manifested as si sounds, chi sounds), current sounds (noises caused by abnormal hardware circuits, manifested as zizizi / wengwengweng sounds), clicking sounds (manifested as cila, click, stuttering sounds), popping sounds (referred to as clip, generated due to sound explosion, clipping, etc., and also easily generated after mixing of human voices and accompaniments), ambient noise, stuttering (manifested as short interruptions of human voices, poor connection before and after, or obvious voice swallowing during network transmission). The audio defect detection model can identify the audio defects in the initial training audio, specifically, the number of types of audio defects, the number of occurrences of each type of audio defect, etc. Based on the defect detection result, the second score label can be generated. In one embodiment, the full score can be set to 5 points, and the score can be deducted as appropriate according to the defect detection result to obtain the second score label, and the mean opinion score label can be generated using the first score label and the second score label.
[0097] S102: Perform non-human voice filtering processing on the initial training audio based on voice activity detection to obtain the training audio.
[0098] It can be understood that when the reader records the initial training audio, multiple sentences are usually spoken in it, and there is usually a time interval of voice blank or other non-human voice audio between different sentences. When performing MOS scoring, the user only evaluates the voice part and ignores the non-human voice audio. Therefore, during the training process of the audio quality detection model, the non-human voice audio in the initial training audio should be removed to avoid interference with the training of the audio quality detection model. In this embodiment, the voice activity detection method can be used to identify the start time position and end time position of the human voice, and then the non-human voice part in the initial training audio can be filtered out to obtain the training audio. Voice activity detection, that is, Voice Activity Detection, VAD, can identify the silent period of the human voice in the audio signal.
[0099] Specifically, in one embodiment, the generation process of the training audio includes:
[0100] Step 11: Perform voice activity detection on the initial training audio to obtain the voice endpoint moments.
[0101] Step 12: Segment the initial training audio according to the voice endpoint moments to obtain multiple audio segments, and remove the non-human voice frequency bands in the audio segments to obtain the human voice frequency bands.
[0102] Step 13: Concatenate the human voice frequency bands to obtain the training audio.
[0103] In this embodiment, voice endpoint detection can identify the start moment and the end moment of the voice audio (i.e., the human voice audio), and these two are the voice endpoint moments. Segment the initial training audio according to the voice endpoint moments to obtain audio segments, which include human voice frequency bands and non-human voice frequency bands, and the two types of audio segments appear alternately. Remove the non-human voice frequency bands therein and retain the human voice frequency bands. Exemplarily, the human voice frequency bands and the non-human voice frequency bands appear alternately, so the start moment and the end moment in the voice endpoint moments also appear alternately. Along the chronological order, determine the audio segments between adjacent start moments and end moments as human voice frequency bands, and determine the audio segments between adjacent end moments and start moments as non-human voice frequency bands. After the classification of the audio segments is completed, remove the non-human voice frequency bands and retain the human voice frequency bands, and then concatenate them to obtain the final training audio. Among them, the non-human voice frequency bands can be blank audio segments, or can be audio segments recording non-human voices such as background sounds.
[0104] S103: Perform feature extraction processing on the training audio to obtain training features.
[0105] In order to enable the initial model to more efficiently learn how to perform audio quality detection, perform feature extraction on the training audio to obtain corresponding training features, so as to better characterize the quality characteristics of the training audio. The specific method of feature extraction is not limited. In one embodiment, the training features can be in the form of images. For example, spectrogram, mel-frequency spectrum or other spectrum forms can be used. In another embodiment, in order to evaluate the signal in a wider frequency band range, the training audio can also be resampled to make its signal frequency better.
[0106] Specifically, the generation process of the training features may include:
[0107] Step 21: Resample the training audio based on the maximum sampling rate perceptible by the human ear to obtain intermediate data.
[0108] Step 22: Perform sliding window framing on the intermediate data based on a preset window length to obtain multiple audio frames.
[0109] Step 23: Perform feature extraction processing on each audio frame to obtain training features.
[0110] The frequency range of signals that can be heard by the human ear is relatively fixed, usually in the range of 20 Hz - 20000 Hz. The sampling frequency is usually twice the signal frequency, so the maximum sampling rate perceptible by the human ear can be determined. It can be understood that since the hearing ranges of different people are different, the upper limit of the human ear frequency range for some people can reach 22000 Hz. Therefore, the maximum sampling rate perceptible by the human ear can be higher than the sampling rate at the upper limit of the ordinary human ear frequency range, and the specific size is not limited. Exemplarily, 48 kHz can be selected as the maximum sampling rate perceptible by the human ear. Through resampling, the frequency domain range of the audio can be broadened during training to obtain corresponding intermediate data.
[0111] Sliding window framing refers to the process of sampling audio frames in chronological order using a preset analysis window. The analysis window can specifically be a Hann window, Hamming window, Blackman - Harris window, etc. The preset window length is the width of the analysis window, and the specific value is not limited. For example, when the sampling frequency is 48 kHz, the preset window length can be 21.3 ms. After each sampling of the analysis window, it slides backward by a certain distance for the next sampling, and this sliding distance is the frame shift. The specific size of the frame shift is not limited. For example, it can be half of the window length.
[0112] After obtaining the audio frames, perform feature extraction processing such as Mel - spectrum extraction and spectrogram extraction on the audio signal to obtain the audio - frame features corresponding to individual audio frames, and splice the respective audio - frame features in chronological order to obtain the corresponding training features.
[0113] S104: Input the training features into the initial model to obtain the corresponding training audio - quality detection result.
[0114] S105: Calculate the loss value using the training audio - quality detection result and the mean opinion score label, and use the loss value to adjust the model parameters of the initial model.
[0115] A comprehensive description of the above two steps is given.
[0116] The initial model refers to a model that has not been fully trained. After sufficient training and parameter adjustment, the adjusted initial model can be used as an audio - quality detection model. This embodiment does not limit the specific form and category of the initial model. For example, it can be a convolutional neural network model, or it can be a combination of a convolutional neural network and a recurrent neural network. After the processing model processes the training features, it can obtain the training audio - quality detection result of the training audio based on the current learning and parameter - tuning situation.
[0117] When not trained sufficiently, there is a certain gap between the training sound quality detection results obtained by the initial model and the truly correct results (i.e., MOS labels). By calculating the loss value and adjusting the model parameters of the initial model based on the loss value, the model can learn how to correctly perform sound quality detection and give the correct sound quality detection results.
[0118] Specifically, in one implementation, the initial model includes a convolutional neural network, a long short-term memory network, a fully connected layer, and an average pooling layer. Among them, the convolutional neural network is used to perform convolutional calculations on the input training features to extract effective local audio features. The long short-term memory network (Long Short-Term Memory, LSTM), specifically, can be a bidirectional long short-term memory network (Bi-directional Long-Short Term Memory, BLSTM), which is used to extract the temporal relationship between local audio features and learn the correlation between adjacent frames. The fully connected layer is used to predict the sound quality detection result corresponding to each frame in units of frames. The average pooling layer is used to synthesize the sound quality detection results of each frame to obtain the final training sound quality detection result.
[0119] Correspondingly, the process of inputting the training features into the initial model to obtain the corresponding training sound quality detection results can include:
[0120] Step 31: Input the training features into the convolutional neural network to obtain training intermediate features.
[0121] Step 32: Input the training intermediate features into the long short-term memory network to obtain the training initial detection results.
[0122] Step 33: Input the training initial detection results into the fully connected layer to obtain the training intermediate detection results.
[0123] Step 34: Input the training intermediate detection results into the average pooling layer to obtain the training sound quality detection results.
[0124] Among them, the training intermediate features are the features obtained after convolutional calculations, and the training initial detection results are the data obtained after the long short-term memory network extracts the temporal relationship. The training intermediate detection results are the sound quality detection results corresponding to each frame in the training features.
[0125] It should be noted that the specific calculation method of the loss value is not limited and can be calculated according to the initial model category, the manifestation form of the training sound quality detection result, the training focus direction, etc. For example, a square loss function, an exponential loss function, a cross-entropy loss function, etc. can be used. In a specific implementation manner, the loss value can be calculated by comprehensively considering the training sound quality detection results corresponding to each initial training audio within a training batch. In this case, when obtaining the initial training audio and the corresponding MOS label, a batch of multiple initial training audios and the mean opinion score labels can be obtained from the training dataset according to a preset batch size. The preset batch size refers to the number of initial training audios obtained in each training batch. Since the frequency of parameter adjustment is the same as the frequency of loss value generation, in this implementation manner, after all the initial training audios in each batch are processed, a loss value calculation and parameter adjustment are performed once.
[0126] Correspondingly, the process of calculating the loss value using the training sound quality detection result and the mean opinion score label may include:
[0127] Step 41: When the training sound quality detection results corresponding to all the initial training audios in a batch are obtained, the loss value is obtained by using the training sound quality detection results, the mean opinion score label, and the training intermediate detection results within this batch.
[0128] In this implementation manner, by comprehensively calculating the loss value using all the training sound quality detection results and the mean opinion score labels in a batch, the parameter adjustment can be performed by comprehensively considering the training situation of the entire batch. In addition, calculating the loss value using the MOS label and the training intermediate detection results can reflect the sound quality evaluation situation of each frame in the loss value.
[0129] Specifically, it can be obtained according to
[0130]
[0131] to obtain the loss value.
[0132] Among them, loss is the loss value, S is the preset batch size, T S is the number of frames corresponding to the training feature, that is, the audio frames divided when the training feature is generated, M S is the mean opinion score label, is the training sound quality detection result, is the value corresponding to the t-th frame in the training feature in the training intermediate detection result, and α is a preset weight, and its specific size is not limited. For example, it can be 1.
[0133] S106: When it is detected that the training completion condition is met, the adjusted initial model is determined as the sound quality detection model.
[0134] Execute the above steps cyclically, and after each training round is executed and the parameters are adjusted, determine whether the training completion condition is met. The training completion condition refers to the condition indicating that the initial model has been sufficiently trained, which can specifically be a condition that limits the training process or a condition that limits the performance of the initial model. For example, it can be the number of training round condition, or it can be the loss value interval condition. Exemplarily, it can be determined whether the loss value is less than a preset threshold. If it is less than the preset threshold, it is determined that the training completion condition is met. The specific size of the preset threshold is not limited and can be set as needed.
[0135] Apply the audio quality detection model training method provided by the embodiments of the present application, and use the mean opinion score (MOS) as the label of the initial training audio and the training audio. The mean opinion score is evaluated by a large number of listeners for the quality of the audio of sentences read aloud by male or female speakers transmitted through a communication circuit. The listeners score each sentence according to the following criteria: (1) very poor (2) poor (3) fair (4) good (5) very good. MOS is the arithmetic method of all listeners' individual scores, ranging from 1 (worst) to 5 (best). The obtained MOS label can accurately characterize the quality of the sentence audio in the initial training audio. Through voice activity detection, the non-human voice part of the audio can be filtered out as interference, and only the human voice part is retained as the training audio. After obtaining the training features through feature extraction, use the initial model to perform quality detection, and calculate the loss value using the obtained training audio quality detection results and the mean opinion score label. Use the loss value to adjust the initial model so that the initial model learns the correct way to evaluate the audio quality, and the obtained training audio quality detection results are as close as possible to the MOS label. After the training is completed, the adjusted initial model can be determined as the audio quality detection model. The obtained audio quality detection model can accurately evaluate the quality of the audio to be tested without having pure clean audio.
[0136] Based on the above embodiments, after obtaining the audio quality detection model, the audio quality of the initial audio to be tested without the MOS label can be detected. Please refer to Figure 4 , Figure 4 This is a flowchart of an audio quality detection method provided by the embodiments of the present application, which specifically includes the following steps:
[0137] S201: Obtain the initial audio to be tested.
[0138] S202: Perform non-human voice filtering processing on the initial audio to be tested based on voice activity detection to obtain the audio to be tested.
[0139] S203: Perform feature extraction processing on the audio to be tested to obtain the features to be tested.
[0140] S204: Input the feature to be measured into the voice quality detection model to obtain the voice quality detection result corresponding to the initial audio to be measured.
[0141] Among them, the voice quality detection model is obtained by using the above voice quality detection model training method. The non-human voice filtering process and the feature extraction process are the same as those in the voice quality detection model training process. Specifically, in one implementation, the voice quality detection model includes a convolutional neural network, a long short-term memory network, a fully connected layer, and an average pooling layer.
[0142] Correspondingly, inputting the feature to be measured into the voice quality detection model to obtain the voice quality detection result corresponding to the initial audio to be measured includes:
[0143] Step 51: Input the feature to be measured into the convolutional neural network to obtain the intermediate feature to be measured.
[0144] Step 52: Input the intermediate feature to be measured into the long short-term memory network to obtain the initial detection result.
[0145] Step 53: Input the initial detection result into the fully connected layer to obtain the intermediate detection result.
[0146] Step 54: Input the intermediate detection result into the average pooling layer to obtain the voice quality detection result.
[0147] Among them, the processes of Step 51 to Step 54 can refer to the processes of Step 31 to Step 34, with the difference being the data being processed.
[0148] Further, please refer to Figure 5 , Figure 5 which is another flowchart of the voice quality detection method provided by the embodiment of the present application. After the audio signal (i.e., the initial audio to be measured) is obtained, the voice activity detection (VAD) algorithm is used to perform voice detection on it, and the signal with human voice (i.e., the audio to be measured) is output. The VAD algorithm can specifically be the webrtc-vad algorithm. Voice detection filters out the silent and non-human voice parts in the audio signal. Then it is resampled to 48 kHz. Audio features in the form of Mel spectrogram (i.e., the feature to be measured) are extracted, and the audio features are input into the voice quality detection model. When extracting the audio features, the Blackman-Harris window is used, and the frame shift size is 10.7 ms.
[0149] In this application, the audio detection model includes a 3-layer CNN network, a 2-layer BLSTM network, a fully connected layer, and an average pooling layer. Among them, the convolutional layer in the CNN network is a 2D convolutional layer, the convolutional kernel is 3*3, and the output filter lengths of the three convolutional layers are 16, 32, and 64 in sequence. Normalization is performed after each layer of CNN, and the ReLU activation function is used. Through the 3-layer CNN, the local features of the audio can be effectively learned. Subsequently, it is sent to a 2-layer bidirectional LSTM network, and its hidden units are all set to 256. The main purpose is to extract the temporal relationship of the local features and learn the correlation between the front and back frames. After cascading the 2-layer bidirectional LSTM, it is sent to the fully connected layer to predict the audio quality score for each frame level (i.e., the intermediate detection result). Finally, through the average pooling layer, the final objective evaluation score corresponding to the initial audio to be tested (i.e., the audio quality detection result) is output.
[0150] Next, the computer-readable storage medium provided by the embodiments of the present application will be introduced. The computer-readable storage medium described below can be mutually corresponding and referred to with the audio quality detection model training method described above.
[0151] This application also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the above-mentioned audio quality detection model training method are implemented.
[0152] The computer-readable storage medium may include: various media that can store program codes such as USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs.
[0153] In this specification, the various embodiments are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple. For the related parts, refer to the description in the method part.
[0154] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but this implementation should not be considered to exceed the scope of this application.
[0155] The steps of the methods or algorithms described in connection with the embodiments disclosed herein may be implemented directly in hardware, in software modules executed by a processor, or in a combination thereof. The software modules may be placed in a random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art.
[0156] Finally, it should also be noted that in this document, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variation thereof are intended to cover non-exclusive inclusion, such that a process, method, article or apparatus that includes a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or apparatus.
[0157] Specific examples are used in this document to illustrate the principles and implementation modes of the present application. The descriptions of the above embodiments are only for helping to understand the method and its core idea of the present application; at the same time, for those of ordinary skill in the art, according to the idea of the present application, there will be changes in the specific implementation modes and application scopes. In summary, the content of this specification should not be construed as a limitation to the present application.
Claims
1. A method for training a sound quality detection model, characterized in that, Including: Obtain an initial training audio and a corresponding mean opinion score label; The mean opinion score label is used to represent an average audio quality evaluation parameter obtained after multiple evaluation objects perform audio quality evaluation on the initial training audio; Perform non-human voice filtering processing on the initial training audio based on voice activity detection to obtain a training audio; Perform feature extraction processing on the training audio to obtain training features; Input the training features into an initial model to obtain a corresponding training audio quality detection result; Calculate a loss value using the training audio quality detection result and the mean opinion score label, and use the loss value to adjust the model parameters of the initial model; When it is detected that the training completion condition is satisfied, determine the adjusted initial model as an audio quality detection model; Among them, the obtaining process of the mean opinion score label includes: Perform averaging processing on each initial audio quality data to obtain a first score label; the initial audio quality data is obtained after each evaluation object performs audio quality evaluation on the initial training audio; Input the initial training audio into an audio defect detection model to obtain a defect detection result; among them, the audio defect detection model is used to detect audio defects that can affect the auditory perception; Generate a second score label based on the defect detection result; Generate the mean opinion score label using the first score label and the second score label.
2. The method for training a sound quality detection model according to claim 1, wherein The obtaining process of the mean opinion score label further includes: Play the initial training audio to each of the evaluation objects; Receive the initial audio quality data obtained after each evaluation object performs audio quality evaluation on the initial training audio.
3. The method for training a sound quality detection model according to claim 1, wherein The performing feature extraction processing on the training audio to obtain training features includes: Resample the training audio based on the maximum sampling rate perceptible by the human ear to obtain intermediate data; Perform sliding window framing on the intermediate data based on a preset window length to obtain a plurality of audio frames; Perform feature extraction processing on each of the audio frames to obtain the training features.
4. The method for training a sound quality detection model according to claim 1, wherein The performing non-human voice filtering processing on the initial training audio based on voice activity detection to obtain a training audio includes: Perform voice activity detection on the initial training audio to obtain voice endpoint moments; Segment the initial training audio according to the voice endpoint moments to obtain a plurality of audio segments, and remove the non-human voice frequency bands in the audio segments to obtain human voice frequency bands; Concatenate the human voice frequency bands to obtain the training audio.
5. The method for training a sound quality detection model according to claim 1, wherein The initial model includes a convolutional neural network, a long short-term memory network, a fully connected layer, and an average pooling layer; The inputting the training features into the initial model to obtain a corresponding training audio quality detection result includes: Input the training features into the convolutional neural network to obtain training intermediate features; Input the training intermediate features into the long short-term memory network to obtain a training initial detection result; Input the training initial detection result into the fully connected layer to obtain a training intermediate detection result; Input the training intermediate detection result into the average pooling layer to obtain the training audio quality detection result.
6. The method for training the sound quality detection model according to claim 5, wherein The obtaining the initial training audio and the corresponding mean opinion score label includes: Obtain a plurality of the initial training audios and the mean opinion score labels of a batch from the training dataset according to a preset batch size; Correspondingly, the calculating of the loss value by using the training sound quality detection result and the mean opinion score label includes: When obtaining the training sound quality detection results corresponding to all the initial training audios in a batch, obtain the loss value by using the training sound quality detection results, mean opinion score labels and training intermediate detection results within this batch.
7. The method for training the sound quality detection model according to claim 6, wherein The obtaining of the loss value by using the training sound quality detection results, mean opinion score labels and training intermediate detection results within this batch includes: According to Obtain the loss value; wherein, the loss is the loss value, the S is the preset batch size, the T S is the number of frames corresponding to the training feature, the M S is the mean opinion score label, the is the training speech quality detection result, the is the value corresponding to the t-th frame in the training feature in the training intermediate detection result, and the α is the preset weight.
8. The method for training a sound quality detection model according to claim 1, wherein The detecting of meeting the training completion condition includes: Judge whether the loss value is less than a preset threshold; If so, determine that the training completion condition is met.
9. A sound quality detection method, characterized in that, Include: Obtain an initial audio to be measured; Perform non-human voice filtering processing based on voice activity detection on the initial audio to be measured to obtain the audio to be measured; Perform feature extraction processing on the audio to be measured to obtain the feature to be measured; Input the feature to be measured into the sound quality detection model to obtain the sound quality detection result corresponding to the initial audio to be measured; wherein, the sound quality detection model is obtained by using the sound quality detection model training method according to any one of claims 1 to 7.
10. The sound quality detection method according to claim 9, characterized in that, The sound quality detection model includes a convolutional neural network, a long short-term memory network, a fully connected layer and an average pooling layer; The inputting of the feature to be measured into the sound quality detection model to obtain the sound quality detection result corresponding to the initial audio to be measured includes: Input the feature to be measured into the convolutional neural network to obtain an intermediate feature to be measured; Input the intermediate feature to be measured into the long short-term memory network to obtain an initial detection result; Input the initial detection result into the fully connected layer to obtain an intermediate detection result; Input the intermediate detection result into the average pooling layer to obtain the sound quality detection result.
11. An electronic device, characterized in that, Include a memory and a processor, wherein: The memory is used to store a computer program; The processor is used to execute the computer program to implement the sound quality detection model training method according to any one of claims 1 to 8, and / or the sound quality detection method according to any one of claims 9 to 10.
12. A computer-readable storage medium, characterized in that, Used to store a computer program, wherein the computer program, when executed by the processor, implements the sound quality detection model training method according to any one of claims 1 to 8, and / or the sound quality detection method according to any one of claims 9 to 10.
Citation Information
Patent Citations
Speech quality detection model training method and speech quality detection method
CN112967735A
Biological sound event detection model training method and sound event detection method
CN113724733A
Voiceprint recognition method and device, electronic equipment and storage medium
CN114141252A