Voice quality inspection method and device, equipment and storage medium

By three-classifying and dividing voice, the problem that traditional voice quality inspection methods cannot accurately check after the on-time is solved, and efficient voice quality inspection is achieved, saving costs.

CN120220733APending Publication Date: 2025-06-27SF TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311810058.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-26
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

Traditional voice quality inspection methods cannot perform quality inspection after the precise judgment of the on-time, resulting in low voice quality inspection efficiency.

Method used

By extracting the first audio feature signal of the voice to be quality-tested, using a pre-trained classification model for three classifications, obtaining the classification results of ringtones, mutes and vocals, and segmenting and decoding the voice according to the vocal classification results. If the decoding result is empty, a quality inspection warning will be generated.

Benefits of technology

It realizes accurate quality inspection after the on-time, skipping invalid data, improving quality inspection efficiency, saving time and labor costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120220733A_ABST
    Figure CN120220733A_ABST
Patent Text Reader

Abstract

The invention provides a voice quality inspection method and system and a storage medium, and the method comprises the steps: extracting a first audio feature signal of voice to be subjected to quality inspection, and classifying the first audio feature signal according to a pre-trained first classification model to obtain a classification result; the classification result comprises a ringtone result, a mute result and a human voice result; clustering the human voice results to obtain human voice classification results corresponding to different speaking subjects, and segmenting the voice to be subjected to quality inspection according to the human voice classification results to obtain a plurality of voice segments; decoding the plurality of voice segments in sequence, and if a decoding result is null, generating a quality inspection warning; the quality inspection warning is used for indicating that quality inspection does not need to be carried out on the voice segments of which the decoding results are null; according to the invention, the quality inspection efficiency can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of speech recognition, and in particular, to a method, device, equipment and storage medium for speech quality inspection. Background Art

[0002] When conducting quality inspection on customer service through speech recognition technology, in actual scenarios, in order to recognize the statements of the customer service in the call, recording is usually started from the moment the call is dialed, and the conversation content between the customer and the customer service is verified based on the obtained recording. However, regardless of whether the call is answered or not, before speech quality inspection, voice data to be quality-inspected is obtained by starting recording after the call is dialed, and the ringing tone recorded before the call, the voice of the customer before the call, and other interfering sounds are all quality-inspected as voice data. Therefore, in the traditional method, quality inspection cannot be carried out after accurately determining the connection moment, resulting in low speech quality inspection efficiency. Summary of the Invention

[0003] The purpose of the present invention is to address the deficiencies of the above-mentioned traditional related technologies, and propose a method, device, equipment and storage medium for speech quality inspection, which can improve the efficiency of speech quality inspection.

[0004] In a first aspect, the present invention provides a method for speech quality inspection, including:

[0005] Extracting a first audio feature signal of the voice data to be quality-inspected, and classifying the first audio feature signal according to a pre-trained first classification model to obtain a classification result; wherein, the classification result includes: a ringing result, a silence result, and a human voice result; the human voice result does not include the human voice in the ringing result;

[0006] Clustering the human voice result to obtain a human voice classification result corresponding to different speakers, and segmenting the voice data to be quality-inspected according to the human voice classification result to obtain a plurality of voice segments;

[0007] Decoding the plurality of voice segments in sequence, and if the decoding result is empty, generating a quality inspection warning; the quality inspection warning is used to indicate that there is no need to conduct quality inspection on the voice segment with an empty decoding result.

[0008] The present invention classifies the first audio feature signal into three categories, and divides the voice data to be quality-inspected according to the human voice classification result of the second classification. Since the human voice classification result does not include the human voice in the ringing result, when decoding the obtained voice segments using the human voice classification result, the human voice in the ringing tone can be accurately classified as the ringing result of an unanswered call, so that quality inspection can be carried out after accurately determining the connection moment, skipping the ringing result or the silence result for quality inspection, thereby improving the quality inspection efficiency and saving the time cost and labor cost of quality inspection.

[0009] In combination with the first aspect, in a possible implementation manner, the training steps of the first classification model include:

[0010] Concatenate the voice feature signal with the pre-constructed ringtone feature signal to obtain a second audio feature signal, and denote the labels of the ringtone segment, voice segment, and silent segment in the second audio feature signal as ringtone, voice, and silent respectively;

[0011] Perform multi-classification training on the initial second classification model according to the label and the second audio feature signal to obtain the first classification model.

[0012] The present invention adopts concatenating the ringtone feature signal with the voice feature signal, which can restore the real call scenario and construct a dataset that conforms to the actual situation, facilitating the second classification model to accurately learn the features of the ringtone feature signal with the interference of the voice feature signal, so as to facilitate the rapid convergence of the three-class second classification model and improve the generalization ability of the model.

[0013] In combination with the first aspect and the above possible implementation manner, in another possible implementation manner, before the step of concatenating the voice feature signal with the pre-constructed ringtone feature signal to obtain a second audio feature signal, it further includes:

[0014] Output a variety of acoustic feature signals by a pre-trained generation network, and use the variety of acoustic feature signals as the ringtone feature signal; wherein, the generation network is trained by taking the acoustic feature signal of the real ringtone as the real sample for input, outputting the generated sample, and discriminating the generated sample and the real sample by a discriminant network.

[0015] In combination with the first aspect and the above possible implementation manner, in another possible implementation manner, the step of performing multi-classification training on the initial second classification model according to the label and the second audio feature signal to obtain the first classification model includes:

[0016] Input the second audio feature signal into the initial second classification model, and construct the initial first decoding graph of the second classification model according to the state transition probability of the second audio feature signal;

[0017] Output the training result of the second audio feature signal according to the first decoding graph, and train the second classification model according to the training result and the label to obtain the first classification model;

[0018] Wherein, the state transition probability is the probability of mutual conversion between different states; the states include: silent state, voice state, and ringtone state.

[0019] Combined with the first aspect and the above possible implementation manners, in another possible implementation manner, constructing the initial first decoding graph of the second classification model according to the state transition probability of the second audio feature signal includes:

[0020] Constructing the initial first decoding graph of the second classification model according to the preset longest duration and shortest duration of the state, the transition probability of the initial state of the second audio feature signal, the probability of returning to its own state, and the mutual transition probabilities between different states.

[0021] The present invention constructs a decoding graph through the state transition probabilities of the mutual transitions of three states, the longest duration and the shortest duration of maintaining its own state, which can improve the interpretability of using the second audio feature information as the second classification model and outputting the classification results of three classifications, and can also improve the anti-attack ability and generalization ability of the model.

[0022] Combined with the first aspect and the above possible implementation manners, in another possible implementation manner, splicing the human voice feature signal with the pre-constructed ringtone feature signal to obtain a second audio feature signal, and respectively denoting the labels of the ringtone segment, the human voice segment, and the silent segment in the second audio feature signal as ringtone, human voice, and silent, includes:

[0023] Overlapping and splicing the human voice feature signal with the pre-constructed ringtone feature signal according to a preset length to obtain a second audio feature signal, denoting the label of the overlapping feature signal in the obtained second audio feature signal as ringtone, and respectively denoting the labels of the human voice segment and the non-human voice segment in the remaining non-overlapping feature signals as human voice and silent.

[0024] Combined with the first aspect and the above possible implementation manners, in another possible implementation manner, splicing the human voice feature signal with the pre-constructed ringtone feature signal to obtain a second audio feature signal, and respectively denoting the labels of the ringtone segment, the human voice segment, and the silent segment in the second audio feature signal as ringtone, human voice, and silent, includes:

[0025] Non-overlappingly splicing the human voice feature signal with the pre-constructed ringtone feature signal to obtain a second audio feature signal, and respectively denoting the labels of the ringtone segment and the human voice segment in the second audio feature signal as ringtone and human voice, and denoting the label of the remaining segment except the ringtone segment and the human voice segment as silent.

[0026] Combined with the first aspect and the above possible implementation manners, in another possible implementation manner, clustering the human voice results to obtain the human voice classification results corresponding to different speakers includes:

[0027] Use the output of the fully connected layer of the last layer in the first classification model as the vocal feature, obtain the distance matrix of all vocal features, and continuously aggregate the vocal features with the closest distances in the distance matrix until the number of clusters in the distance matrix is equal to the number of speakers, to obtain the vocal classification results corresponding to different speakers one by one.

[0028] In the present invention, after performing three-class classification on the first classification model, feature extraction is also performed according to the three-class classification to classify the vocal results according to different speakers, that is, classification and feature extraction are performed simultaneously through one classification model, so as to improve the processing speed of the overall quality inspection. And the feature extraction depends on the classification result. Compared with the existing related technologies that use the secondary clustering algorithm, the present invention uses one classification model to perform classification and feature extraction simultaneously, which can improve the relevance of the internal logic processing of the model, and thus improve the quality inspection efficiency.

[0029] Combined with the first aspect and the above possible implementation manners, in another possible implementation manner, the method further includes:

[0030] If the decoded data obtained by decoding is not empty, no quality inspection warning is generated, so that the backend can perform quality inspection on the second voice segment.

[0031] In a second aspect, the present invention provides a voice quality inspection device, including: a first classification unit, a second classification unit, and a quality inspection unit; wherein,

[0032] The first classification unit is configured to extract the first audio feature signal of the voice to be quality inspected, and classify the first audio feature signal according to a pre-trained first classification model to obtain a classification result; wherein, the classification result includes: a ringtone result, a mute result, and a vocal result; the vocal result does not include the vocals in the ringtone result;

[0033] The second classification unit is configured to cluster the vocal result to obtain the vocal classification results corresponding to different speakers, and segment the voice to be quality inspected according to the vocal classification results to obtain a plurality of voice segments;

[0034] The quality inspection unit is configured to decode the plurality of voice segments in sequence. If the decoding result is empty, a quality inspection warning is generated; the quality inspection warning is used to indicate that there is no need to perform quality inspection on the voice segment with an empty decoding result.

[0035] In a third aspect, the present invention provides an electronic device, the electronic device includes: a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the computer program, the steps of the voice quality inspection method described in the first aspect are implemented.

[0036] When the voice quality inspection method described in the first aspect is integrated into an electronic device, the present invention can perform rapid on-site quality inspection and feedback of on-site quality inspection through various electronic devices, and has stronger scalability.

[0037] In a fourth aspect, the present invention provides a readable computer storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the voice quality inspection method described in the first aspect and the above possible implementation manners are implemented.

[0038] When the voice quality inspection method described in the first aspect is stored in the storage medium in the form of a program, the present invention can perform operations such as rapid quality inspection on the voice to be quality inspected on different operating systems and terminals by running or reading the executable program in the storage medium, is applicable to more operating systems and different application platforms, and has stronger scalability. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] Figure 1 is a quality inspection schematic diagram of the related art in the prior art given in the embodiment of the present application;

[0040] Figure 2 is a flowchart of a voice quality inspection method provided by an embodiment of the present application;

[0041] Figure 3 is a flowchart of extracting the first audio feature signal provided by an embodiment of the present application;

[0042] Figure 4 is a flowchart of obtaining the first classification model by training the second classification model provided by an embodiment of the present application;

[0043] Figure 5 is a flowchart of constructing a ringtone feature signal by a generative adversarial model provided by an embodiment of the present application;

[0044] Figure 6 is a schematic diagram of the second audio feature signal obtained by performing overlapping splicing provided by an embodiment of the present application;

[0045] Figure 7 is a schematic diagram of the second audio feature signal obtained by performing non-overlapping splicing provided by an embodiment of the present application;

[0046] Figure 8 is a schematic diagram of the format of the training data set provided by an embodiment of the present application;

[0047] Figure 9 is a schematic diagram of probability conversion provided by an embodiment of the present application;

[0048] Figure 10 is a schematic diagram of the network structure of TDNN provided by an embodiment of the present application;

[0049] Figure 11 It is a schematic diagram of voice quality inspection provided by an embodiment of the present application;

[0050] Figure 12 It is a schematic diagram of a complete quality inspection process provided by an embodiment of the present application;

[0051] Figure 13 It is a test schematic diagram of the first classification model provided by an embodiment of the present application;

[0052] Figure 14 It is a schematic structural diagram of a voice quality inspection device provided by an embodiment of the present invention;

[0053] Figure 15 It is a schematic diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners

[0054] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0055] Refer to Figure 1 , which is a quality inspection schematic diagram of the existing related technology given by an embodiment of the present application. In the figure, the recording of the voice call of two speakers is subjected to quality inspection. When the voice to be quality inspected is recognized and segmented, since before the call starts, a speaker has already started speaking, and the existing related technology does not discriminate the voice before the actual call starts, resulting in the voice before the actual call being regarded as the target voice for quality inspection, reducing the quality inspection efficiency.

[0056] It should be noted that when recording the voice to be quality inspected, the customer's authorization will be obtained before the call is connected for recording. The voice to be quality inspected in each embodiment of the present application is the voice recorded after the customer's authorization and does not involve privacy information.

[0057] Based on this, the present invention conducts more detailed classification and detects the data before the formal call. When the data before the call is detected, a quality inspection warning will be generated, so that the backend skips the invalid data and does not perform quality inspection processing, thereby improving the quality inspection efficiency. To more clearly illustrate the technical solutions of the present invention, specific embodiments will be shown.

[0058] Embodiment 1

[0059] Refer to Figure 2, which is a schematic flowchart of a voice quality inspection method provided by an embodiment of the present application, includes steps S11 to S13, specifically as follows:

[0060] Step S11: Extract the first audio feature signal of the voice to be quality inspected, and classify the first audio feature signal according to a pre-trained first classification model to obtain a classification result; wherein, the classification result includes: a ringtone result, a mute result, and a human voice result; the human voice result does not include the human voice in the ringtone result.

[0061] In some embodiments of the present application, the first classification model is trained according to a second audio feature signal and the label of the second audio feature signal; the second audio feature signal is obtained by splicing a human voice feature signal and a constructed ringtone feature signal.

[0062] It should be noted that the first audio feature signal is the audio feature signal of the voice recording to be quality inspected after all models are trained; all models include: a classification model and a Generative Adversarial Network (GAN) model; specifically, the classification model is the first classification model after training converges, and the classification model is the second classification model before training converges; the Generative Adversarial Network model includes: a generative network and a discriminant network; the first human voice feature signal is the human voice feature signal that does not include the ringtone and contains the human voice recording used by the second classification model during training, the ringtone feature signal is the ringtone feature signal that does not include the human voice and contains the ringtone used by the second classification model during training, and at least one of the human voice feature signal and the ringtone feature signal contains the mute feature signal when there is mute.

[0063] In some embodiments of the present application, extracting the first audio feature signal of the voice to be quality inspected includes: obtaining the voice to be quality inspected, performing pre-emphasis on the voice to be quality inspected in sequence, windowing the obtained high-frequency components according to a preset window and a preset step length, performing discrete Fourier transform after windowing is completed to convert the time-domain information into frequency-domain information, processing the frequency-domain information through a filter and logarithmic calculation to eliminate harmonics, and performing discrete cosine transform on the obtained envelope information, and using the obtained acoustic feature signal as the first audio feature signal.

[0064] In an exemplary embodiment, the acoustic feature signal may be an MFCC feature signal. The MFCC feature signal is Mel-scale Frequency Cepstral Coefficients (MFCC), which is a commonly used feature signal in speech recognition.

[0065] In an embodiment of the present application, for extracting the first audio feature signal of the voice to be quality inspected, refer to Figure 3, is a schematic flowchart for extracting the first audio feature signal provided by an embodiment of the present application, including sub-steps S111 to S116; among them, sub-steps S111 to S112 are preprocessing operations, and sub-steps S113 to S116 are acoustic feature signal extraction operations, specifically:

[0066] Sub-step S111, pre-emphasis. Specifically, pre-emphasize the speech signal of the speech to be quality-inspected to enhance the high-frequency components in the speech signal.

[0067] Sub-step S112, windowing. Specifically, to prevent spectral leakage, divide the sound into segments with a window of 25 ms and a step size of 10 ms, and use a Hamming window for windowing.

[0068] In other embodiments of the present application, the Hamming window can be replaced with a rectangular window or a Hann window.

[0069] Sub-step S113, Discrete Fourier Transform (DTF). Specifically, perform a convolution operation according to the Hamming window and then perform a Discrete Fourier Transform (DTF) to convert the time domain to the frequency domain.

[0070] Sub-step S114, filter filtering. Specifically, calculate the filter by simulating the perception of the human ear.

[0071] In a preferred embodiment of the present application, a Mel filter is used for filtering.

[0072] In an alternative embodiment of the present application, an equal-area Mel filter is used to obtain the frequency domain signal of the human voice.

[0073] It should be noted that using an equal-area Mel filter can simulate the scenario where the human ear perceives lower frequencies higher and is more sensitive to low-frequency signals than high-frequency signals. That is, the equal-area Mel filter can simulate the process of the human ear processing human voice signals for the human voice in the first audio feature signal.

[0074] In another alternative embodiment of the present application, a constant-height Mel filter is used to process the high-frequency information of non-human voices in the first audio feature signal.

[0075] Sub-step S115, logarithmic calculation. Among them, signal separation is performed through logarithmic calculation to retain the envelope signal.

[0076] It should be noted that through logarithmic calculation, the envelope signal can be retained. The envelope signal is the timbre, and the pitch is removed to obtain the target feature for speech recognition. Therefore, through filter filtering and logarithmic calculation, harmonics are eliminated and envelope information is retained.

[0077] Sub-step S116: Perform a Discrete Cosine Transform (DCT) to obtain an acoustic feature signal, and use the acoustic feature signal as the first audio feature signal.

[0078] It should be noted that DCT processing can remove the correlations among different Mel frequency bands, thereby ensuring low correlations among the features of the obtained first audio feature signal, which exactly meets the requirements of the first classification model for the input. Therefore, it is beneficial to improve the classification accuracy of the first classification model. Through the above 6 sub-steps, the time-domain information can be converted into frequency-domain information with low correlations among features to meet the requirements of the first classification model for the input.

[0079] In some embodiments of the present application, the first classification model is trained based on the second audio feature signal obtained by splicing the human voice feature signal and the constructed ringtone feature signal and the label of the second audio feature signal, including: sub-steps S121 to S123; wherein, the second audio feature signal is obtained by splicing the human voice feature signal and the constructed ringtone feature signal and is used for training the second classification model. Refer to Figure 4 , which is a schematic flowchart of the process for the second classification model provided by the embodiments of the present application to obtain the first classification model through training, specifically:

[0080] Sub-step S121: Learn the feature distribution of real ringtones through a generative adversarial model, and use the constructed acoustic feature signals of multiple ringtones as the ringtone feature signals.

[0081] In some embodiments of the present application, before splicing the human voice feature signal and the pre-constructed ringtone feature signal to obtain the second audio feature signal, learning the feature distribution of real ringtones through a generative adversarial model and using the constructed acoustic feature signals of multiple ringtones as the ringtone feature signals includes: outputting multiple acoustic feature signals by a pre-trained generation network, and using the multiple acoustic feature signals as the ringtone feature signals; wherein, the generation network is trained by using the acoustic feature signal of the real ringtone as the real sample for input, outputting the generated sample, and discriminating the generated sample and the real sample by a discriminant network.

[0082] In some embodiments of the present application, the generation network is trained by taking the acoustic feature signal of a real ringtone as a real sample for input, outputting a generated sample, and then discriminating between the generated sample and the real sample through a discrimination network, including: inputting the acoustic feature signal of the real ringtone as a real sample into the generation network, the generation network constructing a generated sample that approximates the acoustic feature signal of the real ringtone, and using the generated sample and the real sample as inputs to the discrimination network, the discrimination network outputting a discrimination result for the generated sample; training the generation network and the discrimination network in sequence according to the discrimination result, the generated sample, and the real sample, respectively obtaining a generation network and a discrimination network that have learned the feature distribution of the real ringtone; forming a generative adversarial network with the trained generation network and the trained discrimination network, and using the multiple acoustic feature signals output by the generative adversarial network that approximate the real ringtone as ringtone feature signals.

[0083] In one embodiment of the present application, refer to Figure 5 , which is a schematic flowchart of constructing ringtone feature signals by the generative adversarial model provided in the embodiments of the present application. By learning the distribution of real ringtones through the generative adversarial model, various acoustic feature signals that approximate real ringtones are constructed. In the figure, the acoustic feature signal of the real ringtone is used as the real sample x, G is the generation network, using the hidden vector z to approximate the real sample x to obtain the generated sample G(z); D is the discrimination network, receiving the real sample x and the generated sample G(z), and outputting the probability that the generated sample G(z) belongs to the real sample x.

[0084] It should be noted that the training objective of the entire generative adversarial network is to make the generated sample G(z) closer to the real sample x, while the discrimination network D can better discriminate the authenticity of the generated sample.

[0085] In a preferred embodiment of the present application, after preprocessing the voice signal of the ringtone, a time-domain feature signal is obtained, and the time-domain feature signal is used as a convolutional neural-deconvolutional network based on generative adversarial to directly obtain a frequency-domain signal, obtaining multiple acoustic feature signals that approximate real ringtones, and using the multiple acoustic feature signals as ringtone feature signals.

[0086] It should be noted that since the ringtone data in calls is diverse and it is difficult to collect all ringtone data for data recognition, a generative adversarial network is used to construct ringtone data.

[0087] Sub-step S122: Concatenate the voice feature signal with the pre-constructed ringtone feature signal to obtain a second audio feature signal, and denote the labels of the ringtone segment, voice segment, and silent segment in the second audio feature signal as ringtone, voice, and silent, respectively.

[0088] In an embodiment of the present application, the steps of extracting the first human voice feature signal are the same as those of extracting the first audio feature signal, specifically: obtaining the human voice to be trained, performing pre-emphasis on the human voice in sequence, windowing the obtained high-frequency components according to a preset window and a preset step length, and performing discrete Fourier transform after the windowing is completed to convert the time-domain information into frequency-domain information, processing the frequency-domain information through a filter and logarithmic calculation to eliminate harmonics, and performing discrete cosine transform on the obtained envelope information, and using the obtained acoustic feature signal as the first human voice feature signal.

[0089] In an embodiment of the present application, splicing the human voice feature signal with a pre-constructed ringtone feature signal to obtain a second audio feature signal, and respectively denoting the labels of the ringtone segment, the human voice segment, and the silent segment in the second audio feature signal as ringtone, human voice, and silent, includes: overlapping and splicing the human voice feature signal with the pre-constructed ringtone feature signal according to a preset length to obtain a second audio feature signal, and denoting the label of the overlapping feature signal in the obtained second audio feature signal as ringtone, and respectively denoting the labels of the human voice segment and the non-human voice segment in the remaining non-overlapping feature signals as human voice and silent.

[0090] See Figure 6 , which is a schematic diagram of the second audio feature signal obtained by overlapping splicing provided by the embodiment of the present application. In the figure, the acoustic feature signal of the ringtone is overlapped and spliced with the acoustic feature signal of the human voice to obtain a second audio feature signal. When tagging, when the ringtone appears and the human voice also appears at the same time, it is considered that the human voice is interference to the ringtone. When the label of the simultaneous appearance of the human voice and the ringtone is ringtone, the labels of the human voice segment and the non-human voice segment in the remaining non-overlapping feature signals are respectively denoted as human voice and silent.

[0091] In another embodiment of the present application, splicing the human voice feature signal with a pre-constructed ringtone feature signal to obtain a second audio feature signal, and respectively denoting the labels of the ringtone segment, the human voice segment, and the silent segment in the second audio feature signal as ringtone, human voice, and silent, includes: non-overlappingly splicing the human voice feature signal with the pre-constructed ringtone feature signal to obtain a second audio feature signal, and respectively denoting the labels of the ringtone segment and the human voice segment in the second audio feature signal as ringtone and human voice, and denoting the label of the remaining segment after removing the ringtone segment and the human voice segment as silent.

[0092] In an embodiment of the present application, see Figure 7, which is a schematic diagram of the second audio feature signal obtained by non-overlapping splicing provided by an embodiment of the present application. In the figure, the acoustic feature signal of the ringtone is non-overlappingly spliced with the acoustic feature signal of the human voice to obtain the second audio feature signal. When labeling, the labels of the ringtone segment and the human voice segment in the second audio feature signal are respectively denoted as ringtone and human voice, and the label of the remaining segment excluding the ringtone segment and the human voice segment is denoted as silence.

[0093] The present invention uses partial overlapping splicing and non-overlapping splicing of the ringtone feature signal and the human voice feature signal, which can restore the real call scenario, construct a practical training data set, and use a generative adversarial model to construct various ringtone feature signals, which can greatly enrich the samples of the training data set, facilitate the second classification model to accurately learn the features of the ringtone feature signal with the interference of the human voice feature signal as interference, so as to facilitate the rapid convergence of the three-class second classification model and improve the generalization ability of the model.

[0094] In another preferred embodiment of the present application, the ringtone feature signal and the human voice feature signal are simultaneously overlapped and spliced according to a preset length, and the ringtone feature signal and the human voice feature signal are non-overlappingly spliced, and the labels of ringtone, human voice and silence are respectively marked.

[0095] It should be noted that each sub-segment in the second audio feature signal needs to be labeled to obtain complete label data.

[0096] In some embodiments of the present application, after obtaining the label and the second audio feature signal, data preprocessing is performed to obtain a training data set in a preset format. Specifically, the training data set contains four columns of data. The first column is the label of each sub-segment in the second audio feature signal, the second column is the file name, the third column is the start time of the sub-segment, and the fourth column is the end time of the sub-segment. The name format is "file name - start time - end time - label"; if it is a human voice, the label is speech, if it is a ringtone, the label is garbage, and the rest of the sound is a non-human voice segment (such as silence).

[0097] See Figure 8 , which is a schematic diagram of the format of the training data set provided by an embodiment of the present application. In the figure, according to the name format "file name - start time - end time - label", there are 10 sub-segments of the second audio feature signal, and according to the naming format of the training data set shown, so as to perform multi-classification training on the second classification model.

[0098] Sub-step S123: Perform multi-classification training on the initial second classification model according to the label and the second audio feature signal to obtain the first classification model.

[0099] In some embodiments of the present application, multi-classification training is performed on an initial second classification model according to the label and the second audio feature signal to obtain the first classification model, including: inputting the second audio feature signal into the initial second classification model, and constructing an initial first decoding graph of the second classification model according to the state transition probability of the second audio feature signal; outputting a training result of the second audio feature signal according to the first decoding graph, and training the second classification model according to the training result and the label to obtain the first classification model; wherein, the state transition probability is the probability of mutual conversion between different states; the states include: a mute state, a human voice state, and a ringtone state.

[0100] It should be noted that the first decoding graph is a decoding graph of state transition constructed by the second classification model according to the parameters when not converged during training.

[0101] In some embodiments of the present application, constructing an initial first decoding graph of the second classification model according to the state transition probability of the second audio feature signal includes: constructing an initial first decoding graph of the second classification model according to the preset longest duration and shortest duration of the state, the transition probability of the initial state of the second audio feature signal, the probability of returning to its own state, and the mutual conversion probability between different states.

[0102] In some embodiments of the present application, the longest duration is that when the number of conversions that are always in the same state after state conversion reaches the first conversion number threshold, the next state is different from the current state; the shortest duration is that during multiple state transitions, it can only be in the same state, and when the number of conversions reaches the second conversion number threshold, the next state is different from the current state.

[0103] In some embodiments of the present application, by presetting the shortest duration and longest duration of the ringtone state, the recognition of the ringtone is accurately improved. Specifically, the longest duration of the preset ringtone state is Tmax. After Tmax state transitions and still being in the ringtone state, the next state will be other states; the shortest duration of the preset ringtone state is Tmin. If within Tmin, the state can only remain in the ringtone state and can only turn to another state after more than Tmin conversions; wherein, both Tmax and Tmin are positive integers.

[0104] It should be noted that a decoding graph is constructed through the transition probability from the initial state to the three states, the three states include: a mute state, a human voice state, and a ringtone state; the probability of returning to its own state, the mutual conversion probability between different states and the probability of returning to its own state, as well as the shortest duration and longest duration of the state.

[0105] See Figure 9, which is a schematic diagram of probability conversion provided by an embodiment of the present application. In the figure, it includes the conversion from the mute state to the voice state and the ringing state, the conversion from the voice state to the mute state and the ringing state, the conversion from the ringing state to the voice state and the mute state, as well as the mute state maintaining its own state, the voice state maintaining its own state, and the ringing state maintaining its own state, a total of 9 state probabilities.

[0106] Exemplarily, the three states are denoted as the mute state (1), the voice state (2), and the ringing state (3), then there are a total of 9 conversion probabilities, namely p11, p12, p13, p21, p22, p23, p31, p32, and p33. pij refers to the conversion probability from the i state to the j state. At the same time, the shortest duration (Tmin) and the longest duration (Tmax) of each state are set.

[0107] The present invention constructs a decoding graph through the conversion probabilities of the mutual conversion of three states and the preset shortest and longest durations of maintaining its own state, which can improve the interpretability of the classification result of outputting three classifications by using the second audio feature information as the second classification model, and can also improve the anti-attack ability and generalization ability of the model.

[0108] In an embodiment of the present application, a time delay neural network (TDNN)-long short-term memory (LSTM) model is used as the initial second classification model, and the second audio feature signal is used as the input of the TDNN-LSTM network model for supervised multi-classification training.

[0109] In an embodiment of the present application, refer to Figure 10 , which is a schematic diagram of the network structure of the TDNN provided by an embodiment of the present application. In the figure, it includes 4 layers of networks. The multi-hidden layer has stronger feature extraction ability. Since adjacent neurons in each layer have overlapping information, each layer of neurons extracts features through some neurons rather than all neurons to reduce the classification model and improve the training rate of three classifications (including: mute, voice, and ringing).

[0110] In an embodiment of the present application, two fully connected layers are connected after the TDNN to obtain the input feature vector.

[0111] It should be noted that in step S11, the classification results of the audio feature signal have been obtained, including the ringing result, the mute result, and the voice result, and the classification results of the first audio feature signal corresponding to the ringing result, the mute result, and the voice result, rather than the classification result of the speech to be quality inspected. And, the voice result includes the audio feature signals of multiple speakers. Therefore, it is also necessary to classify the audio feature signals of multiple speakers.

[0112] Step S12: Cluster the voice results to obtain voice classification results corresponding to different speakers, and segment the voice to be quality-inspected according to the voice classification results to obtain a number of voice segments.

[0113] In an embodiment of the present application, the number of voice segments includes: a first voice segment corresponding to the ringtone result or the mute result.

[0114] In an embodiment of the present application, by extracting the voice features (x-vector features) of different speakers, and then classifying the voice features through hierarchical clustering to obtain voice classification results corresponding to different speakers. In order to extract the x-vector network, the TDNN network is still used. The input of the network is the acoustic feature signal, and two fully connected layers are connected after the TDNN network.

[0115] In an embodiment of the present application, the output of the fully connected layer of the TDNN in the TDNN-LSTM network model is used as the voice feature.

[0116] In an embodiment of the present application, clustering the voice results to obtain voice classification results corresponding to different speakers includes: using the output of the fully connected layer of the last layer in the first classification model as the voice feature, obtaining the distance matrix of all voice features, and continuously aggregating the voice features with the closest distance in the distance matrix until the number of clusters in the distance matrix is equal to the number of speakers, to obtain voice classification results corresponding one by one to different speakers.

[0117] In an embodiment of the present application, by obtaining the output of the fully connected layer of the last layer of the TDNN in the TDNN-LSTM network model, extracting the voice features of the speakers of each sub-segment in the voice results, denoted as x-vector. There are N sub-segments in total, so there are N-dimensional x-vectors. By calculating the distances between each pair, an N*N distance matrix is formed, and then through hierarchical clustering, continuously merge the points with the closest distance. In the customer service scenario, assume there are two speakers, the customer service and the customer. So the number of clusters is set to 2. Hierarchical clustering continuously merges the data until the number of clusters is equal to 2, and the clustering ends.

[0118] The present invention adopts triple classification of the first classification model, and also performs feature extraction according to the triple classification to classify the voice results according to different speakers, that is, a classification model is used for classification and feature extraction at the same time, so as to improve the processing speed of the overall quality inspection. And the feature extraction depends on the classification results, so as to improve the relevance of the internal logic processing of the model, and further improve the quality inspection efficiency.

[0119] Step S13: Decode the several voice segments in sequence. If the decoding result is empty, a quality inspection warning is generated; the quality inspection warning is used to indicate that there is no need to perform quality inspection on the voice segment with an empty decoding result.

[0120] In some embodiments of the present application, if the decoded data obtained by decoding is not empty, no quality inspection warning is generated, so that the backend can perform quality inspection on the voice segment.

[0121] In some embodiments of the present application, the several voice segments further include a second voice segment with non-empty decoded data ; If the decoding result is empty, a quality inspection warning is generated. Specifically: when the first decoded data obtained by decoding is empty, the first voice segment is detected and a quality inspection warning is generated; when the decoded data obtained by decoding is not empty, the second voice segment corresponding to the voice classification result is detected and no quality inspection warning is generated, so that the backend can perform quality inspection on the second voice segment; wherein, the several voice segments further include the second voice segment.

[0122] It should be noted that the first voice segment refers to the voice segment with an empty decoded text among the several voice segments of the voice to be quality inspected. When the voice segment is a ringtone or silent, the decoded text is empty; the second voice segment refers to the voice segment with a non-empty decoded text among the several voice segments to be quality inspected. Only when a non-ringtone and non-silent human voice is detected, the decoded text is non-empty.

[0123] In an embodiment of the present application, the several voice segments of the voice to be quality inspected are decoded through decoding. When an empty text is obtained, the first voice segment corresponding to the ringtone result or the silent result is detected, and a quality inspection warning is generated and sent to the backend of the quality inspection system, so that the quality inspection system skips the first voice segment; when a non-empty text is obtained, the second voice segment corresponding to the voice classification result is obtained, and no quality inspection warning interruption is performed on the quality inspection system.

[0124] It should be noted that the decoded text is analyzed, and text role classification is performed through text analysis. In order to ensure that only the connected data is quality inspected, a prompt needs to be output to the backend quality inspection system. If the decoding is empty, the first classification model fails to recognize the human voice or the human voice appears before the call is connected (when the call is not connected before the ringtone ends, the sounds in the voice are all invalid data). At this time, a quality inspection warning (Warning) needs to be generated for the backend quality inspection system. The backend quality inspection system does not perform quality inspection on this voice segment after capturing the Warning.

[0125] It should be noted that the invalid data that is not quality inspected includes: voice segments that do not recognize the human voice (such as silence) or voice segments where the human voice appears before the call is connected (such as ringtones).

[0126] It should be noted that the first classification model classifies the first audio feature signal of the voice to be quality inspected, and the obtained classification result is for the first audio feature signal. The voice result of the first audio feature signal is classified a second time according to the speaker, and the obtained result is still for the first audio feature signal. During the quality inspection process, it is for the voice to be quality inspected, and the quality inspection is carried out in sequence according to the recording start time of the voice to be quality inspected. Therefore, although the ringtone result, silent result, and voice result of the first audio feature signal, as well as several voice segments of different speakers in the voice result, are obtained, during the quality inspection, the quality inspection will still be carried out in sequence according to the voice to be quality inspected, including the quality inspection of the ringtone result, silent result, and several voice segments of different speakers. Since the decoded data corresponding to the ringtone result and silent result is empty, a quality inspection warning can be generated during the quality inspection to skip the segments of the voice to be quality inspected with empty decoded data and perform quality inspection on the segments of the voice to be quality inspected with non-empty decoded data.

[0127] The present invention uses the second audio feature signal obtained by splicing the voice feature signal and the ringtone feature signal to train the second classification model for three-class classification, which can greatly enrich the training data set with ringtones. Therefore, during training, the generalization ability and classification accuracy of the second classification model can be improved, and more accurate classification can be made according to the pre-trained first classification model; the voice to be quality inspected is classified into three classes according to the first classification model, and the voice to be quality inspected is divided according to the voice classification result of the second classification. Since the voice classification result does not include the voice classification result in the ringtone, when the obtained voice segments are decoded using the voice classification result, the voice in the ringtone can be accurately classified as the ringtone result of an unanswered call, so that the quality inspection can be carried out after accurately determining the connection moment, skipping the ringtone result or silent result for quality inspection, thereby improving the quality inspection efficiency and saving the quality inspection time cost and labor cost.

[0128] Embodiment 2

[0129] See Figure 11, which is a schematic diagram of voice quality inspection provided by an embodiment of the present application. In the figure, first, the audio feature signals of the voice to be inspected are classified to obtain the audio feature signals corresponding to the ringtone result, the audio feature signals corresponding to the mute result, and the audio feature signals corresponding to the human voice result. Then, the audio feature signals of the human voice result are classified again according to different speakers, and the voice to be inspected is segmented according to the voice classification result of the human voice, obtaining the voice classification results of two speakers; then, the inspection is carried out according to the time sequence of the voice to be inspected. During the inspection process, the audio feature signals corresponding to the ringtone result, the audio feature signals corresponding to the mute result, and the audio feature signals corresponding to the human voice result are all decoded. When the decoding is empty, a quality inspection warning is generated, so that the segment of the voice to be inspected corresponding to the audio feature signal of the quality inspection warning is skipped by the backend; when the decoding is not empty, it does not affect the quality inspection, and the backend continues the quality inspection.

[0130] It should be noted that since the segments of the voice to be inspected with empty decoding data are skipped, in fact, the quality inspection processes of multiple segments of the voice to be inspected with non-empty decoding data are continuous.

[0131] It should be noted that the present application focuses on the quality inspection of the voice segments of the speaker, and skips the quality inspection of the voice segments of non-speakers. Therefore, the voice to be inspected is segmented using the voice classification result of the speaker, and the obtained several voice segments include: the second voice segment corresponding to the separate human voice result, and the first voice segment corresponding to the ringtone result or the mute result. It should be noted that the obtained first voice segments include: the voice segment of the separate ringtone result, the voice segment of the separate mute result, and the voice segment connecting the ringtone result and the mute result. The present application does not care which of the first voice segments are the voice segments of the ringtone result and which are the voice segments of the mute result, because as long as it is a first voice segment, its decoding data must be empty, and the voice segments corresponding to the ringtone result and the voice segments corresponding to the mute result are both invalid voice data, and the present application skips the inspection of the invalid voice data. By classifying the voice segments with human voices in the ringtone as the first voice segments of non-speakers, the present application ensures that the voice of the speaker must be the second voice segment during the call, and thus can accurately identify the effective second voice segment after connection, and only performs quality inspection on the second voice segment of the speaker.

[0132] In Figure 11Among them, the voice quality inspection voice to be inspected is segmented according to the voice classification results of the speaking subject 1 and the speaking subject 2, and multiple voice segments of the speaking subject 1, multiple voice segments of the speaking subject 2, and the remaining voice segments corresponding to the ringtone results or the mute results can be obtained. The voice segments of the ringtone results and the mute results are used as invalid data, and the voice segments of the speaking subject 1 and the speaking subject 2 are used as valid data. The decoding of the invalid data is empty, and a quality inspection warning is issued without performing quality inspection on the invalid data, but only performing quality inspection on the voice segments of the speaking subject 1 and the speaking subject 2.

[0133] The present invention classifies the voice to be inspected into three categories according to the first classification model, and divides the voice to be inspected according to the voice classification results of the second classification. Since the voice classification results do not include the voice classification results in the ringtone, when decoding the obtained voice segments using the voice classification results, the voice in the ringtone can be accurately classified as the ringtone result of an unanswered call, so that quality inspection can be performed after accurately determining the connection moment, so as to skip the ringtone results or the mute results for quality inspection, enabling the backend to continuously perform quality inspection on the valid data, thereby improving the quality inspection efficiency and saving the quality inspection time cost and labor cost.

[0134] Embodiment 3

[0135] See Figure 12 , which is a schematic diagram of the complete quality inspection process provided by the embodiments of the present application. In the figure, the steps of quality inspecting the voice to be inspected are shown, including steps S301 to S306, specifically:

[0136] Step S301, obtain the voice to be inspected, and extract the audio feature signal of the voice to be inspected.

[0137] Step S302, voice classification. Among them, the audio feature signal is used as the input of the first classification model, and the three-classification result of the voice to be inspected is output, including: ringtone result, mute result, and voice result.

[0138] Specifically, in this embodiment, there are 2 speaking subjects, the customer and the customer service. During the test, after extracting the audio feature signal of the voice to be inspected, it is directly used as the input of the first classification model, and the first classification model outputs the classification result corresponding to each segment of the voice to be inspected, that is, the ringtone result is the set of audio feature signals of all ringtone classifications, the mute result is the combination of audio feature signals of all mute classifications, and the voice result is the set of audio feature signals of all voice classifications.

[0139] Step S303, voice classification of the speaking subject. Specifically, the voice result is classified according to different speaking subjects to obtain the sets of audio feature information of the two speaking subjects, and the voice to be inspected is divided into multiple voice segments according to the audio feature information of the two speaking subjects.

[0140] Step S304, decoding. Since the present application classifies the audio feature information where the ringtone and the human voice appear simultaneously as the ringtone result, that is, the non-human voice result, when before a formal call, the human voice of the speaking subject can be detected, and the human voice is classified as the ringtone result instead of the human voice result. Therefore, during quality inspection, decoding is performed from the starting position of the voice to be quality inspected.

[0141] Step S305, post-processing. When the obtained decoded data is empty, the current decoded data corresponds to the ringtone result, and a quality inspection warning needs to be generated so that the backend can skip the corresponding segment of the voice to be quality inspected according to the grabbed quality inspection warning, thereby improving the quality inspection efficiency; if a mute result is detected and the decoded data is still empty, a quality inspection warning also needs to be generated to skip this segment.

[0142] Step S306, quality inspection. Only when non-empty decoded data is obtained through decoding, the target voice for quality inspection is acquired, and no quality inspection warning is generated, so that the backend can continue with quality inspection.

[0143] When the present invention decodes the obtained voice segment using the human voice classification result, it can accurately classify the human voice in the ringtone as the ringtone result of an unanswered call, so that quality inspection can be performed after accurately determining the connection moment, skipping the ringtone result or the mute result for quality inspection, enabling the backend to continuously perform quality inspection on the valid data, thereby improving the quality inspection efficiency.

[0144] Embodiment 4

[0145] See Figure 13 , which is a training schematic diagram of the first classification model provided by the embodiment of the present application. In this embodiment, when training the second classification model based on TDNN-LSTM, the method of Figure 3 is adopted to extract the human voice feature signal, and at the same time, the method of overlapping and non-overlapping the ringtone feature signal obtained according to the generative adversarial model with the human voice feature signal is adopted to obtain the second audio feature signal. The audio feature signal is used as the input of the second classification model, and according to the corresponding label and the output of the second classification model, the second classification model performs supervised learning. When the loss function of TDNN-LSTM converges, the pre-trained first classification model is obtained.

[0146] The present invention adopts partial overlapping splicing of the ringtone feature signal and the first human voice feature signal, which can restore the real call scenario, construct a training data set that conforms to the actual situation, and use the generative adversarial model to construct a variety of ringtone feature signals, which can greatly enrich the samples of the training data set, facilitating the second classification model to accurately learn the characteristics of the ringtone feature signal with the interference of the human voice feature signal as interference, thereby facilitating the rapid convergence of the three-class second classification model and improving the generalization ability of the model.

[0147] Embodiment 5

[0148] See Figure 14 , which is a schematic structural diagram of the voice quality inspection device provided by the embodiment of the present invention. Based on the voice quality inspection method as described above, it includes: a first classification unit 41, a second classification unit 42, and a quality inspection unit 43.

[0149] Among them, the first classification unit 41 mainly performs three-classification on the voice to be quality-inspected to obtain classification results, including: ringtone result, mute result, and human voice result, and transmits the classification results to the second classification unit 42; the second classification unit 42 classifies the human voice result according to different speakers based on the received classification results to obtain a human voice classification result, and divides the voice to be quality-inspected into several voice segments using the human voice classification result, and transmits the several voice segments to the quality inspection unit 43; after receiving the several voice segments, the quality inspection unit 43 performs quality inspection.

[0150] The first classification unit 41 is configured to extract a first audio feature signal of the voice to be quality-inspected, and classify the first audio feature signal according to a pre-trained first classification model to obtain a classification result; wherein, the classification result includes: ringtone result, mute result, and human voice result; the human voice result does not include the human voice in the ringtone result.

[0151] In some embodiments of the present application, extracting the first audio feature signal of the voice to be quality-inspected includes: obtaining the voice to be quality-inspected, performing pre-emphasis on the voice to be quality-inspected in sequence, windowing the obtained high-frequency components according to a preset window and a preset step size, and performing discrete Fourier transform after windowing is completed to convert the time-domain information into frequency-domain information, processing the frequency-domain information through a filter and logarithmic calculation to eliminate harmonics, and performing discrete cosine transform on the obtained envelope information, and using the obtained acoustic feature signal as the first audio feature signal.

[0152] In an embodiment of the present application, for extracting the first audio feature signal of the voice to be quality-inspected, see Figure 3 , which is a schematic flow diagram for extracting the first audio feature signal provided by the embodiment of the present application, including sub-steps S111 to S116; among them, sub-steps S111 to S112 are preprocessing operations, and sub-steps S113 to S116 are feature signal extraction operations, specifically:

[0153] Sub-step S111, pre-emphasis. Specifically, pre-emphasize the voice signal of the voice to be quality-inspected to enhance the high-frequency components in the voice signal.

[0154] Sub-step S112, windowing. Specifically, in order to prevent spectral leakage, the sound is segmented according to a window of 25 ms and a step size of 10 ms, and Hamming window is used for windowing.

[0155] In other embodiments of the present application, the Hamming window can be replaced with a rectangular window or a Hann window.

[0156] Sub-step S113: Discrete Fourier Transform (DTF). Specifically, after performing a convolution operation according to the Hamming window, a discrete Fourier transform (DTF) is performed to convert the time domain to the frequency domain.

[0157] Sub-step S114: Filter filtering. Specifically, filter calculations are performed by simulating the perception of the human ear.

[0158] In a preferred embodiment of the present application, a Mel filter is used for filtering.

[0159] In an alternative embodiment of the present application, an equal-area Mel filter is used to obtain the frequency domain signal of the human voice.

[0160] It should be noted that using an equal-area Mel filter can simulate the scenario where the higher the frequency, the lower the perception of the real human ear, and the perception of low-frequency signals is more sensitive than that of high-frequency signals, that is, the equal-area Mel filter can simulate the process of the human ear processing human voice signals for the human voice in the first audio feature signal.

[0161] In another alternative embodiment of the present application, an equal-height Mel filter is used to process the high-frequency information of non-human voices in the first audio feature signal.

[0162] Sub-step S115: Logarithmic calculation. Among them, signal separation is performed through logarithmic calculation to retain the envelope signal.

[0163] It should be noted that through logarithmic calculation, the envelope signal can be retained. The envelope signal is the timbre, while the pitch is removed to obtain the target feature for speech recognition. Therefore, through filter filtering and logarithmic calculation, harmonics are eliminated and envelope information is retained.

[0164] Sub-step S116: Perform a discrete cosine transform (DCT) to obtain an acoustic feature signal, and use the acoustic feature signal as the first audio feature signal.

[0165] It should be noted that DCT processing can remove the correlation between different Mel frequency bands, so as to ensure low correlation between the features of the obtained first audio feature signal, which exactly meets the requirements of the first classification model for the input. Therefore, it is beneficial to improve the classification accuracy of the first classification model. Through the above 6 sub-steps, the time domain information can be converted into frequency domain information with low correlation between features to meet the requirements of the first classification model for the input.

[0166] In some embodiments of the present application, the first classification model is trained based on the second audio feature signal obtained by splicing the human voice feature signal and the constructed ringtone feature signal and the label of the second audio feature signal, including: sub-steps S121 to S123, see Figure 4 , which is a schematic flowchart of the second classification model provided by the embodiments of the present application to obtain the first classification model through training. Specifically:

[0167] Sub-step S121: Learn the feature distribution of real ringtones through the initial first generative adversarial model, and construct acoustic feature signals of multiple ringtones as ringtone feature signals according to the pre-trained second generative adversarial model.

[0168] In some embodiments of the present application, before splicing the human voice feature signal and the pre-constructed ringtone feature signal to obtain the second audio feature signal, the ringtone feature signal is constructed. Specifically, the generative network outputs multiple acoustic feature signals, and the multiple acoustic feature signals are used as ringtone feature signals; wherein, the generative network is trained by using the acoustic feature signal of the real ringtone as the real sample for input, outputting the generated sample, and discriminating the generated sample and the real sample through the discriminative network.

[0169] In some embodiments of the present application, the generative network is trained by using the acoustic feature signal of the real ringtone as the real sample for input, outputting the generated sample, and discriminating the generated sample and the real sample through the discriminative network, including: inputting the acoustic feature signal of the real ringtone as the real sample into the generative network, the generative network constructs a generated sample close to the acoustic feature signal of the real ringtone, and uses the generated sample and the real sample as the input of the discriminative network, and the discriminative network outputs the discriminative result of the generated sample; according to the discriminative result, the generated sample and the real sample, the generative network and the discriminative network are trained in sequence to respectively obtain the generative network and the discriminative network that have learned the feature distribution of the real ringtone; the trained generative network and the trained discriminative network are combined to form a generative adversarial network, and the multiple acoustic feature signals output by the generative adversarial network close to the real ringtone are used as ringtone feature signals.

[0170] In an embodiment of the present application, see Figure 5, which is a schematic flow chart of generating a bell feature signal by a generative adversarial model provided in an embodiment of the present application. The generative adversarial model learns the distribution of real bells and constructs various acoustic feature signals close to real bells. In the figure, the acoustic feature signal of the real bell is used as the real sample x, G is the generative network, and the hidden vector z is used to approximate the real sample x to obtain the generated sample G(z); D is the discriminative network, which receives the real sample x and the generated sample G(z) and outputs the probability that the generated sample G(z) belongs to the real sample x.

[0171] It should be noted that the training objective of the entire generative adversarial network is to make the generated sample G(z) closer to the real sample x, while the discriminative network D can better distinguish the authenticity of the generated sample.

[0172] In a preferred embodiment of the present application, after preprocessing the voice signal of the bell, a time-domain feature signal is obtained, and the time-domain feature signal is used as a convolutional neural-deconvolutional network based on generative adversarial to directly obtain a frequency-domain signal, and an acoustic feature signal of the bell is obtained.

[0173] Sub-step S122: Concatenate the human voice feature signal with the pre-constructed bell feature signal to obtain a second audio feature signal, and mark the labels of the bell segment, human voice segment, and silent segment in the second audio feature signal as bell, human voice, and silent respectively.

[0174] In an embodiment of the present application, the steps of extracting the first human voice feature signal are the same as those of extracting the first audio feature signal, specifically: obtaining the human voice to be trained, performing pre-emphasis on the human voice in sequence, windowing the obtained high-frequency components according to a preset window and a preset step length, and performing discrete Fourier transform after windowing to convert the time-domain information into frequency-domain information, processing the frequency-domain information through a filter and logarithmic calculation to eliminate harmonics, and performing discrete cosine transform on the obtained envelope information, and using the obtained acoustic feature signal as the first human voice feature signal.

[0175] In an embodiment of the present application, concatenating the human voice feature signal with the pre-constructed bell feature signal to obtain a second audio feature signal, and marking the labels of the bell segment, human voice segment, and silent segment in the second audio feature signal as bell, human voice, and silent respectively, includes: overlapping and concatenating the human voice feature signal with the pre-constructed bell feature signal according to a preset length to obtain a second audio feature signal, and marking the label of the overlapping feature signal in the obtained second audio feature signal as bell, and marking the labels of the human voice segment and non-human voice segment in the remaining non-overlapping feature signals as human voice and silent respectively.

[0176] See Figure 6, which is a schematic diagram of the second audio feature signal obtained by overlapping splicing in the embodiments of the present application. In the figure, the acoustic feature signal of the ringtone is overlapped and spliced with the acoustic feature signal of the human voice to obtain the second audio feature signal. When tagging, since the human voice appears simultaneously with the ringtone, it is considered that the human voice is interference to the ringtone. When the tag of the simultaneous appearance of the human voice and the ringtone is the ringtone, the tags of the human voice segment and the non-human voice segment in the remaining non-overlapping feature signals are respectively recorded as human voice and silence.

[0177] In another embodiment of the present application, the human voice feature signal is spliced with a pre-constructed ringtone feature signal to obtain a second audio feature signal, and the tags of the ringtone segment, the human voice segment, and the silence segment in the second audio feature signal are respectively recorded as ringtone, human voice, and silence, including: splicing the human voice feature signal and the pre-constructed ringtone feature signal without overlap to obtain a second audio feature signal, and recording the tags of the ringtone segment and the human voice segment in the second audio feature signal as ringtone and human voice respectively, and recording the tag of the remaining segment after removing the ringtone segment and the human voice segment as silence.

[0178] In another embodiment of the present application, the ringtone feature signal and the human voice feature signal can also be spliced without overlap, and the tags of the ringtone segment, the human voice segment, and the non-human voice segment in the obtained spliced feature signal are respectively recorded as ringtone, human voice, and silence.

[0179] See Figure 7 , which is a schematic diagram of the second audio feature signal obtained by non-overlapping splicing in the embodiments of the present application. In the figure, the acoustic feature signal of the ringtone is non-overlappingly spliced with the acoustic feature signal of the human voice to obtain a second audio feature signal. When tagging,

[0180] The tags of the ringtone segment and the human voice segment in the second audio feature signal are respectively recorded as ringtone and human voice, and the tag of the remaining segment after removing the ringtone segment and the human voice segment is recorded as silence.

[0181] The present invention adopts partial overlapping splicing and non-overlapping splicing of the ringtone feature signal and the human voice feature signal, which can restore the real call scenario, construct a training data set that conforms to the actual situation, and adopt a generative adversarial model to construct a variety of ringtone feature signals, which can greatly enrich the samples of the training data set, facilitate the second classification model to accurately learn the features of the ringtone feature signal with the interference of the human voice feature signal, so as to facilitate the rapid convergence of the three-class second classification model and improve the generalization ability of the model.

[0182] In another preferred embodiment of the present application, the ringtone feature signal and the human voice feature signal are overlapped and spliced according to a preset length, and the ringtone feature signal and the human voice feature signal are spliced without overlap, and tags for the ringtone, human voice, and silence are respectively marked.

[0183] It should be noted that each sub - segment in the second audio feature signal needs to be labeled to obtain complete label data.

[0184] In some embodiments of the present application, after obtaining the labels and the second audio feature signal, data pre - processing is performed to obtain a training data set in a preset format. Specifically, the training data set contains four columns of data. The first column is the label of each sub - segment in the second audio feature signal, the second column is the file name, the third column is the start time of the sub - segment, and the fourth column is the end time of the sub - segment. The name format is "file name - start time - end time - label"; if it is a human voice, the label is speech, if it is a ringtone, the label is garbage, and the rest of the sounds are non - human voice segments (such as silence).

[0185] See Figure 8 , which is a schematic diagram of the format of the training data set provided by the embodiments of the present application. In the figure, according to the name format "file name - start time - end time - label", there are 10 sub - segments of the second audio feature signal, and according to the naming format of the shown training data set, multi - classification training is performed on the second classification model.

[0186] Sub - step S123: Perform multi - classification training on the initial second classification model according to the label and the second audio feature signal to obtain the first classification model.

[0187] In some embodiments of the present application, performing multi - classification training on the initial second classification model according to the label and the second audio feature signal to obtain the first classification model includes: inputting the second audio feature signal into the initial second classification model, and constructing an initial first decoding graph of the second classification model according to the state transition probability of the second audio feature signal; outputting the training result of the second audio feature signal according to the first decoding graph, and training the second classification model according to the training result and the label to obtain the first classification model; where the state transition probability is the probability of mutual conversion between different states; the states include: silence state, human voice state, and ringtone state.

[0188] In some embodiments of the present application, constructing the initial first decoding graph of the second classification model according to the state transition probability of the second audio feature signal includes:

[0189] Construct the initial first decoding graph of the second classification model according to the preset longest duration and shortest duration of the state, the transition probability of the initial state of the second audio feature signal, the probability of returning to its own state, and the mutual transition probabilities between different states.

[0190] In some embodiments of the present application, the longest duration is that when the number of transitions to the same state after state transition reaches the first transition number threshold, the next state is different from the current state; the shortest duration is that during multiple state transitions, it can only be in the same state, and when the number of transitions reaches the second transition number threshold, the next state is different from the current state.

[0191] In some embodiments of the present application, by presetting the shortest duration and the longest duration of the ringtone state, the recognition of the ringtone is accurately improved. Specifically, the longest duration of the preset ringtone state is Tmax. After Tmax state transitions and still in the ringtone state, the next state will be other states; the shortest duration of the preset ringtone state is Tmin. If within Tmin, the state can only remain in the ringtone state, and after more than Tmin transitions, it can turn to another state; where Tmax and Tmin are both positive integers.

[0192] It should be noted that the decoding graph is constructed through the transition probability from the initial state to three states, including: the mute state, the human voice state, and the ringtone state; the probability of returning to its own state, the mutual transition probabilities between different states and the probability of returning to its own state, as well as the shortest duration and the longest duration of the state.

[0193] See Figure 9 , which is a schematic diagram of the decoding graph provided by the embodiments of the present application. In the figure, it includes the transitions from the mute state to the human voice state and the ringtone state, the transitions from the human voice state to the mute state and the ringtone state, the transitions from the ringtone state to the human voice state and the mute state, as well as the mute state remaining in its own state, the human voice state remaining in its own state, and the ringtone state remaining in its own state, a total of 9 state probabilities.

[0194] Exemplarily, the three states are denoted as the mute state (1), the human voice state (2), and the ringtone state (3), then there are a total of 9 transition probabilities: p11, p12, p13, p21, p22, p23, p31, p32, and p33. pij refers to the transition probability from state i to state j. At the same time, the shortest duration (Tmin) and the longest duration (Tmax) of each state are set.

[0195] It should be noted that, in this embodiment, the longest duration means that after Tmax state transitions, it remains in the same state, and then the next state will be another state; the shortest duration means that if within Tmin, the state can only remain in the current state, and it can only transition to another state after exceeding Tmin.

[0196] In an embodiment of the present application, a time delay neural network (TDNN)-long short-term memory (LSTM) model is used as the initial second classification model, and the second audio feature signal is used as the input of the TDNN-LSTM network model for supervised multi-classification training.

[0197] In an embodiment of the present application, refer to Figure 10 , which is a schematic diagram of the network structure of the TDNN provided by the embodiment of the present application. In the figure, it includes 4 layers of networks. The multi-hidden layer has stronger feature extraction capabilities. Since adjacent neurons in each layer have overlapping information, each layer of neurons extracts features through some neurons rather than all neurons to reduce the first classification model and improve the training rate of three-classification (including: silence, human voice, and ringtone).

[0198] In an embodiment of the present application, two fully connected layers are connected after the TDNN to obtain the input feature vector.

[0199] It should be noted that in step S11, the classification results of the audio feature signal have been obtained, including the ringtone result, the silence result, and the human voice result, which are the classification results of the first audio feature signal corresponding to the ringtone result, the silence result, and the human voice result, rather than the classification result of the speech to be quality inspected. And, the human voice result includes the audio feature signals of multiple speakers. Therefore, it is also necessary to classify the audio feature signals of multiple speakers.

[0200] The second classification unit 42 is used to cluster the human voice result to obtain the human voice classification results corresponding to different speakers, and segment the speech to be quality inspected according to the human voice classification results to obtain several speech segments.

[0201] In an embodiment of the present application, the several speech segments include: the first speech segment corresponding to the ringtone result or the silence result.

[0202] In an embodiment of the present application, by extracting the voice features (x-vector features) of different speakers, and then classifying the voice features through hierarchical clustering, the voice classification results corresponding to different speakers are obtained. In order to extract the x-vector network, the TDNN network is still used. The input of the network is the acoustic feature signal, and two fully connected layers are connected after the TDNN network.

[0203] In an embodiment of the present application, the output of the fully connected layer of the TDNN in the TDNN-LSTM network model is used as the voice feature.

[0204] In an embodiment of the present application, clustering is performed on the voice results to obtain the voice classification results corresponding to different speakers, including: using the output of the fully connected layer of the last layer in the first classification model as the voice feature, obtaining the distance matrix of all voice features, and continuously aggregating the voice features with the closest distances in the distance matrix until the number of clusters in the distance matrix is equal to the number of speakers, so as to obtain the voice classification results corresponding one by one to different speakers.

[0205] In an embodiment of the present application, by obtaining the output of the fully connected layer of the last layer of the TDNN in the TDNN-LSTM network model, the voice features of the speakers of each sub-segment in the voice results are extracted, denoted as x-vector. There are N sub-segments in total, so there are N-dimensional x-vectors. By calculating the distances between each pair, an N*N distance matrix is formed, and then through hierarchical clustering, the points with the closest distances are continuously merged. In the customer service scenario, assuming there are two speakers, namely the customer service and the customer, the number of clusters is set to 2. Hierarchical clustering continuously merges the data until the number of clusters is equal to 2, and the clustering ends.

[0206] The quality inspection unit 43 is used to decode the several voice segments in sequence. If the decoding result is empty, a quality inspection warning is generated; the quality inspection warning is used to indicate that there is no need to perform quality inspection on the voice segment with an empty decoding result.

[0207] In some embodiments of the present application, if the decoded data obtained by decoding is not empty, no quality inspection warning is generated, so that the backend can perform quality inspection on the voice segment.

[0208] In some embodiments of the present application, the several voice segments further include a second voice segment; if the decoding result is empty, a quality inspection warning is generated, specifically: when the first decoded data obtained by decoding is empty, the first voice segment is detected and a quality inspection warning is generated; otherwise, if the second decoded data obtained by decoding is not empty, the second voice segment corresponding to the voice classification result is detected, and no quality inspection warning is generated, so that the backend can perform quality inspection on the second voice segment.

[0209] In an embodiment of the present application, several voice segments of the voice to be quality inspected are decoded. When an empty text is obtained, if the first voice segment corresponding to the ringtone result or the mute result is detected, a quality inspection warning is generated and sent to the back-end of the quality inspection system, so that the quality inspection system skips the first voice segment; when a non-empty text is obtained, if the second voice segment corresponding to the human voice classification result is obtained, the quality inspection system is not interrupted by the quality inspection warning.

[0210] It should be noted that the decoded text is analyzed, and the text role classification is performed through text analysis. To ensure that only the connected data is quality inspected, a prompt needs to be output to the back-end quality inspection system. If the decoding is empty, it means that the first classification model fails to recognize the human voice or the human voice appears before the call is connected (before the ringtone ends, that is, when the call is not connected, the sounds in the voice are all invalid data). At this time, a quality inspection warning (Warning) needs to be generated for the back-end quality inspection system. After the back-end quality inspection system captures the Warning, it does not perform quality inspection on this voice segment.

[0211] It should be noted that the invalid data that is not quality inspected includes: voice segments that fail to recognize the human voice or the human voice appears before the call is connected.

[0212] It should be noted that the first classification model classifies the first audio feature signal of the voice to be quality inspected, and the obtained classification result is for the first audio feature signal. The human voice result of the first audio feature signal is classified a second time according to the speaking subject, and the obtained result is still for the first audio feature signal. During the quality inspection process, it is for the voice to be quality inspected, and the quality inspection is performed in sequence according to the recording start time of the voice to be quality inspected. Therefore, although the ringtone result, mute result, and human voice result of the first audio feature signal, as well as several voice segments of different speaking subjects of the human voice result, are obtained, during the quality inspection, the quality inspection is still performed on the voice to be quality inspected in chronological order, including the quality inspection of the ringtone result, mute result, and several voice segments of different speaking subjects. Since the decoded data corresponding to the ringtone result and the mute result is empty, a quality inspection warning can be generated during the quality inspection to skip the segments of the voice to be quality inspected with empty decoded data and perform quality inspection on the segments of the voice to be quality inspected with non-empty decoded data.

[0213] The present invention uses a first classification unit 41 to perform three-class training on a second classification model using a second audio feature signal obtained by splicing a first human voice feature signal and a ringtone feature signal, which can greatly enrich the training data set with ringtones. Therefore, during training, the generalization ability and classification accuracy of the second classification model can be improved, and more accurate classification can be made according to the pre-trained first classification model; the voice to be quality inspected is classified into three categories according to the first classification model, and the first classification unit 41 is used to divide the voice to be quality inspected according to the human voice classification result of the second classification. Since the human voice classification result does not include the human voice classification result in the ringtone, when the quality inspection unit 43 decodes the obtained voice segment using the human voice classification result, the human voice in the ringtone can be accurately classified as the ringtone result of an unanswered call, so that quality inspection can be performed after accurately determining the connection moment, skipping the ringtone result or the mute result for quality inspection, thereby improving the quality inspection efficiency and saving the quality inspection time cost and labor cost.

[0214] Embodiment 6

[0215] It is a readable computer storage medium provided by an embodiment of the present application, on which a computer program is stored. When the computer program is executed by a processor, the voice quality inspection method as described is implemented.

[0216] In the present application, by splicing the ringtone feature signal and the human voice feature signal and training according to the obtained audio feature signal, a three-class classification including the ringtone result is obtained, so as to facilitate identifying invalid data of a ringtone containing a human voice as a ringtone during testing. When the quality inspection method of skipping invalid data to continuously quality inspect valid data is stored in the storage medium in the form of a program, operations such as quickly quality inspecting the voice to be quality inspected can be performed by running or reading the executable program in the storage medium, which is applicable to more operating systems and different application platforms and has stronger scalability.

[0217] Embodiment 7

[0218] See Figure 15 , which is a schematic diagram of an electronic device provided by an embodiment of the present application. The electronic device includes: a memory 71, a processor 72, and a computer program stored on the memory 71 and executable on the processor. When the processor 72 executes the computer program, the steps of the distributed file processing method as described are implemented.

[0219] In some embodiments of the present application, the electronic device further includes: a communication interface 73 and a communication bus 74; wherein, the processor 72, the communication interface 73, and the memory 71 complete communication with each other through the communication bus 74.

[0220] In this application, the ringtone feature signal and the human voice feature signal are spliced, and training is performed according to the obtained audio feature signal to obtain a three-classification including the ringtone result, so as to facilitate the identification of invalid data where the ringtone containing human voice is recognized as a ringtone during testing. When the quality inspection method of skipping invalid data and continuously performing quality inspection on valid data during quality inspection is integrated into an electronic device, rapid on-site quality inspection and feedback of on-site quality inspection can be carried out through various electronic devices, and the scalability is stronger.

[0221] Those skilled in the art should understand that the embodiments of this application may also provide a computer program product. Therefore, this application may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, this application may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0222] This application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of this application. It should be understood that each process and / or block in the flowchart and / or block diagram, and the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the specified functions in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0223] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device implements the specified functions in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0224] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are performed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the specified functions in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0225] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the technical principle of the present invention, several improvements and modifications can be made, and these improvements and modifications should also be regarded as the protection scope of the present invention.

Claims

1. A voice quality inspection method, characterized in that, Including: Extracting a first audio feature signal of the voice to be quality-inspected, and classifying the first audio feature signal according to a pre-trained first classification model to obtain a classification result; wherein, the classification result includes: a ringtone result, a mute result, and a human voice result; the human voice result does not include the human voice in the ringtone result. Clustering the human voice result to obtain a human voice classification result corresponding to different speakers, and segmenting the voice to be quality-inspected according to the human voice classification result to obtain a plurality of voice segments. Successively decoding the plurality of voice segments, and if the decoding result is empty, generating a quality inspection warning; the quality inspection warning is used to indicate that there is no need to perform quality inspection on the voice segment with an empty decoding result.

2. The voice quality inspection method according to claim 1, wherein The training steps of the first classification model include: Splicing a human voice feature signal with a pre-constructed ringtone feature signal to obtain a second audio feature signal, and respectively denoting the labels of the ringtone segment, human voice segment, and mute segment in the second audio feature signal as ringtone, human voice, and mute. Performing multi-classification training on an initial second classification model according to the label and the second audio feature signal to obtain the first classification model.

3. The voice quality inspection method according to claim 2, wherein Before the step of splicing the human voice feature signal with the pre-constructed ringtone feature signal to obtain a second audio feature signal, it further includes: Outputting a plurality of acoustic feature signals by a pre-trained generation network, and using the plurality of acoustic feature signals as ringtone feature signals; wherein, the generation network is trained by taking the acoustic feature signal of a real ringtone as a real sample for input, outputting a generated sample, and discriminating the generated sample and the real sample by a discriminant network.

4. The voice quality inspection method according to claim 2, wherein, The step of performing multi-classification training on an initial second classification model according to the label and the second audio feature signal to obtain the first classification model includes: Inputting the second audio feature signal into the initial second classification model, and constructing an initial first decoding graph of the second classification model according to the state transition probability of the second audio feature signal. Outputting a training result of the second audio feature signal according to the first decoding graph, and training the second classification model according to the training result and the label to obtain the first classification model. Wherein, the state transition probability is the probability of mutual conversion between different states; the states include: a mute state, a human voice state, and a ringtone state.

5. The voice quality inspection method according to claim 4, wherein The step of constructing an initial first decoding graph of the second classification model according to the state transition probability of the second audio feature signal includes: Constructing an initial first decoding graph of the second classification model according to the preset longest duration and shortest duration of the state, the conversion probability of the initial state of the second audio feature signal, the probability of returning to its own state, and the mutual conversion probability between different states.

6. The voice quality inspection method according to claim 2, wherein, The step of splicing the human voice feature signal with the pre-constructed ringtone feature signal to obtain a second audio feature signal, and respectively denoting the labels of the ringtone segment, human voice segment, and mute segment in the second audio feature signal as ringtone, human voice, and mute, includes: Overlap and splice the human voice feature signal and the pre-constructed ringtone feature signal according to a preset length to obtain a second audio feature signal. Denote the label of the overlapping feature signal in the obtained second audio feature signal as ringtone, and denote the labels of the human voice segment and the non-human voice segment in the remaining non-overlapping feature signal as human voice and silence respectively.

7. The voice quality inspection method according to claim 2, wherein The splicing of the human voice feature signal and the pre-constructed ringtone feature signal to obtain a second audio feature signal, and denoting the labels of the ringtone segment, the human voice segment and the silence segment in the second audio feature signal as ringtone, human voice and silence respectively, includes: Non-overlappingly splice the human voice feature signal and the pre-constructed ringtone feature signal to obtain a second audio feature signal, and denote the labels of the ringtone segment and the human voice segment in the second audio feature signal as ringtone and human voice respectively, and denote the label of the remaining segment after removing the ringtone segment and the human voice segment as silence.

8. The voice quality inspection method according to claim 1, characterized in that The clustering of the human voice result to obtain the human voice classification results corresponding to different speakers includes: Take the output of the fully connected layer of the last layer in the first classification model as the human voice feature, obtain the distance matrix of all human voice features, and continuously aggregate the human voice features with the closest distances in the distance matrix until the number of clusters in the distance matrix is equal to the number of speakers, to obtain the human voice classification results corresponding one by one to different speakers.

9. The voice quality inspection method according to claim 1, characterized in that, The method further includes: If the decoded data obtained by decoding is not empty, no quality inspection warning is generated, so that the backend can perform quality inspection on the voice segment.

10. A voice quality inspection device, characterized in that, It includes: A first classification unit, a second classification unit and a quality inspection unit; wherein, The first classification unit is used to extract the first audio feature signal of the voice to be quality inspected, and classify the first audio feature signal according to a pre-trained first classification model to obtain a classification result; wherein, the classification result includes: ringtone result, silence result and human voice result; the human voice in the human voice result does not include the human voice in the ringtone result; The second classification unit is used to cluster the human voice result to obtain the human voice classification results corresponding to different speakers, and segment the voice to be quality inspected according to the human voice classification results to obtain a plurality of voice segments; The quality inspection unit is used to decode the plurality of voice segments in sequence. If the decoding result is empty, a quality inspection warning is generated; the quality inspection warning is used to indicate that there is no need to perform quality inspection on the voice segment with an empty decoding result.

11. An electronic device, characterized in that, The electronic device includes: a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it implements the steps of the voice quality inspection method according to any one of claims 1-9.

12. A readable computer storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the voice quality inspection method according to any one of claims 1-9.