Voice quality inspection method and device, computer equipment and storage medium
Through multimodal feature fusion technology, combined with speech-to-text, contextual semantics and acoustic features, and combined with pre-trained emotion recognition models, the delay and scalability problems of existing speech quality inspection technologies are solved, and higher accuracy speech quality inspection is achieved.
Patent Information
- Application Number
- CN202510865267.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-25
- Publication Date
- 2025-09-23
- Estimated Expiration
- 2045-06-25
AI Technical Summary
Existing speech quality inspection technology cannot achieve subsequent category expansion and has poor scalability. At the same time, the model has computational delays during inference and cannot be monitored in real time.
Using multimodal feature fusion technology, through speech-to-text conversion, contextual semantic recognition, acoustic feature extraction and pre-trained emotion recognition model, semantic, acoustic and emotional feature vectors are obtained, deeply integrated and output as speech quality inspection results.
It significantly improves the accuracy and intelligence of voice quality inspection, and can reflect the voice content more comprehensively and meticulously, meeting the high-quality automatic quality inspection needs in the financial and healthcare fields.
Smart Images

Figure CN120690231A_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the field of artificial intelligence technology, and specifically relates to a speech quality inspection method, apparatus, computer equipment, and storage medium. Background Art
[0002] In quality inspection or voice service operations, call records are often recorded to facilitate subsequent performance evaluations and service level improvements. For call evaluation systems, the emotions of both the customer and the salesperson are crucial factors in measuring service quality, and the volume and pitch of voice often reflect emotions. For example, in financial services or healthcare and elderly care customer service management, during voice quality inspections of conversations between financial sales representatives and customers, if the salesperson speaks in an impatient tone or loudly, this can lead to customer dissatisfaction, impacting both business progress and the customer experience. Furthermore, if the customer is emotionally agitated and their voice is sharp, the service effectiveness can be significantly reduced if the salesperson fails to appease them promptly. Therefore, accurately extracting emotional information from speech is a challenge that needs to be addressed.
[0003] To address the above issues, current solutions and technologies focus on using deep learning networks to extract features and then classify them. However, such methods cannot achieve subsequent category expansion and have poor scalability. At the same time, the model's calculations will be delayed during inference, making real-time monitoring impossible. Summary of the Invention
[0004] The purpose of the embodiments of the present application is to propose a speech quality inspection method, device, computer equipment and storage medium to solve the technical problems of existing speech quality inspection that cannot achieve subsequent category expansion, poor scalability, and the calculation of the model during reasoning will be delayed, making real-time monitoring impossible.
[0005] In order to solve the above technical problems, the present application provides a method for voice quality inspection, which adopts the following technical solutions:
[0006] A voice quality inspection method, comprising:
[0007] Acquire the voice data to be processed, and perform speech-to-text processing on the voice data to be processed to obtain the text to be processed;
[0008] Perform contextual semantic recognition on the text to be processed, and perform feature embedding processing on the text semantic recognition results to obtain a semantic feature vector;
[0009] Extract acoustic features from the speech data to be processed to obtain acoustic feature vectors;
[0010] Based on the acoustic feature vector, emotion recognition is performed using a pre-trained emotion recognition model, and the emotion recognition results are subjected to feature embedding processing to obtain the emotion feature vector;
[0011] Perform multimodal feature fusion on the semantic feature vector, acoustic feature vector and emotional feature vector to obtain a multimodal feature vector;
[0012] The multimodal feature vector is input into the pre-trained speech quality inspection model, and the speech quality inspection result of the speech data to be processed is output.
[0013] In order to solve the above technical problems, the present application also provides a speech quality inspection device, which adopts the following technical solution:
[0014] A speech quality inspection device, comprising:
[0015] The text conversion module is used to obtain the voice data to be processed and perform voice-to-text processing on the voice data to be processed to obtain the text to be processed;
[0016] The semantic recognition module is used to perform contextual semantic recognition on the text to be processed and perform feature embedding processing on the text semantic recognition results to obtain a semantic feature vector;
[0017] The acoustic feature module is used to extract acoustic features from the speech data to be processed and obtain acoustic feature vectors;
[0018] The emotion recognition module is used to perform emotion recognition based on acoustic feature vectors using a pre-trained emotion recognition model, and perform feature embedding processing on the emotion recognition results to obtain emotion feature vectors;
[0019] A multimodal fusion module is used to perform multimodal feature fusion on the semantic feature vector, the acoustic feature vector, and the emotional feature vector to obtain a multimodal feature vector;
[0020] The speech quality inspection module is used to input the multimodal feature vector into the pre-trained speech quality inspection model and output the speech quality inspection results of the speech data to be processed.
[0021] In order to solve the above technical problems, the embodiment of the present application further provides a computer device, which adopts the following technical solution:
[0022] A computer device includes a memory and a processor, wherein the memory stores computer-readable instructions, and when the processor executes the computer-readable instructions, the steps of the speech quality inspection method as described in any one of the above items are implemented.
[0023] In order to solve the above technical problems, the embodiment of the present application further provides a computer-readable storage medium, which adopts the following technical solution:
[0024] A computer-readable storage medium stores computer-readable instructions, which, when executed by a processor, implement the steps of any one of the above-mentioned speech quality inspection methods.
[0025] Compared with the prior art, the embodiments of the present application have the following beneficial effects:
[0026] This application discloses a speech quality inspection method, apparatus, computer equipment, and storage medium, belonging to the field of artificial intelligence technology, and applied to speech quality inspection scenarios in the fields of finance, health care, etc. This application significantly improves the accuracy and intelligence level of speech quality inspection by introducing multimodal feature fusion technology. Based on the traditional speech-to-text conversion, it further extracts the contextual semantic features of the text and the acoustic features of the speech, and combines it with a pre-trained emotion recognition model to obtain emotional features, achieving a comprehensive understanding of speech data from the three dimensions of semantics, acoustics, and emotion. By deeply fusing the three types of feature vectors, a more expressive multimodal feature vector is obtained, which can more fully explore quality issues in speech, including abnormal behaviors in terms of service attitude, standardization of speech, and emotional fluctuations. Compared with traditional methods that rely solely on text or acoustic features, this application combines text or acoustic features with emotional features to achieve multimodal features. It can more comprehensively and meticulously reflect the information contained in the speech content, thereby improving the accuracy and practicality of speech quality inspection, and meeting the needs of high-quality automatic quality inspection in application scenarios such as intelligent customer service and speech auditing in the fields of finance, health care, etc. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] In order to more clearly illustrate the solutions in this application, a brief introduction will be given below to the drawings required for use in the description of the embodiments of this application. Obviously, the drawings described below are some embodiments of this application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0028] Figure 1 shows an exemplary system architecture diagram in which the present application can be applied;
[0029] Figure 2 A flow chart showing an embodiment of a voice quality inspection method according to the present application is shown;
[0030] Figure 3 A schematic structural diagram of an embodiment of a speech quality inspection device according to the present application is shown;
[0031] Figure 4 A schematic structural diagram of an embodiment of a computer device according to the present application is shown. DETAILED DESCRIPTION
[0032] Unless otherwise defined, all technical and scientific terms used herein have the same meanings as commonly understood by those skilled in the art to which this application belongs. The terms used in the specification of the application are for the purpose of describing specific embodiments only and are not intended to limit this application. The terms "including" and "having" and any variations thereof in the specification and claims of this application and the above-mentioned drawings are intended to cover non-exclusive inclusions. The terms "first", "second", etc. in the specification and claims of this application or the above-mentioned drawings are used to distinguish different objects, not to describe a specific order.
[0033] References herein to "embodiments" mean that a particular feature, structure, or characteristic described in connection with the embodiments may be included in at least one embodiment of the present application. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor does it constitute an independent or alternative embodiment that is mutually exclusive of other embodiments. It is understood, both explicitly and implicitly, by those skilled in the art that the embodiments described herein may be combined with other embodiments.
[0034] In order to enable those skilled in the art to better understand the solution of the present application, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the accompanying drawings.
[0035] like Figure 1 As shown, system architecture 100 may include a terminal device 101, a network 102, and a server 103. Terminal device 101 may be a laptop computer 1011, a tablet computer 1012, or a mobile phone 1013. Network 102 is a medium for providing a communication link between terminal device 101 and server 103. Network 102 may include various connection types, such as wired or wireless communication links or fiber optic cables.
[0036] The user can use the terminal device 101 to interact with the server 103 via the network 102 to receive or send messages, etc. Various communication client applications can be installed on the terminal device 101, such as web browser applications, shopping applications, search applications, instant messaging tools, email clients, social platform software, etc.
[0037] The terminal device 101 can be various electronic devices with a display screen and supporting web browsing. In addition to the laptop computer 1011, tablet computer 1012 or mobile phone 1013, the terminal device 101 can also be an e-book reader, an MP3 player (Moving Picture Experts Group Audio Layer III), an MP4 (Moving Picture Experts Group Audio Layer IV) player, a laptop computer and a desktop computer, etc.
[0038] The server 103 may be a server that provides various services, such as a background server that provides support for web pages displayed on the terminal device 101 .
[0039] It should be noted that the voice quality inspection method provided in the embodiment of the present application is generally executed by a server / terminal device, and accordingly, the voice quality inspection device is generally set in the server / terminal device.
[0040] It should be understood that Figure 1 The numbers of terminal devices, networks and servers in the embodiment are merely illustrative. The above system may have any number of terminal devices, networks and servers according to implementation requirements.
[0041] Continue to refer Figure 2 , shows a flow chart of an embodiment of a method for speech quality inspection according to the present application. The speech quality inspection method comprises the following steps:
[0042] S201, obtaining speech data to be processed, and performing speech-to-text processing on the speech data to be processed to obtain text to be processed;
[0043] Specifically, the system first obtains the voice data to be processed from a voice database or real-time audio stream. This data can come from call center agent conversations, conversations between voice robots and users, customer reports, or other service interaction audio. The audio format can be standard formats such as WAV, MP3, and PCM. To ensure the quality of the speech transcription, the speech is preprocessed before recognition, including noise reduction, silence removal, speech rate normalization, and speech enhancement (such as echo cancellation and bandpass filtering). The speech content is then transcribed into text using an automatic speech recognition (ASR) system. The ASR system can be based on an end-to-end deep learning model, such as an RNN-Transducer, CTC decoder, Transformer, or Conformer architecture. This speech recognition model can be optimized in conjunction with vertical language models to better adapt to specific industry terminology (such as "compensation liability," "policy number," and "surveyor"). During the speech-to-text conversion process, the system also annotates timestamp information such as pauses and emotional peaks.
[0044] S202, performing contextual semantic recognition on the text to be processed, and performing feature embedding processing on the text semantic recognition result to obtain a semantic feature vector;
[0045] Specifically, the system uses contextual semantic modeling technology to process the text obtained in S201. Pre-trained language models such as BERT, RoBERTa, and ERNIE are used, combined with syntactic analysis and named entity recognition, to conduct a deep semantic understanding of the text content. The model can identify service intentions (such as reporting, complaining, consulting, and claim collection, etc.), key business information (such as policy number, compensation amount, responsibility division), and logical relationships in the context (such as cause and effect, parallelism, transitions, etc.) in the text, and segment long texts to maintain semantic integrity. After processing is completed, the recognition results are subjected to feature embedding processing, including using Token Embedding, Segment Embedding, and Posit ion Embedding to encode the text structure to form a fixed-dimensional vector expression. This semantic feature vector can fully express the connection between user intentions and service behaviors, and is highly robust to context changes, so as to provide rich semantic basis.
[0046] S203, extracting acoustic features from the speech data to be processed to obtain an acoustic feature vector;
[0047] Specifically, the system extracts quantifiable acoustic parameters from speech signals to transform speech from the audio domain into a feature space. Common features include Mel-Frequency Cepstral Coefficients (MFCCs), Chroma features, spectral centroids, frequency band energy distribution, pitch, speaking rate, and short-term energy. These features can reveal a speaker's articulation, intonation, speech rate variations, and emphasis patterns. For example, higher pitch and energy fluctuations in speech may indicate emotional excitement, while abnormal speaking rate may indicate stress. During the extraction process, the system first performs frame segmentation and windowing on the speech signal. It then transforms the signal into the frequency domain using methods such as FFT and Mel filter banks. Acoustic features of different dimensions are then concatenated or reduced in dimensionality to generate a unified acoustic feature vector. To enhance the model's adaptability to complex acoustic environments, self-supervised acoustic coding models, such as wav2vec2.0 and Hubert, can be introduced to directly model the raw waveform to obtain more robust high-level acoustic representations. The resulting acoustic feature vector is used not only for emotion recognition but also for identifying quality control issues such as abnormal speaking rate and inappropriate pauses.
[0048] S204, based on the acoustic feature vector, using a pre-trained emotion recognition model to perform emotion recognition, and performing feature embedding processing on the emotion recognition result to obtain an emotion feature vector;
[0049] Specifically, the system uses a pre-trained emotion recognition model to further analyze acoustic features to identify emotional states, preferably using the open-source Emotion2vec model. Emotion2vec is a deep learning model that encodes speech signals into high-dimensional emotion vectors, capable of learning the combined characteristics of emotion, semantics, and context in speech. The Emotion2vec model uses a self-supervised learning method, pre-trained on a large amount of unlabeled public emotion data, to learn highly general speech emotion features. This approach does not require a large amount of labeled data, saving costs and providing strong adaptability.
[0050] When using this model, the system first performs noise reduction, framing, and windowing on the original speech before extracting its acoustic features and inputting them into emotion2vec. Through its neural encoder structure, the model maps the sound into an emotional embedding vector that contains emotional intensity and emotional categories (such as anger, happiness, sadness, and neutrality). Unlike traditional classification models, emotion2vec outputs a continuous emotional feature space, supporting fine-grained emotion analysis and exhibiting good transferability. The output emotion vector is not only used in quality inspection to identify key dimensions such as "service attitude issues" and "customer mood swings," but also serves as an important component of multimodal feature fusion, improving the emotional sensitivity and contextual adaptability of the overall analysis.
[0051] S205, performing multimodal feature fusion on the semantic feature vector, the acoustic feature vector, and the emotion feature vector to obtain a multimodal feature vector;
[0052] Specifically, the system integrates the textual information, acoustic features, and emotional cues contained in speech content into a unified representation, improving the model's ability to identify complex speech quality issues. Fusion methods can employ feature-level fusion (early fusion), decision-level fusion (late fusion), or multimodal alignment fusion techniques, such as cross-modal Transformers, gated attention mechanisms, and fusion autoencoders, to achieve temporal alignment and semantic enhancement of vectors from different modalities. During feature fusion, the system considers the complementarity between the three modalities. For example, the semantic modality reflects "what was said," the acoustic modality reflects "how it was said," and the emotional modality reflects "the psychological state at the time of speaking." By fusing these three types of features, a more expressive multimodal feature vector is formed, preserving the independent strengths of each modality while enhancing information accuracy through mutual verification. For example, if the semantic detection indicates "complaint" content and the emotional feature indicates "anger," and the acoustic features indicate rising intonation and accelerated speech rate, the system can more confidently determine that the content is "strongly dissatisfied," enabling precise quality inspection.
[0053] S206: Input the multimodal feature vector into a pre-trained speech quality inspection model, and output a speech quality inspection result of the speech data to be processed.
[0054] Specifically, the system inputs the multimodal feature vector constructed in S205 into the pre-trained voice quality inspection model for comprehensive analysis. The quality inspection model can adopt a multi-layer neural network architecture, such as a multimodal fusion Transformer, a dual-channel BiGRU, or a deep classification network based on an attention mechanism. The model can set multiple categories of outputs according to the quality inspection task objectives, such as "standardization of service terms", "whether interruptions occur", "customer satisfaction", "whether there is a tendency to complain", etc. During the training phase, the model performs end-to-end optimization in the form of multi-label supervision to improve its judgment ability in complex voice interaction scenarios. During the inference phase, the model quickly judges the multimodal input, outputs structured quality inspection labels and scoring results, and supports joint judgment of multiple indicators and risk warnings. At the same time, the model also supports adaptive updates and can be retrained by regularly introducing new corpus to adapt to business changes.
[0055] For example, in a certain health care system, a patient called to inquire about drug dosage. During the communication, the patient spoke anxiously and repeatedly emphasized, "This medicine is too important and must be explained clearly." The operator responded quickly and did not explain the key dosage information in detail. He only said, "Just take it according to the instructions." There was also a brief silence (i.e., communication interruption).
[0056] To address this situation, the quality inspection model adopts a dual-channel BiGRU architecture and inputs semantic, acoustic and emotional features. During training, multiple categories of outputs are set, such as "standardization of service terms", "whether interruptions occur", "customer satisfaction", and "whether there is a tendency to complain", and end-to-end optimization is carried out in the form of multi-label supervision. During the inference stage, the model receives the communication voice and quickly determines: the standardization of service terms is "poor" (key information is not explained in detail), there is a brief interruption, customer satisfaction is "low" (the patient's tone is anxious and needs are not met), and there is a "high" tendency to complain. Output structured quality inspection labels and scores, support joint judgment of multiple indicators, and trigger risk warnings due to failure to meet multiple indicators.
[0057] In the above embodiments, this application significantly improves the accuracy and intelligence level of speech quality inspection by introducing multimodal feature fusion technology. On the basis of traditional speech-to-text conversion, it further extracts the contextual semantic features of the text and the acoustic features of the speech, and combines the pre-trained emotion recognition model to obtain emotional features, thereby achieving a comprehensive understanding of speech data from the three dimensions of semantics, acoustics and emotions. By deeply fusing the three types of feature vectors, a more expressive multimodal feature vector is obtained, which can more fully explore the quality problems in the speech, including abnormal behaviors in service attitude, standardization of speech, emotional fluctuations, etc.
[0058] Compared with traditional methods that rely only on text or acoustic features, this application combines emotional features on the basis of text or acoustic features to achieve multimodal features, which can more comprehensively and meticulously reflect the information contained in the voice content, thereby improving the accuracy and practicality of voice quality inspection, and meeting the needs of high-quality automatic quality inspection in application scenarios such as intelligent customer service and voice auditing in the fields of finance, health care, etc.
[0059] Furthermore, the steps of performing contextual semantic recognition on the text to be processed and performing feature embedding processing on the text semantic recognition result to obtain a semantic feature vector specifically include:
[0060] Perform word segmentation on the text to be processed to obtain text segmentation;
[0061] Use the pre-trained language model to learn the contextual semantic information of each text segmentation in the text to be processed;
[0062] The preset bag-of-words model is used to perform feature embedding processing on the contextual semantic information to obtain a semantic feature vector.
[0063] In this embodiment, in the process of contextual semantic recognition of the text to be processed, the text is first segmented by a word segmentation tool, and word segmentation units with semantic boundaries are obtained according to the language type (such as Chinese uses a word segmentation algorithm based on dictionary + rules, and English can be directly segmented by space). Subsequently, a pre-trained language model (such as BERT, ERNIE or RoBERTa) is used to model the semantic representation of each word in the context, so that the model not only captures the meaning of the word itself, but also understands its semantic role in the sentence. This process can be implemented through a deep Transformer structure to output a context embedding vector for each word. On this basis, a bag-of-words model or a TF-IDF weighted model is introduced to aggregate, reduce dimensionality or pool multiple context embedding vectors to obtain a structured, fixed-length semantic feature vector. This vector can be used for multimodal fusion or directly participate in the input of the quality inspection model to ensure that the model can fully identify content dimensions such as customer demands, speech specifications, emotional tendencies, etc., and improve the accuracy and expression ability of semantic layer analysis.
[0064] Through the above steps, we can achieve a deep understanding and structured expression of text semantics, and enhance the voice quality inspection model's ability to understand service content.
[0065] Furthermore, the pre-trained emotion recognition model is an Emotion2Vec model, which includes an acoustic feature extractor and a sentiment analyzer. The steps of extracting acoustic features from the speech data to be processed and obtaining an acoustic feature vector specifically include:
[0066] Extracting the acoustic spectrum corresponding to the speech data to be processed;
[0067] Extracting speech acoustic features from the acoustic spectrum using an acoustic feature extractor;
[0068] Perform feature embedding processing on the speech acoustic features to obtain the acoustic feature vector.
[0069] In this embodiment, the Emotion2Vec model is used as a pre-trained emotion recognition model, and the acoustic feature extractor included in its structure can efficiently extract emotion-related features from speech signals. When performing acoustic feature extraction on speech data, the system first converts the continuous speech waveform signal into an acoustic spectrum in the time-frequency domain through short-time Fourier transform (STFT), forming a two-dimensional spectrogram that represents the frequency distribution characteristics over different time periods. Subsequently, the acoustic feature extractor analyzes the spectrogram based on a deep neural network (such as a CNN or Transformer structure) and automatically learns the underlying acoustic information such as rhythm, intonation, energy distribution, and resonance peaks in the speech, which are key parameters reflecting the emotional state. In this process, there is no need to manually set the feature dimensions, but the model autonomously captures discriminative features. Afterwards, the extracted multidimensional acoustic features are converted into structured acoustic feature vectors through a feature embedding layer (such as a fully connected layer or a projection transformation layer). Since the Emotion2Vec model is pre-trained on large-scale speech emotion data, its acoustic features have strong emotion discrimination capabilities and downstream task transfer capabilities.
[0070] Through the above steps, emotion-related acoustic features in speech can be extracted efficiently and accurately, providing a stable and reliable feature basis for emotion recognition and speech quality inspection.
[0071] Furthermore, based on the acoustic feature vector, emotion recognition is performed using a pre-trained emotion recognition model, and feature embedding processing is performed on the emotion recognition result to obtain the emotion feature vector, which specifically includes the following steps:
[0072] Performing feature mask processing on the acoustic feature vector to obtain a feature mask vector;
[0073] Map the feature mask vector to the feature space preset by the sentiment analyzer;
[0074] In the feature space preset by the sentiment analyzer, the feature mask vector is matched with the emotion pattern built into the sentiment analyzer to obtain the emotion recognition result;
[0075] Perform feature embedding processing on the emotion recognition results to obtain the emotion feature vector.
[0076] In this embodiment, the emotion recognition process utilizes a sentiment analyzer based on the Emotion2Vec model. The input acoustic feature vector is first subjected to feature masking. The purpose of feature masking is to enhance the model's focus on key information while suppressing any noise or non-emotion-related redundant features that may be present in the speech. This step uses a masking method similar to the self-attention mechanism in the Transformer to assign different weights to different dimensions in the feature vector, resulting in a feature mask vector representing key emotional regions. Next, this vector is mapped to a high-dimensional emotion feature space preset by the sentiment analyzer. This feature space is trained using large amounts of emotion-labeled speech data and is capable of clearly distinguishing a variety of basic emotions (such as anger, happiness, calmness, anxiety, etc.). Within this space, the system performs similarity calculations or vector matching on the input vector with preset emotion category prototypes (i.e., emotion pattern vectors), and outputs the corresponding emotion recognition results through maximum matching or multi-class regression. Finally, the system further encodes the recognition results into a unified format emotion feature vector through an emotion embedding layer, making it easier to integrate with other modal features or input into a quality inspection model. This approach not only improves recognition accuracy but also offers good robustness and scalability.
[0077] Through the above steps, the emotional state of speech can be efficiently identified and structured emotional features can be generated, thereby improving the emotion perception ability and intelligence level of the speech quality inspection system.
[0078] Furthermore, in the feature space preset by the sentiment analyzer, the feature mask vector is matched with the sentiment pattern built into the sentiment analyzer to obtain the emotion recognition result, which specifically includes:
[0079] Perform feature embedding processing on the emotional pattern to obtain the emotional pattern vector;
[0080] In the feature space preset by the sentiment analyzer, the cosine similarity between the feature mask vector and the sentiment pattern vector is calculated;
[0081] According to the cosine similarity, the feature mask vector and the emotion pattern are matched to obtain the emotion recognition result.
[0082] In this embodiment, emotion pattern matching during emotion recognition relies on the preset emotion feature space in the emotion analyzer. First, the various emotion patterns built into the system (e.g., anger, joy, sadness, calmness, etc.) are converted into vector form, that is, the emotion patterns are subjected to feature embedding processing to obtain emotion pattern vectors. This embedding process can be based on statistical modeling or deep learning, mapping the labeled emotion categories into discriminative vector representations in a multidimensional feature space. After obtaining the emotion pattern vectors, the system uses the cosine similarity method to calculate the similarity between the input feature mask vector and each emotion pattern vector. Cosine similarity measures the directional similarity between two vectors and is suitable for discerning the relative positional relationship of emotional states. The system matches based on the maximum similarity and classifies the feature mask vector into the most similar emotion category. This method avoids misjudgments caused by amplitude differences and focuses on the consistency of feature expression direction, thereby improving the stability and accuracy of emotion recognition. In addition, this matching process has good compatibility and scalability for newly added or refined emotion categories, and can support multi-granularity and hierarchical emotion recognition needs.
[0083] Through the above steps, accurate classification of the emotional state implicit in speech can be achieved.
[0084] Furthermore, the step of performing multimodal feature fusion on the semantic feature vector, the acoustic feature vector, and the emotional feature vector to obtain a multimodal feature vector specifically includes:
[0085] Perform feature vector dimension alignment on semantic feature vectors, acoustic feature vectors, and emotional feature vectors respectively;
[0086] Calculate the attention weights for semantic feature vectors, acoustic feature vectors, and emotional feature vectors respectively;
[0087] According to the attention weight, the dimensionally aligned semantic feature vector, acoustic feature vector and emotional feature vector are weighted and summed to obtain a multimodal feature vector.
[0088] In this embodiment, in order to fully integrate the information from the three modalities of semantics, acoustics and emotions, the system first performs dimensional alignment on the three types of feature vectors. Due to the different sources of the original features, the vectors differ in dimensionality and information density. Therefore, they are uniformly mapped to the same dimensional space through linear transformation (such as a fully connected layer) or projection operation, so that the fusion operation is comparable and computationally consistent. After the alignment is completed, the system introduces an attention mechanism to evaluate the importance of each modal feature in the current speech sample. The calculation of the attention weight can be based on the similarity between each feature vector and the fusion context, or the self-attention module can be used to automatically learn its information contribution. For example, in a customer conversation expressing strong emotions, the attention weight of the emotional feature may be higher than that of the semantic content. Finally, according to the calculated attention weight, the aligned vectors of the three modalities are weighted and summed to generate a multimodal feature vector with the ability to express the fusion of contextual semantics, acoustic characteristics and emotional state.
[0089] Through the above steps, the complementary fusion of modal information can be achieved, significantly improving the recognition accuracy and robustness of the speech quality inspection model for abnormal speech, illegal language, and emotional outbursts.
[0090] Furthermore, the speech quality inspection model is a multi-classification model, and the steps of inputting the multimodal feature vector into the pre-trained speech quality inspection model and outputting the speech quality inspection result of the speech data to be processed specifically include:
[0091] Get the original confidence level corresponding to each quality inspection label built into the speech quality inspection model;
[0092] According to the preset category index, the multimodal feature vector is mapped to each quality inspection label;
[0093] The original confidence level is updated using the normalization function built into the speech quality inspection model, and the probability distribution of each quality inspection label is generated based on the updated original confidence level.
[0094] The speech quality inspection result of the speech data to be processed is determined based on the probability distribution of each quality inspection label.
[0095] In this embodiment, the speech quality inspection model is a multi-classification structure that can comprehensively distinguish the input multimodal feature vector and output the corresponding quality inspection label. First, the system obtains the original confidence associated with each preset quality inspection label (such as "speaking too fast", "emotional loss of control", "sensitive word violation", "poor service attitude", etc.) through the category discrimination unit inside the model. This confidence reflects the strength of the model's preliminary judgment on whether the current sample belongs to each category of labels. Subsequently, the system maps the fused multimodal feature vector to the vector representation space corresponding to each label according to the category index defined in the training phase. In order to ensure the interpretability and stability of the output, the model further applies a normalization function to the original confidence value. Common methods include the Softmax function or the normalization function after temperature adjustment to balance the differences between the labels and avoid the phenomenon that the probability of a certain type of label is extremely high or low. Finally, based on the updated label probability distribution, the system selects the label corresponding to the maximum probability as the final quality inspection result of the speech data. It can also output the Top-N probability result according to the task requirements to support multi-label recognition. During the deployment process, the model can process large amounts of voice data in real time and quickly output classification judgments, adapting to the efficient quality inspection needs of scenarios such as corporate call centers and customer service supervision.
[0096] Raw confidence, also known as classification logits (or unnormalized classification scores), is the raw numerical value output by a neural network in classification tasks. It typically represents the model's "raw confidence" for each category. Raw confidence is the output of the model's last layer (usually a fully connected layer) and has not yet undergone normalization processing such as Softmax or Sigmoid.
[0097] Logits can be converted into probability distribution through the Softmax function:
[0098]
[0099] Among them, z i is the Logit value of class i. Calculating the loss directly on Logits (such as cross entropy loss) is more stable than calculating the probability, and avoids the numerical overflow problem of Softmax (especially when the Logits value is large).
[0100] The formula for cross entropy loss is:
[0101]
[0102] Through the above steps, it is possible to accurately classify various voice quality issues and enhance the intelligence level and service coverage capabilities of the voice quality inspection system.
[0103] In the above embodiments, the present application discloses a speech quality inspection method, which belongs to the field of artificial intelligence technology and is applied to speech quality inspection scenarios in the fields of finance, health care, etc. This application significantly improves the accuracy and intelligence level of speech quality inspection by introducing multimodal feature fusion technology. Based on the traditional speech-to-text conversion, it further extracts the contextual semantic features of the text and the acoustic features of the speech, and combines it with a pre-trained emotion recognition model to obtain emotional features, achieving a comprehensive understanding of the speech data from the three dimensions of semantics, acoustics, and emotion. By deeply fusing the three types of feature vectors, a more expressive multimodal feature vector is obtained, which can more fully explore quality issues in the speech, including abnormal behaviors such as service attitude, standardization of speech, and emotional fluctuations. Compared with traditional methods that rely solely on text or acoustic features, this application combines text or acoustic features with emotional features to achieve multimodal features. It can more comprehensively and meticulously reflect the information contained in the speech content, thereby improving the accuracy and practicality of speech quality inspection, and meeting the needs of high-quality automatic quality inspection in application scenarios such as intelligent customer service and speech auditing in the fields of finance, health care, etc.
[0104] In this embodiment, the electronic device on which the voice quality inspection method is running (eg Figure 1 The server shown in the figure) can receive instructions or obtain data through a wired connection or a wireless connection. It should be noted that the above-mentioned wireless connection method may include but is not limited to 3G / 4G connection, WiFi connection, Bluetooth connection, WiMAX connection, Zigbee connection, UWB (ultra wideband) connection, and other wireless connection methods currently known or to be developed in the future.
[0105] It should be emphasized that in order to further ensure the privacy and security of the above-mentioned voice data information, the above-mentioned voice data information can also be stored in a node of a blockchain.
[0106] The blockchain referred to in this application is a new application model for computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanisms, and encryption algorithms. Blockchain is essentially a decentralized database, a series of data blocks generated using cryptographic methods. Each data block contains information about a batch of network transactions, which is used to verify the validity of this information (to prevent counterfeiting) and generate the next block. Blockchain can include the underlying blockchain platform, the platform product service layer, and the application service layer.
[0107] The embodiments of the present application can acquire and process relevant data based on artificial intelligence technology. Artificial Intelligence (AI) is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to achieve optimal results.
[0108] Fundamental AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interaction systems, and mechatronics. AI software technologies primarily encompass computer vision, robotics, biometrics, speech processing, natural language processing, and machine learning / deep learning.
[0109] Those skilled in the art will appreciate that all or part of the processes in the above-described method embodiments can be implemented by instructing related hardware via computer-readable instructions. The computer-readable instructions can be stored in a computer-readable storage medium, and when the program is executed, it can include the processes in the above-described method embodiments. The aforementioned storage medium can be a non-volatile storage medium such as a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM).
[0110] It should be understood that although the steps in the flowcharts of the accompanying drawings are shown in sequence as indicated by the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some of the steps in the flowcharts of the accompanying drawings may include multiple sub-steps or multiple stages, and these sub-steps or stages are not necessarily executed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be executed in turn or alternately with other steps or at least a portion of the sub-steps or stages of other steps.
[0111] Further references Figure 3 , as a response to the above Figure 2 The present application provides an embodiment of a speech quality inspection device, which is similar to the embodiment of the present invention. Figure 2 Corresponding to the method embodiment shown, the device can be specifically applied to various electronic devices.
[0112] like Figure 3 As shown, the speech quality inspection device 300 described in this embodiment includes:
[0113] The text conversion module 301 is used to obtain the speech data to be processed and perform speech-to-text processing on the speech data to be processed to obtain the text to be processed;
[0114] Semantic recognition module 302, used to perform contextual semantic recognition on the text to be processed, and perform feature embedding processing on the text semantic recognition result to obtain a semantic feature vector;
[0115] The acoustic feature module 303 is used to extract acoustic features from the speech data to be processed and obtain an acoustic feature vector;
[0116] The emotion recognition module 304 is configured to perform emotion recognition based on the acoustic feature vector using a pre-trained emotion recognition model, and perform feature embedding processing on the emotion recognition result to obtain an emotion feature vector;
[0117] A multimodal fusion module 305 is configured to perform multimodal feature fusion on the semantic feature vector, the acoustic feature vector, and the emotional feature vector to obtain a multimodal feature vector;
[0118] The speech quality inspection module 306 is used to input the multimodal feature vector into the pre-trained speech quality inspection model and output the speech quality inspection result of the speech data to be processed.
[0119] Furthermore, the semantic recognition module 302 specifically includes:
[0120] A word segmentation processing unit is used to perform word segmentation on the text to be processed to obtain text segmentation;
[0121] The semantic recognition unit is used to learn the contextual semantic information of each text segment in the text to be processed using a pre-trained language model;
[0122] The semantic embedding unit is used to perform feature embedding processing on contextual semantic information using a preset bag-of-words model to obtain a semantic feature vector.
[0123] Furthermore, the pre-trained emotion recognition model is an Emotion2Vec model, which includes an acoustic feature extractor and a sentiment analyzer. The acoustic feature module 303 specifically includes:
[0124] A spectrum extraction unit, used to extract the acoustic spectrum corresponding to the speech data to be processed;
[0125] an acoustic feature unit, configured to extract speech acoustic features from the acoustic spectrum using an acoustic feature extractor;
[0126] The acoustic feature embedding unit is used to perform feature embedding processing on the speech acoustic features to obtain the acoustic feature vector.
[0127] Furthermore, the emotion recognition module 304 specifically includes:
[0128] A feature mask unit, used for performing feature mask processing on the acoustic feature vector to obtain a feature mask vector;
[0129] A first vector mapping unit, configured to map the feature mask vector to a feature space preset by the sentiment analyzer;
[0130] The pattern matching unit is used to match the feature mask vector with the emotion pattern built into the emotion analyzer within the feature space preset by the emotion analyzer to obtain the emotion recognition result;
[0131] The emotion feature embedding unit is used to perform feature embedding processing on the emotion recognition results to obtain the emotion feature vector.
[0132] Furthermore, the pattern matching unit specifically includes:
[0133] The pattern feature embedding subunit is used to perform feature embedding processing on the emotion pattern to obtain the emotion pattern vector;
[0134] A cosine similarity calculation subunit is used to calculate the cosine similarity between the feature mask vector and the emotion pattern vector in the feature space preset by the emotion analyzer;
[0135] The pattern matching subunit is used to match the feature mask vector and the emotion pattern according to the cosine similarity to obtain the emotion recognition result.
[0136] Furthermore, the multimodal fusion module 305 specifically includes:
[0137] A dimension alignment unit, used to align the feature vector dimensions of the semantic feature vector, the acoustic feature vector, and the emotional feature vector respectively;
[0138] An attention weight unit, used to calculate the attention weights of the semantic feature vector, acoustic feature vector, and emotional feature vector respectively;
[0139] The weighted summation unit is used to perform weighted summation on the dimensionally aligned semantic feature vector, acoustic feature vector, and emotional feature vector according to the attention weight to obtain a multimodal feature vector.
[0140] Furthermore, the speech quality inspection model is a multi-classification model, and the speech quality inspection module 306 specifically includes:
[0141] A confidence acquisition unit is used to obtain the original confidence corresponding to each quality inspection label built into the speech quality inspection model;
[0142] A second vector mapping unit, configured to map the multimodal feature vectors to respective quality inspection labels according to a preset category index;
[0143] A confidence update unit, which is used to update the original confidence using the normalization function built into the speech quality inspection model and generate a probability distribution for each quality inspection label based on the updated original confidence;
[0144] The speech quality inspection unit is used to determine the speech quality inspection result of the speech data to be processed based on the probability distribution of each quality inspection label.
[0145] In the above embodiments, the present application discloses a speech quality inspection device, which belongs to the field of artificial intelligence technology and is applied to speech quality inspection scenarios in the fields of finance, health care, etc. This application significantly improves the accuracy and intelligence level of speech quality inspection by introducing multimodal feature fusion technology. On the basis of traditional speech-to-text conversion, it further extracts the contextual semantic features of the text and the acoustic features of the speech, and combines it with a pre-trained emotion recognition model to obtain emotional features, achieving a comprehensive understanding of speech data from the three dimensions of semantics, acoustics, and emotion. By deeply fusing the three types of feature vectors, a more expressive multimodal feature vector is obtained, which can more fully explore quality issues in speech, including abnormal behaviors in terms of service attitude, standardization of speech, and emotional fluctuations. Compared with traditional methods that rely solely on text or acoustic features, this application combines text or acoustic features with emotional features to achieve multimodal features, which can more comprehensively and meticulously reflect the information contained in the speech content, thereby improving the accuracy and practicality of speech quality inspection, and meeting the needs of high-quality automatic quality inspection in application scenarios such as intelligent customer service and voice auditing in the fields of finance, health care, etc.
[0146] To solve the above technical problems, the present application also provides a computer device. Figure 4 , Figure 4 This is a basic structural block diagram of the computer device in this embodiment.
[0147] The computer device 4 includes a memory 41, a processor 42, and a network interface 43 that are interconnected through a system bus. It should be noted that the figure only shows a computer device 4 with a memory 41, a processor 42, and a network interface 43, but it should be understood that it is not required to implement all the components shown, and more or fewer components can be implemented instead. Among them, those skilled in the art can understand that the computer device here is a device that can automatically perform numerical calculations and / or information processing according to pre-set or stored instructions, and its hardware includes but is not limited to microprocessors, application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), embedded devices, etc.
[0148] The computer device may be a desktop computer, notebook computer, PDA, cloud server, etc. The computer device may interact with the user via a keyboard, mouse, remote control, touchpad, or voice control device.
[0149] The memory 41 includes at least one type of readable storage medium, including flash memory, hard disk, multimedia card, card-type memory (e.g., SD or DX memory), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the memory 41 can be an internal storage unit of the computer device 4, such as the hard disk or memory of the computer device 4. In other embodiments, the memory 41 can also be an external storage device of the computer device 4, such as a plug-in hard disk equipped on the computer device 4, a smart memory card (SMC), a secure digital (SD) card, a flash memory card, etc. Of course, the memory 41 can also include both the internal storage unit of the computer device 4 and its external storage device. In this embodiment, the memory 41 is generally used to store the operating system and various application software installed on the computer device 4, such as computer-readable instructions of the voice quality inspection method. In addition, the memory 41 can also be used to temporarily store various types of data that have been output or are to be output.
[0150] In some embodiments, the processor 42 may be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chip. The processor 42 is generally used to control the overall operation of the computer device 4. In this embodiment, the processor 42 is used to execute computer-readable instructions stored in the memory 41 or process data, such as computer-readable instructions for executing the speech quality inspection method.
[0151] The network interface 43 may include a wireless network interface or a wired network interface. The network interface 43 is generally used to establish a communication connection between the computer device 4 and other electronic devices.
[0152] The present application also provides an embodiment, namely, a computer device, comprising a memory and a processor, wherein the memory stores computer-readable instructions, and when the processor executes the computer-readable instructions, the steps of the above-mentioned voice quality inspection method are implemented, namely, the following steps are implemented:
[0153] A voice quality inspection method, comprising:
[0154] Acquire the voice data to be processed, and perform speech-to-text processing on the voice data to be processed to obtain the text to be processed;
[0155] Perform contextual semantic recognition on the text to be processed, and perform feature embedding processing on the text semantic recognition results to obtain a semantic feature vector;
[0156] Extract acoustic features from the speech data to be processed to obtain acoustic feature vectors;
[0157] Based on the acoustic feature vector, emotion recognition is performed using a pre-trained emotion recognition model, and the emotion recognition results are subjected to feature embedding processing to obtain the emotion feature vector;
[0158] Perform multimodal feature fusion on the semantic feature vector, acoustic feature vector and emotional feature vector to obtain a multimodal feature vector;
[0159] The multimodal feature vector is input into the pre-trained speech quality inspection model, and the speech quality inspection result of the speech data to be processed is output.
[0160] The present application also provides another embodiment, namely, providing a computer-readable storage medium, wherein the computer-readable storage medium stores computer-readable instructions, and the computer-readable instructions can be executed by at least one processor to cause the at least one processor to perform the steps of the above-mentioned voice quality inspection method, namely, to achieve:
[0161] A voice quality inspection method, comprising:
[0162] Acquire the voice data to be processed, and perform speech-to-text processing on the voice data to be processed to obtain the text to be processed;
[0163] Perform contextual semantic recognition on the text to be processed, and perform feature embedding processing on the text semantic recognition results to obtain a semantic feature vector;
[0164] Extract acoustic features from the speech data to be processed to obtain acoustic feature vectors;
[0165] Based on the acoustic feature vector, emotion recognition is performed using a pre-trained emotion recognition model, and the emotion recognition results are subjected to feature embedding processing to obtain the emotion feature vector;
[0166] Perform multimodal feature fusion on the semantic feature vector, acoustic feature vector and emotional feature vector to obtain a multimodal feature vector;
[0167] The multimodal feature vector is input into the pre-trained speech quality inspection model, and the speech quality inspection result of the speech data to be processed is output.
[0168] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus the necessary general hardware platform, and of course can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for enabling a terminal device (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in each embodiment of the present application.
[0169] The present application can be used in many general or special computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronic devices, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, etc. The present application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present application can also be practiced in distributed computing environments, in which tasks are performed by remote processing devices connected via a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media, including storage devices.
[0170] It should be noted that the non-Company's software tools or components appearing in the various embodiments of this application are merely examples and do not represent actual use.
[0171] Obviously, the embodiments described above are only some of the embodiments of the present application, rather than all of the embodiments. The preferred embodiments of the present application are given in the accompanying drawings, but they do not limit the patent scope of the present application. The present application can be implemented in many different forms. On the contrary, the purpose of providing these embodiments is to make the understanding of the disclosure of the present application more thorough and comprehensive. Although the present application has been described in detail with reference to the aforementioned embodiments, for those skilled in the art, it is still possible to modify the technical solutions described in the aforementioned specific embodiments, or to make equivalent replacements for some of the technical features therein. Any equivalent structure made using the contents of the present application specification and the accompanying drawings, directly or indirectly used in other related technical fields, is also within the scope of patent protection of the present application.
Claims
1. A voice quality inspection method, characterized in that: include: Acquire speech data to be processed, and perform speech-to-text processing on the speech data to be processed to obtain text to be processed; Performing contextual semantic recognition on the text to be processed, and performing feature embedding processing on the text semantic recognition result to obtain a semantic feature vector; Extracting acoustic features from the speech data to be processed to obtain an acoustic feature vector; Based on the acoustic feature vector, using a pre-trained emotion recognition model to perform emotion recognition, and performing feature embedding processing on the emotion recognition result to obtain an emotion feature vector; Performing multimodal feature fusion on the semantic feature vector, the acoustic feature vector, and the emotion feature vector to obtain a multimodal feature vector; The multimodal feature vector is input into a pre-trained speech quality inspection model, and a speech quality inspection result of the speech data to be processed is output.
2. The speech quality inspection method according to claim 1, wherein: The step of performing contextual semantic recognition on the text to be processed and performing feature embedding processing on the text semantic recognition result to obtain a semantic feature vector specifically includes: Performing word segmentation processing on the text to be processed to obtain text segmentation; Using a pre-trained language model to learn contextual semantic information of each of the text segmentations in the text to be processed; The contextual semantic information is subjected to feature embedding processing using a preset bag-of-words model to obtain the semantic feature vector.
3. The speech quality inspection method according to claim 1, wherein: The pre-trained emotion recognition model is an Emot ion2Vec model, which includes an acoustic feature extractor and a sentiment analyzer. The step of extracting acoustic features from the speech data to be processed to obtain an acoustic feature vector specifically includes: Extracting the acoustic spectrum corresponding to the speech data to be processed; extracting speech acoustic features from the acoustic spectrum using the acoustic feature extractor; Perform feature embedding processing on the speech acoustic features to obtain the acoustic feature vector.
4. The speech quality inspection method according to claim 3, wherein: The step of performing emotion recognition based on the acoustic feature vector using a pre-trained emotion recognition model and performing feature embedding processing on the emotion recognition result to obtain the emotion feature vector specifically includes: Performing feature mask processing on the acoustic feature vector to obtain a feature mask vector; Mapping the feature mask vector to a feature space preset by the sentiment analyzer; Matching the feature mask vector with the emotion pattern built into the emotion analyzer in a feature space preset by the emotion analyzer to obtain the emotion recognition result; Perform feature embedding processing on the emotion recognition results to obtain the emotion feature vector.
5. The speech quality inspection method according to claim 4, wherein: The step of matching the feature mask vector with the emotion pattern built into the emotion analyzer within the feature space preset by the emotion analyzer to obtain the emotion recognition result specifically includes: Performing feature embedding processing on the emotion pattern to obtain an emotion pattern vector; Calculating the cosine similarity between the feature mask vector and the emotion pattern vector in a feature space preset by the emotion analyzer; The feature mask vector and the emotion pattern are matched according to the cosine similarity to obtain the emotion recognition result.
6. The speech quality inspection method according to claim 1, wherein: The step of performing multimodal feature fusion on the semantic feature vector, the acoustic feature vector, and the emotion feature vector to obtain a multimodal feature vector specifically includes: performing feature vector dimension alignment on the semantic feature vector, the acoustic feature vector, and the emotion feature vector respectively; Calculating attention weights for the semantic feature vector, the acoustic feature vector, and the emotional feature vector respectively; According to the attention weight, a weighted sum is performed on the dimensionally aligned semantic feature vector, the acoustic feature vector, and the emotional feature vector to obtain the multimodal feature vector.
7. The speech quality inspection method according to claim 1, wherein: The speech quality inspection model is a multi-classification model. The step of inputting the multimodal feature vector into the pre-trained speech quality inspection model and outputting the speech quality inspection result of the speech data to be processed specifically includes: Obtaining the original confidence corresponding to each quality inspection label built into the speech quality inspection model; Mapping the multimodal feature vector to each of the quality inspection labels according to a preset category index; Updating the original confidence level using a normalization function built into the speech quality inspection model, and generating a probability distribution of each quality inspection label based on the updated original confidence level; The speech quality inspection result of the speech data to be processed is determined based on the probability distribution of each of the quality inspection labels.
8. A speech quality inspection device, characterized in that: include: A text conversion module is used to obtain the voice data to be processed and perform voice-to-text processing on the voice data to be processed to obtain the text to be processed; A semantic recognition module is used to perform contextual semantic recognition on the text to be processed, and perform feature embedding processing on the text semantic recognition result to obtain a semantic feature vector; An acoustic feature module, configured to extract acoustic features from the speech data to be processed to obtain an acoustic feature vector; An emotion recognition module is configured to perform emotion recognition based on the acoustic feature vector using a pre-trained emotion recognition model, and perform feature embedding processing on the emotion recognition result to obtain an emotion feature vector; a multimodal fusion module, configured to perform multimodal feature fusion on the semantic feature vector, the acoustic feature vector, and the emotional feature vector to obtain a multimodal feature vector; The speech quality inspection module is used to input the multimodal feature vector into a pre-trained speech quality inspection model and output a speech quality inspection result of the speech data to be processed.
9. A computer device, characterized in that: The method comprises a memory and a processor, wherein the memory stores computer-readable instructions, and the processor implements the steps of the speech quality inspection method according to any one of claims 1 to 7 when executing the computer-readable instructions.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer-readable instructions, which, when executed by a processor, implement the steps of the speech quality inspection method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Customer service voice quality inspection method, device and equipment and storage medium
CN118035411A
User emotion recognition method based on AI and voice data
CN120148561A
Speech-based user classification method and device, computer apparatus, and storage medium
WO2020244073A1
Cited By
Intelligent quality inspection system and method based on voice transcription and emotion recognition
CN121284161A
Voice quality evaluation system and method
CN121747615A
A speech quality assessment system and method
CN121747615B
Deep learning-based Chinese language audio phoneme segmentation method and device, and medium
CN122157642A