A voice quality inspection method, apparatus, computer equipment, and storage medium

CN120690231BActive Publication Date: 2026-09-01PING AN TECH (BEIJING) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510865267.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-25
Publication Date
2026-09-01
Estimated Expiration
2045-06-25

AI Technical Summary

Technical Problem

[0004]本申请实施例的目的在于提出一种语音质检方法、装置、计算机设备及存储介质,以解决现有语音质检存在的不能实现后续的类别拓展,可延展性较差,同时模型在推理时的计算会有延迟现象,无法进行实时的监控的技术问题

Benefits of technology

[0026]本申请公开一种语音质检方法、装置、计算机设备及存储介质,属于人工智能技术领域,应用于金融、健康医疗等领域的语音质检场景。本申请通过引入多模态特征融合技术,显著提升了语音质检的准确性和智能化水平,在传统语音转文本基础上,进一步提取文本的上下文语义特征和语音的声学特征,并结合预训练情绪识别模型获取情绪特征,实现从语义、声学和情绪三个维度对语音数据的全面理解,通过对三类特征向量进行深度融合,获得更具表达力的多模态特征向量,能够更充分地挖掘语音中的质量问题,包括服务态度、话术规范性、情绪波动等方面的异常行为。相较于传统仅依赖文本或声学特征的方法,本申请在文本或声学特征的基础上结合情绪特征实现多模态特征,能够更全面、细致地反映语音内容所蕴含的信息,从而提升语音质检的准确率与实用性,满足金融、健康医疗等领域的智能客服、语音审计等应用场景对高质量自动质检的需求。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120690231B_ABST
    Figure CN120690231B_ABST
Patent Text Reader

Abstract

This application discloses a voice quality inspection method, apparatus, computer equipment, and storage medium, belonging to the field of artificial intelligence technology, and applied to voice quality inspection scenarios in finance, healthcare, and elderly care. This application introduces multimodal feature fusion technology, extracting contextual semantic features from text and acoustic features from speech, and combining this with a pre-trained emotion recognition model to obtain emotional features. This achieves a comprehensive understanding of speech data from three dimensions: semantics, acoustics, and emotion. By deeply fusing these three types of feature vectors, it can more fully uncover quality issues in speech. Compared to traditional methods that rely solely on text or acoustic features, this application combines emotional features with text or acoustic features to achieve multimodal features, which can more comprehensively and meticulously reflect the information contained in the speech content, thereby improving the accuracy and practicality of voice quality inspection and meeting the needs of high-quality automated quality inspection in application scenarios such as intelligent customer service and voice auditing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of artificial intelligence technology, specifically relating to a voice quality inspection method, device, computer equipment, and storage medium. Background Technology

[0002] In quality inspection or voice service operations, call information is typically recorded to facilitate subsequent performance evaluation and service level improvement. For call evaluation systems, the emotions of both the customer and the sales staff are crucial aspects of measuring service quality, and the volume and tone of voice often reflect a person's emotions. For example, in customer service management in financial services or healthcare and elderly care, during voice quality inspection scenarios involving conversations between financial sales representatives and customers, if the customer service representative's tone is impatient or their volume is too high, it may cause customer dissatisfaction, affecting business progress and customer experience; conversely, if the customer is emotionally agitated and their tone is shrill, failure to promptly soothe them will significantly diminish the service effectiveness. Therefore, accurately extracting the speaker's emotional information from voice recordings is a problem that needs to be addressed.

[0003] To address the aforementioned issues, current solutions and technologies focus on using deep learning networks to extract features before classification. However, this approach cannot support subsequent category expansion, has poor scalability, and suffers from computational delays during inference, making real-time monitoring impossible. Summary of the Invention

[0004] The purpose of this application is to propose a voice quality inspection method, apparatus, computer equipment, and storage medium to solve the technical problems of existing voice quality inspection methods, such as the inability to achieve subsequent category expansion, poor scalability, and the inability to perform real-time monitoring due to computational delays during model inference.

[0005] To address the aforementioned technical problems, this application provides a voice quality inspection method, employing the following technical solution:

[0006] A voice quality inspection method, comprising:

[0007] Acquire the speech data to be processed, and perform speech-to-text processing on the speech data to obtain the text to be processed;

[0008] Contextual semantic recognition is performed on the text to be processed, and the semantic recognition results are then processed by feature embedding to obtain semantic feature vectors;

[0009] Acoustic feature extraction is performed on the speech data to be processed to obtain acoustic feature vectors;

[0010] Based on acoustic feature vectors, a pre-trained emotion recognition model is used to perform emotion recognition, and the emotion recognition results are processed by feature embedding to obtain emotion feature vectors.

[0011] Multimodal feature fusion is performed on semantic feature vectors, acoustic feature vectors, and emotion feature vectors to obtain multimodal feature vectors;

[0012] The multimodal feature vectors are input into a pre-trained speech quality inspection model, which outputs the speech quality inspection results of the speech data to be processed.

[0013] To address the aforementioned technical problems, this application also provides a voice quality inspection device, which employs the following technical solution:

[0014] A voice quality inspection device, comprising:

[0015] The text conversion module is used to acquire the speech data to be processed and to perform speech-to-text processing on the speech data to obtain the text to be processed.

[0016] The semantic recognition module is used to perform contextual semantic recognition on the text to be processed, and to perform feature embedding processing on the text semantic recognition results to obtain semantic feature vectors;

[0017] The acoustic feature module is used to extract acoustic features from the speech data to be processed, and obtain acoustic feature vectors.

[0018] The emotion recognition module is used to perform emotion recognition based on acoustic feature vectors using a pre-trained emotion recognition model, and to perform feature embedding processing on the emotion recognition results to obtain emotion feature vectors.

[0019] The multimodal fusion module is used to perform multimodal feature fusion on semantic feature vectors, acoustic feature vectors, and emotion feature vectors to obtain multimodal feature vectors;

[0020] The speech quality inspection module is used to input multimodal feature vectors into a pre-trained speech quality inspection model and output the speech quality inspection results of the speech data to be processed.

[0021] To address the aforementioned technical problems, this application also provides a computer device that employs the following technical solution:

[0022] A computer device includes a memory and a processor, the memory storing computer-readable instructions, the processor executing the computer-readable instructions to implement the steps of the voice quality inspection method as described in any of the preceding claims.

[0023] To address the aforementioned technical problems, this application also provides a computer-readable storage medium, employing the technical solution described below:

[0024] A computer-readable storage medium storing computer-readable instructions, which, when executed by a processor, implement the steps of the voice quality inspection method as described in any one of the preceding descriptions.

[0025] Compared with the prior art, the embodiments of this application have the following main advantages:

[0026] This application discloses a voice quality inspection method, apparatus, computer equipment, and storage medium, belonging to the field of artificial intelligence technology, and applied to voice quality inspection scenarios in finance, healthcare, and other fields. This application significantly improves the accuracy and intelligence level of voice quality inspection by introducing multimodal feature fusion technology. Based on traditional speech-to-text conversion, it further extracts the contextual semantic features of the text and the acoustic features of the speech, and combines this with a pre-trained emotion recognition model to obtain emotional features, achieving a comprehensive understanding of speech data from three dimensions: semantics, acoustics, and emotion. By deeply fusing these three types of feature vectors, a more expressive multimodal feature vector is obtained, which can more fully uncover quality problems in speech, including abnormal behaviors such as service attitude, speech format, and emotional fluctuations. Compared to traditional methods that rely solely on text or acoustic features, this application combines emotional features with text or acoustic features to achieve multimodal features, which can more comprehensively and meticulously reflect the information contained in the speech content, thereby improving the accuracy and practicality of voice quality inspection and meeting the needs of high-quality automated quality inspection in application scenarios such as intelligent customer service and voice auditing in finance, healthcare, and other fields. Attached Figure Description

[0027] To more clearly illustrate the solutions in this application, the accompanying drawings used in the description of the embodiments of this application will be briefly introduced below. Obviously, the accompanying drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0028] Figure 1 An exemplary system architecture diagram is shown, in which this application can be applied;

[0029] Figure 2 A flowchart of one embodiment of the voice quality inspection method according to this application is shown;

[0030] Figure 3 A schematic diagram of the structure of one embodiment of the voice quality inspection device according to this application is shown;

[0031] Figure 4 A schematic diagram of the structure of one embodiment of a computer device according to this application is shown. Detailed Implementation

[0032] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains; the terminology used herein in the specification of the application is for the purpose of describing particular embodiments only and is not intended to be limiting of the application; the terms "comprising" and "having," and any variations thereof, in the specification, claims, and foregoing drawings of this application, are intended to cover non-exclusive inclusion. The terms "first," "second," etc., in the specification, claims, or foregoing drawings of this application are used to distinguish different objects, not to describe a particular order.

[0033] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.

[0034] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings.

[0035] like Figure 1 As shown, system architecture 100 may include terminal device 101, network 102, and server 103. Terminal device 101 may be a laptop 1011, tablet 1012, or mobile phone 1013. Network 102 is used as a medium to provide a communication link between terminal device 101 and server 103. Network 102 may include various connection types, such as wired, wireless communication links, or fiber optic cables.

[0036] Users can use terminal device 101 to interact with server 103 via network 102 to receive or send messages, etc. Various communication client applications can be installed on terminal device 101, such as web browser applications, shopping applications, search applications, instant messaging tools, email clients, social media platform software, etc.

[0037] Terminal device 101 can be various electronic devices with a display screen and support web browsing. In addition to laptops 1011, tablets 1012, or mobile phones 1013, terminal device 101 can also be e-book readers, MP3 players (Moving Picture Experts Group Audio Layer III), MP4 players (Moving Picture Experts Group Audio Layer IV), laptops, and desktop computers.

[0038] Server 103 can be a server that provides various services, such as a backend server that provides support for the pages displayed on terminal device 101.

[0039] It should be noted that the voice quality inspection method provided in this application embodiment is generally executed by a server / terminal device, and correspondingly, the voice quality inspection device is generally set in the server / terminal device.

[0040] It should be understood that Figure 1 The number of terminal devices, networks, and servers shown is merely illustrative; the system can have any number of terminal devices, networks, and servers depending on implementation needs.

[0041] Continue to refer to Figure 2 A flowchart of an embodiment of the voice quality inspection method according to this application is shown. The voice quality inspection method includes the following steps:

[0042] S201, acquire the voice data to be processed, and perform speech-to-text processing on the voice data to be processed to obtain the text to be processed;

[0043] Specifically, the system first acquires the voice data to be processed from a voice database or real-time audio stream. This data can originate from call center agent conversations, dialogues between voice robots and users, customer reports, or other service interaction audio. The audio format can be standard formats such as WAV, MP3, and PCM. To ensure the quality of speech-to-text transcription, the speech is preprocessed before recognition, including noise reduction, silence removal, speech rate standardization, and speech enhancement (such as echo cancellation and bandpass filtering). Then, an Automatic Speech Recognition (ASR) system transcribes the speech content into text. The ASR system can be based on an end-to-end deep learning model, such as an RNN-Transducer, CTC decoder, Transformer, or Conformer architecture. This speech recognition model can be optimized by combining it with vertical language models to better adapt to specific industry terminology (such as "claims liability," "policy number," and "claimant"). During the speech-to-text process, the system also marks pauses, emotional peaks, and other timestamp information.

[0044] S202, perform contextual semantic recognition on the text to be processed, and perform feature embedding processing on the text semantic recognition results to obtain semantic feature vectors;

[0045] Specifically, the system employs contextual semantic modeling technology to process the text obtained from S201. Utilizing pre-trained language models such as BERT, RoBERTa, and ERNIE, combined with syntactic analysis and named entity recognition, it performs deep semantic understanding of the text content. The model can identify service intent (such as reporting a crime, complaint, inquiry, or claim collection), key business information (such as policy number, compensation amount, and liability division), and logical relationships within the context (such as causality, parallelism, and contrast). Long texts are segmented to maintain semantic integrity. After processing, the recognition results undergo feature embedding, including encoding the text structure using Token Embedding, Segment Embedding, and Position Embedding to form a fixed-dimensional vector representation. This semantic feature vector can fully express the connection between user intent and service behavior and is highly robust to contextual changes, providing rich semantic evidence.

[0046] S203, perform acoustic feature extraction on the speech data to be processed to obtain acoustic feature vectors;

[0047] Specifically, by extracting quantifiable acoustic parameters from the speech signal, the speech is transformed from the audio domain to a feature space. Commonly used features include Mel-frequency cepstral coefficients (MFCC), chroma features, spectral centroid, frequency band energy distribution, pitch, speech rate, and short-time energy. These features can reveal the speaker's articulation patterns, intonation variations, speech rate changes, and emphasis patterns. For example, high pitch and energy fluctuations in speech may indicate emotional excitement; abnormal speech rate may suggest a state of stress. During the extraction process, the system first performs frame segmentation and windowing on the speech signal, then uses methods such as FFT and Mel filter banks for frequency domain transformation, and concatenates or reduces the dimensionality of acoustic features from different dimensions to generate a unified acoustic feature vector. To enhance the model's adaptability to complex acoustic backgrounds, self-supervised acoustic coding models, such as wav2vec2.0 and Hubert, can be introduced to directly model the original waveform to obtain a more robust high-level acoustic representation. The final acoustic feature vector is used not only for emotion recognition but also for identifying quality control items such as abnormal speech rate and inappropriate pauses.

[0048] S204, based on acoustic feature vectors, uses a pre-trained emotion recognition model to perform emotion recognition, and performs feature embedding processing on the emotion recognition results to obtain emotion feature vectors;

[0049] Specifically, the system employs a pre-trained emotion recognition model to further analyze acoustic features to identify emotional states, preferentially using the open-source Emotion2vec model. Emotion2vec is a deep learning model that encodes speech signals into high-dimensional emotion vectors, capable of learning joint features of emotion, semantics, and context in speech. The Emotion2vec model uses a self-supervised learning method, pre-training on a large amount of unlabeled public sentiment data to learn highly generalized speech emotion features. This method requires no large amount of labeled data, saving costs and demonstrating strong adaptability.

[0050] When using this model, the system first performs noise reduction, frame segmentation, and windowing on the original speech, then extracts its acoustic features and inputs them into emotion2vec. Through its neural encoder structure, the model maps sound into emotion embedding vectors that include emotional intensity and emotion category (such as anger, happiness, sadness, neutrality, etc.). Unlike traditional classification models, emotion2vec outputs a continuous emotion feature space, supporting fine-grained emotion analysis and exhibiting good transferability. The output emotion vectors are not only used in quality inspection to identify key dimensions such as "service attitude problems" and "customer emotional fluctuations," but also serve as an important component of multimodal feature fusion, improving the overall emotional sensitivity and contextual adaptability of the analysis.

[0051] S205, performs multimodal feature fusion on semantic feature vector, acoustic feature vector and emotion feature vector to obtain multimodal feature vector;

[0052] Specifically, the system integrates textual information, acoustic features, and emotional cues contained in speech content into a unified representation, improving the model's ability to recognize complex speech quality issues. Fusion methods can employ early fusion, late fusion, or multimodal alignment fusion techniques, such as using cross-modal Transformers, gated attention mechanisms, and fusion autoencoders, to achieve temporal alignment and semantic enhancement of vectors from different modalities. During feature fusion, the system considers the complementarity between the three modalities; for example, the semantic modality reflects "what was said," the acoustic modality reflects "how it was said," and the emotional modality reflects "the psychological state during speech." By fusing these three types of features, a more expressive multimodal feature vector can be formed, preserving the independent advantages of each modality and enhancing information accuracy through mutual verification. For example, when the semantic detection is "complaint" and the emotional feature is "anger," and the acoustic features show increased pitch and faster speech rate, it can more confidently determine "strong customer dissatisfaction," achieving accurate quality inspection.

[0053] S206 inputs the multimodal feature vectors into the pre-trained speech quality inspection model and outputs the speech quality inspection results of the speech data to be processed.

[0054] Specifically, the system inputs the multimodal feature vectors constructed in S205 into a pre-trained speech quality inspection model for comprehensive analysis. This quality inspection model can employ a multi-layer neural network architecture, such as a multimodal fusion Transformer, a dual-channel BiGRU, or a deep classification network based on an attention mechanism. The model can be configured with multiple output categories based on the quality inspection task objectives, such as "standardization of service language," "whether interruptions occurred," "customer satisfaction," and "whether there is a tendency to complain." During the training phase, the model undergoes end-to-end optimization using multi-label supervision to improve its judgment ability in complex speech interaction scenarios. During the inference phase, the model quickly judges the multimodal inputs, outputting structured quality inspection labels and scoring results, supporting joint judgment of multiple indicators and risk warnings. Simultaneously, the model also supports adaptive updates, allowing for retraining by periodically introducing new corpora to adapt to business changes.

[0055] For example, in a certain healthcare system, a patient called to inquire about the dosage of a medication. During the conversation, the patient sounded anxious and repeatedly emphasized, "This medication is very important, it must be explained clearly." However, the operator responded quickly and did not explain the key dosage information in detail, only saying, "Just take it according to the instructions." There was also a brief silence (i.e., the communication was interrupted).

[0056] To address this issue, the quality inspection model employs a dual-channel BiGRU architecture, taking semantic, acoustic, and emotional features as input. During training, multiple output categories are set, including "standardization of service language," "interruptions," "customer satisfaction," and "tendency to complain," with end-to-end optimization through multi-label supervision. During inference, the model receives the communication audio and quickly determines: "poor" standardization of service language (lack of detailed explanation of key information), "brief interruption," "low" customer satisfaction (patient's tone is anxious, needs unmet), and "high" tendency to complain. The model outputs structured quality inspection labels and scores, supporting joint judgment of multiple indicators, triggering risk warnings when multiple indicators fail to meet standards.

[0057] In the above embodiments, this application significantly improves the accuracy and intelligence level of speech quality inspection by introducing multimodal feature fusion technology. Based on traditional speech-to-text conversion, it further extracts the contextual semantic features of the text and the acoustic features of the speech, and combines them with a pre-trained emotion recognition model to obtain emotion features. This enables a comprehensive understanding of speech data from three dimensions: semantics, acoustics, and emotion. By deeply fusing the three types of feature vectors, more expressive multimodal feature vectors are obtained, which can more fully uncover quality problems in speech, including abnormal behaviors such as service attitude, speech standardization, and emotional fluctuations.

[0058] Compared to traditional methods that rely solely on text or acoustic features, this application combines emotional features with text or acoustic features to achieve multimodal features. This enables a more comprehensive and detailed reflection of the information contained in the speech content, thereby improving the accuracy and practicality of speech quality inspection and meeting the needs of high-quality automatic quality inspection in application scenarios such as intelligent customer service and voice auditing in fields such as finance and healthcare.

[0059] Furthermore, the steps of performing contextual semantic recognition on the text to be processed and embedding the semantic recognition results into semantic feature vectors specifically include:

[0060] The text to be processed is segmented into words to obtain text segments.

[0061] Use a pre-trained language model to learn the contextual semantic information of each text segment in the text to be processed;

[0062] The contextual semantic information is embedded using a pre-defined bag-of-words model to obtain semantic feature vectors.

[0063] In this embodiment, during the contextual semantic recognition of the text to be processed, the text is first segmented using a word segmentation tool. Based on the language type (e.g., Chinese uses a dictionary-based + rule-based segmentation algorithm, while English can be segmented directly by spaces), word units with semantic boundaries are obtained. Subsequently, a pre-trained language model (e.g., BERT, ERNIE, or RoBERTa) is used to model the semantic representation of each word in the context, enabling the model to not only capture the meaning of the word itself but also understand its semantic role in the sentence. This process can be implemented using a deep Transformer structure, outputting a context embedding vector for each word. Building upon this, a bag-of-words model or a TF-IDF weighted model is introduced to aggregate, reduce dimensionality, or pool multiple context embedding vectors to obtain a structured, fixed-length semantic feature vector. This vector can be used for multimodal fusion or directly as input to a quality inspection model, ensuring that the model can fully identify content dimensions such as customer demands, wording conventions, and sentiment tendencies, thereby improving the accuracy and expressive power of semantic layer analysis.

[0064] Through the above steps, a deep understanding and structured expression of text semantics can be achieved, enhancing the voice quality inspection model's ability to understand service content.

[0065] Furthermore, the pre-trained emotion recognition model is the Emotion2Vec model. The Emotion2Vec model includes an acoustic feature extractor and an emotion analyzer. The steps for extracting acoustic features from the speech data to be processed to obtain acoustic feature vectors specifically include:

[0066] Extract the acoustic spectrum corresponding to the speech data to be processed;

[0067] Use an acoustic feature extractor to extract speech acoustic features from the acoustic spectrum;

[0068] The acoustic features of the speech are embedded to obtain acoustic feature vectors.

[0069] In this embodiment, the Emotion2Vec model serves as a pre-trained emotion recognition model. Its acoustic feature extractor efficiently extracts emotion-related features from speech signals. When extracting acoustic features from speech data, the system first converts the continuous speech waveform signal into a time-frequency domain acoustic spectrum using a Short-Time Fourier Transform (STFT), forming a two-dimensional spectrogram representing the frequency distribution characteristics at different time intervals. Subsequently, the acoustic feature extractor analyzes the spectrogram based on a deep neural network (such as a CNN or Transformer structure), automatically learning underlying acoustic information such as rhythm, intonation, energy distribution, and formants—key parameters reflecting emotional states. During this process, no manual setting of feature dimensions is required; instead, the model autonomously captures discriminative features. The extracted multi-dimensional acoustic features are then transformed into structured acoustic feature vectors through a feature embedding layer (such as a fully connected layer or a projection transform layer). Because the Emotion2Vec model is pre-trained on large-scale speech emotion data, its acoustic features possess strong emotion discrimination capabilities and downstream task transfer capabilities.

[0070] Through the above steps, emotion-related acoustic features in speech can be extracted efficiently and accurately, providing a stable and reliable feature foundation for emotion recognition and speech quality inspection.

[0071] Furthermore, based on acoustic feature vectors, a pre-trained emotion recognition model is used to perform emotion recognition, and the emotion recognition results are processed by feature embedding to obtain emotion feature vectors. The specific steps include:

[0072] The acoustic feature vectors are processed by feature masking to obtain the feature mask vectors;

[0073] Map the feature mask vector to the feature space preset by the sentiment analyzer;

[0074] Within the feature space preset by the sentiment analyzer, the feature mask vector is matched with the sentiment patterns built into the sentiment analyzer to obtain the emotion recognition result;

[0075] The emotion recognition results are processed by feature embedding to obtain an emotion feature vector.

[0076] In this embodiment, the emotion recognition process employs an emotion analyzer based on the Emotion2Vec model. First, the input acoustic feature vector undergoes feature masking. The purpose of feature masking is to enhance the model's focus on key information while suppressing noise or non-emotion-related redundant features in the speech. This step uses a masking method similar to the self-attention mechanism in Transformers, assigning different weights to different dimensions of the feature vector to obtain feature mask vectors representing key emotional regions. Next, this vector is mapped to a high-dimensional emotion feature space pre-defined by the emotion analyzer. This feature space is trained using large-scale emotion-labeled speech data and has the ability to clearly distinguish various basic emotions (such as anger, happiness, calmness, anxiety, etc.). In this space, the system performs similarity calculations or vector matching between the input vector and pre-defined emotion category prototypes (i.e., emotion pattern vectors), outputting the corresponding emotion recognition result through maximum matching or multi-class regression. Finally, the system further encodes the recognition result into a unified format emotion feature vector through an emotion embedding layer, facilitating fusion with other modal features or input to a quality control model. This approach not only improves recognition accuracy but also possesses good robustness and scalability.

[0077] Through the above steps, the emotional state of speech can be efficiently identified and structured emotional features can be generated, thereby improving the emotional perception ability and intelligence level of the speech quality inspection system.

[0078] Furthermore, the step of matching the feature mask vector with the built-in sentiment patterns in the feature space of the sentiment analyzer to obtain the sentiment recognition result specifically includes:

[0079] The sentiment pattern is processed by feature embedding to obtain the sentiment pattern vector;

[0080] Within the feature space predefined by the sentiment analyzer, the cosine similarity between the feature mask vector and the sentiment pattern vector is calculated.

[0081] Based on cosine similarity, the feature mask vector and the sentiment pattern are matched to obtain the sentiment recognition result.

[0082] In this embodiment, the emotion pattern matching in the emotion recognition process relies on the pre-defined emotion feature space in the emotion analyzer. First, the various built-in emotion patterns (e.g., anger, joy, sadness, calmness, etc.) are converted into vector form, i.e., feature embedding is performed on the emotion patterns to obtain emotion pattern vectors. This embedding process can be based on statistical modeling or deep learning, mapping labeled emotion categories to discriminative vector representations in a multi-dimensional feature space. After obtaining the emotion pattern vectors, the system uses the cosine similarity method to calculate the similarity between the input feature mask vector and each emotion pattern vector. Cosine similarity measures the directional similarity between two vectors and is suitable for determining the relative positional relationship of emotional states. The system matches based on the maximum similarity value, classifying the feature mask vector into the most similar emotion category. This method avoids misjudgments caused by amplitude differences, focusing on the consistency of feature expression direction, thereby improving the stability and accuracy of emotion recognition. Furthermore, this matching process has good compatibility and scalability for newly added or refined emotion categories, supporting multi-granularity and multi-level emotion recognition needs.

[0083] By following the steps above, accurate classification of the implied emotional states in speech can be achieved.

[0084] Furthermore, the step of fusing semantic feature vectors, acoustic feature vectors, and emotion feature vectors to obtain multimodal feature vectors specifically includes:

[0085] Align the feature vector dimensions of semantic feature vectors, acoustic feature vectors, and emotion feature vectors respectively;

[0086] Calculate the attention weights for the semantic feature vector, acoustic feature vector, and emotion feature vector respectively;

[0087] Based on the attention weights, the semantic feature vector, acoustic feature vector, and emotion feature vector after dimension alignment are weighted and summed to obtain the multimodal feature vector.

[0088] In this embodiment, to fully integrate information from the semantic, acoustic, and emotional modalities, the system first performs dimensional alignment on the three types of feature vectors. Since the original features originate from different sources, the vectors differ in dimensionality and information density. Therefore, through linear transformations (such as fully connected layers) or projection operations, they are uniformly mapped to the same dimensional space, making the fusion operation comparable and computationally consistent. After alignment, the system introduces an attention mechanism to evaluate the importance of each modal feature in the current speech sample. The attention weights can be calculated based on the similarity between each feature vector and the fusion context, or by automatically learning its information contribution using a self-attention module. For example, in a customer dialogue expressing strong emotions, the attention weight of emotional features may be higher than that of semantic content. Finally, based on the calculated attention weights, the aligned vectors of the three modalities are weighted and summed to generate a multimodal feature vector with the ability to fuse contextual semantics, acoustic characteristics, and emotional state.

[0089] Through the above steps, modal information can be complemented and fused, significantly improving the accuracy and robustness of the voice quality inspection model in recognizing abnormal voices, illegal speech, and emotional instability.

[0090] Furthermore, the speech quality inspection model is a multi-classification model. The steps of inputting multi-modal feature vectors into the pre-trained speech quality inspection model and outputting the speech quality inspection results of the speech data to be processed specifically include:

[0091] Obtain the original confidence scores corresponding to each quality inspection label built into the voice quality inspection model;

[0092] Based on the preset category index, the multimodal feature vectors are mapped to each quality inspection label;

[0093] The original confidence level is updated using the normalization function built into the voice quality inspection model, and the probability distribution of each quality inspection label is generated based on the updated original confidence level.

[0094] The voice quality inspection result of the voice data to be processed is determined based on the probability distribution of each quality inspection label.

[0095] In this embodiment, the speech quality inspection model is a multi-classification structure, capable of comprehensively discriminating the input multimodal feature vectors and outputting corresponding quality inspection labels. First, the system uses its internal category discrimination unit to obtain the original confidence scores associated with each preset quality inspection label (such as "speech too fast," "emotional instability," "sensitive word violation," "poor service attitude," etc.). This confidence score reflects the model's initial judgment strength regarding the current sample's classification into each label. Subsequently, based on the category index defined during training, the system maps the fused multimodal feature vectors to the vector representation space corresponding to each label. To ensure the interpretability and stability of the output, the model further applies a normalization function to the original confidence scores. Common methods include the Softmax function or a temperature-adjusted normalization function, used to balance the differences between labels and avoid extreme high or low probabilities for any particular label. Finally, based on the updated label probability distribution, the system selects the label with the highest probability as the final quality inspection result for the speech data. Alternatively, it can output Top-N probability results to support multi-label recognition based on task requirements. During deployment, the model can process large amounts of voice data in real time and quickly output classification judgments, adapting to the high-efficiency quality inspection needs of enterprise call centers, customer service supervision and other scenarios.

[0096] Raw confidence scores, also known as classification logits (or unnormalized classification scores), are the raw numerical values ​​output by a neural network in a classification task. They typically represent the model's "raw confidence" in each class. Raw confidence scores are the output of the last layer of the model (usually a fully connected layer) and have not yet undergone normalization processing such as Softmax or Sigmoid.

[0097] The Softmax function can be used to convert Logits into a probability distribution:

[0098]

[0099] Among them, z i It is the Logit value of the i-th class. Calculating the loss directly on the Logits (such as cross-entropy loss) is more stable than calculating it on the probability, avoiding the numerical overflow problem of Softmax (especially when the Logits value is very large).

[0100] The formula for cross-entropy loss is:

[0101]

[0102] By following the steps above, various voice quality problems can be accurately classified, thereby enhancing the intelligence level and business coverage of the voice quality inspection system.

[0103] In the above embodiments, this application discloses a voice quality inspection method, belonging to the field of artificial intelligence technology, and applied to voice quality inspection scenarios in finance, healthcare, and other fields. This application significantly improves the accuracy and intelligence level of voice quality inspection by introducing multimodal feature fusion technology. Based on traditional speech-to-text conversion, it further extracts the contextual semantic features of the text and the acoustic features of the speech, and combines this with a pre-trained emotion recognition model to obtain emotional features. This achieves a comprehensive understanding of speech data from three dimensions: semantics, acoustics, and emotion. By deeply fusing these three types of feature vectors, a more expressive multimodal feature vector is obtained, which can more fully uncover quality problems in the speech, including abnormal behaviors such as service attitude, speech format, and emotional fluctuations. Compared to traditional methods that rely solely on text or acoustic features, this application combines emotional features with text or acoustic features to achieve multimodal features, which can more comprehensively and meticulously reflect the information contained in the speech content, thereby improving the accuracy and practicality of voice quality inspection and meeting the needs of high-quality automatic quality inspection in application scenarios such as intelligent customer service and voice auditing in finance, healthcare, and other fields.

[0104] In this embodiment, the voice quality inspection method operates on an electronic device (e.g., Figure 1 The server shown can receive instructions or acquire data via wired or wireless connection. It should be noted that the aforementioned wireless connection methods may include, but are not limited to, 3G / 4G connections, WiFi connections, Bluetooth connections, WiMAX connections, Zigbee connections, UWB (ultra-wideband) connections, and other currently known or future wireless connection methods.

[0105] It should be emphasized that, to further ensure the privacy and security of the aforementioned voice data, the voice data can also be stored in a blockchain node.

[0106] The blockchain referred to in this application is a novel application model of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanisms, and encryption algorithms. Essentially, a blockchain is a decentralized database, a chain of data blocks linked together using cryptographic methods. Each data block contains information about a batch of network transactions, used to verify the validity of the information (anti-counterfeiting) and generate the next block. A blockchain can include an underlying blockchain platform, a platform product service layer, and an application service layer.

[0107] The embodiments of this application can acquire and process relevant data based on artificial intelligence technology. Artificial intelligence (AI) is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to obtain optimal results.

[0108] Foundational technologies for artificial intelligence generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies mainly encompass computer vision, robotics, biometrics, speech processing, natural language processing, and machine learning / deep learning.

[0109] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by instructing related hardware with computer-readable instructions. These computer-readable instructions can be stored in a computer-readable storage medium. When executed, the program can include the processes of the embodiments of the above methods. The aforementioned storage medium can be a non-volatile storage medium such as a magnetic disk, optical disk, or read-only memory (ROM), or random access memory (RAM).

[0110] It should be understood that although the steps in the flowcharts of the accompanying figures are shown sequentially as indicated by the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the accompanying figures may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the sub-steps or stages of other steps.

[0111] Further reference Figure 3 As a response to the above Figure 2 To implement the method shown, this application provides an embodiment of a voice quality inspection device, which is similar to... Figure 2 Corresponding to the method embodiments shown, this device can be specifically applied to various electronic devices.

[0112] like Figure 3 As shown, the voice quality inspection device 300 described in this embodiment includes:

[0113] The text conversion module 301 is used to acquire the speech data to be processed and to perform speech-to-text processing on the speech data to be processed to obtain the text to be processed.

[0114] The semantic recognition module 302 is used to perform contextual semantic recognition on the text to be processed, and to perform feature embedding processing on the text semantic recognition results to obtain a semantic feature vector;

[0115] The acoustic feature module 303 is used to extract acoustic features from the speech data to be processed, and obtain an acoustic feature vector.

[0116] The emotion recognition module 304 is used to perform emotion recognition based on acoustic feature vectors using a pre-trained emotion recognition model, and to perform feature embedding processing on the emotion recognition results to obtain emotion feature vectors.

[0117] The multimodal fusion module 305 is used to perform multimodal feature fusion on semantic feature vector, acoustic feature vector and emotion feature vector to obtain multimodal feature vector;

[0118] The speech quality inspection module 306 is used to input multimodal feature vectors into a pre-trained speech quality inspection model and output the speech quality inspection results of the speech data to be processed.

[0119] Furthermore, the semantic recognition module 302 specifically includes:

[0120] The word segmentation unit is used to segment the text to be processed into words to obtain text words.

[0121] The semantic recognition unit is used to learn the contextual semantic information of each text segment in the text to be processed using a pre-trained language model;

[0122] The semantic embedding unit is used to perform feature embedding processing on contextual semantic information using a pre-defined bag-of-words model to obtain semantic feature vectors.

[0123] Furthermore, the pre-trained emotion recognition model is the Emotion2Vec model, which includes an acoustic feature extractor and an emotion analyzer. The acoustic feature module 303 specifically includes:

[0124] The spectrum extraction unit is used to extract the acoustic spectrum corresponding to the speech data to be processed;

[0125] Acoustic feature unit, used to extract speech acoustic features from the acoustic spectrum using an acoustic feature extractor;

[0126] The acoustic feature embedding unit is used to perform feature embedding processing on the acoustic features of speech to obtain acoustic feature vectors.

[0127] Furthermore, the emotion recognition module 304 specifically includes:

[0128] The feature masking unit is used to perform feature masking processing on the acoustic feature vector to obtain the feature mask vector;

[0129] The first vector mapping unit is used to map the feature mask vector to the feature space preset by the sentiment analyzer.

[0130] The pattern matching unit is used to match the feature mask vector with the built-in sentiment patterns of the sentiment analyzer within the preset feature space of the sentiment analyzer to obtain the sentiment recognition result.

[0131] The emotion feature embedding unit is used to perform feature embedding processing on the emotion recognition results to obtain the emotion feature vector.

[0132] Furthermore, the pattern matching unit specifically includes:

[0133] The pattern feature embedding subunit is used to perform feature embedding processing on the sentiment pattern to obtain the sentiment pattern vector.

[0134] The cosine similarity calculation subunit is used to calculate the cosine similarity between the feature mask vector and the sentiment pattern vector within the preset feature space of the sentiment analyzer.

[0135] The pattern matching subunit is used to match the feature mask vector and the sentiment pattern based on cosine similarity to obtain the sentiment recognition result.

[0136] Furthermore, the multimodal fusion module 305 specifically includes:

[0137] The dimension alignment unit is used to perform feature vector dimension alignment on the semantic feature vector, acoustic feature vector, and emotion feature vector respectively.

[0138] Attention weighting unit, used to calculate attention weights for semantic feature vector, acoustic feature vector and emotion feature vector respectively;

[0139] The weighted summation unit is used to perform weighted summation on the dimension-aligned semantic feature vector, acoustic feature vector, and emotion feature vector according to the attention weights, so as to obtain a multimodal feature vector.

[0140] Furthermore, the voice quality inspection model is a multi-classification model, and the voice quality inspection module 306 specifically includes:

[0141] The confidence acquisition unit is used to acquire the original confidence scores corresponding to each quality inspection label built into the speech quality inspection model.

[0142] The second vector mapping unit is used to map multimodal feature vectors to various quality inspection labels according to a preset category index.

[0143] The confidence update unit is used to update the original confidence using the normalization function built into the speech quality inspection model, and to generate the probability distribution of each quality inspection label based on the updated original confidence.

[0144] The voice quality inspection unit is used to determine the voice quality inspection result of the voice data to be processed based on the probability distribution of each quality inspection label.

[0145] In the above embodiments, this application discloses a voice quality inspection device, belonging to the field of artificial intelligence technology, applied to voice quality inspection scenarios in finance, healthcare, and other fields. This application significantly improves the accuracy and intelligence level of voice quality inspection by introducing multimodal feature fusion technology. Based on traditional speech-to-text conversion, it further extracts the contextual semantic features of the text and the acoustic features of the speech, and combines this with a pre-trained emotion recognition model to obtain emotional features. This achieves a comprehensive understanding of speech data from three dimensions: semantics, acoustics, and emotion. By deeply fusing these three types of feature vectors, a more expressive multimodal feature vector is obtained, which can more fully uncover quality problems in the speech, including abnormal behaviors such as service attitude, speech format, and emotional fluctuations. Compared to traditional methods that rely solely on text or acoustic features, this application combines emotional features with text or acoustic features to achieve multimodal features, which can more comprehensively and meticulously reflect the information contained in the speech content, thereby improving the accuracy and practicality of voice quality inspection and meeting the needs of high-quality automatic quality inspection in application scenarios such as intelligent customer service and voice auditing in finance, healthcare, and other fields.

[0146] To address the aforementioned technical problems, embodiments of this application also provide a computer device. Please refer to [link / reference] for details. Figure 4 , Figure 4 This is a basic structural block diagram of the computer device in this embodiment.

[0147] The computer device 4 includes a memory 41, a processor 42, and a network interface 43 that are interconnected via a system bus. It should be noted that only a computer device 4 with a memory 41, a processor 42, and a network interface 43 is shown in the figure; however, it should be understood that it is not required to implement all the components shown, and more or fewer components can be implemented alternatively. Those skilled in the art will understand that the computer device described here is a device capable of automatically performing numerical calculations and / or information processing according to pre-set or stored instructions, and its hardware includes, but is not limited to, microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), embedded devices, etc.

[0148] The computer device can be a desktop computer, laptop, handheld computer, or cloud server, etc. The computer device can interact with the user via a keyboard, mouse, remote control, touchpad, or voice control.

[0149] The memory 41 includes at least one type of readable storage medium, including flash memory, hard disk, multimedia card, card-type memory (e.g., SD or DX memory), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the memory 41 may be an internal storage unit of the computer device 4, such as the hard disk or memory of the computer device 4. In other embodiments, the memory 41 may also be an external storage device of the computer device 4, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc., equipped on the computer device 4. Of course, the memory 41 may include both the internal storage unit and its external storage device of the computer device 4. In this embodiment, the memory 41 is typically used to store the operating system and various application software installed on the computer device 4, such as computer-readable instructions for voice quality inspection methods. In addition, the memory 41 can also be used to temporarily store various types of data that have been output or will be output.

[0150] In some embodiments, the processor 42 may be a central processing unit (CPU), controller, microcontroller, microprocessor, or other data processing chip. The processor 42 is typically used to control the overall operation of the computer device 4. In this embodiment, the processor 42 is used to execute computer-readable instructions stored in the memory 41 or to process data, for example, to execute computer-readable instructions for the voice quality inspection method.

[0151] The network interface 43 may include a wireless network interface or a wired network interface, which is typically used to establish communication connections between the computer device 4 and other electronic devices.

[0152] This application also provides an implementation method, namely, a computer device including a memory and a processor. The memory stores computer-readable instructions, and the processor, when executing the computer-readable instructions, implements the steps of the above-described voice quality inspection method, that is, it implements:

[0153] A voice quality inspection method, comprising:

[0154] Acquire the speech data to be processed, and perform speech-to-text processing on the speech data to obtain the text to be processed;

[0155] Contextual semantic recognition is performed on the text to be processed, and the semantic recognition results are then processed by feature embedding to obtain semantic feature vectors;

[0156] Acoustic feature extraction is performed on the speech data to be processed to obtain acoustic feature vectors;

[0157] Based on acoustic feature vectors, a pre-trained emotion recognition model is used to perform emotion recognition, and the emotion recognition results are processed by feature embedding to obtain emotion feature vectors.

[0158] Multimodal feature fusion is performed on semantic feature vectors, acoustic feature vectors, and emotion feature vectors to obtain multimodal feature vectors;

[0159] The multimodal feature vectors are input into a pre-trained speech quality inspection model, which outputs the speech quality inspection results of the speech data to be processed.

[0160] This application also provides another embodiment, namely, a computer-readable storage medium storing computer-readable instructions that can be executed by at least one processor to cause the at least one processor to perform the steps of the above-described voice quality inspection method, i.e., to achieve:

[0161] A voice quality inspection method, comprising:

[0162] Acquire the speech data to be processed, and perform speech-to-text processing on the speech data to obtain the text to be processed;

[0163] Contextual semantic recognition is performed on the text to be processed, and the semantic recognition results are then processed by feature embedding to obtain semantic feature vectors;

[0164] Acoustic feature extraction is performed on the speech data to be processed to obtain acoustic feature vectors;

[0165] Based on acoustic feature vectors, a pre-trained emotion recognition model is used to perform emotion recognition, and the emotion recognition results are processed by feature embedding to obtain emotion feature vectors.

[0166] Multimodal feature fusion is performed on semantic feature vectors, acoustic feature vectors, and emotion feature vectors to obtain multimodal feature vectors;

[0167] The multimodal feature vectors are input into a pre-trained speech quality inspection model, which outputs the speech quality inspection results of the speech data to be processed.

[0168] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of this application.

[0169] This application can be used in a wide variety of general-purpose or special-purpose computer system environments or configurations. Examples include: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics devices, network PCs, minicomputers, mainframe computers, and distributed computing environments including any of the above systems or devices. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform specific tasks or implement specific abstract data types. This application can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.

[0170] It should be noted that the software tools or components not belonging to this company that appear in the various embodiments of this application are merely illustrative examples and do not represent actual use.

[0171] Obviously, the embodiments described above are only some embodiments of this application, not all embodiments. The accompanying drawings show preferred embodiments of this application, but do not limit the patent scope of this application. This application can be implemented in many different forms; rather, the purpose of providing these embodiments is to provide a more thorough and comprehensive understanding of the disclosure of this application. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing specific embodiments, or make equivalent substitutions for some of the technical features. Any equivalent structures made using the content of this application's specification and drawings, directly or indirectly applied to other related technical fields, are similarly within the scope of patent protection of this application.

Claims

1. A voice quality inspection method, characterized in that, include: Acquire the voice data to be processed, and perform speech-to-text processing on the voice data to obtain the text to be processed; The text to be processed is subjected to contextual semantic recognition, and the text semantic recognition results are subjected to feature embedding processing to obtain a semantic feature vector; Acoustic feature extraction is performed on the speech data to be processed to obtain an acoustic feature vector; Based on the acoustic feature vector, a pre-trained emotion recognition model is used to perform emotion recognition, and the emotion recognition results are processed by feature embedding to obtain an emotion feature vector. Multimodal feature fusion is performed on the semantic feature vector, the acoustic feature vector, and the emotion feature vector to obtain a multimodal feature vector; The multimodal feature vectors are input into a pre-trained speech quality inspection model, and the speech quality inspection results of the speech data to be processed are output. The step of fusing the semantic feature vector, the acoustic feature vector, and the emotion feature vector to obtain a multimodal feature vector specifically includes: The semantic feature vector, the acoustic feature vector, and the emotion feature vector are respectively aligned in terms of feature vector dimension. Calculate the attention weights for the semantic feature vector, the acoustic feature vector, and the emotion feature vector, respectively; Based on the attention weights, the semantic feature vector, the acoustic feature vector, and the emotion feature vector after dimension alignment are weighted and summed to obtain the multimodal feature vector; The speech quality inspection model is a multi-class classification model. The step of inputting the multi-modal feature vector into the pre-trained speech quality inspection model and outputting the speech quality inspection result of the speech data to be processed specifically includes: Obtain the original confidence level corresponding to each quality inspection label built into the voice quality inspection model; Based on the preset category index, the multimodal feature vector is mapped to each quality inspection label; The original confidence level is updated using the normalization function built into the voice quality inspection model, and a probability distribution of each quality inspection label is generated based on the updated original confidence level. The voice quality inspection result of the voice data to be processed is determined based on the probability distribution of each of the quality inspection labels.

2. The voice quality inspection method as described in claim 1, characterized in that, The steps of performing contextual semantic recognition on the text to be processed and performing feature embedding processing on the text semantic recognition results to obtain semantic feature vectors specifically include: The text to be processed is segmented into words to obtain text words; The pre-trained language model is used to learn the contextual semantic information of each of the text segments in the text to be processed; The contextual semantic information is embedded using a pre-defined bag-of-words model to obtain the semantic feature vector.

3. The voice quality inspection method as described in claim 1, characterized in that, The pre-trained emotion recognition model is an Emotion2Vec model, which includes an acoustic feature extractor and an emotion analyzer. The step of extracting acoustic features from the speech data to be processed to obtain acoustic feature vectors specifically includes: Extract the acoustic spectrum corresponding to the speech data to be processed; The acoustic feature extractor is used to extract speech acoustic features from the acoustic spectrum; The acoustic features of the speech are subjected to feature embedding processing to obtain the acoustic feature vector.

4. The voice quality inspection method as described in claim 3, characterized in that, The steps of performing emotion recognition using a pre-trained emotion recognition model based on the acoustic feature vector, and then performing feature embedding processing on the emotion recognition results to obtain the emotion feature vector, specifically include: The acoustic feature vector is subjected to feature masking to obtain a feature mask vector; Map the feature mask vector to the preset feature space of the sentiment analyzer; Within the preset feature space of the sentiment analyzer, the feature mask vector is matched with the sentiment patterns built into the sentiment analyzer to obtain the sentiment recognition result. The emotion recognition results are processed by feature embedding to obtain an emotion feature vector.

5. The voice quality inspection method as described in claim 4, characterized in that, The step of matching the feature mask vector with the built-in emotion patterns of the emotion analyzer within the preset feature space of the emotion analyzer to obtain the emotion recognition result specifically includes: The emotional pattern is subjected to feature embedding processing to obtain an emotional pattern vector; Within the preset feature space of the sentiment analyzer, the cosine similarity between the feature mask vector and the sentiment pattern vector is calculated; The feature mask vector and the emotion pattern are matched based on the cosine similarity to obtain the emotion recognition result.

6. A voice quality inspection device, characterized in that, include: The text conversion module is used to acquire the speech data to be processed and to perform speech-to-text processing on the speech data to obtain the text to be processed. The semantic recognition module is used to perform contextual semantic recognition on the text to be processed, and to perform feature embedding processing on the text semantic recognition results to obtain a semantic feature vector; An acoustic feature module is used to extract acoustic features from the speech data to be processed to obtain an acoustic feature vector; The emotion recognition module is used to perform emotion recognition based on the acoustic feature vector using a pre-trained emotion recognition model, and to perform feature embedding processing on the emotion recognition results to obtain an emotion feature vector. A multimodal fusion module is used to perform multimodal feature fusion on the semantic feature vector, the acoustic feature vector, and the emotion feature vector to obtain a multimodal feature vector; The speech quality inspection module is used to input the multimodal feature vector into a pre-trained speech quality inspection model and output the speech quality inspection result of the speech data to be processed. Furthermore, the multimodal fusion module specifically includes: A dimension alignment unit is used to perform feature vector dimension alignment on the semantic feature vector, the acoustic feature vector, and the emotion feature vector, respectively. An attention weighting unit is used to calculate the attention weights of the semantic feature vector, the acoustic feature vector, and the emotion feature vector, respectively. The weighted summation unit is used to perform a weighted summation of the dimension-aligned semantic feature vector, the acoustic feature vector, and the emotion feature vector according to the attention weights to obtain the multimodal feature vector; The voice quality inspection model is a multi-classification model, and the voice quality inspection module specifically includes: The confidence level acquisition unit is used to acquire the original confidence level corresponding to each quality inspection label built into the speech quality inspection model. The second vector mapping unit is used to map the multimodal feature vector to each quality inspection label according to a preset category index; The confidence update unit is used to update the original confidence using the normalization function built into the speech quality inspection model, and generate the probability distribution of each quality inspection label based on the updated original confidence. The voice quality inspection unit is used to determine the voice quality inspection result of the voice data to be processed based on the probability distribution of each of the quality inspection tags.

7. A computer device, characterized in that, The device includes a memory and a processor, wherein the memory stores computer-readable instructions, and the processor executes the computer-readable instructions to implement the steps of the voice quality inspection method as described in any one of claims 1 to 5.

8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-readable instructions, which, when executed by a processor, implement the steps of the voice quality inspection method as described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Customer service voice quality inspection method, device and equipment and storage medium

    CN118035411A

  • User emotion recognition method based on AI and voice data

    CN120148561A