Voice quality inspection method and device, electronic equipment and storage medium
By dividing the voice signals into text data, combining multimodal feature extraction and industry policy database for quality inspection, the problem of time-consuming and labor-intensive artificial quality inspection is solved, and efficient and accurate voice quality inspection is achieved.
Patent Information
- Application Number
- CN202510783473.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-12
- Publication Date
- 2025-08-15
- Estimated Expiration
- 2045-06-12
AI Technical Summary
In the prior art, artificial quality inspection methods are time-consuming and labor-intensive, and cannot fully and accurately detect voice calls, resulting in low voice quality inspection coverage and missed dialogue inspections.
By obtaining voice signals for voice tracking and identity identification, using the preset voice recognition model to convert it into text data, and combining the preset voice quality inspection model for multimodal feature extraction and quality inspection, combining the industry knowledge base and external policy database for comprehensive quality inspection to determine the target quality inspection results.
A comprehensive and timely quality inspection of the session to be inspected when the call is completed is achieved, which improves the efficiency and accuracy of voice quality inspection and reduces the need for manual quality inspection.
Smart Images

Figure CN120496574A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present invention relate to the field of artificial intelligence technology, and in particular to a speech quality inspection method, device, electronic device, and storage medium. Background Art
[0002] With the development of artificial intelligence, remote voice communication is becoming increasingly popular across various fields to provide customer service. For example, banking service personnel often answer customer questions, provide product recommendations, and conduct sales over the phone. However, as the number of customers increases, the volume of voice communication between service personnel and customers also increases. To further identify service issues, mitigate risks, and optimize business processes, voice quality inspection is used to monitor and test call conversations.
[0003] Currently, manual quality control is commonly used to inspect voice calls. However, this method is time-consuming and labor-intensive, often failing to address all aspects that require inspection. It may even focus only on a single aspect, such as whether the service agent's response violates regulations. This results in incomplete and inaccurate voice quality control. Furthermore, this method often relies on spot checks, resulting in low coverage and the possibility of overlooking illegal conversations. Summary of the Invention
[0004] The embodiments of the present invention provide a voice quality inspection method, device, electronic device and storage medium, which can timely perform comprehensive quality inspection on the quality inspection session of the ended call, and perform timeliness quality inspection on the policy content involved in the quality inspection session, without the need for manual quality inspection, thereby improving the efficiency and accuracy of voice quality inspection.
[0005] In a first aspect, an embodiment of the present invention provides a method for speech quality inspection, comprising:
[0006] Acquire a voice signal of the conversation to be quality inspected, and perform voice track segmentation on the voice signal to determine voice data containing the identity identifier;
[0007] Performing format conversion on the voice data based on a preset voice recognition model to determine the text data corresponding to the conversation to be inspected;
[0008] Performing multimodal feature extraction and voice quality inspection based on the first preset instruction, the preset voice quality inspection model, the voice data, and the text data to determine a first quality inspection result of the conversation to be inspected;
[0009] Performing a quality inspection based on a second preset instruction, the text data, the preset voice quality inspection model, an industry knowledge base, and an external policy library to determine a second quality inspection result of the session to be quality inspected;
[0010] A target quality inspection result of the session to be quality inspected is determined based on the first quality inspection result and the second quality inspection result.
[0011] Optionally, the method further includes: performing noise reduction processing on the voice signal for steady-state noise and non-steady-state noise based on a preset environmental noise reduction method, and determining a voice signal that retains a clear human voice; performing voice track segmentation and role labeling on the voice signal that retains a clear human voice based on voiceprint recognition technology, and determining voice data containing an identity identifier.
[0012] Optionally, the method also includes: the preset speech recognition model is constructed by loading a preset banking business-specific dictionary on the basis of a general speech recognition model; the voice data is formatted based on the preset speech recognition model to determine the text data to be detected; the text data to be detected is error-corrected and modified based on contextual semantics to determine the corresponding text data of the session to be quality inspected.
[0013] Optionally, the method also includes: inputting a first preset instruction, the voice data and the text data into a preset voice quality inspection model; in the preset voice quality inspection model, performing feature extraction on the voice data based on the first operation instruction in the first preset instruction, and determining the emotional features and speech speed features corresponding to each dialogue in the session to be inspected; performing feature extraction on the text data based on the second operation instruction in the first preset instruction, and determining the keywords and intentions corresponding to each dialogue in the session to be inspected; performing voice quality inspection based on the first preset instruction and the emotional features, speech speed features, keywords and intentions corresponding to each dialogue in the session to be inspected, and determining the first quality inspection result of the session to be inspected.
[0014] Optionally, the method also includes: based on the third operation instruction in the first preset instruction, feature-associate the emotional features, speech speed features, keywords and intentions corresponding to each conversation to determine the conversation features of each conversation, and determine the conversation status of each conversation based on the conversation features of each conversation; based on the fourth operation instruction in the first preset instruction and the conversation status, perform progressive quality inspection on each conversation to determine the first quality inspection result of the conversation to be quality inspected.
[0015] Optionally, the method further includes: performing single-conversation quality inspection on each conversation based on the fourth operation instruction in the first preset instruction to determine the individual quality inspection result of each conversation; performing correlation calculation on each conversation state to determine the semantic correlation between each two conversations, and determining a conversation set for joint quality inspection based on the preset correlation threshold and each semantic correlation; performing conversation joint quality inspection on the conversation set to determine the joint quality inspection result of the conversation set; and determining the first quality inspection result of the session to be quality inspected based on the individual quality inspection result and the joint quality inspection result.
[0016] Optionally, the method also includes: inputting a second preset instruction, the text data, an industry knowledge base and an external policy base into the preset voice quality inspection model; in the preset voice quality inspection model, based on the second preset instruction and in combination with the industry knowledge base, performing information identification on the text data to determine the industry-specific professional information in the conversation to be quality inspected and the limiting information corresponding to the professional information; searching the professional information in combination with the external policy base to determine the standard information corresponding to the professional information; performing quality inspection on the limiting information and the standard information to determine the second quality inspection result of the conversation to be quality inspected.
[0017] In a second aspect, an embodiment of the present invention further provides a speech quality inspection device, the device comprising:
[0018] A voice data determination module, configured to obtain a voice signal of a conversation to be inspected, and perform voice track separation on the voice signal to determine voice data containing an identity identifier;
[0019] A text data determination module, configured to convert the format of the voice data based on a preset voice recognition model to determine the text data corresponding to the conversation to be inspected;
[0020] a first quality inspection result determination module, configured to perform multimodal feature extraction and voice quality inspection based on a first preset instruction, a preset voice quality inspection model, the voice data, and the text data, and determine a first quality inspection result of the conversation to be inspected;
[0021] A second quality inspection result determination module, configured to perform quality inspection based on a second preset instruction, the text data, the preset voice quality inspection model, the industry knowledge base, and the external policy library, and determine a second quality inspection result of the to-be-inspected session;
[0022] A target quality inspection result determination module is configured to determine a target quality inspection result of the session to be quality inspected based on the first quality inspection result and the second quality inspection result.
[0023] In a third aspect, an embodiment of the present invention further provides an electronic device, comprising:
[0024] one or more processors;
[0025] a memory for storing one or more programs;
[0026] When the one or more programs are executed by the one or more processors, the one or more processors implement the speech quality inspection method provided by any embodiment of the present invention.
[0027] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the speech quality inspection method provided by any embodiment of the present invention.
[0028] In a fifth aspect, an embodiment of the present invention provides a computer program product, including a computer program, which, when executed by a processor, implements the speech quality inspection method provided by any embodiment of the present invention.
[0029] The technical solution of the embodiment of the present invention obtains the voice signal of the conversation to be quality inspected, and performs voice track separation on the voice signal to determine the voice data containing the identity identifier; performs format conversion on the voice data based on a preset voice recognition model to determine the corresponding text data of the conversation to be quality inspected; performs multimodal feature extraction and voice quality inspection based on a first preset instruction, a preset voice quality inspection model, the voice data and the text data to determine a first quality inspection result of the conversation to be quality inspected; performs quality inspection based on a second preset instruction, the text data, the preset voice quality inspection model, an industry knowledge base and an external policy base to determine a second quality inspection result of the conversation to be quality inspected; determines a target quality inspection result of the conversation to be quality inspected based on the first quality inspection result and the second quality inspection result, so that a comprehensive quality inspection can be performed on the conversation to be quality inspected after the call is ended in a timely manner, and a timeliness quality inspection can be performed on the policy content involved in the conversation to be quality inspected without the need for manual quality inspection, thereby improving the efficiency and accuracy of voice quality inspection.
[0030] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present invention, nor is it intended to limit the scope of the present invention. Other features of the present invention will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0032] Figure 1 This is a flow chart of a voice quality inspection method provided by Example 1 of the present invention;
[0033] Figure 2 This is a flow chart of a voice quality inspection method provided by Embodiment 2 of the present invention;
[0034] Figure 3 This is a structural diagram of a speech quality inspection device provided in Embodiment 3 of the present invention;
[0035] Figure 4 The figure is a schematic diagram of the structure of an electronic device for implementing the voice quality inspection method according to an embodiment of the present invention. DETAILED DESCRIPTION
[0036] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.
[0037] It should be noted that the terms "first", "second", etc. in the description and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the numbers used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0038] Example 1
[0039] Figure 1 A flow chart of a voice quality inspection method is provided for the first embodiment of the present invention. This embodiment is applicable to the case where a large model is used to inspect the quality of a call. The method can be executed by a voice quality inspection device, which can be implemented in the form of hardware and / or software and can be configured in an electronic device. Figure 1 As shown, the method includes:
[0040] S110: Acquire a voice signal of a conversation to be inspected, and perform voice track separation on the voice signal to determine voice data containing an identity identifier.
[0041] Among them, the conversation awaiting quality inspection may refer to a conversation whose content needs to be checked for violations. In this embodiment, the conversation awaiting quality inspection may be a conversation generated during the course of banking services, and the conversation content needs to be checked for compliance with the bank's internal requirements. For example, the conversation awaiting quality inspection may be, but is not limited to, a conversation between a customer and a customer service representative, a conversation between a customer and an account manager, or a conference conversation. The voice signal may refer to a sound wave signal that records the content of the conversation. The voice signal may include, but is not limited to, a human voice and environmental sounds. In this embodiment, the voice signal may be, but is not limited to, a call recording or a conference recording. The identity identifier may refer to the identity identifier of the person participating in the conversation. The identity identifier can be used to distinguish the conversation participants in the conversation awaiting quality inspection. In this embodiment, the identity identifier may include, but is not limited to, the customer and the customer service representative. The voice data may refer to the voice data in which all participants in the conversation awaiting quality inspection have been distinguished. In this embodiment, the voice data may be the voice data obtained by labeling the voice signal with the identity of the participant to whom the voice signal belongs.
[0042] Specifically, call records uploaded by customers are received to obtain voice signals from conversations awaiting quality inspection. Voice signals from conversations awaiting quality inspection can also be periodically retrieved from a call record repository. This call record repository can serve as a random inspection database for bank service personnel to assess service compliance. Ambient noise reduction and voiceprint recognition are performed on the voice signals to generate voice data with an identity tag along a complete timeline. For example, one time period can be labeled as customer voice, while another time period can be labeled as customer service voice.
[0043] On the basis of the above technical solution, "segmenting the voice signal into voice tracks to determine the voice data containing the identity identification" may include: performing noise reduction processing on the voice signal for steady-state noise and non-steady-state noise based on a preset environmental noise reduction method to determine the voice signal that retains a clear human voice; performing voice segmentation and role labeling on the voice signal that retains a clear human voice based on voiceprint recognition technology to determine the voice data containing the identity identification.
[0044] The preset environmental noise reduction method may refer to a pre-set environmental noise removal method. In this embodiment, the preset environmental noise reduction method can be used to remove typical bank noise, such as the sounds of counting money and keyboards. The preset environmental noise reduction method may be a noise reduction method that combines a noise reduction algorithm and a deep neural network. The noise reduction algorithm is used to reduce steady-state noise. Deep neural networks (DNNs) are used to reduce non-steady-state noise. In this embodiment, the noise reduction algorithm may be, but is not limited to, spectral subtraction. Steady-state noise may refer to noise that is continuous, regular, or spectrally stable. For example, steady-state noise may include, but is not limited to, the low-frequency hum of an air conditioner and the sound of a server fan. Non-steady-state noise may refer to noise that is sudden, irregular, or has dynamically changing spectra. For example, non-steady-state noise may include, but is not limited to, the sounds of customers queuing, keyboard tapping, and counting money. Voiceprint recognition technology may refer to a pre-set technology that identifies the speaker's identity by analyzing biometric features in speech, such as vocal cord structure and pronunciation habits.
[0045] Specifically, spectral subtraction or linear filtering is used to quickly eliminate steady-state noise, while a noise reduction model pre-trained for banking scenarios is used to remove non-stationary noise. This approach eliminates both steady-state and non-stationary noise, resulting in a voice signal that retains a clear human voice, thereby improving the effectiveness of the voice signal noise reduction process. A pre-trained voiceprint recognition model is used to extract biometric features and perform voiceprint comparison on the voice signal that retains a clear human voice, generating voice data containing identity identifiers. This allows for dialogue character labeling while maintaining voice quality, ultimately outputting clean, dual-channel voice data with identity tags.
[0046] In this embodiment, in a bank customer service scenario, the process of implementing voice tracking and character labeling based on voiceprint recognition technology may include: framing the voice signal that retains clear human voice (20-30ms / frame) and extracting MFCC features to obtain static features (fundamental frequency, formant frequency) and dynamic features (speech rate, intonation). Based on energy mutation detection or silence segmentation, the speech segments of different speakers (such as alternating conversations between customer service and customer) are separated. The speech of the distinguished speakers is compared with a pre-registered voiceprint database. For example, the customer service voiceprint has been pre-entered into the system (ID: Agent_001), and the customer voiceprint is registered in real time (ID: User_12345). Speaker switching points are marked on the timeline, and a voice stream with identity tags is generated (e.g., [00:12-00:30] Agent_001: "What service do I need to handle?"), thereby obtaining voice data containing identity identifiers.
[0047] S120: Convert the format of the voice data based on a preset voice recognition model to determine the corresponding text data of the conversation to be quality inspected.
[0048] The preset speech recognition model may refer to a pre-trained deep learning model for converting speech into text. In this embodiment, the preset speech recognition model is a joint model including an acoustic model and a language model. Text data may refer to data that is converted from speech content and displayed in text form.
[0049] Specifically, speech data is fed into a pre-set speech recognition model. An acoustic model (such as a Hidden Markov Model (HMM)) is used to convert the speech data into phonemes or character sequences. A language model is then used to assess the probability of these character sequences, ensuring semantic coherence in the recognition results. The language model combines context and grammatical rules to optimize the acoustic model output, making the converted text data more reasonable.
[0050] Based on the above technical solution, the preset speech recognition model is constructed by loading a preset banking-specific dictionary onto a general speech recognition model. "Converting the speech data format based on the preset speech recognition model to determine the corresponding text data for the conversation to be inspected" may include: converting the speech data format based on the preset speech recognition model to determine the text data to be inspected; and performing error correction and modification on the text data to be inspected based on contextual semantics to determine the corresponding text data for the conversation to be inspected.
[0051] The preset banking-specific dictionary can refer to a preconfigured dictionary specifically containing banking-related terminology. This dictionary can be used to enable the model to accurately translate and understand specialized financial terms. For example, the word "bandwidth" in text data to be tested can be modified to "loan" using the preset banking-specific dictionary internalized by the model, thereby improving the model's recognition accuracy for banking-specific terms such as "wealth management products" and "cross-border remittances."
[0052] Specifically, a deep learning-based speech-to-text framework was pre-built, integrating a banking-specific adaptive module into the basic Automatic Speech Recognition (ASR) engine. This effectively improves the accuracy of professional vocabulary recognition by loading pre-set banking-specific dictionaries (such as a wealth management product terminology library and a compliance terminology dictionary). Typos in the ASR results are automatically corrected based on contextual semantics, generating the corresponding text data for the conversations to be inspected.
[0053] In this embodiment, during the inference phase (e.g., the decoding phase), the probability distribution of the language model can be adjusted to prioritize matching words in a preset banking-specific dictionary, thereby correcting errors in the text data to be inspected and determining the corresponding text data for the conversation to be inspected. A higher weight can also be assigned to words in the preset banking-specific dictionary, making them more likely to be selected during decoding.
[0054] S130: Perform multimodal feature extraction and voice quality inspection based on the first preset instruction, the preset voice quality inspection model, the voice data, and the text data to determine a first quality inspection result of the conversation to be inspected.
[0055] The first preset instruction may refer to a prompt for the preset voice quality inspection model to perform operations on voice data and text data. For example, the first preset instruction may be an end-to-end operation instruction, enabling the large model to implement end-to-end processing from semantic parsing to problem location by building a fully automated quality inspection pipeline, effectively replacing traditional manual rule configuration and voice quality inspection processes. The preset voice quality inspection model may be a large language model. The preset voice quality inspection model may be obtained by efficiently fine-tuning the parameters of the large language model using a sample library. The sample library may include training samples labeled with the presence or absence of violations, such as customer service conversations and complaint recordings. In this embodiment, the sample library can be updated after correcting misjudgments in the violation results marked by manual review, thereby periodically fine-tuning the preset voice quality inspection model using the sample library. The first quality inspection result may be the quality inspection result determined by the preset voice quality inspection model based on the non-professional content of the conversation to determine whether there is a violation. For example, the first quality inspection result may include, but is not limited to, determining whether there is a violation based on the bank service employee's tone and attitude, as well as whether the answer is irrelevant to the question.
[0056] For example, efficient fine-tuning of the parameters of a pre-set speech quality inspection model can be achieved using a sample library and a newly added low-rank matrix, allowing the fine-tuned model to adapt to the semantic understanding and compliance judgment requirements of quality inspection scenarios. Specifically, only the newly added low-rank matrix can be trained, freezing most of the original model parameters. For example, an adaptation layer can be inserted into a large language model, and the weights can be updated using quality inspection conversation data (sample library).
[0057] In this embodiment, the first preset instruction may include: extracting emotional and speech rate features from the input voice data; extracting banking-related keywords from the input text data and understanding the customer's intent through semantic analysis of the text data; and using the extracted emotional, speech rate, keywords, and intent to determine whether there are any violations in the conversation. The first preset instruction gradually guides the preset voice quality inspection model to perform specific operations on the voice and text data, thereby obtaining a first quality inspection result for the conversation to be inspected.
[0058] In this embodiment, the input of the preset voice quality inspection model may also include operation examples. The operation examples can be used to enable the preset voice quality inspection model to understand how to correctly execute the first preset instruction. The operation examples may include: emotion feature extraction examples, speech rate feature extraction examples, keyword extraction examples, intention extraction examples, and examples of quality inspection using multimodal features. The more examples there are, the more accurate the features extracted by the model. For example, the emotion feature extraction examples include: voice data samples, feature extraction results for voice data samples, and emotion feature labels for feature extraction results. On this basis, each step in the first preset instruction is added with "perform operations (such as extraction and quality inspection) with reference to the provided XX examples."
[0059] S140: Perform quality inspection based on the second preset instruction, text data, a preset voice quality inspection model, an industry knowledge base, and an external policy library to determine a second quality inspection result of the session to be quality inspected.
[0060] Among them, the second preset instruction may refer to the prompt words of the preset voice quality inspection model for the operation on text data, industry knowledge base and external policy base. Similar to the first preset instruction, the second preset instruction can be input into the preset voice quality inspection model together with the example, or it can be input into the preset voice quality inspection model separately. The industry knowledge base may refer to a knowledge base that stores business service processes and reported speech rules and is composed of a vector database. The external policy library may refer to a database composed of the latest compliance requirements issued by an external regulatory platform (such as the China Securities Association policy library in the financial industry). In this embodiment, the external policy library can be connected through the API interface, and the latest compliance requirements can be dynamically obtained. The second quality inspection result may refer to the quality inspection result of whether there is any violation in the conversation determined by the preset voice quality inspection model based on the professional content in the conversation. For example, the second quality inspection result can be to judge whether there is any violation based on the accuracy of the answer of the bank service staff.
[0061] Specifically, the preset voice quality inspection model performs a quality inspection based on the second preset instruction, combining the industry knowledge base and the external policy database, to produce a second quality inspection result. For example, when a customer service representative mentions "loan interest discounts" in the industry knowledge base, the model searches the external policy database for the latest Loan Prime Rate (LPR) requirements of the central bank to determine whether the customer service representative's statement contains any misleading statements. If so, no violation has occurred.
[0062] S150: Determine a target quality inspection result for the session to be quality inspected based on the first quality inspection result and the second quality inspection result.
[0063] The target quality inspection result may refer to the final quality inspection result that combines the first quality inspection result and the second quality inspection result. The target quality inspection result may include the specific conversation in which the violation occurred and the specific violation. Specifically, if at least one of the first quality inspection result and the second quality inspection result contains a violation, the target quality inspection result also contains a violation. The first quality inspection result and the second quality inspection result are merged, and each violation is sorted in descending order of priority based on the preset violation level to obtain the target quality inspection result of the conversation to be inspected.
[0064] The technical solution of the embodiment of the present invention obtains the voice signal of the conversation to be quality inspected, and performs voice track separation on the voice signal to determine the voice data containing the identity identifier; performs format conversion on the voice data based on a preset voice recognition model to determine the corresponding text data of the conversation to be quality inspected; performs multimodal feature extraction and voice quality inspection based on a first preset instruction, a preset voice quality inspection model, voice data and text data to determine a first quality inspection result of the conversation to be quality inspected; performs quality inspection based on a second preset instruction, text data, a preset voice quality inspection model, an industry knowledge base and an external policy base to determine a second quality inspection result of the conversation to be quality inspected; determines a target quality inspection result of the conversation to be quality inspected based on the first quality inspection result and the second quality inspection result, so that a comprehensive quality inspection can be performed on the conversation to be quality inspected in a timely manner after the call is ended, and a timeliness quality inspection can be performed on the policy content involved in the conversation to be quality inspected without the need for manual quality inspection, thereby improving the efficiency and accuracy of voice quality inspection.
[0065] Example 2
[0066] Figure 2 This is a flowchart of a speech quality inspection method provided in the second embodiment of the present invention. This embodiment describes in detail the process of determining the first quality inspection result based on the above embodiment. The explanations of the terms that are the same or corresponding to the above embodiments are not repeated here. Figure 2 As shown, the method includes:
[0067] S210: Acquire a voice signal of the conversation to be quality-checked, and perform voice track separation on the voice signal to determine voice data containing an identity identifier.
[0068] S220: Convert the format of the voice data based on a preset voice recognition model to determine the corresponding text data of the conversation to be quality inspected.
[0069] Exemplarily, the text data after format conversion is pre-processed by cleaning, word segmentation, and stop word removal to facilitate the subsequent processing of the large model. Specifically, the noise in the text data is cleaned and removed, such as extra spaces, tabs, line breaks, and invisible special characters. The cleaned text data is divided into independent vocabulary units according to preset rules (such as spaces, punctuation marks, etc.), that is, word segmentation, and the vocabulary units are used to form the text data input into the preset speech quality inspection model.
[0070] S230: Input the first preset instruction, voice data, and text data into a preset voice quality inspection model.
[0071] In this embodiment, the preset voice quality inspection model can implement the operations of each step in the preset instruction by calling an external model in a hybrid solution, thereby utilizing a dedicated small model to perform voice quality inspection, further improving the accuracy of voice quality inspection. The external model can be pre-trained using bank sample data, and the training process can also be implemented using the preset voice quality inspection model.
[0072] S240: In a preset voice quality inspection model, feature extraction is performed on the voice data based on the first operation instruction in the first preset instruction to determine the emotional features and speech speed features corresponding to each dialogue in the conversation to be inspected.
[0073] The first operation instruction is used to instruct the preset voice quality inspection model to call an external emotion feature extraction model and an external speech rate feature extraction model to extract features from the voice data. Emotional features can refer to the emotion type of each character in each conversation in the conversation to be inspected. A conversation is a question and answer between a customer and a bank representative. Emotional features can include, but are not limited to, happiness or anger. Speech rate features can refer to the average speaking speed of each character in each conversation. For example, the speech rate feature can be 1x speed or 1.2x speed.
[0074] Specifically, in the preset voice quality inspection model, the preset voice quality inspection model is directed to perform an action based on the first operation instruction in the first preset instruction, such as calling an external emotion feature extraction model to perform an emotion feature extraction operation on the input voice data, and calling an external speech rate feature extraction model to perform a speech rate feature extraction operation on the input voice data. Based on the output results of the external emotion feature extraction model, the emotion features corresponding to each dialogue in the conversation to be inspected are determined. Based on the output results of the external speech rate feature extraction model, the speech rate features corresponding to each dialogue in the conversation to be inspected are determined.
[0075] S250: Perform feature extraction on the text data based on the second operation instruction in the first preset instruction to determine the keywords and intentions corresponding to each dialogue in the conversation to be inspected.
[0076] Keywords can refer to words related to banking services. For example, keywords include but are not limited to "account" and "deposit." Intent can refer to the intention of each role in each conversation. For example, in a conversation, the customer's intention is to inquire about the deposit method with the highest interest rate, while the customer service representative's intention is to discuss financial management methods with high interest rates but also high risks.
[0077] Specifically, in the preset voice quality inspection model, the second operation instruction in the first preset instruction instructs the preset voice quality inspection model to perform an action, such as calling an external keyword extraction model to extract keywords related to banking services from the input text data, and calling an external intent recognition model to recognize intents on the input text data. Based on the output of the external keyword extraction model, the keywords corresponding to each conversation in the conversation to be inspected are determined. Based on the output of the external intent recognition model, the intent corresponding to each conversation in the conversation to be inspected is determined.
[0078] S260: Perform voice quality inspection based on the first preset instruction and the emotional features, speech speed features, keywords, and intentions corresponding to each dialogue in the conversation to be quality inspected, and determine a first quality inspection result of the conversation to be quality inspected.
[0079] Specifically, in the preset speech quality inspection model, an external multimodal speech quality inspection model is called based on the first preset instruction, and the called external multimodal speech quality inspection model is used to perform comprehensive speech quality inspection in combination with the emotional characteristics, speaking speed characteristics, keywords and intentions corresponding to each dialogue in the conversation to be inspected to determine the first quality inspection result of the conversation to be inspected.
[0080] On the basis of the above technical solution, "performing voice quality inspection based on the first preset instruction and the emotional characteristics, speech speed characteristics, keywords and intentions corresponding to each dialogue in the conversation to be quality inspected, and determining the first quality inspection result of the conversation to be quality inspected" may include: based on the third operation instruction in the first preset instruction, feature-correlating the emotional characteristics, speech speed characteristics, keywords and intentions corresponding to each dialogue to determine the conversation characteristics of each dialogue, and determining the conversation status of each dialogue based on the conversation characteristics of each dialogue; based on the fourth operation instruction in the first preset instruction and the conversation status, performing progressive quality inspection on each dialogue to determine the first quality inspection result of the conversation to be quality inspected.
[0081] Among them, the third operation instruction can be used to guide the preset voice quality inspection model to call the external cross-modal model, and use the external cross-modal model to associate multimodal features, and determine the conversation state based on the conversation features. Conversation features may refer to features obtained after the multimodal features are associated. The conversation state may refer to the result of the structured integration of conversation features. The conversation state can be used to represent a "snapshot" or "portrait" of the conversation. For example, high speaking speed + angry emotion + keyword "fault" → the customer is likely to be in an "urgent repair" state. The fourth operation instruction can be used to guide the preset voice quality inspection model to perform quality inspection operations on similar conversations together.
[0082] Specifically, in the preset voice quality inspection model, an external cross-modal model is called based on the third operation instruction in the first preset instruction, and the external cross-modal model is used to perform feature association of the emotional features, speech rate features, keywords and intentions corresponding to each conversation to determine the conversation features of each conversation. Based on the third operation instruction in the first preset instruction, the conversation features of each conversation are probabilistically classified, and the classification result with the highest probability is determined as the conversation state of each conversation. The probabilistic classification operation can also use a pre-trained classification model. In this embodiment, the associated features are mapped to structured states through rules or models (such as state machines, deep learning). For example, input: emotion = negative, intention = complaint, keyword = "delay" → output status label "complaint in progress". Based on the fourth operation instruction in the first preset instruction, similar conversations are subjected to quality inspection operations together, that is, progressive quality inspection, so as to determine the first quality inspection result of the conversation to be quality inspected.
[0083] On the basis of the above technical solution, "performing progressive quality inspection on each conversation based on the fourth operation instruction in the first preset instruction and the conversation state to determine the first quality inspection result of the conversation to be quality inspected" may include: performing single conversation quality inspection on each conversation based on the fourth operation instruction in the first preset instruction to determine the individual quality inspection result of each conversation; performing correlation calculation on each conversation state to determine the semantic correlation between each two conversations, and determining a conversation set for joint quality inspection based on the preset correlation threshold and each semantic correlation; performing conversation joint quality inspection on the conversation set to determine the joint quality inspection result of the conversation set; and determining the first quality inspection result of the conversation to be quality inspected based on the individual quality inspection results and the joint quality inspection result.
[0084] The individual quality inspection result may refer to the quality inspection result of a single conversation to determine whether there are any violations. Semantic relevance may refer to the similarity between the content of two conversations, i.e., semantic similarity. The preset relevance threshold may refer to the minimum similarity threshold for determining that the content of two conversations is similar. There must be at least two individual conversations in the conversation set. The joint quality inspection result may refer to the quality inspection result of the conversations as a whole.
[0085] Specifically, a single conversation quality inspection is performed on each conversation based on the fourth operation instruction in the first preset instruction, and an individual quality inspection result is determined for each conversation. For example, in a scenario where the intended purpose is to make a complaint, customer service anger is determined to be a violation. The basis for determining whether a violation exists in a single conversation is the sample library and the examples entered with the instruction. Based on different business needs and different judgment criteria, users can customize the sample library and examples so that the preset voice quality inspection model can perform conversation quality inspection at different scales to meet user needs. A correlation is calculated for each conversation state to determine the semantic correlation between each two conversations. In this embodiment, the conversation states and correlation features of all conversations can be represented using a conversation feature graph, allowing the semantic correlation between each conversation to be calculated using a graph neural network. Based on a preset correlation threshold and the individual semantic correlations, conversations with semantic correlations greater than the preset correlation threshold are combined to form a conversation set for joint quality inspection. The conversations in the conversation set are then quality inspected as a whole to determine the joint quality inspection result for the conversation set. The individual quality inspection results and the joint quality inspection result are merged to determine the first quality inspection result for the conversation to be inspected.
[0086] It should be noted that the progressive quality inspection method in this embodiment enables the large language model to capture key information nodes across dozens of rounds of conversations, effectively identify deep-seated problems such as logical breaks, topic jumping, and information omissions in the response process, and further improve the large language model's contextual understanding ability and voice quality inspection accuracy during the voice quality inspection process.
[0087] S270: Perform quality inspection based on the second preset instruction, text data, preset voice quality inspection model, industry knowledge base, and external policy library to determine a second quality inspection result of the session to be quality inspected.
[0088] On the basis of the above technical solution, "performing quality inspection based on the second preset instruction, text data, preset voice quality inspection model, industry knowledge base and external policy library to determine the second quality inspection result of the session to be quality inspected" may include: inputting the second preset instruction, text data, industry knowledge base and external policy library into the preset voice quality inspection model; in the preset voice quality inspection model, based on the second preset instruction and combined with the industry knowledge base, performing information identification on the text data to determine the industry-specific professional information and the limiting information corresponding to the professional information in the session to be quality inspected; searching the professional information in combination with the external policy library to determine the standard information corresponding to the professional information; performing quality inspection on the limiting information and the standard information to determine the second quality inspection result of the session to be quality inspected.
[0089] Specialized information can refer to information with specific limitations within an industry. Qualified information can refer to qualified information regarding specialized information provided by bank personnel during a conversation. Standard information can refer to standard information describing specialized information. For example, if the specialized information is loan interest discounts, qualified information can refer to content related to loan interest discounts provided by bank personnel during the conversation, such as explanatory content regarding loan interest discount policies. Standard information can refer to standard descriptions of loan interest discounts from an external policy repository, such as the original text of the loan interest discount policy.
[0090] Specifically, the second preset instruction, text data, industry knowledge base, and external policy database are input into a preset voice quality inspection model. Within the preset voice quality inspection model, information recognition is performed on the text data based on the second preset instruction and the industry knowledge base. When the conversation mentions "loan interest discounts" (equivalent to specialized information) from the industry knowledge base, the external policy database is searched for the latest Loan Prime Rate (LPR) requirements (equivalent to standard information) from the central bank. This information is then used to determine whether the relevant content provided by the customer service representative (equivalent to qualified information) contains any misleading statements. If so, no violation has occurred.
[0091] It should be noted that in this embodiment, the large language model is called twice to perform two quality inspection results determination processes respectively. Through deep semantic understanding, context association analysis and dynamic knowledge enhancement, complex speech and cross-round violations can be accurately identified, manual processing can be reduced, and the unity of high accuracy and high efficiency can be achieved.
[0092] S280: Determine a target quality inspection result for the session to be quality inspected based on the first quality inspection result and the second quality inspection result.
[0093] The technical solution of the embodiment of the present invention is as follows: a first preset instruction, voice data and text data are input into a preset voice quality inspection model; in the preset voice quality inspection model, feature extraction is performed on the voice data based on the first operation instruction in the first preset instruction, and the emotional features and speech speed features corresponding to each dialogue in the session to be inspected are determined; feature extraction is performed on the text data based on the second operation instruction in the first preset instruction, and the keywords and intentions corresponding to each dialogue in the session to be inspected are determined; voice quality inspection is performed based on the first preset instruction and the emotional features, speech speed features, keywords and intentions corresponding to each dialogue in the session to be inspected, and the first quality inspection result of the session to be inspected is determined, so that the dialogue content quality inspection can be comprehensively performed on each dialogue in the session to be inspected from multiple aspects, thereby further improving the accuracy of voice quality inspection.
[0094] The following is an embodiment of a speech quality inspection device provided by an embodiment of the present invention. The device and the speech quality inspection methods of the above embodiments belong to the same inventive concept. For details not fully described in the embodiment of the speech quality inspection device, please refer to the embodiment of the above speech quality inspection method.
[0095] Example 3
[0096] Figure 3 This is a structural diagram of a speech quality inspection device provided by the third embodiment of the present invention. Figure 3 As shown, the apparatus includes: a voice data determination module 310 , a text data determination module 320 , a first quality inspection result determination module 330 , a second quality inspection result determination module 340 and a target quality inspection result determination module 350 .
[0097] Among them, the voice data determination module 310 is used to obtain the voice signal of the conversation to be quality inspected, and perform voice track separation on the voice signal to determine the voice data containing the identity identifier; the text data determination module 320 is used to convert the format of the voice data based on the preset voice recognition model to determine the corresponding text data of the conversation to be quality inspected; the first quality inspection result determination module 330 is used to perform multimodal feature extraction and voice quality inspection based on the first preset instruction, the preset voice quality inspection model, the voice data and the text data to determine the first quality inspection result of the conversation to be quality inspected; the second quality inspection result determination module 340 is used to perform quality inspection based on the second preset instruction, the text data, the preset voice quality inspection model, the industry knowledge base and the external policy base to determine the second quality inspection result of the conversation to be quality inspected; the target quality inspection result determination module 350 is used to determine the target quality inspection result of the conversation to be quality inspected based on the first quality inspection result and the second quality inspection result.
[0098] The technical solution of the embodiment of the present invention obtains the voice signal of the conversation to be quality inspected, and performs voice track separation on the voice signal to determine the voice data containing the identity identifier; performs format conversion on the voice data based on a preset voice recognition model to determine the corresponding text data of the conversation to be quality inspected; performs multimodal feature extraction and voice quality inspection based on a first preset instruction, a preset voice quality inspection model, voice data and text data to determine a first quality inspection result of the conversation to be quality inspected; performs quality inspection based on a second preset instruction, text data, a preset voice quality inspection model, an industry knowledge base and an external policy base to determine a second quality inspection result of the conversation to be quality inspected; determines a target quality inspection result of the conversation to be quality inspected based on the first quality inspection result and the second quality inspection result, so that a comprehensive quality inspection can be performed on the conversation to be quality inspected in a timely manner after the call is ended, and a timeliness quality inspection can be performed on the policy content involved in the conversation to be quality inspected without the need for manual quality inspection, thereby improving the efficiency and accuracy of voice quality inspection.
[0099] Based on the above technical solution, the voice data determination module 310 is specifically used to: perform noise reduction processing on the voice signal for steady-state noise and non-steady-state noise based on a preset environmental noise reduction method, and determine the voice signal that retains a clear human voice; perform voice track segmentation and role labeling on the voice signal that retains a clear human voice based on voiceprint recognition technology, and determine the voice data containing identity identification.
[0100] On the basis of the above technical solution, the preset speech recognition model is constructed by loading a preset banking business-specific dictionary on the basis of the general speech recognition model.
[0101] The text data determination module 320 is specifically used to: perform format conversion on the speech data based on a preset speech recognition model to determine the text data to be detected; perform error correction and modification on the text data to be detected based on context semantics to determine the corresponding text data of the conversation to be quality checked.
[0102] Based on the above technical solution, the first quality inspection result determination module 330 may include:
[0103] A data input submodule, configured to input the first preset instruction, voice data, and text data into a preset voice quality inspection model;
[0104] A first feature extraction submodule is configured to perform feature extraction on the voice data based on the first operation instruction in the first preset instruction in a preset voice quality inspection model, and determine the emotional features and speech rate features corresponding to each dialogue in the conversation to be inspected;
[0105] A second feature extraction submodule is used to extract features from the text data based on the second operation instruction in the first preset instruction, and determine the keywords and intentions corresponding to each dialogue in the conversation to be inspected;
[0106] The first quality inspection result determination submodule is used to perform voice quality inspection based on the first preset instruction and the emotional characteristics, speech speed characteristics, keywords and intentions corresponding to each dialogue in the conversation to be inspected, and determine the first quality inspection result of the conversation to be inspected.
[0107] Based on the above technical solution, the first quality inspection result determination submodule may include:
[0108] a conversation state determining unit, configured to perform feature correlation based on the third operation instruction in the first preset instruction with the emotion feature, speech rate feature, keyword, and intention corresponding to each conversation to determine the conversation feature of each conversation, and determine the conversation state of each conversation based on the conversation feature of each conversation;
[0109] The first quality inspection result determining unit is configured to perform a progressive quality inspection on each dialogue based on the fourth operation instruction in the first preset instruction and the dialogue state, and determine a first quality inspection result of the dialogue to be quality inspected.
[0110] Based on the above technical solution, the first quality inspection result determination unit is specifically used to: perform single conversation quality inspection on each conversation based on the fourth operation instruction in the first preset instruction, and determine the individual quality inspection result of each conversation; perform correlation calculation on each conversation state, determine the semantic correlation between each two conversations, and determine the conversation set for joint quality inspection based on the preset correlation threshold and each semantic correlation; perform conversation joint quality inspection on the conversation set, and determine the joint quality inspection result of the conversation set; determine the first quality inspection result of the conversation to be quality inspected based on the individual quality inspection result and the joint quality inspection result.
[0111] Based on the above technical solution, the second quality inspection result determination module 340 is specifically used to: input the second preset instruction, text data, industry knowledge base and external policy library into the preset voice quality inspection model; in the preset voice quality inspection model, based on the second preset instruction and combined with the industry knowledge base, the text data is identified to determine the industry-specific professional information and the limiting information corresponding to the professional information in the conversation to be quality inspected; the professional information is retrieved in combination with the external policy library to determine the standard information corresponding to the professional information; the limiting information and the standard information are quality inspected to determine the second quality inspection result of the conversation to be quality inspected.
[0112] The speech quality inspection device provided in the embodiment of the present invention can execute the speech quality inspection method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of executing the speech quality inspection method.
[0113] It is worth noting that in the above-mentioned embodiment of voice quality inspection, the various units and modules included are only divided according to functional logic, but are not limited to the above-mentioned division, as long as the corresponding functions can be achieved; in addition, the specific names of the functional units are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of the present invention.
[0114] Example 4
[0115] Figure 4 A schematic diagram of the structure of an electronic device 10 that can be used to implement an embodiment of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or claimed herein.
[0116] like Figure 4As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc., which is communicatively connected to the at least one processor 11. The memory stores a computer program that can be executed by the at least one processor. The processor 11 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 into the random access memory (RAM) 13. Various programs and data required for the operation of the electronic device 10 can also be stored in the RAM 13. The processor 11, ROM 12, and RAM 13 are connected to each other via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0117] Multiple components in the electronic device 10 are connected to the I / O interface 15, including an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a magnetic disk, an optical disk, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0118] The processor 11 can be any general-purpose and / or specialized processing component with processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The processor 11 executes the various methods and processes described above, such as the voice quality inspection method.
[0119] In some embodiments, the speech quality inspection method can be implemented as a computer program, which is tangibly contained in a computer-readable storage medium, such as the storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed on the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded into the RAM 13 and executed by the processor 11, one or more steps of the speech quality inspection method described above can be performed. Alternatively, in other embodiments, the processor 11 can be configured to execute the speech quality inspection method in any other appropriate manner (for example, by means of firmware).
[0120] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0121] Computer programs for implementing the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the computer program is executed by the processor, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The computer program may be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0122] In the context of the present invention, computer-readable storage media can be tangible media that can contain or store a computer program for use with an instruction execution system, device or equipment or used in combination with an instruction execution system, device or equipment. Computer-readable storage media can include but are not limited to electronic, magnetic, optical, electromagnetic, infrared or semiconductor systems, devices or equipment, or any suitable combination of the foregoing. Alternatively, computer-readable storage media can be machine-readable signal media. More specific examples of machine-readable storage media can include electrical connections based on one or more lines, portable computer disks, hard disks, random access memories (RAM), read-only memories (ROM), erasable programmable read-only memories (EPROM or flash memory), optical fibers, portable compact disk read-only memories (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0123] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0124] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.
[0125] A computing system may include clients and servers. The clients and servers are typically remote from each other and typically interact via a communication network. This client-server relationship arises through computer programs running on the respective computers, creating a client-server relationship. The server may be a cloud server, also known as a cloud computing server or cloud host. This server is a hosting product within the cloud computing service ecosystem that addresses the management difficulties and limited scalability of traditional physical hosting and VPS services.
[0126] An embodiment of the present invention further provides a computer program product, including a computer program, which, when executed by a processor, implements the speech quality inspection method provided in any embodiment of the present application.
[0127] During the implementation of the computer program product, the computer program code for performing the operation of the present invention can be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, C++, and also conventional procedural programming languages such as "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as an independent software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (for example, using an Internet service provider to connect through the Internet). The program product and the speech quality inspection method disclosed in each embodiment of the present application belong to the same inventive concept, so they are not described here.
[0128] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in the present invention can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present invention can be achieved. This is not limited herein.
[0129] The above specific embodiments do not limit the scope of protection of the present invention. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention are intended to be included within the scope of protection of the present invention.
Claims
1. A voice quality inspection method, characterized in that: include: Acquire a voice signal of the conversation to be quality inspected, and perform voice track segmentation on the voice signal to determine voice data containing the identity identifier; Performing format conversion on the voice data based on a preset voice recognition model to determine the text data corresponding to the conversation to be inspected; Performing multimodal feature extraction and voice quality inspection based on the first preset instruction, the preset voice quality inspection model, the voice data, and the text data to determine a first quality inspection result of the conversation to be inspected; Performing a quality inspection based on a second preset instruction, the text data, the preset voice quality inspection model, an industry knowledge base, and an external policy library to determine a second quality inspection result of the session to be quality inspected; A target quality inspection result of the session to be quality inspected is determined based on the first quality inspection result and the second quality inspection result.
2. The method according to claim 1, characterized in that The performing voice track separation on the voice signal to determine the voice data containing the identity identifier includes: Performing noise reduction processing on the speech signal based on a preset environmental noise reduction method for steady-state noise and non-steady-state noise, and determining a speech signal that retains a clear human voice; Based on voiceprint recognition technology, the voice signal retaining clear human voice is subjected to voice track segmentation and character labeling to determine voice data containing identity identification.
3. The method according to claim 1, characterized in that The preset speech recognition model is constructed by loading a preset banking business-specific dictionary on the basis of a general speech recognition model; The converting the format of the voice data based on a preset voice recognition model to determine the text data corresponding to the conversation to be inspected includes: Performing format conversion on the voice data based on a preset voice recognition model to determine text data to be detected; Error correction and modification are performed on the text data to be detected based on context semantics to determine the corresponding text data of the session to be quality checked.
4. The method according to claim 1, wherein The performing multimodal feature extraction and voice quality inspection based on the first preset instruction, the preset voice quality inspection model, the voice data, and the text data to determine a first quality inspection result of the conversation to be inspected includes: Inputting the first preset instruction, the voice data and the text data into a preset voice quality inspection model; In the preset voice quality inspection model, feature extraction is performed on the voice data based on the first operation instruction in the first preset instruction to determine the emotional features and speech speed features corresponding to each dialogue in the conversation to be inspected; Performing feature extraction on the text data based on the second operation instruction in the first preset instruction to determine the keywords and intentions corresponding to each dialogue in the conversation to be inspected; Based on the first preset instruction and the emotional characteristics, speech speed characteristics, keywords and intentions corresponding to each dialogue in the conversation to be quality checked, a voice quality check is performed to determine a first quality check result of the conversation to be quality checked.
5. The method according to claim 4, characterized in that The performing voice quality inspection based on the first preset instruction and the emotion characteristics, speech speed characteristics, keywords, and intention corresponding to each dialogue in the conversation to be quality inspected, and determining a first quality inspection result of the conversation to be quality inspected, includes: Based on the third operation instruction in the first preset instruction, the emotional characteristics, speech speed characteristics, keywords and intentions corresponding to each dialogue are feature-correlated to determine the dialogue characteristics of each dialogue, and the dialogue state of each dialogue is determined according to the dialogue characteristics of each dialogue; Based on the fourth operation instruction in the first preset instruction and the dialog state, progressive quality inspection is performed on each dialog to determine a first quality inspection result of the dialog to be quality inspected.
6. The method according to claim 5, characterized in that The performing progressive quality inspection on each dialogue based on the fourth operation instruction in the first preset instruction and the dialogue state to determine a first quality inspection result of the dialogue to be inspected includes: Performing a single-conversation quality inspection on each conversation based on the fourth operation instruction in the first preset instruction, and determining a separate quality inspection result for each conversation; Calculate the correlation between each conversation state to determine the semantic correlation between each two conversations, and determine the conversation set for joint quality inspection based on the preset correlation threshold and each semantic correlation; Performing a joint quality inspection on the conversation set to determine a joint quality inspection result of the conversation set; A first quality inspection result of the session to be quality inspected is determined based on the individual quality inspection result and the joint quality inspection result.
7. The method according to claim 1, characterized in that The performing of quality inspection based on the second preset instruction, the text data, the preset voice quality inspection model, the industry knowledge base, and the external policy library to determine a second quality inspection result of the session to be quality inspected includes: Inputting the second preset instruction, the text data, the industry knowledge base, and the external policy library into the preset voice quality inspection model; In the preset voice quality inspection model, information recognition is performed on the text data based on the second preset instruction in combination with the industry knowledge base to determine the industry-specific professional information in the conversation to be inspected and the limiting information corresponding to the professional information; Searching the professional information in combination with the external policy database to determine the standard information corresponding to the professional information; Performing quality inspection on the limited information and the standard information to determine a second quality inspection result of the session to be quality inspected.
8. A speech quality inspection device, characterized in that: The device comprises: A voice data determination module, configured to obtain a voice signal of a conversation to be inspected, and perform voice track separation on the voice signal to determine voice data containing an identity identifier; A text data determination module, configured to convert the format of the voice data based on a preset voice recognition model to determine the text data corresponding to the conversation to be inspected; a first quality inspection result determination module, configured to perform multimodal feature extraction and voice quality inspection based on a first preset instruction, a preset voice quality inspection model, the voice data, and the text data, and determine a first quality inspection result of the conversation to be inspected; A second quality inspection result determination module, configured to perform quality inspection based on a second preset instruction, the text data, the preset voice quality inspection model, the industry knowledge base, and the external policy library, and determine a second quality inspection result of the to-be-inspected session; A target quality inspection result determination module is configured to determine a target quality inspection result of the session to be quality inspected based on the first quality inspection result and the second quality inspection result.
9. An electronic device, characterized in that: The electronic device comprises: one or more processors; a memory for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the speech quality inspection method according to any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the speech quality inspection method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Intelligent voice customer service quality inspection method and device based on multi-modal large model
CN118631939A
Customer service verbal skill quality inspection result determination method and device, equipment, medium and product
CN119294885A
Multi-language audio quality inspection method and system based on voice processing and large language model
CN119889355A
Information processing method, computer equipment and computer program product
CN119903851A
Quality assurance systems based on speaker intent detection
WO2024163482A1
Cited By
Dialogue abnormity real-time detection method and system based on multi-dimensional feature fusion
CN121905211A