A voice quality inspection method and device, electronic equipment and storage medium

By performing track-by-track and text conversion on the speech signal, combined with multimodal feature extraction and policy database quality inspection, the problem of time-consuming and labor-intensive manual quality inspection is solved, and efficient and accurate speech quality inspection is achieved.

CN120496574BActive Publication Date: 2026-02-10JIANGSU JIANGNAN RURAL COMMERCIAL BANK CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510783473.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-12
Publication Date
2026-02-10
Estimated Expiration
2045-06-12

AI Technical Summary

Technical Problem

In existing technologies, manual quality inspection is time-consuming and labor-intensive, and cannot fully cover call voice, resulting in inaccurate voice quality inspection and low coverage, with the problem of missing illegal dialogues.

Method used

By acquiring the voice signal of the conversation to be inspected, voice tracking and identification are performed. The voice is converted into text data using a preset voice recognition model. Multimodal feature extraction and quality inspection are performed in conjunction with a preset voice quality inspection model. Comprehensive quality inspection is carried out in conjunction with industry knowledge base and external policy base to determine the target quality inspection result.

Benefits of technology

It enables comprehensive and timely quality inspection of pending call sessions that have ended, improving the efficiency and accuracy of voice quality inspection and reducing the need for manual quality inspection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120496574B_ABST
    Figure CN120496574B_ABST
Patent Text Reader

Abstract

Embodiments of the present application disclose a kind of voice quality inspection method, device, electronic equipment and storage medium involved.The method comprises: obtaining the voice signal of the session to be inspected, and the voice signal is carried out speech track determination and contains the voice data of identity identification;Based on the preset speech recognition model, the format conversion is carried out to voice data, and the corresponding text data of the session to be inspected is determined;Based on the first preset instruction, the preset voice quality inspection model, voice data and text data are carried out multi-modal feature extraction and voice quality inspection, and the first inspection result of the session to be inspected is determined;Based on the second preset instruction, text data, the preset voice quality inspection model, industry knowledge base and external policy base are carried out inspection, and the second inspection result of the session to be inspected is determined;Based on the first inspection result and the second inspection result, the target inspection result of the session to be inspected is determined, so that the session to be inspected of ending call can be comprehensively inspected in time, and then the efficiency and accuracy of voice quality inspection are improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] Embodiments of the present application relate to the technical field of artificial intelligence, and particularly relate to a voice quality inspection method and device, electronic equipment and storage medium. BACKGROUND

[0002] With the development of artificial intelligence, remote voice communication is widely used in various fields to serve customers. For example, service personnel in the banking field answer questions, recommend products and provide sales services to customers through telephone. However, as the number of customers increases, the number of communication voices between service personnel and customers also increases. In order to further find service problems, avoid risks and optimize business processes, voice quality inspection is used to monitor and detect the communication voice.

[0003] At present, the communication voice is usually inspected by artificial quality inspection. However, this artificial quality inspection method is time-consuming and laborious, and cannot pay attention to all aspects that need to be inspected, or only pays attention to one aspect, such as whether the reply content of the service personnel is illegal, resulting in that the voice quality inspection is not comprehensive and the quality inspection result is not accurate. Moreover, this artificial quality inspection method usually uses sampling inspection, resulting in low quality inspection coverage and missing illegal conversation. SUMMARY

[0004] Embodiments of the present application provide a voice quality inspection method, device, electronic equipment and storage medium, which can comprehensively inspect the to-be-inspected conversation that ends the communication in time, and perform time-sensitive quality inspection on the policy content involved in the to-be-inspected conversation, without artificial quality inspection, thereby improving the efficiency and accuracy of voice quality inspection.

[0005] In a first aspect, embodiments of the present application provide a voice quality inspection method, comprising:

[0006] obtaining a voice signal of a to-be-inspected conversation, and determining voice data containing an identity identifier by performing voice track splitting on the voice signal;

[0007] performing format conversion on the voice data based on a preset voice recognition model to determine corresponding text data of the to-be-inspected conversation;

[0008] performing multi-modal feature extraction and voice quality inspection based on a first preset instruction, a preset voice quality inspection model, the voice data and the text data to determine a first quality inspection result of the to-be-inspected conversation;

[0009] performing quality inspection based on a second preset instruction, the text data, the preset voice quality inspection model, an industry knowledge base and an external policy base to determine a second quality inspection result of the to-be-inspected conversation;

[0010] determining a target quality inspection result of the to-be-inspected conversation based on the first quality inspection result and the second quality inspection result.

[0011] Optionally, the method further comprises: performing noise reduction processing on the voice signal based on a preset environmental noise reduction mode to determine a voice signal with clear human voice reserved; performing voice track separation and role labeling on the voice signal with clear human voice reserved based on a voiceprint recognition technology to determine voice data containing an identity label.

[0012] Optionally, the method further comprises: the preset voice recognition model is constructed by loading a preset bank business exclusive dictionary on the basis of a general voice recognition model; performing format conversion on the voice data based on the preset voice recognition model to determine text data to be detected; performing error correction and modification on the text data to be detected based on context semantics to determine corresponding text data of the conversation to be quality inspected.

[0013] Optionally, the method further comprises: inputting a first preset instruction, the voice data and the text data into a preset voice quality inspection model; in the preset voice quality inspection model, performing feature extraction on the voice data based on a first operation instruction in the first preset instruction to determine emotion features and speech rate features corresponding to each dialogue in the conversation to be quality inspected; performing feature extraction on the text data based on a second operation instruction in the first preset instruction to determine keywords and intents corresponding to each dialogue in the conversation to be quality inspected; performing voice quality inspection based on the first preset instruction and the emotion features, the speech rate features, the keywords and the intents corresponding to each dialogue in the conversation to be quality inspected to determine a first quality inspection result of the conversation to be quality inspected.

[0014] Optionally, the method further comprises: based on a third operation instruction in the first preset instruction, associating the emotion features, the speech rate features, the keywords and the intents corresponding to each dialogue to determine dialogue features of each dialogue, and determining dialogue states of each dialogue according to the dialogue features of each dialogue; based on a fourth operation instruction in the first preset instruction and the dialogue states, performing progressive quality inspection on each dialogue to determine the first quality inspection result of the conversation to be quality inspected.

[0015] Optionally, the method further comprises: based on the fourth operation instruction in the first preset instruction, performing single-dialogue quality inspection on each dialogue to determine individual quality inspection results of each dialogue; performing correlation degree calculation on each dialogue state to determine semantic correlation degrees between each two dialogues, and based on a preset correlation degree threshold and each semantic correlation degree, determining a dialogue set for joint quality inspection; performing dialogue joint quality inspection on the dialogue set to determine a joint quality inspection result of the dialogue set; based on the individual quality inspection results and the joint quality inspection result, determining the first quality inspection result of the conversation to be quality inspected.

[0016] Optionally, the method further comprises: inputting a second preset instruction, the text data, an industry knowledge base and an external policy base into the preset voice quality inspection model; in the preset voice quality inspection model, based on the second preset instruction and the industry knowledge base, information recognition is performed on the text data to determine professional information specially defined in an industry in the to-be-inspected conversation and limited information corresponding to the professional information; based on the external policy base, the professional information is searched to determine standard information corresponding to the professional information; the limited information and the standard information are inspected to determine a second inspection result of the to-be-inspected conversation.

[0017] In a second aspect, an embodiment of the present application further provides a voice quality inspection device, which comprises:

[0018] A voice data determination module is configured to acquire voice signals of a to-be-inspected conversation and perform voice track determination on the voice signals to determine voice data containing an identity.

[0019] A text data determination module is configured to perform format conversion on the voice data based on a preset voice recognition model to determine corresponding text data of the to-be-inspected conversation.

[0020] A first inspection result determination module is configured to perform multi-modal feature extraction and voice quality inspection based on a first preset instruction, a preset voice quality inspection model, the voice data and the text data to determine a first inspection result of the to-be-inspected conversation.

[0021] A second inspection result determination module is configured to perform inspection based on a second preset instruction, the text data, the preset voice quality inspection model, an industry knowledge base and an external policy base to determine a second inspection result of the to-be-inspected conversation.

[0022] A target inspection result determination module is configured to determine a target inspection result of the to-be-inspected conversation based on the first inspection result and the second inspection result.

[0023] In a third aspect, an embodiment of the present application further provides an electronic device, which comprises:

[0024] One or more processors;

[0025] A memory configured to store one or more programs;

[0026] When the one or more programs are executed by the one or more processors, the one or more processors implement the voice quality inspection method provided by any embodiment of the present application.

[0027] In a fourth aspect, an embodiment of the present application provides a computer readable storage medium, which stores a computer program, and the computer program is executed by a processor to implement the voice quality inspection method provided by any of the embodiments of the present application.

[0028] In a fifth aspect, an embodiment of the present application provides a computer program product, which comprises a computer program, and the computer program is executed by a processor to implement the voice quality inspection method provided by any of the embodiments of the present application.

[0029] The technical scheme of the embodiment of the present application comprises the following steps: obtaining a voice signal of a to-be-inspected session, and determining voice data containing an identity identifier by performing voice track splitting on the voice signal; performing format conversion on the voice data based on a preset voice recognition model to determine corresponding text data of the to-be-inspected session; performing multi-modal feature extraction and voice quality inspection based on a first preset instruction, a preset voice quality inspection model, the voice data and the text data to determine a first inspection result of the to-be-inspected session; performing inspection based on a second preset instruction, the text data, the preset voice quality inspection model, an industry knowledge base and an external policy base to determine a second inspection result of the to-be-inspected session; and determining a target inspection result of the to-be-inspected session based on the first inspection result and the second inspection result, so that the to-be-inspected session that ends the call can be comprehensively inspected in time, and the policy content involved in the to-be-inspected session can be timely inspected, without manual inspection, thereby improving the efficiency and accuracy of voice quality inspection.

[0030] It should be understood that the content described in this part is not intended to identify key or important features of the embodiments of the present application, nor is it used to limit the scope of the present application. Other features of the present application will become apparent from the following description. BRIEF DESCRIPTION OF DRAWINGS

[0031] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings needed in the embodiment description. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can also be obtained by those skilled in the art without creative labor.

[0032] Figure 1 is a flowchart of a voice quality inspection method provided by an embodiment one of the present application;

[0033] Figure 2 is a flowchart of a voice quality inspection method provided by an embodiment two of the present application;

[0034] Figure 3 is a structural schematic diagram of a voice quality inspection device provided by an embodiment three of the present application;

[0035] Figure 4 is a structural schematic diagram of an electronic device for implementing a voice quality inspection method according to an embodiment of the present application. DETAILED DESCRIPTION

[0036] In order to make the personnel in the art better understand the present application, the technical solutions in the embodiments of the present application will be described clearly and completely below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work should belong to the scope of protection of the present application.

[0037] It should be noted that the terms "first", "second", and the like in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily have to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or device that includes a series of steps or units does not have to be limited to those steps or units clearly listed, but can include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0038] Embodiment one

[0039] Figure 1 A flowchart of a voice quality inspection method is provided for the first embodiment of the present application. The present embodiment can be applied to the case of using a large model to inspect calls. The method can be performed by a voice quality inspection device, which can be realized in the form of hardware and / or software, and can be configured in an electronic device. As shown in the figure, the method comprises: Figure 1

[0040] S110, obtaining a voice signal of a session to be inspected, and determining voice data containing an identity identifier by voice splitting.

[0041] ​The to-be-inspected session can be a session that needs to be detected for whether the session content has a violation. In this embodiment, the to-be-inspected session can be a session generated in the process of a bank service business, which needs to be detected for whether the session content meets the internal requirements of the bank. For example, the to-be-inspected session can be, but is not limited to, a session between a customer and a customer service representative, a session between a customer and a customer manager, or a conference session. The voice signal can be a sound wave signal recording the session content. The voice signal can include human voice and environmental sound. In this embodiment, the voice signal can be, but is not limited to, a call recording or a conference recording. The identity identifier can be an identity identifier of a participant in the session. The identity identifier can be used to distinguish the session participants in the to-be-inspected session. In this embodiment, the identity identifier can include, but is not limited to, a customer and a customer service representative. The voice data can be voice data in which all participants in the to-be-inspected session have been distinguished. In this embodiment, the voice data can be voice data obtained by performing identity annotation on the voice signal for the voice signal.

[0042] Specifically, a call recording uploaded by a customer is received to obtain a voice signal of a to-be-inspected session. The voice signal of the to-be-inspected session can also be obtained from a call recording repository at regular intervals. The call recording repository can be a sampling repository for whether the service provided by a bank service personnel is compliant. The voice signal is subjected to environmental noise reduction and voiceprint recognition, thereby obtaining voice data in which an identity identifier is annotated on a completed time axis. For example, a certain time period is annotated as customer voice, and another time period is annotated as customer service representative voice.

[0043] On the basis of the above technical solution, the voice signal can be subjected to voice track determination to determine voice data containing an identity identifier, which can include: performing noise reduction processing on the voice signal based on a preset environmental noise reduction manner to determine a voice signal in which clear human voice is retained; performing voice track determination and role annotation on the voice signal in which clear human voice is retained based on a voiceprint recognition technology to determine voice data containing an identity identifier.

[0044] The preset environmental noise reduction mode can be a pre-set environmental noise removal mode. In this embodiment, the preset environmental noise reduction mode can be applied to remove typical noises in a bank, such as the sound of counting money and the sound of a keyboard. The preset environmental noise reduction mode can be a noise reduction mode combining a noise reduction algorithm and a deep neural network. The noise reduction algorithm is used to reduce noise for steady-state noise. The deep neural network (Deep Neural Networks, DNN) is used to reduce noise for non-steady-state noise. In this embodiment, the noise reduction algorithm can be, but is not limited to, spectral subtraction. The steady-state noise can be continuous, regular or frequency-stable noise. For example, the steady-state noise can include, but is not limited to, the low-frequency humming of an air conditioner and the sound of a server fan. The non-steady-state noise can be burst, irregular or frequency-dynamic noise. For example, the non-steady-state noise can include, but is not limited to, customer queuing conversations, keyboard tapping sounds and the sound of counting money. The voiceprint recognition technology can be a pre-set technology for identifying the identity of a speaker by analyzing biological features in the voice, such as vocal cord structure and pronunciation habits.

[0045] Specifically, the steady-state noise is quickly eliminated by spectral subtraction or linear filter, and the non-steady-state noise is removed by using a pre-trained noise reduction model for a bank scene, so as to obtain a voice signal retaining clear human voice by eliminating the steady-state noise and the non-steady-state noise respectively, thereby improving the effectiveness of the voice signal noise reduction processing. The voice signal retaining clear human voice is subjected to biological feature extraction and voiceprint comparison by using a pre-trained voiceprint recognition model, to obtain voice data containing an identity label, so as to complete the dialogue role labeling while ensuring the voice quality, and finally output pure two-channel voice data with an identity label.

[0046] In this embodiment, in a bank customer service scene, the process of voice splitting and role labeling based on the voiceprint recognition technology can include: framing (20-30 ms / frame) the voice signal retaining clear human voice, and extracting MFCC features to obtain static features (fundamental frequency, formant frequency) and dynamic features (speech rate, tone). Based on energy mutation detection or silent section division, the voice sections of different speakers (such as the alternating dialogue of the customer service and the customer) are separated. The voice of the distinguished speaker is compared with the pre-registered voiceprint library. For example, the customer service voiceprint has been pre-recorded into the system (ID: Agent_001), and the customer voiceprint is registered in real time (ID: User_12345). The speaker switching point is marked on the time axis, and the voice stream with an identity label (such as [00:12-00:30] Agent_001: “What business do you need to handle?”) is generated, so as to obtain voice data containing an identity label.

[0047] S120, format conversion is performed on the voice data based on a pre-set voice recognition model, to determine the corresponding text data of the to-be-inspected conversation.

[0048] The preset voice recognition model can be a pre-trained deep learning model for converting voice into text. In this embodiment, the preset voice recognition model is a joint model including an acoustic model and a language model. The text data can be data for displaying voice content in a literal form.

[0049] Specifically, the voice data is input into the preset voice recognition model, the voice data is converted into a phoneme or character sequence by using an acoustic model (such as a hidden Markov model HMM), and then a language model is used to evaluate the probability of the character sequence to ensure that the recognition result is semantically smooth. The language model combines context and grammar rules to optimize the output of the acoustic model, so that the converted text data is more reasonable.

[0050] On the basis of the above technical solution, the preset voice recognition model is constructed by loading a preset bank business exclusive dictionary based on a general voice recognition model. The format conversion of the voice data based on the preset voice recognition model to determine the corresponding text data of the to-be-inspected conversation can include: format conversion of the voice data based on the preset voice recognition model to determine the text data to be detected; error correction and modification of the text data to be detected based on context semantics to determine the corresponding text data of the to-be-inspected conversation.

[0051] The preset bank business exclusive dictionary can be a pre-configured dictionary specially recording bank business related terms. The preset bank business exclusive dictionary can be used to enable the model to accurately translate and understand professional vocabulary in the financial field. For example, “bandwidth” in the text data to be detected can be modified to “loan” by the preset bank business exclusive dictionary internalized by the model, thereby improving the recognition accuracy of the model for bank business specific vocabulary (such as “financial products” and “cross-border remittance”).

[0052] Specifically, a deep learning-based voice-to-text framework is pre-constructed, and a bank business field adaptive module is integrated in a basic automatic speech recognition (ASR) engine, so as to effectively improve the recognition accuracy of professional vocabulary by loading a preset bank business exclusive dictionary (such as a financial product term library and a compliance language dictionary and other professional corpus resources). The automatic speech recognition (ASR) result is automatically corrected based on context semantics, so as to obtain the corresponding text data of the to-be-inspected conversation.

[0053] In this embodiment, in the inference stage (such as the decoding stage), the probability distribution of the language model can be adjusted to preferentially match the words in the preset bank business exclusive dictionary, so as to realize error correction and modification of the text data to be detected, and determine the corresponding text data of the to-be-inspected conversation. Higher weights can also be set for the words in the preset bank business exclusive dictionary, so that the words in the dictionary are more likely to be selected during decoding.

[0054] S130, based on the first preset instruction, the preset voice inspection model, the voice data and the text data, performing multi-modal feature extraction and voice inspection to determine a first inspection result of the to-be-inspected conversation.

[0055] The first preset instruction can be a prompt word for the operation of the preset voice inspection model on the voice data and the text data. For example, the first preset instruction can be an end-to-end operation instruction, so that the large model can realize end-to-end processing from semantic analysis to problem positioning by constructing a fully automated inspection pipeline, effectively replacing traditional manual rule configuration and voice inspection links. The preset voice inspection model can be a large language model. The preset voice inspection model can be obtained by efficiently fine-tuning the large language model using a sample library. The sample library can include training samples with or without a violation label, such as customer service dialogues and complaint audio texts. In this embodiment, the sample library can be updated after the misjudged cases are corrected in the manually reviewed and marked violation results, so that the sample library is periodically fine-tuned to the preset voice inspection model. The first inspection result can be an inspection result of whether the conversation has a violation according to the non-professional content in the conversation. For example, the first inspection result can be, but is not limited to, judging whether there is a violation from the tone, attitude and whether the answer is not the question of the bank service personnel.

[0056] For example, efficient parameter fine-tuning of the preset voice inspection model can be achieved using a sample library and a newly added low-rank matrix, so that the fine-tuned model adapts to the semantic understanding and compliance judgment requirements of the inspection scene. Specifically, only the newly added low-rank matrix can be trained, and most of the parameters of the original model are frozen. For example, an adaptive layer is inserted in the large language model, and the weights are updated using inspection dialogue data (sample library).

[0057] In this embodiment, the first preset instruction can include: extracting emotion features and speech rate features from the input voice data; extracting keywords related to bank business from the input text data, and understanding the customer's intention by performing semantic analysis on the text data; and using the extracted emotion features, speech rate features, keywords and intention to determine whether there is a violation in the conversation. Through the first preset instruction, the preset voice inspection model is guided to perform specific operations on the voice data and the text data step by step, so as to obtain the first inspection result of the to-be-inspected conversation.

[0058] In this embodiment, the input of the preset voice quality inspection model can further include operation examples. The operation examples can be used to enable the preset voice quality inspection model to understand how to correctly execute the first preset instruction. The operation examples can include: emotion feature extraction examples, speech rate feature extraction examples, keyword extraction examples, intent extraction examples, and examples of quality inspection using multi-modal features. The more examples, the more accurate the features extracted by the model. For example, the emotion feature extraction examples include: a voice data sample, a feature extraction result for the voice data sample, and an emotion feature label of the feature extraction result. On this basis, each step in the first preset instruction is added with "operate according to the provided XX example (such as extraction and quality inspection)".

[0059] S140, based on the second preset instruction, the text data, the preset voice quality inspection model, the industry knowledge base and the external policy base, performing quality inspection to determine a second quality inspection result of the to-be-inspected conversation.

[0060] The second preset instruction can be a prompt word for the operation of the preset voice quality inspection model on the text data, the industry knowledge base and the external policy base. Similar to the first preset instruction, the second preset instruction can be input into the preset voice quality inspection model together with the examples, or can be input into the preset voice quality inspection model alone. The industry knowledge base can be a knowledge base storing business service processes and reported speech rules, and composed in the form of a vector database. The external policy base can be a database composed of the latest compliance requirements released by an external regulatory platform (such as the policy base of the China Securities Regulatory Commission in the financial industry). In this embodiment, the external policy base can be connected through an API interface, and the latest compliance requirements can be dynamically obtained. The second quality inspection result can be a quality inspection result of whether the conversation has a violation behavior, which is determined by the preset voice quality inspection model according to the professional content in the conversation. For example, the second quality inspection result can be to determine whether there is a violation behavior from the accuracy of the answer content of the bank service personnel.

[0061] Specifically, the preset voice quality inspection model performs quality inspection according to the second preset instruction, in combination with the industry knowledge base and the external policy base, to obtain the second quality inspection result. For example, when the customer service mentions "loan interest discount" in the industry knowledge base, the latest Loan Prime Rate (LPR) requirement in the external policy base is searched to determine whether there is a misleading statement in the related content described by the customer service. If there is, there is no violation behavior.

[0062] S150, based on the first quality inspection result and the second quality inspection result, determining a target quality inspection result of the to-be-inspected conversation.

[0063] The target quality inspection result can be a final quality inspection result that comprehensively considers the first quality inspection result and the second quality inspection result. The target quality inspection result can include specific conversations and specific violation behaviors that exist in the target quality inspection result. Specifically, if at least one of the first quality inspection result and the second quality inspection result contains a violation behavior, the target quality inspection result also contains a violation behavior. The first quality inspection result and the second quality inspection result are combined, and each violation behavior is arranged in a descending order of priority based on a preset violation behavior level, to obtain the target quality inspection result of the to-be-inspected conversation.

[0064] The technical solution of the embodiment of the application obtains a voice signal of a to-be-inspected conversation, determines voice data containing an identity identifier by performing voice track splitting on the voice signal, converts the voice data into text data corresponding to the to-be-inspected conversation based on a preset voice recognition model, performs multi-modal feature extraction and voice quality inspection based on a first preset instruction, a preset voice quality inspection model, the voice data, and the text data to determine a first quality inspection result of the to-be-inspected conversation, performs quality inspection based on a second preset instruction, the text data, the preset voice quality inspection model, an industry knowledge base, and an external policy base to determine a second quality inspection result of the to-be-inspected conversation, and determines a target quality inspection result of the to-be-inspected conversation based on the first quality inspection result and the second quality inspection result. Therefore, the to-be-inspected conversation that ends a call can be comprehensively inspected in a timely manner, and policy content involved in the to-be-inspected conversation can be inspected in a timely manner. Manual inspection is not required, and therefore the efficiency and accuracy of voice quality inspection are improved.

[0065] Embodiment Two

[0066] Figure 2 A flowchart of a voice quality inspection method provided by Embodiment Two of the application is described in detail below. The same or corresponding terms as those in the above embodiments are not described again. As shown in FIG. 2, the method includes the following steps. Figure 2

[0067] S210, a voice signal of a to-be-inspected conversation is obtained, and voice data containing an identity identifier is determined by performing voice track splitting on the voice signal.

[0068] S220, the voice data is converted into text data corresponding to the to-be-inspected conversation based on a preset voice recognition model.

[0069] ​Exemplarily, the text data after format conversion is preprocessed, such as cleaning, word segmentation, and removal of stop words, so as to facilitate subsequent processing of the large model. Specifically, noise in the text data, such as extra spaces, tabs, line breaks, and invisible special characters, is cleaned up. The cleaned text data is segmented into independent lexical units according to a preset rule (such as space, punctuation, etc.), that is, word segmentation, and the lexical units are used to form text data input into the preset voice quality inspection model.

[0070] S230, input the first preset instruction, the voice data and the text data into the preset voice quality inspection model.

[0071] In this embodiment, the preset voice quality inspection model can realize the operation of each step in the preset instruction by calling the hybrid scheme of the external model, so as to realize voice quality inspection by using a special small model, and further improve the accuracy of voice quality inspection. The external model can be pre-trained using bank sample data, and the training process can also use the preset voice quality inspection model.

[0072] S240, in the preset voice quality inspection model, feature extraction is performed on the voice data based on the first operation instruction in the first preset instruction, and emotion features and speech rate features corresponding to each dialogue in the to-be-inspected conversation are determined.

[0073] The first operation instruction is used to guide the preset voice quality inspection model to call an external emotion feature extraction model and an external speech rate feature extraction model to realize feature extraction on the voice data. The emotion feature can be the emotion type of each role in each dialogue in the to-be-inspected conversation. One question and answer between the customer and the bank service personnel is one dialogue. The emotion feature can be, but is not limited to, joy or anger. The speech rate feature can be the average value of the speaking speed of each role in each dialogue. For example, the speech rate feature can be 1 times speed or 1.2 times speed.

[0074] Specifically, in the preset voice quality inspection model, based on the first operation instruction in the first preset instruction, the preset voice quality inspection model is guided to act, such as calling an external emotion feature extraction model to perform emotion feature extraction operation on the input voice data, and calling an external speech rate feature extraction model to perform speech rate feature extraction operation on the input voice data. The emotion feature corresponding to each dialogue in the to-be-inspected conversation is determined based on the output result of the external emotion feature extraction model. The speech rate feature corresponding to each dialogue in the to-be-inspected conversation is determined based on the output result of the external speech rate feature extraction model.

[0075] S250, based on the second operation instruction in the first preset instruction, feature extraction is performed on the text data, and keywords and intents corresponding to each dialogue in the to-be-inspected conversation are determined.

[0076] Keywords can refer to terms related to banking services. For example, keywords may include, but are not limited to, accounts and deposits. Intent can refer to the intention of each character in each dialogue. For example, in a dialogue, the customer's intent is to inquire about the deposit method with the highest interest rate, while the customer service representative's intent is to explain a financial management method with a high interest rate but also high risk.

[0077] Specifically, in the preset voice quality inspection model, the second operation instruction in the first preset command guides the preset voice quality inspection model to perform actions, such as calling an external keyword extraction model to extract keywords related to banking services from the input text data, and calling an external intent recognition model to perform intent recognition on the input text data. The keywords corresponding to each dialogue in the session to be inspected are determined based on the output of the external keyword extraction model. The intent corresponding to each dialogue in the session to be inspected is determined based on the output of the external intent recognition model.

[0078] S260. Based on the first preset instruction and the emotional characteristics, speech rate characteristics, keywords and intentions corresponding to each dialogue in the session to be inspected, perform voice quality inspection and determine the first quality inspection result of the session to be inspected.

[0079] Specifically, in the preset speech quality inspection model, an external multimodal speech quality inspection model is called based on the first preset instruction. The external multimodal speech quality inspection model is then used to perform a comprehensive speech quality inspection by combining the emotional features, speech rate features, keywords and intentions of each dialogue in the session to be inspected, and the first quality inspection result of the session to be inspected is determined.

[0080] Based on the above technical solution, "performing voice quality inspection based on the first preset instruction and the emotional features, speech rate features, keywords and intentions corresponding to each dialogue in the session to be inspected, and determining the first quality inspection result of the session to be inspected" may include: using the third operation instruction in the first preset instruction to perform feature association on the emotional features, speech rate features, keywords and intentions corresponding to each dialogue to determine the dialogue features of each dialogue, and determining the dialogue state of each dialogue based on the dialogue features of each dialogue; performing progressive quality inspection on each dialogue based on the fourth operation instruction in the first preset instruction and the dialogue state, and determining the first quality inspection result of the session to be inspected.

[0081] The third operation instruction can be used to guide the preset voice quality inspection model to call an external cross-modal model, use the external cross-modal model to associate multimodal features, and determine the dialogue state based on the dialogue features. Dialogue features can refer to the features obtained after associating multimodal features. Dialogue state can refer to the structured integration result of dialogue features. Dialogue state can be used to represent a "snapshot" or "portrait" of the dialogue. For example, high speech speed + anger + keyword "fault" → the customer is likely in an "urgent repair" state. The fourth operation instruction can be used to guide the preset voice quality inspection model to perform quality inspection operations on similar dialogues together.

[0082] Specifically, in the preset speech quality inspection model, an external cross-modal model is invoked based on the third operation instruction in the first preset instruction. This external cross-modal model is then used to correlate the emotion features, speech rate features, keywords, and intent corresponding to each dialogue to determine the dialogue features of each dialogue. Based on the third operation instruction in the first preset instruction, the dialogue features of each dialogue are probabilistically classified, and the classification result with the highest probability is determined as the dialogue state of each dialogue. The probabilistic classification operation can also use a pre-trained classification model. In this embodiment, the associated features are mapped to structured states through rules or models (such as state machines or deep learning). For example, input: emotion = negative, intent = complaint, keyword = "delay" → output state label "complaint being processed". Based on the fourth operation instruction in the first preset instruction, similar dialogues are subjected to a joint quality inspection operation, i.e., progressive quality inspection, thereby determining the first quality inspection result of the dialogue to be inspected.

[0083] Based on the above technical solution, "performing progressive quality checks on each dialogue based on the fourth operation instruction in the first preset instruction and the dialogue state, and determining the first quality check result of the session to be checked" may include: performing single-dialogue quality checks on each dialogue based on the fourth operation instruction in the first preset instruction, and determining the individual quality check result of each dialogue; calculating the correlation degree of each dialogue state, determining the semantic correlation degree between each pair of dialogues, and determining the dialogue set for joint quality checks based on the preset correlation degree threshold and each semantic correlation degree; performing joint quality checks on the dialogue set, and determining the joint quality check result of the dialogue set; and determining the first quality check result of the session to be checked based on the individual quality check results and the joint quality check results.

[0084] Individual quality inspection results can refer to the quality inspection results of whether a single dialogue contains violations. Semantic relevance can refer to the similarity between the content of two dialogues, i.e., semantic similarity. The preset relevance threshold can refer to the minimum similarity threshold for confirming that the content of two dialogues is similar. There must be at least two individual dialogues in the dialogue set. Joint quality inspection results can refer to the quality inspection results of combining the dialogues as a whole.

[0085] Specifically, based on the fourth operation instruction in the first preset instruction, a single-dialogue quality inspection is performed on each dialogue to determine the individual quality inspection result for each dialogue. For example, in a scenario where the intent is to complain, the customer service representative's angry expression is determined to be a violation. The criteria for judging whether a single dialogue contains a violation are the sample library and the examples input with the instruction. According to different business needs and different judgment scales, users can customize the sample library and examples so that the preset voice quality inspection model can perform dialogue quality inspection at different scales according to user needs. The correlation degree of each dialogue state is calculated to determine the semantic correlation degree between each pair of dialogues. In this embodiment, the dialogue state and correlation features of all dialogues can be represented using dialogue feature maps, so that the semantic correlation degree between each dialogue can be calculated through graph neural networks. Based on the preset correlation degree threshold and the comparison of each semantic correlation degree, dialogues corresponding to semantic correlation degrees greater than the preset correlation threshold are integrated to form a dialogue set for joint quality inspection. The dialogues in the dialogue set are treated as a whole for quality inspection to determine the joint quality inspection result of the dialogue set. Based on the merging of the individual quality inspection results and the joint quality inspection results, the first quality inspection result of the dialogue to be inspected is determined.

[0086] It should be noted that the progressive quality inspection method in this embodiment enables the large language model to capture key information nodes across dozens of rounds of dialogue, effectively identify deep-seated problems such as logical breaks, topic jumps, and information omissions in the response process, and further improve the contextual understanding ability and accuracy of the large language model in the speech quality inspection process.

[0087] S270. Based on the second preset instruction, text data, preset voice quality inspection model, industry knowledge base and external policy base, perform quality inspection and determine the second quality inspection result of the session to be inspected.

[0088] Based on the above technical solution, "performing quality inspection based on a second preset instruction, text data, a preset voice quality inspection model, an industry knowledge base, and an external policy database to determine the second quality inspection result of the session to be inspected" can include: inputting the second preset instruction, text data, an industry knowledge base, and an external policy database into a preset voice quality inspection model; in the preset voice quality inspection model, performing information recognition on the text data based on the second preset instruction and the industry knowledge base to determine the industry-specific professional information and the corresponding limiting information in the session to be inspected; retrieving the professional information in conjunction with the external policy database to determine the standard information corresponding to the professional information; and performing quality inspection on the limiting information and the standard information to determine the second quality inspection result of the session to be inspected.

[0089] Specifically, "professional information" can refer to information with specific limitations within an industry. "Limited information" can refer to the specific details about the professional information presented by the bank staff during the conversation. "Standard information" can refer to the standard information describing the professional information. For example, if the professional information is about loan interest rate discounts, then the limited information would be the content related to loan interest rate discounts presented by the bank staff during the conversation, such as explanatory content about loan interest rate discount policies, and the standard information would be the standard descriptions related to loan interest rate discounts from external policy databases, such as the original text of the loan interest rate discount policies.

[0090] Specifically, the second preset instruction, text data, industry knowledge base, and external policy database are input into a preset voice quality inspection model. Within this model, information recognition is performed on the text data based on the second preset instruction and the industry knowledge base. When the dialogue mentions "loan interest rate discounts" (equivalent to professional information) from the industry knowledge base, the latest Loan Prime Rate (LPR) requirement from the central bank (equivalent to standard information) is retrieved from the external policy database. The system then assesses whether there are any misleading statements in the customer service representative's presentation (equivalent to limiting information). If so, no violation has been detected.

[0091] It should be noted that in this embodiment, the process of calling the large language model twice to determine two different quality inspection results is carried out. Through deep semantic understanding, contextual association analysis and dynamic knowledge enhancement, complex speech and cross-round violations are accurately identified, reducing manual processing and achieving a balance between high accuracy and high efficiency.

[0092] S280. Based on the first quality inspection result and the second quality inspection result, determine the target quality inspection result for the session to be inspected.

[0093] The technical solution of this invention involves inputting a first preset instruction, voice data, and text data into a preset voice quality inspection model. Within the preset voice quality inspection model, features are extracted from the voice data based on the first operation instruction in the first preset instruction to determine the emotional and speech rate features corresponding to each dialogue in the session to be inspected. Features are extracted from the text data based on the second operation instruction in the first preset instruction to determine the keywords and intentions corresponding to each dialogue in the session to be inspected. Voice quality inspection is then performed based on the first preset instruction and the emotional, speech rate, keywords, and intentions corresponding to each dialogue in the session to be inspected, determining the first quality inspection result for the session to be inspected. This allows for comprehensive quality inspection of the dialogue content of each dialogue in the session to be inspected from multiple perspectives, further improving the accuracy of voice quality inspection.

[0094] The following are embodiments of the voice quality inspection device provided in this invention. This device and the voice quality inspection methods in the above embodiments belong to the same inventive concept. For details not described in detail in the embodiments of the voice quality inspection device, please refer to the embodiments of the above voice quality inspection methods.

[0095] Example 3

[0096] Figure 3 This is a schematic diagram of a voice quality inspection device provided in Embodiment 3 of the present invention. Figure 3 As shown, the device includes: a voice data determination module 310, a text data determination module 320, a first quality inspection result determination module 330, a second quality inspection result determination module 340, and a target quality inspection result determination module 350.

[0097] The system includes the following modules: a voice data determination module 310, which acquires the voice signal of the session to be inspected and performs voice track division to determine the voice data containing the identity identifier; a text data determination module 320, which performs format conversion on the voice data based on a preset voice recognition model to determine the corresponding text data of the session to be inspected; a first quality inspection result determination module 330, which performs multimodal feature extraction and voice quality inspection based on a first preset instruction, a preset voice quality inspection model, voice data, and text data to determine the first quality inspection result of the session to be inspected; a second quality inspection result determination module 340, which performs quality inspection based on a second preset instruction, text data, a preset voice quality inspection model, an industry knowledge base, and an external policy base to determine the second quality inspection result of the session to be inspected; and a target quality inspection result determination module 350, which determines the target quality inspection result of the session to be inspected based on the first and second quality inspection results.

[0098] The technical solution of this invention acquires the voice signal of a session to be inspected and performs voice track segmentation to determine voice data containing identity identifiers; performs format conversion on the voice data based on a preset voice recognition model to determine the corresponding text data of the session to be inspected; performs multimodal feature extraction and voice quality inspection based on a first preset instruction, a preset voice quality inspection model, voice data, and text data to determine a first quality inspection result of the session to be inspected; performs quality inspection based on a second preset instruction, text data, a preset voice quality inspection model, an industry knowledge base, and an external policy database to determine a second quality inspection result of the session to be inspected; and determines a target quality inspection result of the session to be inspected based on the first and second quality inspection results. This allows for timely and comprehensive quality inspection of terminated sessions and timely quality inspection of policy content involved in the sessions, eliminating the need for manual quality inspection and thus improving the efficiency and accuracy of voice quality inspection.

[0099] Based on the above technical solution, the voice data determination module 310 is specifically used to: perform noise reduction processing on the voice signal for steady-state noise and non-steady-state noise based on a preset environmental noise reduction method, and determine the voice signal that retains clear human voice; and perform voice track separation and role labeling on the voice signal that retains clear human voice based on voiceprint recognition technology, and determine the voice data containing identity identifiers.

[0100] Based on the above technical solution, the preset speech recognition model is constructed by loading a preset banking business-specific dictionary onto a general speech recognition model.

[0101] The text data determination module 320 is specifically used for: converting the format of speech data based on a preset speech recognition model to determine the text data to be detected; correcting and modifying the text data to be detected based on contextual semantics to determine the corresponding text data of the session to be inspected.

[0102] Based on the above technical solution, the first quality inspection result determination module 330 may include:

[0103] The data input submodule is used to input the first preset command, voice data and text data into the preset voice quality inspection model;

[0104] The first feature extraction submodule is used to extract features from the speech data based on the first operation instruction in the first preset instruction in the preset speech quality inspection model, and to determine the emotion features and speech rate features corresponding to each dialogue in the conversation to be inspected.

[0105] The second feature extraction submodule is used to extract features from text data based on the second operation instruction in the first preset instruction, and to determine the keywords and intents corresponding to each dialogue in the session to be inspected.

[0106] The first quality inspection result determination submodule is used to perform voice quality inspection based on the first preset instruction and the emotional characteristics, speech rate characteristics, keywords and intentions corresponding to each dialogue in the session to be inspected, and to determine the first quality inspection result of the session to be inspected.

[0107] Based on the above technical solution, the first quality inspection result determination submodule may include:

[0108] The dialogue state determination unit is used to determine the dialogue features of each dialogue by performing feature association on the emotional features, speech rate features, keywords and intentions corresponding to each dialogue based on the third operation instruction in the first preset instruction, and to determine the dialogue state of each dialogue based on the dialogue features of each dialogue.

[0109] The first quality inspection result determination unit is used to perform progressive quality inspection on each dialogue based on the fourth operation instruction in the first preset instruction and the dialogue state, and determine the first quality inspection result of the dialogue to be inspected.

[0110] Based on the above technical solution, the first quality inspection result determination unit is specifically used for: performing single-dialogue quality inspection on each dialogue based on the fourth operation instruction in the first preset instruction, and determining the individual quality inspection result of each dialogue; calculating the correlation degree of each dialogue state, determining the semantic correlation degree between each pair of dialogues, and determining the dialogue set for joint quality inspection based on the preset correlation degree threshold and each semantic correlation degree; performing joint quality inspection on the dialogue set, and determining the joint quality inspection result of the dialogue set; and determining the first quality inspection result of the session to be inspected based on the individual quality inspection results and the joint quality inspection results.

[0111] Based on the above technical solution, the second quality inspection result determination module 340 is specifically used for: inputting the second preset instruction, text data, industry knowledge base, and external policy base into the preset voice quality inspection model; in the preset voice quality inspection model, based on the second preset instruction and the industry knowledge base, performing information recognition on the text data to determine the industry-specific professional information and the corresponding limitation information in the session to be inspected; retrieving the professional information in conjunction with the external policy base to determine the standard information corresponding to the professional information; and performing quality inspection on the limitation information and standard information to determine the second quality inspection result of the session to be inspected.

[0112] The voice quality inspection device provided in the embodiments of the present invention can execute the voice quality inspection method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects for executing the voice quality inspection method.

[0113] It is worth noting that in the above-described embodiments of voice quality inspection, the various units and modules included are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be achieved; in addition, the specific names of each functional unit are only for easy differentiation and are not used to limit the scope of protection of this invention.

[0114] Example 4

[0115] Figure 4 A schematic diagram of an electronic device 10 that can be used to implement embodiments of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (e.g., helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.

[0116] like Figure 4As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded from storage unit 18 into the RAM 13. The RAM 13 may also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.

[0117] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0118] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as voice quality inspection methods.

[0119] In some embodiments, the voice quality inspection method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the voice quality inspection method described above may be performed. Alternatively, in other embodiments, processor 11 may be configured to perform the voice quality inspection method by any other suitable means (e.g., by means of firmware).

[0120] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0121] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0122] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0123] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0124] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or computing systems that include middleware components (e.g., application servers), or computing systems that include frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.

[0125] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.

[0126] This invention also provides a computer program product, including a computer program that, when executed by a processor, implements the voice quality inspection method provided in any embodiment of this application.

[0127] In implementing the computer program product, computer program code for performing the operations of this invention can be written in one or more programming languages ​​or a combination thereof. Programming languages ​​include object-oriented programming languages ​​such as Java, Smalltalk, and C++, as well as conventional procedural programming languages ​​such as C or similar languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider). This program product belongs to the same inventive concept as the voice quality inspection method disclosed in the embodiments of this application, and therefore will not be described further here.

[0128] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.

[0129] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.

Claims

1. A voice quality inspection method, characterized in that, include: Acquire the voice signal of the session to be inspected, and perform voice track segmentation on the voice signal to determine the voice data containing the identity identifier; The voice data is converted into a format based on a preset speech recognition model to determine the corresponding text data of the session to be inspected. Based on the first preset instruction, the preset voice quality inspection model, the voice data and the text data, multimodal feature extraction and voice quality inspection are performed to determine the first quality inspection result of the session to be inspected; Based on the second preset instruction, the text data, the preset voice quality inspection model, the industry knowledge base, and the external policy base, a quality inspection is performed to determine the second quality inspection result of the session to be inspected; Based on the first quality inspection result and the second quality inspection result, the target quality inspection result of the session to be inspected is determined; The step of performing multimodal feature extraction and voice quality inspection based on a first preset instruction, a preset voice quality inspection model, the voice data, and the text data to determine the first quality inspection result of the session to be inspected includes: inputting the first preset instruction, the voice data, and the text data into the preset voice quality inspection model; in the preset voice quality inspection model, performing feature extraction on the voice data based on the first operation instruction in the first preset instruction to determine the emotion features and speech rate features corresponding to each dialogue in the session to be inspected; performing feature extraction on the text data based on the second operation instruction in the first preset instruction to determine the keywords and intentions corresponding to each dialogue in the session to be inspected; and performing voice quality inspection based on the first preset instruction and the emotion features, speech rate features, keywords, and intentions corresponding to each dialogue in the session to be inspected to determine the first quality inspection result of the session to be inspected.

2. The method according to claim 1, characterized in that, The step of performing voice track segmentation on the voice signal to determine voice data containing identity identifiers includes: The speech signal is subjected to noise reduction processing based on a preset environmental noise reduction method to reduce steady-state noise and non-steady-state noise, thereby determining the speech signal that retains clear human voice. Based on voiceprint recognition technology, the speech signal that retains a clear human voice is divided into audio tracks and labeled with roles to determine the speech data containing identity identifiers.

3. The method according to claim 1, characterized in that, The preset speech recognition model is constructed by loading a preset banking business-specific dictionary onto a general speech recognition model; The step of converting the voice data according to a preset speech recognition model to determine the corresponding text data of the session to be inspected includes: The speech data is converted into a format based on a preset speech recognition model to determine the text data to be detected; The text data to be detected is corrected and modified based on contextual semantics to determine the corresponding text data of the session to be inspected.

4. The method according to claim 1, characterized in that, The process of performing voice quality inspection based on the first preset instruction and the emotional features, speech rate features, keywords, and intent corresponding to each dialogue in the session to be inspected, and determining the first quality inspection result of the session to be inspected, includes: Based on the third operation instruction in the first preset instruction, the emotional characteristics, speech rate characteristics, keywords and intentions corresponding to each dialogue are correlated to determine the dialogue characteristics of each dialogue, and the dialogue state of each dialogue is determined according to the dialogue characteristics of each dialogue. Based on the fourth operation instruction in the first preset instruction and the dialogue state, progressive quality inspection is performed on each dialogue to determine the first quality inspection result of the dialogue to be inspected.

5. The method according to claim 4, characterized in that, The progressive quality inspection of each dialogue based on the fourth operation instruction in the first preset instruction and the dialogue state, and the determination of the first quality inspection result of the dialogue to be inspected, includes: Based on the fourth operation instruction in the first preset instruction, perform single dialogue quality inspection on each dialogue and determine the individual quality inspection result of each dialogue. The correlation degree of each dialogue state is calculated to determine the semantic correlation degree between each pair of dialogues, and the set of dialogues for joint quality inspection is determined based on the preset correlation degree threshold and each semantic correlation degree. Perform joint quality inspection on the dialogue set and determine the joint quality inspection result of the dialogue set; Based on the individual quality inspection results and the joint quality inspection results, the first quality inspection result of the session to be inspected is determined.

6. The method according to claim 1, characterized in that, The quality inspection based on the second preset instruction, the text data, the preset voice quality inspection model, the industry knowledge base, and the external policy database, to determine the second quality inspection result of the session to be inspected, includes: The second preset instruction, the text data, the industry knowledge base, and the external policy base are input into the preset voice quality inspection model; In the preset voice quality inspection model, the text data is identified based on the second preset instruction and the industry knowledge base to determine the industry-specific professional information and the corresponding limitation information in the conversation to be inspected. By combining the external policy database to retrieve the professional information, the standard information corresponding to the professional information is determined; The limiting information and the standard information are subjected to quality inspection to determine the second quality inspection result of the session to be inspected.

7. A voice quality inspection device, characterized in that, The device includes: The voice data determination module is used to acquire the voice signal of the session to be inspected, and to perform voice track division on the voice signal to determine the voice data containing the identity identifier; The text data determination module is used to perform format conversion on the voice data based on a preset speech recognition model to determine the corresponding text data of the session to be inspected. The first quality inspection result determination module is used to perform multimodal feature extraction and voice quality inspection based on the first preset instruction, the preset voice quality inspection model, the voice data and the text data, and determine the first quality inspection result of the session to be inspected; The second quality inspection result determination module is used to perform quality inspection based on the second preset instruction, the text data, the preset voice quality inspection model, the industry knowledge base and the external policy base, and determine the second quality inspection result of the session to be inspected; The target quality inspection result determination module is used to determine the target quality inspection result of the session to be inspected based on the first quality inspection result and the second quality inspection result. The first quality inspection result determination module includes: The data input submodule is used to input the first preset instruction, the voice data, and the text data into the preset voice quality inspection model; The first feature extraction submodule is used to extract features from the voice data based on the first operation instruction in the first preset instruction in the preset voice quality inspection model, and determine the emotion features and speech rate features corresponding to each dialogue in the conversation to be inspected. The second feature extraction submodule is used to extract features from the text data based on the second operation instruction in the first preset instruction, and to determine the keywords and intents corresponding to each dialogue in the session to be inspected. The first quality inspection result determination submodule is used to perform voice quality inspection based on the first preset instruction and the emotional features, speech rate features, keywords and intentions corresponding to each dialogue in the session to be inspected, and to determine the first quality inspection result of the session to be inspected.

8. An electronic device, characterized in that, The electronic device includes: One or more processors; Memory, used to store one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the voice quality inspection method as described in any one of claims 1-6.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the voice quality inspection method as described in any one of claims 1-6.

Citation Information

Patent Citations

  • Intelligent voice customer service quality inspection method and device based on multi-modal large model

    CN118631939A

  • Customer service verbal skill quality inspection result determination method and device, equipment, medium and product

    CN119294885A