Voice quality inspection method and device, equipment, storage medium and program product
By combining emotion recognition and speech recognition with text quality inspection model to process emotion recognition results and speech text data of speech data, the problem of low accuracy of speech quality inspection in the prior art is solved, and efficient recognition and accurate quality inspection of complex semantic violation speech data is achieved.
Patent Information
- Application Number
- CN202510728438.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-30
- Publication Date
- 2025-09-02
AI Technical Summary
In the prior art, the accuracy of speech data quality inspection is low, especially when user-defined sensitive words are not comprehensive or universal enough, it is difficult to identify illegal speech data in complex semantics, resulting in false detection and missed detection.
Through emotion recognition and speech recognition, the pre-trained text quality inspection model is used to process the emotional recognition results and speech text data of the speech data, and prompt words are constructed to realize speech quality inspection. The first prompt word is constructed based on the emotion recognition results and speech text data, and the text quality inspection model is used to perform speech quality inspection.
The accuracy of speech data quality inspection is improved, especially when identifying complex semantic violation speech data, the accuracy and recall rate of quality inspection is improved, solving the shortcomings of traditional sensitive thesaurus and rule matching solutions.
Smart Images

Figure CN120581040A_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the technical field of artificial intelligence (AI), and in particular relates to a speech quality inspection method, apparatus, device, storage medium, and program product. Background Art
[0002] In related technologies, the quality inspection of voice data is implemented by pre-setting a sensitive word library, then matching the voice text data through regular expressions, and then judging whether the voice data is compliant based on the matching results; in the above-mentioned scheme for quality inspection of voice data, users are required to pre-define the sensitive word library and matching rules. When the user-defined sensitive words are not comprehensive or universal, the problem of low quality inspection accuracy is likely to occur; and the above-mentioned scheme for quality inspection of voice data based on the sensitive word library and matching rules is difficult to identify illegal voice data with complex semantics that does not contain specific sensitive words, which reduces the quality inspection accuracy of voice data to a certain extent. Summary of the Invention
[0003] In order to solve the problem of low quality inspection accuracy of voice data in related technologies, the embodiments of the present application propose a voice quality inspection method, apparatus, device, storage medium and program product.
[0004] This embodiment of the present application proposes a method for speech quality inspection, the method comprising:
[0005] Get voice data;
[0006] Performing emotion recognition on the voice data based on the acoustic features of the voice data to obtain an emotion recognition result; performing speech recognition on the voice data to obtain voice text data;
[0007] A first prompt word is constructed according to the emotion recognition result and the speech text data, and the first prompt word is processed using a pre-trained text quality inspection model to obtain a speech quality inspection result.
[0008] In some embodiments, the emotion recognition is performed on the speech data based on the acoustic features of the speech data to obtain the emotion recognition result, including: using a pre-trained emotion recognition model to process the acoustic features of the speech data to obtain the emotion recognition result; the emotion recognition model is trained based on audio files of multiple emotion categories, and the emotion recognition result includes a label of the emotion category corresponding to the speech data.
[0009] In some embodiments, constructing the first prompt word based on the emotion recognition result and the voice text data includes: obtaining a violation risk point, where the violation risk point represents the violation risk information present in the voice data; and constructing the first prompt word based on the emotion recognition result, the voice text data, and the violation risk point.
[0010] In some embodiments, the obtaining of the violation risk points includes: obtaining the violation risk points from a knowledge base, or extracting the violation risk points from the voice text data.
[0011] In some embodiments, after extracting the violation risk points from the voice text data, the method further includes: entering the extracted violation risk points into the knowledge base.
[0012] In some embodiments, extracting the violation risk point from the voice text data includes: constructing a second prompt word for extracting the violation risk point; and generating the violation risk point based on the second prompt word and the voice text data.
[0013] In some embodiments, entering the extracted violation risk points into the knowledge base includes: if the extracted violation risk points pass manual review, entering the extracted violation risk points into the knowledge base.
[0014] The present invention also provides a speech quality inspection device, comprising:
[0015] An acquisition module, used to acquire voice data;
[0016] A first processing module is configured to perform emotion recognition on the speech data based on acoustic features of the speech data to obtain an emotion recognition result; and perform speech recognition on the speech data to obtain speech text data;
[0017] The second processing module is used to construct a first prompt word according to the emotion recognition result and the voice text data, and process the first prompt word using a pre-trained text quality inspection model to obtain a voice quality inspection result.
[0018] An embodiment of the present application also provides an electronic device, comprising a processor and a memory for storing a computer program that can be run on the processor; wherein the processor is used to run the computer program to perform any of the above-mentioned speech quality inspection methods.
[0019] An embodiment of the present application further provides a computer storage medium on which a computer program is stored. When the computer program is executed by a processor, any of the above-mentioned speech quality inspection methods is implemented.
[0020] An embodiment of the present application further provides a computer program product, including a computer program, which implements any of the above-mentioned speech quality inspection methods when executed by a processor.
[0021] In the solution of the related art that performs quality inspection of voice data based on the sensitive word library and matching rules, when the user-defined sensitive words are not comprehensive or universal, the problem of low quality inspection accuracy is easy to occur; however, the embodiment of the present application can perform emotion recognition and voice recognition on the voice data, and use the text quality inspection model to process the prompt words constructed based on the emotion recognition results of the voice data and the voice text data, so as to obtain the voice quality inspection results. Therefore, the embodiment of the present application does not need to perform quality inspection on the voice data through pre-defined sensitive words. Compared with the solution of the related art that performs quality inspection of voice data based on the sensitive word library and matching rules, the quality inspection accuracy is improved to a certain extent. Moreover, the embodiment of the present application can perform quality inspection on the voice data based on the emotion recognition results of the voice data. The emotion recognition results of the voice data help to understand the illegal voice data with complex semantics. Therefore, the quality inspection accuracy of the voice data can be further improved to a certain extent. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] Figure 1 Flowchart of the voice quality inspection method according to the embodiment of the present application
[0023] Figure 2 A schematic diagram of the architecture of the emotion representation model provided in the embodiment of the present application;
[0024] Figure 3 A schematic diagram comparing the emotion recognition results provided in the embodiment of the present application with the actual emotions;
[0025] Figure 4 A flowchart for quality inspection of recording files provided in an embodiment of the present application;
[0026] Figure 5 This is a structural diagram of a speech quality inspection device according to an embodiment of the present application;
[0027] Figure 6 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0028] In related technologies, the quality inspection of voice data is implemented as follows: by pre-setting a sensitive word library, and then matching the voice text data input by the user through regular expressions, and then judging whether the voice data is compliant based on the matching results; in the above-mentioned scheme for quality inspection of voice data, the user is required to pre-define the sensitive word library and matching rules. For example, in a call scenario, the user pre-defines the sensitive word library and matching rules for the call scenario, first performs scene recognition on the call content, and then uses regular expressions to match the scene sensitive words one by one on the call text, and judges whether the voice text data is in violation of regulations based on the matching rules.
[0029] In the aforementioned voice data quality inspection solution based on a sensitive word library and matching rules, sensitive words can be matched against user-entered text messages based on pre-set regular expressions and the sensitive word library. This solution primarily relies on keywords and regular expression technology, requiring quality inspectors to pre-configure a large number of sensitive words that violate regulations. However, this approach struggles to identify violations in complex semantic scenarios. For example, in loan collection, customer service representatives don't directly use sensitive words to insult or threaten users to repay their loans. Instead, they use threatening expressions that don't contain sensitive words, such as "I know your home address" or "You don't want us to visit you directly," to achieve collection purposes. Traditional keyword detection cannot effectively identify such violations, leading to false detections and missed detections. Furthermore, the complete semantic information of Chinese is strongly correlated with the call context, making single-sentence detection less effective in understanding semantics.
[0030] It can be seen that when the user-defined sensitive words are not comprehensive or universal, the problem of low quality inspection accuracy is likely to occur; in addition, the above-mentioned solution of voice data quality inspection based on the sensitive word library and matching rules is difficult to identify illegal voice data with complex semantics that does not contain specific sensitive words, which to a certain extent reduces the quality inspection accuracy of voice data.
[0031] Other related art solutions can introduce large text quality inspection models, leveraging their ability to understand complex semantics and perform compliance checks on call text content based on quality inspection questions pre-stored in a database. Related art quality inspection methods based on image recognition, speech recognition, and text compliance testing can use image technology to extract textual information from images during interactions and use regular expressions to match sensitive words against the aggregated text. These methods are essentially text-based regular expression matching.
[0032] Although the large text quality inspection model solves the problem of understanding complex semantics, it still only performs detection through a single text dimension. During actual calls, audio files still contain other semantic information in addition to the text dimension. Therefore, the quality inspection solution based solely on a single text dimension has the problem of low quality inspection accuracy.
[0033] In order to solve the technical problems existing in the related technologies, the technical solutions of the embodiments of the present application are proposed.
[0034] The following is a further detailed description of the embodiments of the present application in conjunction with the accompanying drawings and examples. It should be understood that the embodiments provided herein are merely for explaining the embodiments of the present application and are not intended to limit the embodiments of the present application. In addition, the embodiments provided below are partial embodiments for implementing the present application, rather than providing all embodiments for implementing the present application. In the absence of conflict, the technical solutions described in the embodiments of the present application can be implemented in any combination.
[0035] It should be noted that, in the embodiments of the present application, the terms "comprise", "include" or any other variants thereof are intended to cover non-exclusive inclusion, so that a method or apparatus comprising a series of elements includes not only the elements explicitly stated, but also other elements not explicitly listed, or also includes elements inherent to the implementation of the method or apparatus. In the absence of further restrictions, an element defined by the sentence "comprises a ..." does not exclude the presence of other related elements (such as steps in the method or units in the apparatus, for example, a unit may be a portion of a circuit, a portion of a processor, a portion of a program or software, etc.) in the method or apparatus comprising the element.
[0036] The speech quality inspection method provided in the embodiment of the present application includes a series of steps, but the speech quality inspection method provided in the embodiment of the present application is not limited to the recorded steps. Similarly, the speech quality inspection device provided in the embodiment of the present application includes a series of modules, but the device provided in the embodiment of the present application is not limited to including the modules explicitly recorded, and may also include modules required to obtain relevant information or perform processing based on the information.
[0037] Figure 1 This is a flow chart of the voice quality inspection method according to an embodiment of the present application. Figure 1 As shown, the process includes:
[0038] Step 101: Acquire voice data.
[0039] Here, the voice data may be a call recording file or other types of voice data. In practical applications, the voice data may be obtained in a variety of ways; for example, the voice data may be obtained from the storage space of a local computer device, downloaded from the network, or recorded using a recording device.
[0040] Step 102: Perform emotion recognition on the speech data based on the acoustic features of the speech data to obtain an emotion recognition result; perform speech recognition on the speech data to obtain speech text data.
[0041] In an embodiment of the present application, the acoustic features of the speech data may be voiceprint information, speaker tone, and other information. The voiceprint information may be Mel Frequency Cepstrum Coefficient (MFCC), pitch, energy, spectrum, and other information. By analyzing the voiceprint information, the speaker tone of the speech data can be determined.
[0042] When performing emotion recognition on speech data based on its acoustic features, in some embodiments, the speech data may first be preprocessed. For example, the speech data may be subjected to necessary noise reduction processing and volume normalization on the noise-reduced speech data. Mel-spectrogram cepstral coefficients, pitch, energy, spectrum, and other information may then be extracted from the volume-normalized speech data. After the speech data is preprocessed, emotion recognition may be performed on the speech data based on the preprocessed acoustic features to obtain an emotion recognition result.
[0043] For example, when the voice data is a call recording file, the technical solution of the embodiment of the present application can be adopted to use the emotion recognition model to perform emotion recognition on the call content, voiceprint information, and speaker tone, to determine whether the speaker has illegal characteristics such as threats, intimidation, and insults to others, and to combine the voice text data to perform quality inspection and scoring of the voice data in multiple dimensions, thereby identifying potential violation risks and effectively improving the accuracy of voice compliance quality inspection.
[0044] When performing speech recognition on speech data, the speech data can be preprocessed. For example, audio features of the speech data can be extracted based on a fixed audio sampling rate and a log-Mel spectrogram algorithm. The fixed audio sampling rate can be 16kHz, 24kHz, etc. After the audio features of the speech data are extracted, speech recognition can be performed based on the audio features. Specifically, the audio features can be input into an encoder, which extracts high-level audio features. The high-level audio features extracted by the encoder are then input into a decoder. The decoder processes the audio features to continuously generate text until an end marker is generated, at which point speech text data can be output. Here, the speech text data includes text content converted from the speech data. The speech recognition process can achieve conversion from speech data to text content. For example, the speech text data can include punctuated text and a timestamp. The timestamp is used to indicate the time when the speech text data was generated. In practical applications, the current time can be added to the speech text data as a timestamp when the speech text data is generated.
[0045] Step 103: construct a first prompt word according to the emotion recognition result and the speech text data, and process the first prompt word using a pre-trained text quality inspection model to obtain a speech quality inspection result.
[0046] Here, the first prompt word is prompt information input into the text quality inspection model through natural language. For example, the first prompt word can be a question, a description, or other forms of input information, which is used to help the text quality inspection model understand and perform the quality inspection task of the voice data. The voice quality inspection result can be prompt information that there is no violation risk in the voice data, or it can be prompt information that there is a specific violation risk in the voice data. The specific violation risk can be, for example, fraud risk, information leakage risk, or other violation risks. The text quality inspection model can be a large language model. In this way, the embodiment of the present application can use the complex semantic understanding ability of the large language model to process the first prompt word more accurately.
[0047] In practical applications, steps 101 to 103 may be implemented based on a processor, and the processor may be at least one of an application specific integrated circuit (ASIC), a digital signal processor (DSP), a digital signal processing device (DSPD), a programmable logic device (PLD), a field programmable gate array (FPGA), a central processing unit (CPU), a controller, a microcontroller, and a microprocessor.
[0048] It can be seen that in the solution of the related art that performs voice data quality inspection based on the sensitive word library and matching rules, when the user-defined sensitive words are not comprehensive or universal, the problem of low quality inspection accuracy is likely to occur; while the embodiment of the present application can perform emotion recognition and voice recognition on the voice data, and use the text quality inspection model to process the prompt words constructed based on the emotion recognition results of the voice data and the voice text data, so as to obtain the voice quality inspection results. Therefore, the embodiment of the present application does not need to perform quality inspection on the voice data through pre-defined sensitive words. Compared with the solution of the related art that performs voice data quality inspection based on the sensitive word library and matching rules, the quality inspection accuracy is improved to a certain extent. Moreover, the embodiment of the present application can perform quality inspection on the voice data based on the emotion recognition results of the voice data. The emotion recognition results of the voice data help to understand the illegal voice data with complex semantics. Therefore, the quality inspection accuracy of the voice data can be further improved to a certain extent.
[0049] In the related art scheme of performing voice data quality inspection based on sensitive word libraries and matching rules, regardless of whether a large model is introduced, the speaker's emotional dimension in the audio file is directly ignored. In Chinese semantics, a sentence expressed with different tones and emotions sometimes conveys completely opposite meanings. By introducing an emotion recognition model to identify the speaker's emotional information, contextual text information can be constructed for input into the text quality inspection model, thereby effectively improving the quality inspection accuracy and recall rate. For example, the quality inspection accuracy rate can be increased to 91%, and the recall rate can be increased to 97%. The embodiment of the present application can perform full-text semantic understanding of the current voice data based on the text quality inspection model. By identifying deeper illegal content, it can effectively capture potential risk behaviors, and there is no need to pre-configure a large number of similar questions, solving the occurrence of missed detections due to incomplete sensitive word settings, thereby improving the accuracy of text quality inspection.
[0050] Compared with the technical solutions of the related art, the technical solutions of the embodiments of the present application have at least the following advantages:
[0051] 1) Complex Semantic Understanding: Compared to related technical solutions, our text quality inspection model can understand and process the nuances and complex semantics of natural language. It goes beyond literal meaning and captures context, linguistic context, and implicit meaning, thereby identifying violations with cryptic expressions and effectively avoiding missed detections. It also improves its ability to understand polysemous words and long sentences, maintaining high accuracy in complex conversations. It performs particularly well in speech detection in the financial sector. The advantages of complex semantic recognition help enhance the accuracy of speech data quality inspection.
[0052] 2) Multimodal Recognition: Compared to related techniques that rely solely on call text for quality inspection, the emotion recognition model implemented in this application incorporates speaker emotion analysis. This effectively captures non-textual information such as the speaker's emotions and tone, identifies potential risky emotions, and proactively identifies potential risks in scenarios involving inappropriate tone and intense emotions. The inclusion of the emotion dimension makes the speech quality inspection process more context-aware and situation-sensitive, enabling more intelligent and comprehensive quality inspection capabilities.
[0053] The embodiment of the present application can perform quality inspection on various types of voice data. For example, when the voice data is a call recording file, the technical solution of the embodiment of the present application can be used to perform a relatively accurate quality inspection on the call recording file.
[0054] In order to further improve the accuracy of emotion recognition, in some embodiments, an emotion recognition model may be pre-trained, and then the pre-trained emotion recognition model may be used to process the acoustic features of the speech data to obtain an emotion recognition result.
[0055] Among them, the emotion recognition model is trained based on audio files of multiple emotion categories, and the emotion recognition results include labels of emotion categories corresponding to the speech data.
[0056] For example, the multiple emotion types may include at least two of the following: neutral, calm, happy, sad, angry, fearful, disgusted, and surprised. The emotion recognition model may employ a model architecture such as a Bidirectional Long Short-Term Memory (BiLSTM) network. The training data for the emotion recognition model may include industry-specific characteristics. Therefore, in industry-specific call scenarios, a pre-trained emotion recognition model can more accurately identify emotions.
[0057] The training data for the emotion recognition model can be set according to actual requirements. For example, the training data for the emotion recognition model contains 8,000 audio files. The emotion categories corresponding to these training data include neutral, calm, happy, sad, angry, fearful, disgusted, and surprised. The number of audio files corresponding to each emotion category is 1,000. Each folder stores audio data of one category, and the length of each audio data is about 3 seconds. For example, the audio file corresponding to the anger category can be stored in a folder with the directory "such as dataset / audio / angry / ".
[0058] In the emotion recognition model, the emotion representation (Emotion2Vec) model or other models can be used for feature extraction. For example, the input of the Emotion2Vec model is a 16kHz audio file; the output is an emotion representation vector in the numpy (Numerical Python) format.
[0059] Reference Figure 2 In the Emotion2Vec model, knowledge transfer can be performed using a teacher network and a student network. The teacher network can include a feature extractor and a backbone network. The student network also includes a feature extractor and a backbone network. Based on the Exponential Moving Average (EMA) mechanism, the backbone network of the student network can transfer parameters to the backbone network of the teacher network, thereby updating the parameters of the teacher network. The feature extractor of the teacher network can also copy features from the feature extractor of the student network. For example, the loss function used in training the Emotion2Vec model can include a sentence-level loss function and a frame-level loss function. The sentence-level loss function aims to learn global emotional information, and the frame-level loss function is designed as a preset frame-by-frame task for learning emotional information in context.
[0060] The emotion recognition model uses a BiLSTM architecture, which includes a linear layer, a long short-term memory (LSTM) layer, a Tanh activation function, a dropout layer, a Reluctant Unified Unit (ReLU) layer, and multiple linear layers, with a total of 2,104,072 parameters. The training data consists of 8,000 audio files with a sampling rate of 16 kHz. The maximum number of training epochs is 60. The training process uses the Adam optimizer and the WarmupCosineSchedulerLR learning rate scheduler.
[0061] It can be seen that the embodiment of the present application not only recognizes complex contextual semantic information based on a large model, but also introduces an emotion recognition model to perform emotion recognition on call content, voiceprint information, and speaker tone, and determines the speaker's neutral, calm, happy, sad, angry, fearful, disgusted, and surprised emotion type score, and combines the voice text data to perform quality inspection and scoring on the voice data in multiple dimensions, thereby identifying potential violation risks and effectively improving the accuracy of voice compliance quality inspection.
[0062] In the embodiment of the present application, by combining the emotion recognition model, emotion recognition is performed on the speech text data, voiceprint information, and speaker tone, which solves the pain point problem of losing important emotion dimensions in the quality inspection of pure text dimensions. By inputting the emotion recognition results and speech text data as context information into the text quality inspection model, the violation risk of the speech data is comprehensively judged, and the information expression outside the audio text modality is effectively captured, thereby improving the quality inspection accuracy.
[0063] To more accurately reflect the details of the emotion recognition results, the emotion recognition results can also include the probability corresponding to the emotion category label. For example, if the voice data is a call recording text, the emotion recognition result can be recorded as SPECIAL_TOKEN, and the format of the emotion recognition result can be (call recording text, emotion category label, probability value).
[0064] In practical applications, the accuracy of an emotion recognition model can be evaluated using test data. This test data can include speech data and the actual emotion category labels corresponding to the speech data. After performing emotion recognition on the test data, an emotion recognition result containing the predicted emotion category labels can be obtained. For example, the recognition results of an emotion recognition model can be evaluated using 1000 pieces of test data, and a confusion matrix can be plotted to represent the difference between the actual emotion labels corresponding to the speech data and the predicted emotion category labels corresponding to the speech data.
[0065] exist Figure 3 In the figure, the horizontal axis represents the predicted emotion category label corresponding to the speech data, the vertical axis represents the actual emotion category label corresponding to the speech data, and the numbers in the box represent the probability of the predicted emotion category label. Figure 3 It can be seen that the emotion recognition model performs well in identifying emotions such as anger, fear, disgust, sadness, and surprise, and is more suitable for emotion recognition in illegal outbound call scenarios.
[0066] In an embodiment of the present application, when constructing the first prompt word, other information in addition to the emotion recognition results and voice text data can also be combined. For example, the process of constructing the first prompt word based on the emotion recognition results and voice text data includes: obtaining violation risk points, where the violation risk points represent the violation risk information existing in the voice data; then, constructing the first prompt word based on the emotion recognition results, voice text data, and violation risk points.
[0067] In practical applications, violation risk points can be obtained from a pre-built knowledge base, where various types of violation risk points can be pre-set based on prior knowledge. Violation risk points can also be extracted from speech and text data. In this way, the first prompt word used in the text quality inspection model can be constructed based on the extracted violation risk points.
[0068] For example, the voice text data is recorded as A, the emotion recognition result is recorded as B, and the violation risk point is recorded as C. A can be "The above content is a bank collection call", B can include the following information: the customer's emotion is SPECIAL_TOKEN (call text, emotion classification label, probability value), and the customer service's emotion is SPECIAL_TOKEN (call text, emotion classification label, probability value); C can include the following information: whether the customer is at risk of telephone fraud during the call, if so, please answer yes and give the specific risk points.
[0069] It can be seen that the embodiment of the present application can construct prompt words that reflect richer information based on emotion recognition results, voice text data and violation risk points. Therefore, by processing the prompt words through the text quality inspection model, more accurate voice quality inspection results can be obtained.
[0070] In an embodiment of the present application, new violation risk points may be added to the knowledge base; illustratively, the extracted violation risk points may be entered into the knowledge base.
[0071] It can be understood that after the extracted violation risk points are entered into the knowledge base, prompt words can be constructed based on the extracted violation risk points, so that the prompt words of the text quality inspection model can be continuously optimized, which is conducive to improving the accuracy of the speech quality inspection results output by the text quality inspection model.
[0072] In the scenario of quality inspection of call recording files, the quality inspection solution based on keyword matching in related technologies requires users to pre-enter industry-specific violation risk knowledge and a large number of similar problems, which invisibly increases the user's usage threshold and maintenance costs and has poor scalability. The embodiment of the present application can propose a violation risk knowledge extraction capability based on a text quality inspection model. As the number of quality inspection recording files increases, the independently accumulated field risk knowledge will also be improved, lowering the user's usage threshold while forming data assets.
[0073] When extracting violation risk points from speech and text data, other information can also be combined to achieve this extraction. For example, a second prompt word can be constructed for extracting violation risk points. Then, based on the second prompt word and the speech and text data, violation risk points are generated. The second prompt word is prompt information input into the target macro model via natural language. The target macro model is used to generate a macro model of violation risk points. For example, the second prompt word can be a question, a description, or other form of input information to help the target macro model understand and execute the task of generating violation risk points.
[0074] Specifically, violation risk points can be generated based on the second prompt word and voice text data and using the target large model; here, the reason for generating violation risk points by the target large model is: since the target large model cannot directly output all violation risk points, the combination of actual users' voice text data and prompt words can achieve the effect of discovering new risks. For example: directly asking the target large model what behaviors of violent debt collection are, the direct reply of the target large model cannot cover all potential violent debt collection behaviors, but if combined with voice text data and prompt words, the target large model can summarize and extract violation risk points based on specific voice text data.
[0075] For example, in the process of generating violation risk points using the target large model, the structure of the prompt words input into the target large model may be: summarize and refine the above-mentioned violation risk points, with the summary word count not exceeding 100 words.
[0076] It can be seen that when generating violation risk points, not only the prompt words are considered, but also the voice text data. The voice text data can provide relatively rich information about the violation risk. Therefore, based on the second prompt words and voice text data, new violation risk points can be generated more accurately.
[0077] In order to further improve the accuracy of the violation risk points in the knowledge base, in some embodiments, the extracted violation risk points may be entered into the knowledge base if the extracted violation risk points pass manual review.
[0078] In the embodiment of the present application, after the violation risk points are generated, manual review of the violation risk points can also be performed through human-computer interaction. If the violation risk point fails the manual review, the process of the voice quality inspection method can be directly terminated; if the violation risk point passes the manual review, the violation risk point is entered into the knowledge base.
[0079] Here, the main role of manual review is to fine-tune the in-context-learning of the text quality inspection model. The violation risk points extracted by the target large model are sometimes not directly applicable to the samples input into the knowledge base in the context learning. When the violation risk points pass the manual review, it is necessary to determine whether the violation risk points need to be manually modified and deposited into the corresponding industry knowledge base. For the deposited violation risk points, there is no need for additional manual supplementation of specific knowledge details (this is the difference from related technologies). For example, in the fraud detection scenario, it is only necessary to clearly push financial products with excessively high yields, without specifying specific details such as "pushing financial products with an annualized yield of more than 10%" or "pushing products with a monthly yield of more than 15%" (such specific instructions are easily affected by timeliness and industry, and have weak knowledge generalization capabilities). This reduces the user's system usage cost and avoids false detection and missed detection caused by insufficient user prior knowledge or knowledge timeliness, thereby improving the accuracy of voice data.
[0080] If the violation risk points pass manual review, they can be manually annotated and fed back to the knowledge base. This allows for the creation of prompts based on the violation risk points in the knowledge base during subsequent quality checks on voice data. This creates a "data flywheel," where new violation risk points identified during quality checks on currently acquired voice data can be used to conduct quality checks on the next batch of acquired voice data.
[0081] In an embodiment of the present application, the violation risk points that have passed manual review are entered into the knowledge base. Based on the violation risk points in the knowledge base, prompt words can be accurately constructed for quality inspection of subsequently acquired voice data, thereby improving the quality inspection accuracy of the text quality inspection model on subsequently acquired voice data to a certain extent.
[0082] Based on the above records, it can be seen that in the embodiment of the present application, the complex semantic understanding ability of the fine-tuned text quality inspection model can be used to solve the problems of low quality inspection accuracy and weak generalization ability of the keyword matching quality inspection method in the related technology, effectively avoid the occurrence of illegal outbound call personnel bypassing the loopholes of the traditional quality inspection system to implement telephone fraud, violent collection and other behaviors, identify potential risks, and improve the quality inspection accuracy of voice data. By combining the emotion recognition model, the pain point problem of the loss of important emotional dimensions in the pure text dimension quality inspection is solved. By inputting the results of speech recognition as context information into the text-to-text model, the information of the prompt words input into the text quality inspection model can be enriched, and then the violation risks existing in the voice data can be comprehensively judged, and the information expression outside the audio text modality can be effectively captured, thereby improving the quality inspection accuracy of voice data. The text quality inspection model is used to perform full-text semantic understanding and violation risk summary on the voice data. By comparing with the knowledge base, the large model summarizes new violation risk points, and the violation risk points that have been manually reviewed are entered into the corresponding knowledge base to achieve data element precipitation, solve the problem of inaccurate voice quality inspection results due to low training data quality of the text quality inspection model, and realize data flywheel.
[0083] Figure 4 The flowchart of the quality inspection of the recording file provided in the embodiment of the present application is as follows: Figure 4 As shown, the process includes:
[0084] Step 41: Get the call recording file.
[0085] Here, the call recording file can be automatically imported from the outbound call system. For example, the voice quality inspection method of the embodiment of the present application can be implemented by a voice quality inspection device, which is usually used in combination with the customer's outbound call system.
[0086] Step 42: Perform voice recognition on the call recording file to obtain voice text data.
[0087] Step 43: Perform emotion recognition on the call recording file to obtain an emotion recognition result.
[0088] In the embodiment of the present application, the execution order of step 42 and step 43 is not limited. Step 42 and step 43 can be executed simultaneously, or step 42 and step 43 can be executed in sequence.
[0089] Step 44: Construct a first prompt word based on the emotion recognition result, the voice text data, and the violation risk point.
[0090] Here, violation risk points in natural language form can be obtained from the knowledge base, which is different from the construction of sensitive word libraries in related technologies. The sensitive word libraries in related technologies are mainly constructed through keywords. For example, the user defines an affirmative intent and needs to fill in all possible keywords in the affirmative intent, such as: yes, good, um, ok; it can be seen that the maintenance cost of the sensitive word library is high and the generalization ability is weak, and it is difficult to migrate to other application scenarios. However, the embodiment of the present application can use the fine-tuned text quality inspection model to effectively identify complex semantics and realize the construction of prompt words through violation risks expressed in natural language.
[0091] Step 45: Use the text quality inspection model to process the first prompt word and output the speech quality inspection result.
[0092] For example, the voice quality inspection result may be: Yes, there is a high risk of telephone fraud in the above conversation, and the customer service staff has violated the regulations in guiding users.
[0093] Step 46: Based on the voice quality inspection results, determine whether there are new violation risk points in the voice data. If so, execute step 47; if not, end the process.
[0094] In some embodiments, prompt words can be constructed in combination with the risk points existing in the voice quality inspection results, so that the target large model can summarize the new violation risk points through the prompt words. The structure of the prompt words can be: In addition to the above-mentioned risk points, are there any new violation risk points? If so, please give reasons. If not, return no. After obtaining the output results of the large model, use regular expressions to judge the matching status of the output results of the large model with the keywords. When the output results of the target large model match the keyword "no", it is determined that there are no new violation risk points in the voice data, and the process is terminated directly; when the output results of the target large model match the keyword "no", it is determined that there are new violation risk points in the voice data. At this time, a second prompt word can be constructed to extract the new violation risk points. The second prompt word is processed by the target large model to generate a new violation risk point, and then step 47 is executed.
[0095] Step 47: Determine whether the new violation risk point has passed manual review. If the new violation risk point has passed manual review, execute step 48. If the new violation risk point has not passed manual review, end the process directly.
[0096] Step 48: Enter the new violation risk point into the knowledge base.
[0097] Steps 46 to 48 can be used to solve the hallucination problem of the text quality inspection model. In an embodiment of the present application, the complex semantic understanding ability and industry knowledge of the text quality inspection model are utilized. For each call recording file, a higher-dimensional large model is used to summarize new violation risk points, and the text quality inspection model is fine-tuned through context learning through manual review to achieve a data flywheel. That is, the embodiment of the present application can achieve context learning fine-tuning of the text quality inspection model through actual call files and manual review processes without changing the parameters of the text quality inspection model, thereby reducing the hallucination problem of the large model.
[0098] After step 48, for subsequent call recording files obtained again, the updated violation risk points in the knowledge base can be used to perform voice quality inspection. In this way, the actual user's call recording files can be used to continuously refine new violation risk points, thereby continuously improving the quality inspection accuracy of the text quality inspection model.
[0099] In summary, in the embodiment of the present application, the call recording files of the outbound call system can be automatically imported, and various complex emotions can be recognized through voice recognition and acoustic feature analysis. Then, prompt words are constructed to recognize complex semantics, and the text quality inspection model is called to perform violation risk analysis. Through manual review, context learning is fine-tuned and new violation risk points are summarized. The knowledge base is optimized, the data flywheel is realized, the quality inspection accuracy of the voice data is improved, and efficient recognition in the illegal outbound call scenario is ensured. Finally, a complete voice quality inspection process is formed to effectively identify potential violation risks.
[0100] Those skilled in the art will understand that in the above-mentioned method of the specific implementation method, the writing order of each step does not mean a strict execution order and does not constitute any limitation on the implementation process. The specific execution order of each step should be determined by its function and possible internal logic.
[0101] Figure 5 This is a structural diagram of the speech quality inspection device according to an embodiment of the present application. Figure 5 As shown, the device includes:
[0102] Acquisition module 501, used to acquire voice data;
[0103] The first processing module 502 is configured to perform emotion recognition on the speech data based on the acoustic features of the speech data to obtain an emotion recognition result; and perform speech recognition on the speech data to obtain speech text data;
[0104] The second processing module 503 is configured to construct a first prompt word according to the emotion recognition result and the speech text data, and process the first prompt word using a pre-trained text quality inspection model to obtain a speech quality inspection result.
[0105] In some embodiments, the first processing module 502 performs emotion recognition on the speech data based on the acoustic features of the speech data to obtain an emotion recognition result, including: using a pre-trained emotion recognition model to process the acoustic features of the speech data to obtain the emotion recognition result; the emotion recognition model is trained based on audio files of multiple emotion categories, and the emotion recognition result includes a label of the emotion category corresponding to the speech data.
[0106] In some embodiments, the second processing module 503 constructs a first prompt word based on the emotion recognition result and the voice text data, including: obtaining a violation risk point, where the violation risk point represents the violation risk information present in the voice data; and constructing the first prompt word based on the emotion recognition result, the voice text data, and the violation risk point.
[0107] In some embodiments, the second processing module 503 , obtaining the violation risk points includes: obtaining the violation risk points from a knowledge base, or extracting the violation risk points from the voice text data.
[0108] In some embodiments, the second processing module 503 is further configured to, after extracting the violation risk points from the speech text data, enter the extracted violation risk points into the knowledge base.
[0109] In some embodiments, the second processing module 503 extracts the violation risk point from the voice text data, including: constructing a second prompt word for extracting the violation risk point; and generating the violation risk point based on the second prompt word and the voice text data.
[0110] In some embodiments, the second processing module 503 enters the extracted violation risk points into the knowledge base, including: if the extracted violation risk points pass manual review, entering the extracted violation risk points into the knowledge base.
[0111] In practical applications, the acquisition module 501 , the first processing module 502 and the second processing module 503 may be implemented based on a processor.
[0112] It should be noted that the description of the above device embodiment is similar to the description of the above method embodiment and has similar beneficial effects as the method embodiment. For technical details not disclosed in the device embodiment of this application, please refer to the description of the method embodiment of this application for understanding.
[0113] It should be noted that, in the embodiment of the present application, if the above method is implemented in the form of a software function module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiment of the present application is essentially or the part that contributes to the prior art can be embodied in the form of a software product, which is stored in a storage medium and includes a number of instructions for enabling a computer device (which can be a terminal, server, etc.) to execute all or part of the methods described in each embodiment of the present application. The aforementioned storage medium includes various media that can store program codes, such as a U disk, a mobile hard disk, a read-only memory (ROM), a magnetic disk or an optical disk. In this way, the embodiment of the present application is not limited to any specific combination of hardware and software.
[0114] Correspondingly, an embodiment of the present application further provides a computer program product, which includes computer-executable instructions, and the computer-executable instructions are used to implement any one of the speech quality inspection methods provided in the embodiments of the present application.
[0115] Accordingly, an embodiment of the present application further provides a computer storage medium, on which computer executable instructions are stored. The computer executable instructions are used to implement any one of the speech quality inspection methods provided in the above embodiments.
[0116] An embodiment of the present application also provides an electronic device. Figure 6 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application is shown in FIG. Figure 6 As shown, the electronic device 60 may include:
[0117] Memory 601, used to store executable instructions;
[0118] The processor 602 is configured to implement any of the above-mentioned speech quality inspection methods when executing the executable instructions stored in the memory 601.
[0119] The processor 602 may be at least one of an ASIC, a DSP, a DSPD, a PLD, an FPGA, a CPU, a controller, a microcontroller, and a microprocessor.
[0120] The above-mentioned computer-readable storage medium and memory 601 can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a magnetic random access memory (FRAM), a flash memory (Flash Memory), a magnetic surface memory, an optical disc, or a compact disc read-only memory (CD-ROM); it can also be various terminals including one or any combination of the above-mentioned memories, such as mobile phones, computers, tablet devices, personal digital assistants, etc.
[0121] In some embodiments, the functions or modules included in the device provided in the embodiments of the present application can be used to execute the method described in the above method embodiments. The specific implementation can refer to the description of the above method embodiments. For the sake of brevity, it will not be repeated here.
[0122] The above description of the various embodiments tends to emphasize the differences between the various embodiments. The same or similar aspects can be referenced with each other and will not be repeated herein for the sake of brevity.
[0123] The methods disclosed in the various method embodiments provided in this application can be arbitrarily combined without conflict to obtain new method embodiments.
[0124] The features disclosed in the various product embodiments provided in this application can be arbitrarily combined without conflict to obtain new product embodiments.
[0125] The features disclosed in the various method or device embodiments provided in this application can be arbitrarily combined without conflict to obtain new method embodiments or device embodiments.
[0126] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus the necessary general hardware platform, and of course can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for enabling a terminal (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in each embodiment of the present application.
[0127] The embodiments of the present application are described above in conjunction with the accompanying drawings, but the present application is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the guidance of this application, ordinary technicians in this field can also make many forms without departing from the purpose of this application and the scope of protection of the claims, all of which are protected by this application.
Claims
1. A method for speech quality inspection, characterized in that: The method comprises: Get voice data; Performing emotion recognition on the voice data based on the acoustic features of the voice data to obtain an emotion recognition result; performing speech recognition on the voice data to obtain voice text data; A first prompt word is constructed according to the emotion recognition result and the speech text data, and the first prompt word is processed using a pre-trained text quality inspection model to obtain a speech quality inspection result.
2. The method according to claim 1, characterized in that The performing emotion recognition on the speech data based on the acoustic features of the speech data to obtain an emotion recognition result includes: The acoustic features of the speech data are processed using a pre-trained emotion recognition model to obtain the emotion recognition result; the emotion recognition model is trained based on audio files of multiple emotion categories, and the emotion recognition result includes a label of the emotion category corresponding to the speech data.
3. The method according to claim 1 or 2, characterized in that The step of constructing a first prompt word according to the emotion recognition result and the voice text data includes: Obtaining a violation risk point, where the violation risk point represents violation risk information present in the voice data; The first prompt word is constructed according to the emotion recognition result, the voice text data and the violation risk point.
4. The method according to claim 3, characterized in that The obtaining of the violation risk points includes: obtaining the violation risk points from a knowledge base, or extracting the violation risk points from the voice text data.
5. The method according to claim 4, characterized in that After extracting the violation risk points from the voice text data, the method further includes: entering the extracted violation risk points into the knowledge base.
6. The method according to claim 4, characterized in that The extracting the violation risk point from the voice and text data includes: Constructing a second prompt word for extracting violation risk points; The violation risk point is generated based on the second prompt word and the voice text data.
7. The method according to claim 5, characterized in that The step of entering the extracted violation risk points into the knowledge base includes: When the extracted violation risk points pass manual review, the extracted violation risk points are entered into the knowledge base.
8. A speech quality inspection device, characterized in that: The device comprises: An acquisition module, used to acquire voice data; A first processing module is configured to perform emotion recognition on the speech data based on acoustic features of the speech data to obtain an emotion recognition result; and perform speech recognition on the speech data to obtain speech text data; The second processing module is used to construct a first prompt word according to the emotion recognition result and the voice text data, and process the first prompt word using a pre-trained text quality inspection model to obtain a voice quality inspection result.
9. An electronic device, characterized in that: The electronic device comprises a processor and a memory for storing a computer program that can be run on the processor; wherein, The processor is configured to run the computer program to execute the speech quality inspection method according to any one of claims 1 to 7.
10. A computer storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the speech quality inspection method according to any one of claims 1 to 7 is implemented.
11. A computer program product comprising a computer program, characterized in that When executed by a processor, the computer program implements the speech quality inspection method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Collection method and system based on emotion recognition, computer equipment and storage medium
CN115249481A
Voice detection method, device and equipment and readable storage medium
CN115602153A
Customer service dialogue quality inspection method and device, computer equipment and storage medium
CN119179777A
Processing method and device for dialogue type audio, equipment and storage medium
CN119943049A
Cited By
Intelligent voice quality evaluation method and device for multi-mode intelligent terminal
CN120913599A