Voice quality inspection method and device, electronic equipment and storage medium
By adopting a text translation model of a bidirectional long and short-term memory network and a multi-head attention mechanism in the voice collection system, combined with the matching models of CNN, S-LSTM and DNN, the problems of poor recognition efficiency and lack of sentiment analysis in the voice collection system are solved, and more efficient and accurate voice quality inspection is achieved.
Patent Information
- Application Number
- CN202411949267.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-27
- Publication Date
- 2025-05-06
AI Technical Summary
The existing voice collection system has problems such as poor recognition efficiency, inapplicability of complex scenarios, and lack of sentiment analysis, resulting in low quality inspection accuracy and inefficiency.
By obtaining voice data, text translation is performed using the first model of the bidirectional long and short-term memory network and the multi-head attention mechanism to extract emotional characteristics and scene characteristics; then, based on the second model of CNN, S-LSTM and DNN, the translated text and preset quality inspection rules are matched to obtain voice quality inspection results. At the same time, the original voice, the speaker's separate voice and the voice with the annotated keywords are used for manual re-checking, and the feedback guidance model training is used.
It improves the accuracy and efficiency of voice quality inspection, enhances the system's adaptability to complex scenarios, improves the accuracy of sentiment analysis, and saves human resources.
Smart Images

Figure CN119943092A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of voice debt collection and voice quality inspection, and in particular to a voice quality inspection method, device, electronic device, and storage medium. Background Art
[0002] Voice collection is a management system that uses modern communication technology and artificial intelligence technology to automatically notify and collect overdue bills. It is usually used in banks, financial institutions and other companies or institutions that need to collect overdue payments.
[0003] The bank voice collection intelligent quality inspection system in the related art matches the voice-converted text with the set keywords and returns the number and content of rules corresponding to the hit keywords. However, the voice collection system still has problems such as poor recognition efficiency, unsuitability for complex scenarios, and lack of sentiment analysis. Summary of the invention
[0004] The embodiments of the present application provide a voice quality inspection method, device, electronic device, and storage medium to improve the accuracy of quality inspection during voice collection.
[0005] The present application embodiment adopts the following technical solutions:
[0006] In a first aspect, an embodiment of the present application provides a method for speech quality inspection, wherein the method comprises:
[0007] Get voice data;
[0008] Inputting the speech data into a first model to obtain translated text content;
[0009] Based on the second model, the translated text content is matched with the preset quality inspection rules to obtain the speech quality inspection result and the matching keyword text;
[0010] According to the matching keyword text, obtaining the speech marked with the keyword;
[0011] According to the voice data, obtaining the original voice and the speaker separated voice;
[0012] Based on the speech quality inspection result, the speech with the keyword annotated, the original speech and the speaker separated speech are used to perform speech quality inspection to obtain a final speech quality inspection result.
[0013] In some embodiments, inputting the speech data into a first model to obtain translated text content includes:
[0014] The speech data is pre-processed and then input into the first model;
[0015] According to the first model, outputting the translated text content including the sentiment features and the scene features;
[0016] in,
[0017] The first model includes a bidirectional long short-term memory network and a multi-head attention mechanism.
[0018] In some embodiments, the step of matching the translated text content with a preset quality inspection rule based on the second model to obtain a speech quality inspection result and a matching keyword text includes:
[0019] According to the second model, after matching with the preset quality inspection rules, the matching keyword text related to the translated text content containing the emotional features and the scene features is obtained;
[0020] Using the output of the second model as the speech quality inspection result;
[0021] Among them, the second model includes a CNN network, an S-LSTM network and a DNN network.
[0022] In some embodiments, based on the speech quality inspection result, the speech with the keyword marked, the original speech and the speaker separated speech are used to perform speech quality inspection to obtain the final speech quality inspection result, including:
[0023] The original speech, the speaker separated speech and the speech assistance with the marked keywords are used to perform manual re-inspection to obtain a final speech quality inspection result.
[0024] In some embodiments, the method further comprises:
[0025] Based on the quality inspection data of the manual review, feedback is provided to guide the training of the second model.
[0026] In some embodiments, obtaining a speech marked with a keyword according to the matching keyword text includes:
[0027] Get the voice segment of the matching keyword text:
[0028] According to the speech segment of the matching keyword text, the keyword text obtained by intelligent matching is mapped to the original speech frame by frame to obtain the speech with the keyword marked.
[0029] In some embodiments, obtaining the original voice and the speaker-separated voice according to the voice data includes:
[0030] Segmenting the speech data to obtain segmented speech segments;
[0031] Extracting speaker features from the segmented speech segments;
[0032] According to the speaker characteristics, the speech segments are segmented and clustered by multiple speakers to obtain speech segmentation and clustering results;
[0033] According to the speech segmentation and clustering results, different speakers are separated to obtain speaker-separated speech.
[0034] In a second aspect, an embodiment of the present application further provides a speech quality inspection device, wherein the device comprises:
[0035] An acquisition module, used for acquiring voice data;
[0036] A text translation module, used for inputting the speech data into a first model to obtain translated text content;
[0037] A matching module, used to match the translated text content with preset quality inspection rules based on the second model to obtain a speech quality inspection result and a matching keyword text;
[0038] A marking module, used for obtaining speech marked with keywords according to the matching keyword text;
[0039] A processing module, used for obtaining original speech and speaker-separated speech according to the speech data;
[0040] The quality inspection module is used to perform voice quality inspection based on the voice quality inspection result by using the voice marked with keywords, the original voice and the speaker separated voice to obtain a final voice quality inspection result.
[0041] In a third aspect, an embodiment of the present application further provides an electronic device, comprising: a processor; and a memory arranged to store computer executable instructions, wherein the executable instructions, when executed, cause the processor to perform the above method.
[0042] In a fourth aspect, an embodiment of the present application further provides a computer-readable storage medium, which stores one or more programs. When the one or more programs are executed by an electronic device including multiple application programs, the electronic device executes the above method.
[0043] At least one of the above technical solutions adopted in the embodiments of the present application can achieve the following beneficial effects: by acquiring voice data; inputting the voice data into the first model to obtain the translated text content. Based on the second model, the translated text content is matched with the preset quality inspection rules to obtain the voice quality inspection results and the matching keyword text. Finally, based on the voice quality inspection results, the voice marked with keywords, the original voice and the separated voice of the speaker are used to perform voice quality inspection to obtain the final voice quality inspection results. Through the above method, not only a quality inspection model with higher accuracy is provided, but also manual re-inspection is performed through three kinds of voices, which can improve the efficiency of manual re-inspection and save manpower. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:
[0045] Figure 1 This is a schematic diagram of the overall implementation process of the voice quality inspection method in the embodiment of the present application;
[0046] Figure 2 This is a flow chart of the speech quality inspection method in the embodiment of the present application;
[0047] Figure 3 This is a schematic diagram of the structure of a speech quality inspection device in an embodiment of the present application;
[0048] Figure 4 This is a schematic diagram of the network structure of the first model of the speech quality inspection method in an embodiment of the present application;
[0049] Figure 5 This is a schematic diagram of the network structure of the second model of the voice quality inspection method in an embodiment of the present application;
[0050] Figure 6 This is a schematic diagram of the structure of an electronic device in an embodiment of the present application. DETAILED DESCRIPTION
[0051] In order to make the purpose, technical solution and advantages of the present application clearer, the technical solution of the present application will be clearly and completely described below in combination with the specific embodiments of the present application and the corresponding drawings. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present application.
[0052] The existing intelligent quality inspection system for bank voice collection matches the voice-converted text with the set keywords and returns the number and content of rules corresponding to the hit keywords. However, there are the following problems:
[0053] (1) Voice quality inspection is more dependent on voice data. If the data quality is not high, such as unclear voice, too fast or too slow speech speed, etc., the recognition rate will be very poor.
[0054] (2) In complex scenarios, the recognition rate may be affected. For example, problems such as multi-person conversations and dialects will have a significant impact on the recognition rate.
[0055] (3) In the process of sentiment analysis, inaccuracies may occur. For example, incorrect analysis of certain emotional sentences will cause errors and misjudgments in subsequent detection and analysis.
[0056] (4) Quality inspection experience cannot be effectively recorded and fed back. Since the current system recognition rate is poor, each intelligent quality inspection result needs to be manually reviewed if it hits the corresponding rules to ensure the accuracy of the overall quality inspection result. At the same time, since manual review can only listen to the entire speech, it is more time-consuming and labor-intensive.
[0057] In response to the above problems, in some solutions, voice data to be inspected is received, and the voice data to be inspected is converted into text to obtain text data to be inspected; the text data to be inspected is identified to determine whether it has various preset quality inspection items to obtain a target recognition result; and the voice data to be inspected is evaluated based on the target recognition result. Specifically, a feature matrix is extracted from the text data using a convolutional layer and a pooling layer, and then an activation layer is used to predict the feature matrix to obtain a predicted set of quality inspection element labels. Although the convolutional neural network used has good local perception capabilities for text data, it is difficult to effectively process time series data.
[0058] In other schemes, a voice file to be inspected is obtained; the voice file to be inspected is converted into a text to be inspected by automatic speech recognition technology; the file to be inspected corresponding to the selected detection dimension is determined for analysis and calculation to obtain quality inspection result data, wherein the file to be inspected includes the voice file to be inspected and the text to be inspected. The voice file to be inspected is converted into a text to be inspected by automatic speech recognition technology, the sentence vector is obtained by using the LTSM network model, and supervised training is performed using a machine learning classification model until the model converges. Although the use of the LSTM model can better capture longer-distance dependencies, it cannot encode information from back to front.
[0059] like Figure 1 FIG. 1 is a schematic diagram of an implementation process of an intelligent speech quality inspection method based on intelligent speech recognition in an embodiment of the present application, which specifically includes the following implementation process:
[0060] First, the input speech signal is preprocessed, including but not limited to speech signal noise removal, speech signal enhancement, and the like.
[0061] Secondly, a model that combines the bidirectional long short-term memory network Bi-LSTM and the Transformer multi-head attention mechanism is used to train the preprocessed speech data, emotional data, and scene data to obtain the translated text containing emotional and scene features.
[0062] Finally, the CNN+S-LSTM+DNN model is used to intelligently match the translated text with the set quality inspection rules to obtain the quality inspection results and the text matching the keywords.
[0063] Based on the above quality inspection results, when sampling for manual re-inspection, the original voice signal is first separated by speaker to obtain the voice parts of the debt collector and the customer. Then the text of the matching keywords obtained by intelligent matching is mapped frame by frame to the original voice to obtain the voice marked with keywords. Finally, the original voice, the voice separated by speaker, and the voice marked with keywords are used to assist in manual quality inspection to obtain the final manual quality inspection results. In addition, the quality inspection experience of manual quality inspection is used as feedback to guide the training of the intelligent matching model. In order to obtain more accurate quality inspection results.
[0064] The technical solutions provided by various embodiments of the present application are described in detail below in conjunction with the accompanying drawings.
[0065] The present application embodiment provides a method for voice quality inspection, such as Figure 2 As shown, a flow chart of a method for speech quality inspection in an embodiment of the present application is provided, and the method at least includes the following steps S210 to S260:
[0066] Step S210, acquiring voice data.
[0067] The acquired voice data mainly includes but is not limited to the original voice signal, which can usually be pre-processed to obtain normalized data, so as to facilitate input into the model. In particular, the voice data includes the voice signal of voice debt collection.
[0068] Step S220: input the speech data into the first model to obtain translated text content.
[0069] The speech data is input into the first model, and the translated text content corresponding to the speech data is obtained through the first model.
[0070] It should be noted that the first model is obtained by training one or more networks to translate voice data into translated text content, that is, converting voice signals into text content.
[0071] Preferably, if Figure 1 As shown, the translated text content includes emotional features and scene features. It can be understood that the emotional features are used to represent the emotional information in the translated text content, such as recording emotions such as happiness, frustration, and excitement. The scene features are used to represent the scene information in the translated text content, such as recording environments such as indoors and outdoors, in a car, and in an office.
[0072] Step S230, based on the second model, matching the translated text content with preset quality inspection rules to obtain speech quality inspection results and matching keyword texts.
[0073] The second model is also obtained by training one or more networks to match the translated text content with the preset quality inspection rules, and obtain a voice quality inspection result while matching the keyword text. The matching keyword text is used to characterize the degree of matching with the preset quality inspection rules. It can be understood that the higher the matching degree, the more it meets the preset quality inspection rules. Similarly, the lower the matching degree, the less it meets the preset quality inspection rules.
[0074] Step S240, obtaining the speech marked with the keyword according to the matching keyword text.
[0075] The text of the matching keywords obtained by intelligent matching is mapped frame by frame to the original speech to obtain the speech marked with the keywords.
[0076] Step S250, obtaining the original voice and the speaker separated voice according to the voice data.
[0077] The original voice can be obtained according to the voice data, and the speaker separation operation is performed on the original voice to obtain the speaker separation voice. It should be noted that the execution order of step S250 is not required to be performed after step S240, and can be performed after step S210 or in parallel with step S240.
[0078] Step S260: Based on the speech quality inspection result, speech quality inspection is performed using the speech with the keyword annotated, the original speech and the speaker separated speech to obtain a final speech quality inspection result.
[0079] Based on the speech quality inspection results output by the model, manual quality inspection is performed using the original speech, speaker-separated speech, and speech-assisted speech with keyword annotations to obtain the final manual quality inspection results.
[0080] Preferably, the model is optimized again according to the manual quality inspection results to improve the matching effect of the model.
[0081] Through the above method, the first model is used to obtain the translated text content, and based on the second model, the translated text content is matched with the preset quality inspection rules to obtain the voice quality inspection results and the matching keyword text. Through the first model and the second model, the process of converting voice data to translated text content and then matching the translated text content is realized, thereby improving the accuracy of the voice quality inspection process.
[0082] Through the above method, based on the speech quality inspection result, the speech with the keyword annotated, the original speech and the speaker separated speech are used to perform speech quality inspection to obtain the final speech quality inspection result. By providing three kinds of speech, namely the original speech, the speaker separated speech and the speech with the keyword annotated to assist in manual re-inspection, the efficiency of manual re-inspection can be improved and manpower can be saved. And the manual re-inspection result is fed back to the intelligent matching model training to improve the model matching accuracy.
[0083] Different from the related art, voice quality inspection is more dependent on voice data. If the data quality is not high, the recognition rate will be very poor. Through the above method, the voice data is input into the first model to obtain the translated text content, and the output of the first model contains the translated text with emotional features and scene features.
[0084] Different from related technologies, in complex scenarios, there may be incompatibility. And inaccuracies may occur when performing sentiment analysis. Through the above method, the enhanced voice signal is used to extract emotional features, extract the tone, intonation, speech speed and other feature information of the voice signal, and then perform feature analysis to extract the emotional label of the voice signal; the scene features are used to distinguish the recording environments such as indoors and outdoors, in the car, and in the office, so as to enhance the system's adaptability to complex scenarios.
[0085] Different from the problem in related technologies that quality inspection experience cannot be effectively recorded and fed back, the above method feeds back the manual re-inspection results to the intelligent matching model training to improve the model matching accuracy.
[0086] In one embodiment of the present application, the step of inputting the speech data into a first model to obtain translated text content includes: the speech data is preprocessed and then input into the first model; according to the first model, the translated text content including emotional features and scene features is output; wherein the first model includes a bidirectional long short-term memory network and a multi-head attention mechanism.
[0087] like Figure 4 As shown in the figure, Bi-LSTM is used when obtaining text. While capturing short-term memory dependencies, it can better capture bidirectional semantic dependencies while taking into account past and future information. S-LSTM is used for intelligent matching. S-LSTM adds global nodes and uses the same gate rules as LSTM to update the information of global nodes. The information of global nodes is also used to exchange information with each word node in the sentence.
[0088] like Figure 4As shown in the figure, during the model training phase, a model that integrates the bidirectional long short-term memory network Bi-LSTM and the Transformer multi-head attention mechanism is used to train the preprocessed speech data, emotional data, and scene data, and obtains the translated text containing emotional and scene features.
[0089] It should be noted that BiLSTM (Bidirectional Long Short-Term Memory Network) is a commonly used neural network model for processing time series data and has good prediction performance. In addition to capturing the local dependencies of the sequence, the introduction of the gating mechanism can also better capture the long-term dependencies in the sequence.
[0090] like Figure 4 As shown in Figure 2, Transformer is a neural network model based on the attention mechanism, which is widely used in the field of natural language processing. Transformer consists of Encoder and Decoder, which are used to process input sequences and output sequences respectively. It uses the self-attention mechanism to perform linear transformations in multiple subspaces to obtain multiple representations.
[0091] Preferably, Bi-LSTM, as a bidirectional LSTM, can better capture bidirectional semantic dependencies while capturing short-term memory dependencies, while Transformer enhances the ability to capture long-term historical dependencies. Therefore, after the two are combined, they can better obtain long-term and short-term historical memories, thereby improving the accuracy of the translated text. The preprocessed speech, emotion data, and scene data are input into the Bi-LSTM+Transformer model, and after multiple rounds of training, the translated text containing emotion features and scene features is obtained.
[0092] Through the above method, the input data used to train the model has richer dimensions. It not only takes into account the emotional information of the customer's voice, but also uses scene features to distinguish between recording environments such as indoors and outdoors, in cars, and offices, enhancing the system's adaptability to complex scenes. Artificial intelligence voice quality inspection can detect the service quality level and emotional changes of customer service representatives through word count (word count / time), volume, channel, number of fluctuations, and call silence. Artificial intelligence voice quality inspection can segment the conversation between customer service and customers into scenes to conduct data analysis.
[0093] In addition, the above network model can also extract scene features to distinguish recording environments such as indoors and outdoors, in cars, and in offices, enhancing the system's adaptability to complex scenes, such as multi-person conversations, dialect accents, etc. That is, the environmental features of debt collection.
[0094] The above network model can also combine sentiment lexicon and semantic analysis to improve the accuracy and reliability of sentiment analysis and more accurately identify sentiment tendencies and emotional states, i.e., the emotional characteristics during debt collection.
[0095] In one embodiment of the present application, based on the second model, the translated text content is matched with preset quality inspection rules to obtain speech quality inspection results and matching keyword texts, including: according to the second model, after matching with preset quality inspection rules, the matching keyword text related to the translated text content containing emotional features and scene features is obtained; the output of the second model is used as the speech quality inspection result; wherein the second model includes a CNN network, an S-LSTM network and a DNN network.
[0096] The matching translated text content is input into the trained CNN+S-LSTM+DNN network model, and the quality inspection results and the text matching the keywords are output in combination with the quality inspection rules.
[0097] Figure 5 As shown in the figure, in the intelligent matching stage, CNN (convolutional neural network) is used to extract more adaptive features and input them into S-LSTM (Sentence-State LSTM, S-long short-term memory network). S-LSTM combines the original time series information to process the high-level features of CNN input. Finally, DNN (deep neural network) increases the depth between the hidden layer and the output layer, and performs deeper processing on the features processed by S-LSTM, thereby obtaining stronger prediction capabilities.
[0098] Considering that LSTM solves the problem of long-term dependency, it cannot utilize the context information of the text. BiLSTM (bidirectional long short-term memory network) also considers the context, but is limited to sequential text. S-LSTM does not read the word sequence incrementally, but transmits information layer by layer according to the distance relationship of words in continuous iteration. That is, S-LSTM adds global nodes to LSTM and uses the same gate rules as LSTM to update the information of global nodes. The information of global nodes is also used to exchange information with each word node in the sentence, which ultimately reduces the complexity of LSTM while improving performance.
[0099] The above network structure is designed based on the representation model. Represent-base Model is an important type of model in natural language processing. Its main idea is to perform downstream tasks by learning the mapping from raw language input to semantic vector representation.
[0100] It should be noted that the quality inspection rules include but are not limited to call quality, that is, whether the collection personnel's call is professional, polite, and patient, whether the speaking speed is too fast or too slow, whether any inappropriate collection methods are used, etc.
[0101] It should be noted that quality inspection rules include but are not limited to case handling: whether the debt collector understands the circumstances of the case, whether the correct debt collection process has been adopted, whether the company and legal regulations have been complied with, etc.
[0102] In one embodiment of the present application, based on the speech quality inspection result, speech quality inspection is performed using the speech with the marked keywords, the original speech and the speaker separated speech to obtain the final speech quality inspection result, including: manual re-inspection is performed using the original speech, the speaker separated speech and the speech with the marked keywords to assist in obtaining the final speech quality inspection result.
[0103] The original speech, the speech separated from the speaker, and the speech assisted by the keyword annotation are used for manual quality inspection to obtain the final manual quality inspection result.
[0104] In one embodiment of the present application, the method further includes: providing feedback to guide the training of the second model based on the quality inspection data of the manual review.
[0105] The quality inspection experience of manual quality inspection is used to guide the training of the intelligent matching model, so as to obtain more accurate quality inspection results. The feedback of manual re-inspection experience is used to guide the training of the quality inspection model.
[0106] In one embodiment of the present application, obtaining the speech marked with the keyword based on the matching keyword text includes: obtaining the speech segment of the matching keyword text: based on the speech segment of the matching keyword text, using intelligent matching to obtain the keyword text and mapping it to the original speech frame by frame to obtain the speech marked with the keyword.
[0107] Obtain the keyword speech segment, use the keyword text obtained by intelligent matching, map it to the original speech frame by frame, and obtain the speech segment marked with the keyword.
[0108] In one embodiment of the present application, the method of obtaining original speech and speaker-separated speech based on the speech data includes: segmenting the speech data to obtain segmented speech segments; extracting speaker features in the segmented speech segments; performing multi-speaker segmentation and clustering on the speech segments based on the speaker features to obtain speech segmentation and clustering results; and separating different speakers based on the speech segmentation and clustering results to obtain speaker-separated speech.
[0109] The process of speaker separation is as follows: S1, voice activity detection (VAD), using the DNN-VAD (deep neural network-based voice activity detection) algorithm to segment the speech segment. S2, speaker feature extraction, using Transformer to extract speaker features from the segmented speech segment. S3, clustering to distinguish speakers, using the K-means clustering algorithm to segment and cluster the speech segments into multiple speakers. S4, speaker label separation, using the DNN deep model to separate speaker labels, perform speaker identification and segmentation, and finally obtain the separated speech segments of the debt collector and the customer.
[0110] Through the above method, the speaker separation technology is used to separate the "collector" voice from the "customer" voice. Finally, three types of voice are provided for manual re-examination, including complete voice, complete voice with keyword annotations, and speaker-separated voice. The above method improves the recognition rate of quality inspection voice, and a small number of quality inspection results can be extracted for manual re-examination.
[0111] The present application embodiment also provides a speech quality inspection device 300, such as Figure 3 As shown, a structural schematic diagram of a speech quality inspection device in an embodiment of the present application is provided, wherein the speech quality inspection device 300 at least includes: an acquisition module 310, a text translation module 320, a matching module 330, a marking module 340, a processing module 350 and a quality inspection module 360, wherein:
[0112] In one embodiment of the present application, the acquisition module 310 is specifically used to: acquire voice data.
[0113] The acquired voice data mainly includes but is not limited to the original voice signal, which can usually be pre-processed to obtain normalized data, so as to facilitate input into the model. In particular, the voice data includes the voice signal of voice debt collection.
[0114] In one embodiment of the present application, the text translation module 320 is specifically used to: input the voice data into a first model to obtain translated text content.
[0115] The speech data is input into the first model, and the translated text content corresponding to the speech data is obtained through the first model.
[0116] It should be noted that the first model is obtained by training one or more networks to translate voice data into translated text content, that is, converting voice signals into text content.
[0117] Preferably, if Figure 1As shown, the translated text content includes emotional features and scene features. It can be understood that the emotional features are used to represent the emotional information in the translated text content, such as recording emotions such as happiness, frustration, and excitement. The scene features are used to represent the scene information in the translated text content, such as recording environments such as indoors and outdoors, in a car, and in an office.
[0118] In one embodiment of the present application, the matching module 330 is specifically used to: match the translation text content with preset quality inspection rules based on the second model to obtain speech quality inspection results and matching keyword text.
[0119] The second model is also obtained by training one or more networks to match the translated text content with the preset quality inspection rules, and obtain a voice quality inspection result while matching the keyword text. The matching keyword text is used to characterize the degree of matching with the preset quality inspection rules. It can be understood that the higher the matching degree, the more it meets the preset quality inspection rules. Similarly, the lower the matching degree, the less it meets the preset quality inspection rules.
[0120] In one embodiment of the present application, the annotation module 340 is specifically used to obtain the speech annotated with the keyword according to the matching keyword text.
[0121] The text of the matching keywords obtained by intelligent matching is mapped frame by frame to the original speech to obtain the speech marked with the keywords.
[0122] In one embodiment of the present application, the processing module 350 is specifically used to obtain the original voice and the speaker separated voice according to the voice data.
[0123] The original voice can be obtained according to the voice data, and the speaker separation operation is performed on the original voice to obtain the speaker separated voice.
[0124] In one embodiment of the present application, the quality inspection module 360 is specifically used to: based on the speech quality inspection result, use the speech marked with keywords, the original speech and the speaker separated speech to perform speech quality inspection to obtain a final speech quality inspection result.
[0125] Based on the speech quality inspection results output by the model, manual quality inspection is performed using the original speech, speaker-separated speech, and speech-assisted speech with keyword annotations to obtain the final manual quality inspection results.
[0126] Preferably, the model is optimized again according to the manual quality inspection results to improve the matching effect of the model.
[0127] In one embodiment of the present application, the text translation module 320 is also used to
[0128] The speech data is pre-processed and then input into the first model;
[0129] According to the first model, outputting the translated text content including the sentiment features and the scene features;
[0130] in,
[0131] The first model includes a bidirectional long short-term memory network and a multi-head attention mechanism.
[0132] In one embodiment of the present application, the matching module 330 is also used to
[0133] According to the second model, after matching with the preset quality inspection rules, the matching keyword text related to the translated text content containing the emotional features and the scene features is obtained;
[0134] Using the output of the second model as the speech quality inspection result;
[0135] Among them, the second model includes a CNN network, an S-LSTM network and a DNN network.
[0136] In one embodiment of the present application, the quality inspection module 360 is also used to
[0137] The original speech, the speaker separated speech and the speech assistance with the marked keywords are used to perform manual re-inspection to obtain a final speech quality inspection result.
[0138] In one embodiment of the present application, a feedback module is also included for
[0139] Based on the quality inspection data of the manual review, feedback is provided to guide the training of the second model.
[0140] In one embodiment of the present application, the marking module 340 is also used to
[0141] Get the voice segment of the matching keyword text:
[0142] According to the speech segment of the matching keyword text, the keyword text obtained by intelligent matching is mapped to the original speech frame by frame to obtain the speech with the keyword marked.
[0143] In one embodiment of the present application, the processing module 350 is also used to
[0144] Segmenting the speech data to obtain segmented speech segments;
[0145] Extracting speaker features from the segmented speech segments;
[0146] According to the speaker characteristics, the speech segments are segmented and clustered by multiple speakers to obtain speech segmentation and clustering results;
[0147] According to the speech segmentation and clustering results, different speakers are separated to obtain speaker-separated speech.
[0148] It can be understood that the above-mentioned speech quality inspection device can implement each step of the speech quality inspection method provided in the above-mentioned embodiment, and the relevant explanations about the speech quality inspection method are applicable to the speech quality inspection device and will not be repeated here.
[0149] Figure 6 This is a schematic diagram of the structure of an electronic device according to an embodiment of the present application. Figure 6 At the hardware level, the electronic device includes a processor, and optionally also includes an internal bus, a network interface, and a memory. The memory may include a memory, such as a high-speed random access memory (RAM), and may also include a non-volatile memory (non-volatile memory), such as at least one disk storage. Of course, the electronic device may also include hardware required for other services.
[0150] The processor, network interface and memory can be interconnected through an internal bus, which can be an ISA (Industry Standard Architecture) bus, a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 6 Only one bidirectional arrow is used in the diagram, but this does not mean that there is only one bus or only one type of bus.
[0151] The memory is used to store the program. Specifically, the program may include a program code, and the program code includes a computer operation instruction. The memory may include a memory and a non-volatile memory, and provides instructions and data to the processor.
[0152] The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it, forming a voice quality inspection device at the logical level. The processor executes the program stored in the memory and is specifically used to perform the following operations:
[0153] Get voice data;
[0154] Inputting the speech data into a first model to obtain translated text content;
[0155] Based on the second model, matching the translated text content with preset quality inspection rules to obtain speech quality inspection results and matching keyword texts;
[0156] According to the matching keyword text, obtaining the speech marked with the keyword;
[0157] According to the voice data, obtaining the original voice and the speaker separated voice;
[0158] Based on the speech quality inspection result, the speech with the keyword annotated, the original speech and the speaker separated speech are used to perform speech quality inspection to obtain a final speech quality inspection result.
[0159] The above application Figure 2 The method performed by the speech quality inspection device disclosed in the illustrated embodiment can be applied to a processor or implemented by a processor. The processor may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by an integrated logic circuit of hardware in the processor or instructions in the form of software. The above processor can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components. The methods, steps and logic block diagrams disclosed in the embodiments of the present application can be implemented or executed. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. The steps of the method disclosed in the embodiments of the present application can be directly embodied as a hardware decoding processor for execution, or a combination of hardware and software modules in the decoding processor for execution. The software module can be located in a storage medium mature in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. The storage medium is located in the memory, and the processor reads the information in the memory and completes the steps of the above method in combination with its hardware.
[0160] The electronic device may also perform Figure 2 The method is implemented by the speech quality inspection device in Figure 2 The functions of the illustrated embodiment will not be described in detail in the embodiments of the present application.
[0161] The present application also provides a computer-readable storage medium, which stores one or more programs, wherein the one or more programs include instructions, which, when executed by an electronic device including multiple application programs, enable the electronic device to execute Figure 2 The method performed by the handover speech quality inspection device in the embodiment shown is specifically used to perform:
[0162] Get voice data;
[0163] Inputting the speech data into a first model to obtain translated text content;
[0164] Based on the second model, matching the translated text content with preset quality inspection rules to obtain speech quality inspection results and matching keyword texts;
[0165] According to the matching keyword text, obtaining the speech marked with the keyword;
[0166] According to the voice data, obtaining the original voice and the speaker separated voice;
[0167] Based on the speech quality inspection result, the speech with the keyword annotated, the original speech and the speaker separated speech are used to perform speech quality inspection to obtain a final speech quality inspection result.
[0168] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, systems, or computer program products. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0169] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0170] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0171] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process in the computer or other programmable device. Figure 1 A process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0172] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.
[0173] The memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.
[0174] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.
[0175] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.
[0176] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment or an embodiment in combination with software and hardware. Moreover, the present application may adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.
[0177] The above is only an embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application should be included in the scope of the claims of the present application.
Claims
1. A method for speech quality inspection, wherein: The method comprises: Get voice data; Inputting the speech data into a first model to obtain translated text content; Based on the second model, the translated text content is matched with the preset quality inspection rules to obtain the speech quality inspection result and the matching keyword text; According to the matching keyword text, obtaining the speech marked with the keyword; According to the voice data, obtaining the original voice and the speaker separated voice; Based on the speech quality inspection result, the speech with the keyword annotated, the original speech and the speaker separated speech are used to perform speech quality inspection to obtain a final speech quality inspection result.
2. The method of claim 1, wherein: The step of inputting the speech data into a first model to obtain translated text content includes: The speech data is pre-processed and then input into the first model; According to the first model, outputting the translated text content including the sentiment features and the scene features; in, The first model includes a bidirectional long short-term memory network and a multi-head attention mechanism.
3. The method of claim 2, wherein: The step of matching the translated text content with preset quality inspection rules based on the second model to obtain a speech quality inspection result and a matching keyword text includes: According to the second model, after matching with the preset quality inspection rules, the matching keyword text related to the translated text content containing the emotional features and the scene features is obtained; Using the output of the second model as the speech quality inspection result; Among them, the second model includes a CNN network, an S-LSTM network and a DNN network.
4. The method of claim 1, wherein: The method of performing speech quality inspection based on the speech quality inspection result by using the speech with the keyword marked, the original speech and the speaker separated speech to obtain a final speech quality inspection result includes: The original speech, the speaker separated speech and the speech assistance with the marked keywords are used to perform manual re-inspection to obtain a final speech quality inspection result.
5. The method according to claim 4, further comprising: Based on the quality inspection data of the manual review, feedback is provided to guide the training of the second model.
6. The method according to any one of claims 1 to 5, wherein: The step of obtaining a speech marked with a keyword according to the matching keyword text includes: Get the voice segment of the matching keyword text: According to the speech segment of the matching keyword text, the keyword text obtained by intelligent matching is mapped to the original speech frame by frame to obtain the speech with the keyword marked.
7. The method according to any one of claims 1 to 5, wherein: The step of obtaining the original voice and the speaker-separated voice according to the voice data comprises: Segmenting the speech data to obtain segmented speech segments; Extracting speaker features from the segmented speech segments; According to the speaker characteristics, the speech segments are segmented and clustered by multiple speakers to obtain speech segmentation and clustering results; According to the speech segmentation and clustering results, different speakers are separated to obtain speaker-separated speech.
8. A speech quality inspection device, wherein: The device comprises: An acquisition module, used for acquiring voice data; A text translation module, used for inputting the speech data into a first model to obtain translated text content; A matching module, used to match the translated text content with preset quality inspection rules based on the second model to obtain a speech quality inspection result and a matching keyword text; A marking module, used for obtaining speech marked with keywords according to the matching keyword text; A processing module, used for obtaining original speech and speaker-separated speech according to the speech data; The quality inspection module is used to perform voice quality inspection based on the voice quality inspection result by using the voice marked with keywords, the original voice and the speaker separated voice to obtain a final voice quality inspection result.
9. An electronic device, comprising: processor; as well as A memory arranged to store computer executable instructions, which when executed cause the processor to perform the method of any one of claims 1 to 7.
10. A computer-readable storage medium storing one or more programs, which, when executed by an electronic device including a plurality of application programs, causes the electronic device to execute any one of the methods of claims 1 to 7.