Offline voice quality inspection method, system, electronic device and storage medium
By combining offline speech quality inspection methods with speech translation, speaker separation, audio analysis and quality inspection models, the problem of inaccurate recognition of polysemy and emerging network vocabulary in existing technologies is solved, and the accuracy of speech recognition and quality inspection effects are improved.
Patent Information
- Application Number
- CN202210230837.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-10
- Publication Date
- 2025-09-12
- Estimated Expiration
- 2042-03-10
AI Technical Summary
The existing speech quality inspection methods are inaccurate in identifying user behaviors in complex or specific scenarios, especially the low recognition rate for polysemy in Chinese words and newly emerged network words, which affects the quality inspection effect.
An offline speech quality inspection method is adopted to identify the conversation content between users and customer service personnel through a combination of speech translation, speaker separation, audio analysis, text regularization model and quality inspection model. The trained quality inspection model is then used for matching and correction. Combined with the expansion of the corpus, model training is optimized to improve recognition accuracy.
It improves the recognition accuracy of polysemous words and newly emerged network words, makes up for the omissions and misidentification of user behavior, enhances the quality inspection effect, and facilitates the assessment of customer service personnel.
Smart Images

Figure CN114783442B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of speech processing technology, and in particular to an offline speech quality inspection method, system, electronic device, and storage medium. Background Art
[0002] As artificial intelligence technology becomes increasingly mature, some products have added semantic recognition, which can perform voice quality inspection on user information provided.
[0003] Although current products can perform voice quality inspection, their usage level is relatively shallow. When faced with complex or specific scenarios, user behavior may not be recognized or may be incorrectly recognized, especially in Chinese expressions. For example, for a word with multiple meanings or newly emerging Internet vocabulary, using only regular matching will not be able to flexibly identify it, and the recognition accuracy is low, which seriously affects the quality inspection effect.
[0004] Currently, there is no effective solution to the problem that voice quality inspection in related technologies cannot provide flexible recognition of user information, has low recognition accuracy, and seriously affects the quality inspection effect. Summary of the Invention
[0005] The embodiments of the present application provide an offline voice quality inspection method, system, electronic device and storage medium to at least solve the problem in the related art that the voice information provided by the user cannot be flexibly recognized, the recognition accuracy is low, and the quality inspection effect is seriously affected.
[0006] In a first aspect, an embodiment of the present application provides an offline voice quality inspection method, the method comprising the following steps:
[0007] Get the audio file;
[0008] Performing voice translation processing on the voice file to obtain a processed text file;
[0009] Inputting the text file into a speaker separation processor to separate the conversation contents between the user and the customer service staff, thereby obtaining a separated text file;
[0010] Inputting the differentiated text file into an audio analysis processor for audio analysis processing to obtain a quality inspection text;
[0011] Input the quality inspection text into the text regularization model to identify preset keywords and obtain regularization recognition results;
[0012] The quality inspection text and the regular recognition result are respectively input into the trained quality inspection model for matching processing to obtain the quality inspection result.
[0013] In some embodiments, the quality inspection text and the regular recognition result are respectively input into a trained quality inspection model for matching processing to obtain a quality inspection result, which includes:
[0014] Inputting the quality inspection text into the trained quality inspection model to obtain a model detection result;
[0015] Matching the semantic content of the model detection result and the regular recognition result;
[0016] If the match is successful, the model detection result and the regular recognition result are taken as the union to obtain the quality inspection result; if the match fails, the model detection result is used as the quality inspection result.
[0017] In some embodiments, the quality inspection text and the regular recognition result are respectively input into a trained quality inspection model for matching processing. After obtaining the quality inspection result, the method further includes:
[0018] If the quality inspection result is found to be erroneous through manual verification, the quality inspection text is added to the corpus to obtain an expanded corpus;
[0019] The quality inspection model is retrained based on the expanded corpus to obtain a retrained quality inspection model.
[0020] In some embodiments, the adding of the quality inspection text to the corpus to obtain the expanded corpus includes:
[0021] Analyze the user's actual voice content based on the quality inspection text;
[0022] Determining multiple different text expressions corresponding to the real voice content;
[0023] Obtaining the network vocabulary corresponding to the real voice content;
[0024] Different text expression contents and the network vocabulary are respectively set as training corpora, and the training corpora are added to the corpus.
[0025] In some embodiments, after inputting the differentiated text file into an audio analysis processor for audio analysis processing to obtain a quality inspection text, the method further includes:
[0026] When the text content input by the user is detected, the text content input by the user is fused with the quality inspection text to obtain a fused quality inspection text.
[0027] In some embodiments, the quality inspection text and the regular recognition result are respectively input into a trained quality inspection model for matching processing. After obtaining the quality inspection result, the method further includes:
[0028] The quality inspection result is input into the result processor for scoring to obtain the score corresponding to the quality inspection result.
[0029] In a second aspect, an embodiment of the present application provides an offline speech quality inspection system, the system comprising:
[0030] Acquisition module, used to obtain voice files;
[0031] A speech translation module, configured to perform speech translation processing on the speech file to obtain a processed text file;
[0032] A text differentiation module is used to input a text file into a speaker separation processor to differentiate the conversation contents between the user and the customer service staff, thereby obtaining a differentiated text file;
[0033] An audio processing module is used to input the differentiated text file into an audio analysis processor for audio analysis processing to obtain a quality inspection text;
[0034] A regular expression recognition module is used to input the quality inspection text into a text regular expression model to recognize preset keywords and obtain a regular expression recognition result;
[0035] The quality inspection module is used to input the quality inspection text and the regular recognition result into the trained quality inspection model for matching processing to obtain the quality inspection result.
[0036] In some embodiments, the system further comprises:
[0037] A model detection result module is used to input the quality inspection text into the trained quality inspection model to obtain a model detection result;
[0038] A matching module, which matches the semantic content of the model detection result with the regular recognition result;
[0039] The quality inspection result module is used to take the union of the model detection result and the regular recognition result to obtain the quality inspection result if the match is successful; if the match fails, the model detection result is used as the quality inspection result.
[0040] In a third aspect, an embodiment of the present application provides an electronic device, comprising a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to perform the above-mentioned offline speech quality inspection method.
[0041] In a fourth aspect, an embodiment of the present application provides a storage medium, in which a computer program is stored, wherein the computer program is configured to execute the above-mentioned offline speech quality inspection method when running.
[0042] Based on the above technical solution, the offline speech quality inspection method provided in the embodiment of the present application, on the one hand, inputs the quality inspection text into the text regular model to identify preset keywords, obtains the regular recognition result, thereby realizing the recognition of the surface meaning of the keyword, and on the other hand, inputs the quality inspection text and the regular recognition result into the trained quality inspection model for matching processing to obtain the quality inspection result. Compared with the previous situation where only regular matching is used for polysemous words or newly emerging network words, it is no longer possible to flexibly identify them, the recognition accuracy is low, and the quality inspection effect is seriously affected. It can not only make up for the situation where the user has a certain behavior but no keywords appear, but also correct the scenario where keywords appear but it is not user behavior, thereby improving the accuracy of speech recognition, and conveniently implements the assessment of customer service personnel based on the quality inspection result, thereby improving the quality inspection effect, and solving the problem in the related technology that the voice information provided by the user cannot be flexibly recognized, the recognition accuracy is low, and the quality inspection effect is seriously affected. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:
[0044] Figure 1 This is a first flow chart of the offline voice quality inspection method according to an embodiment of the present application;
[0045] Figure 2 This is a flow chart of the training steps of the quality inspection model according to an embodiment of the present application;
[0046] Figure 3 2 is a schematic diagram of a second flow chart of the offline voice quality inspection method according to an embodiment of the present application;
[0047] Figure 4 is a structural block diagram of an offline voice quality inspection system according to an embodiment of the present application;
[0048] Figure 5 is a structural block diagram of a quality inspection module according to an embodiment of the present application;
[0049] Figure 6 Schematic diagram of the internal structure of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION
[0050] In order to make the purpose, technical solutions and advantages of the present application clearer, the present application is described and illustrated below in conjunction with the accompanying drawings and examples. It should be understood that the specific embodiments described herein are only used to explain the present application and are not intended to limit the present application. Based on the embodiments provided in the present application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present application. In addition, it can also be understood that although the efforts made in this development process may be complex and lengthy, for ordinary technicians in the field related to the contents disclosed in the present application, some changes such as design, manufacturing or production based on the technical contents disclosed in the present application are only conventional technical means and should not be understood as the contents disclosed in the present application being insufficient.
[0051] References to "embodiments" in this application mean that a particular feature, structure, or characteristic described in connection with the embodiment may be included in at least one embodiment of the application. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor does it refer to independent or alternative embodiments that are mutually exclusive of other embodiments. It is understood, both explicitly and implicitly, by those skilled in the art that the embodiments described in this application may be combined with other embodiments unless there is a conflict.
[0052] Unless otherwise defined, the technical or scientific terms used in this application should have the ordinary meaning understood by a person of ordinary skill in the technical field to which this application belongs. The words "one", "a", "the" and the like used in this application do not indicate a limit on quantity and may indicate the singular or plural. The terms "include", "comprise", "have" and any variations thereof used in this application are intended to cover non-exclusive inclusions; for example, a process, method, system, product or device that includes a series of steps or modules (units) is not limited to the listed steps or units, but may also include steps or units that are not listed, or may also include other steps or units that are inherent to these processes, methods, products or devices. The words "connect", "connected", "coupled" and the like used in this application are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. The word "multiple" used in this application means greater than or equal to two. "And / or" describes the association relationship of associated objects, indicating that three relationships can exist. For example, "A and / or B" can mean: A exists alone, A and B exist at the same time, and B exists alone. The terms "first", "second", "third" and the like involved in this application are merely used to distinguish similar objects and do not represent a specific ordering of the objects.
[0053] The present invention provides an off-line speech quality inspection method. Figure 1This is a first flow chart of the offline voice quality inspection method according to an embodiment of the present application. Figure 1 As shown, in one embodiment of the present invention, the offline voice quality inspection method proposed by the present invention is applied to a scenario where a customer service representative of a bank conducts a voice conversation with a user. Of course, in other embodiments, it is also applicable to a scenario where a marketing representative conducts voice marketing conversations with a user. The method includes the following steps:
[0054] Step S101, obtaining a voice file;
[0055] In step S102, the voice file is subjected to speech translation processing to obtain a processed text file. The speech translation processing in this embodiment is implemented using ASR (Automatic Speech Recognition). Since those skilled in the art know that ASR is a technology for converting human speech into text, it will not be described in detail here.
[0056] In step S103, the text file is input into a speaker separation processor to distinguish the conversation content (text, i.e., text content) between the user and the customer service staff, thereby obtaining a distinguished text file. This facilitates distinguishing the user's speech content from the customer service staff's speech content. In this embodiment, the speaker separation processor is implemented using an existing speaker separation processor. Since those skilled in the art know that the working principle of the speaker separation processor is to distinguish the conversation content (text, i.e., text content) between the user and the customer service staff and obtain a distinguished text file, they will not be described in detail here. Of course, in some other embodiments, the speaker separation processor can also be implemented using some existing discourse distinction models, as long as it can distinguish the text content of the user and the customer service staff, and no specific limitation is made here.
[0057] In step S104, the differentiated text file is input into an audio analysis processor for audio analysis processing to obtain a quality inspection text; the quality inspection text includes at least one or more combinations of silence detection results, volume detection results, talk-over detection results, and emotion recognition results. In this embodiment, inputting the differentiated text file into an audio analysis processor for audio analysis processing facilitates identification of the user's emotions. Specifically, step S104 in this embodiment includes the following steps:
[0058] The differentiated text file is input into an audio analysis processor for silence detection processing to obtain a silence result. The silence detection result includes the time interval between each conversation (i.e., the time interval between each conversation between users). Based on the silence result, people can easily infer the user's intention, such as thinking and intention.
[0059] The differentiated text file is input into the audio analysis processor for volume detection processing to obtain the volume detection result (i.e., the volume of the user's speech). In this way, it is convenient for people to use the volume detection result to infer the user's emotion recognition result (such as nervousness, disgust, etc.) as a reference to facilitate the identification of the user's true intention;
[0060] The differentiated text files are input into the audio analysis processor for interruption detection processing to obtain interruption detection results. For example, if the listener (user) starts talking before the speaker (customer service representative) finishes speaking, the interruption detection results can be used as a reference for judging the user's personality and facilitating the identification of the user's true intentions.
[0061] The differentiated text file is input into the audio analysis processor for speech speed detection to obtain a speech speed detection result, which is then used as a reference for judging the user's anxious emotions and facilitating the identification of the user's true intentions;
[0062] The differentiated text file is input into the audio analysis processor for emotion recognition to obtain the emotion recognition result. The emotion recognition result is connected to Baidu's open source emotion recognition model, and the above detection results are used to easily derive the user's final intention.
[0063] Step S105: Input the quality inspection text into the text regularization model to identify preset keywords and obtain regularization recognition results. This facilitates keyword recognition and thus obtains the user's true behavioral intention. The text regularization model is an algorithm that directly uses regular expressions (Regular Expression, abbreviated as Regex) to realize keyword recognition.
[0064] In step S106, the quality inspection text and the regular recognition result are respectively input into the trained quality inspection model for matching processing to obtain the quality inspection result; wherein, the quality inspection model can be implemented using any existing neural network model, wherein the quality inspection model in this embodiment is implemented using the intent matching-NIP algorithm model. Compared with the previous embodiment, which only uses regular matching for polysemy or newly emerging network words and cannot flexibly identify them, the recognition accuracy is low, which seriously affects the quality inspection effect. This embodiment can not only make up for the situation where the user has a certain behavior but no keywords appear, but also correct the scenario where keywords appear but are not user behaviors, thereby improving the accuracy of speech recognition, and conveniently implement the assessment of customer service personnel based on the quality inspection result, thereby improving the quality inspection effect; for example, when the quality inspection text includes the network word "YYDS", through step S106, it can be judged that the user is using it to express that something or someone is excellent and amazing like a god, that is, an affirmative attitude, thereby making up for the situation where there is a praise behavior but no praise or affirmation keywords appear, thereby improving the accuracy of speech recognition.
[0065] Through the above steps S101 to S106, this embodiment, on the one hand, inputs the quality inspection text into the text regular model to identify preset keywords and obtains regular recognition results, thereby realizing the recognition of the surface meaning (shallow meaning) of the keywords. On the other hand, the quality inspection text and the regular recognition results are respectively input into the trained quality inspection model for matching processing to obtain quality inspection results. Compared with the previous situation where only regular matching is used for polysemous words or newly emerging network words, it is no longer possible to flexibly identify them, the recognition accuracy is low, and the quality inspection effect is seriously affected. This can not only make up for the situation where the user has a certain behavior but no keywords appear, but also correct the scenario where keywords appear but are not user behaviors, thereby improving the accuracy of voice recognition, and conveniently implement the assessment of customer service personnel based on the quality inspection results, thereby improving the quality inspection effect, and solving the problem in related technologies that the voice information provided by users cannot be flexibly recognized, the recognition accuracy is low, and the quality inspection effect is seriously affected.
[0066] In some embodiments, the quality inspection text and the regular recognition result are respectively input into the trained quality inspection model for matching processing, and obtaining the quality inspection result includes the following steps:
[0067] Input the quality inspection text into the trained quality inspection model to obtain the model detection results;
[0068] Match the semantic content of the model detection results and the regular recognition results;
[0069] If the match is successful, the model detection result and the regular recognition result are taken as the union to obtain the quality inspection result. If the match fails, the model detection result is used as the quality inspection result. In other words, when the trained quality inspection model detects a discrepancy between the model detection result and the regular recognition result for the same sentence, the model detection result is used as the quality inspection result, which makes the quality inspection result more accurate.
[0070] Figure 2 This is a flow chart of the training steps of the quality inspection model of the embodiment of the present application. Figure 2 As shown, in some embodiments, the training process of the quality inspection model includes the following steps:
[0071] First, create a quality inspection model;
[0072] Then, the industry corpus is obtained and pre-processed, specifically including: pre-processing the industry corpus to remove duplicate corpus; then, extracting corpus features from the pre-processed industry corpus and grouping it into training and test sets; wherein, 90% of the corpus after corpus feature extraction is used for training (i.e., training set), and 10% of the corpus is used for testing (i.e., test set);
[0073] The quality inspection model is pre-loaded with the XLNet network (natural regression training model) through the training set to obtain the trained quality inspection model. Finally, the generated quality inspection model is repeatedly tested on the test set (continuously adjusting parameters, etc.) to obtain the trained quality inspection model (that is, the model that finally meets the expectations is obtained).
[0074] Figure 3 : is a second flow chart of the offline voice quality inspection method according to an embodiment of the present application. Figure 3 In actual application, in order to continuously improve the quality inspection model and further enhance the quality inspection effect, in some embodiments, the quality inspection text and the regular recognition result are respectively input into the trained quality inspection model for matching processing. After obtaining the quality inspection result, the method further includes the following steps:
[0075] If errors are found in the quality inspection results manually, the quality inspection text is added to the corpus to obtain an expanded corpus;
[0076] The quality inspection model is retrained based on the expanded corpus to obtain a retrained quality inspection model. It is easy for those skilled in the art to understand that the expanded corpus can be a regularly modified algorithm model, and the quality inspection model is trained based on the algorithm model to obtain a retrained quality inspection model to improve the training effect of the model.
[0077] In some embodiments, adding the quality-checked text to the corpus to obtain the expanded corpus includes the following steps:
[0078] Analyze the user's actual voice content based on the quality inspection text;
[0079] Determine the multiple different text expressions corresponding to the real voice content; that is, obtain the different text expressions with the same semantics corresponding to the quality inspection text. For example, if the quality inspection text analysis shows that the user's real language content is "affirmative", it can be determined that the multiple different text expressions corresponding to the real voice content may be words such as "good", "OK", and "agree";
[0080] The network vocabulary corresponding to the real voice content is obtained; for example, when the real voice content of the user is "affirmative", the network vocabulary obtained may be "YYDS", "Oli Gei", "Jue Jue Zi", etc.
[0081] Different text expressions and network words are set as training corpus respectively, and the training corpus is added to the corpus. This embodiment trains the quality inspection model with the expanded corpus, which not only makes up for the situation where the model training effect is poor due to the lack of training data, but also helps to improve the recognition accuracy of the quality inspection model.
[0082] In some embodiments, after inputting the differentiated text file into an audio analysis processor for audio analysis processing to obtain a quality inspection text, the method further includes the following steps:
[0083] When the text content input by the user is detected, the text content input by the user is merged with the quality inspection text to obtain the merged quality inspection text. In this way, compatibility with plain text chat content is achieved.
[0084] In some embodiments, the quality inspection text and the regular expression recognition result are respectively input into the trained quality inspection model for matching processing. After the quality inspection result is obtained, the method further includes the following steps:
[0085] The quality inspection results are input into the result processor for scoring processing to obtain the score corresponding to the quality inspection result. In this way, it is convenient to implement the assessment of customer service personnel according to the score; for example, the high or low score is used as part of the customer service personnel assessment. Of course, in some other embodiments, hot word analysis can be performed before inputting the result processor to facilitate the subsequent scoring processing by the result processor; the result processor of this embodiment can be implemented by using an existing result processor or other methods, and is not specifically limited here; in addition, the result processor of this embodiment supports labeling of quality inspection results to determine whether the work of customer service personnel is standardized, for example, to determine whether the customer service personnel introduced themselves during the voice communication with the user. If so, one point is added to the quality inspection result, and if not, one point is subtracted from the quality inspection result. In addition, the user's label can identify the user's intention satisfaction and can also extract the user's information to facilitate the subsequent classification of users.
[0086] It should be noted that the steps shown in the above process or the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0087] Figure 4 is a structural block diagram of an offline voice quality inspection system according to an embodiment of the present application. Figure 4 As shown, the system includes:
[0088] An acquisition module 41 is used to acquire a voice file;
[0089] The speech translation module 42 is used to perform speech translation processing on the speech file to obtain a processed text file;
[0090] A text distinguishing module 43 is used to input a text file into a speaker separation processor to distinguish the conversation contents between the user and the customer service staff, thereby obtaining a distinguished text file;
[0091] The audio processing module 44 is used to input the differentiated text file into the audio analysis processor for audio analysis processing to obtain a quality inspection text;
[0092] Regular expression recognition module 45, used to input the quality inspection text into the text regular expression model to recognize preset keywords and obtain regular expression recognition results;
[0093] The quality inspection module 46 is used to input the quality inspection text and the regular recognition result into the trained quality inspection model for matching processing to obtain the quality inspection result. On the one hand, this embodiment inputs the quality inspection text into the text regular model to identify the preset keywords and obtain the regular recognition result, thereby realizing the recognition of the surface meaning of the keywords. On the other hand, it is convenient to input the quality inspection text and the regular recognition result into the trained quality inspection model for matching processing to obtain the quality inspection result. Compared with the previous situation where only regular matching is used for polysemous words or newly emerging network words, it is no longer flexible to identify them, and the recognition accuracy is low, which seriously affects the quality inspection effect. This can not only make up for the situation where the user has a certain behavior but no keywords appear, but also correct the situation where keywords appear but are not user behaviors, thereby improving the accuracy of voice recognition. In addition, it is convenient to implement the assessment of customer service personnel based on the quality inspection result, thereby improving the quality inspection effect, and solving the problem in the related technology that the voice information provided by the user cannot be flexibly recognized, the recognition accuracy is low, and the quality inspection effect is seriously affected.
[0094] Figure 5 is a structural block diagram of the quality inspection module according to an embodiment of the present application, such as Figure 5 As shown, in some embodiments, the quality inspection module 46 includes:
[0095] The model detection result module 51 is used to input the quality inspection text into the trained quality inspection model to obtain the model detection result;
[0096] Matching module 52, matches the semantic content of the model detection result and the regular recognition result;
[0097] The quality inspection result module 53 is used to combine the model detection result with the regular recognition result to obtain the quality inspection result if the match is successful. If the match fails, the model detection result is used as the quality inspection result. In other words, when the trained quality inspection model detects a discrepancy between the model detection result and the regular recognition result for the same sentence, the model detection result is used as the quality inspection result.
[0098] It should be noted that the above modules can be functional modules or program modules, and can be implemented through software or hardware. For modules implemented through hardware, the above modules can be located in the same processor; or the above modules can be located in different processors in any combination.
[0099] This embodiment further provides an electronic device, including a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to execute the steps in any one of the above method embodiments.
[0100] Optionally, the electronic device may further include a transmission device and an input / output device, wherein the transmission device is connected to the processor, and the input / output device is connected to the processor.
[0101] Optionally, in this embodiment, the processor may be configured to execute the following steps through a computer program:
[0102] Step S101, obtaining a voice file;
[0103] Step S102, performing speech translation processing on the speech file to obtain a processed text file;
[0104] Step S103: Input the text file into a speaker separation processor to separate the conversation contents between the user and the customer service staff, thereby obtaining a separated text file;
[0105] Step S104: input the differentiated text file into an audio analysis processor for audio analysis processing to obtain a quality inspection text;
[0106] Step S105: Input the quality inspection text into the text regularization model to identify preset keywords and obtain regularization recognition results;
[0107] Step S106: input the quality inspection text and the regular expression recognition result into the trained quality inspection model for matching processing to obtain the quality inspection result.
[0108] It should be noted that the specific examples in this embodiment can refer to the examples described in the above embodiments and optional implementation modes, and this embodiment will not be repeated here.
[0109] In addition, in conjunction with the offline voice quality inspection method in the above embodiments, the present application embodiment can provide a storage medium for implementation. The storage medium stores a computer program; when the computer program is executed by a processor, it implements any of the offline voice quality inspection methods in the above embodiments.
[0110] In one embodiment, a computer device is provided, which may be a terminal. The computer device includes a processor, memory, a network interface, a display screen, and an input device connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external terminal via a network connection. When executed by the processor, the computer program implements an offline voice quality inspection method. The display screen of the computer device may be a liquid crystal display or an electronic ink display screen. The input device of the computer device may be a touch layer covering the display screen, or may be buttons, a trackball, or a touchpad provided on the computer device housing, or may be an external keyboard, touchpad, or mouse.
[0111] In one embodiment, Figure 6 is a schematic diagram of the internal structure of an electronic device according to an embodiment of the present application, such as Figure 6 As shown, an electronic device is provided, which may be a server, and its internal structure diagram may be as shown in FIG. Figure 6 As shown. The electronic device includes a processor, a network interface, an internal memory, and a non-volatile memory connected via an internal bus, wherein the non-volatile memory stores an operating system, a computer program, and a database. The processor is used to provide computing and control capabilities, the network interface is used to communicate with external terminals via a network connection, the internal memory is used to provide an environment for the operation of the operating system and the computer program. When the computer program is executed by the processor, it implements an offline voice quality inspection method, and the database is used to store data.
[0112] Those skilled in the art will understand that Figure 6 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the electronic device to which the solution of the present application is applied. The specific electronic device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.
[0113] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiment methods can be implemented by instructing the relevant hardware through a computer program, which can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).
[0114] Those skilled in the art should understand that the various technical features of the above-described embodiments can be combined arbitrarily. In order to make the description concise, not all possible combinations of the various technical features in the above-described embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0115] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that a person skilled in the art could make various modifications and improvements without departing from the spirit of the present application, all of which fall within the scope of protection of the present application. Therefore, the scope of protection of the present patent application shall be determined by the appended claims.
Claims
1. An offline voice quality inspection method, characterized in that: The method comprises: Get the audio file; Performing voice translation processing on the voice file to obtain a processed text file; Inputting the text file into a speaker separation processor to separate the conversation contents between the user and the customer service staff, thereby obtaining a separated text file; Inputting the differentiated text file into an audio analysis processor for audio analysis processing to obtain a quality inspection text; Input the quality inspection text into the text regularization model to identify preset keywords and obtain regularization recognition results; The quality inspection text is input into the trained quality inspection model to obtain the model detection result; the semantic content of the model detection result and the regular recognition result are matched; if the match is successful, the model detection result and the regular recognition result are taken as the union to obtain the quality inspection result; if the match fails, the model detection result is used as the quality inspection result.
2. The method according to claim 1, characterized in that The quality inspection text and the regular recognition result are respectively input into the trained quality inspection model for matching processing. After obtaining the quality inspection result, the method further includes: If the quality inspection result is found to be erroneous through manual verification, the quality inspection text is added to the corpus to obtain an expanded corpus; The quality inspection model is retrained based on the expanded corpus to obtain a retrained quality inspection model.
3. The method according to claim 2, characterized in that The quality inspection text is added to the corpus, and the expanded corpus includes: Analyze the user's actual voice content based on the quality inspection text; Determining multiple different text expressions corresponding to the real voice content; Obtaining the network vocabulary corresponding to the real voice content; Different text expression contents and the network vocabulary are respectively set as training corpora, and the training corpora are added to the corpus.
4. The method according to claim 1, wherein After inputting the differentiated text file into an audio analysis processor for audio analysis processing to obtain a quality inspection text, the method further includes: When the text content input by the user is detected, the text content input by the user is fused with the quality inspection text to obtain a fused quality inspection text.
5. The method according to claim 1, wherein The quality inspection text and the regular recognition result are respectively input into the trained quality inspection model for matching processing. After obtaining the quality inspection result, the method further includes: The quality inspection result is input into the result processor for scoring to obtain the score corresponding to the quality inspection result.
6. An offline voice quality inspection system, characterized in that: The system comprises: Acquisition module, used to obtain voice files; A speech translation module, configured to perform speech translation processing on the speech file to obtain a processed text file; A text differentiation module is used to input a text file into a speaker separation processor to differentiate the conversation contents between the user and the customer service staff, thereby obtaining a differentiated text file; An audio processing module is used to input the differentiated text file into an audio analysis processor for audio analysis processing to obtain a quality inspection text; A regular expression recognition module is used to input the quality inspection text into a text regular expression model to recognize preset keywords and obtain a regular expression recognition result; The quality inspection module is used to input the quality inspection text into the trained quality inspection model to obtain the model detection result; match the semantic content of the model detection result with the regular recognition result; if the match is successful, take the union of the model detection result and the regular recognition result to obtain the quality inspection result; if the match fails, use the model detection result as the quality inspection result.
7. An electronic device comprising a memory and a processor, characterized in that: A computer program is stored in the memory, and the processor is configured to run the computer program to perform the offline speech quality inspection method according to any one of claims 1 to 5.
8. A storage medium, characterized in that: The storage medium stores a computer program, wherein the computer program is configured to execute the offline speech quality inspection method according to any one of claims 1 to 5 when running.
Citation Information
Patent Citations
Speech quality inspection method, apparatus and device based on artificial intelligence and storage medium
CN109658923A
Voice quality inspection method and device, electronic equipment and storage medium
CN114157765A