A voice detection method, device, equipment and readable storage medium

By performing segmented processing and emotional analysis on speech, combined with text recognition and emotion recognition, the problem of inaccurate speech detection results in the prior art is solved, and a comprehensive evaluation and accurate evaluation of customer service quality is achieved.

CN115602153BActive Publication Date: 2025-08-15MASHANG CONSUMER FINANCE CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202110771202.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-07-08
Publication Date
2025-08-15
Estimated Expiration
2041-07-08

AI Technical Summary

Technical Problem

Existing voice detection technology can only recognize the text information of voice, resulting in inaccurate detection results and the inability to comprehensively evaluate the service quality of customer service.

Method used

By obtaining the setting parameters of the voice to be detected, the voice is processed in segments, combined with text recognition and emotion analysis, the text information and emotional information of the voice segments are obtained, and the process quality inspection information and keyword detection results are obtained.

Benefits of technology

The comprehensive evaluation of voice detection results is achieved, the accuracy and objectivity of the detection are improved, and the service quality of customer service can be better evaluated.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115602153B_ABST
    Figure CN115602153B_ABST
Patent Text Reader

Abstract

This application discloses a speech detection method, apparatus, device, and readable storage medium, relating to the field of speech processing technology, to improve the accuracy of speech detection results. The method comprises: obtaining setting parameters for a speech to be detected and the speech to be detected; segmenting the speech to be detected to obtain at least one segmented speech; processing the segmented speech according to the setting parameters to obtain text information and emotional information for each segmented speech; and obtaining process quality inspection information and keyword detection results for the speech to be detected based on the text information and emotional information. Embodiments of the present application can improve the accuracy of speech detection results.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of speech processing technology, and in particular to a speech detection method, apparatus, device, and readable storage medium. Background Art

[0002] With the rapid development of the telecommunications industry, telephones have become a part of almost everyone's daily lives, and people from all walks of life can now use them to communicate with customers. Therefore, illegal or non-compliant operations are inevitable in this process. Traditionally, manual spot checks have been used to verify communications between customers and agents. This method is not only time-consuming and labor-intensive, but also incomplete. With the rapid development of AI (artificial intelligence) technology, intelligent voice quality inspection has emerged, replacing traditional manual verification of legal and compliant communications between customers and agents.

[0003] The existing technology can detect speech by the emotional information expressed in the speech. However, the detection results obtained by this method are relatively one-sided, resulting in inaccurate speech detection results. Summary of the Invention

[0004] The embodiments of the present application provide a speech detection method, apparatus, device, and readable storage medium to improve the accuracy of speech detection results.

[0005] In a first aspect, an embodiment of the present application provides a voice detection method, comprising:

[0006] Obtaining setting parameters of the voice to be detected and the voice to be detected;

[0007] Segmenting the speech to be detected to obtain at least one segmented speech;

[0008] Processing the segmented speech according to the setting parameters to obtain text information and emotional information of each segmented speech;

[0009] According to the text information and the emotion information, process quality inspection information and keyword detection results of the speech to be detected are obtained.

[0010] In a second aspect, an embodiment of the present application further provides a speech detection device, comprising:

[0011] A first acquisition module, configured to acquire setting parameters of a voice to be detected and the voice to be detected;

[0012] A first segmentation module, configured to segment the speech to be detected to obtain at least one segmented speech;

[0013] A second acquisition module is used to process the segmented speech according to the setting parameters to obtain text information and emotional information of each segmented speech;

[0014] The third acquisition module is used to obtain process quality inspection information and keyword detection results of the speech to be detected based on the text information and the emotion information.

[0015] In a third aspect, an embodiment of the present application further provides an electronic device comprising: a transceiver, a memory, a processor, and a program stored in the memory and executable on the processor, wherein when the processor executes the program, the steps in the voice detection method as described above are implemented.

[0016] In a fourth aspect, an embodiment of the present application further provides a readable storage medium, on which a program is stored, and when the program is executed by a processor, the steps in the voice detection method as described above are implemented.

[0017] In the embodiment of the present application, the setting parameters of the speech to be detected are obtained to perform text recognition and emotion analysis on the speech to be detected, and process quality inspection analysis and keyword detection are obtained based on the results of text recognition and emotion analysis, thereby obtaining text information, emotion information, process quality inspection information and keyword detection results of the speech to be detected. Therefore, compared with the existing technology, the solution of the embodiment of the present application can obtain more comprehensive information about the speech to be detected, including text information, emotion information, process quality inspection information and keyword detection results, so that the recognition result of the speech to be detected obtained based on this information is more accurate. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 This is one of the flow charts of the voice detection method provided in the embodiment of the present application;

[0019] Figure 2 This is the second flow chart of the voice detection method provided in the embodiment of the present application;

[0020] Figure 3 is a schematic diagram of performing speech detection using the system of an embodiment of the present application;

[0021] Figure 4 It is a structural diagram of the speech detection device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0022] In the embodiments of this application, the term "and / or" describes the association relationship between associated objects, indicating that three relationships can exist. For example, A and / or B can represent three situations: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally indicates that the associated objects are in an "or" relationship.

[0023] In the embodiments of the present application, the term "plurality" refers to two or more than two, and other quantifiers are similar.

[0024] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0025] In the process of quality inspection of customer service voice files, conventional voice detection methods can usually only identify one type of information of the voice file to be detected, such as the text information of the voice. Based on this type of information of the identified customer service voice text information, the customer service service quality is determined and the customer service quality inspection result is obtained. Therefore, this method cannot accurately reflect the information included in the voice to be detected, and it is also impossible to obtain accurate recognition results. To this end, it is necessary to propose a method that can obtain multiple information of the voice file to be detected and obtain the recognition result of the voice to be detected based on the multiple information. In the method of the embodiment of the present application, a voice detection method is proposed, which is mainly used in the voice quality inspection of intelligent customer service. In this method, the voice to be detected is segmented according to the set parameters, and text information and emotional information are obtained based on the segmentation results. Then, process quality inspection information and keyword detection results of the voice to be detected are obtained, and the service quality of the customer service is evaluated by combining the text information, emotional information, process quality inspection information and keyword detection results. This makes the recognition result of the voice to be detected obtained based on this information more accurate and the evaluation of the service quality of the customer service more objective.

[0026] See also Figure 1 , Figure 1 This is a flow chart of the voice detection method provided by the embodiment of the present application. Figure 1 As shown, the following steps are included:

[0027] Step 101: Acquire setting parameters of a speech to be detected and the speech to be detected.

[0028] The speech to be detected can be any speech to be quality-checked. The setting parameters include one or more of the following: the identification of the speech to be detected, the identification information of the application scenario of the speech to be detected, the acquisition attribute information of the speech to be detected, the vocal channel information of the speech to be detected, the speech emotion model information of the speech to be detected, etc. The acquisition attribute information includes one or more of the following: the recording address of the speech to be detected, the audio sampling rate, and the format of the speech to be detected (such as MP3, WAV, etc.).

[0029] Specifically, as shown in Table 1, the setting parameters of the voice to be detected may include:

[0030] Table 1

[0031]

[0032] In Table 1 above, the meanings of the parameters are as follows:

[0033] requestId, used to indicate the identifier of the voice to be detected;

[0034] src, used to indicate the source of the voice to be detected, such as which application or client it comes from;

[0035] asrType, used to indicate the application scenario of the speech to be detected, which may include the four scenarios shown in Table 1;

[0036] audioUrl, used to indicate the recording address of the voice to be detected;

[0037] rate, which is used to indicate the audio sampling rate of the speech to be detected;

[0038] soundtrack, used to indicate whether the speech to be detected is binaural or monophonic. If this information is not included, channel differentiation technology can be used to distinguish whether the speech to be detected is binaural or monophonic.

[0039] suffix, used to indicate the audio file suffix of the speech to be detected, i.e. the format;

[0040] modelTypex is used to represent the speech emotion model information of the speech to be detected. It includes binary classification and four-category classification. Binary classification can identify the two emotions shown in Table 1, and four-category classification can identify the four emotions shown in Table 1. If Table 1 does not include this parameter, binary or four-category classification is used by default.

[0041] In an embodiment of the present application, the setting parameters stored in the first message queue can be obtained by monitoring the first message queue. In this embodiment of the present application, the first message queue includes a RabbitMq queue, and the setting parameters can be stored in the form of a RabbitMq queue. In this way, the setting parameters are stored in the form of a queue. When it is monitored that there are setting parameters in the first message queue, the setting parameters can be obtained and each service can be called for processing, so that the utilization rate of each service can be maximized. When there are multiple processing tasks, multiple groups of setting parameters are stored in the form of queues. Similarly, when it is monitored that there are multiple groups of setting parameters in the first message queue, each service can be called for processing for each group of setting parameters, which solves the high concurrency problem well.

[0042] Step 102: Segment the speech to be detected to obtain at least one segmented speech.

[0043] In this step, the voice to be detected is subjected to channel recognition to obtain first channel voice and second channel voice. For example, SOX (Sound eXchange) can be used to perform channel recognition on the voice to be detected to obtain first channel voice and second channel voice. Afterwards, the first channel voice and the second channel voice can be segmented by VAD (Voice Activity Detection) technology to obtain at least one first channel segmented voice and at least one second channel segmented voice. The first channel segmented voice can be the segmented voice of the agent, and the second channel segmented voice can be the segmented voice of the customer; or vice versa.

[0044] Step 103: Process the segmented speech according to the setting parameters to obtain text information and emotional information of each segmented speech.

[0045] In an embodiment of the present application, the setting parameters include identification information of the application scenario of the speech to be detected (such as name, identifier, etc.) and speech emotion model information of the speech to be detected. Specifically, in this step, according to the identification information of the application scenario of the speech to be detected in the setting parameters, the segmented speech is converted into text using speech recognition technology to obtain the text information of the segmented speech, and according to the speech emotion model information of the speech to be detected in the setting parameters, the segmented speech is subjected to emotion recognition using a speech emotion recognition algorithm to obtain the emotion information of the segmented speech. Afterwards, the text information of the segmented speech and the emotion information of the segmented speech are stored in the second message queue.

[0046] Specifically, in the process of obtaining the text information of each segmented speech, the first channel segmented speech and the second channel segmented speech can be converted into text respectively using speech recognition technology according to the application scenario of the speech to be detected in the setting parameters, thereby obtaining the text information of the first channel segmented speech and the text information of the second channel segmented speech. By distinguishing different application scenarios, the obtained text information can be made more accurate. Specifically, for different application scenarios, text conversion can be performed through different services, or text conversion can be performed using processing methods corresponding to different application scenarios in the same service.

[0047] In the process of obtaining the emotional information of the segmented speech, a speech emotion recognition algorithm can be used to perform emotion recognition on the first channel segmented speech and the second channel segmented speech, respectively, based on the speech emotion model information of the speech to be detected in the setting parameters, to obtain the emotional information of the first channel segmented speech and the emotional information of the second channel segmented speech. If the speech emotion model information represents a binary classification, the emotional information can include positive and negative; if the speech emotion model information represents a four-category classification, the emotional information can include angry, happy, sad, and neutral.

[0048] In practical applications, emotional information recognition can be performed first and then text conversion can be performed.

[0049] In the process of storing the text information of the segmented speech and the emotional information of the segmented speech in the second message queue, the text information of the first channel segmented speech, the text information of the second channel segmented speech, the emotional information of the first channel segmented speech and the emotional information of the second channel segmented speech can be stored in the second message queue.

[0050] During storage, in order to facilitate subsequent process quality inspection and thereby improve processing efficiency, in an embodiment of the present application, the text information of the first channel segmented speech, the text information of the second channel segmented speech, the emotional information of the first channel segmented speech, and the emotional information of the second channel segmented speech can be spliced in chronological order to obtain a splicing result, and then the splicing result can be stored in the second message queue.

[0051] The time sequence can be from early to late or from late to early. The time refers to the time at which the segmented speech is located in the entire speech to be detected. For example, if one segmented speech is at the first minute of the entire speech to be detected and another segmented speech is at the first minute and fifth second of the entire speech to be detected, then the time corresponding to the former is earlier than the time corresponding to the latter.

[0052] Suppose the text information for the first channel's segmented speech is A, B, and C, and the text information for the second channel's segmented speech is D, E, and F. For example, in order from morning to night, the text information is sorted as A, D, B, E, C, and F. Then, after A, the emotional information of A is recorded, followed by D and the emotional information of D; then, B and the emotional information corresponding to B are recorded. And so on, ultimately forming the concatenation result.

[0053] For example, the text information of the first channel segmented speech, the text information of the second channel segmented speech, the emotional information of the first channel segmented speech, and the emotional information of the second channel segmented speech can be spliced to form a result in XML (eXtensible Markup Language) format, and then the result is stored in the second message queue.

[0054] In addition to storing the splicing results, the second message queue may also store the identifier and source of the voice to be detected. Specific details are shown in Table 2:

[0055] Table 2

[0056] parameter type illustrate Is it necessary to transmit requestId String Request ID is unique yes src String source no contentXml String Text information, emotional information yes

[0057] In Table 2 above, the meanings of the parameters are as follows:

[0058] requestId, used to indicate the identifier of the voice to be detected;

[0059] src, used to indicate the source of the voice to be detected, such as which software or client it comes from;

[0060] contentXml, used for the concatenation of text information and sentiment information.

[0061] Step 104: Obtain process quality inspection information and keyword detection results of the speech to be detected based on the text information and the emotion information.

[0062] In this step, the text information and emotional information in the second message queue are compared with the preset process text through the quality inspection process algorithm to obtain process quality inspection information, which is the process standard text corresponding to the application scenario of the voice to be detected; the text information and the preset keywords are compared through the keyword detection algorithm to obtain the keyword detection result. The process quality inspection information is used to evaluate whether a certain text information matches the process text required by the scenario in which the text information is applied. For example, suppose the application scenario is balance inquiry. Corresponding to the balance inquiry, there is a corresponding process standard text between the customer and the agent, which stipulates which sentences the agent needs to use to communicate with the customer. Then, here, the process quality inspection information is information that evaluates the degree of matching between the obtained text information and the process standard text. Among them, the degree of matching can be reflected in the form of a score value or the like.

[0063] The quality inspection process algorithm can be an NLP (Natural Language Processing) quality inspection process algorithm. Different scenarios require specific questions or responses from agents, resulting in standardized process texts for each scenario. Therefore, in this step, the obtained text information can be compared with the pre-set process text to obtain process quality inspection information. Different process quality inspection scores can be obtained based on the degree of match between the obtained text information and the pre-set process text.

[0064] The keyword detection algorithm may be an NLP keyword detection algorithm. The purpose of keyword detection is to identify whether the text message contains preset keywords. Different keywords may be set based on different application scenarios. These keywords may include sensitive words, such as words expressing anger, such as "idiot."

[0065] In an embodiment of the present application, text recognition and emotion analysis are performed on the speech to be detected respectively by obtaining the setting parameters of the speech to be detected, and process quality inspection analysis and keyword detection are obtained based on the results of text recognition and emotion analysis, thereby obtaining text information, emotion information, process quality inspection information and keyword detection results of the speech to be detected. Therefore, compared with the prior art, the solution of the embodiment of the present application can obtain more comprehensive information about the speech to be detected (including text information, emotion information, process quality inspection information and keyword detection results), thereby making the recognition result of the speech to be detected obtained based on this information more accurate.

[0066] In addition, in order to facilitate the user to understand the recognition result of the speech to be detected, the text information, the emotional information, the process quality inspection information and the keyword detection result can also be displayed.

[0067] See also Figure 2 , Figure 2 This is a flow chart of the voice detection method provided by the embodiment of the present application. Figure 2 As shown, the following steps are included:

[0068] Step 201: Acquire setting parameters of the speech to be detected and the speech to be detected.

[0069] The speech detection system provided in the embodiment of the present application may include multiple servers, such as a first server, a second server, and a third server.

[0070] Combine Figure 3As shown, it is a schematic diagram of voice detection using the system of an embodiment of the present application. The first server is used to pull the address of the audio file to be inspected (voice to be detected) from the data warehouse, and store the setting parameters of the audio file to be inspected in RabbitMq queue 1 (first message queue). The second server is used to monitor RabbitMq queue 1 and identify the text information and emotional information of the audio file to be inspected. The third server is used to monitor RabbitMq queue 2 (second message queue), obtain the parameters (splicing results) in the RabbitMq queue, call the NLP process quality inspection algorithm and the NLP keyword detection algorithm, and identify the process quality inspection information and keyword detection results of the audio file to be inspected.

[0071] The setting parameters may be as shown in Table 1.

[0072] Step 202: Segment the speech to be detected to obtain at least one segmented speech, and process the segmented speech according to the set parameters to obtain text information and emotional information of each segmented speech.

[0073] In this step, the second server Service2 identifies the text information and emotional information of the audio file to be quality checked.

[0074] Specifically, the second server Service2 listens to RabbitMq Queue 1, obtains the parameters set in RabbitMq Queue 1, and downloads the audio file to be inspected locally. Afterwards, the audio file to be inspected is subjected to channel recognition through SOX, and the audio file to be inspected after channel recognition is segmented by VAD. Afterwards, the ASR parsing algorithm service and the speech emotion algorithm service are called to obtain the text information and emotional information of each segmented speech. Finally, the text information and emotional information of each segment are spliced into XML text, and the parameters of the XML text are stored in RabbitMq Queue 2. Among them, the parameters stored by Service2 in RabbitMq Queue 2 are shown in Table 2.

[0075] Step 203: Obtain process quality inspection information and keyword detection results of the speech to be detected based on the text information and the emotion information.

[0076] In this step, the third server Service3 identifies the process quality inspection information and keyword detection results of the audio file to be quality inspected.

[0077] Specifically, the third server, Service3, listens to RabbitMq Queue 2, obtains the XML text parameters, and invokes the NLP process quality inspection algorithm and the NLP keyword detection algorithm. The NLP process quality inspection algorithm uses the text content in the contentXml parameter to determine whether the sentence complies with the specification, assigns a score, and returns the score result. The NLP keyword detection algorithm returns the identified keywords via the keyword parameter. The results are ultimately stored and returned to the web page in JSON (JavaScript Object Notation) format for quality inspection.

[0078] The parameters output by Service3 are shown in Table 3:

[0079] Table 3

[0080] parameter type illustrate Is it necessary to transmit requestId String Request ID is unique yes src String source no contentXml String Text information, emotional information yes score Float Process quality inspection score yes keyword String Query keywords yes

[0081] In Table 3 above, the meanings of the parameters are as follows:

[0082] requestId, used to indicate the identifier of the voice to be detected;

[0083] src, used to indicate the source of the voice to be detected, such as which software or client it comes from;

[0084] contentXml, used to represent the concatenation of text information and sentiment information.

[0085] score, used to indicate the process quality inspection score;

[0086] keyword, used to indicate the keyword found in the query.

[0087] It can be seen from the above description that compared with the existing technology, the solution of the embodiment of the present application can obtain more comprehensive information about the voice to be detected, including text information, emotional information, process quality inspection information and keyword detection results of the voice to be detected. The above information is combined to evaluate the service quality of customer service, so that the recognition results of the voice to be detected obtained based on this information are more accurate, and the evaluation of the service quality of customer service is more objective. The above steps can be completed on one system, further saving labor costs and resources.

[0088] The present application also provides a speech detection device. Figure 4 As shown, the speech detection device 400 includes:

[0089] The first acquisition module 401 is used to obtain the setting parameters of the speech to be detected and the speech to be detected; the first segmentation module 402 is used to segment the speech to be detected to obtain at least one segmented speech; the second acquisition module 403 is used to process the segmented speech according to the setting parameters to obtain text information and emotional information of each segmented speech; the third acquisition module 404 is used to obtain process quality inspection information and keyword detection results of the speech to be detected based on the text information and the emotional information.

[0090] Optionally, the first acquisition module 401 is used to monitor the first message queue and acquire the setting parameters stored in the first message queue.

[0091] Optionally, the setting parameters include identification information of the application scenario of the speech to be detected and speech emotion model information of the speech to be detected; the second acquisition module 403 includes:

[0092] A first acquisition submodule is configured to convert the segmented speech into text using speech recognition technology according to identification information of the application scenario of the speech to be detected in the setting parameters, so as to obtain text information of the segmented speech;

[0093] A second acquisition submodule is configured to perform emotion recognition on the segmented speech using a speech emotion recognition algorithm according to the speech emotion model information of the speech to be detected in the setting parameters, so as to obtain emotion information of the segmented speech;

[0094] The first storage submodule is configured to store the text information of the segmented speech and the emotional information of the segmented speech in a second message queue.

[0095] Optionally, the first segmentation module 402 includes:

[0096] The first recognition submodule is used to perform channel recognition on the speech to be detected to obtain first-channel speech and second-channel speech; the first segmentation submodule is used to segment the first-channel speech and the second-channel speech respectively through speech endpoint detection technology to obtain at least one first-channel segmented speech and at least one second-channel segmented speech.

[0097] The first acquisition submodule is configured to convert the first channel segmented speech and the second channel segmented speech into text using speech recognition technology according to the application scenario of the speech to be detected in the setting parameters, thereby obtaining text information of the first channel segmented speech and text information of the second channel segmented speech;

[0098] The second acquisition submodule is configured to perform emotion recognition on the first channel segmented speech and the second channel segmented speech respectively using a speech emotion recognition algorithm according to the speech emotion model information of the speech to be detected in the setting parameters, to obtain emotion information of the first channel segmented speech and emotion information of the second channel segmented speech;

[0099] The first storage submodule is used to store the text information of the first channel segmented speech, the text information of the second channel segmented speech, the emotional information of the first channel segmented speech, and the emotional information of the second channel segmented speech in the second message queue.

[0100] Optionally, the first storage submodule includes:

[0101] The first splicing unit is used to splice the text information of the first channel segmented speech, the text information of the second channel segmented speech, the emotional information of the first channel segmented speech, and the emotional information of the second channel segmented speech in chronological order to obtain a splicing result; and the storage unit is used to store the splicing result in a second message queue.

[0102] Optionally, the third obtaining module 404 includes:

[0103] A first processing submodule is configured to compare the text information and the emotional information in the second message queue with a preset process text through a quality inspection process algorithm to obtain process quality inspection information, wherein the preset process text is a process standard text corresponding to the application scenario of the voice to be detected;

[0104] The second processing submodule is configured to compare the text information with preset keywords through a keyword detection algorithm to obtain a keyword detection result.

[0105] Optionally, the device further includes:

[0106] The display module is used to display the text information, the emotional information, the process quality inspection information and the keyword detection results.

[0107] The device provided in the embodiment of the present application can execute the above method embodiment, and its implementation principle and technical effects are similar, so this embodiment will not be repeated here.

[0108] An embodiment of the present application also provides an electronic device, comprising: a memory, a processor, and a program stored in the memory and executable on the processor; the processor is configured to read the program in the memory to implement the various processes of the above-mentioned speech detection method embodiment.

[0109] The embodiment of the present application also provides a readable storage medium, on which a program is stored. When the program is executed by the processor, each process of the above-mentioned speech detection method embodiment is implemented, and the same technical effect is achieved. To avoid repetition, it is not repeated here. Wherein, the readable storage medium can be any available medium or data storage device that can be accessed by the processor, including but not limited to magnetic storage (such as floppy disk, hard disk, magnetic tape, magneto-optical disk (MO), etc.), optical storage (such as CD, DVD, BD, HVD, etc.), and semiconductor storage (such as ROM, EPROM, EEPROM, non-volatile memory (NAND FLASH), solid-state drive (SSD)), etc.

[0110] It should be noted that, in this document, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or apparatus comprising the element.

[0111] Through the description of the above embodiments, those skilled in the art can clearly understand that the above embodiment methods can be implemented by means of software plus the necessary general hardware platform, and of course can also be implemented by hardware, but in many cases the former is a better embodiment. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, disk, CD-ROM), and includes a number of instructions for enabling a terminal (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in each embodiment of the present application.

[0112] The embodiments of the present application are described above in conjunction with the accompanying drawings, but the present application is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the guidance of this application, ordinary technicians in this field can also make many forms without departing from the purpose of this application and the scope of protection of the claims, all of which are within the protection of this application.

Claims

1. A speech detection method, characterized in that: include: Acquire setting parameters of a speech to be detected and the speech to be detected, wherein the setting parameters include identification information of an application scenario of the speech to be detected and speech emotion model information of the speech to be detected; Segmenting the speech to be detected to obtain at least one segmented speech; Processing the segmented speech according to the identification information of the application scenario of the speech to be detected and the speech emotion model information of the speech to be detected to obtain text information and emotion information of each segmented speech; According to the text information and the emotion information, process quality inspection information and keyword detection results of the speech to be detected are obtained.

2. The method according to claim 1, characterized in that The step of obtaining the setting parameters of the voice to be detected includes: The first message queue is monitored to obtain the setting parameters stored in the first message queue.

3. The method according to claim 1, characterized in that The processing of the segmented speech according to the identification information of the application scenario of the speech to be detected and the speech emotion model information of the speech to be detected to obtain text information and emotion information of each segmented speech includes: According to the identification information of the application scenario of the speech to be detected, the segmented speech is converted into text using speech recognition technology to obtain text information of the segmented speech; According to the speech emotion model information of the speech to be detected, using a speech emotion recognition algorithm to perform emotion recognition on the segmented speech to obtain emotion information of the segmented speech; The text information of the segmented speech and the emotional information of the segmented speech are stored in a second message queue.

4. The method according to claim 3, characterized in that The segmenting of the speech to be detected to obtain at least one segmented speech includes: Performing channel recognition on the speech to be detected to obtain a first channel speech and a second channel speech; The first channel speech and the second channel speech are segmented and processed respectively by using speech endpoint detection technology to obtain at least one first channel segmented speech and at least one second channel segmented speech.

5. The method according to claim 4, characterized in that According to the identification information of the application scenario of the speech to be detected, the segmented speech is converted into text using speech recognition technology to obtain text information of the segmented speech, including: According to the application scenario of the speech to be detected, using speech recognition technology to convert the first channel segmented speech and the second channel segmented speech into text, respectively, to obtain text information of the first channel segmented speech and text information of the second channel segmented speech; According to the speech emotion model information of the speech to be detected, using a speech emotion recognition algorithm to perform emotion recognition on the segmented speech to obtain emotion information of the segmented speech, including: According to the speech emotion model information of the speech to be detected in the setting parameters, using a speech emotion recognition algorithm to perform emotion recognition on the first channel segmented speech and the second channel segmented speech respectively, to obtain emotion information of the first channel segmented speech and emotion information of the second channel segmented speech; The storing of the text information of the segmented speech and the emotional information of the segmented speech in the second message queue includes: The text information of the first channel segmented speech, the text information of the second channel segmented speech, the emotional information of the first channel segmented speech, and the emotional information of the second channel segmented speech are stored in the second message queue.

6. The method according to claim 5, characterized in that The storing the text information of the first channel segmented speech, the text information of the second channel segmented speech, the emotional information of the first channel segmented speech, and the emotional information of the second channel segmented speech into the second message queue includes: splicing the text information of the first channel segmented speech, the text information of the second channel segmented speech, the emotional information of the first channel segmented speech, and the emotional information of the second channel segmented speech in chronological order to obtain a splicing result; The splicing result is stored in the second message queue.

7. The method according to claim 1, characterized in that The obtaining, based on the text information and the emotion information, process quality inspection information and keyword detection results of the speech to be detected includes: Comparing the text information and the emotional information in the second message queue with a preset process text through a quality inspection process algorithm to obtain process quality inspection information, wherein the preset process text is a process standard text corresponding to the application scenario of the voice to be detected; The text information is compared with preset keywords through a keyword detection algorithm to obtain a keyword detection result.

8. A speech detection device, characterized in that: include: A first acquisition module is used to acquire setting parameters of the speech to be detected and the speech to be detected, wherein the setting parameters include identification information of the application scenario of the speech to be detected and speech emotion model information of the speech to be detected; A first segmentation module, configured to segment the speech to be detected to obtain at least one segmented speech; A second acquisition module is used to process the segmented speech according to the identification information of the application scenario of the speech to be detected and the speech emotion model information of the speech to be detected, and obtain text information and emotion information of each segmented speech; The third acquisition module is used to obtain process quality inspection information and keyword detection results of the speech to be detected based on the text information and the emotion information.

9. An electronic device comprising: A memory, a processor, and a program stored in the memory and executable on the processor; wherein the processor is configured to read the program in the memory to implement the steps of the speech detection method as described in any one of claims 1 to 7.

10. A readable storage medium for storing a program, characterized in that: When the program is executed by a processor, the steps of the speech detection method according to any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Voice quality inspection method and device

    CN111916110A

  • Telephone traffic quality inspection method and device

    CN112580367A

  • Customer service call voice quality inspection method and device, electronic equipment and storage medium

    CN112804400A

  • Voice quality inspection method and system and storage medium

    CN112885332A