Voice response processing method and device, electronic equipment and storage medium
By identifying the voiceprint and semantic features of voice data in the customer service center and selecting appropriate task processing resources based on the voiceprint blacklist information, the customer service center's identification and processing problems of blacklist customers are solved, and resource utilization efficiency and customer satisfaction are improved.
Patent Information
- Application Number
- CN202510659095.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-21
- Publication Date
- 2025-08-15
AI Technical Summary
In the prior art, it is difficult for customer service centers to effectively identify and intercept blacklist customers, resulting in low morale among agents and decreased customer satisfaction, and unable to meet the needs of efficient management and risk prevention and control.
By determining the vocal print characteristics and semantic features of the voice data in the interactive voice response stage, combining the vocal print blacklist information, selecting appropriate task processing resources to process voice interaction tasks, and improving resource utilization efficiency and problem solving speed.
It realizes accurate identification and effective processing of blacklisted customers, reduces waiting time, and improves the experience during voice interaction and the operation efficiency of customer service centers.
Smart Images

Figure CN120496580A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of voice processing technology, and in particular to a voice response processing method, device, electronic device and storage medium. Background Art
[0002] In a call center's operations system, agents are a core resource, and their work status directly determines the center's overall performance. Frequent customer complaints and harassment can severely dampen agent enthusiasm, lower morale, and ultimately weaken the center's operational stability and efficiency. Furthermore, improperly handling customer complaints can lead to decreased customer satisfaction.
[0003] Currently, the use of technology to identify and filter blacklisted customers primarily relies on agents manually counting and labeling problematic customers, and then entering information such as their calling numbers and ID numbers into a blacklist database. When a customer calls the interactive voice response system, the system compares the calling number and ID number to determine if they are on the blacklist. However, this method has significant flaws. Customers can easily bypass blacklist verification by changing their calling device number or logging in using someone else's card information. This makes it impossible to effectively intercept blacklisted customers, making it difficult to meet the actual needs of customer service centers for efficient management and risk prevention. Summary of the Invention
[0004] The present invention provides a voice response processing method, device, electronic device and storage medium to improve the efficiency of resolving customer demands and avoid unnecessary interference or loss to the customer service center.
[0005] According to one aspect of the present invention, a method for processing a voice response is provided, wherein the method comprises:
[0006] Determining a first voiceprint feature of first voice data, where the first voice data includes voice data generated by performing a first voice interaction task during an interactive voice response phase;
[0007] Determining voiceprint blacklist information, the voiceprint blacklist information including a plurality of second voiceprint features and semantic features associated with the second voiceprint features, where the second voiceprint features are voiceprint features of second voice data generated when performing a second voice interaction task, each second voice interaction task being a voice interaction task that has been performed before performing the first voice interaction task, and the semantic features associated with each second voiceprint feature are used to indicate appeal information conveyed through voice expression during the execution of the second voice interaction task and an emotional feature contained in the appeal information conveyed through voice expression;
[0008] Based on the first voiceprint feature and the voiceprint blacklist information, a task processing resource matching the first voice interaction task is determined from multiple task processing resources to assist in processing the first voice interaction task that continues to be executed after the interactive voice response phase ends.
[0009] According to another aspect of the present invention, a voice response processing device is provided, wherein the device comprises:
[0010] a determining module configured to determine a first voiceprint feature of first voice data, wherein the first voice data includes voice data generated by performing a first voice interaction task during an interactive voice response phase;
[0011] The determination module is further configured to determine voiceprint blacklist information, the voiceprint blacklist information including a plurality of second voiceprint features and semantic features associated with the second voiceprint features, the second voiceprint features being voiceprint features of second voice data generated when performing a second voice interaction task, each second voice interaction task being a voice interaction task that has been performed before performing the first voice interaction task, and the semantic features associated with each second voiceprint feature being used to indicate appeal information conveyed through voice expression during the execution of the second voice interaction task and an emotional feature contained in the appeal information conveyed through voice expression;
[0012] A processing module is used to determine a task processing resource that matches the first voice interaction task from multiple task processing resources based on the first voiceprint feature and the voiceprint blacklist information, and to assist in processing the first voice interaction task that continues to be executed after the interactive voice response stage ends.
[0013] According to another aspect of the present invention, an electronic device is provided, comprising:
[0014] at least one processor; and
[0015] a memory communicatively connected to the at least one processor; wherein,
[0016] The memory stores a computer program that can be executed by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the voice response processing method described in any embodiment of the present invention.
[0017] According to another aspect of the present invention, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the voice response processing method described in any embodiment of the present invention when executed.
[0018] The technical solution of the embodiment of the present invention, when executing the first voice interaction task in the interactive voice response stage, by determining the first voiceprint feature of the first voice data, the ownership of the first voice data can be accurately identified according to the uniqueness of the voiceprint. The voiceprint blacklist information not only includes the voiceprint feature, but also is associated with the semantic feature. The semantic feature indicates the demand information during the execution of the voice interaction task, and the emotional feature reflects the emotional state when expressing the demand through language. This association helps to match the appropriate semantic feature according to the voiceprint feature, that is, according to the voiceprint feature and the semantic and emotional information reflected in the previous voice interaction task, find the semantic feature related to the first voice data, and then select a more suitable resource from multiple task processing resources according to the semantic feature related to the first voice data to assist in processing the first voice interaction task that continues to be executed after the interactive voice response stage ends, thereby improving the utilization efficiency of the task processing resources. Moreover, by reasonably allocating task processing resources, problems arising in the voice interaction task can be quickly and effectively solved, the waiting time can be reduced, and the experience in the voice interaction process can be improved.
[0019] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present invention, nor is it intended to limit the scope of the present invention. Other features of the present invention will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0021] Figure 1 is a flow chart of a voice response processing method provided according to an embodiment of the present invention;
[0022] Figure 2 is a diagram of a voice response processing architecture applicable to an embodiment of the present invention;
[0023] Figure 3 This is a flow chart of a voice response processing method provided by an embodiment of the present invention;
[0024] Figure 4 2 is a schematic structural diagram of a voice response processing device provided according to an embodiment of the present invention;
[0025] Figure 5 The figure is a schematic diagram of the structure of an electronic device for implementing the voice response processing method according to an embodiment of the present invention. DETAILED DESCRIPTION
[0026] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.
[0027] It should be noted that the terms "first", "second", etc. in the description and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the numbers used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0028] Figure 1 A flowchart of a voice response processing method is provided for an embodiment of the present invention. This embodiment is applicable to situations where appropriate task processing resources can be allocated in a timely manner when executing a first voice interaction task in an interactive voice response phase so that appropriate task processing resources can be allocated after the interactive voice response phase ends to enter the manual response phase for voice response processing. The method can be executed by a voice response processing device, which can be implemented in the form of hardware and / or software, and can be configured in any electronic device with network communication function.
[0029] like Figure 1 As shown, the voice response processing method provided in this embodiment may include the following process:
[0030] S110. Determine a first voiceprint feature of first voice data, where the first voice data includes voice data generated by performing a first voice interaction task in an interactive voice response phase.
[0031] At the initial stage of call access, the system immediately enters the interactive voice response (IVR) stage and simultaneously performs the following first voice interaction tasks: automatically plays a standardized voice navigation menu to guide users to select service categories through keystrokes or voice commands; recognizes and analyzes the voice information input by users in real time, and intelligently matches the corresponding service process; and completes interactive tasks such as information query, business processing, or transfer to human agents based on operating instructions, thus achieving efficient human-computer voice dialogue.
[0032] Among them, interactive voice response (IVR) is an automatic voice service system that uses pre-recorded voice prompts and voice recognition technology for voice interaction, guiding customers to complete a series of operations through keystrokes or voice commands, such as querying information, handling business, selecting service types, etc. The entire process does not require direct human participation, and customers can complete related operations independently according to the system's voice prompts.
[0033] The first voice data may be audio information input to the interactive voice response (IVR) system via voice when performing a first voice interaction task with the interactive voice response (IVR) system during the interactive voice response phase. For example, the first voice data may include voice data generated during the voice interaction with the interactive voice response (IVR) system and containing at least one of a number, a keyword, a phrase, or a complete sentence. The first voice data may be a response to a prompt of the interactive voice response (IVR) system during the interactive voice response phase, or a question or request input to the interactive voice response (IVR) system during the interactive voice response phase. For example, the response information generated in response to a prompt issued by the interactive voice response (IVR) system indicates a request to perform a specific task.
[0034] Voiceprint features are unique and stable characteristics contained in speech data. They can describe at least one of the following: frequency, amplitude, and pitch. Extracting and analyzing this information through acoustic analysis and signal processing techniques can serve as an important basis for identity verification and voice recognition. Voiceprint features are a set of acoustic parameters extracted from speech data that can reflect the physiological and behavioral characteristics of the voice in the speech data, thereby distinguishing different sound sources. The first voiceprint feature is an acoustic parameter feature extracted from the first speech data that reflects the physiological and behavioral characteristics of the voice in the first speech data.
[0035] As an optional but non-limiting implementation, determining the first voiceprint feature of the first voice data includes the following steps:
[0036] Acquire first voice data generated when the first voice interaction task is executed; and extract a first voiceprint feature from the first voice data using a feature extraction method of Mel-frequency cepstral coefficients.
[0037] During the execution of the first voice interaction task in the interactive voice response phase, when voice interaction is performed with the interactive voice response system, first voice data input to the interactive voice response system through voice will be recorded during the execution of the voice interaction task.
[0038] Mel-frequency cepstral coefficients (MFCC) are a feature extraction algorithm commonly used in speech signal processing. They have good sound differentiation capabilities and can reflect important speech features such as pitch and timbre. They are widely used in speech recognition and identification of the sound source to which speech belongs.
[0039] The specific implementation process of the feature extraction method of Mel-frequency cepstral coefficients includes: first, framing the first voice data. The first voice data is a signal that changes with time. Through framing, the first voice data can be regarded as a series of short-time stable signals for processing; then, fast Fourier transform (FFT) is performed on each frame of voice data of the first voice data, and each frame of voice data of the first voice data is converted into the frequency domain to obtain a spectrum; then, the spectrum of each frame of voice data of the first voice data is filtered according to the Mel-frequency scale, and the spectrum is filtered through a group of Mel filter groups to obtain a Mel-spectrum; finally, the logarithm of the Mel-spectrum is taken and discrete cosine transform (DCT) is performed to obtain Mel-frequency cepstral coefficients. These coefficients can effectively describe the first voiceprint features of the first voice data, and have a strong characterization ability for the timbre, pronunciation method, etc. of the voice.
[0040] Voiceprints are unique voice features of each person, just like fingerprints, and have individual differences. After processing the first voice data using the Mel-frequency cepstral coefficient method, the resulting MFCC features contain characteristic information related to the speaker, which can be used to characterize the speaker's voiceprint. Specifically, due to physiological factors such as vocal cord length, thickness, and oral shape, as well as psychological factors such as pronunciation habits, different people will have different Mel-frequency cepstral coefficients of their voice signals when uttering the same voice. By extracting these coefficients, these differences can be quantified as voiceprint features for subsequent applications such as voiceprint recognition and voice authentication.
[0041] As an optional but non-limiting implementation solution, before extracting the first voiceprint feature from the first voice data, the following steps are further included:
[0042] A first preprocessing operation is performed on the first speech data, where the first preprocessing operation includes at least one of the following: noise reduction, gain, framing, windowing, pre-emphasis, endpoint detection, and sampling rate normalization.
[0043] Noise reduction on the first voice data can be used to remove background noise contained in the first voice data, making the first semantic data clearer. For example, voice data recorded in a noisy environment may contain various environmental noises, such as wind and machine sounds. Noise reduction is designed to minimize the interference of these noises on the voice signal. Noise reduction on the first voice data can improve the quality and intelligibility of the voice data, reduce the impact of noise on voice processing tasks, and improve recognition accuracy.
[0044] Applying gain to the first voice data can be used to adjust the amplitude of the voice signal of the first voice data, including increasing or decreasing the strength of the voice signal. If the voice signal of the first voice data is generally weak, the gain operation can amplify it to an appropriate amplitude range, ensuring that the energy of the voice signal is within the appropriate range. This avoids information loss due to a weak signal or distortion due to an overly strong signal, thereby improving the performance of the voice processing system.
[0045] Framing the first voice data can be used to divide the continuous first voice data signal into several shorter frames, typically tens of milliseconds in length. Because the first voice data signal exhibits short-term stationary characteristics and its features do not change significantly over a short period of time, framing facilitates more detailed analysis and processing of the voice signal. For example, when extracting voice features, calculations are typically performed on a frame-by-frame basis, helping to improve feature extraction accuracy.
[0046] Windowing the first speech data can be performed by multiplying each frame of the first speech data by a window function based on framing. The window function is used to smoothly transition the speech signals at both ends of the frame to zero, thereby reducing edge effects caused by framing. For example, the window function can be at least one of a Hamming window and a Hanning window. Windowing the first speech data can reduce spectral leakage, make spectral analysis more accurate, and improve the stability and reliability of speech features.
[0047] Pre-emphasis on the first voice data can boost the energy of the high-frequency portion of the voice signal. Because the high-frequency portion of a voice signal is relatively weak and easily attenuated during transmission, pre-emphasis can enhance this high-frequency information. Pre-emphasis on the first voice data can highlight high-frequency details in the voice signal, improving voice clarity and intelligibility. This is particularly true for phonemes with high high-frequency components, such as fricatives, enabling better recognition of these phonemes in tasks such as speech recognition.
[0048] Endpoint detection on the first voice data can be used to determine the starting and ending points of the voice signal, distinguishing speech from silence. This removes silence before and after speech, reducing unnecessary data processing. Performing endpoint detection on the first voice data improves voice processing efficiency, reduces the processing of invalid data, and helps improve the accuracy of tasks such as speech recognition by preventing silence from interfering with recognition results.
[0049] Sampling rate normalization of the first voice data can be used to convert voice signals with different sampling rates to a uniform sampling rate. This is because different recording devices or acquisition environments may result in different sampling rates for voice signals, and subsequent voice processing algorithms generally require a fixed sampling rate. Sampling rate normalization of the first voice data enables unified processing and analysis of the voice data, avoiding algorithm errors or performance degradation caused by inconsistent sampling rates.
[0050] S120. Determine the voiceprint blacklist information, where the voiceprint blacklist information includes multiple second voiceprint features and semantic features associated with the second voiceprint features. The second voiceprint feature is the voiceprint feature of the second voice data generated when executing the second voice interaction task. Each second voice interaction task is a voice interaction task that has been executed before executing the first voice interaction task. The semantic features associated with each second voiceprint feature are used to indicate the appeal information conveyed through voice expression during the execution of the second voice interaction task and the emotional features contained in the appeal information conveyed through voice expression.
[0051] Think of a voiceprint blacklist as a database or record collection that stores specific information, including multiple second voiceprint features and semantic features associated with them. The second voiceprint features included in the voiceprint blacklist are the unique acoustic characteristics of voice data identified as requiring restriction or special processing in voice interaction scenarios. The second voiceprint features in the voiceprint blacklist are used to identify voice data with the second voiceprint features and to restrict or perform special processing during voice interaction.
[0052] The second voice interaction task may be a voice interaction task that has been executed before the first voice interaction task is executed. The second voice interaction task includes voice interaction tasks executed during the interactive voice response phase and the manual response phase after the interactive voice response phase. The second voiceprint feature is a voiceprint feature extracted from the voice data generated during the execution of the second voice interaction task.
[0053] The appeal conveyed through speech refers to the purpose or information desired to be achieved through speech expression, focusing more on the subjective intention and appeal of the language. For example, the appeal conveyed through speech expression can be asking about the weather, seeking advice, expressing a certain emotion, issuing a command, etc.
[0054] The emotional characteristics conveyed by voice when conveying a demand reflect the urgency of the demand and the degree of expectation of its satisfaction. These emotional characteristics can be characterized by intonation, speech rate, volume, and timbre. For example, emotional characteristics might be characterized by anxiety, calmness, or anger.
[0055] The voiceprint blacklist information is set based on the previous voice interaction situation. When executing the first voice interaction task, the voiceprint and semantic features generated by the previous second voice interaction task can be referred to to determine whether the conditions of the voiceprint blacklist are met. For example, some combinations of voiceprint features and semantic features can be used as the basis for determining whether a voiceprint is added to the blacklist, or special processing can be performed on specific voiceprints and semantic situations when executing the first voice interaction task.
[0056] S130. Determine, based on the first voiceprint feature and the voiceprint blacklist information, a task processing resource that matches the first voice interaction task from a plurality of task processing resources, for assisting in processing the first voice interaction task that continues to be executed after the interactive voice response phase ends.
[0057] The interactive voice response phase is the initial stage of executing the first voice interaction task. During this phase, the interactive voice response (IVR) system comes into play. The IVR system conducts voice interaction using pre-set voice prompts. Sometimes, the IVR system cannot meet the needs of the first voice interaction task. In this case, the manual response phase begins. Direct communication through an agent provides more flexible and personalized service, resolving complex issues that are difficult for the IVR system to handle, ensuring that the user's problem is effectively resolved.
[0058] When the following situations occur in the interactive voice response stage, you can enter the manual response stage from the interactive voice response stage: during the interaction with the IVR system, it is found that the problem cannot be solved through the IVR system's automatic service, so you choose to transfer to manual customer service to enter the manual response stage; for some complex business demands or special situations, the IVR system cannot accurately understand or process the user's instructions, so it automatically enters the manual response stage; in some cases, there may be emotional excitement or emotional support required during the voice interaction, and the IVR system is difficult to provide such humane service, so it automatically enters the manual response stage.
[0059] The first voiceprint feature refers to the voiceprint information extracted from the currently ongoing first voice interaction task. The voiceprint blacklist information is a collection of second voiceprint features. These second voiceprint features are usually associated with situations that require special treatment. For example, voiceprint features with a history of bad behavior or violations of regulations will be recorded in the blacklist. By comparing the first voiceprint feature with each second voiceprint feature recorded in the voiceprint blacklist information, it can be determined whether the first voiceprint feature belongs to the voiceprint feature recorded in the voiceprint blacklist information.
[0060] Task processing resources refer to various resources capable of processing voice interaction tasks, such as different processing algorithms, specific processing equipment, and different agents. Different task processing resources may have different characteristics and capabilities, and their task processing performance and effectiveness may vary when processing the first voice interaction task. For example, some task processing resources may be more effective for certain types of voice interaction tasks, while other task processing resources may be less effective for certain types of voice interaction tasks.
[0061] Each second voiceprint feature in the voiceprint blacklist information is associated with a semantic feature. The semantic feature associated with each second voiceprint feature is used to indicate the demand information conveyed through voice expression during the execution of the second voice interaction task and the emotional features contained when the demand information is conveyed through voice expression. When the first voiceprint feature hits the second voiceprint feature in the voiceprint blacklist information, it means that the first voice interaction task may involve some special circumstances. When entering the manual response stage, task processing resources cannot be randomly selected, and appropriate task processing resources need to be selected for processing. Specifically, the task processing resources matching the first voice interaction task can be determined from multiple task processing resources pre-associated with the voiceprint blacklist information based on the semantic features associated with the first voiceprint feature and the second voiceprint feature hit by the first voiceprint feature.
[0062] The task processing resources matched with the first voice interaction task are used to assist in processing the first voice interaction task that continues to be executed after the interactive voice response stage. The interactive voice response stage is the voice interaction stage corresponding to the interactive voice response (IVR) system, and the interactive voice response (IVR) system will give a response result. At the end of the interactive voice response stage, entering the manual response stage will continue the unfinished first voice interaction task. At this time, the task processing resources matched with the first voice interaction task can be used to continue to assist in processing the first voice interaction task. Appropriate task processing resources can improve the efficiency and accuracy of task processing and better meet the demands of voice interaction.
[0063] The technical solution of the embodiment of the present invention, when executing the first voice interaction task in the interactive voice response stage, by determining the first voiceprint feature of the first voice data, the ownership of the first voice data can be accurately identified according to the uniqueness of the voiceprint. The voiceprint blacklist information not only includes the voiceprint feature, but also is associated with the semantic feature. The semantic feature indicates the demand information during the execution of the voice interaction task, and the emotional feature reflects the emotional state when expressing the demand through language. This association helps to match the appropriate semantic feature according to the voiceprint feature, that is, according to the voiceprint feature and the semantic and emotional information reflected in the previous voice interaction task, find the semantic feature related to the first voice data, and then select a more suitable resource from multiple task processing resources according to the semantic feature related to the first voice data to assist in processing the first voice interaction task that continues to be executed after the interactive voice response stage ends, thereby improving the utilization efficiency of the task processing resources. Moreover, by reasonably allocating task processing resources, problems arising in the voice interaction task can be quickly and effectively solved, the waiting time can be reduced, and the experience in the voice interaction process can be improved.
[0064] Based on the above embodiment, optionally, the voiceprint blacklist information is generated in the following manner:
[0065] Acquire second voice data, and perform semantic recognition on the second voice data to obtain semantic features contained in the second voice data, where the semantic features contained in the second voice data include appeal information conveyed through voice expression during the execution of the second voice interaction task and emotional features contained when the appeal information is conveyed through voice expression; when the semantic features contained in the second voice data hit the semantic feature category associated with the voiceprint blacklist information, extract the second voiceprint feature from the second voice data; determine the semantic features contained in the second voice data as semantic features associated with the second voiceprint feature, and use them to generate voiceprint blacklist information.
[0066] The voiceprint blacklist information is associated with the voiceprint features of voice data identified as requiring restriction or special processing in voice interaction scenarios. Although different voiceprint features are included in the scope of restriction or special processing, the semantic characteristics exhibited by voice data with different voiceprint features may vary. Therefore, the task processing resources required for voice interaction tasks for voice data with different voiceprint features will vary. Therefore, it is necessary to select appropriate task processing resources based on the semantic feature classification contained in the voice data to which the voiceprint feature belongs.
[0067] The semantic feature category associated with the voiceprint blacklist information is a pre-set semantic feature category of voice data that is identified as requiring restriction or special processing in a voice interaction scenario. In the case where the semantic feature contained in the second voice data hits the semantic feature category associated with the voiceprint blacklist information, it indicates that the second voice data is identified as voice data that requires restriction or special processing in the voice interaction scenario. Therefore, the voiceprint feature of the second voice data needs to be saved so that the next time the same voiceprint feature appears, the task processing resources selected from the task processing resources associated with the voiceprint blacklist information can be used to process the voice interaction task in the manual response stage, avoiding the problem of poor voice interaction task processing performance caused by random selection of task processing resources in the manual response stage.
[0068] If the semantic features contained in the second voice data match the semantic feature category associated with the voiceprint blacklist information, the voiceprint features are extracted from the second voice data, and the second voiceprint features extracted from the second voice data are associated with the semantic features of the second voice data and stored. Optionally, extracting the second voiceprint features from the second voice data includes extracting the second voiceprint features from the second voice data using a feature extraction method based on Mel-frequency cepstral coefficients.
[0069] By identifying and associating the semantic features of voice data, voiceprint blacklists no longer rely solely on voiceprint features. They also consider the demands conveyed through voice and the emotional characteristics inherent in those demands. This allows the blacklist to more accurately capture the semantic characteristics of individuals who demonstrate negative behavior or potential risks. When the voice interaction system detects that the semantic features of voice data match a blacklisted category, it extracts the voiceprint features and updates the blacklist, helping the system promptly identify and prevent potential security threats.
[0070] See also Figure 2The speech processing unit includes a semantic understander and a voiceprint feature extractor. The speech processing unit operates on two processes. One process is used for the second speech data. The semantic understander performs semantic recognition on the second speech data to obtain the semantic features contained in the second speech data. The semantic features contained in the second speech data include the appeal information conveyed through speech expression during the execution of the second speech interaction task and the emotional features contained when the appeal information is conveyed through speech expression. When the semantic features contained in the second speech data hit the semantic feature category associated with the voiceprint blacklist information, the second voiceprint feature is extracted from the second speech data; the semantic features contained in the second speech data are determined as the semantic features associated with the second voiceprint feature, and the second voiceprint feature and the semantic features associated with the second voiceprint feature are transmitted to the information management unit to generate voiceprint blacklist information for management. The other process is to directly use the voiceprint feature extractor to extract the first voiceprint feature of the second speech data, and also transmit the first voiceprint feature to the information management unit.
[0071] Optionally, the first voice data is the recording data generated by the recording system of the telephone customer service center when performing a first voice interaction task with the interactive voice response system during the interactive response phase, and the second voice data is the recording data generated by the recording system of the telephone customer service center when performing a second voice interaction task with the interactive voice response system, and each second voice interaction task is a historical voice interaction task that has been performed before executing the first voice interaction task.
[0072] As an optional but non-limiting implementation scheme, performing semantic recognition on the second voice data to obtain semantic features contained in the second voice data includes but is not limited to the following steps:
[0073] Determine the text data of the second voice data, and input the text data of the second voice data into a semantic recognition model for semantic recognition; the semantic recognition model is a multi-task neural network model for parsing and recognizing the semantic information expressed by the voice signal; and output the semantic features contained in the second voice data through the semantic recognition model.
[0074] A multi-task neural network model is a neural network architecture that can simultaneously handle multiple related or unrelated tasks. It is used in fields such as image recognition and segmentation, natural language processing, and other fields. For example, it can simultaneously perform image classification and object detection, or text classification and sentiment analysis. Semantic recognition is the process of parsing and identifying the semantic information expressed in speech signals. This process can identify and understand information such as intent, emotion, and knowledge in speech. Multi-task neural network models for semantic recognition can be built based on the recurrent neural network (RNN) framework.
[0075] As an optional but non-limiting implementation solution, before inputting the text data of the second voice data into the semantic recognition model for semantic recognition, the following steps are further included:
[0076] A second preprocessing operation is performed on the text data of the second speech data, where the second preprocessing operation includes at least one of the following: text cleaning and text segmentation.
[0077] As an optional but non-limiting implementation, the semantic recognition model is generated in the following way:
[0078] Acquire training sample data for a semantic recognition training task, where the training sample data includes text data of the third voice data and semantic feature labels of the text data of the third voice data; input the training sample data into a semantic recognition model to perform recognition and prediction of semantic features to obtain semantic feature prediction results; use a cross-entropy loss function to calculate the loss difference between the semantic feature prediction results and the semantic feature labels, and backpropagate the gradients of the parameters in the semantic recognition model based on the loss difference to update the parameters of the semantic recognition model.
[0079] See also Figure 2 , obtain the text data of the third voice data, use the data annotator of the model training unit to label the text data of the third voice data to obtain the semantic feature label of the text data of the third voice data, and form the text data of the third voice data and the semantic feature label of the text data of the third voice data into training sample data. After building the semantic recognition model using the recurrent neural network RNN framework, the network trainer of the model training unit extracts sample data from the training sample data and inputs it into the semantic recognition model for forward propagation, calculates and outputs the semantic feature prediction result, and uses the cross entropy loss function to calculate the loss difference between the output semantic feature prediction result and the semantic feature label. According to the loss difference calculated by the loss function, the gradient of the model parameters of the semantic recognition model is back-propagated, the semantic recognition model parameters are updated, and the above process is repeated until the semantic recognition model meets the requirements.
[0080] See also Figure 2 The model training unit includes a data annotator and a network trainer. The data annotator receives the text data of the third voice data maintained by the information management unit and annotates the text data of the third voice data with semantic feature labels, such as data annotations for behaviors such as harassment, malicious complaints, and fraud during semantic interaction tasks. The network trainer is used to train the multi-task neural network model used for semantic recognition. It collects and pre-processes the text data of the third voice data that matches the set of semantic feature categories associated with the voiceprint blacklist information as training sample data.
[0081] Figure 3A flow chart of another voice response processing method provided for an embodiment of the present invention. The technical solution of this embodiment further optimizes the process of determining the task processing resource matching the first voice interaction task from multiple task processing resources based on the first voiceprint feature and voiceprint blacklist information in the aforementioned embodiment on the basis of the technical solution of the aforementioned embodiment. This embodiment can be combined with various optional schemes in one or more of the above embodiments.
[0082] like Figure 3 As shown, the voice response processing method of the embodiment of the present invention may include the following process:
[0083] S310: Determine a first voiceprint feature of first voice data, where the first voice data includes voice data generated when performing a first voice interaction task in an interactive voice response phase.
[0084] S320. Determine the voiceprint blacklist information. The voiceprint blacklist information includes multiple second voiceprint features and semantic features associated with the second voiceprint features. The second voiceprint feature is the voiceprint feature of the second voice data generated when executing the second voice interaction task. Each second voice interaction task is a voice interaction task that has been executed before executing the first voice interaction task. The semantic features associated with each second voiceprint feature are used to indicate the appeal information conveyed through voice expression during the execution of the second voice interaction task and the emotional features contained in the appeal information conveyed through voice expression.
[0085] S330. When there is a third voiceprint feature that matches the first voiceprint feature in the voiceprint blacklist information, determine the semantic feature associated with the third voiceprint feature based on the voiceprint blacklist information, where the third voiceprint feature is the second voiceprint feature that matches the first voiceprint feature and is selected from the voiceprint blacklist information.
[0086] The first voiceprint feature currently obtained is compared with the voiceprint features in the voiceprint blacklist information. If a third voiceprint feature matching the first voiceprint feature is found in the voiceprint blacklist information, the third voiceprint feature is selected. The third voiceprint feature here already exists in the voiceprint blacklist information and is the second voiceprint feature already recorded in the voiceprint blacklist information. The feature similarity between the third voiceprint feature and the first voiceprint feature is greater than the feature similarity between the second voiceprint feature in the voiceprint blacklist information excluding the third voiceprint feature and the first voiceprint feature. The feature similarity between voiceprint features can be calculated using cosine similarity or Euclidean distance.
[0087] The voiceprint blacklist information not only includes the second voiceprint feature, but also has associated semantic features. Once a matching third voiceprint feature is found, the semantic features associated with the third voiceprint feature are retrieved from the voiceprint blacklist information. For example, the semantic features corresponding to the third voiceprint feature may include the appeal information conveyed through voice expression during the execution of the second voice interaction task, as well as the emotional features contained in the voice expression when conveying the appeal information. The second voice interaction task corresponding to the third voiceprint feature and the first voice interaction task come from the same sound source.
[0088] S340. According to the reference information of the third voiceprint feature, determine the task processing resource that matches the first voice interaction task from multiple task processing resources, and use it to assist in processing the first voice interaction task that continues to be executed after the interactive voice response stage ends. The reference information of the third voiceprint feature includes the semantic feature category to which the semantic feature associated with the third voiceprint feature belongs and the number of hits of the third voiceprint feature. The number of hits of the third voiceprint feature can be used to indicate the number of times the third voiceprint feature in the voiceprint blacklist information is matched or accessed during the query operation. The task processing resource that matches the first voice interaction task is associated with the reference information associated with the third voiceprint feature.
[0089] The reference information of the third voiceprint feature includes the semantic feature category and the number of times the third voiceprint feature was matched or accessed during the previous query operation. Furthermore, based on the reference information of the third voiceprint feature, a task processing resource that matches the first voice interaction task can be found from multiple optional task processing resources (such as different agents, different processing processes, etc.) pre-associated with the voiceprint blacklist information. For example, if the semantic feature category associated with the third voiceprint feature is "complaint" and the number of hits is large (indicating that the speaker has made multiple complaints), the agent with more experience in handling complaints will be determined as the matching task processing resource.
[0090] The task processing resource matched to the first voice interaction task is correlated with the reference information associated with the third voiceprint feature, indicating that the task processing resource was determined based on the reference information of the third voiceprint feature. Different reference information will result in different task processing resources being selected to better handle the first voice interaction task. For example, for users who are emotionally agitated and have filed multiple complaints, a task processing resource that is more adept at handling such situations will be selected to respond, thereby improving processing efficiency and user satisfaction.
[0091] Voiceprint matching and semantic feature analysis can more accurately understand user behavior patterns and appeal characteristics, thereby matching the most appropriate task processing resources for the first voice interaction task based on this information. Determining task processing resources based on reference information from the third voiceprint feature can achieve reasonable resource allocation, avoid wasting resources on inappropriate tasks, and improve resource utilization efficiency. The use of voiceprint blacklist information helps to identify potential risk users. By matching them with corresponding task processing resources, these users can be more effectively managed and monitored.
[0092] See also Figure 2 The information management unit is primarily used to process and manage the second voiceprint feature and the semantic feature information associated with the second voiceprint feature transmitted by the voice processing unit, perform routing allocation of traffic resources based on the first voiceprint feature and voiceprint blacklist information, and determine the task processing resource that matches the first voice interaction task from multiple task processing resources. The information management unit includes a semantic memory, a semantic manager, a voiceprint memory, a voiceprint manager, a voiceprint comparator, and a routing distributor.
[0093] As an optional but non-limiting implementation scheme, determining a task processing resource matching the first voice interaction task from multiple task processing resources based on reference information of the third voiceprint feature includes but is not limited to the following steps A1-A2:
[0094] Step A1. If it is determined based on the reference information of the third voiceprint feature that the hit number of the third voiceprint feature is greater than the preset hit number, then the task processing resources matching the first voice interaction task are determined from multiple task processing resources, and the task processing performance and / or task processing effect of the task processing resources matching the first voice interaction task when assisting in processing the first voice interaction task is greater than the task processing performance and / or task processing effect of the remaining task processing resources in the multiple task processing resources except the task processing resources matching the first voice interaction task.
[0095] Step A2: Based on the semantic feature category to which the semantic feature associated with the third voiceprint feature in the reference information associated with the third voiceprint feature belongs, select a task processing resource that matches the semantic feature category to which the semantic feature associated with the third voiceprint feature belongs from multiple task processing resources, and determine the task processing resource that matches the first voice interaction task, wherein different task processing resources among at least some of the multiple task processing resources are associated with different semantic feature categories.
[0096] The number of hits of the third voiceprint feature (i.e., the number of times the third voiceprint feature was matched or accessed in the previous query operation) is compared with a preset number of hits (the preset number of hits can be determined based on factors such as business needs and experience). If the number of hits of the third voiceprint feature exceeds the preset number of hits, it means that the sound source corresponding to the third voiceprint feature has a high-frequency record in the voiceprint blacklist information, and a task processing resource matching the first voice interaction task will be selected from multiple available task processing resources (for example, the multiple available task processing resources can be at least one of different agents, different processing algorithms, or different processing processes for handling customer requests).
[0097] The task processing resource matched to the first voice interaction task, when continuing to execute the first voice interaction task after the auxiliary processing interactive voice response phase, must have better processing performance (for example, which can be characterized by but not limited to at least one of processing speed, resource utilization, and response time) and / or processing effect (for example, which can be characterized by but not limited to at least one of problem resolution rate and user satisfaction) on the first voice interaction task than other unselected task processing resources. For example, for users with high frequency interactions, the system will select agents with more experience and stronger processing capabilities as matching task processing resources to better meet user needs.
[0098] The categories to which the semantic features associated with the third voiceprint features belong (such as complaint category, consultation category, business processing category, etc.) will be analyzed. These semantic feature categories are a way to classify the content of voice interaction and can reflect the type of task demands in voice interaction tasks. Based on the analysis of the above semantic feature categories, task processing resources that are suitable for the semantic feature category will be selected from multiple task processing resources. For example, if the semantic feature category is complaint category, an agent who is good at handling complaint issues or a specific processing flow will be selected as a matching task processing resource, and the selected task processing resource will be determined as a resource that matches the first voice interaction task. Moreover, among multiple task processing resources, at least some task processing resources will be associated with different semantic feature categories, which provides a basis for selecting appropriate task processing resources according to semantic feature categories.
[0099] See also Figure 2, semantic memory and voiceprint memory: responsible for entering the voiceprint blacklist information from the voice processing unit into the system. Semantic manager and voiceprint manager: review, confirm and delete the entered information to ensure accuracy and validity; classify the voiceprint blacklist information according to the reference information of the second voiceprint feature in the voiceprint blacklist information; perform statistical analysis on the voiceprint blacklist information to understand the relevant situation and trends, and adjust the labeled data of the model training unit according to the analysis results. The voiceprint comparator can match the extracted voiceprint features with the existing voiceprint features to authenticate the voiceprint blacklist information. The routing distributor can determine the task processing resource that matches the first voice interaction task from multiple task processing resources based on the first voiceprint feature and the voiceprint blacklist information.
[0100] By selecting task processing resources based on the number of hits and semantic feature categories, more appropriate resources can be allocated to the corresponding tasks. For situations with a high number of hits, selecting task processing resources with better performance can process tasks more quickly and reduce user waiting time. For tasks with different semantic feature categories, selecting matching task processing resources can more efficiently solve problems, improve the speed and quality of task processing, and thus enhance overall task processing efficiency. Matching users with more appropriate task processing resources can better meet their needs, and this method of determining task processing resources based on voiceprint feature reference information can rationally allocate system resources. This avoids wasted resources and focuses them on tasks and users that need them most.
[0101] The technical solution of the embodiment of the present invention, when executing the first voice interaction task in the interactive voice response stage, by determining the first voiceprint feature of the first voice data, the ownership of the first voice data can be accurately identified according to the uniqueness of the voiceprint. The voiceprint blacklist information not only includes the voiceprint feature, but also is associated with the semantic feature. The semantic feature indicates the demand information during the execution of the voice interaction task, and the emotional feature reflects the emotional state when expressing the demand through language. This association helps to match the appropriate semantic feature according to the voiceprint feature, that is, according to the voiceprint feature and the semantic and emotional information reflected in the previous voice interaction task, find the semantic feature related to the first voice data, and then select a more suitable resource from multiple task processing resources according to the semantic feature related to the first voice data to assist in processing the first voice interaction task that continues to be executed after the interactive voice response stage ends, thereby improving the utilization efficiency of the task processing resources. Moreover, by reasonably allocating task processing resources, problems arising in the voice interaction task can be quickly and effectively solved, the waiting time can be reduced, and the experience in the voice interaction process can be improved.
[0102] Moreover, the present invention only utilizes voiceprint features to authenticate blacklisted customers. Voiceprint authentication is unique and difficult to forge, so the present invention has higher accuracy and reliability in identifying blacklisted customers. The present invention utilizes a multi-task neural network model for semantic understanding, which can identify scenarios in a variety of blacklisted customer behavior sets. At the same time, the information management unit has a feedback mechanism for the model training unit, which can adjust the annotation strategy of the training data according to new data features, which helps to improve adaptability and accuracy. The semantic recognition function of the present invention processes multiple historical recordings of customers, and directly obtains historical recording data from the recording quality inspection system that is prevalent in customer service centers. It has a simple structure and does not require a separate recording device. In addition, semantic recognition is not performed on the current call, only the voiceprint is extracted, and a decision can be made as soon as the customer opens his mouth, which is conducive to a rapid response to the subsequent allocation of the current call, while reducing occasional misjudgments.
[0103] Figure 4 A structural schematic diagram of a voice response processing device is provided for an embodiment of the present invention. This embodiment is applicable to situations where appropriate task processing resources can be allocated in a timely manner when executing a first voice interaction task in an interactive voice response phase so that appropriate task processing resources can be allocated after the interactive voice response phase ends to enter the manual response phase for voice response processing. The voice response processing device can be implemented in the form of hardware and / or software, and the voice response processing device can be configured in any electronic device with network communication function.
[0104] like Figure 4 As shown, the voice response processing device provided in this embodiment may include the following:
[0105] A determination module 410 is configured to determine a first voiceprint feature of first voice data, where the first voice data includes voice data generated during a first voice interaction task performed during an interactive voice response phase;
[0106] The determination module 410 is further configured to determine voiceprint blacklist information, the voiceprint blacklist information including a plurality of second voiceprint features and semantic features associated with the second voiceprint features, the second voiceprint features being voiceprint features of second voice data generated when performing a second voice interaction task, each second voice interaction task being a voice interaction task that has been performed before performing the first voice interaction task, and the semantic features associated with each second voiceprint feature being used to indicate appeal information conveyed through voice expression during the execution of the second voice interaction task and the emotional features contained in the appeal information conveyed through voice expression;
[0107] The processing module 420 is used to determine the task processing resource that matches the first voice interaction task from multiple task processing resources based on the first voiceprint feature and the voiceprint blacklist information, and is used to assist in processing the first voice interaction task that continues to be executed after the interactive voice response stage ends.
[0108] Based on the above embodiment, optionally, determining the first voiceprint feature of the first voice data includes:
[0109] Acquire first voice data generated when the first voice interaction task is executed;
[0110] The first voiceprint feature is extracted from the first speech data by using a feature extraction method of Mel-frequency cepstral coefficients.
[0111] Based on the above embodiment, optionally, the voiceprint blacklist information is generated in the following manner:
[0112] Acquiring second voice data and performing semantic recognition on the second voice data to obtain semantic features contained in the second voice data, where the semantic features contained in the second voice data include appeal information conveyed through voice expression during the execution of the second voice interaction task and emotional features contained in the appeal information conveyed through voice expression;
[0113] extracting a second voiceprint feature from the second voice data when the semantic feature contained in the second voice data matches the semantic feature category associated with the voiceprint blacklist information;
[0114] The semantic feature contained in the second voice data is determined as the semantic feature associated with the second voiceprint feature, and is used to generate voiceprint blacklist information.
[0115] Based on the above embodiment, optionally, performing semantic recognition on the second voice data to obtain semantic features contained in the second voice data includes:
[0116] Determining text data of the second voice data, and inputting the text data of the second voice data into a semantic recognition model for semantic recognition; the semantic recognition model is a multi-task neural network model for parsing and recognizing semantic information expressed by the voice signal;
[0117] The semantic features contained in the second speech data are outputted through the semantic recognition model.
[0118] Based on the above embodiment, optionally, the semantic recognition model is generated in the following manner:
[0119] Acquire training sample data for a semantic recognition training task, wherein the training sample data includes text data of the third speech data and semantic feature labels of the text data of the third speech data;
[0120] Inputting the training sample data into the semantic recognition model to perform semantic feature recognition prediction to obtain a semantic feature prediction result;
[0121] The cross entropy loss function is used to calculate the loss difference between the semantic feature prediction result and the semantic feature label, and the gradient of the parameters in the semantic recognition model is back-propagated according to the loss difference to update the parameters of the semantic recognition model.
[0122] Based on the above embodiment, optionally, determining a task processing resource matching the first voice interaction task from a plurality of task processing resources according to the first voiceprint feature and the voiceprint blacklist information includes:
[0123] If a third voiceprint feature matching the first voiceprint feature exists in the voiceprint blacklist information, determining a semantic feature associated with the third voiceprint feature according to the voiceprint blacklist information, where the third voiceprint feature is the second voiceprint feature that matches the first voiceprint feature and is selected from the voiceprint blacklist information;
[0124] According to the reference information of the third voiceprint feature, the task processing resource matching the first voice interaction task is determined from multiple task processing resources, the reference information of the third voiceprint feature includes the number of hits between the semantic feature category to which the semantic feature associated with the third voiceprint feature belongs and the third voiceprint feature, the number of hits of the third voiceprint feature can be used to indicate the number of times the third voiceprint feature in the voiceprint blacklist information is matched or accessed during the query operation, and the task processing resource matching the first voice interaction task is associated with the reference information associated with the third voiceprint feature.
[0125] Based on the above embodiment, optionally, determining a task processing resource matching the first voice interaction task from a plurality of task processing resources according to the reference information of the third voiceprint feature includes:
[0126] If it is determined according to the reference information of the third voiceprint feature that the hit number of the third voiceprint feature is greater than the preset hit number, then a task processing resource matching the first voice interaction task is determined from multiple task processing resources, and the task processing performance and / or task processing effect of the task processing resource matching the first voice interaction task when assisting in processing the first voice interaction task is greater than the task processing performance and / or task processing effect of the remaining task processing resources in the multiple task processing resources except the task processing resource matching the first voice interaction task;
[0127] According to the semantic feature category to which the semantic feature associated with the third voiceprint feature in the reference information associated with the third voiceprint feature belongs, a task processing resource that matches the semantic feature category to which the semantic feature associated with the third voiceprint feature belongs is selected from multiple task processing resources, and a task processing resource that matches the first voice interaction task is determined, and different task processing resources among at least some of the multiple task processing resources are associated with different semantic feature categories.
[0128] The voice response processing device provided in the embodiment of the present invention can execute the voice response processing method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of executing the voice response processing method.
[0129] It is worth noting that the various units and modules included in the above-mentioned device are only divided according to functional logic, but are not limited to the above-mentioned division, as long as the corresponding functions can be achieved; in addition, the specific names of the functional units are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of the embodiments of the present invention.
[0130] Figure 5 A schematic diagram of the structure of an electronic device 10 that can be used to implement an embodiment of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or claimed herein.
[0131] like Figure 5 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc., which is communicatively connected to the at least one processor 11. The memory stores a computer program that can be executed by the at least one processor. The processor 11 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 into the random access memory (RAM) 13. Various programs and data required for the operation of the electronic device 10 can also be stored in the RAM 13. The processor 11, ROM 12, and RAM 13 are connected to each other via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0132] Multiple components in the electronic device 10 are connected to the I / O interface 15, including an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a magnetic disk, an optical disk, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0133] The processor 11 can be any general-purpose and / or specialized processing component with processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The processor 11 executes the various methods and processes described above, such as the voice response processing method.
[0134] In particular, according to an embodiment of the present invention, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present invention includes a computer program product that includes a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via the communication unit 19, or installed from the storage unit 18, or installed from the ROM 12. When the computer program is executed by the processor 11, the above-mentioned functions defined in the method of the embodiment of the present invention are performed.
[0135] In some embodiments, the voice response processing method can be implemented as a computer program, which is tangibly contained in a computer-readable storage medium, such as the storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed on the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded into the RAM 13 and executed by the processor 11, one or more steps of the voice response processing method described above can be performed. Alternatively, in other embodiments, the processor 11 can be configured to execute the voice response processing method in any other appropriate manner (for example, by means of firmware).
[0136] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0137] Computer programs for implementing the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the computer program is executed by the processor, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The computer program may be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0138] In the context of the present invention, computer-readable storage media can be tangible media that can contain or store a computer program for use with an instruction execution system, device or equipment or used in combination with an instruction execution system, device or equipment. Computer-readable storage media can include but are not limited to electronic, magnetic, optical, electromagnetic, infrared or semiconductor systems, devices or equipment, or any suitable combination of the foregoing. Alternatively, computer-readable storage media can be machine-readable signal media. More specific examples of machine-readable storage media can include electrical connections based on one or more lines, portable computer disks, hard disks, random access memories (RAM), read-only memories (ROM), erasable programmable read-only memories (EPROM or flash memory), optical fibers, portable compact disk read-only memories (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0139] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0140] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.
[0141] A computing system may include clients and servers. The clients and servers are typically remote from each other and typically interact via a communication network. This client-server relationship arises through computer programs running on the respective computers, creating a client-server relationship. The server may be a cloud server, also known as a cloud computing server or cloud host. This server is a hosting product within the cloud computing service ecosystem that addresses the management difficulties and limited scalability of traditional physical hosting and VPS services.
[0142] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in the present invention can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present invention can be achieved. This is not limited herein.
[0143] The above specific embodiments do not limit the scope of protection of the present invention. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention are intended to be included within the scope of protection of the present invention.
Claims
1. A voice response processing method, characterized in that: The method comprises: Determining a first voiceprint feature of first voice data, where the first voice data includes voice data generated by performing a first voice interaction task during an interactive voice response phase; Determining voiceprint blacklist information, the voiceprint blacklist information including a plurality of second voiceprint features and semantic features associated with the second voiceprint features, where the second voiceprint features are voiceprint features of second voice data generated when performing a second voice interaction task, each second voice interaction task being a voice interaction task that has been performed before performing the first voice interaction task, and the semantic features associated with each second voiceprint feature are used to indicate appeal information conveyed through voice expression during the execution of the second voice interaction task and an emotional feature contained in the appeal information conveyed through voice expression; Based on the first voiceprint feature and the voiceprint blacklist information, a task processing resource matching the first voice interaction task is determined from multiple task processing resources to assist in processing the first voice interaction task that continues to be executed after the interactive voice response phase ends.
2. The method according to claim 1, characterized in that Determining a first voiceprint feature of the first voice data includes: Acquire first voice data generated when the first voice interaction task is executed; The first voiceprint feature is extracted from the first speech data by using a feature extraction method of Mel-frequency cepstral coefficients.
3. The method according to claim 1, characterized in that The voiceprint blacklist information is generated in the following way: Acquiring second voice data and performing semantic recognition on the second voice data to obtain semantic features contained in the second voice data, where the semantic features contained in the second voice data include appeal information conveyed through voice expression during the execution of the second voice interaction task and emotional features contained in the appeal information conveyed through voice expression; extracting a second voiceprint feature from the second voice data when the semantic feature contained in the second voice data matches the semantic feature category associated with the voiceprint blacklist information; The semantic feature contained in the second voice data is determined as the semantic feature associated with the second voiceprint feature, and is used to generate voiceprint blacklist information.
4. The method according to claim 3, characterized in that Performing semantic recognition on the second voice data to obtain semantic features contained in the second voice data includes: Determining text data of the second voice data, and inputting the text data of the second voice data into a semantic recognition model for semantic recognition; the semantic recognition model is a multi-task neural network model for parsing and recognizing semantic information expressed by the voice signal; The semantic features contained in the second speech data are outputted through the semantic recognition model.
5. The method according to claim 4, characterized in that The semantic recognition model is generated in the following way: Acquire training sample data for a semantic recognition training task, wherein the training sample data includes text data of the third speech data and semantic feature labels of the text data of the third speech data; Inputting the training sample data into the semantic recognition model to perform semantic feature recognition prediction to obtain a semantic feature prediction result; The cross entropy loss function is used to calculate the loss difference between the semantic feature prediction result and the semantic feature label, and the gradient of the parameters in the semantic recognition model is back-propagated according to the loss difference to update the parameters of the semantic recognition model.
6. The method according to claim 1, characterized in that Determining, based on the first voiceprint feature and the voiceprint blacklist information, a task processing resource that matches the first voice interaction task from a plurality of task processing resources, including: If a third voiceprint feature matching the first voiceprint feature exists in the voiceprint blacklist information, determining a semantic feature associated with the third voiceprint feature according to the voiceprint blacklist information, where the third voiceprint feature is the second voiceprint feature that matches the first voiceprint feature and is selected from the voiceprint blacklist information; According to the reference information of the third voiceprint feature, the task processing resource matching the first voice interaction task is determined from multiple task processing resources, the reference information of the third voiceprint feature includes the number of hits between the semantic feature category to which the semantic feature associated with the third voiceprint feature belongs and the third voiceprint feature, the number of hits of the third voiceprint feature can be used to indicate the number of times the third voiceprint feature in the voiceprint blacklist information is matched or accessed during the query operation, and the task processing resource matching the first voice interaction task is associated with the reference information associated with the third voiceprint feature.
7. The method according to claim 6, characterized in that Determining, based on the reference information of the third voiceprint feature, a task processing resource that matches the first voice interaction task from a plurality of task processing resources, including: If it is determined according to the reference information of the third voiceprint feature that the hit number of the third voiceprint feature is greater than the preset hit number, then a task processing resource matching the first voice interaction task is determined from multiple task processing resources, and the task processing performance and / or task processing effect of the task processing resource matching the first voice interaction task when assisting in processing the first voice interaction task is greater than the task processing performance and / or task processing effect of the remaining task processing resources in the multiple task processing resources except the task processing resource matching the first voice interaction task; According to the semantic feature category to which the semantic feature associated with the third voiceprint feature in the reference information associated with the third voiceprint feature belongs, a task processing resource that matches the semantic feature category to which the semantic feature associated with the third voiceprint feature belongs is selected from multiple task processing resources, and a task processing resource that matches the first voice interaction task is determined, and different task processing resources among at least some of the multiple task processing resources are associated with different semantic feature categories.
8. A voice response processing device, characterized in that: The device comprises: a determining module configured to determine a first voiceprint feature of first voice data, wherein the first voice data includes voice data generated by performing a first voice interaction task during an interactive voice response phase; The determination module is further configured to determine voiceprint blacklist information, the voiceprint blacklist information including a plurality of second voiceprint features and semantic features associated with the second voiceprint features, the second voiceprint features being voiceprint features of second voice data generated when performing a second voice interaction task, each second voice interaction task being a voice interaction task that has been performed before performing the first voice interaction task, and the semantic features associated with each second voiceprint feature being used to indicate appeal information conveyed through voice expression during the execution of the second voice interaction task and an emotional feature contained in the appeal information conveyed through voice expression; A processing module is used to determine a task processing resource that matches the first voice interaction task from multiple task processing resources based on the first voiceprint feature and the voiceprint blacklist information, and to assist in processing the first voice interaction task that continues to be executed after the interactive voice response stage ends.
9. An electronic device, characterized in that: The electronic device comprises: at least one processor; and a memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to perform the voice response processing method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the voice response processing method according to any one of claims 1 to 7 when executed.