Voice emotion recognition method and device, computer equipment and storage medium

By collecting and cacheing voice data during the call between customer service and customers, using target voice activity detection and feature fusion technology to generate accurate emotion recognition results, the problem of insufficient accuracy and timeliness of voice emotion recognition in the prior art is solved, and a more efficient voice emotion recognition effect is achieved.

CN120472943APending Publication Date: 2025-08-12PING AN TECH (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510664686.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-22
Publication Date
2025-08-12

AI Technical Summary

Technical Problem

The existing voice emotion recognition technology has low accuracy and timeliness in customer service complaint handling scenarios in the financial field. It is difficult to achieve high-precision and low-latency real-time responses when complex voice signals and emotions change dynamically, resulting in high misjudgment rates and affecting service quality and customer satisfaction.

Method used

By collecting voice data during the call between customer service and customers and buffering it to the buffer, the target voice activity detection algorithm is used to split the voice fragments, and the speech feature vector and steady-state emotion vector are extracted by combining the preset extraction model. After the feature fusion is performed, the emotion recognition model is used for emotion recognition to generate accurate emotion recognition results.

Benefits of technology

It realizes faster and more accurate emotion recognition, improves the accuracy and timeliness of voice emotion recognition, optimizes the accuracy and timeliness of customer service response, and improves service experience and operation efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120472943A_ABST
    Figure CN120472943A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of artificial intelligence, and relates to a voice emotion recognition method and device, computer equipment and a storage medium, and the method comprises the steps: collecting the voice data of a customer in a call process between a customer service and the customer; caching the voice data into a buffer area; segmenting the voice data in the buffer area based on a target voice activity detection algorithm to obtain voice segments; performing feature extraction on the voice segments based on the extraction model to obtain voice feature vectors; obtaining customer call data in a call process, and extracting a steady-state emotion vector from the customer call data; performing feature fusion on the voice feature vector and the steady-state emotion vector to obtain a target feature vector; and performing emotion recognition on the target feature vector based on an emotion recognition model to generate an emotion recognition result of the voice segment. In addition, the emotion recognition result can be stored in the block chain. The method can be applied to voice emotion recognition scenes in the financial field, and the accuracy and timeliness of emotion recognition are effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence technology and can be applied to the field of financial technology, and in particular to speech emotion recognition methods, devices, computer equipment and storage media. Background Art

[0002] In the traditional customer service model, speech emotion recognition mainly relies on traditional machine learning models or basic deep learning architectures (such as CNN, RNN), which classify emotions by extracting acoustic features of speech (such as fundamental frequency, energy, MFCC) or simple time series patterns. Although such methods can achieve basic emotion recognition in standardized scenarios, they have significant limitations when processing complex speech signals and dynamic changes in emotions. Specifically, traditional models lack the ability to deeply analyze the semantics of speech context, and are limited by computational efficiency, making it difficult to meet both high precision and low latency requirements in real-time call scenarios. This technical defect results in a high misjudgment rate in customer emotion monitoring in existing systems, low accuracy and timeliness of emotion recognition, and even missed opportunities for intervention due to processing delays, which directly affects service quality and customer satisfaction.

[0003] For example, in customer service complaint handling scenarios in the financial sector, traditional systems might simply identify "anger" based on a customer's rising voice pitch or faster speech rate, without conducting a comprehensive analysis of the interaction history. If the customer's actual request is for technical assistance rather than emotional outburst, this misjudgment could trigger unnecessary manual intervention, resulting in wasted service resources. Conversely, if customer sentiment continues to deteriorate due to unresolved issues, delayed system responses could exacerbate the risk of customer churn. These scenarios highlight the limited adaptability of existing technologies to complex business scenarios.

[0004] Therefore, there is an urgent need for an efficient and accurate voice emotion recognition technology to improve the accuracy and timeliness of customer service responses, thereby optimizing service experience and operational efficiency. Summary of the Invention

[0005] The purpose of the embodiments of the present application is to propose a speech emotion recognition method, apparatus, computer equipment and storage medium to solve the technical problems of low accuracy and timeliness of existing speech emotion recognition methods.

[0006] In a first aspect, a method for speech emotion recognition is provided, comprising:

[0007] During a call between a customer service representative and a customer, voice data of the customer is collected;

[0008] Cache the voice data into a preset buffer;

[0009] Segmenting the voice data in the buffer based on a target voice activity detection algorithm to obtain corresponding voice segments; wherein the target voice activity detection algorithm is an algorithm obtained by adjusting a threshold of a preset initial voice activity detection algorithm;

[0010] Performing feature extraction on the speech segment based on a preset extraction model to obtain a speech feature vector corresponding to the speech segment;

[0011] Acquiring customer call data corresponding to the call process, and extracting a steady-state emotion vector corresponding to the customer from the customer call data;

[0012] Performing feature fusion on the speech feature vector and the steady-state emotion vector to obtain a corresponding target feature vector;

[0013] Emotion recognition is performed on the target feature vector based on a preset emotion recognition model to generate an emotion recognition result corresponding to the speech segment.

[0014] In a second aspect, a speech emotion recognition device is provided, comprising:

[0015] The collection module is used to collect the voice data of the customer during the conversation between the customer service and the customer;

[0016] A cache module, configured to cache the voice data into a preset buffer;

[0017] a segmentation module configured to segment the voice data in the buffer based on a target voice activity detection algorithm to obtain corresponding voice segments; wherein the target voice activity detection algorithm is an algorithm obtained by adjusting a threshold of a preset initial voice activity detection algorithm;

[0018] An extraction module, configured to perform feature extraction on the speech segment based on a preset extraction model to obtain a speech feature vector corresponding to the speech segment;

[0019] A first acquisition module is configured to acquire customer call data corresponding to the call process, and extract a steady-state emotion vector corresponding to the customer from the customer call data;

[0020] A fusion module, configured to perform feature fusion on the speech feature vector and the steady-state emotion vector to obtain a corresponding target feature vector;

[0021] The recognition module is used to perform emotion recognition on the target feature vector based on a preset emotion recognition model to generate an emotion recognition result corresponding to the speech segment.

[0022] In a third aspect, a computer device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the above-mentioned speech emotion recognition method when executing the computer program.

[0023] In a fourth aspect, a computer-readable storage medium is provided, which stores a computer program. When the computer program is executed by a processor, the steps of the above-mentioned speech emotion recognition method are implemented.

[0024] In the scheme implemented by the above-mentioned voice emotion recognition method, device, computer equipment and storage medium, first, during the conversation between the customer service and the customer, the customer's voice data is collected; and the voice data is cached in a preset buffer; then the voice data in the buffer is segmented based on the target voice activity detection algorithm to obtain corresponding voice segments; wherein the target voice activity detection algorithm is an algorithm obtained by threshold adjustment of the preset initial voice activity detection algorithm; then, feature extraction is performed on the voice segment based on a preset extraction model to obtain a voice feature vector corresponding to the voice segment; and customer call data corresponding to the call process is obtained, and a steady-state emotion vector corresponding to the customer is extracted from the customer call data; subsequently, feature fusion is performed on the voice feature vector and the steady-state emotion vector to obtain the corresponding target feature vector; finally, emotion recognition is performed on the target feature vector based on a preset emotion recognition model to generate an emotion recognition result corresponding to the voice segment. This application collects the customer's voice data and caches it in a preset buffer during the call between the customer service and the customer, then segments the voice data in the buffer to obtain voice segments based on the use of a target voice activity detection algorithm, and extracts features from the voice segments based on the use of an extraction model to obtain voice feature vectors of the voice segments, then obtains the customer call data corresponding to the call process, extracts the customer's steady-state emotion vector from the customer call data, subsequently performs feature fusion on the voice feature vector and the steady-state emotion vector to obtain a target feature vector, and finally performs emotion recognition on the target feature vector based on the use of an emotion recognition model to generate an emotion recognition result corresponding to the voice segment. This application segments and extracts features from the customer's voice data based on the use of a target voice activity detection algorithm and an extraction model to obtain corresponding voice segments, and extracts the customer's steady-state emotion vector from the current call process, and subsequently performs feature fusion on the voice feature vector and the steady-state emotion vector to obtain a target feature vector, and then performs emotion recognition on the target feature vector based on the use of an emotion recognition model. This feature fusion method can more comprehensively reflect the customer's emotional state, thereby enabling faster and more accurate generation of the customer's emotion recognition results, effectively improving the accuracy and timeliness of voice emotion recognition. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] In order to more clearly illustrate the solutions in this application, a brief introduction will be given below to the drawings required for use in the description of the embodiments of this application. Obviously, the drawings described below are some embodiments of this application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0026] Figure 1 is an exemplary system architecture diagram to which the present application may be applied;

[0027] Figure 2 is a flow chart of an embodiment of a method for speech emotion recognition according to the present application;

[0028] Figure 3 is a structural diagram of an embodiment of a speech emotion recognition device according to the present application;

[0029] Figure 4 It is a structural diagram of an embodiment of a computer device according to the present application. DETAILED DESCRIPTION

[0030] Unless otherwise defined, all technical and scientific terms used herein have the same meanings as commonly understood by those skilled in the art to which this application belongs. The terms used in the specification of the application are for the purpose of describing specific embodiments only and are not intended to limit this application. The terms "including" and "having" and any variations thereof in the specification and claims of this application and the above-mentioned drawings are intended to cover non-exclusive inclusions. The terms "first", "second", etc. in the specification and claims of this application or the above-mentioned drawings are used to distinguish different objects, not to describe a specific order.

[0031] References herein to "embodiments" mean that a particular feature, structure, or characteristic described in connection with the embodiments may be included in at least one embodiment of the present application. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor does it constitute an independent or alternative embodiment that is mutually exclusive of other embodiments. It is understood, both explicitly and implicitly, by those skilled in the art that the embodiments described herein may be combined with other embodiments.

[0032] In order to enable those skilled in the art to better understand the solution of the present application, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the accompanying drawings.

[0033] like Figure 1As shown, system architecture 100 may include a terminal device 101, a network 102, and a server 103. Terminal device 101 may be a laptop computer 1011, a tablet computer 1012, or a mobile phone 1013. Network 102 is a medium for providing a communication link between terminal device 101 and server 103. Network 102 may include various connection types, such as wired or wireless communication links or fiber optic cables.

[0034] The user can use the terminal device 101 to interact with the server 103 via the network 102 to receive or send messages, etc. Various communication client applications can be installed on the terminal device 101, such as web browser applications, shopping applications, search applications, instant messaging tools, email clients, social platform software, etc.

[0035] The terminal device 101 can be various electronic devices with a display screen and supporting web browsing. In addition to the laptop computer 1011, tablet computer 1012 or mobile phone 1013, the terminal device 101 can also be an e-book reader, an MP3 player (Moving Picture Experts Group Audio Layer III), an MP4 (Moving Picture Experts Group Audio Layer IV) player, a laptop computer and a desktop computer, etc.

[0036] The server 103 may be a server that provides various services, such as a background server that provides support for web pages displayed on the terminal device 101 .

[0037] It should be noted that the speech emotion recognition method provided in the embodiments of the present application is generally executed by a server / terminal device, and accordingly, the speech emotion recognition device is generally set in the server / terminal device.

[0038] It should be understood that Figure 1 The number of terminal devices, networks and servers in the embodiment is merely illustrative. Any number of terminal devices, networks and servers may be provided as required.

[0039] Continue to refer Figure 2, shows a flow chart of an embodiment of the speech emotion recognition method according to the present application. According to different needs, the order of the steps in the flow chart can be changed, and some steps can be omitted. The speech emotion recognition method provided in the embodiment of the present application can be applied to any scenario where speech emotion recognition is required, and the speech emotion recognition method can be applied to products in these scenarios, for example, speech emotion recognition in the financial and insurance fields. The speech emotion recognition method comprises the following steps:

[0040] Step S201: During a conversation between a customer service representative and a customer, voice data of the customer is collected.

[0041] In this embodiment, the electronic device on which the speech emotion recognition method is executed (eg Figure 1 The server / terminal device shown in the figure) can obtain the customer's voice data through a wired connection or a wireless connection. It should be noted that the above-mentioned wireless connection methods may include but are not limited to 3G / 4G / 5G connection, Wi-Fi connection, Bluetooth connection, WiMAX connection, Zigbee connection, UWB (ultra wideband) connection, and other wireless connection methods currently known or to be developed in the future. The execution subject of this application may specifically be a customer service system, which may be referred to as the system for short. This application can be applied to customer service business scenarios or voice quality inspection and analysis scenarios in the financial services field. Among them, during the conversation between customer service and customers, the customer's voice data is collected in real time, and the clarity and integrity of the voice signal are ensured. Specifically, a voice acquisition card or software interface is deployed in the voice communication link of the customer service system to capture the customer's voice signal in real time, and perform preliminary processing on the collected voice signal, such as noise reduction and gain adjustment, to ensure the clarity and integrity of the voice signal, and then convert the analog voice signal into a digital format (such as PCM encoding) to obtain voice data for subsequent processing.

[0042] Step S202: Cache the voice data into a preset buffer.

[0043] In this embodiment, the buffer can be a ring buffer in memory for temporarily storing the real-time collected voice data. The buffer size can be pre-set (e.g., 5 seconds of voice data) to ensure data continuity during voice processing delays.

[0044] Step S203 , segmenting the voice data in the buffer based on a target voice activity detection algorithm to obtain corresponding voice segments; wherein the target voice activity detection algorithm is an algorithm obtained by adjusting a threshold of a preset initial voice activity detection algorithm.

[0045] In this embodiment, the construction process of the target voice activity detection algorithm is described in further detail in subsequent specific embodiments of this application and is not elaborated on here. The target voice activity detection algorithm can be used to segment the voice data in the buffer, detecting energy changes and silence segments in the voice signal to segment the recording into independent voice segments, ensuring that each voice segment corresponds to a relatively complete emotional expression unit.

[0046] Step S204 : performing feature extraction on the speech segment based on a preset extraction model to obtain a speech feature vector corresponding to the speech segment.

[0047] In this embodiment, the extraction model can specifically employ a pre-trained WavLM model. The WavLM model can be used to extract features from the speech segment to obtain a speech feature vector for the segment. The WavLM model is a deep learning model based on the Transformer architecture that can capture subtle changes in speech, providing a rich feature foundation for subsequent sentiment analysis.

[0048] Step S205 , obtaining customer call data corresponding to the call process, and extracting a steady-state emotion vector corresponding to the customer from the customer call data.

[0049] In this embodiment, since the customer's emotions are generally stable during a call, all designated recorded voice segments included in the call, such as the first few voice segments of the current call, are obtained as the customer's call data. Feature extraction and cluster analysis are then performed on the customer's call data to obtain a steady-state emotion vector corresponding to the customer. Feature extraction can be performed on the customer's call data using the extraction model to obtain the corresponding designated feature vector. A clustering method such as the K-Means algorithm or the DBSCAN algorithm is then used to select the class with the most members as the customer's steady-state speech feature set, and the central feature of the steady-state speech feature set is calculated (e.g., the average value) to serve as the customer's steady-state emotion vector.

[0050] Step S206 , performing feature fusion on the speech feature vector and the steady-state emotion vector to obtain a corresponding target feature vector.

[0051] In this embodiment, the specific implementation process of the above-mentioned feature fusion of the speech feature vector and the steady-state emotion vector to obtain the corresponding target feature vector will be further described in detail in subsequent specific embodiments of this application and will not be elaborated on here.

[0052] Step S207 , performing emotion recognition on the target feature vector based on a preset emotion recognition model to generate an emotion recognition result corresponding to the speech segment.

[0053] In this embodiment, real-time emotion recognition can be performed by inputting the target feature vector into the emotion recognition model. The emotion recognition model processes the target feature vector and outputs a corresponding emotion category as the emotion recognition result for the speech segment. The construction process of the emotion recognition model will be further described in detail in subsequent specific embodiments of this application and will not be elaborated on here.

[0054] For example, a customer at a financial institution's customer service center inquires about the voice emotion recognition process involved in the insurance claims process. The system collects the customer's voice data in real time and temporarily stores it in a buffer. When the customer asks, "What documents are needed for a claim?" the target voice activity detection algorithm detects and segments the voice segment. The extraction model then extracts the feature vector of the voice segment. The customer's steady-state emotion vector (assuming it's "calm") is also extracted from the current call. The real-time feature vector and the steady-state emotion vector are then concatenated and fed into the emotion recognition model, which outputs the emotion category "calm." The system then reports the "calm" emotion to the customer service representative and logs it. Subsequent analysis shows that the customer remained emotionally stable throughout the call and expressed satisfaction with the representative's response.

[0055] This application first collects the customer's voice data during the conversation between the customer service and the customer; and caches the voice data in a preset buffer; then segments the voice data in the buffer based on a target voice activity detection algorithm to obtain corresponding voice segments; wherein the target voice activity detection algorithm is an algorithm obtained by threshold adjustment of a preset initial voice activity detection algorithm; then, based on a preset extraction model, feature extraction is performed on the voice segment to obtain a voice feature vector corresponding to the voice segment; and customer call data corresponding to the call process is obtained, and a steady-state emotion vector corresponding to the customer is extracted from the customer call data; subsequently, feature fusion is performed on the voice feature vector and the steady-state emotion vector to obtain a corresponding target feature vector; finally, emotion recognition is performed on the target feature vector based on a preset emotion recognition model to generate an emotion recognition result corresponding to the voice segment. This application collects the customer's voice data and caches it in a preset buffer during the call between the customer service and the customer, then segments the voice data in the buffer to obtain voice segments based on the use of a target voice activity detection algorithm, and extracts features from the voice segments based on the use of an extraction model to obtain voice feature vectors of the voice segments, then obtains the customer call data corresponding to the call process, extracts the customer's steady-state emotion vector from the customer call data, subsequently performs feature fusion on the voice feature vector and the steady-state emotion vector to obtain a target feature vector, and finally performs emotion recognition on the target feature vector based on the use of an emotion recognition model to generate an emotion recognition result corresponding to the voice segment. This application segments and extracts features from the customer's voice data based on the use of a target voice activity detection algorithm and an extraction model to obtain corresponding voice segments, and extracts the customer's steady-state emotion vector from the current call process, and subsequently performs feature fusion on the voice feature vector and the steady-state emotion vector to obtain a target feature vector, and then performs emotion recognition on the target feature vector based on the use of an emotion recognition model. This feature fusion method can more comprehensively reflect the customer's emotional state, thereby enabling faster and more accurate generation of the customer's emotion recognition results, effectively improving the accuracy and timeliness of voice emotion recognition.

[0056] In some optional implementations, before step S202, the electronic device may further perform the following steps:

[0057] Get real-time network status information.

[0058] In this embodiment, the following network status indicators can be collected in real time as corresponding network status information. The network status indicators may include at least: round-trip time (RTT), which reflects network latency; packet loss rate, which directly reflects the degree of congestion; and available bandwidth, which dynamically evaluates the current available bandwidth.

[0059] Get real-time voice processing latency information.

[0060] In this embodiment, the following information can be collected in real time as the corresponding voice processing delay information. The voice processing delay information may include at least: encoding / decoding delay: the processing time of the voice codec; jitter buffer delay: the delay introduced by the current buffer.

[0061] Invoke the preset buffer adjustment strategy.

[0062] In this embodiment, the processing logic of the above-mentioned buffer adjustment strategy includes: 1. Initial buffer size setting. Set the default value (such as 50-100ms of voice data volume) based on historical data or initial detection (such as the network status in the previous few seconds). 2. Dynamic adjustment logic. Congestion detection: When the packet loss rate increases or the RTT increases significantly, it is determined to be network congestion. Adjustment action: Gradually increase the buffer (such as increasing the capacity by 20% each time) until the packet loss rate is lower than the threshold (such as 1%). Smooth detection: When a low packet loss rate (such as <0.1%) is continuously monitored and the RTT is stable, it is determined to be a smooth network. Adjustment action: Gradually reduce the buffer (such as reducing the capacity by 15% each time) until the minimum safety threshold (such as 30ms data volume) is reached. Jitter processing: If a sudden jitter is detected (such as a sudden surge in RTT), temporarily expand the buffer to absorb the jitter, and then slowly recover. 3. Smooth transition mechanism. Incremental adjustment: Avoid sudden changes that cause sound quality problems. The adjustment range is limited to ±20% each time. Delay constraints: Set upper and lower limits for the buffer (e.g., 20-200ms) to prevent extreme situations from affecting real-time performance.

[0063] The buffer adjustment strategy is used to adjust the capacity of the buffer accordingly according to the network status information and the voice processing delay information.

[0064] In this embodiment, the buffer capacity can be adjusted accordingly according to the processing strategy corresponding to the above-mentioned buffer adjustment strategy and the above-mentioned real-time network status information and voice processing delay information to generate a final adjusted capacity value of the buffer. The capacity of the buffer is then dynamically adjusted according to the final adjusted capacity value, for example, by expanding / shrinking the ring buffer or adjusting the read / write pointer. Among them, the buffer can also be optimized and fault-tolerant, including: Adaptive threshold: Dynamically adjust the congestion / unblocked judgment threshold based on historical data. Abnormal fallback: If the packet loss rate increases after adjustment, quickly fall back to the previous state. Multi-objective trade-off: Prioritize stability in a weak network environment (allowing higher latency), and prioritize low latency in a high-quality network.

[0065] This application obtains real-time network status information and real-time voice processing delay information, then invokes a preset buffer adjustment strategy. Subsequently, the buffer adjustment strategy is used to adjust the buffer capacity accordingly based on the network status and voice processing delay information. This application can automatically and intelligently adjust the buffer capacity dynamically based on real-time network status and voice processing delay, effectively improving the system's robustness to network fluctuations and ensuring the continuity and real-time nature of voice data.

[0066] In some optional implementations of this embodiment, before step S203, the electronic device may further perform the following steps:

[0067] The energy and the zero-crossing rate of the speech data in a preset sliding window are monitored, and a current initial noise level is generated based on the energy and the zero-crossing rate.

[0068] In this embodiment, during a call between a customer service representative and a client, the system continuously monitors the energy and zero-crossing rate of the client's voice data. Specifically, a sliding window can be used to calculate the energy and zero-crossing rate within the current sliding window in real time. A shorter sliding window (e.g., 100 milliseconds) can be set based on actual business needs to calculate the current initial noise level in real time based on the energy and zero-crossing rate within the sliding window. The current initial noise level includes the initial noise energy level and the initial noise zero-crossing rate level.

[0069] The current initial noise level is smoothed based on a preset smoothing algorithm to obtain a corresponding current noise energy level and a current noise zero-crossing rate level.

[0070] In this embodiment, the smoothing algorithm may employ an exponential smoothing algorithm or other time series smoothing methods. The current initial noise level may be smoothed according to the selected smoothing algorithm to avoid frequent threshold adjustments due to transient noise fluctuations, thereby obtaining the corresponding current noise energy level and current noise zero-crossing rate level. Specifically, the current noise level may be expressed as: current noise level = α × current measurement value + (1-α) × historical noise level. Where α is a smoothing factor (e.g., 0.1) used to control the impact of the current measurement value on the historical noise level.

[0071] The initial voice activity detection algorithm is invoked.

[0072] In this embodiment, the above-mentioned initial voice activity detection algorithm is specifically a voice activation detection (VAD) algorithm, whose purpose is to detect whether the current voice signal contains a voice signal, that is, to judge the input signal, distinguish the voice signal from various background noise signals, and adopt different processing methods for the two signals.

[0073] An energy threshold of the initial voice activity detection algorithm is adjusted based on the current noise energy level, and a zero-crossing rate threshold of the initial voice activity detection algorithm is adjusted based on the current noise zero-crossing rate level to obtain an adjusted voice activity detection algorithm.

[0074] In this embodiment, the energy threshold of the initial voice activity detection algorithm can be dynamically adjusted based on the current noise energy level. For example, the energy threshold can be set as a multiple of the noise energy baseline (e.g., 3 times). Furthermore, the zero-crossing rate threshold of the initial voice activity detection algorithm can be dynamically adjusted based on the current noise zero-crossing rate level. For example, the zero-crossing rate threshold can be set as a multiple of the noise zero-crossing rate baseline (e.g., 2 times).

[0075] Among them, the process of noise estimation and initial baseline calculation includes: Silence segment detection: At the beginning of a call, the system first detects the silence segment (i.e., the period when the customer is not speaking). The energy level of the voice signal can be detected. When the energy is lower than a certain initial threshold, it is determined to be a silence segment. Noise baseline calculation: During the silence segment, the system calculates the energy baseline and zero-crossing rate baseline of the voice signal. The energy baseline can be obtained by calculating the root mean square (RMS) value of the silence segment signal, and the zero-crossing rate baseline can be obtained by counting the number of times the signal crosses the zero point. In addition, the calculated initial noise baseline (noise energy baseline and noise zero-crossing rate baseline) can be stored in the memory as a reference point for subsequent dynamic adjustment.

[0076] The adjusted voice activity detection algorithm is used as the target voice activity detection algorithm.

[0077] In this embodiment, the dynamically adjusted energy threshold and zero-crossing rate threshold in the adjusted voice activity detection algorithm can be used to detect activity in the current voice signal. If the signal energy exceeds the adjusted energy threshold or the zero-crossing rate exceeds the adjusted zero-crossing rate threshold, voice activity is determined. The detection results of the target voice activity detection algorithm can be verified, for example, by checking whether the detection results of multiple consecutive windows are consistent. If the detection results are inconsistent, threshold fine-tuning or re-estimation of noise can be triggered.

[0078] Additionally, a minimum and maximum threshold range can be pre-set to avoid overly aggressive or conservative threshold adjustments. For example, the minimum and maximum energy thresholds can be set to 2 and 5 times the noise energy baseline, respectively. Threshold updates can be triggered periodically (e.g., every second) or based on the magnitude of the noise level change (e.g., if the noise level changes by more than 10%).

[0079] The present application monitors the energy and zero-crossing rate of the voice data within a preset sliding window and generates a current initial noise level based on the energy and the zero-crossing rate; then smoothes the current initial noise level based on a preset smoothing algorithm to obtain a corresponding current noise energy level and current noise zero-crossing rate level; then calls the initial voice activity detection algorithm; and adjusts the energy threshold of the initial voice activity detection algorithm based on the current noise energy level and the zero-crossing rate threshold of the initial voice activity detection algorithm based on the current noise zero-crossing rate level to obtain an adjusted voice activity detection algorithm; and subsequently uses the adjusted voice activity detection algorithm as the target voice activity detection algorithm. The present application monitors the energy and zero-crossing rate of speech data within a preset sliding window, generates a current initial noise level based on the obtained energy and zero-crossing rate, then smoothes the current initial noise level based on a preset smoothing algorithm, and intelligently and dynamically adjusts the threshold of the initial voice activity detection algorithm based on the obtained current noise energy level and current noise zero-crossing rate level, so as to efficiently and accurately construct the required target voice activity detection algorithm. Subsequently, by using the target voice activity detection algorithm, high-precision voice activity detection can be maintained in different noise environments, thereby facilitating improved accuracy of subsequent emotion recognition.

[0080] In some optional implementations, step S206 includes the following steps:

[0081] Get the preset stitching strategy.

[0082] In this embodiment, the concatenation strategy includes concatenation and dimensionality matching. Specifically, the concatenation operation involves concatenating the real-time extracted speech feature vector with the steady-state emotion vector to form a fused feature vector. Dimension matching involves ensuring the dimensionality of the two vectors is compatible, performing padding or truncation as necessary.

[0083] The speech feature vector and the steady-state emotion vector are spliced based on the splicing strategy to obtain a corresponding first feature vector.

[0084] In this embodiment, the splicing process of the speech feature vector and the steady-state emotion vector may be performed according to the strategy content of the splicing strategy, thereby obtaining a spliced first feature vector.

[0085] The first eigenvector is normalized to obtain a corresponding second eigenvector.

[0086] In this embodiment, the first eigenvector may be normalized by Z-score normalization or Min-Max normalization, and normalized using a pre-calculated mean and standard deviation to ensure consistency of feature scales, thereby obtaining a corresponding second eigenvector.

[0087] The second feature vector is used as the target feature vector.

[0088] This application obtains a preset splicing strategy; then splices the speech feature vector and the steady-state emotion vector based on the splicing strategy to obtain a corresponding first feature vector; then normalizes the first feature vector to obtain a corresponding second feature vector; and subsequently uses the second feature vector as the target feature vector. This application obtains a first feature vector by splicing the speech feature vector and the steady-state emotion vector based on the use of a splicing strategy, and then normalizes the first feature vector, thereby achieving efficient and accurate fusion processing of the feature fusion of the speech feature vector and the steady-state emotion vector, ensuring the accuracy and standardization of the generated target feature vector, and the generated target feature vector can provide more comprehensive information for emotion recognition, which is beneficial for the subsequent emotion recognition processing of the generated target feature vector based on the use of an emotion recognition model, thereby achieving more accurate recognition of customer emotions.

[0089] In some optional implementations, before step S207, the electronic device may further perform the following steps:

[0090] Acquire pre-collected historical speech data, and perform emotion annotation on the historical speech data to obtain corresponding speech annotation data.

[0091] In this embodiment, the historical voice data collection process includes collecting customer voice recordings over a historical time period. This recording covers the complete voice performance of the customer during a call, ensuring data integrity and continuity to fully reflect the customer's emotional state. The customer voice recordings are then segmented into independent voice segments based on a target voice activity detection algorithm. All independent voice segments are then integrated to generate the historical voice data. The historical time period is not specifically defined and can be set based on actual business needs, for example, within the past three years.

[0092] Furthermore, the aforementioned emotion annotation methods include those based on manual annotation or existing emotion classification standards to ensure accuracy and consistency. Annotated emotion categories may include "happy," "angry," "sad," "calm," and so on. The aforementioned historical speech data can be emotion-annotated using the aforementioned emotion annotation methods to obtain corresponding speech annotated data.

[0093] The speech annotation data is subjected to feature fusion processing to obtain corresponding speech sample data.

[0094] In this embodiment, the aforementioned extraction model is used to extract features from the annotated speech data to obtain a corresponding feature vector. A steady-state emotion vector of the corresponding customer that matches the feature vector is then obtained. The feature vector is then concatenated with the steady-state emotion vector of the corresponding customer to obtain the corresponding speech sample data.

[0095] The speech sample data is divided into a training set and a test set.

[0096] In this embodiment, the voice sample data may be divided into a training set and a test set according to a preset division ratio. There is no specific limitation on the numerical value of the division ratio, which may be set to 7:3, for example.

[0097] Call the preset deep learning model.

[0098] In this embodiment, the above-mentioned deep learning model can specifically adopt a deep learning model based on a Transformer structure, that is, the deep learning model based on a Transformer structure is used as a judgment model for emotion recognition.

[0099] Based on a preset optimization training strategy, the deep learning model is trained using the training set to obtain a corresponding first generation model.

[0100] In this embodiment, the model training can be performed by using the above-mentioned training set as the input of the above-mentioned deep learning model. By adjusting the hyperparameters of the deep learning model (such as learning rate, number of layers, number of hidden units, etc.) and optimizing the training strategy (such as using Adam optimizer, learning rate scheduling, etc.), the performance of the deep learning model can be further improved to construct a corresponding first generation model.

[0101] The first generation model is evaluated and optimized based on the test set to obtain a second generation model that meets the preset construction requirements.

[0102] In this embodiment, the trained first generative model can be evaluated using a test set to calculate metrics such as precision, recall, and F1 score. Based on the evaluation results, the model's hyperparameters can be further optimized to improve model performance, ultimately yielding a second generative model with performance that meets the requirements. If the model performance is still unsatisfactory, consider introducing more training data or trying other deep learning models.

[0103] The second generative model is used as the emotion recognition model.

[0104] This application obtains pre-collected historical voice data and performs emotion annotation on the historical voice data to obtain corresponding voice annotation data; then performs feature fusion processing on the voice annotation data to obtain corresponding voice sample data; and divides the voice sample data into a training set and a test set; then calls a preset deep learning model; subsequently, based on a preset optimization training strategy, uses the training set to train the deep learning model to obtain a corresponding first generation model; further evaluates and optimizes the first generation model based on the test set to obtain a second generation model that meets the preset construction requirements; and finally, uses the second generation model as the emotion recognition model. This application obtains pre-collected historical speech data, and performs emotion annotation on the historical speech data to obtain speech annotation data, then performs feature fusion processing on the speech annotation data to obtain speech sample data, and divides the speech sample data into a training set and a test set. Then, based on the optimization training strategy, the training set is used to train the deep learning model to obtain a first generation model, and then the first generation model is evaluated and optimized based on the use of the test set to obtain a second generation model that meets the construction requirements and serves as the final emotion recognition model, thereby achieving efficient and accurate completion of the model construction processing of the emotion recognition model, improving the construction efficiency of the emotion recognition model, and effectively improving the emotion recognition performance of the emotion recognition model.

[0105] In some optional implementations of this embodiment, after the step of using the second generated model as the emotion recognition model, the electronic device may further perform the following steps:

[0106] Call the preset monitoring tool.

[0107] In this embodiment, the monitoring tool is an automated tool capable of real-time monitoring of model performance. There is no specific limitation on the selection of the monitoring tool, which can be determined based on actual business needs.

[0108] The performance of the emotion recognition model is monitored based on the monitoring tool to obtain corresponding performance monitoring data.

[0109] In this embodiment, key performance indicators (KPIs) are pre-set, such as accuracy, recall, F1 score, inference latency, etc. Then, the performance of the emotion recognition model is monitored using the monitoring tool to collect performance monitoring data corresponding to the KPIs in real time.

[0110] Get the preset model optimization strategy.

[0111] In this embodiment, the above-mentioned model optimization strategy includes the following: Data augmentation: If it is found that the recognition effect of certain emotion categories is poor, relevant training data is added (such as more samples of the emotion "anger"). Model tuning: Adjust model hyperparameters (such as learning rate and batch size), or try different model structures (such as using a deeper Transformer). A / B testing: A / B testing is performed on the new and old models to compare their performance in real-world environments and select the optimal model for deployment.

[0112] Through the model optimization strategy, the emotion recognition model is subjected to corresponding model tuning processing according to the performance monitoring data.

[0113] In this embodiment, the emotion recognition model may be tuned accordingly using the strategy content of the model optimization strategy based on the obtained performance monitoring data.

[0114] This application calls a preset monitoring tool; then monitors the performance of the emotion recognition model based on the monitoring tool to obtain corresponding performance monitoring data; then obtains a preset model optimization strategy; and subsequently uses the model optimization strategy to perform corresponding model tuning processing on the emotion recognition model according to the performance monitoring data. After completing the construction of the emotion recognition model, this application will also automatically monitor the performance of the emotion recognition model based on the use of the monitoring tool to obtain performance monitoring data, and then intelligently perform corresponding model tuning processing on the emotion recognition model based on the performance monitoring data based on the use of the model optimization strategy, thereby being able to continuously optimize the performance of the emotion recognition model, which is beneficial to improving customer service quality and providing strong decision support for managers.

[0115] In some optional implementations of this embodiment, after step S207, the electronic device may further perform the following steps:

[0116] Obtain the call text and customer service operation records corresponding to the call process.

[0117] In this embodiment, the call text refers to the textual content of the call recorded during the call between the customer service representative and the customer. The customer service operation record refers to the key operations and call transfer records triggered by the customer service representative during the call between the customer service representative and the customer.

[0118] The emotion recognition result, the call text, and the customer service operation record are integrated to generate a voice quality inspection report corresponding to the voice data.

[0119] In this embodiment, the above-mentioned content integration processing includes data association, key event annotation, comprehensive analysis, and report generation. Among them, data association includes: associating emotion recognition results with call texts and customer service operation records. Key event annotation includes: marking key points of emotional changes in call texts (such as clips where customers suddenly become angry). Comprehensive analysis includes: combining emotion recognition results and customer service operations to analyze whether the customer service staff's response strategies are effective, and generate the desired operation analysis results. Report generation includes: designing a unified quality inspection report template, which includes key event analysis, comprehensive analysis and other parts. By filling the emotion recognition results, call texts, customer service operation records, generated key event annotation data and operation analysis results into the quality inspection report template, a corresponding voice quality inspection report is generated.

[0120] Call the preset push tool.

[0121] In this embodiment, there is no specific limitation on the selection of the above-mentioned push tool. For example, an email tool or an internal message system, etc. can be used.

[0122] Based on the push tool, the voice quality inspection report is sent to relevant management personnel.

[0123] In this embodiment, the generated voice quality inspection report can be automatically pushed to the mailbox or workstation of the relevant manager according to the selected push tool, so as to be used to evaluate the emotional management ability of the customer service staff.

[0124] This application obtains the call text and customer service operation records corresponding to the call process; then integrates the emotion recognition results, the call text and the customer service operation records to generate a voice quality inspection report corresponding to the voice data; then calls a preset push tool; and subsequently sends the voice quality inspection report to relevant managers based on the push tool. This application obtains the call text and customer service operation records corresponding to the call process, and automatically integrates the emotion recognition results, the call text and the customer service operation records to generate a voice quality inspection report corresponding to the voice data, thereby improving the generation efficiency and intelligence of the voice quality inspection report. In addition, the generated voice quality inspection report will be automatically sent to the relevant managers based on the use of the push tool, thereby providing strong decision-making support for managers to evaluate the customer service's emotional management capabilities, thereby improving the managers' work efficiency and user experience.

[0125] In some optional implementations, the user information obtained is obtained with the user's consent and complies with relevant laws and policies.

[0126] In addition, any software tools or components not provided by our company that appear in the embodiments of this application are merely examples and do not represent actual use.

[0127] Furthermore, this application extracts steady-state emotion feature vectors from the current call and utilizes a deep learning model based on the Transformer architecture for emotion recognition, enabling more accurate analysis of customer emotions. This enables comprehensive emotion analysis of customer service calls after the fact, providing data support for customer service staff training and performance evaluation. Furthermore, by combining emotion recognition with voice quality control, this method can more comprehensively assess the service quality of customer service staff, optimize customer service processes, reduce customer complaints, enhance the competitiveness of financial companies in the financial services sector, and provide strong support for business development.

[0128] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0129] It should be emphasized that in order to further ensure the privacy and security of the above-mentioned emotion recognition results, the above-mentioned emotion recognition results can also be stored in a node of a blockchain.

[0130] The blockchain referred to in this application refers to a new application model for computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanisms, and encryption algorithms. Blockchain is essentially a decentralized database, a series of data blocks generated using cryptographic methods. Each data block contains information about a batch of network transactions, which is used to verify the validity of the information (to prevent counterfeiting) and generate the next block. Blockchain can include the blockchain underlying platform, the platform product service layer, and the application service layer.

[0131] The embodiments of the present application can acquire and process relevant data based on artificial intelligence technology. Artificial Intelligence (AI) is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to achieve optimal results.

[0132] Those skilled in the art will appreciate that all or part of the processes in the above-described method embodiments can be implemented by instructing related hardware via computer-readable instructions. The computer-readable instructions can be stored in a computer-readable storage medium, and when the program is executed, it can include the processes in the above-described method embodiments. The aforementioned storage medium can be a non-volatile storage medium such as a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM).

[0133] It should be understood that although the steps in the flowcharts of the accompanying drawings are shown in sequence as indicated by the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some of the steps in the flowcharts of the accompanying drawings may include multiple sub-steps or multiple stages, and these sub-steps or stages are not necessarily executed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be executed in turn or alternately with other steps or at least a portion of the sub-steps or stages of other steps.

[0134] Further references Figure 3 , as a response to the above Figure 2 The present application provides an embodiment of a speech emotion recognition device, which is similar to the embodiment of the present invention. Figure 2 Corresponding to the method embodiment shown, the device can be specifically applied to various electronic devices.

[0135] like Figure 3As shown, the speech emotion recognition device 300 of this embodiment includes: a collection module 301, a buffer module 302, a segmentation module 303, an extraction module 304, a first acquisition module 305, a fusion module 306 and a recognition module 307. Among them:

[0136] The collection module 301 is used to collect the voice data of the customer during the conversation between the customer service and the customer;

[0137] The buffer module 302 is used to buffer the voice data into a preset buffer;

[0138] a segmentation module 303 configured to segment the voice data in the buffer based on a target voice activity detection algorithm to obtain corresponding voice segments; wherein the target voice activity detection algorithm is an algorithm obtained by adjusting a threshold of a preset initial voice activity detection algorithm;

[0139] An extraction module 304 is configured to perform feature extraction on the speech segment based on a preset extraction model to obtain a speech feature vector corresponding to the speech segment;

[0140] A first acquisition module 305 is configured to acquire customer call data corresponding to the call process, and extract a steady-state emotion vector corresponding to the customer from the customer call data;

[0141] A fusion module 306 is configured to perform feature fusion on the speech feature vector and the steady-state emotion vector to obtain a corresponding target feature vector;

[0142] The recognition module 307 is configured to perform emotion recognition on the target feature vector based on a preset emotion recognition model, and generate an emotion recognition result corresponding to the speech segment.

[0143] In this embodiment, the operations performed by the above modules or units correspond one-to-one to the steps of the voice emotion recognition method in the aforementioned embodiment, and are not described in detail here.

[0144] In some optional implementations of this embodiment, the speech emotion recognition apparatus further includes:

[0145] The second acquisition module is used to obtain real-time network status information;

[0146] A third acquisition module is used to obtain real-time voice processing delay information;

[0147] A first calling module is used to call a preset buffer adjustment strategy;

[0148] The first adjustment module is configured to use the buffer adjustment strategy to adjust the capacity of the buffer accordingly according to the network status information and the voice processing delay information.

[0149] In some optional implementations of this embodiment, the speech emotion recognition device further includes:

[0150] a monitoring module, configured to monitor the energy and zero-crossing rate of the speech data within a preset sliding window, and generate a current initial noise level based on the energy and the zero-crossing rate;

[0151] a smoothing module, configured to smooth the current initial noise level based on a preset smoothing algorithm to obtain a corresponding current noise energy level and a current noise zero-crossing rate level;

[0152] A second calling module, configured to call the initial voice activity detection algorithm;

[0153] a second adjustment module, configured to adjust an energy threshold of the initial voice activity detection algorithm based on the current noise energy level, and adjust a zero-crossing rate threshold of the initial voice activity detection algorithm based on the current noise zero-crossing rate level, to obtain an adjusted voice activity detection algorithm;

[0154] The first determining module is configured to use the adjusted voice activity detection algorithm as the target voice activity detection algorithm.

[0155] In some optional implementations of this embodiment, the fusion module 306 includes:

[0156] Get submodule, used to obtain the preset splicing strategy;

[0157] a splicing submodule, configured to splice the speech feature vector and the steady-state emotion vector based on the splicing strategy to obtain a corresponding first feature vector;

[0158] a processing submodule, configured to perform normalization processing on the first eigenvector to obtain a corresponding second eigenvector;

[0159] A determination submodule is configured to use the second feature vector as the target feature vector.

[0160] In some optional implementations of this embodiment, the speech emotion recognition apparatus further includes:

[0161] A first processing module is used to obtain pre-collected historical speech data and perform emotion annotation on the historical speech data to obtain corresponding speech annotation data;

[0162] A second processing module is used to perform feature fusion processing on the speech annotation data to obtain corresponding speech sample data;

[0163] A division module, configured to divide the speech sample data into a training set and a test set;

[0164] The third calling module is used to call the preset deep learning model;

[0165] A training module, configured to train the deep learning model using the training set based on a preset optimization training strategy to obtain a corresponding first generation model;

[0166] An optimization module, configured to evaluate and optimize the first generation model based on the test set to obtain a second generation model that meets preset construction requirements;

[0167] The second determining module is configured to use the second generation model as the emotion recognition model.

[0168] In some optional implementations of this embodiment, the speech emotion recognition apparatus further includes:

[0169] The fourth calling module is used to call the preset monitoring tool;

[0170] A monitoring module, configured to monitor the performance of the emotion recognition model based on the monitoring tool and obtain corresponding performance monitoring data;

[0171] The fourth acquisition module is used to obtain a preset model optimization strategy;

[0172] A tuning module is used to perform corresponding model tuning processing on the emotion recognition model according to the performance monitoring data through the model optimization strategy.

[0173] In some optional implementations of this embodiment, the speech emotion recognition apparatus further includes:

[0174] A fifth acquisition module is used to obtain the call text and customer service operation records corresponding to the call process;

[0175] A generation module, configured to integrate the emotion recognition result, the call text, and the customer service operation record to generate a voice quality inspection report corresponding to the voice data;

[0176] The fifth calling module is used to call the preset push tool;

[0177] The sending module is used to send the voice quality inspection report to relevant management personnel based on the push tool.

[0178] To solve the above technical problems, the present application also provides a computer device. Figure 4 , Figure 4 This is a basic structural block diagram of the computer device in this embodiment.

[0179] The computer device 4 includes a memory 41, a processor 42, and a network interface 43 that are interconnected through a system bus. It should be noted that the figure only shows a computer device 4 with components 41-43, but it should be understood that it is not required to implement all the components shown, and more or fewer components can be implemented instead. Among them, those skilled in the art can understand that the computer device here is a device that can automatically perform numerical calculations and / or information processing according to pre-set or stored instructions, and its hardware includes but is not limited to microprocessors, application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), embedded devices, etc.

[0180] The computer device may be a desktop computer, notebook computer, PDA, cloud server, etc. The computer device may interact with the user via a keyboard, mouse, remote control, touchpad, or voice control device.

[0181] The memory 41 includes at least one type of readable storage medium, including flash memory, hard disk, multimedia card, card-type memory (e.g., SD or DX memory), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the memory 41 can be an internal storage unit of the computer device 4, such as a hard disk or memory of the computer device 4. In other embodiments, the memory 41 can also be an external storage device of the computer device 4, such as a plug-in hard disk equipped on the computer device 4, a smart memory card (SMC), a secure digital (SD) card, a flash card, etc. Of course, the memory 41 can also include both the internal storage unit of the computer device 4 and its external storage device. In this embodiment, the memory 41 is generally used to store the operating system and various application software installed on the computer device 4, such as computer-readable instructions of the speech emotion recognition method. In addition, the memory 41 can also be used to temporarily store various types of data that have been output or are to be output.

[0182] In some embodiments, the processor 42 may be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chip. The processor 42 is generally used to control the overall operation of the computer device 4. In this embodiment, the processor 42 is used to execute computer-readable instructions or process data stored in the memory 41, such as computer-readable instructions for executing the speech emotion recognition method.

[0183] The network interface 43 may include a wireless network interface or a wired network interface. The network interface 43 is generally used to establish a communication connection between the computer device 4 and other electronic devices.

[0184] Compared with the prior art, the embodiments of the present application have the following beneficial effects:

[0185] In an embodiment of the present application, the customer's voice data is segmented and feature extracted based on the use of a target voice activity detection algorithm and an extraction model to obtain corresponding voice segments, and the customer's steady-state emotion vector is extracted from the current call process. Subsequently, the voice feature vector and the steady-state emotion vector are feature fused to obtain a target feature vector, and then emotion recognition is performed on the target feature vector based on the use of an emotion recognition model. This feature fusion method can more comprehensively reflect the customer's emotional state, thereby enabling faster and more accurate generation of the customer's emotion recognition results, effectively improving the accuracy and timeliness of voice emotion recognition.

[0186] The present application also provides another embodiment, namely, providing a computer-readable storage medium, which stores computer-readable instructions, and the computer-readable instructions can be executed by at least one processor to enable the at least one processor to perform the steps of the speech emotion recognition method as described above.

[0187] Compared with the prior art, the embodiments of the present application have the following beneficial effects:

[0188] In an embodiment of the present application, the customer's voice data is segmented and feature extracted based on the use of a target voice activity detection algorithm and an extraction model to obtain corresponding voice segments, and the customer's steady-state emotion vector is extracted from the current call process. Subsequently, the voice feature vector and the steady-state emotion vector are feature fused to obtain a target feature vector, and then emotion recognition is performed on the target feature vector based on the use of an emotion recognition model. This feature fusion method can more comprehensively reflect the customer's emotional state, thereby enabling faster and more accurate generation of the customer's emotion recognition results, effectively improving the accuracy and timeliness of voice emotion recognition.

[0189] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus the necessary general hardware platform, and of course can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for enabling a terminal device (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in each embodiment of the present application.

[0190] Obviously, the embodiments described above are only some of the embodiments of the present application, rather than all of the embodiments. The preferred embodiments of the present application are given in the accompanying drawings, but they do not limit the patent scope of the present application. The present application can be implemented in many different forms. On the contrary, the purpose of providing these embodiments is to make the understanding of the disclosure of the present application more thorough and comprehensive. Although the present application has been described in detail with reference to the aforementioned embodiments, for those skilled in the art, it is still possible to modify the technical solutions described in the aforementioned specific embodiments, or to make equivalent replacements for some of the technical features therein. Any equivalent structure made using the contents of the present application specification and the accompanying drawings, directly or indirectly used in other related technical fields, is also within the scope of patent protection of the present application.

Claims

1. A method for speech emotion recognition, characterized in that: The steps include: During a call between a customer service representative and a customer, voice data of the customer is collected; Cache the voice data into a preset buffer; Segmenting the voice data in the buffer based on a target voice activity detection algorithm to obtain corresponding voice segments; wherein the target voice activity detection algorithm is an algorithm obtained by adjusting a threshold of a preset initial voice activity detection algorithm; Performing feature extraction on the speech segment based on a preset extraction model to obtain a speech feature vector corresponding to the speech segment; Acquiring customer call data corresponding to the call process, and extracting a steady-state emotion vector corresponding to the customer from the customer call data; Performing feature fusion on the speech feature vector and the steady-state emotion vector to obtain a corresponding target feature vector; Emotion recognition is performed on the target feature vector based on a preset emotion recognition model to generate an emotion recognition result corresponding to the speech segment.

2. The speech emotion recognition method according to claim 1, wherein Before the step of caching the voice data in a preset buffer, the method further includes: Get real-time network status information; Get real-time voice processing delay information; Call the preset buffer adjustment strategy; The buffer adjustment strategy is used to adjust the capacity of the buffer accordingly according to the network status information and the voice processing delay information.

3. The speech emotion recognition method according to claim 1, wherein Before the step of segmenting the voice data in the buffer based on the target voice activity detection algorithm to obtain corresponding voice segments, the method further includes: monitoring the energy and the zero-crossing rate of the speech data within a preset sliding window, and generating a current initial noise level based on the energy and the zero-crossing rate; Smoothing the current initial noise level based on a preset smoothing algorithm to obtain a corresponding current noise energy level and a current noise zero-crossing rate level; Invoking the initial voice activity detection algorithm; adjusting an energy threshold of the initial voice activity detection algorithm based on the current noise energy level, and adjusting a zero-crossing rate threshold of the initial voice activity detection algorithm based on the current noise zero-crossing rate level, to obtain an adjusted voice activity detection algorithm; The adjusted voice activity detection algorithm is used as the target voice activity detection algorithm.

4. The speech emotion recognition method according to claim 1, wherein The step of performing feature fusion on the speech feature vector and the steady-state emotion vector to obtain a corresponding target feature vector specifically includes: Get the preset splicing strategy; splicing the speech feature vector and the steady-state emotion vector based on the splicing strategy to obtain a corresponding first feature vector; Normalizing the first eigenvector to obtain a corresponding second eigenvector; The second feature vector is used as the target feature vector.

5. The speech emotion recognition method according to claim 1, wherein: Before the step of performing emotion recognition on the target feature vector based on a preset emotion recognition model to generate an emotion recognition result corresponding to the speech segment, the method further includes: Acquire pre-collected historical speech data, and perform emotion annotation on the historical speech data to obtain corresponding speech annotation data; Performing feature fusion processing on the speech annotation data to obtain corresponding speech sample data; Dividing the speech sample data into a training set and a test set; Call the preset deep learning model; Based on a preset optimization training strategy, the deep learning model is trained using the training set to obtain a corresponding first generation model; Evaluate and optimize the first generation model based on the test set to obtain a second generation model that meets preset construction requirements; The second generative model is used as the emotion recognition model.

6. The speech emotion recognition method according to claim 5, characterized in that After the step of using the second generative model as the emotion recognition model, the method further includes: Call the preset monitoring tool; Performing performance monitoring on the emotion recognition model based on the monitoring tool to obtain corresponding performance monitoring data; Get the preset model optimization strategy; Through the model optimization strategy, the emotion recognition model is subjected to corresponding model tuning processing according to the performance monitoring data.

7. The speech emotion recognition method according to claim 1, wherein: After the step of performing emotion recognition on the target feature vector based on a preset emotion recognition model to generate an emotion recognition result corresponding to the speech segment, the method further includes: Obtain the call text and customer service operation records corresponding to the call process; Integrate the emotion recognition result, the call text, and the customer service operation record to generate a voice quality inspection report corresponding to the voice data; Call the preset push tool; Based on the push tool, the voice quality inspection report is sent to relevant management personnel.

8. A speech emotion recognition device, characterized in that: include: The collection module is used to collect the voice data of the customer during the conversation between the customer service and the customer; A cache module, configured to cache the voice data into a preset buffer; a segmentation module configured to segment the voice data in the buffer based on a target voice activity detection algorithm to obtain corresponding voice segments; wherein the target voice activity detection algorithm is an algorithm obtained by adjusting a threshold of a preset initial voice activity detection algorithm; An extraction module, configured to perform feature extraction on the speech segment based on a preset extraction model to obtain a speech feature vector corresponding to the speech segment; A first acquisition module is configured to acquire customer call data corresponding to the call process, and extract a steady-state emotion vector corresponding to the customer from the customer call data; A fusion module, configured to perform feature fusion on the speech feature vector and the steady-state emotion vector to obtain a corresponding target feature vector; The recognition module is used to perform emotion recognition on the target feature vector based on a preset emotion recognition model to generate an emotion recognition result corresponding to the speech segment.

9. A computer device, characterized in that: The method comprises a memory and a processor, wherein the memory stores computer-readable instructions, and the processor implements the steps of the speech emotion recognition method according to any one of claims 1 to 7 when executing the computer-readable instructions.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer-readable instructions, which, when executed by a processor, implement the steps of the speech emotion recognition method according to any one of claims 1 to 7.