Artificial customer service voice emotion optimization method and device
By performing noise reduction and emotion scoring on the voice signals of human customer service representatives, and then converting them into optimized text to generate optimized voice output, the service quality issues caused by the emotional fluctuations of human customer service representatives are resolved, thereby improving the user experience.
Patent Information
- Application Number
- CN202511041950.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-28
- Publication Date
- 2025-11-04
AI Technical Summary
The quality of human customer service is significantly affected by individual emotional states, leading to problems such as harsh tone and lack of warmth in wording, which negatively impacts user experience.
The speech signal is collected and noise is removed. An emotion score is obtained using an LSTM speech emotion recognition model. When the emotion score exceeds a threshold, the speech signal is converted into text for optimization. Finally, TTS speech synthesis is used to generate the optimized speech output.
It enables real-time monitoring and adjustment of customer service staff's emotions, improving service quality, reducing communication problems caused by emotional fluctuations, and enhancing user experience.
Smart Images

Figure FT_1 
Figure FT_2
Abstract
Description
Technical Field
[0001] This invention belongs to the field of computer technology, specifically relating to a method and apparatus for optimizing the emotional expression of human customer service voice. Background Technology
[0002] In the context of the rapid development of digital services, human customer service remains irreplaceable as the core link for deep interaction between businesses and customers. Although AI-powered intelligent customer service systems have become the primary service providers due to their standardized responses and efficient processing capabilities, human customer service has unique advantages in conveying empathy, providing emotional communication, and flexibly resolving complex business issues. For example, in scenarios such as handling complaints in the financial industry and after-sales disputes in e-commerce, users not only need problem-solving but also expect emotional reassurance and humanized care, which are precisely the key capabilities of human customer service that cannot be replaced by technology.
[0003] However, the service quality of human customer service is significantly affected by individual emotional states. High workloads and diverse customer complaint scenarios can easily lead to emotional fluctuations in customer service personnel, resulting in issues such as harsh tone and impersonal language. For example, when a customer service representative speaks faster due to busyness, it may be perceived as impatience by the user; using phrases like "that's how the system stipulates" can easily provoke user resentment. Therefore, it is crucial to utilize intelligent technology to monitor and adjust the emotional state of human customer service personnel in real time to reduce the impact of workload and customer complaint scenarios on service quality, avoid problems such as harsh tone and impersonal language caused by emotional fluctuations, and ultimately improve the user experience. Summary of the Invention
[0004] The purpose of this invention is to provide a method and apparatus for optimizing the voice emotion of human customer service, which solves the problem of poor user experience when users have in-depth interactions with telephone customer service.
[0005] To achieve the above objectives, the technical solution adopted by the present invention is as follows: This invention provides a method for optimizing the emotional tone of human customer service voice, comprising the following steps: The acquired speech signal is denoised to obtain the processed speech signal. Obtain the emotion score of the processed speech signal; The processed speech signal is then processed based on the emotion score, including: If the emotion score is greater than or equal to the preset threshold, the processed speech signal will be converted into text, the text content will be optimized, and the optimized text will be synthesized into speech after being confirmed and adjusted by customer service to generate optimized speech output. If the emotion score is less than the preset threshold, the processed voice signal will be output directly.
[0006] Preferably, the acquired speech signal is denoised using the ENC environmental noise reduction algorithm to obtain the processed speech signal.
[0007] Preferably, the emotion score of the processed speech signal is obtained, specifically through the following method: Construct an LSTM-based speech emotion recognition model; The processed speech signal is used as input to an LSTM-based speech emotion recognition model, and the output is an emotion score.
[0008] Preferably, the text content is optimized, specifically by: The text is optimized by replacing negative words in a pre-defined negative word replacement library.
[0009] Preferably, TTS speech synthesis is used to synthesize the optimized text into speech, generating optimized speech for direct output.
[0010] Secondly, the present invention provides a voice emotion optimization device for human customer service, comprising: The voice input module is used to collect voice signals and perform noise reduction processing to obtain the processed voice signal; The sentiment analysis module is used to obtain the sentiment score of the processed speech signal; The semantic polishing module is used to process the processed speech signal based on the emotion score, wherein: If the emotion score is greater than or equal to the preset threshold, the processed speech signal will be converted into text, the text content will be optimized, and the optimized text will be output to the speech synthesis module after being confirmed by customer service. If the emotion score is less than the preset threshold, the processed speech signal will be output to the speech synthesis module. The speech synthesis module is used to synthesize the text or processed speech signal after customer service confirmation to generate optimized speech output.
[0011] Thirdly, the present invention provides an electronic device including a processor and a memory, wherein the memory stores computer instructions, and when the computer instructions are executed by the processor, the electronic device performs the method described thereon.
[0012] Fourthly, the present invention provides a computing device cluster, comprising at least one computing device, each computing device including a processor and a memory; The processor of the at least one computing device is used to execute instructions stored in the memory of the at least one computing device to cause the cluster of computing devices to perform the method.
[0013] Fifthly, the present invention provides a computer program product, the computer program product including computer-executable instructions, which, when executed, implement the method described.
[0014] In a sixth aspect, the present invention provides a computer-readable storage medium storing computer-executable instructions that, when executed by a processor, implement the method described herein.
[0015] Compared with the prior art, the beneficial effects of the present invention are: This invention provides a method for optimizing the emotional tone of customer service voices. By learning from the voice data of customer service personnel during their service process, it can promptly identify normal and emotional voice data, analyze tone, speaking speed, and negative vocabulary in real time, output an emotion score, and refine the voice data based on the emotion score to produce optimized voice output, which greatly assists customer service personnel in providing professional services.
[0016] This invention provides a voice emotion optimization device for human customer service. The device is small and portable, easily building a bridge for emotion optimization between customer service personnel and computers / telephones. It processes the responses of customer service personnel with emotion, reducing the risk of customer service quality decline and work performance being evaluated due to occasional emotional changes. It reflects the humanistic care of technology for people and further improves customer service quality.
[0017] In summary, this invention can promptly identify and eliminate emotional content in customer service voice messages, transforming the original customer service voice messages into a rational, gentle, and empathetic tone and semantic approach before outputting them to the user, thereby improving the quality of customer service voice services and reducing communication problems caused by emotional language. Attached Figure Description
[0018] Figure 1 This is a structural diagram of the device of the present invention; Figure 2 This is a schematic diagram illustrating the method of using the device of the present invention. Detailed Implementation
[0019] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods have been omitted so as not to obscure the description of this application with unnecessary detail.
[0020] It should be understood that, when used in this application specification and the appended claims, the term "comprising" indicates the presence of the described features, integrals, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or a collection thereof.
[0021] It should also be understood that the term “and / or” as used in this application specification and the appended claims means any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.
[0022] As used in this application specification and the appended claims, the term "if" may be interpreted, depending on the context, as "when," "once," "in response to determination," or "in response to detection." Similarly, the phrase "if determined" or "if detected [the described condition or event]" may be interpreted, depending on the context, as meaning "once determined," "in response to determination," "once detected [the described condition or event]," or "in response to detection [the described condition or event]."
[0023] Furthermore, in the description of this application and the appended claims, the terms "first," "second," "third," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.
[0024] References to "one embodiment" or "some embodiments" as described in this specification mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.
[0025] Example 1 This embodiment provides a method for optimizing the emotional tone of human customer service voice, including the following steps: The acquired speech signal is denoised to obtain the processed speech signal. Obtain the emotion score of the processed speech signal; The processed speech signal is then processed based on the emotion score, including: If the emotion score is greater than or equal to the preset threshold, the processed speech signal is converted into text, the text content is optimized, the optimized text is then confirmed by customer service, and then speech synthesis is performed to generate optimized speech output. If the emotion score is less than the preset threshold, the processed voice signal will be output directly.
[0026] Example 2 Based on Example 1, this example provides a method for optimizing the emotional expression of human customer service voice. The method uses the ENC environmental noise reduction algorithm to denoise the acquired voice signal and obtain the processed voice signal.
[0027] Example 3 Based on Example 1, this example provides a method for optimizing the emotion of human customer service voice, used to obtain an emotion score for the processed voice signal. The specific method is as follows: Construct an LSTM-based speech emotion recognition model; The processed speech signal is used as input to an LSTM-based speech emotion recognition model, and the output is an emotion score.
[0028] Example 4 This embodiment provides a voice emotion optimization device for human customer service, comprising: The voice input module is used to collect voice signals and perform noise reduction processing to obtain the processed voice signal; The sentiment analysis module is used to obtain the sentiment score of the processed speech signal; The semantic polishing module is used to process the processed speech signal based on the emotion score, wherein: If the emotion score is greater than or equal to the preset threshold, the processed speech signal will be converted into text, the text content will be optimized, and the optimized text will be output to the speech synthesis module after being confirmed by customer service. If the emotion score is less than the preset threshold, the processed speech signal will be output to the speech synthesis module. The speech synthesis module is used to synthesize the text or processed speech signal after customer service confirmation to generate optimized speech output.
[0029] Example 5 Figure 1 This is a schematic diagram of the architecture of a voice emotion optimization device for human customer service, which is applicable to environments requiring telephone customer service, such as e-commerce, system operation and maintenance, and financial services.
[0030] This embodiment provides a voice emotion optimization device for human customer service, which includes a voice input module 1, an emotion analysis module 2, a semantic polishing module 3, a voice synthesis module 4, a text monitoring module 5, a voice output module 6, a Wi-Fi / Bluetooth adapter module 7, and a local storage module 8, wherein: The voice input module 1 includes a microphone interface and a noise reduction unit, wherein: The microphone interface is compatible with a standard 3.5mm audio interface or a USB interface, and is used to collect voice signals from customer service personnel.
[0031] The noise reduction unit incorporates the ENC environmental noise reduction algorithm to filter background noise and obtain the processed speech signal.
[0032] The emotion analysis module 2 is used to obtain an emotion score by taking the processed speech signal as input and outputting it based on the LSTM-based speech emotion recognition model.
[0033] The semantic polishing module 3 is used to process the processed speech signal according to the emotion score, wherein: If the emotion score is greater than or equal to the preset threshold, the processed speech signal is converted into text, the text content is optimized, and the optimized text is output to the speech synthesis module after being confirmed by customer service.
[0034] The speech synthesis module 4 is used to synthesize optimized text or processed speech signals into speech using TTS speech synthesis, and output the optimized speech.
[0035] The text monitoring module 5 is a monitoring module with a 4-6 inch display screen. It receives voice data from the speech synthesis module 4 via built-in Wi-Fi or Bluetooth, converts it into text information, and displays it on the screen for customer service to manually confirm and adjust the converted text information.
[0036] The voice output module 6 includes a headphone jack and a real-time monitoring button; the headphone jack connects to the customer service headset and outputs optimized voice; the real-time monitoring button allows the customer service representative to manually switch between the original voice and the optimized voice, facilitating intervention in emergency situations.
[0037] The Wi-Fi / Bluetooth adapter module 7 is used to connect the voice editing module 3 and the text monitoring module 5.
[0038] The local storage module 8 is used to store historical dialogue data and optimization logs.
[0039] Example 6 Based on Example 5, this example provides a voice emotion optimization device for human customer service, wherein the emotion analysis module includes: The real-time feature extractor is used to analyze the processed speech signal in real time and extract prosodic features (physical features) such as pitch, speech rate and energy value, while linking them with the emotional vocabulary of the text (semantic features) to form a hybrid feature vector of "speech prosody + text semantics". Voice emotion recognition model: The hybrid feature vector is used as input to an LSTM or CNN neural network to output an emotion score.
[0040] Example 7 Based on Example 5, this example provides a voice emotion optimization device for human customer service, wherein the semantic polishing module includes: Negative word replacement library: used to store preset awkward words and their corresponding positive expressions. In this embodiment, the awkward word is "cannot be processed", and the corresponding positive expression is "we will actively coordinate". Sentence optimization engine: The speech-to-text engine, which integrates BERT or GPT models, converts the processed speech signal into text information. Then, it replaces and optimizes negative words in the text information to generate empathetic sentences, thus obtaining the optimized text.
[0041] Example 8 Based on Example 5, this example provides a voice emotion optimization device for human customer service, wherein the voice synthesis unit includes: Emotional TTS module: used to generate an emotional parameter set based on emotion scores, the emotional parameter set including intonation parameters and speech parameters.
[0042] Speech waveform generator: Used to convert text into synthesized speech that matches the original speech timbre based on a set of emotion parameters and a preset original timbre model.
[0043] In this embodiment, the optimized text and emotion parameter set are used as input to the speech waveform generator. The text is converted into a speech waveform using a vocoder (such as WaveNet or Parallel WaveGAN). At the same time, based on a preset original timbre model, timbre transfer is performed on the generated waveform to ensure that the synthesized speech matches the original timbre of the customer service representative (for example, preserving the representative's dialect accent or vocal characteristics) and avoids a robotic feel.
[0044] Example 9 This embodiment describes a method for using a voice emotion optimization device for human customer service, including: (1) Device installation: Connect the device in series between the customer service microphone and the telephone system, and turn on the power; (2) Device configuration: After the device is started, set the default voice style of the device, configure the Wi-Fi connection of the device, and update the algorithm library and device firmware; (3) Voice acquisition and preprocessing: Environmental noise is filtered through the noise reduction unit to generate a clean voice signal; (4) Sentiment analysis and text generation: speech is converted into text, text and speech features are analyzed, and sentiment scores and polished text are generated; (5) Optimize voice output: Generate empathetic voice through the TTS module and output it to the user terminal.
[0045] Example 10 This embodiment also provides a computing device. The computing device includes a bus, a processor, a memory, and a communication interface. The processor, memory, and communication interface communicate with each other via the bus. The computing device can be a server or a terminal device. It should be understood that this application does not limit the number of processors and memory in the computing device.
[0046] A bus can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of representation, a bus can include a path for transmitting information between various components of a computing device (e.g., memory, processor, communication interfaces).
[0047] The processor may include any one or more of the following: central processing unit (CPU), graphics processing unit (GPU), tensor processing unit (TPU), application specific integrated circuit (ASIC), field-programmable gate array (FPGA), microprocessor (MP), or digital signal processor (DSP).
[0048] The memory may include volatile memory, such as random access memory (RAM). The processor may also include non-volatile memory, such as read-only memory (ROM), flash memory, hard disk drive (HDD), or solid state drive (SSD).
[0049] The memory stores executable program code, which the processor executes to implement the functions of the aforementioned modules, thereby achieving, for example, the method described in Embodiment 1. That is, the memory may store instructions for the methods and functions relating to the computing device in any of the above embodiments.
[0050] The communication interface uses transceiver modules such as, but not limited to, network interface cards and transceivers to enable communication between computing devices and other devices or communication networks.
[0051] Example 11 This embodiment also provides a computing device cluster. The computing device cluster includes at least one computing device. The computing device can be a server, such as a central server, an edge server, or a local server in a local data center. In some embodiments, the computing device can also be a terminal device such as a desktop computer, a laptop computer, or a smartphone.
[0052] The computing device cluster includes at least one computing device. The memory of one or more computing devices in the computing device cluster may store the same instructions for performing the methods and functions related to the computing devices in any of the above embodiments.
[0053] In some possible implementations, the memory of one or more computing devices in the computing device cluster may also store partial instructions for performing the methods and functions of the computing devices involved in any of the above embodiments. In other words, a combination of one or more computing devices can jointly execute the instructions for performing the methods and functions of the computing devices.
[0054] It should be noted that the memory in different computing devices within a computing device cluster can store different instructions, which are used to execute parts of the device's functions.
[0055] In some possible implementations, one or more computing devices in a computing device cluster can be connected via a network. This network can be a wide area network (WAN) or a local area network (LAN), etc. Two computing devices are connected to each other via the network. Specifically, they connect to the network through communication interfaces in each computing device.
[0056] Embodiments of this disclosure also provide a computer program product containing instructions that, when run on a computer, cause the computer to perform the methods and functions related to a computing device in any of the above embodiments.
[0057] Example 12 This embodiment also provides a computer-readable storage medium storing computer instructions that, when executed by a processor, cause the processor to perform the methods and functions of the computing device involved in any of the above embodiments.
[0058] Generally, the various embodiments of this disclosure can be implemented in hardware or dedicated circuitry, software, logic, or any combination thereof. Some aspects can be implemented in hardware, while others can be implemented in firmware or software, which can be executed by a controller, microprocessor, or other computing device. Although various aspects of the embodiments of this disclosure are shown and described as block diagrams, flowcharts, or represented using some other illustration, it should be understood that the blocks, apparatuses, systems, techniques, or methods described herein can be implemented as, as non-limiting examples, in hardware, software, firmware, dedicated circuitry or logic, general-purpose hardware or controllers or other computing devices, or some combination thereof.
[0059] Example 13 This embodiment provides at least one computer program product tangibly stored on a non-transitory computer-readable storage medium. The computer program product includes computer-executable instructions, such as instructions included in program modules, which execute in a device on a target real or virtual processor to perform the processes / methods as described above with reference to the accompanying drawings. Typically, program modules include routines, programs, libraries, objects, classes, components, data structures, etc., that perform specific tasks or implement specific abstract data types. In various embodiments, the functionality of program modules can be combined or divided among program modules as needed. The machine-executable instructions for the program modules can execute within a local or distributed device. In a distributed device, the program modules can reside in both local and remote storage media.
[0060] Computer program code used to implement the methods of this disclosure may be written in one or more programming languages. This computer program code may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus, such that when executed by the computer or other programmable data processing apparatus, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be performed. The program code may be executed entirely on a computer, partially on a computer, as a stand-alone software package, partially on a computer and partially on a remote computer, or entirely on a remote computer or server.
[0061] In the context of this disclosure, computer program code or related data may be carried on any suitable carrier to enable a device, apparatus, or processor to perform the various processes and operations described above. Examples of carriers include signals, computer-readable media, and so on. Examples of signals may include electrical, optical, radio, sound, or other forms of propagation signals, such as carrier waves, infrared signals, etc.
[0062] Computer-readable media can be any tangible medium that contains or stores programs for or relating to an instruction execution system, apparatus, or device, or a data storage device such as a data center containing one or more available media. Computer-readable media can be computer-readable signal media or computer-readable storage media. Computer-readable media can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. More detailed examples of computer-readable storage media include electrical connections with one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0063] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.
Claims
1. A method for optimizing the emotional tone of human customer service voice, characterized in that, Includes the following steps: The acquired speech signal is denoised to obtain the processed speech signal. Obtain the emotion score of the processed speech signal; The processed speech signal is then processed based on the emotion score, including: If the emotion score is greater than or equal to the preset threshold, the processed speech signal will be converted into text, the text content will be optimized, and the optimized text will be synthesized into speech after being confirmed and adjusted by customer service to generate optimized speech output. If the emotion score is less than the preset threshold, the processed voice signal will be output directly.
2. The method for optimizing the emotional tone of human customer service voice according to claim 1, characterized in that, The acquired speech signal is denoised using the ENC environmental noise reduction algorithm to obtain the processed speech signal.
3. The method for optimizing the emotional tone of human customer service voice according to claim 1, characterized in that, The specific method for obtaining the emotion score of the processed speech signal is as follows: Construct an LSTM-based speech emotion recognition model; The processed speech signal is used as input to an LSTM-based speech emotion recognition model, and the output is an emotion score.
4. The method for optimizing the emotional tone of human customer service voice according to claim 1, characterized in that, The text content can be optimized in the following ways: The text is optimized by replacing negative words in a pre-defined negative word replacement library.
5. The method for optimizing the emotional tone of human customer service voice according to claim 1, characterized in that, The optimized text is synthesized using TTS speech synthesis to generate optimized speech for direct output.
6. A voice emotion optimization device for human customer service representatives, characterized in that, include: The voice input module is used to collect voice signals and perform noise reduction processing to obtain the processed voice signal; The sentiment analysis module is used to obtain the sentiment score of the processed speech signal; The semantic polishing module is used to process the processed speech signal based on the emotion score, wherein: If the emotion score is greater than or equal to the preset threshold, the processed speech signal will be converted into text, the text content will be optimized, and the optimized text will be output to the speech synthesis module after being confirmed by customer service. If the emotion score is less than the preset threshold, the processed speech signal will be output to the speech synthesis module. The speech synthesis module is used to synthesize the text or processed speech signal after customer service confirmation to generate optimized speech output.
7. An electronic device, characterized in that, It includes a processor and a memory, the memory storing computer instructions that, when executed by the processor, cause the electronic device to perform the method of any one of claims 1 to 5.
8. A computing device cluster, characterized in that, It includes at least one computing device, each computing device including a processor and memory; The processor of the at least one computing device is configured to execute instructions stored in the memory of the at least one computing device to cause the cluster of computing devices to perform the method according to any one of claims 1 to 5.
9. A computer program product, characterized in that, The computer program product includes computer-executable instructions that, when executed, implement the method according to any one of claims 1 to 5.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions that, when executed by a processor, implement the method according to any one of claims 1 to 5.
Citation Information
Cited By
Voice interaction method and device, electronic equipment and computer program product
CN121459794A