Speech processing device, speech processing method, and, program

The voice processing system addresses the issue of unsuppressed speaker emotions by generating a synthesized voice with suppressed emotions, reducing listener stress and enabling appropriate responses.

JP2025106595AActive Publication Date: 2025-07-15SOFTBANK CORPORATION
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2025070071
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-04-22
Publication Date
2025-07-15
Estimated Expiration
2041-11-25

AI Technical Summary

Technical Problem

Existing voice conversion systems fail to sufficiently suppress the emotions of a speaker, leading to increased stress and inappropriate responses from listeners.

Method used

A voice processing system that includes an acquisition unit for speech signals, a voice recognition unit to generate text data, an emotion recognition unit to identify speaker emotions, and a voice synthesis unit to generate a synthesized voice with suppressed emotions, which is then output to reduce listener stress and facilitate appropriate responses.

Benefits of technology

The system effectively reduces listener stress and enables appropriate responses by generating a synthesized voice that sufficiently suppresses the emotions of the speaker, allowing the listener to recognize and respond to the speaker's emotions accurately.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025106595000001_ABST
    Figure 2025106595000001_ABST
Patent Text Reader

Abstract

To enable reduction of stress on a listener.SOLUTION: A speech processing system 1 comprises: an acquisition unit for acquiring a speech signal that is a signal of a first user's speech; a speech recognition unit that inputs feature quantities extracted based on the speech signal into a speech recognition model to generate text data including a word sequence consisting of one or more words; a speech synthesis unit that inputs feature quantities extracted based on the text data to generate a synthesized speech signal that is a signal of synthesized speech; and a speech output unit that outputs the synthesized speech signal to a second user.SELECTED DRAWING: Figure 4
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an audio processing system, an audio processing apparatus, and an audio processing method.

Background Art

[0002] Conventionally, various call centers have been operated in which operators respond by phone to customer complaints and the like in order to improve customer satisfaction (CS). In such customer service operations, "customer harassment," in which customers make intimidating remarks or unreasonable demands to operators, has been regarded as a problem because it can cause mental disorders among operators or increase the turnover rate of operators.

[0003] In recent years, voice conversion systems for protecting employees, namely operators, from such customer harassment have also been studied. For example, in Patent Document 1, the volume and pitch fluctuation amounts are calculated from an input voice signal, and when the volume and pitch fluctuation amounts exceed a predetermined value, the volume and pitch are controlled to be converted and output so that the volume and pitch fluctuation amounts are within a predetermined range.

[0004]

Prior Art Documents

Patent Documents

[0004]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0005] However, for example, simply converting the speech voice of a speaker by the method described in Patent Document 1In this case, the speaker's (first user's) emotions are not sufficiently suppressed, and the listener's (second user's) stress is increased. On the other hand, in order to reduce the stress of the listener, By converting the speaker's speech output to the hand, the listener can fully recognize the speaker's emotions. There is also a risk that the listener will not be able to respond appropriately.

[0006] Therefore, the present invention provides a method for sufficiently reducing the stress of a listener and / or for enabling the listener to respond appropriately. The present invention provides a voice processing system, a voice processing device, and a voice processing method that enable the above. [Means for solving the problem]

[0007] A voice processing system according to one embodiment of the present invention includes: An acquisition unit that acquires an utterance voice signal, and a feature amount extracted based on the utterance voice signal are extracted as a speech signal. A sound is input to a recognition model to generate text data that includes a word sequence of one or more words. A voice recognition unit and a feature quantity extracted based on the text data are input to a voice synthesis model. a speech synthesis unit for generating a synthetic speech signal, which is a signal of a synthetic speech, for a second user; and a voice output unit that outputs the synthetic voice.

[0008] According to this aspect, generating text data based on a speech signal of a first user, A synthetic voice generated based on the text data is output to the second user. The synthetic voice in which the emotion of the customer contained in the speech of the first user is sufficiently suppressed is sent to the second user. The second user's stress caused by the emotional utterance of the first user can be heard by the second user. can be sufficiently reduced.

[0009] In the above voice processing system, the emotion recognition unit uses, as input, the uttered voice signal, the feature quantity extracted from the uttered voice signal, the text data generated from the uttered voice signal, the feature quantity extracted from the text data, or a combination of at least two of these, and outputs the emotion information of the speaker of the uttered voice signal to an emotion recognition model that is machine-learned. By inputting the uttered voice signal acquired by the acquisition unit, the voice feature quantity extracted from the uttered voice signal, the text data generated from the uttered voice signal, the text feature quantity corresponding to the text data, or a combination of at least two of these, the emotion information of the first user corresponding to the uttered voice signal acquired by the acquisition unit may be generated. the feature quantity extracted therefrom, the text data generated from the uttered voice signal, the feature quantity extracted from the text data, or a combination of at least two of these, and outputs the emotion information of the speaker of the uttered voice signal to an emotion recognition model that is machine-learned. By inputting the uttered voice signal acquired by the acquisition unit, the voice feature quantity extracted from the uttered voice signal, the text data generated from the uttered voice signal, the text feature quantity corresponding to the text data, or a combination of at least two of these, the emotion information of the first user corresponding to the uttered voice signal acquired by the acquisition unit may be generated. the emotion information of the speaker of the uttered voice signal to an emotion recognition model that is machine-learned. By inputting the uttered voice signal acquired by the acquisition unit, the voice feature quantity extracted from the uttered voice signal, the text data generated from the uttered voice signal, the text feature quantity corresponding to the text data, or a combination of at least two of these, the emotion information of the first user corresponding to the uttered voice signal acquired by the acquisition unit may be generated. the text data generated from the uttered voice signal, the text feature quantity corresponding to the text data, or a combination of at least two of these, and outputs the emotion information of the speaker of the uttered voice signal to an emotion recognition model that is machine-learned. By inputting the uttered voice signal acquired by the acquisition unit, the voice feature quantity extracted from the uttered voice signal, the text data generated from the uttered voice signal, the text feature quantity corresponding to the text data, or a combination of at least two of these, the emotion information of the first user corresponding to the uttered voice signal acquired by the acquisition unit may be generated. the emotion information of the first user corresponding to the uttered voice signal acquired by the acquisition unit may be generated.

Brief Description of the Drawings

[0010]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5A

Figure 5B

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Mode for Carrying Out the Invention

[0011] Embodiments of the present invention will be described with reference to the accompanying drawings. In each figure, those with the same reference numerals have the same or similar configurations.

[0012] Hereinafter, the voice processing system according to this embodiment will be described assuming its use in customer response operations such as call centers. However, the application form of the present invention is not limited to this. This embodiment is applicable to any scenario where voice generated by performing predetermined processing on the signal of the speech voice of the first user (hereinafter referred to as the "speech voice signal") is output to the second user. Hereinafter, it is assumed that the first user is a customer and the second user is an operator, but it is not limited to this.

[0013] (Configuration of Voice Processing System) <Overall Configuration> FIG. 1 is a diagram showing an example of the outline of a voice processing system 1 according to this embodiment. As shown in FIG. 1, the voice processing system 1 includes a voice processing device 10, a terminal (hereinafter referred to as the "operator terminal") 20 used by a second user (hereinafter referred to as the "operator"), and a terminal (hereinafter referred to as the "customer terminal") 30 used by a first user (hereinafter referred to as the "customer"). ​​​​

[0014] The voice processing device 10 converts an utterance voice signal acquired by a customer terminal 30 into a The network 40 may be an external network such as the Internet. It is also possible to use external networks and internal networks such as Local Access Networks (LANs). The voice processing device 10 performs a predetermined process on the voice signal of the customer. The voice processing device 10 transmits the voice to the operator terminal 20. It may be composed of a server.

[0015] The operator terminal 20 may be, for example, a telephone, a smartphone, a personal computer, a The operator terminal 20 receives the voice generated by the voice processing device 10 through a predetermined process. Based on the voice signal or the speech signal from the customer terminal 30, a voice is output to the operator. .

[0016] The customer terminal 30 may be, for example, a telephone, a smartphone, a personal computer, a tablet, The customer terminal 30 picks up the customer's voice by a microphone and converts the voice into A speech signal, which is a signal, is transmitted to the speech processing device 10 .

[0017] <Physical configuration> FIG. 2 shows an example of a physical configuration of each device constituting the speech processing system 1 according to this embodiment. Each device (for example, a voice processing device 10, an operator terminal 20, and a customer terminal 30) ) includes a processor 10a corresponding to a calculation unit and a RAM (Random Access Memory) corresponding to a storage unit. A ROM (Read Only Memory) 10b, which corresponds to a storage unit, and a communication unit 10 d, an input unit 10e, a display unit 10f, a camera 10g, an audio input unit 10h, and an audio output unit 10i. These components are connected to be able to transmit and receive data to and from each other via a bus It should be noted that the configuration shown in FIG. 2 is an example, and each device may have a configuration other than these and may not have some of these configurations.

[0018] The processor 10a is, for example, a CPU (Central Processing Unit). The processor 10a is a control unit that controls various processes in each device by executing a program stored in the RAM 10b or the ROM 10c The processor 10a realizes the functions of each device and controls the execution of processing through cooperation with other configurations provided in the device and the program. The processor 10a receives various data from the input unit 10e or the communication unit 10d, and displays the calculation result of the data on the display unit 10f or stores it in the RAM 10b.

[0019] The RAM 10b and the ROM 10c are storage units that store data necessary for various processes and data of processing results. In addition to the RAM 10b and the ROM 10c, each device may be provided with a large-capacity storage unit such as a hard disk drive. The RAM 10b and the ROM 10c may be composed of, for example , semiconductor memory elements.

[0020] The communication unit 10d is an interface that connects each device to other devices. The communication unit 10d communicates with other devices. The input unit 10e is a device for receiving data input from the user or a device for inputting data from the outside of each device. The input unit 10e may include, for example, a keyboard, a mouse, and a touch panel. The display unit 10f is a pro It is a device that displays information according to the control by sensor 10a. The display unit 10f may be configured by, for example, an LCD (Liquid Crystal Display).

[0021] The camera 10g includes an image sensor that captures a still image or a moving image, and generates a captured image (for example, a still image or a moving image) of a predetermined area. The voice input unit 10h is a device that picks up voice, for example, a microphone. The voice output unit 10i is a device that outputs voice, for example, a speaker. For example, a still image or a moving image) is generated. The voice input unit 10h is a device that picks up voice, for example, a microphone. The voice output unit 10i is a device that outputs voice, for example, a speaker. is a device that picks up voice, for example, a microphone. The voice output unit 10i is a device that outputs voice, for example, a speaker. is a device that outputs voice, for example, a speaker.

[0022] The program for operating each device may be stored and provided in a storage medium readable by a computer such as RAM 10b or ROM 10c, or may be provided via a network 40 connected by the communication unit 10d. In each device, various operations for controlling each device are realized by the processor 10 a executing the program. Note that these physical configurations are examples and do not necessarily have to be independent configurations. For example, each device may include an LSI (Large-Scale Integration) in which the processor 10a, RAM 10b, and ROM 10c are integrated. a and RAM 10b or ROM 10c may be integrated. a executing the program. Note that these physical configurations are examples and do not necessarily have to be independent configurations. For example, each device may include an LSI (Large-Scale Integration) in which the processor 10a, RAM 10b, and ROM 10c are integrated. are realized. Note that these physical configurations are examples and do not necessarily have to be independent configurations. For example, each device may include an LSI (Large-Scale Integration) in which the processor 10a, RAM 10b, and ROM 10c are integrated. a and RAM 10b or ROM 10c may be integrated. SI (Large-Scale Integration).

[0023] <Functional Configuration> <<Voice Processing Device>> FIG. 3 is a diagram showing an example of the functional configuration of the voice processing device 10 according to the present embodiment. The voice processing device 10 includes a storage unit 101, a transmission / reception unit 102, a voice recognition unit 103, a removal unit 104, a voice synthesis unit 105, an emotion recognition unit 106, a stress recognition unit 107, a control unit 108, and a learning unit 109. is included.

[0024] The memory unit 101 stores various information, programs, algorithms, models, operation logs, etc. Specifically, the memory unit 101 stores the speech recognition model 101a, speech synthesis model 101b, emotion recognition model 101c, stress recognition model 101d, emotion suppression switching model 101e, etc. as described later.

[0025] The transceiver unit 102 transmits and / or receives various information and / or signals to and / or from the operator terminal 20 and / or the customer terminal 30. For example, the transceiver unit 102 (acquisition unit) acquires a speech signal of the customer's speech voice collected at the customer terminal 30, which is a speech voice signal. The transceiver unit 1 02 transmits a synthesized speech signal and / or a speech voice signal to the operator terminal 20. In addition, the transceiver unit 102 may acquire an operation log by the operator from the operator terminal 20. The operation log may include information related to the subjective evaluation of the customer's emotion by the operator (hereinafter, referred to as "subjective evaluation information"), the "degree of stress" described later, and the "manual switching history data" described later. In addition, the transceiver unit 102 may transmit information related to the emotion of the customer (hereinafter, referred to as "emotion information") etc. to the operator terminal 20.

[0026] The speech recognition unit 103 inputs a feature quantity (hereinafter, referred to as "speech feature quantity") extracted based on the speech voice signal acquired by the transceiver unit 102 into the speech recognition model 101a, and generates text data including a word sequence composed of one or more words. Specifically, the speech recognition unit 103 may generate a word sequence from the above speech feature quantity using the acoustic model of the speech recognition model 101a, and generate the above text data according to the analysis result of the word sequence using the language model. Speech The recognition unit 103 may perform preprocessing (for example, digitization of an analog signal, noise removal, Fourier transform, etc.) on the uttered voice signal to extract voice feature quantities.

[0027] The voice recognition model 101a is an algorithm for estimating the content of voice based on a voice signal and exists. The voice recognition model 101a may include an acoustic model that models how a certain word is likely to appear as a sound, and / or a language model that models the probability of appearance of a certain word sequence in a specific language. As the acoustic model, for example, a Hidden Markov Model (HMM) and / or a Deep Neural Network (DNN) may be used. As the language model, for example, a probabilistic language model such as an n-gram language model may be used. For example, a Hidden Markov Model (HMM) and / or a Deep Neural Network (DNN) may be used. For example, a Hidden Markov Model (HMM) and / or a Deep Neural Network (DNN) may be used. For example, a probabilistic language model such as an n-gram language model may be used.

[0028] The removal unit 104 detects a specific word sequence included in the text data generated by the voice recognition unit 103, generates text data in which the specific word sequence is removed or the specific word sequence is replaced with another word sequence, and outputs it to the voice synthesis unit 105. If no specific word sequence is detected in the text data generated by the voice recognition unit 103, the removal unit 104 may output the text data to the voice synthesis unit 105. If no specific word sequence is detected in the text data generated by the voice recognition unit 103, the removal unit 104 may output the text data to the voice synthesis unit 105.

[0029] The specific word sequence may be, for example, one or more words that give a psychological adverse effect to the listener, such as insulting the listener, denying the personality of the listener, making the listener uncomfortable, etc. Here, each word may be at least one part of speech such as a noun, verb, adverb, particle, adjective, auxiliary verb, etc. Here, each word may be at least one part of speech such as a noun, verb, adverb, particle, adjective, auxiliary verb, etc. The part of speech may include a sound change. For example, a specific word string may be "You, kill me!" It can be a sentence like "I'll do it" or "tsutsun" like "Konaruttsutsutenno" Alternatively, the removal unit 104 may remove a part of a sentence that indicates rough language. It is a text data in which only specific word strings detected in the text data are replaced with other word strings. Alternatively, the entire sentence including the particular word string may be output to the speech synthesis unit 105. The text data replaced with the other word string may be output to the speech synthesis unit 105. , or blank.

[0030] The removal unit 104 removes the text data based on a specific word string stored in advance in the storage unit 101. Detection of specific word strings in the data and / or replacement with other word strings may also be performed.

[0031] Alternatively, the removal unit 104 may remove the text data based on a model learned by machine learning. Detecting specific word sequences in the data and / or replacing them with other word sequences that have a less semantic sentiment. For example, a specific word string "omae" in the text data may be converted to "anata". Based on a machine learning-based model, certain words in the text data may be replaced. String detection and / or replacement with other word strings may be performed.

[0032] When a specific word string is detected in the text data, the removal unit 104 The system may generate information regarding the detection of the word string (hereinafter, "detection information"). The information may be, for example, information indicating that the particular word string was detected (for example, "NG word" " or "NG word detected"), information indicating the specific word string, and It may include at least one piece of information regarding a warning (hereinafter referred to as "warning information"). This warning information may be, for example, information for notifying that the content of the customer's speech to the operator may be subject to a criminal complaint such as insult or defamation of character. The detection information may be transmitted to the operator terminal 20 by the transceiver unit 102. When the detection information is generated, the voice processing device 10 may cause the customer terminal 30 to output warning information (for example, "There is a risk of insult or the like to our operator. Although we may be at fault, we would appreciate your cooperation as it may cause an excessive burden on our operator."). Such warning information can be used as a prior notice for customer harassment.

[0033] The voice synthesis unit 105 inputs the feature amount (hereinafter referred to as "text feature amount") extracted based on the text data input from the removal unit 104 into the voice synthesis model 101b to generate a synthesized voice signal (hereinafter referred to as "synthesized voice signal"). Specifically, the removal unit 104 may predict voice synthesis parameters based on the text feature amount and generate a synthesized voice signal using the predicted voice synthesis parameters. The voice synthesis unit 105 outputs the synthesized voice signal to the transceiver unit 102. The synthesized voice signal can also be said to be a signal of voice reading out the content of the text data.

[0034] The voice synthesis model 101b is an algorithm that takes text data as input and outputs a synthesized voice signal corresponding to the content of the text data. As the voice synthesis model 101b, for example, the above HMM and / or DNN may be used.

[0035] ​​​​​​​​​​ The voice synthesis model 101b may support multiple voice types. The voice synthesis unit 105 selects a voice type to be used for the synthesized voice signal from among the multiple voice types, and inputs the selected voice type and the text data into the voice synthesis model 101b to synthesize a synthesized voice signal of the selected voice type . The multiple voice types may be, for example, at least one of a voice with little intonation, a mechanical voice, a voice of a character, a voice of a celebrity, and a voice of a voice actor, etc. The voice synthesis unit 105 may receive a selection of a voice type from an operator via the operator terminal 20 . .

[0036] FIG. 4 is a diagram showing an example of generation of a synthesized voice signal according to the present embodiment. In FIG. 4, it is assumed that text data T1 to T3 are generated in the speech recognition unit 103 based on the uttered voice signals S1 to S3 acquired by the transceiver unit 102. For example, in FIG. 4, since the removal unit 104 does not detect a specific word sequence in the text data T1, the text data T1 is output to the voice synthesis unit 105 as it is. On the other hand, since the removal unit 104 detects specific word sequences (in T2, "I'll kill you", and in T3, "saying") in the text data T2 and T3, the text data T2' and T3' with the specific word sequences removed or replaced are output to the voice synthesis unit 105. For example, in the text data T2', the specific word sequence in the text data T2 is replaced with a blank (□). Also, in the text data T3', the specific word sequence "saying" in the text data T3 is replaced with "saying that". The voice synthesis unit 105 generates synthesized voice signals S1, S2', and S3' from the text data T1, T2, and T3, respectively . . . . . . . . . . .

[0037] ​The emotion recognition unit 106 generates customer emotion information based on at least one of the speech audio signal acquired by the transmission / reception unit 102, the text data generated by the speech recognition unit 103, and the subjective evaluation information received by the transmission / reception unit 102. The emotion recognition unit 106 may generate customer emotion information based on the voice feature amounts (such as intonation and volume) extracted based on the speech audio signal. The emotion recognition unit 106 may generate customer emotion information based on the fact that a specific word sequence is detected in the text data generated based on the speech audio signal, or the fact that a specific word sequence has not been detected for a predetermined time or more. The emotion recognition unit 106 may generate customer emotion information based on the captured image of the customer acquired by the camera 10g. The emotion recognition unit 106 may generate customer emotion information using the emotion recognition model 101c. The emotion recognition model 101c is a model that takes as input at least two combinations of the speech audio signal, the voice feature amounts extracted from the speech audio signal, the text data generated from the speech audio signal, the text feature amounts, or these, and outputs the emotion information, which is the emotion of the customer corresponding to the speech audio signal. Figure 5A is an explanatory diagram of the learning process of the emotion recognition model 101c. For example, for the learning of the emotion recognition model 101c, a plurality of data sets (hereinafter referred to as "data sets") each including at least one of the voice feature amounts extracted from the speech audio signal, the text feature amounts extracted from the text data, and the "subjective evaluation information" (or the feature amounts extracted from the subjective evaluation information) by the operator may be used. The subjective evaluation information is the evaluation made by the operator on the customer's speech.

[0038]

[0039] ​​​​​​​​​​​​​​​Information obtained by subjectively evaluating a customer's emotion by listening to an audio signal. For example, anger level 1 to 10 The operator may evaluate the customer's anger at multiple levels, such as The dataset for training the emotion recognition model 101c may be generated as follows, for example The operator listens to the raw speech audio signal of the customer and annotates the emotion of the customer estimated from the speech audio signal (that is, assigns "subjective evaluation information" to the speech audio signal). This obtains information in which the speech audio signal and the emotion of the customer estimated from the speech audio signal are associated on the time axis. By multiple operators assigning subjective evaluation information to multiple speech audio signals, a dataset that is such a bundle of information is obtained. The emotion recognition model 101c may be trained with such a dataset using supervised machine learning. Note that the dataset used for training the emotion recognition model 101c may include the speech audio signal in addition to or instead of the audio feature amount, and may include the text data in addition to or instead of the text feature amount.

[0040] FIG. 5B is an explanatory diagram of the estimation process using the emotion recognition model 101c. For example, as shown in FIG. 5B, by inputting the audio feature amount extracted from the speech audio signal S1 and / or the text feature amount extracted from the text data T1 generated from the speech audio signal S1 into the emotion recognition model 101c, an output corresponding to the input, that is, emotion information corresponding to the speech audio signal is obtained. Note that the speech audio signal S1 may be input to the emotion recognition model 101c in addition to or instead of the audio feature amount, and the text data T1 may be input in addition to or instead of the text feature amount.

[0041] ​​​​​​​​​​​​​ The subjective evaluation information may indicate, in numerical values, the degree of one or more emotions (e.g., at least one of "happiness", "surprise", "terror", "anger", "disgust", and "sadness", etc.). Or, the emotion information may indicate a specific emotion (e.g., "anger") that the customer is likely to feel.

[0042] The stress recognition unit 107 generates information regarding the stress situation of the operator (hereinafter referred to as "stress information"). For example, the stress recognition unit 107 may estimate the stress situation of the operator by a conventionally well-known method based on vital data such as the operator's heart rate, sweating amount, and breathing amount, or image information such as the operator's line of sight and facial expression collected using a camera. For example, the stress recognition unit 107 may estimate the stress situation of the operator based on the speech voice of the operator. Specifically, the stress recognition unit 107 may estimate the stress situation of the operator based on changes in the tone and speed of the operator's speech, the appearance of words related to apology, speaking while covering the customer's speech, etc. For example, the stress recognition unit 107 may estimate the stress situation of the operator based on the operation log of the operator terminal 20. Specifically, the stress recognition unit 107 may estimate the stress situation of the operator according to movements of a mouse or the like, and the absence of operation input in a scene where an operation should be performed. The stress recognition unit 107 may generate stress information based on the stress recognition model 101d. The stress recognition model 101d takes as input an utterance voice signal, a voice feature amount extracted from the utterance voice signal, text data generated from the utterance voice signal, a text feature amount, or a combination of at least two of these, and the stress felt by the operator listening to the utterance voice stress. It is a model that outputs an estimated value of stress. For the learning of the stress recognition model 101d, the actual measured value of the stress actually felt by the operator by listening to the customer's speech voice may be used. The dataset for learning the stress recognition model 101d may be generated, for example, as follows. The operator annotates the degree of stress (for example, at a level such as 1 to 10) felt by listening to the customer's speech voice (that is, assigns the "degree of stress" that he / she felt to the speech voice signal). As a result, information in which the speech voice signal and the stress of the operator when listening to the speech voice signal are associated on the time axis can be obtained. By multiple operators assigning the degree of stress to multiple speech voice signals, a dataset that is a bundle of such information can be obtained. The stress recognition model 101d may be learned by supervised machine learning using such a dataset. The control unit 108 performs various controls regarding the voice processing device 10. Specifically, the control unit 108 switches whether to output the synthesized voice generated by the voice synthesis unit 105 or the customer's speech voice at the operator terminal 20 based on the stress information generated in the stress recognition unit 107. The control unit 108 may switch whether to generate a synthesized voice signal based on the speech voice signal based on the stress information. For example, when the stress degree indicated by the stress information is equal to or greater than a predetermined threshold or larger, the control unit 108 may control to output the synthesized voice instead of the customer's speech voice to the operator. On the other hand, when the stress degree indicated by the stress information is smaller than or equal to the predetermined threshold, the control unit 108 Outputs the speech voice. By multiple operators assigning the degree of stress to multiple speech voice signals, a dataset that is a bundle of such information can be obtained. The stress recognition model 101d may be learned by supervised machine learning using such a dataset. The control unit 108 performs various controls regarding the voice processing device 10. Specifically, the control unit 108 switches whether to output the synthesized voice generated by the voice synthesis unit 105 or the customer's speech voice at the operator terminal 20 based on the stress information generated in the stress recognition unit 107. The control unit 108 may switch whether to generate a synthesized voice signal based on the speech voice signal based on the stress information. For example, when the stress degree indicated by the stress information is equal to or greater than a predetermined threshold or larger, the control unit 108 may control to output the synthesized voice instead of the customer's speech voice to the operator.

[0043] The control unit 108 performs various controls regarding the voice processing device 10. Specifically, the control unit 108 switches whether to output the synthesized voice generated by the voice synthesis unit 105 or the customer's speech voice at the operator terminal 20 based on the stress information generated in the stress recognition unit 107. Based on the stress information generated in the stress recognition unit 107, the control unit 108 switches whether to output the synthesized voice generated by the voice synthesis unit 105 or the customer's speech voice at the operator terminal 20. The control unit 108 may switch whether to generate a synthesized voice signal based on the speech voice signal based on the stress information. For example, when the stress degree indicated by the stress information is equal to or greater than a predetermined threshold or larger, the control unit 108 may control to output the synthesized voice instead of the customer's speech voice to the operator. On the other hand, when the stress degree indicated by the stress information is smaller than or equal to the predetermined threshold, the control unit 108 Outputs the speech voice. For example, when the stress degree indicated by the stress information is equal to or greater than a predetermined threshold or larger, the control unit 108 may control to output the synthesized voice instead of the customer's speech voice to the operator. On the other hand, when the stress degree indicated by the stress information is smaller than or equal to the predetermined threshold, the control unit 108 Outputs the speech voice. It may be controlled to output the voice to the operator. The control unit 108 senses from the operator When instruction information regarding automatic switching of the emotion suppression function is input, the above switching may be performed based on the stress information The emotion suppression function is a function of outputting synthesized voice to the operator instead of the customer's uttered voice.

[0044] The control unit 108 may perform the above switching based on the emotion information. The control unit 108 may perform the switching based on the output of the emotion suppression switching model 101e. The emotion suppression switching model 101e is a model that takes an uttered voice signal, voice feature amount, text data, text feature amount, or a combination of at least two of these as inputs and outputs the timing for switching the on / off of the emotion suppression function. The emotion suppression switching model 101e may further take stress information or emotion information as an input. Details of the emotion suppression switching model 101e will be described later.

[0045] Also, the control unit 108 may perform the above switching based on the switching information input by the operator. Here, the switching information is information regarding switching between application (on) or non-application (off) of the customer's emotion suppression function. For example, when the switching information indicates application of the customer's emotion suppression function, the control unit 108 may control to output synthesized voice to the operator. On the other hand, when the switching information indicates non-application of the customer's emotion suppression function, the control unit 108 may control to output the uttered voice to the operator. When instruction information regarding manual switching of the emotion suppression function is input from the operator, the control unit 108 may perform the above switching based on the above switching information.

[0046] ​​​​​​​​​​ The learning unit 109 may perform the learning processes of the emotion recognition model 101c, the stress recognition model 101d, and the emotion suppression switching model 101e.

[0047] The voice processing device 10 associates any one of the information shown in the following 1) to 7), or a combination of at least two pieces of information on the time axis, and may transmit it to the operator terminal 2 0 via the transceiver unit 102. 1) The customer's speech voice signal, 2) The text data generated from the speech voice signal, 3) The text data after passing through the processing of the removal unit 104, 4) The detection information, 5 ) The synthesized voice signal, 6) The customer's emotion information estimated from the customer's speech voice signal, 7) The timing for switching the on / off of the emotion suppression function. When the emotion suppression function is on, the voice processing device 10 does not have to send the customer's speech voice signal to the operator terminal 20. When the emotion suppression function is off, the voice processing device 10 does not have to send the synthesized voice signal to the operator terminal 20 either. Regardless of whether the emotion suppression function is on or off, the voice processing device 10 may send both the customer's speech voice signal and the synthesized voice signal to the operator terminal 20.

[0048] ≪Operator Terminal≫

[0049] FIG. 6 is a diagram showing an example of the functional configuration of the operator terminal according to the present embodiment. The operator terminal 20 includes a transceiver unit 201, an input reception unit 202, and a control unit 203. Note that the functional configuration shown in FIG. 6 is only an example, and it may have other configurations not shown in the figure.

[0050] The transceiver unit 201 transmits and / or receives various information and / or signals between the voice processing device 10 and / or the customer terminal 30. For example, the transceiver unit 201 receives at the customer terminal 30 It may receive a speech voice signal that is the signal of the customer's spoken voice. The transceiver unit 102 may receive a synthesized voice signal from the voice processing device 10. Also, the transceiver unit 201 may transmit subjective evaluation information to the voice processing device 10. Also, the transceiver unit 201 may receive customer emotion information from the voice processing device 10.

[0051] The input reception unit 202 receives inputs of various information based on the operation of the input unit 10e by the operator. For example, the input reception unit 202, as part of the work for generating a dataset for training the emotion recognition model 101c or the stress recognition model 101d, may receive inputs of subjective evaluation information and the degree of stress for the customer's raw speech voice signal. Hereinafter, the work of the operator inputting subjective evaluation information and the degree of stress on the operator terminal 20 is called "annotation work". The annotation work may be positioned as a work separate from the normal call center work. Also, the input reception unit 20 2 may receive an input of switching information for the customer emotion suppression function. Also, the input reception unit 202 may receive an input of instruction information indicating either manual switching or automatic switching of the emotion suppression function. The control unit 203 performs various controls regarding the operator terminal 20. For example, the control unit 20

[0052] 3 controls the display of information and / or images on the display unit 10f. Also, the control unit 203 controls the output of voice at the voice output unit 10i. The control unit 203 may control the output of voice based on the information transmitted from the voice processing device 1 0, or may control the output of voice based on the information received by the input reception unit 202. ​​​​​​​

[0053] The control unit 203 causes the voice output unit 10i to output the synthesized voice based on the synthesized voice signal received from the voice processing device 10. The control unit 203 may cause the voice output unit 10i to output the uttered voice based on the uttered voice signal from the customer terminal 30.

[0054] Further, the control unit 203 may cause the display unit 10f to display the emotion information corresponding to the synthesized voice signal based on the emotion information received from the voice processing device 10. Also, the control unit 203 may cause the display unit 10f to display the text data corresponding to the synthesized voice signal received from the voice processing device 10. For example, the control unit 203 may cause the display unit 10f to display a screen D1 including at least one of emotion information, text data, and detection information. Further, the control unit 203 may cause the display unit 10f to display stress information. For example, the control unit 203 may cause the display unit 10f to display a screen D2 including stress information.

[0055] FIG. 7 is a diagram showing an example of the screen D1 according to the present embodiment. As shown in FIG. 7, on the screen D1, the control unit 203 may cause the display unit 10f to display the emotion information I1 in accordance with the output timing T of the synthesized voice from the voice output unit 10i. By displaying the emotion information I1 every time the synthesized voice is output at the output timing T, the operator can recognize the customer's emotion in real time even when listening to the synthesized voice whose emotion has been suppressed by the emotion suppression function.

[0056] Also, on the screen D1, the control unit 203 may cause the display unit 10f to display the content of the text data I2 corresponding to the synthesized voice in accordance with the output timing T of the synthesized voice. ​​​​​​​​​​​​​​By displaying the content of the text data I2, the operator can visually understand the customer's utterance content as well as through the synthesized voice alone. Without relying solely on the synthesized voice, the operator can also visually grasp the content of the customer's speech.

[0057] Also, on the screen D1, based on the detection information received from the voice processing device 10, instead of displaying the specific word sequence itself, the control unit 203 may cause the display unit 10f to display information I3 indicating the detection of the specific word sequence (for example, "NG word detected"). This function is called the "NG word non-display function". This can avoid the operator from directly recognizing the content of the customer's speech that may have an adverse psychological impact, thus suppressing the operator's stress. Also, since the operator can be notified that such a speech occurred, the operator can appropriately respond to the customer.

[0058] Furthermore, on the screen D1, based on the emotion information from the voice processing device 10, for each output timing T of the synthesized voice, the control unit 203 may cause the display unit 10f to display the level I4 of the customer's specific emotion in chronological order. For example, in FIG. 7, the level I4 of the customer's "anger" for each output timing T of the synthesized voice is shown as a line graph. This allows the operator to easily grasp the transition of the customer's specific emotion (for example, "anger"), thus improving the satisfaction of the operator's response to the customer.

[0059] On the screen D1, the control unit 203 may also cause the display unit 10f to display a selection button I5. The selection button I5 is an interface that allows the operator to select whether to automatically or manually switch between applying (on) or not applying (off) the emotion suppression function. The operator ​​​​By performing operations such as clicking, tapping, or sliding on the selection button I5, the "automatic switching mode" and the "manual switching mode" can be switched. In the automatic switching mode, for example, the on / off of the emotion suppression function is automatically switched based on emotion information, stress information, or the output from the emotion suppression switching model 101e. When the "manual switching mode" is selected, the control unit 203 may display a switching button I6, which is an interface that enables the operator to select the application or non-application of the emotion suppression function, on the display unit 10f. The timing when the operator switches the on / off of the emotion suppression function is associated with the customer's speech voice (and / or various feature quantities extracted based on the speech voice) on the time axis and stored as "manual switching history data" in a storage unit (not shown). The "manual switching history data" may further be associated with the operator's identification information. For example, based on emotion information, stress information, or the output from the emotion suppression switching model 101e, etc. The on / off of the emotion suppression function automatically switches.

[0060] When the "manual switching mode" is selected, the control unit 203 may display a switching button I6, which is an interface that enables the operator to select the application or non-application of the emotion suppression function, on the display unit 10f. The timing when the operator switches the on / off of the emotion suppression function is associated with the customer's speech voice (and / or various feature quantities extracted based on the speech voice) on the time axis and stored as "manual switching history data" in a storage unit (not shown). The "manual switching history data" may further be associated with the operator's identification information. The customer's speech voice (and / or various feature quantities extracted based on the speech voice) on the time axis and stored as "manual switching history data" in a storage unit (not shown). The "manual switching history data" may further be associated with the operator's identification information. The "manual switching history data" may further be associated with the operator's identification information.

[0061] The switching button I7 is a button for switching the on / off of the "NG word non-display function". When the "NG word non-display function" is off, even if a specific word sequence is detected in the text data I2, the text data I2 before the processing by the removal unit 104 is directly displayed on the display unit 10f. When the emotion suppression function is on and the "NG word non-display function" is off, the operator does not directly hear a specific word sequence from the customer, so the stress is reduced. On the other hand, by accurately grasping the customer's speech content, the customer's emotions can be grasped more accurately. When the "NG word non-display function" is off, even if a specific word sequence is detected in the text data I2, the text data I2 before the processing by the removal unit 104 is directly displayed on the display unit 10f. When the emotion suppression function is on and the "NG word non-display function" is off, the operator does not directly hear a specific word sequence from the customer, so the stress is reduced. On the other hand, by accurately grasping the customer's speech content, the customer's emotions can be grasped more accurately. When the emotion suppression function is on and the "NG word non-display function" is off, the operator does not directly hear a specific word sequence from the customer, so the stress is reduced. On the other hand, by accurately grasping the customer's speech content, the customer's emotions can be grasped more accurately. The data set for learning the emotion suppression switching model 101e includes stress information, emotion information

[0062] The data set for learning the emotion suppression switching model 101e includes stress information, emotion information information, speech signal S1, speech features, text data, text features or any of these At least two combinations and the time when the operator switches the emotion suppression function on and off. The emotion suppression switching model 10 may be a bundle of data associated with each other on the time axis. There are various ways to learn 1e, such as those described below in 1) to 3). 1) The emotion suppression switching model 101e may be trained for each operator. The emotion suppression switching model 101e applied to the data is the emotion suppression by the operator. It may be possible to learn based only on the "manual switching history data" of the function. The emotion suppression switching model 101e can suppress emotions at the timing that suits the operator's preference. Or, 2) the operator can select the The emotion suppression switching model 101e is based on the "manual switching history data" of an unspecified number of operators. According to this method, the data available for learning may be Since the number of the emotion inhibition switching models 101e is increased, the emotion inhibition switching model 101e can be learned quickly. Or, 3) the emotion suppression switching model 101e applied to a certain operator is "Manual switching history data" by operators with similar age, gender, and other characteristics to the operator According to this method, the number of learning methods used is smaller than that in the method 1). Since the amount of data that can be used increases, the emotion inhibition switching model 101e can be trained quickly. By comparing this with method 2), you can learn the switching timing that suits you best. It becomes like this.

[0063] FIG. 8 is a diagram showing an example of a screen D2 according to the present embodiment. 203 may display stress information from the voice processing device 10. For example, in FIG. 8 , as stress information, information indicating an estimated value of stress felt by the operator (e.g., "5 6%") and information indicating a relative evaluation value from the operator's normal state (e.g., "8.1% lower than normal") are displayed.

[0064] FIG. 12 is a diagram showing an example of the screen D3 according to the present embodiment. In the screen D3, the control unit 203 may display an interface I8 for the operator to perform annotation work. The operator, for example, while listening to the raw voice (sample voice) of the customer, selects the emotion of the customer felt from the sample voice from the interface I8 each time. In FIG. 1 2, the customer emotion I1 is subjective evaluation information of the customer emotion by the operator. For example , if the operator annotates the emotion of "anger" for the sample voice "Please somehow deliver it by this evening", then, as shown in FIG. 12, the sample voice of "Please somehow deliver it by this evening" and the information of "anger" are associated on the time axis. The annotation may be performed in units of sentences or at predetermined time intervals.

[0065] (Operation of the Voice Processing System) FIG. 9 is a flowchart showing an example of the emotion suppression operation according to the present embodiment. Note that FIG. 9 is merely an illustration, and the order of at least some steps (e.g., step S106) may be swapped, steps not shown may be performed, or some steps may be omitted.

[0066] ​​​The voice processing device 10 acquires a speech voice signal, which is a signal of the customer's speech picked up by the voice input unit 10h of the customer terminal 30 (S101).

[0067] The voice processing device 10 inputs the feature amount extracted based on the speech voice signal acquired in S101 into the speech recognition model 101a to generate text data including a word sequence composed of one or more words (S102).

[0068] The voice processing device 10 determines whether a specific word sequence is included in the text data generated in S102 (S103). When the specific word sequence is included in the text data, the voice processing device 10 generates text data in which the specific word sequence is removed or the specific word sequence is converted into another word sequence (S104).

[0069] The voice processing device 10 inputs the feature amount extracted based on the text data into the voice synthesis model 101b to generate a synthesized voice signal, which is a signal of the synthesized voice (S105).

[0070] The voice processing device 10 inputs the feature amount extracted based on at least one of the speech voice signal acquired in S101, the text data generated in S102, and the subjective evaluation information of the customer's emotion input by the operator into the emotion recognition model 101c to generate customer emotion information (S106).

[0071] The operator terminal 20 outputs the synthesized voice from the voice output unit 10i based on the synthesized voice signal generated in S105, and displays the emotion information corresponding to the synthesized voice on the display unit 10f in accordance with the output timing T of the synthesized voice (S107, for example, FIG. 7).

[0072] ​​​​​​​​​​​​ The voice processing device 10 determines whether to end the processing (S108). If the processing is not ended (S108: NO), the voice processing device 10 re-executes the processes S101 to S107. On the other hand, when ending the voice conversion processing (S108: YES), the voice processing device 10 ends the processing.

[0073] FIG. 10 is a flowchart showing the automatic switching operation of the emotion suppression function according to the present embodiment. Note that FIG. 10 is merely an example, and the order of at least some steps may be changed, steps not shown may be performed, or some steps may be omitted.

[0074] The voice processing device 10 generates stress information of the operator (S201).

[0075] The voice processing device 10 determines whether the stress information satisfies a predetermined condition (S202 ). For example, the predetermined condition may be that the stress degree indicated by the stress information is equal to or greater than a predetermined threshold value.

[0076] When the stress information satisfies the predetermined condition (S202: YES), the voice processing device 10 may apply the emotion suppression function (that is, output the synthesized voice from the operator terminal 20) ( S203). On the other hand, when the stress information does not satisfy the predetermined condition ( S202: NO), the voice processing device 10 may not apply the emotion suppression function (that is, output the customer's speech voice from the operator terminal 20) (S204).

[0077] The voice processing device 10 determines whether to end the processing (S205). If the processing is not ended If not (S205: NO), the voice processing device 10 re-executes processes S201 to S204. On the other hand, when ending the voice conversion process (S205: YES), the voice processing device 10 ends the process. In S201 and S202, the voice processing device 10 may determine whether to apply the emotion suppression function based on the output of the emotion information and the emotion suppression switching model 101e.

[0078] As described above, according to the voice processing system 1 according to the present embodiment, text data is generated based on the customer's spoken voice signal, and the synthesized voice generated based on the text data is output to the operator. Therefore, it is possible to let the operator hear a synthesized voice in which the customer's emotion included in the customer's spoken voice is sufficiently suppressed, and the stress on the operator caused by the customer's emotional speech can be reduced. The inventor of the present invention conducted an experiment in which about 50 subjects were asked to compare four types of voices: 1) the customer's spoken voice itself, 2) the voice with the volume of the customer's spoken voice adjusted, 3) the voice with the voice quality of the customer's spoken voice converted, and 4) the synthesized voice generated after converting the customer's spoken voice into text, and evaluate the degree of anger felt from the voices on a 7-point scale. As a result, compared with 2) and 3), 4) was significantly more effective in reducing the anger conveyed to the subjects.

[0079] Also, according to the voice processing system 1 according to the present embodiment, not only the synthesized voice is output to the operator, but also the customer's emotion information can be notified in accordance with the timing of the synthesized voice output. Therefore, the operator who hears the synthesized voice can recognize the customer's emotion in real time, and can appropriately respond to the customer.

[0080] ​​​​​​​​In addition, according to the voice processing system 1 according to the present embodiment, whether to apply the emotion suppression function (that is, which of the synthesized voice or the uttered voice is output to the operator) is switched based on the stress information of the operator or the emotion information of the customer or the like. Therefore, it is possible to appropriately balance the stress of the operator and the satisfaction of the customer. For the operator, whether to output the synthesized voice or the uttered voice can be switched, so that the balance between the stress of the operator and the satisfaction of the customer can be appropriately achieved.

[0081] (Modified Example) In the voice processing system 1, the voice recognition unit 103 generates text data including a word sequence determined as one or more sentences from the uttered voice signal, but it is not limited to this. The voice recognition unit 103 may generate text data including a word sequence consisting of one or more words (part of speech or morpheme) before the word sequence recognized from the uttered voice signal is determined as one or more sentences. The removal unit 104 removes a specific word sequence in the text data that has not been determined as the sentence, and the voice synthesis unit 105 may generate a synthesized voice signal from the text data that has not been determined as the sentence.

[0082] FIG. 11 is a diagram showing an example of generation of a synthesized voice signal according to a modified example of the present embodiment. In FIG. 11, it is assumed that text data T41 to T43 are generated in the voice recognition unit 103 based on the uttered voice signal S4 acquired by the transmission / reception unit 102. As shown in FIG. 11, the text data T41 to T43 are different from FIG. 4 in that the text data is generated in units of morphemes having meanings ("quickly", "send", "please") before the sentence "Please send it quickly" is determined. The removal unit 104 determines whether each of the text data T41 to T43 includes a specific word sequence, removes the specific word sequence, and then the voice synthesis unit ​​​​​​​​​​Output to 105. The voice synthesis unit 105 synthesizes voice signals S41 to S43 from the text data T41 to T43 respectively. Generate the synthesized voice signals S41 to S43.

[0083] As shown in FIG. 11, before the sentence is finalized, text data is generated in one or more morpheme units and the synthesized voice is output, thereby reducing the response delay of the operator due to the generation of text data. Note that a model or the like for determining whether a plurality of text data (or synthesized voices) in morpheme units is semantically unnatural may be used. In addition, in order to reduce the response delay, before and / or after each of the synthesized voice signals S1 to S3 shown in FIG. 4 and the synthesized voice signals S41 to S43 shown in FIG. 11, for example, fillers such as "ah~", "eh~", "well" may be added. Thereby, the operator can also prevent the satisfaction of the customer from decreasing due to the response delay. In addition, the voice synthesis unit 105 may select a voice synthesis model 101b that matches the customer's emotion from among a plurality of voice synthesis models 101b based on the emotion of the customer estimated by the emotion recognition unit 106. For example, when the emotion of the customer estimated by the emotion recognition unit 106 is "excited", the voice synthesis unit 105 may use a voice synthesis model 101b with a fast pitch and intense intonation. For

[0084] example, when the emotion of the customer estimated by the emotion recognition unit 106 is "wailing", the voice synthesis unit 105 may use a voice synthesis model 101b that outputs a voice like a crying voice. Alternatively, the voice synthesis unit 105 may change the parameters of the voice synthesis model 101b based on the emotion of the customer estimated by the emotion recognition unit 106 and adjust it so that a voice that matches the customer's emotion is output. "ah~", "eh~", "well", etc. may be added. Thereby, the operator can also prevent the satisfaction of the customer from decreasing due to the response delay. In addition, the voice synthesis unit 105 may select a voice synthesis model 101b that matches the customer's emotion from among a plurality of voice synthesis models 101b based on the emotion of the customer estimated by the emotion recognition unit 106. For example, when the emotion of the customer estimated by the emotion recognition unit 106 is "excited", the voice synthesis unit 105 may use a voice synthesis model 101b with a fast pitch and intense intonation. For

[0085] example, when the emotion of the customer estimated by the emotion recognition unit 106 is "wailing", the voice synthesis unit 105 may use a voice synthesis model 101b that outputs a voice like a crying voice. Alternatively, the voice synthesis unit 105 may change the parameters of the voice synthesis model 101b based on the emotion of the customer estimated by the emotion recognition unit 106 and adjust it so that a voice that matches the customer's emotion is output. In addition, the voice synthesis unit 105 may select a voice synthesis model 101b that matches the customer's emotion from among a plurality of voice synthesis models 101b based on the emotion of the customer estimated by the emotion recognition unit 106. For example, when the emotion of the customer estimated by the emotion recognition unit 106 is "excited", the voice synthesis unit 105 may use a voice synthesis model 101b with a fast pitch and intense intonation. For example, when the emotion of the customer estimated by the emotion recognition unit 106 is "wailing", the voice synthesis unit 105 may use a voice synthesis model 101b that outputs a voice like a crying voice. Alternatively, the voice synthesis unit 105 may change the parameters of the voice synthesis model 101b based on the emotion of the customer estimated by the emotion recognition unit 106 and adjust it so that a voice that matches the customer's emotion is output. example, when the emotion of the customer estimated by the emotion recognition unit 106 is "wailing", the voice synthesis unit 105 may use a voice synthesis model 101b that outputs a voice like a crying voice. Alternatively, the voice synthesis unit 105 may change the parameters of the voice synthesis model 101b based on the emotion of the customer estimated by the emotion recognition unit 106 and adjust it so that a voice that matches the customer's emotion is output. example, when the emotion of the customer estimated by the emotion recognition unit 106 is "wailing", the voice synthesis unit 105 may use a voice synthesis model 101b that outputs a voice like a crying voice. Alternatively, the voice synthesis unit 105 may change the parameters of the voice synthesis model 101b based on the emotion of the customer estimated by the emotion recognition unit 106 and adjust it so that a voice that matches the customer's emotion is output. example, when the emotion of the customer estimated by the emotion recognition unit 106 is "wailing", the voice synthesis unit 105 may use a voice synthesis model 101b that outputs a voice like a crying voice. Alternatively, the voice synthesis unit 105 may change the parameters of the voice synthesis model 101b based on the emotion of the customer estimated by the emotion recognition unit 106 and adjust it so that a voice that matches the customer's emotion is output. example, when the emotion of the customer estimated by the emotion recognition unit 106 is "wailing", the voice synthesis unit 105 may use a voice synthesis model 101b that outputs a voice like a crying voice. Alternatively, the voice synthesis unit 105 may change the parameters of the voice synthesis model 101b based on the emotion of the customer estimated by the emotion recognition unit 106 and adjust it so that a voice that matches the customer's emotion is output. 1b and adjust it so that a voice that matches the customer's emotion is output. An operator who directly hears the raw voice when a customer is excited feels extremely strong stress and ends up. On the other hand, in order for the operator to appropriately perform customer service operations, it is necessary to accurately grasp the customer's emotions in real time. By not allowing the operator to directly hear the speech voice, the operator does not feel excessive stress, and by adding the customer's emotions to the synthesized voice, the operator can grasp the customer's emotions in real time through hearing.

[0086] (Other embodiments) In the above embodiment, the speech voice signal of the customer is texturized and the synthesized voice signal is output to the operator but is not limited to this. The voice processing device 10 inputs the voice feature amount extracted based on the speech voice signal of the customer into the voice conversion model, generates a converted voice signal, and may output the converted voice from the operator terminal 20.

[0087] The "voice conversion model" described in the claims includes both a model that once texturizes the speech voice signal and outputs it as a synthesized voice, and a model that converts the voice quality without texturizing the speech voice signal and outputs it. By outputting the synthesized voice or the converted voice to the operator instead of the customer's speech voice, although there is a difference in the degree of effect, the stress felt by the operator can be reduced. On the other hand, for the performance of customer service operations, it is also essential for the operator to grasp the customer's emotions in real time.

[0088] The voice processing device 10 in this modification inputs the voice feature amount extracted based on the speech voice signal of the customer into the voice conversion model and generates a converted voice signal. The voice processing device 10 is 1 Generate information by associating 1) a converted voice signal and 2) customer emotion information estimated from the customer's spoken voice on the time axis, and transmit it to the operator terminal 20. The information transmitted by the voice processing device 10 may include the spoken voice signal, text data generated from the spoken voice signal, text data after passing through the processing of the removal unit 10 4, detection information, and the timing for switching the on / off of the emotion suppression function may be associated.

[0089] The operator terminal 20 outputs the converted voice signal received from the voice processing device 10 from the voice output unit 1 0i, and may display information indicating emotion information on the display unit 10f in accordance with the output timing T of the converted voice from the voice output unit 10i. The operator terminal 20 may further display text data on the display unit 1 0f in accordance with the output timing T of the converted voice from the voice output unit 10i. Such a display mode may be as illustrated in FIG. 7.

[0090] In the voice processing device 10 in this modification example, based on the emotion information, the voice processing device 10 may generate a converted voice signal so that the emotion indicated by the emotion information is reflected in the converted voice. For example, when the emotion indicated by the emotion information is "excitement", a voice conversion model with a fast pitch and intense intonation may be used. For example, when the emotion indicated by the emotion information is "wailing", a voice conversion model that outputs a voice like a crying voice may be used. The voice processing device 10 may generate a converted voice signal so that the emotion indicated by the emotion information is reflected in the converted voice. By not directly letting the operator hear the spoken voice, the operator can grasp the customer's emotion in real time through hearing without feeling excessive stress, by imposing the customer's emotion on the converted voice.

[0091] In the voice processing system 1 in this modification example, the annotation work by the operator can be performed on the converted voice during the normal call center work by the operator. When the operator annotates the converted voice with "angry emotion", based on the result of the annotation, the voice conversion model may be adjusted in real time to output a softer voice.

[0092] In the embodiment described above, it is assumed that the first user is a customer and the second user is an operator in a call center. However, the applicable scenarios of this embodiment are not limited to call centers. For example, it can be applied to any scenario where the voice suppressing the emotion of the first user is output to the second user, such as a Web meeting. That is, this embodiment can be used not only as a measure against customer harassment, but also as a measure by the company side against various harassments such as in-house power harassment.

[0093] In the embodiment described above, the process of "associating the emotion information and the synthesized voice on the time axis" As shown in FIG. 7, as long as it is a realizable mode to display the emotion information estimated from the original uttered voice in accordance with the output timing of the synthesized voice or the converted voice, regardless of its specific mode. The process of "associating on the time axis" in the embodiment described above may be a process of associating based on time information such as hours, minutes, and seconds, or a process of associating based on information such as how many minutes and seconds have elapsed since the start of the uttered voice information, or a process of associating in units of sentences, words, or morphemes.

[0094] In the voice processing system 1 in the embodiment described above, from the customer, his / her own voice may not be known to reach the operator with emotion suppressed. That is , whether the emotion suppression function is on or off may not be grasped by the customer in this way.

[0095] The annotation work may be performed by the operator on the operator terminal 20, or separately, a dedicated application or terminal for the annotation work may be prepared.

[0096] Also, the embodiments described above are for facilitating the understanding of the present invention, and are not for limiting the interpretation of the present invention. Each element included in the embodiments, as well as their arrangements, materials , conditions, shapes, sizes, etc. are not limited to those illustrated and can be changed as appropriate . Also, the configurations shown in different embodiments can be partially replaced or combined with each other . Also, the functions described as the functions of the voice processing device 10 may be provided in the operator terminal 2 0. Also, the functions described as the functions of the operator terminal 20 may be provided in the voice processing device 10.

Explanation of Reference Numerals

[0097] 1... Voice processing system, 10... Voice processing device, 20... Operator terminal, 30... Customer terminal , 10a... Processor, 10b... RAM, 10c... ROM, 10d... Communication unit, 10e... Input unit, 10f... Display unit, 10g... Camera, 10h... Voice input unit, 10i... Voice output unit, 1 01... Storage unit, 102... Transmission / reception unit, 103... Voice recognition unit, 104... Removal unit, 105... Voice synthesis unit, 106... Emotion recognition unit, 107... Stress recognition unit, 108... Control unit, 109... Learning Section, 201... Transmission and reception section, 202... Input reception section, 203... Control section

Claims

1. An acquisition unit that acquires a speech audio signal that is a signal of the speech audio of a first user; Inputting the feature amount extracted based on the speech audio signal into an audio recognition model to generate text data including a word sequence consisting of one or more Words, an audio recognition unit; Inputting the feature amount extracted based on the text data into an audio synthesis model to generate a synthesized audio An audio synthesis unit that is a signal of the voice; An audio output unit that outputs the synthesized voice to a second user; An audio processing system comprising:

2. An emotion recognition unit that generates emotion information of a first user corresponding to the speech audio signal; A display unit that displays the emotion information to the second user, comprising: The display unit is configured to display the emotion information corresponding to the synthesized voice in accordance with the output timing of the synthesized voice by the audio output unit. The audio processing system according to claim 1.

3. An emotion recognition unit that generates emotion information of a first user corresponding to the speech audio signal; A control unit that switches which of the synthesized voice or the speech audio is output from the audio output unit based on the emotion information; The audio processing system according to claim 1 or 2, comprising:

4. The emotion recognition unit inputs a speech audio signal, a feature amount extracted from the speech audio signal, text data generated from the speech audio signal, a feature amount extracted from the text data, or a combination of at least two of these, and inputs the speech audio signal acquired by the acquisition unit, the audio feature amount extracted from the speech audio signal, the text data generated from the speech audio signal, the text feature amount corresponding to the text data, or a combination of at least two of these into a machine-learned emotion recognition model that outputs the emotion information of the speaker of the speech audio signal, thereby generating the emotion information of the first user corresponding to the speech audio signal acquired by the acquisition unit. The audio processing system according to claim 2 or 3.

5. The audio synthesis unit generates the synthesized audio signal based on the emotion information generated by the emotion recognition unit so that the emotion indicated by the emotion information is reflected in the synthesized voice. The audio processing system according to any one of claims 2 to Claim 4.

6. A stress recognition unit that generates stress information regarding the stress situation of the second user; Based on the stress information, which of the synthesized voice or the speech audio is output from the audio output unit ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ A control unit that switches whether to output the flash; The voice processing system according to claim 1 or 2, comprising:

7. Based on the switching information input by the second user, which of the synthesized voice or the uttered voice is output from the voice output unit A control unit that switches; The switching information input by the second user is associated with the uttered voice signal at the time when the switching information is input on the time axis, and based on the information, the uttered voice A learning unit that machine-learns an emotion suppression switching model that takes a signal, a feature amount extracted from the uttered voice signal, text data generated from the uttered voice signal, a feature amount extracted from the text data, or a combination of at least two of these as inputs and outputs the timing for switching between the synthesized voice and the uttered voice. A signal, a feature amount extracted from the uttered voice signal, text data generated from the uttered voice signal, a feature amount extracted from the text data, or a combination of at least two of these; The voice processing system according to any one of claims 1 to 6, wherein the control unit inputs a signal, a feature amount extracted from the uttered voice signal, text data generated from the uttered voice signal, a feature amount extracted from the text data, or a combination of at least two of these to the emotion suppression switching model to generate the timing for switching between the synthesized voice and the uttered voice. And outputs the timing for switching between the synthesized voice and the uttered voice. The voice processing system according to any one of claims 1 to 6, further comprising a learning unit that machine-learns an emotion suppression switching model, where the control unit inputs the uttered voice signal acquired by the acquisition unit, a feature amount extracted from the uttered voice signal, text data generated from the uttered voice signal, a feature amount extracted from the text data, or a combination of at least two of these to the emotion suppression switching model to generate the timing for switching between the synthesized voice and the uttered voice. The control unit inputs the uttered voice signal acquired by the acquisition unit, a feature amount extracted from the uttered voice signal, text data generated from the uttered voice signal, a feature amount extracted from the text data, or a combination of at least two of these to the emotion suppression switching model to generate the timing for switching between the synthesized voice and the uttered voice. A feature amount extracted from the uttered voice signal, text data generated from the uttered voice signal, a feature amount extracted from the text data, or a combination of at least two of these; The voice processing system according to any one of claims 1 to 6, wherein the control unit inputs a signal, a feature amount extracted from the uttered voice signal, text data generated from the uttered voice signal, a feature amount extracted from the text data, or a combination of at least two of these to the emotion suppression switching model to generate the timing for switching between the synthesized voice and the uttered voice. The voice processing system according to any one of claims 1 to 6, wherein the control unit inputs a signal, a feature amount extracted from the uttered voice signal, text data generated from the uttered voice signal, a feature amount extracted from the text data, or a combination of at least two of these to the emotion suppression switching model to generate the timing for switching between the synthesized voice and the uttered voice. The voice processing system according to any one of claims 1 to 6, wherein the control unit inputs a signal, a feature amount extracted from the uttered voice signal, text data generated from the uttered voice signal, a feature amount extracted from the text data, or a combination of at least two of these to the emotion suppression switching model to generate the timing for switching between the synthesized voice and the uttered voice.

8. An acquisition unit that acquires an uttered voice signal that is a signal of the uttered voice of the first user; A voice recognition unit that inputs a feature amount extracted based on the uttered voice signal to a voice recognition model to generate text data including a word sequence consisting of one or more words; A voice synthesis unit that inputs a feature amount extracted based on the text data to a voice synthesis model to generate a synthesized voice signal that is a signal of the synthesized voice output to the second user; A voice processing apparatus comprising: A voice synthesis unit that inputs a feature amount extracted based on the text data to a voice synthesis model to generate a synthesized voice signal that is a signal of the synthesized voice output to the second user. The voice processing apparatus according to claim 8, comprising:

9. An emotion recognition unit that generates emotion information of the first user corresponding to the uttered voice signal; A transmission unit that associates the emotion information with the uttered voice signal and / or the synthesized voice signal corresponding to the emotion information on the time axis and transmits the associated information to an external device. The voice processing apparatus according to claim 8, comprising: The voice processing apparatus according to claim 8, further comprising a transmission unit that associates the emotion information with the uttered voice signal and / or the synthesized voice signal corresponding to the emotion information on the time axis and transmits the associated information to an external device.

10. A step of acquiring an uttered voice signal that is a signal of the uttered voice of the first user; A step of inputting a feature amount extracted based on the uttered voice signal to a voice recognition model to generate text data including a word sequence consisting of one or more words; A step of inputting a feature amount extracted based on the text data to a voice synthesis model to generate a synthesized voice signal that is a signal of the synthesized voice output to the second user. Input the feature amount extracted based on the text data into a voice synthesis model to obtain a synthesized voice A step of generating a synthesized voice signal which is a signal of the synthesized voice A step of outputting the synthesized voice to a second user A voice processing method including the above steps.

11. An acquisition unit that acquires a speech voice signal which is a signal of the speech voice of a first user Input the feature amount extracted based on the speech voice signal into a voice conversion model to obtain a converted voice A voice conversion unit that generates a signal of the converted voice A voice output unit that outputs the converted voice to a second user An emotion recognition unit that generates emotion information of the first user corresponding to the speech voice signal In accordance with the output timing of the converted voice by the voice output unit to the second user A display unit that displays the emotion information corresponding to the converted voice A voice processing system including the above components.

12. An acquisition unit that acquires a speech voice signal which is a signal of the speech voice of a first user Input the feature amount extracted based on the speech voice signal into a voice conversion model to obtain a converted voice A voice conversion unit that generates a signal of the converted voice A voice output unit that outputs the converted voice to a second user An emotion recognition unit that generates emotion information of the first user corresponding to the speech voice signal, and includes 、 Based on the emotion information generated by the emotion recognition unit, the voice conversion unit Generates a signal of the converted voice so that the emotion indicated by the emotion information is reflected in the converted voice. A voice processing system Including the above components.

Citation Information

Patent Citations

  • Voice control system, voice controller, voice control method, and voice control program

    JP2013046088A

  • Speech synthesis system, and prediction model learning method and device thereof

    JP2017049535A

  • Voice synthesis parameter generating device and computer program for the same

    JP2018013721A

  • Information processing system, information processing method, and program

    JP2019110451A

  • Information processing device, information processing method and program

    JP2020021025A