Speech processing device, speech processing method, and program
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- SOFTBANK CORPORATION
- Filing Date
- 2025-04-22
- Publication Date
- 2026-08-03
Smart Images

Figure 0007899390000001 
Figure 0007899390000002 
Figure 0007899390000003
Abstract
Description
Technical Field
[0001] The present invention relates to an audio processing system, an audio processing apparatus, and an audio processing method.
Background Art
[0002] Conventionally, for the purpose of improving customer satisfaction (CS), various call centers have been operated where operators respond to customer complaints and the like by phone. In such customer service operations, "customer harassment" where customers make intimidating remarks or unreasonable demands to operators has been regarded as a problem, which may lead to mental disorders of operators or a high turnover rate of operators. <0^000012>
[0003] In recent years, in order to protect employees who are operators from such customer harassment, voice conversion systems have also been studied by companies. For example, in Patent Document 1, the volume and pitch variation amounts are calculated from an input voice signal, and when the volume and pitch variation amounts exceed a predetermined value, the volume and pitch are controlled to be converted and output so that the volume and pitch variation amounts fall within a predetermined range. <00^00021>
Prior Art Documents
Patent Documents
[0004]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0005] However, for example, by only converting the speech voice of a speaker by the method described in Patent Document 1 This is because the speaker's (first user's) emotions are not sufficiently controlled, and the listener's (second user's) story... There is a risk that the response may not be sufficiently reduced. On the other hand, in order to reduce the listener's stress, By converting the speaker's utterances output to the hand, the listener can fully recognize the speaker's emotions. Furthermore, there is a risk that the listener may not be able to respond appropriately.
[0006] Therefore, the present invention aims to sufficiently reduce the listener's stress and / or enable the listener to respond appropriately. The present invention provides a voice processing system, a voice processing device, and a voice processing method that enable this. [Means for solving the problem]
[0007] A speech processing system according to one aspect of the present invention is a signal of a first user's utterance. An acquisition unit that acquires a spoken speech signal, and a feature quantity extracted based on the spoken speech signal. A sound that is input into a recognition model to generate text data containing a sequence of one or more words. The voice recognition unit and the feature quantities extracted based on the text data are input to the speech synthesis model. The system includes a speech synthesis unit that generates a synthesized speech signal, which is a synthesized speech signal, and a second user. The system includes an audio output unit that outputs the synthesized voice.
[0008] According to this embodiment, text data is generated based on the speech signal of the first user, The synthesized speech generated based on the text data is output to a second user. The first user's speech, with the customer's emotions sufficiently suppressed, is used to create a synthesized voice for the second user. It can be heard by the second user, and the stress on the second user caused by the first user's emotional utterances. This can significantly reduce the problem.
[0009] In the above voice processing system, the emotion recognition unit uses, as input, the uttered voice signal, the feature quantity extracted from the uttered voice signal , the text data generated from the uttered voice signal, the feature quantity extracted from the text data , or a combination of at least two of these, and inputs them into an emotion recognition model that has been machine-learned to output the emotion information of the speaker of the uttered voice signal. By inputting the uttered voice signal acquired by the acquisition unit, the voice feature quantity extracted from the uttered voice signal, the text data generated from the uttered voice signal, the text feature quantity corresponding to the text data, or a combination of at least two of these, it is possible to generate the emotion information of the first user corresponding to the uttered voice signal acquired by the acquisition unit.
Brief Description of the Drawings
[0010] [Figure 1] FIG. 1 is a diagram showing an example of the outline of the voice processing system 1 according to the present embodiment. [Figure 2] FIG. 2 is a diagram showing an example of the physical configuration of each device constituting the voice processing system 1 according to the present embodiment. [Figure 3] FIG. 3 is a diagram showing an example of the functional configuration of the voice processing device 10 according to the present embodiment. [Figure 4] FIG. 4 is a diagram showing an example of the generation of a synthesized voice signal according to the present embodiment. [Figure 5A] FIG. 5 is a diagram showing an example of the generation of customer emotion information according to the present embodiment. [Figure 5B] FIG. 6 is a diagram showing an example of the generation of customer emotion information according to the present embodiment. [Figure 6] FIG. 7 is a diagram showing an example of the functional configuration of the operator terminal 20 according to the present embodiment. [Figure 7] FIG. 8 is a diagram showing an example of the screen D1 according to the present embodiment. [Figure 8] FIG. 9 is a diagram showing an example of the screen D2 according to the present embodiment. [Figure 9] It is a flowchart showing an example of an emotion suppression operation according to the present embodiment. [Figure 10] It is a flowchart showing an automatic switching operation of an emotion suppression function according to the present embodiment. [Figure 11] It is a diagram showing an example of generation of a synthesized voice signal according to a modification example of the present embodiment. [Figure 12] It is a diagram showing an example of screen D3 according to the present embodiment.
Mode for Carrying Out the Invention
[0011] Embodiments of the present invention will be described with reference to the accompanying drawings. In each figure, those with the same reference numerals have the same or similar configurations.
[0012] Hereinafter, the voice processing system according to the present embodiment will be described assuming its use in customer response operations such as a call center. However, the application form of the present invention is not limited to this. The present embodiment is applicable to any scenario where voice generated by performing predetermined processing on a signal of the speech voice of a first user (hereinafter referred to as "speech voice signal") is output to a second user. Hereinafter, it is assumed that the first user is a customer and the second user is an operator, but it is not limited to this.
[0013] (Configuration of Voice Processing System) <Overall Configuration> FIG. 1 is a diagram showing an example of the outline of a voice processing system 1 according to the present embodiment. As shown in FIG. 1, the voice processing system 1 includes a voice processing device 10, a terminal (hereinafter referred to as "operator terminal") 20 used by a second user (hereinafter referred to as "operator"), and a terminal (hereinafter referred to as "customer terminal") 30 used by a first user (hereinafter referred to as "customer").
[0014] The voice processing device 10 receives the spoken voice signal acquired by the customer terminal 30 via the network 40. It receives via. Network 40 is an external network such as the Internet. Yes, external networks and internal networks such as Local Access Networks (LANs) It may include "ku". The voice processing device 10 performs predetermined processing on the customer's spoken voice signal. The voice is transmitted to the operator terminal 20. The voice processing device 10 can send one or more voices to the operator terminal 20. It may be composed of bars.
[0015] The operator terminal 20 is, for example, a telephone, smartphone, personal computer, etc. Bullets, etc. The operator terminal 20 is generated by the voice processing device 10 through predetermined processing. Based on the voice signal or the spoken voice signal from the customer terminal 30, the system outputs voice to the operator. .
[0016] Customer terminal 30 is, for example, a telephone, smartphone, personal computer, tablet The customer terminal 30 picks up the customer's spoken voice with a microphone and processes the spoken voice. The speech signal, which is a signal, is transmitted to the speech processing device 10.
[0017] <Physical configuration> Figure 2 shows an example of the physical configuration of each device constituting the voice processing system 1 according to this embodiment. This is a diagram showing each device (for example, an audio processing device 10, an operator terminal 20, and a customer terminal 30). ) consists of a processor 10a which corresponds to the arithmetic unit and RAM (Random Access) which corresponds to the storage unit. Memory) 10b, ROM (Read Only Memory) 10c which corresponds to the storage unit, and communication unit 10 d, input unit 10e, display unit 10f, camera 10g, audio input unit 10h, audio output It has a section 10i and a component. Each of these components is connected to each other via a bus so that data can be sent and received from each other. The configuration shown in Figure 2 is just one example, and each device may have other configurations. Furthermore, it is not necessary to have some of these components.
[0018] Processor 10a is, for example, a CPU (Central Processing Unit). 10a executes a program stored in RAM10b or ROM10c. The processor 10a controls various processes in each device. Through collaboration with other configurations and programs, the functions of each device are realized and processing is executed. Control. The processor 10a receives various data from the input unit 10e and the communication unit 10d. The calculation results of the data are then displayed on the display unit 10f or stored in the RAM 10b.
[0019] RAM10b and ROM10c store data necessary for various processes and data of processing results. It is a memory unit. In addition to RAM10b and ROM10c, each device has a hard disk It may also be equipped with a large-capacity storage unit such as a drive. RAM10b and ROM10c are, for example, It may be composed of semiconductor memory elements.
[0020] The communication unit 10d is an interface that connects each device to other devices. It communicates with other devices. The input unit 10e accepts data input from the user. This is a device for inputting data from outside the device or each device. The input unit 10e is For example, it may include a keyboard, mouse, and touch panel. The display unit 10f is a professional This is a device that displays information according to the control of the separator 10a. The display unit 10f is, for example, For example, it may be composed of an LCD (Liquid Crystal Display).
[0021] The camera 10g includes an image sensor for capturing still or moving images, and captures a predetermined area. The system generates captured images (for example, still images or moving images). The audio input unit 10h receives audio. This is a sound-collecting device, such as a microphone. The audio output unit 10i outputs sound. It is a device, for example, a speaker.
[0022] The programs that run each device are computer programs such as RAM10b and ROM10c. It may be stored and provided on a storage medium readable by the data, or by the communication unit 10d It may be provided via a connected network 40. Each device has a processor 10 When a executes the program, various operations for controlling each device are realized. These physical configurations are examples only and do not necessarily have to be independent configurations. For example, each device has a processor 10a and RAM 10b and ROM 10c integrated into one L It may also include SI (Large-Scale Integration).
[0023] <Functional configuration> ≪Sound Processing Device≫ Figure 3 shows an example of the functional configuration of the audio processing device 10 according to this embodiment. The processing device 10 includes a storage unit 101, a transmitting / receiving unit 102, a voice recognition unit 103, a removal unit 104, and a voice Synthesis unit 105, emotion recognition unit 106, stress recognition unit 107, control unit 108, learning unit 109 Includes.
[0024] The memory unit 101 stores various information, programs, algorithms, models, operation logs, etc. Specifically, the memory unit 101 contains the speech recognition model 101a and the speech synthesis model 1, which will be described later. 01b, Emotion Recognition Model 101c, Stress Recognition Model 101d, Emotion Inhibition Switching Model 1 Store 01e, etc.
[0025] The transmitting / receiving unit 102 transmits various information between the operator terminal 20 and / or the customer terminal 30. It transmits and / or receives signals. For example, the transmitting / receiving unit 102 (acquisition unit) receives customer The terminal 30 acquires the speech signal, which is the signal of the customer's spoken voice. Transceiver 1 02 transmits a synthesized voice signal and / or a spoken voice signal to the operator terminal 20. Furthermore, the transmitting / receiving unit 102 acquires operation logs from the operator terminal 20. It is also acceptable. The operation log may contain information about the operator's subjective evaluation of the customer's emotions (hereinafter, This includes "subjective evaluation information," "stress level" (described later), and "manual switching history" (described later). The data may include "data". Also, the transmitting / receiving unit 102 sends customer information to the operator terminal 20. You may transmit information related to your emotions (hereinafter referred to as "emotional information"), etc.
[0026] The speech recognition unit 103 extracts based on the spoken speech signal acquired by the transmitting / receiving unit 102. The feature vectors (hereinafter referred to as "speech feature vectors") are input to the speech recognition model 101a, and one or more It generates text data containing a sequence of words. Specifically, the speech recognition unit 103 Then, using the acoustic model of the speech recognition model 101a, a word sequence is generated from the above speech features, The above text data may be generated according to the results of word sequence analysis using a word model. The recognition unit 103 performs preprocessing on the spoken audio signal (for example, digitization of the analog signal, Audio features may be extracted by performing noise reduction, Fourier transform, etc.
[0027] The speech recognition model 101a is an algorithm that estimates the content of speech based on the speech signal. Yes. The speech recognition model 101a determines what sounds a given word is likely to appear as. Acoustic models that model this, and / or how many times a certain word sequence is used in a particular language It may include a language model that models the probability of an expression appearing. An example of an acoustic model is... For example, Hidden Markov Models (HMMs) and / or Deep Nucleo A Deep Neural Network (DNN) may be used. For example, probabilistic language models such as n-gram language models may be used.
[0028] The removal unit 104 removes specific words contained in the text data generated by the speech recognition unit 103. The system detects a sequence of words and removes the specific sequence of words or replaces it with another sequence of words. The speech data is generated and output to the speech synthesis unit 105. The removal unit 104 is the speech recognition unit 103 If a specific word sequence is not detected in the text data generated by, then the text data The output may be sent to the speech synthesis unit 105.
[0029] The particular string of words in question may, for example, insult the listener or deny the listener's character. This may include one or more words that cause discomfort or other psychological harm to the listener. Here, each word is a noun, verb, adverb, particle, adjective, auxiliary verb, etc., This may include variations of the part of speech that have undergone sound changes. For example, a certain sequence of words could be "I'll kill you." It could be a sentence like "I'll do it," or "I'm telling you this is troublesome," like "I'm telling you this." It may also be "part of a sentence" that indicates abusive language. The removal unit 104 is text Text data in which only specific word sequences detected within the stock data are replaced with other word sequences. The speech synthesis unit 105 may output the entire sentence containing the specific word sequence to another word The text data, which has been replaced with columns, may be output to the speech synthesis unit 105. The other word sequences are You can also use spaces, etc.
[0030] The removal unit 104 removes text data based on a specific word sequence pre-stored in the storage unit 101. The system may detect and / or replace specific word sequences within the data with other word sequences.
[0031] Alternatively, the removal unit 104 removes text data based on a model learned by machine learning. Detection of specific word sequences within a text, and / or replacement with other word sequences that mitigate semantic sentiment. This may be done. For example, a specific sequence of words in the text data, "omae", can be changed to "anata". It may be replaced. Based on a machine learning-based model, a specific word in the text data may be replaced. Column detection and / or replacement with other word sequences may be performed.
[0032] Furthermore, if a specific sequence of words is detected in the text data, the removal unit 104 will remove the specific Information regarding the detection of a sequence of words (hereinafter referred to as "detection information") may be generated. The information includes, for example, information indicating that a particular word sequence has been detected (e.g., "NG word"). The string "" or "NG word detection"), information indicating the specific word sequence, and the customer It may include at least one piece of information regarding a warning (hereinafter referred to as "warning information"). The warning information in question may include, for example, cases where a customer's statements to an operator constitute insult, defamation, etc. The information may also be used to notify that it may be subject to criminal charges. The detected information is transmitted and received. The information may be transmitted to the operator terminal 20 by the signal unit 102. The voice processing device 10 sends warning information to the customer terminal 30 (for example, "To our operator...") This could lead to charges of insult, etc. We acknowledge that we may have made some mistakes, but please contact our operator. This may place an excessive burden on you, so we would appreciate your cooperation. Good. This type of warning information can be used as advance notice against customer harassment. It is possible.
[0033] The speech synthesis unit 105 extracts based on the text data input from the removal unit 104. The features (hereinafter referred to as "text features") are input into the speech synthesis model 101b, The synthesized voice signal (hereinafter referred to as the "synthesized voice signal") is generated. Specifically, the removal unit 104 It predicts speech synthesis parameters based on text features, and the predicted speech synthesis parameters A synthesized voice signal may be generated using a data processor. The voice synthesis unit 105 sends and receives the synthesized voice signal. Output to signal unit 102. The synthesized speech signal is the speech signal of the content of the text data being read aloud. That could also be said.
[0034] The speech synthesis model 101b takes text data as input and processes the content of that text data. This is an algorithm that outputs a corresponding synthesized speech signal. The speech synthesis model 101b is as follows: For example, the above-mentioned HMM and / or DNN may be used.
[0035] The speech synthesis model 101b may support multiple speech types. Speech synthesis unit 105 This involves selecting the voice type to be used for the synthesized voice signal from among several voice types, and the selected voice type The text data is input into the speech synthesis model 101b, and the selected speech type is synthesized. The signals may be combined. The multiple types of voices may include, for example, voices with little intonation, machine sounds, and k It may be at least one of the following: the voice of a character, the voice of a celebrity, and the voice of a voice actor. The voice synthesis unit 105 receives a selection of voice type from the operator via the operator terminal 20. You can leave it.
[0036] Figure 4 shows an example of the generation of a synthesized speech signal according to this embodiment. In Figure 4, transmission and reception Based on the speech signals S1 to S3 acquired by the signal unit 102, the speech recognition unit 103 It is assumed that text data T1 to T3 are generated. For example, in Figure 4, the removal unit 104 is Since no specific word sequence is detected within the text data T1, the text data T1 is left as is. The output is sent to the speech synthesis unit 105. Meanwhile, the removal unit 104 removes the text data T2 and T3. This detects specific word sequences (e.g., "I'm gonna kill you" in T2, and "I'm telling you" in T3). Therefore, the text data T2' and T3', from which the specific word sequence has been removed or replaced, are used for speech synthesis. Output to section 105. For example, in text data T2', specific information within text data T2 The sequence of words will be replaced with a space (□). Also, in text data T3', the text data The specific word sequence "ttsutten" in T3 is replaced with "to iu". The speech synthesis unit 105, Synthesized speech signals S1, S2', and S3' are generated from text data T1, T2, and T3, respectively. Generate.
[0037] The emotion recognition unit 106 receives the spoken audio signal acquired by the transmission / reception unit 102, and the speech recognition unit 103... The generated text data and subjective evaluation information received by the transmitting / receiving unit 102 are of a small magnitude. Based on this, customer emotion information is generated. The emotion recognition unit 106 receives the spoken voice signal. Based on the extracted audio features (such as intonation and volume), customer emotional information is generated. It is permissible. The emotion recognition unit 106 processes the text data generated based on the spoken audio signal. Whether a specific word sequence was detected, or whether a specific word sequence was not detected for a predetermined period of time or longer. Customer emotion information may be generated based on this. The emotion recognition unit 106 acquires with camera 10g. Emotional information of the customer may be generated based on the captured images of the customer. Emotional recognition unit 106 The emotion recognition model 101c may be used to generate customer emotion information.
[0038] The emotion recognition model 101c uses a spoken speech signal and speech features extracted from that speech signal. , text data, text features, or at least the same as the speech signal generated from the utterance. The input is a combination of two elements, and the emotional information is the customer's emotion corresponding to the spoken audio signal. This is a model that outputs [the output shown].
[0039] Figure 5A is an explanatory diagram of the learning process of emotion recognition model 101c. For example, emotion recognition model For training 101c, speech features extracted from spoken speech signals and text data are used. The text features, and the "subjective evaluation information" (or subjective evaluation information) by the operator. Multiple sets of data (hereinafter) each containing at least one of the features extracted from the report You may use a "dataset". Subjective evaluation information is obtained from the operator's conversation with the customer. This information is based on a subjective evaluation of customer emotions by listening to audio signals. For example, anger levels 1-10. The operator may also assess the customer's anger at multiple levels. The dataset for training the recognition model 101c can be generated as follows, for example. i. The operator listens to the customer's raw speech signal and estimates from that speech signal Annotating customer emotions (i.e., adding "subjective evaluation information" to spoken audio signals) (to assign). This allows the spoken voice signal and the customer's emotions estimated from the said spoken voice signal to be associated with the This provides information that is linked on the time axis. Multiple operators can access multiple speech signals. In contrast, by adding subjective evaluation information, such a collection of information becomes a dataset. This is obtained. The emotion recognition model 101c uses such a dataset to perform supervised machine learning. It may be trained. Note that the dataset used to train the emotion recognition model 101c is In addition to or instead of speech features, the utterance signal may be included, or in addition to text features. Alternatively, text data may be included instead.
[0040] Figure 5B is an explanatory diagram of the estimation process using the emotion recognition model 101c. For example, in Figure 5B As shown, speech features extracted from the utterance speech signal S1, and / or the utterance speech signal Text features extracted from text data T1 generated from S1 are used in the emotion recognition model 10 By inputting to 1c, the output corresponding to the input, i.e., the emotion corresponding to the spoken audio signal, is generated. Information can be obtained. In addition, the emotion recognition model 101c uses speech features in addition to or instead of speech features. A spoken audio signal S1 may be input, or text data may be input in addition to or instead of text features. T1 may be entered.
[0041] Subjective evaluation information includes one or more emotions (e.g., "happiness," "surprise," "fear," "anger"). It may also be a numerical representation of the degree of at least one of the following: "disgust" and "sadness." Alternatively, emotional information can indicate a specific emotion that the customer is likely to be feeling (e.g., "anger"). It may also be something to show.
[0042] The stress recognition unit 107 collects information regarding the operator's stress status (hereinafter referred to as "stress"). It generates information (called "information"). For example, the stress recognition unit 107 generates the operator's heart rate, Vital data such as sweat volume and respiration volume, or the operator's gaze collected using a camera. Based on image information such as facial expressions, the operator's stress level is assessed using conventionally known methods. It is acceptable to make an estimate. For example, the stress recognition unit 107, based on the operator's speech, It may be possible to estimate the operator's stress level. Specifically, the stress recognition unit 107 estimates the operator's stress level. Changes in the tone and speed of the lectern's speech, the appearance of words related to apologies, and overlapping with the customer's statements It is permissible to estimate the operator's stress level based on what they say, etc. For example, stress The stress recognition unit 107 recognizes the operator's stress status based on the operation log of the operator terminal 20. It is possible to estimate this. Specifically, the stress recognition unit 107 will estimate the movement of the mouse, etc., and the operation to be performed. The operator's stress level can be estimated based on factors such as the absence of user input in a given situation. The stress recognition unit 107 generates stress information based on the stress recognition model 101d. Good. The stress recognition model 101d uses a speech voice signal and sounds extracted from that speech voice signal. Voice features, text data generated from the speech signal, text features, or these The input consists of at least two combinations, and the operator listening to the utterance perceives... This is a model that outputs an estimated stress level. Training the stress recognition model 101d involves... You may use the actual measured stress levels that operators experienced when listening to customer speech. The dataset for training the recognition model 101d can be generated as follows, for example. Good. The operator can rate the level of stress they felt after listening to the customer's speech (for example, 1-10). Annotating (at a level like this) (i.e., the " (Assigns a "stress level"). This allows the spoken audio signal and the spoken audio signal to be heard. Information is obtained that correlates the operator's stress at the time with the time axis. The illustrator assigns a stress level to multiple speech signals, thus A dataset containing a collection of such information is obtained. The stress recognition model 101d is as follows: Supervised machine learning may be used with any dataset.
[0043] The control unit 108 performs various controls related to the audio processing device 10. Specifically, the control unit 1 08 is based on the stress information generated in the stress recognition unit 107, and the operator In terminal 20, either the synthesized voice generated by the speech synthesis unit 105 or the customer's spoken voice is selected. The output can be switched on or off. The control unit 108 synthesizes a speech signal based on the spoken speech signal. Whether or not to generate it may be switched based on stress information. For example, the control unit 108 If the stress level indicated by the stress information is above or greater than a predetermined threshold, the customer's speech sounds The control system may be controlled to output synthesized speech to the operator instead of a real voice. Meanwhile, the control unit 108 If the stress level indicated by the stress information is less than or equal to a predetermined threshold, then the utterance The control unit 108 may be controlled to output audio to the operator. If instruction information regarding the automatic switching of the emotional suppression function is entered, based on stress information... The above switching may be performed. The emotion suppression function replaces the customer's spoken voice with synthesized speech. This function outputs information to the operator.
[0044] The control unit 108 may perform the above switching based on emotional information. This switching may be performed based on the output of the emotion suppression switching model 101e. Replacement model 101e is for speech audio signals, speech features, text data, text features or These two or more combinations are used as input to switch the emotion suppression function on and off. This model outputs the timing of when to release. The emotion suppression switching model 101e further reduces stress Information or emotional information may be used as input. Details of the emotion suppression switching model 101e will be provided later. To state.
[0045] Furthermore, the control unit 108 switches based on the switching information input by the operator. You may make a change. Here, the switching information is the application (on) of the customer's emotion suppression function or This is information regarding the switching to non-applicable (off). For example, the control unit 108 receives information regarding the switching If the report indicates the application of the customer's emotional suppression function, the system will control the output of synthesized speech to the operator. This may be done. On the other hand, the control unit 108 indicates that the switching information indicates that the customer's emotion suppression function is not applied. In some cases, the system may be controlled to output the spoken voice to the operator. The control unit 108 controls the operator If instruction information regarding the manual switching of the emotion suppression function is entered from the data, the above switching will occur. You may make the above switch based on the information provided.
[0046] The learning unit 109 includes the emotion recognition model 101c, the stress recognition model 101d, and emotion suppression. You may perform the training process for the switching model 101e.
[0047] The audio processing device 10 receives any of the information shown in 1) to 7) below, or at least two of them. The combination of information is associated on a time axis, and transmitted via the transmitting / receiving unit 102 to the operator terminal 2 It is acceptable to send to 0. 1) Customer's spoken voice signal, 2) Data generated from the spoken voice signal 3) Text data after processing by the removal unit 104, 4) Detection information, 5 ) Synthesized speech signal, 6) Customer emotion information estimated from customer speech signal, 7) Emotional suppression The timing for switching the function on and off. When the emotion suppression function is on, the voice processing function The device 10 does not need to send the customer's spoken voice signal to the operator terminal 20. Emotional suppression function If it is turned off, the voice processing device 10 does not need to send a synthesized voice signal to the operator terminal 20. Good. Regardless of whether the emotion suppression function is on or off, the voice processing device 10 processes the customer's spoken voice signal and Both the synthesized voice signal and the synthesized voice signal may be sent to the operator terminal 20.
[0048] Operator terminal
[0049] Figure 6 shows an example of the functional configuration of the operator terminal according to this embodiment. The terminal 20 includes a transmitting / receiving unit 201, an input receiving unit 202, and a control unit 203. (See Figure 6) The functional configuration shown is merely an example, and other configurations not shown may also be included.
[0050] The transmitting / receiving unit 201 transmits various information between the voice processing device 10 and / or the customer terminal 30. It transmits and / or receives signals. For example, the transmitting / receiving unit 201 receives at the customer terminal 30. The transceiver 102 may receive a speech signal, which is a signal of the customer's spoken voice. The audio processing device 10 may also receive a synthesized speech signal. The transmitting / receiving unit 201 also receives sound Subjective evaluation information may be transmitted to the voice processing device 10. Furthermore, the transmitting / receiving unit 201, The voice processing device 10 may receive customer emotion information.
[0051] The input receiving unit 202 receives various information based on the operator's operation of the input unit 10e. It accepts input. For example, the input receiving unit 202 accepts the emotion recognition model 101c and stress recognition. As part of the work to generate a dataset for training the recognition model 101d, Even if subjective evaluation information and stress levels are input for the raw speech signals of customers... Good. From now on, the operator will use the operator terminal 20 to view subjective evaluation information and stress levels. The process of entering the data is called "annotation work." Annotation work is a normal process. It may be positioned as a separate operation from the call center operations. Also, the input reception unit 20 2 may accept input of information regarding the switching of the customer's emotional suppression function. Also, input reception unit 202 is instruction information that instructs whether to manually or automatically switch the emotional regulation function. You may accept this input.
[0052] The control unit 203 performs various controls related to the operator terminal 20. For example, the control unit 20 3 controls the display of information and / or images in the display unit 10f. Also, control unit 203 The control unit 203 controls the audio output in the audio output unit 10i. The audio output may be controlled based on the information transmitted from 0, or the input receiving unit 202 may receive The audio output may be controlled based on the information provided.
[0053] The control unit 203 generates synthesized speech based on the synthesized speech signal received from the speech processing device 10. The voice output unit 10i outputs the voice. The control unit 203 receives the voice signal from the customer terminal 30. The spoken audio may then be output from the audio output unit 10i.
[0054] Furthermore, the control unit 203 generates synthesized speech based on the emotion information received from the speech processing device 10. The emotional information corresponding to the signal may be displayed on the display unit 10f. In addition, the control unit 203 controls the sound Text data corresponding to the synthesized speech signal received from the voice processing device 10 is displayed on the display unit 10f. It may be shown. For example, the control unit 203 displays a small amount of emotion information, text data and detection information. Even if there is none, screen D1 containing one may be displayed on the display unit 10f. Also, the control unit 203 Stress information may also be displayed on the display unit 10f. For example, the control unit 203 displays stress The screen D2 containing the information may be displayed on the display unit 10f.
[0055] Figure 7 shows an example of screen D1 according to this embodiment. As shown in Figure 7, screen D In step 1, the control unit 203 synchronizes with the output timing T of the synthesized speech from the audio output unit 10i. Additionally, emotional information I1 may be displayed on the display unit 10f. Synthesized speech output timing T By displaying emotional information I1 each time, the operator can control the customer's emotional state through the emotional suppression function. Even when listening to emotionally suppressed synthesized speech, it is possible to recognize the customer's emotions in real time. Cut.
[0056] Furthermore, in screen D1, the control unit 203 adjusts the output timing T of the synthesized speech accordingly. The content of the text data I2 corresponding to the synthesized speech may be displayed on the display unit 10f. i. By displaying the contents of text data I2, the operator can use synthesized speech alone. Furthermore, it becomes possible to understand the customer's speech visually.
[0057] Furthermore, on screen D1, the control unit 203, based on the detection information received from the audio processing device 10, Instead of displaying the specific word sequence itself, information I3 (for example) indicates the detection of the specific word sequence. Alternatively, the message "NG word detected" may be displayed on the display unit 10f. This is called the "hide function." This allows the content of customer statements that have a psychologically negative impact to remain hidden. This reduces operator stress by avoiding the need for the operator to recognize the system. The operator can be notified that such an utterance has occurred, allowing the operator to respond to the customer appropriately. It can be done promptly.
[0058] Furthermore, in screen D1, the control unit 203, based on emotional information from the voice processing device 10, Then, for each output timing T of the synthesized voice, the customer's specific emotional level I4 is displayed in chronological order. It may also be displayed at 10f. For example, in Figure 7, the customer at each output timing T of the synthesized voice. The level of "anger" (level I4) is shown as a line graph. This allows the operator to identify the customer. Because it is easy to grasp the transition of emotions (for example, "anger"), the operator can easily understand the customer's feelings. Customer satisfaction with the service can be improved.
[0059] On screen D1, the control unit 203 may display the selection button I5 on the display unit 10f. . The selection button I5 automatically or manually turns the emotion suppression function on or off. This interface allows the operator to choose whether to switch between the two. By performing operations such as clicking, tapping, or sliding on the selection button I5, "self It is possible to switch between "automatic switching mode" and "manual switching mode". For example, emotional information, stress information, or output from the emotion suppression switching model 101e, etc. Based on this, the emotional regulation function is automatically switched on and off.
[0060] When "Manual Switching Mode" is selected, the control unit 203 applies or deapplies the emotion suppression function. The toggle button I6, which is an interface that allows the operator to select the function, is displayed on the 10f It is acceptable to display it there. The timing at which the operator switches the emotion suppression function on and off is: Customer speech (and / or various features extracted based on speech) and related time axis The data is linked together and stored in a memory unit (not shown) as "manual switching history data". The "historical data" may also be associated with operator identification information.
[0061] Button I7 is used to toggle the "NG word hiding function" on and off. It is Tan. If the "NG word hiding function" is turned off, the text data I2 contains special Even if a specific word sequence is detected, the text data before processing by the removal unit 104 is processed The I2 is displayed directly on display unit 10f. The emotion suppression function is turned on, but the forbidden word is not displayed. If the display function is turned off, the operator will not directly hear any specific word sequences from the customer. This reduces stress, while also allowing for accurate understanding of what the customer is saying, thus improving customer perception. It allows for a more accurate understanding of emotions.
[0062] The dataset for training the emotion inhibition switching model 101e includes stress information and emotional information. Report, spoken speech signal S1, speech features, text data, text features, or a small amount thereof At least two combinations, and the timing when the operator switches the emotion suppression function on and off. The ng and can be a bundle of data associated on a time axis. Emotional Inhibition Switching Model 10 There are various ways to learn 1e, such as those described in 1) to 3) below. 1) The emotion inhibition switching model 101e may be learned for each operator. That is, a certain operator The emotion suppression switching model 101e applied to the user is the emotion suppression by the operator. It may be possible to learn based solely on the "manual switching history data" of the function. According to this method, The emotion suppression switching model 101e allows the operator to suppress emotions at a timing that suits their preferences. It becomes possible to switch between these. Alternatively, 2) It can be applied to a certain operator. The emotion suppression switching model 101e uses "manual switching history data" by an unspecified number of operators. The learning may be based on this. According to this method, the data that can be used for learning is Because the amount of data increases, the emotion inhibition switching model 101e can learn more quickly. Or, 3) an emotion suppression switching model 101e applied to a certain operator is, Manual switching history data by operators with similar age, gender, and other characteristics to the operator. You may learn based on this. This method allows you to learn compared to method 1). Because there is more data available, the emotion inhibition switching model 101e can be trained more quickly. Compared to method 2), you can learn the switching timing that suits your preferences. It becomes u.
[0063] Figure 8 shows an example of screen D2 according to this embodiment. In screen D2, the control unit 203 may display stress information from the voice processing device 10. For example, in Figure 8 , stress information includes information that shows an estimated value of the stress felt by the operator (for example, "5 6%) and information indicating the relative evaluation value from the operator's normal state (for example, The message displayed is "8.1% decrease compared to normal times."
[0064] Figure 12 shows an example of screen D3 according to this embodiment. In screen D3, control Section 203 displays interface I8 for the operator to perform annotation work. They may do so. The operator may, for example, listen to the customer's raw voice (sample voice), The customer's emotions, as perceived from the sample audio, are selected each time from interface I8. (Figure 1) In section 2, customer sentiment I1 is subjective evaluation information of customer sentiment by the operator. For example, The operator responded to the sample voice "Please deliver it somehow by this evening" with "Angry" If we annotate the emotion "ri", then as shown in Figure 12, it means "by this evening The sample audio "Please deliver it somehow" and the information "anger" are linked on a timeline. Annotation can be done on a sentence-by-sentence basis or at predetermined time intervals. stomach.
[0065] (Operation of the voice processing system) Figure 9 is a flowchart illustrating an example of emotion suppression behavior according to this embodiment. 9 is merely an example, and the order of at least some of the steps (for example, step S106) is Steps may be rearranged, steps not shown may be performed, or some steps may be omitted. It may be abbreviated.
[0066] The voice processing device 10 processes the customer's spoken voice picked up by the voice input unit 10h of the customer terminal 30. The speech signal, which is a signal, is acquired (S101).
[0067] The speech processing device 10 extracts features based on the spoken speech signal acquired in S101. The text data containing a sequence of one or more words is input to the speech recognition model 101a. Generate the data (S102).
[0068] The speech processing device 10 contains a specific sequence of words within the text data generated in S102. Determine whether or not (S103). If the text data contains a specific sequence of words. The audio processing device 10 removes the specific word sequence or changes the specific word sequence to another word sequence. The converted text data is generated (S104).
[0069] The speech processing device 10 extracts features based on text data and uses them in the speech synthesis model 1 The signal is input to 01b to generate a synthesized speech signal (S105).
[0070] The speech processing device 10 receives the spoken speech signal acquired in S101 and the text generated in S102. The data, and the subjective evaluation information of customer emotions entered by operators, are of a small size. The features extracted based on one of these are input into the emotion recognition model 101c to understand the customer's emotions. Generate information (S106).
[0071] The operator terminal 20 outputs synthesized speech based on the synthesized speech signal generated in S105. The output is made from the output unit 10i, and the output timing T of the synthesized voice is synchronized with the output timing T. Emotional information corresponding to the synthesized speech is displayed on the display unit 10f (S107, for example, Figure 7).
[0072] The audio processing device 10 determines whether or not to terminate the process (S108). If not (S108: NO), the voice processing device 10 will execute processes S101 to S107 again. On the other hand, when the voice conversion process is terminated (S108: YES), the voice processing device 10 processes End of discussion.
[0073] Figure 10 is a flowchart showing the automatic switching operation of the emotion suppression function according to this embodiment. Yes. Note that Figure 10 is merely an example, and the order of at least some of the steps can be changed. It is also possible that steps not shown be performed, or that some steps be omitted. stomach.
[0074] The voice processing device 10 generates operator stress information (S201).
[0075] The voice processing device 10 determines whether the stress information meets predetermined conditions (S202 ). For example, a given condition is that the stress level indicated by the stress information is above a given threshold or more It can be something big.
[0076] The voice processing device 10, if the stress information meets a predetermined condition (S202: YES), The emotion suppression function may be applied (i.e., synthesized speech may be output from the operator terminal 20). S203). On the other hand, if the stress information does not meet the predetermined conditions, the voice processing device 10 ( S202:NO), emotion suppression function not applied (i.e., customer communication from operator terminal 20) You may output spoken audio (S204).
[0077] The audio processing device 10 determines whether or not to terminate the process (S205). If not (S205: NO), the voice processing device 10 will execute processes S201 to S204 again. On the other hand, when the voice conversion process is terminated (S205:YES), the voice processing device 10 processes The process is terminated. Note that in S201 and S202, the voice processing device 10 receives emotional information and Based on the output of the emotion suppression switching model 101e, it is decided whether or not to apply the emotion suppression function. That's good too.
[0078] As described above, according to the voice processing system 1 of this embodiment, the customer's spoken voice signal is processed Based on this, text data is generated, and based on this text data, synthesized speech is generated. Output to the operator. Therefore, the customer's emotions contained in the customer's spoken voice must be sufficiently suppressed. The synthesized voice can be played to the operator, and the operator can hear the customer's emotional speech. This can reduce stress. The inventor of this invention conducted a study with approximately 50 subjects, 1) Customer 1) The spoken audio itself, 2) Audio with the volume of the customer's spoken voice adjusted, 3) The voice quality of the customer's spoken voice. There are four types: 1) the converted audio, 2) synthesized speech generated after the customer's spoken audio is transcribed into text, and 3) synthesized speech generated from that text. Participants were asked to compare the audio recordings and rate the degree of anger they perceived from the recordings on a 7-point scale. An experiment was conducted to see how much anger was conveyed to the subjects. As a result, compared to 2) and 3), 4) conveyed more anger to the subjects. The degree of reduction was significant.
[0079] Furthermore, according to the voice processing system 1 of this embodiment, synthesized voice is provided to the operator. In addition to outputting data, it also notifies customers of their emotional state in sync with the timing of the synthesized voice output. This allows operators who hear the synthesized voice to recognize the customer's emotions in real time. They can provide appropriate service to customers.
[0080] Furthermore, according to the voice processing system 1 of this embodiment, operator stress information or Based on customer emotional information, etc., whether or not to apply the emotional suppression function (i.e., to the operator) In contrast, the operator can switch between outputting synthesized speech or spoken speech. This allows for an appropriate balance between employee stress and customer satisfaction.
[0081] (Example of change) In the above-described voice processing system 1, the voice recognition unit 103 extracts one or more speech signals from the spoken voice signal. Text data containing, but not limited to, sequences of words that have been confirmed as a sentence has been generated. The voice recognition unit 103 determines that the sequence of words recognized from the spoken audio signal constitutes one or more sentences. Before being processed, text data containing a sequence of words (parts of speech or morphemes) consisting of one or more words. The following may be generated. The removal unit 104 removes text data that has not been determined to be the sentence. The speech synthesis unit 105 removes specific word sequences and then processes the text data that has not been determined to be part of the sentence. A synthesized voice signal may be generated from the data.
[0082] Figure 11 shows an example of the generation of a synthesized speech signal according to a modified version of this embodiment. In step 1, based on the speech voice signal S4 acquired by the transmitting / receiving unit 102, the speech recognition unit 103 receives the signal. Text data T41~T43 will be generated in the following manner. As shown in Figure 11, The Kist data T41-T43 were created before the sentence "Please send it quickly" was finalized, meaning Text data is generated using morphemes ("quickly", "send", "please") that have the following properties. In this respect, it differs from Figure 4. The removal unit 104 is applied to each of the text data T41 to T43. Then, it determines whether a specific word sequence is included, and removes that specific word sequence before proceeding to the speech synthesis unit. Output to 105. The speech synthesis unit 105 combines the text data T41 to T43 respectively. The generated audio signals S41 to S43 are produced.
[0083] As shown in Figure 11, text data is generated in units of one or more morphemes before the sentence is finalized. By outputting synthesized speech, the operator's response delay is reduced due to the generation of text data. This can reduce the problem. Furthermore, multiple text data (or synthesized speech) at the morphological level can be semantically... Models that determine whether something is natural or not may also be used.
[0084] Furthermore, in order to reduce response delay, the synthesized speech signals S1-S3 shown in Figure 4 and Figure 11 are used. Before and / or after each of the synthesized speech signals S41 to S43, for example, "ah", "eh", etc. Filler sounds such as "well" may be added. This also reduces the response delay for the operator. This can prevent a decline in customer satisfaction.
[0085] Furthermore, the speech synthesis unit 105 generates multiple voice synthesis signals based on the customer's emotions estimated by the emotion recognition unit 106. Select the speech synthesis model 101b that best matches the customer's emotions from among the available speech synthesis models 101b. It is also possible to do so. For example, if the emotion recognition unit 106 estimates the customer's emotion to be "furious", the sound The voice synthesis unit 105 may use a voice synthesis model 101b with a fast pitch and strong intonation. For example, if the emotion recognition unit 106 estimates the customer's emotion as "sobbing," the speech synthesis unit 105 Alternatively, a speech synthesis model 101b that outputs a crying-like sound may be used. The synthesis unit 105 creates a speech synthesis model 10 based on the customer's emotions estimated by the emotion recognition unit 106. You may change the parameters of 1b to adjust the output of voice that matches the customer's emotions. Operators who directly hear the raw audio of a customer in a fit of rage experience extremely high levels of stress. On the other hand, operators need to understand the customer's emotions realistically in order to properly perform customer service duties. It is necessary to grasp the timing accurately. By not letting the operator hear the spoken audio directly... Operators do not experience excessive stress and can convey customer emotions through synthesized voices. This allows operators to understand customers' emotions in real time through auditory cues.
[0086] (Other embodiments) In the above embodiment, the customer's spoken voice signal is converted to text, and the synthesized voice signal is sent to the operator. The output is intended to be in this way, but is not limited to this. The voice processing device 10 outputs the customer's spoken voice signal. The extracted speech features are input into the speech conversion model to generate the converted speech signal. The converted voice may be output from the operator terminal 20.
[0087] The "speech conversion model" described in the claims converts the spoken audio signal into text and then... One model outputs as finished speech, while the other outputs by converting the voice quality of the spoken speech signal without converting it to text. It is a concept that encompasses both models, including those that replace the customer's spoken voice with synthesized or converted speech. By outputting audio to the operator, to varying degrees of effectiveness, the operator It can reduce the stress felt. On the other hand, in order to perform customer service tasks, the operator must consider the customer. Understanding customer emotions in real time is also essential.
[0088] In this modified example, the voice processing device 10 extracts voice based on the customer's spoken voice signal. The feature quantities are input to the speech conversion model to generate a converted speech signal. The speech processing device 10 is 1 1) the converted audio signal and 2) the customer's emotional information estimated from the customer's speech are related on the time axis. The voice processing device 10 generates linked information and transmits it to the operator terminal 20. The information to be processed includes the spoken audio signal, the text data generated from the spoken audio signal, and the removal unit 10. After processing step 4, the text data, detection information, and the ability to toggle the emotion suppression function on or off can be accessed. The timing of the action may be associated with it.
[0089] The operator terminal 20 receives the converted audio signal from the audio processing device 10 and outputs it to the audio output unit 1 Output from 0i, and synchronized with the output timing T of the converted audio from the audio output unit 10i. Information indicating emotion may be displayed on the display unit 10f. The operator terminal 20 can also provide voice The text data is displayed on the display unit 1 in accordance with the output timing T of the converted audio from the output unit 10i. It may be displayed as 0f. The manner of such display may be as shown in Figure 7.
[0090] In this modified example, the audio processing device 10 changes the emotion indicated by the emotion information based on the emotion information. The converted speech signal may be generated in such a way that it is reflected in the converted speech. For example, the emotional information that is conveyed When the emotion is "fury," a speech conversion model with a fast pitch and strong intonation may be used. For example... If the emotion information indicates "sobbing," then the voice conversion software will output a sound that resembles crying. You may use Dell. The voice processing device 10 will reflect the emotion indicated by the emotion information in the converted speech. Therefore, it is acceptable to generate a converted speech signal. By not allowing the operator to hear the spoken speech directly... The operators can convey customer emotions through the converted voice without experiencing excessive stress. This allows operators to understand customers' emotions in real time through auditory information.
[0091] In the audio processing system 1 of this modified example, the operator performs annotation work. The work is performed on converted speech during normal call center operations by operators. It is acceptable. If the operator annotates "anger" with the converted speech, Based on the annotation results, the speech conversion model is designed to produce softer speech. It may be adjusted in real time.
[0092] In the embodiments described above, the first user is the customer and the second user is the operator. While this embodiment is envisioned as being applicable to a call center, its application is not limited to call centers. For example, in a web meeting, if the first user's voice, with its suppressed emotions, is transmitted to a second user... It can be applied to any output situation. In other words, this embodiment is applicable to customer harassment. In addition to measures against harassment, we also have a plan to address various forms of harassment, including power harassment within the company. This can be used as a countermeasure by businesses.
[0093] In the embodiments described above, emotional information and synthesized speech are "associated on a time axis". As shown in Figure 7, the processing is synchronized with the output timing of the synthesized or converted speech. In a manner that makes it possible to display emotional information estimated from the original spoken audio, However, the specific form is irrelevant. In the embodiments described above, "association on the time axis" The process of "associating" can be a process that associates based on time information such as what time it is (hours, minutes, and seconds), It would also be acceptable to associate the information based on data such as the number of minutes and seconds elapsed since the start of the spoken audio. Alternatively, the process may involve associating elements at the sentence, word, or morpheme level.
[0094] In the voice processing system 1 of the embodiment described above, the customer can communicate their own voice It is also possible to make it so that the emotions are suppressed and not conveyed to the operator. Customers cannot tell whether the emotional regulation function is turned on or off. You can do that.
[0095] Annotation work can be performed by the operator on the operator terminal 20, or separately. Dedicated applications or terminals may be provided for annotation work.
[0096] Furthermore, the embodiments described above are provided to facilitate understanding of the present invention. This is not intended to limit the interpretation of each element and its arrangement and materials of the embodiment. The conditions, shape, size, etc., are not limited to those exemplified and may be changed as appropriate. This is possible. Furthermore, the configurations shown in different embodiments can be partially substituted or combined. This is possible. In addition, the functions described as those of the voice processing device 10 can be performed on the operator terminal 2 0 may also be present. In addition, the functions described as functions of the operator terminal 20 are voice processing. Device 10 may also be equipped with this. [Explanation of Symbols]
[0097] 1…Voice processing system, 10…Voice processing device, 20…Operator terminal, 30…Customer terminal 10a...Processor, 10b...RAM, 10c...ROM, 10d...Communication Unit, 10e...Input Power unit, 10f...Display unit, 10g...Camera, 10h...Audio input unit, 10i...Audio output unit, 1 01...Storage unit, 102...Transmit / receive unit, 103...Voice recognition unit, 104...Removal unit, 105...Voice Synthesis unit, 106... Emotion recognition unit, 107... Stress recognition unit, 108... Control unit, 109... Learning unit Section 201...Transmitting / receiving section, 202...Input receiving section, 203...Control unit
Claims
1. A voice processing device used for customer service in a call center, An acquisition unit that acquires a speech signal, which is the speech signal of the first user, A speech recognition unit inputs feature quantities extracted based on the aforementioned spoken audio signal into a speech recognition model to generate text data containing a sequence of one or more words. In the aforementioned text data, a specific sequence of words that has a psychologically negative impact on a second user is detected. (A) Replaced text data obtained by replacing the aforementioned specific word sequence with another word sequence that mitigates the semantic sentiment, (B) Removed text data obtained by removing the specific word sequence from the text data. A removal unit that generates, A speech synthesis unit inputs features extracted based on the replacement text data or the removal text data into a speech synthesis model to generate a synthesized speech signal, which is a synthesized speech signal, in such a way as to suppress the emotions of the first user contained in the utterance speech signal. The system comprises a stress recognition unit that generates stress information relating to the stress status of the second user, The synthesized voice or spoken voice is output to the second user. The output of the spoken voice or synthesized voice is controlled to switch based on the generated stress information. Voice processing device.
2. A voice processing device used for customer service in a call center, An acquisition unit that acquires a speech signal, which is the speech signal of the first user, A speech recognition unit inputs feature quantities extracted based on the aforementioned spoken audio signal into a speech recognition model to generate text data containing a sequence of one or more words. In the aforementioned text data, a specific sequence of words that has a psychologically negative impact on a second user is detected. (A) Replaced text data obtained by replacing the aforementioned specific word sequence with another word sequence that mitigates the semantic sentiment, (B) Removed text data obtained by removing the specific word sequence from the text data. A removal unit that generates, A speech synthesis unit inputs features extracted based on the replacement text data or the removal text data into a speech synthesis model to generate a synthesized speech signal, which is a synthesized speech signal, in such a way as to suppress the emotions of the first user contained in the utterance speech signal. Equipped with, The synthesized voice or spoken voice is output to the second user. Based on the switching information input by the second user, the output of either the synthesized voice or the spoken voice is switched, and by inputting the spoken voice signal acquired by the acquisition unit, the feature quantities extracted from the spoken voice signal, the text data generated from the spoken voice signal, the feature quantities extracted from the text data, or at least a combination of these, a timing is generated in which the output of either the synthesized voice or the spoken voice is switched. The system further comprises a learning unit that generates information relating the switching information input by the second user to the spoken speech signal at the time the switching information was input on a time axis, and based on this information, performs machine learning on the emotion suppression switching model, which takes the spoken speech signal, features extracted from the spoken speech signal, text data generated from the spoken speech signal, features extracted from the text data, or at least two combinations thereof as input, and outputs the timing for switching between the synthesized speech and the spoken speech. Voice processing device.
3. A program for causing a computer to function as the audio processing device described in claim 1 or 2.