Audio processing device, audio processing method, and program

Through the acoustic signal processing system, the customer's voice signal is converted into text and a synthetic voice signal is generated, which solves the problem of difficult to suppress customer emotions in the prior art, and effectively reduces operator pressure and improves response.

JP7674798B2Active Publication Date: 2025-05-12SOFTBANK CORPORATION
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2023151074
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2023-09-19
Publication Date
2025-05-12
Estimated Expiration
2041-11-25

AI Technical Summary

Technical Problem

In the prior art, when dealing with customer complaints, it is difficult to effectively suppress the speaker's emotions, resulting in increased pressure on the listener and insufficient response.

Method used

Through the acoustic signal processing system, the acoustic recognition model is used to convert the customer's voice signal into text data, and a synthetic voice signal is generated through the speech synthesis model and output it to the operator, thereby suppressing customer emotions and reducing the pressure on the operator.

Benefits of technology

Effectively reduces the pressure on the operator and ensures that the operator can respond appropriately, improving the quality and efficiency of customer service.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007674798000001
    Figure 0007674798000001
  • Figure 0007674798000002
    Figure 0007674798000002
  • Figure 0007674798000003
    Figure 0007674798000003
Patent Text Reader

Abstract

To enable reduction of stress on a listener.SOLUTION: A speech processing system 1 comprises: an acquisition unit for acquiring a speech signal that is a signal of a first user's speech; a speech recognition unit that inputs feature quantities extracted based on the speech signal into a speech recognition model to generate text data including a word sequence consisting of one or more words; a speech synthesis unit that inputs feature quantities extracted based on the text data to generate a synthesized speech signal that is a signal of synthesized speech; and a speech output unit that outputs the synthesized speech signal to a second user.SELECTED DRAWING: Figure 4
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present invention relates to a voice processing system, a voice processing device, and a voice processing method. [Background technology]

[0002] In the past, in order to improve customer satisfaction (CS), In these cases, various call centers are operated where operators respond to customers by phone. In customer service work, there is a risk of customers making intimidating remarks or unreasonable demands on operators. "Male harassment" can cause mental disorders among operators and lead to high turnover rates among operators. It is becoming a problem that

[0003] In recent years, companies have begun to protect their employees, the operators, from this type of customer harassment. For example, in Patent Document 1, a voice conversion system for converting an input voice signal into a If the volume and pitch fluctuations exceed a predetermined value, the volume And control the volume and pitch to be converted and output so that the pitch fluctuation amount falls within a predetermined range. It is stated that [Prior art documents] [Patent documents]

[0004] [Patent Document 1] JP 2004-252085 A Summary of the Invention [Problem to be solved by the invention]

[0005] However, for example, by simply converting the speaker's speech using the method described in Patent Document 1, In this case, the speaker's (first user's) emotions are not sufficiently suppressed, and the listener's (second user's) stress is increased. On the other hand, in order to reduce the stress of the listener, By converting the speaker's speech output to the hand, the listener can fully recognize the speaker's emotions. There is also a risk that the listener will not be able to respond appropriately.

[0006] Therefore, the present invention provides a method for sufficiently reducing the stress of a listener and / or for enabling the listener to respond appropriately. The present invention provides a voice processing system, a voice processing device, and a voice processing method that enable the above. [Means for solving the problem]

[0007] The voice processing system according to one embodiment of the present invention includes: An acquisition unit that acquires an utterance voice signal, and a feature amount that is extracted based on the utterance voice signal are extracted as a speech signal. A sound is input to a recognition model to generate text data that includes a word sequence of one or more words. A voice recognition unit and a feature quantity extracted based on the text data are input to a voice synthesis model. a speech synthesis unit for generating a synthetic speech signal which is a signal of synthetic speech; and a voice output unit that outputs the synthetic voice.

[0008] According to this aspect, generating text data based on a speech signal of a first user, A synthetic voice generated based on the text data is output to the second user. The synthetic voice in which the emotion of the customer contained in the speech of the first user is sufficiently suppressed is sent to the second user. The second user's stress caused by the emotional utterance of the first user can be heard by the second user. can be sufficiently reduced.

[0009] In the above voice processing system, the emotion recognition unit receives an utterance voice signal, feature vectors extracted from the speech signal, text data generated from the speech signal, The input is a feature quantity extracted from the data, or a combination of at least two of these, and the The emotion recognition model is machine-trained to output emotion information of a speaker of a speech signal. A speech signal acquired by the acquisition unit, a speech feature quantity extracted from the speech signal, and the speech Text data generated from the signal, text features corresponding to the text data, or By inputting at least two of these combinations, the utterance acquired by the acquisition unit Emotion information of the first user corresponding to the audio signal may be generated. [Brief description of the drawings]

[0010] [Figure 1] FIG. 1 is a diagram showing an example of an outline of a speech processing system 1 according to an embodiment of the present invention. [Diagram 2] 1 is a diagram illustrating an example of a physical configuration of each device constituting a speech processing system 1 according to the present embodiment. [Diagram 3] 1 is a diagram illustrating an example of a functional configuration of a voice processing device 10 according to an embodiment of the present invention. [Figure 4] FIG. 4 is a diagram showing an example of generation of a synthetic speech signal according to the present embodiment. [Figure 5A] FIG. 11 is a diagram illustrating an example of generation of customer emotion information according to the embodiment. [Figure 5B] FIG. 11 is a diagram illustrating an example of generation of customer emotion information according to the embodiment. [Figure 6] FIG. 2 is a diagram illustrating an example of a functional configuration of an operator terminal 20 according to the present embodiment. [Figure 7] FIG. 11 is a diagram showing an example of a screen D1 according to the present embodiment. [Figure 8] FIG. 11 is a diagram showing an example of a screen D2 according to the present embodiment. [Figure 9]13 is a flowchart illustrating an example of an emotion suppression operation according to the present embodiment. [Figure 10] 10 is a flowchart showing an automatic switching operation of an emotion suppression function according to the present embodiment. [Figure 11] 13A and 13B are diagrams illustrating an example of generation of a synthetic speech signal according to a modified example of the present embodiment. [Figure 12] FIG. 13 is a diagram showing an example of a screen D3 according to the present embodiment. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0011] An embodiment of the present invention will be described with reference to the accompanying drawings. Those marked with the same symbol have the same or similar configuration.

[0012] Hereinafter, the voice processing system according to this embodiment will be described in detail with reference to the customer service operations of a call center or the like. However, the application of the present invention is not limited to this embodiment. In this embodiment, a predetermined process is performed on a signal of a speech voice of a first user (hereinafter, referred to as an "utterance voice signal"). The present invention can be applied to any situation in which a speech generated by the speech recognition device is output to a second user. In the following, the first user is the customer and the second user is the operator. However, this is not limited to this.

[0013] (Configuration of voice processing system) <Overall composition> FIG. 1 is a diagram showing an example of a schematic configuration of a speech processing system 1 according to the present embodiment. As shown in the figure, the voice processing system 1 includes a voice processing device 10 and a second user (hereinafter, “operator”). a terminal (hereinafter referred to as an "operator terminal") 20 used by the first A terminal (hereinafter referred to as a "Customer Terminal") used by a user (hereinafter referred to as a "Customer") of (c) 30 and

[0014] The voice processing device 10 converts an utterance voice signal acquired by a customer terminal 30 into a The network 40 may be an external network such as the Internet. It is also possible to use external networks and internal networks such as Local Access Networks (LANs). The voice processing device 10 performs a predetermined process on the voice signal of the customer. The voice processing device 10 transmits the voice to the operator terminal 20. It may be composed of a server.

[0015] The operator terminal 20 may be, for example, a telephone, a smartphone, a personal computer, a The operator terminal 20 receives the voice generated by the voice processing device 10 through a predetermined process. Based on the voice signal or the speech signal from the customer terminal 30, a voice is output to the operator. .

[0016] The customer terminal 30 may be, for example, a telephone, a smartphone, a personal computer, a tablet, The customer terminal 30 picks up the customer's voice by a microphone and converts the voice into A speech signal, which is a signal, is transmitted to the speech processing device 10 .

[0017] <Physical configuration> FIG. 2 shows an example of a physical configuration of each device constituting the speech processing system 1 according to this embodiment. Each device (for example, a voice processing device 10, an operator terminal 20, and a customer terminal 30) ) includes a processor 10a corresponding to a calculation unit and a RAM (Random Access Memory) corresponding to a storage unit. A ROM (Read Only Memory) 10b, which corresponds to a storage unit, and a communication unit 10 d, an input unit 10e, a display unit 10f, a camera 10g, an audio input unit 10h, and an audio output Each of these components is connected to each other via a bus so that data can be transmitted and received. Note that the configuration shown in FIG. 2 is an example, and each device may have a configuration other than those shown. However, some of these configurations may not be included.

[0018] The processor 10a is, for example, a CPU (Central Processing Unit). The CPU 10a executes a program stored in the RAM 10b or the ROM 10c. The processor 10a is a control unit that controls various processes in each device. By cooperating with other components and programs, the functions of each device are realized and processing is executed. The processor 10a receives various data from the input unit 10e and the communication unit 10d. The calculation results of the data are displayed on the display unit 10f and stored in the RAM 10b.

[0019] The RAM 10b and the ROM 10c store data necessary for various processes and data of the results of the processes. Each device includes a RAM 10b and a ROM 10c, as well as a hard disk The RAM 10b and the ROM 10c may be, for example, , may be composed of a semiconductor memory element.

[0020] The communication unit 10d is an interface that connects each device to other devices. The input unit 10e is for receiving data input from a user. The input unit 10e is a device for inputting data from a device or an external device of each device. For example, the display unit 10f may include a keyboard, a mouse, a touch panel, and the like. The display unit 10f is a device that displays information according to the control of the processor 10a. For example, it may be configured with an LCD (Liquid Crystal Display).

[0021] The camera 10g includes an image sensor for capturing still or moving images, and is used to capture an image of a predetermined area. The audio input unit 10h receives audio. The audio output unit 10i is a device for collecting audio, such as a microphone. A device, for example a speaker.

[0022] The programs for executing each device are stored in the RAM 10b, the ROM 10c, and other computer memory. The information may be provided by being stored in a storage medium readable by the computer, or may be provided by the communication unit 10d. The information may be provided via a network 40 connected to the processor 10. By executing the program, various operations for controlling each device are realized. Note that these physical configurations are merely examples and do not necessarily have to be independent configurations. For example, each device may be a LSI in which a processor 10a is integrated with a RAM 10b and a ROM 10c. It may also be equipped with SI (Large-Scale Integration).

[0023] <Functional configuration> <Sound processing device> 3 is a diagram showing an example of a functional configuration of the voice processing device 10 according to the present embodiment. The processing device 10 includes a storage unit 101, a transmission / reception unit 102, a voice recognition unit 103, a removal unit 104, and a voice A synthesis unit 105, an emotion recognition unit 106, a stress recognition unit 107, a control unit 108, and a learning unit 109. Includes.

[0024] The storage unit 101 stores various information, programs, algorithms, models, operation logs, etc. Specifically, the storage unit 101 stores a voice recognition model 101a and a voice synthesis model 101b, which will be described later. 01b, emotion recognition model 101c, stress recognition model 101d, emotion inhibition switching model 1 Remember 01e etc.

[0025] The transmission / reception unit 102 transmits various information between the operator terminal 20 and / or the customer terminal 30. For example, the transceiver unit 102 (acquisition unit) transmits and / or receives a customer The terminal 30 acquires a speech signal, which is a signal of the customer's speech, collected by the terminal 30. 02 transmits a synthesized voice signal and / or a spoken voice signal to the operator terminal 20. In addition, the transmitting / receiving unit 102 acquires an operation log by an operator from the operator terminal 20. The operation log may include information on the subjective evaluation of the customer's emotions by the operator (hereinafter, "Subjective evaluation information"), "Stress level" described later, "Manual switching history" described later The transmission / reception unit 102 may transmit the customer data to the operator terminal 20. Information regarding the emotions of the user (hereinafter referred to as "emotion information") may also be transmitted.

[0026] The voice recognition unit 103 extracts a voice signal based on the speech signal acquired by the transmission / reception unit 102. A feature (hereinafter referred to as "speech feature") is input to the speech recognition model 101a, and one or more The speech recognition unit 103 generates text data including a word string consisting of words. Then, a word string is generated from the speech feature quantity using the acoustic model of the speech recognition model 101a, and a linguistic The text data may be generated according to an analysis of a word string using a word model. The recognition unit 103 performs preprocessing on the speech signal (e.g., digitizing an analog signal, Alternatively, speech features may be extracted by performing noise removal, Fourier transform, etc.

[0027] The speech recognition model 101a is an algorithm for estimating the content of a speech based on a speech signal. The speech recognition model 101a is a model that estimates what kind of sound a certain word is likely to appear as. Acoustic models that model how well a word sequence is generated in a particular language and / or The acoustic model may include a language model that models the probability of each word appearing. For example, Hidden Markov Models (HMMs) and / or deep neural networks A Deep Neural Network (DNN) may be used. For example, a probabilistic language model such as an n-gram language model may be used.

[0028] The removal unit 104 removes specific words contained in the text data generated by the speech recognition unit 103. A text obtained by detecting a specific word string, removing the specific word string or replacing the specific word string with another word string, The removal unit 104 generates the sample data and outputs it to the speech synthesis unit 105. If a specific word sequence is not detected in the text data generated by may be output to the speech synthesis unit 105.

[0029] The particular word sequence may, for example, be insulting to the listener or denying the listener's personality. It may be one or more words that have a negative psychological effect on the listener, such as making the listener feel uncomfortable. Here, each word is at least one part of speech such as a noun, verb, adverb, particle, adjective, auxiliary verb, etc. The part of speech may include a sound change. For example, a specific word string may be "You, kill me!" It can be a sentence like "I'll do it" or "tsutsun" like "Konaruttsutsutenno" Alternatively, the removal unit 104 may remove a part of a sentence that indicates rough language. It is a text data in which only specific word strings detected in the text data are replaced with other word strings. Alternatively, the entire sentence including the particular word string may be output to the speech synthesis unit 105. The text data replaced with the other word string may be output to the speech synthesis unit 105. , or blank.

[0030] The removal unit 104 removes the text data based on a specific word string stored in advance in the storage unit 101. Detection of specific word strings in the data and / or replacement with other word strings may also be performed.

[0031] Alternatively, the removal unit 104 may remove the text data based on a model learned by machine learning. Detecting specific word sequences in the data and / or replacing them with other word sequences that have a less semantic sentiment. For example, a specific word string "omae" in the text data may be converted to "anata". Based on a machine learning-based model, certain words in the text data may be replaced. String detection and / or replacement with other word strings may be performed.

[0032] When a specific word string is detected in the text data, the removal unit 104 The system may generate information regarding the detection of the word string (hereinafter, "detection information"). The information may be, for example, information indicating that the particular word string was detected (for example, "NG word" " or "NG word detected"), information indicating the specific word string, and (hereinafter, referred to as "warning information") The warning information may be, for example, information regarding a customer's speech to an operator that is insulting or defamatory. The detection information may be information to notify the sender or receiver that the sender or receiver may be subject to criminal prosecution. The detection information may be transmitted to the operator terminal 20 by the receiving unit 102. The voice processing device 10 then transmits warning information (for example, "To our operator" There is a risk of being charged with defamation. I understand that this is also due to our company's negligence, but This may be an excessive burden on the system, so we would appreciate your cooperation. Such warning information can be used as advance notice against customer harassment. It is possible.

[0033] The voice synthesis unit 105 extracts the voice based on the text data input from the removal unit 104. The feature quantity (hereinafter referred to as "text feature quantity") is input to the speech synthesis model 101b, and Specifically, the removal unit 104 generates a signal of synthesized speech (hereinafter, referred to as a "synthetic speech signal"). predicts speech synthesis parameters based on text features and The voice synthesis unit 105 may generate a synthetic voice signal using the data. The synthesized voice signal is a voice signal that reads out the contents of the text data. It can also be said that.

[0034] The voice synthesis model 101b receives text data as input and performs voice synthesis based on the contents of the text data. The speech synthesis model 101b is an algorithm for outputting a corresponding synthetic speech signal. For example, the above-mentioned HMM and / or DNN may be used.

[0035] The voice synthesis model 101b may correspond to a plurality of voice types. selects a voice type to be used for the synthetic voice signal from among a plurality of voice types, and and text data are input to the voice synthesis model 101b to generate a synthetic voice of the selected voice type. The multiple voice types may be, for example, voices with a flat intonation, mechanical sounds, The voice may be at least one of a character's voice, a celebrity's voice, and a voice actor's voice. The voice synthesis unit 105 receives a selection of a voice type from an operator via the operator terminal 20. It is okay.

[0036] FIG. 4 is a diagram showing an example of generation of a synthetic voice signal according to the present embodiment. Based on the speech signals S1 to S3 acquired by the reception unit 102, the speech recognition unit 103 Assume that text data T1 to T3 are generated. For example, in FIG. 4, the removal unit 104 Since no specific word string is detected in the text data T1, the text data T1 is used as is. The removal unit 104 also outputs the utterances in the text data T2 and T3 to the speech synthesis unit 105. Detect a specific word sequence (T2: "I'm going to kill you," T3: "Ttsutten") Therefore, the text data T2' and T3' in which the specific word string is removed or replaced are used for speech synthesis. For example, in the text data T2′, a specific In the text data T3', the word string is replaced with a space (□). The specific word string "ttsutten" in T3 is replaced with "toiu". Synthesized speech signals S1, S2', and S3' are generated from text data T1, T2, and T3, respectively. Generate.

[0037] The emotion recognition unit 106 recognizes the speech signal acquired by the transmission / reception unit 102 and the voice recognition unit 103. At least the generated text data and the subjective evaluation information received by the transmission / reception unit 102 The emotion recognition unit 106 generates emotion information of the customer based on the speech signal. Based on the extracted voice features (e.g. intonation and volume), customer emotion information is generated. The emotion recognition unit 106 may generate an emotion based on the text data generated based on the speech signal. A specific word string has been detected, or a specific word string has not been detected for a specified period of time. The emotion recognition unit 106 may generate emotion information of the customer based on the emotion acquired by the camera 10g. The emotion recognition unit 106 may generate emotion information of the customer based on the captured image of the customer. may generate emotion information of the customer using the emotion recognition model 101c.

[0038] The emotion recognition model 101c is a speech signal, a speech feature vector extracted from the speech signal, , text data, text feature quantities, or at least these generated from the speech signal The two combinations are input, and emotion information that is the emotion of the customer corresponding to the speech voice signal is obtained. This is a model that outputs

[0039] FIG. 5A is an explanatory diagram of the learning process of the emotion recognition model 101c. For example, The training of 101c involves extracting speech features from speech signals and extracting features from text data. The text features to be analyzed and the "subjective evaluation information" (or "subjective evaluation information") by the operator A set of data (hereinafter referred to as a set of features extracted from the information) each of which contains at least one of the features. The subjective evaluation information may be obtained by using the utterances of customers by the operator. This is information obtained by subjectively evaluating the customer's emotions by listening to the voice signal. For example, anger level 1 to 10 The operator may evaluate the anger of the customer on multiple levels, such as: The data set for training the recognition model 101c may be generated, for example, as follows: The operator listens to the customer's raw speech signal and estimates the customer's speech from the speech signal. Annotate customer emotions (i.e., add "subjective evaluation information" to speech signals) This allows the speech signal and the customer's emotion estimated from the speech signal to be The information obtained is related to the time axis. By adding subjective evaluation information to such a bundle of information, The emotion recognition model 101c uses such a dataset to generate a supervised machine learning model. The dataset used for training the emotion recognition model 101c may be In addition to or instead of the audio features, speech signals may be included, and in addition to the text features, Or alternatively, it may include text data.

[0040] FIG. 5B is an explanatory diagram of an estimation process using the emotion recognition model 101c. For example, in FIG. As shown, the speech feature quantity extracted from the speech signal S1 and / or the speech signal The text features extracted from the text data T1 generated from S1 are used to generate emotion recognition model 10. By inputting the input to 1c, the output corresponding to the input, i.e., the emotion corresponding to the speech signal, is obtained. In addition to or instead of the speech feature, the emotion recognition model 101c can also obtain information. A speech signal S1 may be input, or text data may be input in addition to or instead of the text features. The data T1 may be input.

[0041] The subjective evaluation information may include one or more emotions (e.g., "happiness," "surprise," "fear," "anger" , "disgust," and / or "sadness" may be expressed numerically. Alternatively, emotional information can be used to identify a particular emotion (e.g., "anger") that a customer is likely to be feeling. It may also be something that shows.

[0042] The stress recognition unit 107 recognizes information about the stress state of the operator (hereinafter, “stress For example, the stress recognition unit 107 generates stress information based on the operator's heart rate, Vital data such as sweat rate and breathing rate, or the operator's gaze data collected using a camera The operator's stress state is calculated by a conventional method based on image information such as facial expressions. For example, the stress recognition unit 107 may estimate the stress based on the speech of the operator. The stress recognition unit 107 may estimate the stress state of the operator. The change in tone and speed of the rater's speech, the appearance of words related to apology, and the The stress state of the operator may be estimated based on what the operator says. The stress recognition unit 107 recognizes the stress state of the operator based on the operation log of the operator terminal 20. Specifically, the stress recognition unit 107 may estimate the movement of a mouse or the like, The stress level of the operator may be estimated based on the absence of operational input in the scene. The stress recognition unit 107 generates stress information based on the stress recognition model 101d. The stress recognition model 101d is a speech signal, a sound extracted from the speech signal, Voice feature, text data generated from the speech signal, text feature, or these At least two combinations are input, and an operator who is listening to the speech feels This is a model that outputs an estimated value of stress. The actual stress value felt by the operator after listening to the customer's speech may be used. For example, the data set for training the image recognition model 101d may be generated as follows: The operator should rate the level of stress (for example, 1 to 10) felt by listening to the customer's voice. (i.e., the level of the speech signal that one feels is " This allows the system to determine the degree of stress by listening to the speech signal and the corresponding speech signal. The information obtained is related to the stress of the operator when the operator The stress level is assigned to multiple speech signals by the speech decoder. A data set that is a bundle of such information is obtained. The stress recognition model 101d It may be possible to perform supervised machine learning using such a dataset.

[0043] The control unit 108 performs various controls related to the voice processing device 10. 08, based on the stress information generated by the stress recognition unit 107, Either the synthetic voice generated by the voice synthesis unit 105 or the customer's spoken voice is selected in the terminal 20. The control unit 108 may switch whether to output a synthetic voice signal based on the speech signal. For example, the control unit 108 may switch whether to generate the If the stress level indicated by the stress information is equal to or greater than a predetermined threshold, the customer's speech sound Alternatively, the control unit 108 may be controlled so as to output a synthetic voice to the operator instead of a voice. If the stress level indicated by the stress information is less than or equal to a predetermined threshold, The control unit 108 may control the device so that a voice is output to the operator. If instructions for automatic switching of the emotion suppression function are entered, the function will be switched based on stress information. The emotion suppression function is a function that uses synthetic voice instead of the customer's spoken voice. This function outputs to the operator.

[0044] The control unit 108 may perform the above switching based on emotion information. The switching may be performed based on the output of the emotion suppression switching model 101e. The replacement model 101e is a speech signal, a speech feature, text data, a text feature, or Using at least two of these combinations as input, the emotion suppression function is switched on and off. The model outputs the timing of the emotion suppression switching. The emotion suppression switching model 101e may be input with information or emotion information. State.

[0045] Moreover, the control unit 108 controls the above switching based on switching information input by the operator. Here, the switching information may be the application (on) or non-application (on) of the customer emotion suppression function. For example, the control unit 108 may If the information indicates the application of the customer emotion suppression function, control is performed to output a synthetic voice to the operator. On the other hand, the control unit 108 may detect that the switching information indicates that the customer's emotion suppression function is not applied. In this case, the control unit 108 may control the voice to be output to the operator. If the user inputs instructions on manually switching the emotion suppression function, the above switching will be The switching may be performed based on the information.

[0046] The learning unit 109 includes an emotion recognition model 101c, a stress recognition model 101d, and an emotion suppression model. A learning process may be performed on the switching model 101e.

[0047] The voice processing device 10 receives any one of the following information 1) to 7), or at least two of the following information: The combination of the information is associated on the time axis, and the operator terminal 2 0. 1) A customer's speech voice signal; 2) A text generated from the speech voice signal. 3) the text data after processing by the removal unit 104; 4) the detection information; 5) ) Synthetic speech signal, 6) Customer emotion information estimated from customer speech signal, 7) Emotion suppression When the emotion suppression function is on, the voice processing device The device 10 does not need to send the customer's speech signal to the operator terminal 20. When it is off, the voice processing device 10 does not need to send a synthesized voice signal to the operator terminal 20. Regardless of whether the emotion suppression function is on or off, the voice processing device 10 detects the customer's speech signal and Both the voice signal and the synthesized voice signal may be sent to the operator terminal 20.

[0048] <Operator terminal>

[0049] FIG. 6 is a diagram illustrating an example of a functional configuration of an operator terminal according to the present embodiment. The data terminal 20 includes a transmission / reception unit 201, an input reception unit 202, and a control unit 203. The functional configuration shown is merely an example, and other configurations not shown may be included.

[0050] The transmission / reception unit 201 transmits and receives various information between the voice processing device 10 and / or the customer terminal 30. For example, the transceiver unit 201 receives and transmits a signal at the customer terminal 30. The transceiver 102 may receive a voice signal, which is a signal of the customer's voice that is being transmitted. Alternatively, the transmission / reception unit 201 may receive a synthetic voice signal from the voice processing device 10. The subjective evaluation information may be transmitted to the voice processing device 10. Customer emotion information may be received from the voice processing device 10 .

[0051] The input receiving unit 202 receives various information based on the operation of the input unit 10e by the operator. For example, the input receiving unit 202 receives an emotion recognition model 101c and a stress recognition model 102b. As part of our work to generate a dataset for training the 101d recognition model, Even if we accept subjective evaluation information and stress level input for the customer's live speech signal, After that, the operator can input the subjective evaluation information and the degree of stress on the operator terminal 20. The task of inputting the information is called "annotation work." The input reception unit 20 may be positioned as a separate operation from the call center operation. The input receiving unit 2 may receive switching information for the emotion suppression function of the customer. 202 is instruction information for instructing either manual switching or automatic switching of the emotion suppression function may be accepted.

[0052] The control unit 203 performs various controls related to the operator terminal 20. For example, The control unit 203 controls the display of information and / or images on the display unit 10f. The control unit 203 controls the output of the sound from the sound output unit 10i. The voice output may be controlled based on the information transmitted from the input receiving unit 202. The audio output may be controlled based on the received information.

[0053] The control unit 203 generates a synthetic voice based on the synthetic voice signal received from the voice processing device 10. The control unit 203 outputs the voice from the voice output unit 10i based on the voice signal from the customer terminal 30. Based on this, the speech may be output from the speech output unit 10i.

[0054] In addition, the control unit 203 generates a synthetic voice based on the emotion information received from the voice processing device 10. The control unit 203 may display emotion information corresponding to the signal on the display unit 10f. The text data corresponding to the synthesized voice signal received from the voice processing device 10 is displayed on the display unit 10f. For example, the control unit 203 may display emotion information, text data, and at least one of the detection information. The control unit 203 may display a screen D1 including at least one of the images on the display unit 10f. For example, the control unit 203 may display the stress information on the display unit 10f. A screen D2 including the information may be displayed on the display unit 10f.

[0055] FIG. 7 is a diagram showing an example of a screen D1 according to the present embodiment. As shown in FIG. In step 1, the control unit 203 synchronizes with the output timing T of the synthetic voice from the voice output unit 10i. The emotion information I1 may be displayed on the display unit 10f in conjunction with the synthesized voice output timing T By displaying emotion information I1 every time, the operator can suppress the emotion of the customer by using the emotion suppression function. It is possible to recognize customer emotions in real time, even when listening to synthetic voices with suppressed emotions. Cut.

[0056] In addition, in the screen D1, the control unit 203 outputs the synthesized voice in accordance with the output timing T. The contents of the text data I2 corresponding to the synthesized voice may be displayed on the display unit 10f. By displaying the contents of the text data I2, the operator can understand the message using only the synthesized voice. This makes it possible to visually understand what the customer is saying without having to wait for a moment.

[0057] In addition, in the screen D1, the control unit 203 displays the following information based on the detection information received from the voice processing device 10: Instead of displaying the specific word string itself, information I3 indicating the detection of the specific word string (e.g. For example, "NG word detection" may be displayed on the display unit 10f. This function is called the "hiding function." This function hides the contents of customer utterances that have a negative psychological effect. This reduces the stress on the operator because it is not necessary for the operator to recognize the The operator can be notified of the occurrence of the utterance, so that the operator can respond appropriately to the customer. This can be done honestly.

[0058] In addition, in the screen D1, the control unit 203 displays a screen based on emotion information from the voice processing device 10. Then, the level I4 of a particular emotion of the customer is displayed in time series for each output timing T of the synthetic voice. For example, in FIG. 7, the customer's voice for each output timing T of the synthetic voice may be displayed on the display 10f. The "anger" level I4 is shown in a line graph. This allows the operator to identify the customer. The transition of emotions (e.g., anger) can be easily grasped, so the operator's attitude toward the customer can be This can improve customer satisfaction.

[0059] In the screen D1, the control unit 203 may display a selection button I5 on the display unit 10f. . The selection button I5 is used to automatically or manually apply (ON) or not apply (OFF) the emotion suppression function. This is an interface that allows the operator to select which one to switch between. By clicking, tapping, sliding, etc. on the selection button I5, In automatic switching mode, the camera can be switched between "automatic switching mode" and "manual switching mode". For example, the emotion information, the stress information, or the output from the emotion suppression switching model 101e The emotion suppression function is automatically switched on or off based on this.

[0060] When the "manual switching mode" is selected, the control unit 203 applies or disables the emotion suppression function. The display unit 10f is provided with a selector button I6, which is an interface that allows the operator to select the desired use. The timing when the operator switches the emotion suppression function on and off may be displayed on the The relationship between the customer's speech (and / or various features extracted based on the speech) and the time axis The data is stored in a storage unit (not shown) as "manual switching history data." The history data may further be associated with an operator's identification information.

[0061] The I7 switch button is used to turn on and off the "NG word hiding function." If the "NG word hiding function" is off, the text data I2 contains Even if a certain word string is detected, the text data before the processing by the removal unit 104 is The data I2 is displayed on the display unit 10f as it is. If you turn off the display feature, the operator will not hear the specific word sequence directly from the customer. This reduces stress, while at the same time, by accurately understanding what the customer is saying, the You can get a more accurate picture of the situation.

[0062] The dataset for training the emotion inhibition switching model 101e is stress information, emotion information, and information, speech signal S1, speech features, text data, text features or any of these At least two combinations and the time when the operator switches the emotion suppression function on and off. The emotion suppression switching model 10 may be a bundle of data associated with each other on the time axis. There are various ways to learn 1e, such as those described below in 1) to 3). 1) The emotion suppression switching model 101e may be trained for each operator. The emotion suppression switching model 101e applied to the data is the emotion suppression by the operator. It may be possible to learn based only on the "manual switching history data" of the function. The emotion suppression switching model 101e can suppress emotions at the timing that suits the operator's preference. Or, 2) the operator can select the The emotion suppression switching model 101e is based on the "manual switching history data" of an unspecified number of operators. According to this method, the data available for learning may be Since the number of the emotion inhibition switching models 101e is increased, the emotion inhibition switching model 101e can be learned quickly. Or, 3) the emotion suppression switching model 101e applied to a certain operator is "Manual switching history data" by operators with similar age, gender, and other characteristics to the operator According to this method, the number of learning methods used is smaller than that in the method 1). Since the amount of data that can be used increases, the emotion inhibition switching model 101e can be trained quickly. By comparing this with method 2), you can learn the switching timing that suits you best. It becomes like this.

[0063] FIG. 8 is a diagram showing an example of a screen D2 according to the present embodiment. 203 may display the stress information from the voice processing device 10. For example, in FIG. As the stress information, information indicating the estimated stress felt by the operator (e.g., "5 6%") and information showing the relative evaluation value from the normal state of the operator (for example, The message displayed was, "8.1% decrease from normal."

[0064] FIG. 12 is a diagram showing an example of a screen D3 according to the present embodiment. The unit 203 displays an interface I8 for an operator to perform annotation work. For example, the operator may ask the customer to answer the question while listening to the actual voice (sample voice) of the customer. The customer's emotions sensed from the sample voices are selected each time from the interface I8. Figure 1 In 2, customer sentiment I1 is subjective evaluation information of customer sentiment by an operator. For example, The operator responded to the sample voice, "Please deliver it somehow by this evening." If you annotate the emotion "by this evening," as shown in Figure 12, The sample voice saying "Please deliver it somehow" and the information "anger" are associated on the timeline. Annotation can be done sentence by sentence or at a given time interval. stomach.

[0065] (Operation of the voice processing system) FIG. 9 is a flowchart showing an example of the emotion suppression operation according to the present embodiment. 9 is merely an example, and the order of at least some steps (e.g., step S106) may be different. The steps may be replaced, steps not shown may be performed, or some steps may be omitted. It may be omitted.

[0066] The voice processing device 10 converts the voice of a customer uttered by the voice input unit 10h of the customer terminal 30 into A speech signal is acquired (S101).

[0067] The speech processing device 10 extracts features based on the speech signal acquired in S101. is input to the speech recognition model 101a to generate text data including a word string consisting of one or more words. Then, data is generated (S102).

[0068] The voice processing device 10 detects whether a specific word string is included in the text data generated in S102. If the text data contains a specific word string, the process determines whether the specific word string is included in the text data (S103). The speech processing device 10 then removes the specific word string or changes the specific word string to another word string. Converted text data is generated (S104).

[0069] The speech processing device 10 uses the feature quantity extracted based on the text data as a speech synthesis model 1 01b to generate a synthetic voice signal, which is a signal for synthetic voice (S105).

[0070] The speech processing device 10 receives an utterance speech signal acquired in S101, a text generated in S102, and At least one of the subjective evaluation information of the customer's emotions input by the operator is used. The feature amount extracted based on either one of them is input to the emotion recognition model 101c to recognize the emotion of the customer. The information is generated (S106).

[0071] The operator terminal 20 outputs a synthetic voice based on the synthetic voice signal generated in S105. The output unit 10i outputs the synthesized voice in accordance with the output timing T of the synthesized voice. Emotion information corresponding to the synthetic voice is displayed on the display unit 10f (S107, for example, FIG. 7).

[0072] The voice processing device 10 judges whether or not to end the process (S108). If not (S108: NO), the voice processing device 10 executes the processes S101 to S107 again. On the other hand, when the voice conversion process is to be ended (S108: YES), the voice processing device 10 End the process.

[0073] FIG. 10 is a flowchart showing the automatic switching operation of the emotion suppression function according to this embodiment. Note that FIG. 10 is merely an example, and the order of at least some of the steps may be changed. Alternatively, steps not shown in the figure may be performed, or some steps may be omitted. stomach.

[0074] The voice processing device 10 generates stress information of an operator (S201).

[0075] The voice processing device 10 judges whether the stress information satisfies a predetermined condition (S202 For example, the predetermined condition may be when the stress level indicated by the stress information is equal to or greater than a predetermined threshold. It could be something big.

[0076] When the stress information satisfies a predetermined condition (S202: YES), the voice processing device 10 An emotion suppression function may be applied (i.e., synthetic voice may be output from the operator terminal 20) ( On the other hand, if the stress information does not satisfy the predetermined condition (S203), the voice processing device 10 S202: NO), the emotion suppression function is not applied (i.e., the customer's call is not The speech may be output (S204).

[0077] The voice processing device 10 judges whether or not to end the process (S205). If not (S205: NO), the voice processing device 10 executes the processes S201 to S204 again. On the other hand, when the voice conversion process is to be ended (S205: YES), the voice processing device 10 In S201 and S202, the voice processing device 10 Based on the output of the emotion suppression switching model 101e, it is determined whether or not to apply the emotion suppression function. Good too.

[0078] As described above, according to the voice processing system 1 of the present embodiment, the voice signal uttered by the customer is and generating text data based on the text data. The customer's emotions contained in the customer's voice are sufficiently suppressed. The operator can listen to the synthesized voice generated by the customer, and the operator can identify the emotional responses of the customer to the synthesized voice. The inventor of the present invention conducted a survey of approximately 50 subjects to determine whether 1) the customer's 2) the actual voice of the customer's speech, with the volume of the customer's speech adjusted; and 3) the quality of the customer's voice. 4) Synthetic voice generated by converting the customer's speech into text Participants were asked to compare the two audio recordings and rate the level of anger they felt from the recordings on a 7-point scale. As a result, 4) conveyed more anger to the subjects than 2) and 3). The reduction was remarkable.

[0079] In addition, according to the voice processing system 1 of this embodiment, a synthetic voice is provided to the operator. Not only outputting the voice, but also notifying the customer's emotional information in accordance with the timing of the synthetic voice output. This allows the operator to recognize the customer's emotions in real time when listening to the synthesized voice. Able to respond appropriately to customers.

[0080] In addition, according to the voice processing system 1 of this embodiment, the stress information of the operator or Based on the customer's emotion information, etc., whether or not to apply emotion suppression function (i.e., to the operator) The operator can choose whether to output a synthesized voice or spoken voice. This allows for an appropriate balance between stress and customer satisfaction.

[0081] (Example of change) In the above-described speech processing system 1, the speech recognition unit 103 extracts one or more However, the present invention is not limited to this. The voice recognition unit 103 determines whether a word string recognized from the speech signal is one or more sentences. Text data including a word string consisting of one or more words (parts of speech or morphemes) before The removal unit 104 may generate a sentence in the text data that is not determined as the sentence. The specific word string is removed, and the speech synthesis unit 105 generates a text data that is not determined as the sentence. A synthetic speech signal may be generated from the data.

[0082] FIG. 11 is a diagram showing an example of generation of a synthetic speech signal according to a modification of this embodiment. In step 1, a speech recognition unit 103 recognizes a speech signal S4 acquired by a transmission / reception unit 102. As shown in FIG. 11, text data T41 to T43 are generated. The text data T41 to T43 are used to determine the meaning of the sentence "Please send it quickly." The text data was generated using morpheme units with the following characteristics: "hurry up," "send," and "please" The removal unit 104 performs the following for each of the text data T41 to T43. The speech synthesis unit determines whether a specific word string is included in the speech signal, and removes the specific word string. The speech synthesis unit 105 outputs the synthesized speech from the text data T41 to T43. Then, synthesized audio signals S41 to S43 are generated.

[0083] As shown in Figure 11, text data is generated in units of one or more morphemes before the sentence is finalized. By generating text data, the operator's response delay is reduced by outputting a synthetic voice. In addition, multiple text data (or synthetic voice) in morpheme units can be semantically A model for determining whether an image is unnatural may be used.

[0084] In order to reduce the response delay, the synthetic voice signals S1 to S3 shown in FIG. 4 and the For example, "Ah", "Eh", etc. may be inserted before and / or after each of the synthetic speech signals S41 to S43. Filler sounds such as "Oh well" may be added to the response. This allows the operator to feel at ease even when the response is delayed. This can prevent a decline in customer satisfaction.

[0085] In addition, the voice synthesis unit 105 generates a plurality of voices based on the customer's emotions estimated by the emotion recognition unit 106. Select a voice synthesis model 101b that matches the customer's emotions from among the voice synthesis models 101b. For example, if the emotion of the customer estimated by the emotion recognition unit 106 is "anger," The voice synthesis unit 105 may use a voice synthesis model 101b that has a fast pitch and strong intonation. For example, when the emotion of the customer estimated by the emotion recognition unit 106 is “crying”, the voice synthesis unit 105 Alternatively, the voice synthesis model 101b that outputs a crying-like voice may be used. The synthesis unit 105 generates a voice synthesis model 10 based on the customer's emotion estimated by the emotion recognition unit 106. The parameters of 1b may be changed and adjusted so that a voice that matches the customer's emotions is output. The operators who heard the actual voices of angry customers felt extremely stressed. On the other hand, operators need to be able to understand the customer's emotions in real time in order to carry out customer service tasks appropriately. It is necessary to grasp the time accurately. By not letting the operator hear the voice directly, The operator does not feel excessive stress, and the synthetic voice conveys the customer's emotions. This allows the operator to grasp the customer's emotions in real time through hearing.

[0086] (Other embodiments) In the above embodiment, the customer's speech signal is converted into text, and the synthesized speech signal is transmitted to the operator. However, the present invention is not limited to this. inputting the speech feature quantity extracted based on the above to a speech conversion model to generate a converted speech signal; The converted voice may be output from the operator terminal 20.

[0087] The "voice conversion model" described in the claims is a model that converts a speech signal into text and then synthesizes it. A model that outputs a synthesized voice and a model that converts the voice quality of the speech signal without converting it to text. This concept encompasses both a model that uses synthetic voice or conversion instead of the customer's speech. By outputting voice to the operator, the operator can On the other hand, in order to carry out customer service operations, the operator must It is also essential to understand customer sentiment in real time.

[0088] The voice processing device 10 in this modification is a voice processing device that extracts voice signals based on the speech signals of customers. The feature quantity is input to a voice conversion model to generate a converted voice signal. 1) The converted voice signal and 2) the customer's emotion information estimated from the customer's speech are related on the time axis. The voice processing device 10 generates linked information and transmits it to the operator terminal 20. The information to be processed includes a speech signal, text data generated from the speech signal, and a removal unit 10 Text data after processing in 4, detection information, emotion suppression function on / off switching The timing of the detection may be associated with the detection.

[0089] The operator terminal 20 outputs the converted voice signal received from the voice processing device 10 to the voice output unit 1 0i and in accordance with the output timing T of the converted audio from the audio output unit 10i. The operator terminal 20 may further display information indicating the emotion information on the display unit 10f. The text data is displayed on the display unit 10i in accordance with the output timing T of the converted voice from the output unit 10i. 0f. Such a display may be as shown in FIG.

[0090] The voice processing device 10 in this modification changes the emotion indicated by the emotion information based on the emotion information. For example, a signal of the converted voice may be generated so that the converted voice is reflected in the emotional information. If the emotion is "enraged," a voice conversion model with a fast pitch and strong intonation may be used. For example, if the emotion information indicates "crying," the voice conversion model outputs a voice that sounds like crying. The voice processing device 10 may use a voice-converting function to convert the emotion represented by the emotion information into a converted voice. In this way, a signal of converted voice may be generated. The operator does not feel excessive stress, and the converted voice conveys the customer's emotions. This allows the operator to grasp the customer's emotions in real time through hearing.

[0091] In the speech processing system 1 of this modified example, the operator creates annotations. The transaction is carried out by the operator during normal call center operations on the converted voice. If the operator annotates the converted voice with "anger," Based on the annotation results, the voice conversion model is trained to output softer voices. may be adjusted in real time.

[0092] In the above-described embodiment, the first user is a customer and the second user is an operator. Although a call center is assumed, the application of this embodiment is not limited to call centers. For example, in a web meeting, the voice of a first user with suppressed emotions can be transmitted to a second user. In other words, this embodiment can be applied to any situation where customer harassment occurs. Not only for harassment prevention measures, but also for various types of harassment, such as power harassment in the company. This can be used as a countermeasure for the business side.

[0093] In the embodiment described above, emotion information and synthetic speech are associated on the time axis. As shown in Figure 7, the processing is performed according to the output timing of the synthesized voice or converted voice. If it is possible to display emotion information estimated from the original speech of the user, In the embodiment described above, the "association on the time axis" is not limited to the specific form. The process of "associating" may be a process of associating based on time information such as the hour, minute, and second, or The process of associating the information based on information such as how many minutes and seconds have elapsed since the start of the speech information may also be used. Alternatively, the process may involve association on a sentence-by-sentence, word-by-word, or morpheme-by-morpheme basis.

[0094] In the voice processing system 1 according to the embodiment described above, a customer may input his / her own voice. It is also possible to suppress emotions and make it difficult for the operator to understand that the message is being received. Customers cannot tell whether the emotion suppression function is on or off. You can do so.

[0095] The annotation work may be performed by the operator on the operator terminal 20, or may be performed separately. Dedicated applications and terminals for annotation work may be provided.

[0096] The above-described embodiment is intended to facilitate understanding of the present invention. The elements, arrangements, materials, and configurations of the embodiments are not intended to be interpreted as being limited to the above. The conditions, shapes, sizes, etc. are not limited to those shown in the examples and may be changed as appropriate. In addition, the configurations shown in different embodiments may be partially substituted or combined with each other. In addition, the functions described as the functions of the voice processing device 10 can be implemented in the operator terminal 2. In addition, the functions described as the functions of the operator terminal 20 may be provided in the voice processing The device 10 may include: [Explanation of symbols]

[0097] 1... voice processing system, 10... voice processing device, 20... operator terminal, 30... customer terminal , 10a... processor, 10b... RAM, 10c... ROM, 10d... communication unit, 10e... input 10f... display unit, 10g... camera, 10h... audio input unit, 10i... audio output unit, 01...storage unit, 102...transmission / reception unit, 103...voice recognition unit, 104...removal unit, 105...voice Synthesis unit, 106... emotion recognition unit, 107... stress recognition unit, 108... control unit, 109... learning 201: transmitting and receiving unit, 202: input receiving unit, 203: control unit

Claims

1. A voice processing device used for customer service at a call center, An acquisition unit that acquires a speech signal that is a signal of a speech voice of a first user; a voice conversion unit that converts at least one of a volume and a voice quality of the uttered voice to generate a converted voice so as to suppress the emotion of the first user included in the uttered voice; a stress recognition unit that generates stress information regarding a stress state of the second user; a control unit that switches whether to output the spoken voice or the converted voice from a voice output unit based on the generated stress information; an emotion recognition unit that generates emotion information of the first user corresponding to the speech voice signal; a display unit that displays, for each output timing of the converted voice output to the second user based on the generated converted voice and the emotion information, an emotion level based on the emotion information corresponding to the converted voice in a time series; Equipped with Audio processing device.

2. an emotion recognition unit that generates emotion information of the first user corresponding to the speech voice signal; 2. The speech processing device according to claim 1, wherein the emotion recognition unit generates emotion information of a first user corresponding to the speech signal acquired by the acquisition unit by inputting the speech signal acquired by the acquisition unit, the speech feature extracted from the speech signal, the text data generated from the speech signal, the feature extracted from the text data, or a combination of at least two of these, into an emotion recognition model trained by machine learning to input the speech signal, the feature extracted from the speech signal, the text data generated from the speech signal, the text feature corresponding to the text data, or a combination of at least two of these, and output emotion information of a speaker of the speech signal.

3. a speech recognition unit that inputs the feature amount extracted based on the speech signal into a speech recognition model to generate text data including a word string consisting of one or more words; a removal unit that detects a specific word string included in the text data, and removes the specific word string or replaces the specific word string with another word string to generate text data; The audio processing device of claim 1 , further comprising:

4. The removal unit generates information regarding a warning for the first user when the specific word string is detected; The audio processing device according to claim 3 , wherein the generated information is output in an information processing device operated by the first user.

5. The speech processing device according to claim 3 , wherein the specific word string includes one or more words that have a negative psychological effect on the second user.

6. The speech conversion unit generates the converted speech without converting the spoken speech into text. The audio processing device according to claim 1 .

7. A voice processing device used for customer service at a call center, An acquisition unit that acquires a speech signal that is a signal of a speech voice of a first user; a voice conversion unit that converts at least one of the volume and the voice quality of the uttered voice to generate a converted voice so as to reduce the emotion of the first user conveyed to a second user; a stress recognition unit that generates stress information regarding a stress state of the second user; a control unit that switches whether to output the spoken voice or the converted voice from a voice output unit based on the generated stress information; an emotion recognition unit that generates emotion information of the first user corresponding to the speech voice signal; a display unit that displays, for each output timing of the converted voice output to the second user based on the generated converted voice and the emotion information, an emotion level based on the emotion information corresponding to the converted voice in a time series; Equipped with Audio processing device.

8. The first user emotion is an emotion of anger.

8. The audio processing device according to claim 1 or 7.

9. A voice processing device used for customer service at a call center, An acquisition unit that acquires a speech signal that is a signal of a speech voice of a first user; a voice conversion unit that converts at least one of a volume and a voice quality of the uttered voice to generate a converted voice so as to suppress the emotion of the first user included in the uttered voice; a stress recognition unit that generates stress information regarding a stress state of the second user; an emotion recognition unit that generates emotion information of the first user corresponding to the speech voice signal; a display unit that displays, for each output timing of the converted voice output to the second user based on the generated converted voice and the emotion information, an emotion level based on the emotion information corresponding to the converted voice in a time series; Equipped with The stress recognition unit generates the stress information based on a stress recognition model, The stress recognition model is trained based on information in which the speech voice and the stress of the second user when listening to the speech voice are associated on a time axis. Audio processing device.

10. A voice processing method for use in customer service at a call center, comprising: An acquisition step of acquiring a speech signal which is a signal of a speech voice of a first user; a voice conversion step of converting at least one of the volume and the voice quality of the uttered voice to generate a converted voice so as to suppress the emotion of the first user included in the uttered voice; a stress recognition step of generating stress information relating to a stress situation of the second user; a switching step of switching whether the speech voice or the converted voice is to be output from a voice output unit based on the generated stress information; generating emotion information of the first user corresponding to the speech signal; a display step of displaying, in time series, emotion levels based on the emotion information corresponding to the converted voice at each output timing of the converted voice output to the second user based on the generated converted voice and the emotion information; Equipped with Audio processing methods.

11. A program for causing a computer to function as the voice processing device according to any one of claims 1 to 9.

12. A voice processing device used for customer service at a call center, An acquisition unit that acquires a speech signal that is a signal of a speech voice of a first user; a voice conversion unit that converts at least one of a volume and a voice quality of the uttered voice to generate a converted voice so as to suppress the emotion of the first user included in the uttered voice; a learning unit configured to learn a stress recognition model that outputs stress information indicating an estimated value of stress felt by the second user; an input receiving unit that receives an input of a degree of stress felt due to an emotional utterance of the first user from the second user; an emotion recognition unit that generates emotion information of the first user corresponding to the speech voice signal; a display unit that displays, for each output timing of the converted voice output to the second user based on the generated converted voice and the emotion information, an emotion level based on the emotion information corresponding to the converted voice in a time series; Equipped with The learning unit learns the stress recognition model using a data set in which a speech voice signal and a stress of the second user when listening to the speech voice signal are associated on a time axis, the data set being generated based on the input accepted by the input accepting unit. Audio processing device.

13. the stress recognition unit generates the stress information indicating an estimated value of the stress felt by the second user based on the speech voice signal, a feature amount extracted from the speech voice signal, text data generated from the speech voice signal, a text feature amount extracted from the text data, or a combination of at least two of them.

8. The audio processing device according to claim 1 or 7.

14. The stress recognition model receives as input the speech signal, features extracted from the speech signal, text data generated from the speech signal, text features extracted from the text data, or a combination of at least two of these, and outputs an estimate of the stress felt by the second user.

13. A voice processing device according to claim 9 or 12.

Citation Information

Patent Citations

  • System and program for voice conversion

    JP2004252085A

  • Voice control system, voice controller, voice control method, and voice control program

    JP2013046088A

  • Telephone call answering job support system and method of the same

    JP2013157666A

  • Information processing system, information processing method, and program

    JP2019110451A

  • Speech summary generation apparatus, speech summary generation method, and program

    JP2020071676A