Device and method
The apparatus and method address suboptimal communication by analyzing speaker states and adjusting voice inputs, enhancing communication smoothness and reducing psychological burdens in audio terminal interactions.
Patent Information
- Application Number
- PCT/JP2024/014418
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-04-09
- Publication Date
- 2025-10-16
AI Technical Summary
Existing communication technologies fail to support smooth conversation between speakers via audio terminals, as the state of one speaker can affect the speech and perception of the other, leading to psychological burdens and suboptimal communication experiences.
An apparatus and method that includes an acquisition unit for voice and biometric information, detection units for analyzing speaker states, and an adjustment unit to modify voice inputs based on detected emotional and biometric data to enhance communication smoothness.
The solution enables smooth conversation by adjusting voice inputs based on speaker states, reducing psychological burdens and improving communication quality through targeted voice processing.
Smart Images

Figure JP2024014418_16102025_PF_FP_ABST
Abstract
Description
Apparatus and method
[0001] One aspect of the present disclosure relates to an apparatus and a method.
[0002] Patent Document 1 discloses a technology in which, if a change in emotion is detected in a customer's voice during a conversation after a telephone call has begun, the representative's voice is changed in real time to correspond to the change in emotion of the customer, thereby performing voice conversion tuning.
[0003] Japanese Patent Application Laid-Open No. 2014-095753
[0004] In communication between speakers via audio terminals, the state of one speaker can affect the speech uttered by that speaker and the impression of that speech. Furthermore, the state of the other speaker listening to the speech can affect how the speech uttered by the first speaker is perceived. The other speaker may wish to ensure smooth communication, at least for the other speaker, regardless of the state of the first speaker. Therefore, there is a demand for technology that supports smooth communication by applying appropriate processing according to the state of each speaker.
[0005] Therefore, an object of the present disclosure is to provide an apparatus and method that can realize smooth conversation between speakers via audio terminals.
[0006] The device of the present disclosure comprises an acquisition unit that acquires the voice of a first speaker and biometric information of a second speaker who is conversing with the first speaker via a voice terminal, a first detection unit that detects information regarding the state of the first speaker based on the voice of the first speaker, a second detection unit that detects information regarding the state of the second speaker based on the biometric information of the second speaker, and an adjustment unit that adjusts the voice based on the information regarding the state of the first speaker and the information regarding the state of the second speaker.
[0007] According to one aspect of the present disclosure, smooth conversation between speakers via audio terminals can be realized.
[0008] Fig. 1 is a block diagram showing a configuration of a processing system including an apparatus according to an embodiment of the present disclosure. Fig. 2 is a diagram showing an example of the configuration of information about emotions, which is an example of information about a state. Fig. 3 is a diagram showing information about reactions, which is an example of information about a state. Fig. 4 is a flowchart showing the steps of an example of a processing method by the processing system. Fig. 5 is a diagram showing an example of the hardware configuration of an apparatus according to an embodiment of the present disclosure.
[0009] The present disclosure will be described with reference to the accompanying drawings. Whenever possible, the same parts are designated by the same reference numerals and redundant description will be omitted.
[0010] Fig. 1 is a block diagram showing the configuration of a processing system including an apparatus according to an embodiment of the present disclosure. The processing system 1 shown in Fig. 1 includes an apparatus 10, a first terminal 11, a second terminal 12, a detection device 13, and a user state database 14, which are configured to be able to communicate with each other via a network including a wireless communication network and a fixed communication network. The processing system 1 is used in a conversation between a first speaker using the first terminal 11, which is an example of a voice terminal, and a second speaker who converses via the second terminal 12, which is also an example of a voice terminal. The first speaker is, for example, a customer of a certain product or service. The second speaker is an operator belonging to an organization that is the recipient of an inquiry regarding the product or service.
[0011] The device 10 adjusts at least one of the voice input to the first terminal 11 and the voice input to the second terminal 12 based on the state of a first speaker who is uttering voice toward the first terminal 11 and the state of a second speaker who is uttering voice toward the second terminal 12. The device 10 of this embodiment relays the voice of the first speaker from the first terminal 11 to the second terminal 12, and relays the voice of the second speaker from the second terminal 12 to the first terminal 11. The device 10 adjusts the voice input to the first terminal 11 and the voice input to the second terminal 12. The device 10 adjusts the voice based on information obtained from the detection device 13 and the user state database 14. Each component will be described in detail below.
[0012] The first terminal 11 and the second terminal 12 are audio terminals used when a first speaker and a second speaker intend to have a conversation. The first terminal 11 and the second terminal 12 are devices such as personal computers, smartphones, tablet terminals, feature phones, server devices, game consoles, etc. Note that while FIG. 1 illustrates only one terminal each as the first terminal 11 and the second terminal 12, the processing system 1 may include any number of first terminals 11 and second terminals 12, two or more.
[0013] The first terminal 11 acquires the voice of the first speaker (hereinafter referred to as the first voice). The second terminal 12 acquires the voice of the second speaker (hereinafter referred to as the second voice). The first terminal 11 and the second terminal 12 each output the acquired voice to the device 10. The first terminal 11 acquires the voice of the second speaker adjusted by the device 10 (hereinafter referred to as the second adjusted voice). The second terminal 12 acquires the voice of the first speaker adjusted by the device 10 (hereinafter referred to as the first adjusted voice). Note that if the device 10 does not adjust the second voice, the device 10 does not relay the second voice from the second terminal 12, and the second terminal 12 may output the second voice directly to the first terminal 11 via the network.
[0014] The detection device 13 detects biometric information of the second speaker. The biometric information of the second speaker detected by the detection device 13 includes at least one of information on the voice of the second speaker, the biometric signal of the second speaker, and a moving image of the second speaker. The detection device 13 includes, for example, a wearable device, a camera, etc. The detection device 13 is not limited to these. The wearable device of the detection device 13 is worn on the body of the second speaker. For example, when the first speaker and the second speaker are conversing via the processing system 1, the wearable device of the detection device 13 detects a value related to at least one index of the second speaker's heart rate, breathing rate, activity amount (amount of movement), and body temperature as biometric information of the second speaker.
[0015] The camera of the detection device 13 captures an image of at least a part of the second speaker's body. The camera of the detection device 13 may also acquire the voice of the second speaker. For example, when the first speaker and the second speaker are conversing via the processing system 1, the camera of the detection device 13 detects information regarding the movement of each part of the second speaker's body, such as the amount and direction of movement per unit time, and the second voice as biometric information of the second speaker. The movement of each part includes, for example, eye movement and whole-body movement. The detection device 13 outputs the detected biometric information of the second speaker. The detection device 13 stores the detected biometric information of the second speaker. Note that the device 10 may have the configuration and functions of the detection device 13.
[0016] The user state database 14 stores a table having indicators and thresholds that can be compared with information on the state of the first speaker (hereinafter referred to as first state information) and information on the state of the second speaker (hereinafter referred to as second state information). Details of the information stored in the user state database 14 will be described later.
[0017] The device 10 is configured to include, as functional components, an acquisition unit 20, a detection unit 30, and an adjustment unit 40. The device 10 adjusts the first audio acquired from the first terminal 11. The device 10 outputs the first adjusted audio to the second terminal 12. The device 10 adjusts the second audio acquired from the second terminal 12. The device 10 outputs the second adjusted audio to the first terminal 11. The functions of each functional unit of the device 10 will be described in detail below.
[0018] The acquisition unit 20 acquires the first speech and the biometric information of the second speaker. The acquisition unit 20 has a speech acquisition unit 21 and a state acquisition unit 22. The speech acquisition unit 21 acquires the first speech input to the first terminal 11 by the first speaker from the first terminal 11. The speech acquisition unit 21 acquires the first speech by accepting the first speech transmitted from the first terminal 11. The speech acquisition unit 21 acquires the second speech input to the second terminal 12 by the second speaker from the second terminal 12. The speech acquisition unit 21 acquires the second speech by accepting the second speech transmitted from the second terminal 12.
[0019] The status acquisition unit 22 acquires biometric information of the second speaker. The status acquisition unit 22 acquires the biometric information of the second speaker from the detection device 13. The status acquisition unit 22 may acquire information (second status information) about the state of the second speaker after the first voice or the first adjusted voice reaches the second terminal. When the voice acquisition unit 21 acquires the first voice or the second voice, the status acquisition unit 22 may output a signal to the detection device 13 to acquire the biometric information of the second speaker. The detection device 13 may acquire the biometric information of the second speaker using the signal as a trigger.
[0020] The detection unit 30 detects first status information and second status information. The first status information and second status information include information regarding the emotions of the first speaker and the second speaker, respectively. The detection unit 30 has a first detection unit 31 and a second detection unit 32. The first detection unit 31 detects the first status information based on the first voice acquired by the voice acquisition unit 21. The first detection unit 31 analyzes an index indicating the voice quality indicated by the first voice. The first detection unit 31 analyzes at least one index of the volume, pitch, pitch, and clarity indicated by the first voice, as an example of the index indicating the voice quality. The first detection unit 31 of the present embodiment analyzes, for example, the volume, pitch, pitch, and clarity indicated by the first voice.
[0021] The first detection unit 31 acquires information on emotions, which is an example of information on the state, from the user state database 14. The first detection unit 31 detects the emotion indicated by the first voice as first state information, based on the indicators indicated by the first voice and the information on emotions acquired from the user state database 14. The first detection unit 31 of this embodiment detects the emotion indicated by the first voice by comparing the indicators of volume, pitch, pitch, and clarity indicated by the first voice with the indicators (rules) corresponding to predetermined emotions.
[0022] FIG. 2 is a diagram showing an example of the structure of information about emotions, which is an example of information about a state. In the example shown in FIG. 2, the information about emotions acquired from the user state database 14 may be an emotion table showing the correspondence between emotions and voice features (indicators). Each emotion, such as "joy," "anger," or "sadness," is correlated with an index indicating voice quality (for example, at least one index of "volume," "pitch," "pitch," and "clarity" indicated by the voice). The numerical values shown in FIG. 2 indicate reference values (%) when the maximum value of each index is 100 and the minimum value is 0.
[0023] For example, if the first speaker is speaking rapidly in a loud, low voice, the first detection unit 31 analyzes the "volume," "pitch," "pitch," and "clarity" indicated by the first voice as 70, 20, 75, and 40, respectively. The first detection unit 31 searches the emotion table acquired from the user state database 14 to determine which emotion corresponds to "volume," "pitch," "pitch," and "clarity" when they are 70, 20, 75, and 40. Based on the emotion table acquired from the user state database 14, the first detection unit 31 estimates that the emotion of the first speaker includes anger. Therefore, the first detection unit 31 detects anger from the first voice as information related to the emotion of the first state information. In this way, the first detection unit 31 detects information related to the emotion of the first state information based on the values of each index indicated by the first voice and the emotion table.
[0024] Note that the first detection unit 31 does not have to detect the emotion indicated by the first voice by comparing an index indicating the voice quality of the first voice with each index (rule) corresponding to a predetermined emotion. For example, the first detection unit 31 may input each index indicated by the first voice to a machine learning model within the device 10 or external to the device 10, and output information related to the emotion of the first status information. The machine learning model may include, for example, at least one of a generative artificial intelligence (AI) model, a discrimination / determination type AI model, and a model combining these. Below, an example in which a generative AI model is used as the machine learning model will be described. The generative AI model outputs, as the first status information, an emotion estimated to be indicated by the first voice based on an index indicating the voice quality of the input first voice. The first detection unit 31 may acquire the first status information output by the generative AI model.
[0025] Furthermore, the first detection unit 31 may not analyze an index indicating the voice quality of the first voice. In this case, for example, the first detection unit 31 may input the first voice to a generative AI model to detect an index indicating the voice quality of the first voice. The generative AI model may directly output, as first state information, an emotion estimated to be indicated by the first voice based on the first voice input from the first detection unit 31. Note that the index indicating voice quality does not need to include any of the indexes of volume, pitch, pitch, and clarity. The index indicating voice quality may include, for example, at least one of timbre, resonance, warmth, and nasality.
[0026] The second detection unit 32 detects second status information based on the second voice acquired by the voice acquisition unit 21 and the biometric information of the second speaker acquired by the status acquisition unit 22. The second detection unit 32 may detect an emotion indicated by the second voice as the second status information by processing similar to that of the first detection unit 31 described above. In this case, the second detection unit 32 analyzes an index indicating the voice quality indicated by the second voice. As an example of the index indicating the voice quality, the second detection unit 32 analyzes at least one index of the volume, pitch, pitch, and clarity indicated by the second voice. The second detection unit 32 acquires information regarding emotions, which is an example of information regarding the status, from the user status database 14. The second detection unit 32 detects the emotion indicated by the second voice as the second status information based on each index indicated by the second voice and the information regarding emotions acquired from the user status database 14. The second detection unit 32 of this embodiment detects the emotion indicated by the second voice by comparing the volume, pitch, pitch, and clarity indicators indicated by the second voice with each indicator (rule) related to the voice corresponding to a predetermined emotion.
[0027] The second detection unit 32 acquires information on a reaction, which is an example of information on a state, from the user state database 14. The second detection unit 32 detects a reaction indicated by the second voice as second state information, based on each indicator indicated by the biometric information of the second speaker and the information on the reaction acquired from the user state database 14. The second detection unit 32 of the present embodiment detects a reaction indicated by the second voice by comparing indicators of volume, pitch, pitch, and clarity indicated by the second voice with each indicator (rule) corresponding to a predetermined reaction.
[0028] FIG. 3 is a diagram showing information about reactions, which is an example of information about a state. In the example shown in FIG. 3, the information about reactions acquired from the user state database 14 may be a reaction table showing a correspondence between states (reactions) and indicators of biometric information. Each possible reaction of the second speaker to a stimulus (here, the first voice), such as "excitement," "impatience," or "shock," is correlated with at least one indicator of "heart rate," "eye movement," "breathing rate," and "whole body movement" indicated by the biometric information. The "heart rate," "eye movement," and "breathing rate" shown in FIG. 3 represent the number of times of each indicator, and the feature value of "whole body movement" represents a reference value (%) with a maximum value of 100 and a minimum value of 0. The "heart rate," "eye movement," "breathing rate," and "whole body movement" shown in FIG. 3 are examples of biometric information. The biometric information may include information other than the "heart rate," "eye movement," "breathing rate," and "whole body movement."
[0029] For example, if the second speaker appears distraught, scratching his head and looking in various directions, the second detection unit 32 analyzes the second speaker's biometric information to determine whether the "heart rate," "eye movement," "breathing rate," and "whole body movement" are 95, 20, 25, and 40, respectively. The second detection unit 32 searches the reaction table acquired from the user state database 14 to determine which emotion corresponds to the "heart rate," "eye movement," "breathing rate," and "whole body movement" when they are 95, 20, 25, and 40. Based on the reaction table acquired from the user state database 14, the second detection unit 32 estimates that the second speaker's reaction includes impatience. Therefore, the second detection unit 32 detects an impatient reaction from the second voice as information regarding the reaction of the second state information. In this way, the second detection unit 32 detects information regarding the reaction of the second state information based on the values of each index indicated by the second speaker's biometric information and the reaction table.
[0030] Note that the second detection unit 32 does not have to detect at least one of the emotion and the reaction indicated by the second voice by comparing an index indicating the voice quality of the second voice with each index (rule) corresponding to predetermined emotions and reactions. For example, the second detection unit 32 may input values of each index indicated by the biometric information of the second speaker to a generative AI model (an example of a machine learning model) within the device 10 or external to the device 10, and output at least one of information related to the emotion and information related to the reaction in the second status information. The generative AI model outputs at least one of the emotion and the reaction estimated to be indicated by the second voice as second status information based on the input index indicating the voice quality of the second voice. The second detection unit 32 may acquire the second status information output by the generative AI model.
[0031] Furthermore, the second detection unit 32 does not need to analyze an index indicating the voice quality of the second voice. In this case, for example, the second detection unit 32 may input the second voice to a generative AI model to detect an index indicating the voice quality of the second voice. The generative AI model may directly output, as second state information, at least one of an emotion and a reaction estimated to be indicated by the second voice, based on the second voice input from the second detection unit 32.
[0032] The second detection unit 32 estimates the stress level of the second speaker based on the second state information. The stress level of the second speaker is an index that quantifies the level of stress of the second speaker. The second detection unit 32 estimates the stress level of the second speaker based on, for example, each index of the second voice, including at least one of volume, pitch, pitch, and clarity, and at least one of emotional information and reaction information. The second detection unit 32 estimates that the stress level is higher the greater the degree of deviation between each index of the second voice, including at least one of volume, pitch, pitch, and clarity, and a reference value in the emotion table corresponding to the emotion of the second speaker detected by the second detection unit 32. The degree of deviation here is, for example, the average value of the difference between each index of the second voice and each reference value in the emotion table corresponding to the emotion of the second speaker, divided by each reference value. The degree of deviation here may be, for example, the sum of the differences between each index of the second voice and each reference value in the emotion table corresponding to the emotion of the second speaker.
[0033] The second detection unit 32 estimates that the stress level is higher when the degree of deviation between each index of the second speaker's biometric information, including, for example, "heart rate," "eye movement," "breathing rate," "whole body movement," etc., and the reference value in the reaction table corresponding to the second speaker's reaction detected by the second detection unit 32 is greater. The degree of deviation here is, for example, the average value of the difference between each index of the second voice and each reference value in the reaction table corresponding to the second speaker's reaction, divided by each reference value. The degree of deviation here may also be, for example, the sum of the differences between each index of the second voice and each reference value in the reaction table corresponding to the second speaker's reaction.
[0034] The second detection unit 32 does not have to detect the stress level of the second speaker by comparing the volume, pitch, pitch, and clarity indicators indicated by the second voice with reference values indicated in a predetermined emotion table and reaction table. For example, the second detection unit 32 may input each indicator indicated by the second speaker's biometric information and second state information (at least one of information related to emotions and information related to reactions) into a generative AI model (an example of a machine learning model) within the device 10 or external to the device 10, and output the stress level of the second speaker. The generative AI model estimates the stress level of the second speaker based on the input information and outputs the stress level. The second detection unit 32 may also acquire the stress level output by the generative AI model.
[0035] The adjustment unit 40 adjusts the audio based on the first status information and the second status information. The audio adjusted by the adjustment unit 40 is the audio of at least one of the first speaker and the second speaker. The adjustment unit 40 of this embodiment adjusts both the first audio and the second audio.
[0036] The adjustment unit 40 includes a voice conversion unit 41 and a voice output unit 42. The voice conversion unit 41 adjusts the first voice based on the first status information and the stress value of the second speaker. For example, when the information on the emotion of the first speaker in the first status information includes negative emotions such as "anger" and "sadness," the voice conversion unit 41 adjusts the first voice so that the degree of deviation between each of the indices of the first voice, including volume, pitch, pitch, and clarity, and each of the reference values in the emotion table corresponding to the emotion of the first speaker detected by the first detection unit 31 becomes smaller. In particular, the higher the stress value of the second speaker is above a predetermined threshold, or when the stress value of the second speaker gradually increases, the voice conversion unit 41 adjusts the first voice so that the degree of deviation between each of the indices of the first voice, including volume, pitch, pitch, and clarity, and each of the indices in the emotion table corresponding to the emotion of the first speaker detected by the first detection unit 31 becomes smaller.
[0037] The voice conversion unit 41 may adjust the first voice even if the information on the emotion of the first speaker in the first status information does not include a negative emotion. The voice conversion unit 41 may adjust each index of the first voice, including volume, pitch, pitch, and clarity, to a value that does not correspond to any emotion item in the emotion table. The voice conversion unit 41 may adjust each index of the first voice in accordance with the stress value of the second speaker, regardless of the information on the emotion of the first speaker in the first status information.
[0038] The voice conversion unit 41 adjusts the second voice based on the first status information and the second status information. For example, if the information on the emotion of the first speaker in the first status information includes negative emotions such as "anger" and "sadness," the voice conversion unit 41 adjusts each indicator of the second voice, including volume, pitch, pitch, and clarity, so as to soothe the emotion of the first speaker detected by the first detection unit 31. At this time, the voice conversion unit 41 may adjust each indicator of the second voice, including volume, pitch, pitch, and clarity, so as to produce a calm second adjusted voice. Furthermore, the voice conversion unit 41 may adjust each indicator of the second voice, including volume, pitch, pitch, and clarity, so as to reduce the degree of deviation between each indicator of the second voice, including volume, pitch, pitch, and clarity, and each reference value in the emotion table corresponding to the emotion of the second speaker detected by the second detection unit 32.
[0039] Furthermore, for example, if the information regarding the emotions of the second speaker in the second status information includes negative emotions such as "anger" and "sadness," the voice conversion unit 41 adjusts each indicator of the second voice, including volume, pitch, pitch, and clarity, so that the degree of deviation between the indicators and the reference values in the emotion table corresponding to the emotions of the second speaker detected by the second detection unit 32 becomes small.
[0040] Furthermore, for example, if the information regarding the second speaker's reaction in the second status information includes a negative reaction such as "impatience," the voice conversion unit 41 adjusts the indicators of the second voice, including volume, pitch, pitch, and clarity, so that the degree of deviation between them and the reference values in the reaction table corresponding to the second speaker's reaction detected by the second detection unit 32 becomes smaller.
[0041] The voice conversion unit 41 may adjust the second voice when the information on the emotion of the second speaker in the second status information does not include a negative emotion and when the information on the reaction of the second speaker does not include a negative emotion.The voice conversion unit 41 may adjust each index of the second voice, including volume, pitch, and clarity, to a value that does not correspond to each emotion item in the emotion table.The voice conversion unit 41 may adjust each index of the second voice, including volume, pitch, and clarity, to a value that does not correspond to each reaction item in the reaction table.The voice conversion unit 41 may adjust each index of the first voice according to the stress value of the second speaker.
[0042] The voice conversion unit 41 adjusts an index indicating the voice quality indicated by the first voice and the second voice. The voice conversion unit 41 adjusts, for example, at least one index of the volume, pitch, pitch, and clarity indicated by the first voice and the second voice. The voice conversion unit 41 adjusts at least one index included in the first voice and the second voice according to a predetermined priority. The priority may be predetermined, for example, by the second speaker. The priority may be, for example, in descending order of the degree of deviation of each index from the index of the voice in the emotion table indicated by the detected emotion. The voice conversion unit 41 stores the adjusted value of the at least one index.
[0043] The voice conversion unit 41 generates a first adjusted voice by adjusting an index of the voice quality indicated by the first voice. Note that the voice conversion unit 41 may, for example, adjust the first voice to generate the first adjusted voice such that it sounds like a person other than the first speaker is speaking, as with a voice changer. The voice output unit 42 outputs the first adjusted voice adjusted by the voice conversion unit 41 to the second terminal 12. The second terminal 12 receives the first adjusted voice.
[0044] The voice conversion unit 41 generates the second adjusted voice by adjusting the voice quality index indicated by the second voice. Note that the voice conversion unit 41 may, for example, adjust the second voice to generate the second adjusted voice such that it sounds like a person different from the second speaker is speaking, like a voice changer. The voice output unit 42 outputs the second adjusted voice adjusted by the voice conversion unit 41 to the first terminal 11. The first terminal 11 receives the second adjusted voice.
[0045] The voice conversion unit 41 may, for example, input at least one of the first voice and the second voice to a generative AI model (an example of a machine learning model) within the device 10 or external to the device 10, and output the first adjusted voice and the second adjusted voice. The generative AI model may have learned at least one of the above-mentioned voice quality indicators, emotion table, and reaction table. For example, based on the input voice, the generative AI model adjusts the voice so that the degree of deviation from each indicator of any emotion shown in the emotion table becomes small, and outputs the adjusted voice. The adjustment unit 40 may acquire at least one of the first adjusted voice and the second adjusted voice output by the generative AI model.
[0046] The processing procedure by the processing system 1 and device 10 configured as described above, i.e., the flow of the processing method according to this embodiment, will be described. Fig. 4 is a flowchart showing the procedure of an example of the processing method by the processing system. The processing method shown in Fig. 4 is started, for example, when a call start signal is received from a first terminal 11 operated by a first speaker to a second terminal 12 operable by a second speaker, and the device 10 receives information from the second terminal 12 indicating that the signal has been received. Note that the processing method may also be started at any timing during a conversation between the first speaker and the second speaker, when the device 10 receives a signal generated by an input operation by the second speaker to the second terminal 12.
[0047] In the processing method, first, in step S1, the speech acquisition unit 21 of the acquisition unit 20 acquires the speech of the first speaker. The speech acquisition unit 21 acquires, from the first terminal 11, the first speech input to the first terminal 11 by the first speaker.
[0048] Next, in step S2, the state acquisition unit 22 of the acquisition unit 20 acquires biometric information of the second speaker. When the first voice is acquired by the voice acquisition unit 21, the state acquisition unit 22 outputs a signal to the detection device 13 to acquire biometric information of the second speaker. The detection device 13 acquires the biometric information of the second speaker using the signal as a trigger. The state acquisition unit 22 acquires the biometric information of the second speaker while the first speaker is speaking.
[0049] Next, in step S3, the first detection unit 31 of the detection unit 30 detects first state information. The first detection unit 31 acquires, for example, information on emotions, which is an example of state information, from the user state database 14. The first detection unit 31 detects the emotion indicated by the first voice as first state information based on the indicators indicated by the first voice and the information on emotions acquired from the user state database 14.
[0050] Next, in step S4, the second detection unit 32 of the detection unit 30 detects second state information. The second detection unit 32, for example, acquires at least one of information on emotions and information on reactions, which are examples of state information, from the user state database 14. The second detection unit 32 may detect the emotion indicated by the second voice as the first state information based on each indicator indicated by the second voice and the information on emotions acquired from the user state database 14. The second detection unit 32 may detect the reaction indicated by the second speaker's biometric information as the second state information based on each indicator indicated by the second speaker's biometric information and the information on reactions acquired from the user state database 14. The second detection unit 32 estimates the stress level of the second speaker based on the second state information.
[0051] Next, in step S5, the speech conversion unit 41 of the adjustment unit 40 selects an index of the speech to be adjusted. The speech conversion unit 41 selects at least one index included in the first speech in accordance with a preset priority.
[0052] Next, in step S6, the voice conversion unit 41 of the adjustment unit 40 adjusts the voice (first voice) of the first speaker based on the first state information and the second state information, or the first state information and the stress value.
[0053] Next, in step S7, the audio output unit 42 of the adjustment unit 40 outputs the adjusted audio of the first speaker (first adjusted audio). The audio output unit 42 outputs the first adjusted audio to the second terminal 12. The second speaker listens to the first adjusted audio via the second terminal 12. The first adjusted audio has values for the indicators selected in step S5 changed from the first audio. The content of the first adjusted audio is the same as the content of the first audio.
[0054] In step S8, the voice acquisition unit 21 of the acquisition unit 20 determines whether or not the voice of the second speaker (second voice) has been acquired from the second terminal 12. For example, if the voice acquisition unit 21 has not acquired the second voice from the second terminal 12 for a predetermined period, such as when the call ends, the voice acquisition unit 21 determines that the second voice has not been acquired from the second terminal 12 (step S8: NO), and ends the flowchart shown in FIG. 4 , which is the processing method by the processing system 1 and the device 10.
[0055] If the voice acquisition unit 21 acquires the second voice from the second terminal 12 during a predetermined period, such as when the second speaker answers and the call continues, the voice acquisition unit 21 determines that the second voice has been acquired from the second terminal 12 (step S8: YES) and proceeds to step S9.
[0056] If it is determined that the second voice has been acquired by the voice acquisition unit 21 (step S8: YES), the adjustment unit 40 determines whether to adjust the voice of the second speaker (second voice) in step S9. For example, if the emotion of the second speaker detected in step S4 is not a negative emotion, or if the reaction of the second speaker is not a negative reaction, the adjustment unit 40 determines that there is no need to adjust the second voice (step S9: NO), and the voice output unit 42 of the adjustment unit 40 outputs the second voice to the first terminal 11 as is without adjusting it. After the second voice is output, the flowchart shown in FIG. 4, which is a processing method by the processing system 1 and the device 10, ends.
[0057] For example, if the emotion of the second speaker detected in step S4 is detected to be a negative emotion, or if the reaction of the second speaker is detected to be a negative reaction, the adjustment unit 40 determines that the second voice needs to be adjusted (step S9: YES) and proceeds to step S10.
[0058] If the adjustment unit 40 determines to adjust the second voice (step S9: YES), the voice conversion unit 41 of the adjustment unit 40 adjusts the voice of the second speaker (second voice) at step S10. The voice conversion unit 41 adjusts the second voice based on the first state information and the second state information, or the first state information and the stress value.
[0059] Next, in step S11, the audio output unit 42 of the adjustment unit 40 outputs the adjusted audio of the second speaker (second adjusted audio). The audio output unit 42 outputs the second adjusted audio to the first terminal 11. The first speaker listens to the second adjusted audio via the first terminal 11. The second adjusted audio has values for the indices selected in step S5 changed from the second audio. The content of the second adjusted audio is the same as the content of the second audio. Note that, before step S10, the audio conversion unit 41 of the adjustment unit 40 may select an audio indices to adjust for the second audio. When step S10 is completed, the flowchart shown in FIG. 4, which is a processing method by the processing system 1 and the device 10, is terminated.
[0060] Next, the effects of the device and method of the present disclosure will be described with reference to an example of a conventional problem. For example, in a workplace where voice terminals are used, such as a call center or a front desk, an operator may experience negative emotions from various customers or may have to deal with problems brought to them by customers. In such cases, the psychological burden on the operator dealing with customers is very large, and therefore a technology for reducing the psychological burden on the operator is desired.
[0061] Furthermore, in communication between speakers via audio terminals, the state of one speaker can affect the speech uttered by that speaker and the impression of that speech, and the state of the other speaker listening to that speech can also affect how the speech uttered by that speaker is perceived. The other speaker may wish to achieve smooth communication at least for the other speaker, regardless of the state of the first speaker. Therefore, there is a demand for technology that supports smooth communication by performing appropriate processing according to the state of the speakers.
[0062] The device 10 of the present disclosure comprises an acquisition unit 20 that acquires the voice of a first speaker and biometric information of a second speaker who is conversing with the first speaker via a voice terminal (first terminal 11 and second terminal 12), a first detection unit 31 that detects information related to the state based on the voice of the first speaker (first voice), a second detection unit 32 that detects information related to the state based on the biometric information of the second speaker, and an adjustment unit 40 that adjusts the voice based on the information related to the state of the first speaker (first state information) and the information related to the state of the second speaker (second state information).
[0063] The method disclosed herein also includes steps of acquiring the voice of a first speaker and biometric information of a second speaker who is conversing with the first speaker via an audio terminal (steps S1 and S2), detecting status-related information from the voice of the first speaker (first voice) (step S3), detecting status-related information from the biometric information of the second speaker (step S4), and adjusting the voice based on the status-related information of the first speaker (first status information) and the status-related information of the second speaker (second status information) (steps S5 to S7, S10, and S11).
[0064] The device 10 and method of the present disclosure adjust at least one of the first voice and the second voice based on the detected first state information and second state information. By performing appropriate processing on the voice depending on the state of the speakers, the device 10 and method can support and realize smooth conversation (communication) between the speakers via the voice terminals.
[0065] Furthermore, in the device 10 of the present disclosure, the voice adjusted by the adjustment unit 40 is the voice of at least one of the first speaker and the second speaker. In this case, for example, adjusting the first voice can reduce the psychological burden on the second speaker based on first state information, such as negative emotions from the first speaker. For example, adjusting the second voice can adjust the second voice based on second state information including at least one of the emotions and reactions of the second speaker, who has experienced psychological burden based on first state information, such as negative emotions from the first speaker. Therefore, the first speaker can receive the second adjusted voice, or the second speaker can receive the first adjusted voice, thereby realizing smooth conversation between speakers via the voice terminal.
[0066] In the device 10 of the present disclosure, the biometric information of the second speaker includes at least one of the following information: the voice of the second speaker, the biometric signal of the second speaker, and a moving image of the second speaker. In this case, the second state information can be detected from the biometric information of the second speaker through various media and various indicators, and at least one of the first voice and the second voice can be appropriately adjusted.
[0067] In the device 10 of the present disclosure, the acquisition unit 20 acquires information about the state of the second speaker after listening to the speech of the first speaker. In this case, the second detection unit 32 can detect changes in the emotions and reactions of the second speaker who listened to the first speech. Therefore, the adjustment unit 40 can appropriately adjust at least one of the first speech and the second speech based on the changes in the emotions and reactions of the second speaker.
[0068] Furthermore, in the device 10 of the present disclosure, the second detection unit 32 estimates the stress level of the second speaker based on information about the second speaker's state, and the adjustment unit 40 adjusts the voice based on the information about the first speaker's state and the stress level. In this case, the adjustment unit 40 can reflect the first state information and the stress level of the second speaker in its adjustment, thereby appropriately adjusting at least one of the first voice and the second voice.
[0069] Furthermore, in the device 10 of the present disclosure, the adjustment unit 40 adjusts at least one indicator of the volume, pitch, pitch, and clarity of the voice. In this case, the adjustment unit 40 can adjust at least one of the first voice and the second voice with respect to at least one of the indicators described above. The adjustment unit 40 is not limited to adjusting at least one indicator of the volume, pitch, pitch, and clarity, and may adjust at least one indicator indicating voice quality.
[0070] In the device 10 of the present disclosure, the adjustment unit 40 adjusts at least one indicator included in the audio according to a preset priority. In this case, the adjustment unit 40 can select an appropriate indicator according to the priority and adjust at least one of the first audio and the second audio.
[0071] In the device 10 of the present disclosure, the information about the state of the first speaker includes information about the emotion of the first speaker. In this case, the emotion of the first speaker can be estimated from the first voice, and at least one of the first voice and the second voice can be adjusted with respect to at least one of the above-described indicators in accordance with the emotion associated with the first voice.
[0072] The device and method of the present disclosure have the following configuration.
[0073] [1] A device comprising: an acquisition unit that acquires the voice of a first speaker and biometric information of a second speaker who is conversing with the first speaker via a voice terminal; a first detection unit that detects information related to the state of the first speaker based on the voice of the first speaker; a second detection unit that detects information related to the state of the second speaker based on the biometric information of the second speaker; and an adjustment unit that adjusts the voice based on the information related to the state of the first speaker and the information related to the state of the second speaker.
[0074] [2] The device according to [1], wherein the voice adjusted by the adjustment unit is the voice of at least one of the first speaker and the second speaker.
[0075] [3] The device described in [1] or [2] above, wherein the biometric information of the second speaker includes at least one of information on the voice of the second speaker, the biometric signal of the second speaker, and a video image of the second speaker.
[0076] [4] The device according to any one of [1] to [3], wherein the acquisition unit acquires information about the state of the second speaker after listening to the voice of the first speaker.
[0077] [5] The device according to any one of [1] to [4], wherein the adjustment unit estimates a stress value of the second speaker based on information about the second speaker's state, and adjusts the voice based on the information about the first speaker's state and the stress value.
[0078] [6] The device according to any one of [1] to [5] above, wherein the adjustment unit adjusts at least one indicator of volume, pitch, pitch, and clarity indicated by the sound.
[0079] [7] The device according to [6], wherein the adjustment unit adjusts the at least one indicator included in the audio according to a preset priority.
[0080] [8] The device according to any one of [1] to [7] above, wherein the information about the state of the first speaker includes information about the emotion of the first speaker.
[0081] [9] A method comprising the steps of: acquiring the voice of a first speaker and biometric information of a second speaker who is conversing with the first speaker via a voice terminal; detecting information about the state from the voice of the first speaker; detecting information about the state from the biometric information of the second speaker; and adjusting the voice based on the information about the state of the first speaker and the information about the state of the second speaker.
[0082] The block diagrams used to explain the above embodiments show functional blocks. These functional blocks (components) are realized by any combination of hardware and / or software. Furthermore, the method for realizing each functional block is not particularly limited. That is, each functional block may be realized using a single device that is physically or logically coupled, or may be realized using two or more physically or logically separated devices that are connected directly or indirectly (e.g., via wire, wirelessly, etc.) and these multiple devices. The functional block may also be realized by combining the single device or multiple devices with software.
[0083] Functions include, but are not limited to, judgment, determination, assessment, calculation, computation, processing, derivation, investigation, search, confirmation, reception, transmission, output, access, resolution, selection, selection, establishment, comparison, assumption, expectation, consideration, broadcasting, notifying, communicating, forwarding, configuring, reconfiguring, allocating, mapping, and assignment. For example, a functional block (component) that performs transmission is called a transmitting unit or transmitter. As mentioned above, there are no particular limitations on how these functions are implemented.
[0084] For example, the device 10 constituting the conversion system according to an embodiment of the present disclosure may function as a computer that performs processing of the control method of the present disclosure. FIG. 5 is a diagram illustrating an example of the hardware configuration of the device 10 according to an embodiment of the present disclosure. The above-described device 10 may be physically configured as a computer including a processor 1001, a memory 1002, a storage 1003, a communication device 1004, an input device 1005, an output device 1006, a bus 1007, and the like. The device 10 may be configured as a computer including at least one processor such as a CPU or a GPU, or may be configured as a computer including multiple processors or may include multiple computer devices. The first terminal 11, the second terminal 12, the detection device 13, and the like may also have a similar hardware configuration.
[0085] In the following description, the term "apparatus" can be interpreted as a circuit, a device, a unit, etc. The hardware configuration of apparatus 10 may be configured to include one or more of the apparatuses shown in the drawings, or may be configured to exclude some of the apparatuses.
[0086] Each function of the device 10 is realized by loading specified software (programs) onto hardware such as the processor 1001 and memory 1002, causing the processor 1001 to perform calculations, control communication via the communication device 1004, and control at least one of reading and writing data in the memory 1002 and storage 1003.
[0087] The processor 1001 controls the entire computer by running, for example, an operating system. The processor 1001 may be configured by a central processing unit (CPU) including an interface with peripheral devices, a control device, an arithmetic unit, a register, etc. For example, the above-mentioned acquisition unit 20, detection unit 30, adjustment unit 40, etc. may be realized by the processor 1001.
[0088] The processor 1001 also reads programs (program codes), software modules, data, etc. from at least one of the storage 1003 and the communication device 1004 into the memory 1002 and executes various processes in accordance with these programs. The programs used are those that cause a computer to execute at least some of the operations described in the above-described embodiments. For example, the acquisition unit 20, the detection unit 30, and the adjustment unit 40 may be implemented by a control program stored in the memory 1002 and running on the processor 1001, and similar implementations may be used for other functional blocks. While the above-described various processes have been described as being executed by a single processor 1001, they may also be executed simultaneously or sequentially by two or more processors 1001. The processor 1001 may be implemented by one or more chips. The programs may also be transmitted from a network via a telecommunications line.
[0089] The memory 1002 is a computer-readable recording medium and may be configured, for example, by at least one of a read-only memory (ROM), an erasable programmable ROM (EPROM), an electrically erasable programmable ROM (EEPROM), a random access memory (RAM), etc. The memory 1002 may also be called a register, a cache, a main memory (primary storage device), etc. The memory 1002 can store executable programs (program codes), software modules, etc. for implementing a control method according to an embodiment of the present disclosure.
[0090] Storage 1003 is a computer-readable recording medium, and may be composed of at least one of, for example, an optical disk such as a CD-ROM (Compact Disc ROM), a hard disk drive, a flexible disk, a magneto-optical disk (e.g., a compact disk, a digital versatile disk, a Blu-ray (registered trademark) disk), a smart card, a flash memory (e.g., a card, a stick, a key drive), a floppy (registered trademark) disk, a magnetic strip, etc. Storage 1003 may also be referred to as an auxiliary storage device. The above-mentioned storage medium may be, for example, a database, a server, or other appropriate medium including at least one of memory 1002 and storage 1003.
[0091] The communication device 1004 is hardware (transmission / reception device) for communicating between computers via at least one of a wired network and a wireless network, and is also referred to as, for example, a network device, a network controller, a network card, a communication module, etc. The communication device 1004 may be configured to include a high-frequency switch, a duplexer, a filter, a frequency synthesizer, etc. to realize at least one of frequency division duplex (FDD) and time division duplex (TDD). For example, the above-mentioned acquisition unit 20, detection unit 30, adjustment unit 40, etc. may be realized by the communication device 1004.
[0092] The input device 1005 is an input device (e.g., a keyboard, a mouse, a microphone, a switch, a button, a sensor, etc.) that accepts input from the outside. The output device 1006 is an output device (e.g., a display, a speaker, an LED lamp, etc.) that outputs to the outside. Note that the input device 1005 and the output device 1006 may be integrated into one device (e.g., a touch panel).
[0093] Furthermore, each device, such as the processor 1001 and the memory 1002, is connected by a bus 1007 for communicating information. The bus 1007 may be configured using a single bus, or may be configured using different buses between each device.
[0094] The device 10 may also be configured to include hardware such as a microprocessor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a programmable logic device (PLD), or a field programmable gate array (FPGA), and some or all of the functional blocks may be realized by the hardware. For example, the processor 1001 may be implemented using at least one of these pieces of hardware.
[0095] The notification of information is not limited to the aspects / embodiments described in the present disclosure and may be performed using other methods. For example, the notification of information may be performed by physical layer signaling (e.g., Downlink Control Information (DCI) and Uplink Control Information (UCI)), higher layer signaling (e.g., Radio Resource Control (RRC) signaling, Medium Access Control (MAC) signaling, broadcast information (Master Information Block (MIB) and System Information Block (SIB))), other signals, or a combination thereof. Furthermore, the RRC signaling may be referred to as an RRC message, and may be, for example, an RRC Connection Setup message, an RRC Connection Reconfiguration message, or the like.
[0096] The order of the procedures, sequences, flowcharts, etc. of each aspect / embodiment described in this disclosure may be changed unless it is consistent. For example, the methods described in this disclosure present elements of various steps using an example order, and are not limited to the particular order presented.
[0097] Input and output information may be stored in a specific location (for example, memory) or may be managed using a management table. Input and output information may be overwritten, updated, or added to. Output information may be deleted. Input information may be sent to another device.
[0098] The determination may be made based on a value represented by one bit (0 or 1), a Boolean value (true or false), or a numerical comparison (e.g., comparison with a predetermined value).
[0099] The aspects / embodiments described in this disclosure may be used alone, in combination, or switched depending on the implementation. Notification of predetermined information (e.g., notification that "X is true") is not limited to explicit notification, but may be implicit (e.g., not notifying the predetermined information).
[0100] Although the present disclosure has been described in detail above, it is clear to those skilled in the art that the present disclosure is not limited to the embodiments described herein. The present disclosure can be implemented in modified and altered forms without departing from the spirit and scope of the present disclosure as defined by the claims. Therefore, the description of the present disclosure is intended to be illustrative and does not have any limiting meaning on the present disclosure.
[0101] Software shall be construed broadly to mean instructions, instruction sets, code, code segments, program code, programs, subprograms, software modules, applications, software applications, software packages, routines, subroutines, objects, executable files, threads of execution, procedures, functions, etc., whether referred to as software, firmware, middleware, microcode, hardware description language, or otherwise.
[0102] Software, instructions, information, etc. may also be transmitted or received over a transmission medium. For example, if software is transmitted from a website, server, or other remote source using wired technologies (such as coaxial cable, fiber optic cable, twisted pair, Digital Subscriber Line (DSL)), and / or wireless technologies (such as infrared, microwave), then these wired and / or wireless technologies are included within the definition of transmission media.
[0103] The information, signals, etc. described in this disclosure may be represented using any of a variety of different technologies. For example, data, instructions, commands, information, signals, bits, symbols, chips, etc. that may be referred to throughout the above description may be represented by voltages, currents, electromagnetic waves, magnetic fields or magnetic particles, optical fields or photons, or any combination thereof.
[0104] Note that terms described in this disclosure and terms necessary for understanding this disclosure may be replaced with terms having the same or similar meanings. For example, at least one of a channel and a symbol may be a signal (signaling). Furthermore, a signal may be a message. Furthermore, a component carrier (CC) may be called a carrier frequency, a cell, a frequency carrier, etc.
[0105] Furthermore, the information, parameters, etc. described in the present disclosure may be expressed using absolute values, may be expressed using relative values from a predetermined value, or may be expressed using other corresponding information. For example, a radio resource may be indicated by an index.
[0106] The names used for the above-described parameters are not intended to be limiting in any way. Furthermore, the mathematical expressions using these parameters may differ from those explicitly disclosed in this disclosure. The various channels (e.g., PUCCH, PDCCH, etc.) and information elements may be identified by any suitable names, and therefore the various names assigned to these various channels and information elements are not intended to be limiting in any way.
[0107] In this disclosure, the terms "Mobile Station (MS)," "user terminal," "User Equipment (UE)," "terminal," and the like may be used interchangeably.
[0108] A mobile station may also be referred to by those skilled in the art as a subscriber station, mobile unit, subscriber unit, wireless unit, remote unit, mobile device, wireless device, wireless communication device, remote device, mobile subscriber station, access terminal, mobile terminal, wireless terminal, remote terminal, handset, user agent, mobile client, client, or some other suitable terminology.
[0109] As used in this disclosure, the terms "determining" and "determining" may encompass a wide variety of actions. "Determining" and "determining" may include, for example, judging, calculating, computing, processing, deriving, investigating, looking up, searching, inquiring (e.g., searching in a table, database, or other data structure), ascertaining, and the like. "Determining" and "determining" may also include receiving (e.g., receiving information), transmitting (e.g., sending information), input, output, accessing (e.g., accessing data in memory), and the like. Furthermore, "judgment" and "decision" can include regarding resolving, selecting, choosing, establishing, comparing, etc. as having been "judged" or "decided." In other words, "judgment" and "decision" can include regarding some action as having been "judged" or "decided." Furthermore, "judgment (decision)" can be interpreted as "assuming," "expecting," "considering," etc.
[0110] The terms "connected," "coupled," or any variation thereof, refer to any direct or indirect connection or coupling between two or more elements, and may include the presence of one or more intermediate elements between two elements that are "connected" or "coupled" to each other. The coupling or connection between elements may be physical, logical, or a combination thereof. For example, "connected" may be read as "access." As used in this disclosure, two elements may be considered to be "connected" or "coupled" to each other using one or more wires, cables, and / or printed electrical connections, as well as electromagnetic energy having wavelengths in the radio frequency range, microwave range, and optical (both visible and invisible) range, as some non-limiting and non-exhaustive examples.
[0111] As used in this disclosure, the phrase "based on" does not mean "based only on," unless expressly stated otherwise. In other words, the phrase "based on" means both "based only on" and "based at least on."
[0112] As used in this disclosure, any reference to an element using a designation such as "first," "second," etc. does not generally limit the quantity or order of those elements. These designations may be used in this disclosure as a convenient method of distinguishing between two or more elements. Thus, a reference to a first and a second element does not imply that only two elements may be employed or that the first element must in some way precede the second element.
[0113] When the terms "include," "including," and variations thereof are used in this disclosure, these terms are intended to be inclusive, similar to the term "comprising." Furthermore, when the term "or" is used in this disclosure, it is not intended to be an exclusive or.
[0114] In this disclosure, where articles are added by translation, such as a, an, and the in English, the disclosure may include that the nouns following these articles are in the plural form.
[0115] In the present disclosure, the term "A and B are different" may mean "A and B are different from each other." The term may also mean "A and B are each different from C." Terms such as "separate" and "coupled" may also be interpreted in the same way as "different."
[0116] 1...processing system, 10...device, 11...first terminal, 12...second terminal, 13...detection device, 14...user status database, 20...acquisition unit, 21...voice acquisition unit, 22...status acquisition unit, 30...detection unit, 31...first detection unit, 32...second detection unit, 40...adjustment unit, 41...voice conversion unit, 42...voice output unit.
Claims
1. A device comprising: an acquisition unit that acquires the voice of a first speaker and biometric information of a second speaker who is conversing with the first speaker via a voice terminal; a first detection unit that detects information related to the state of the first speaker based on the voice of the first speaker; a second detection unit that detects information related to the state of the second speaker based on the biometric information of the second speaker; and an adjustment unit that adjusts the voice based on the information related to the state of the first speaker and the information related to the state of the second speaker.
2. The device according to claim 1, wherein the voice adjusted by the adjustment unit is the voice of at least one of the first speaker and the second speaker.
3. The device described in claim 1, wherein the biometric information of the second speaker includes at least one of information on the voice of the second speaker, biometric signals of the second speaker, and video images of the second speaker.
4. The device according to claim 1, wherein the acquisition unit acquires information about the state of the second speaker after listening to the speech of the first speaker.
5. The device described in claim 1, wherein the second detection unit estimates the stress value of the second speaker based on information about the state of the second speaker, and the adjustment unit adjusts the voice based on the information about the state of the first speaker and the stress value.
6. The device according to claim 1, wherein the adjustment unit adjusts at least one indicator of volume, pitch, tone, and clarity of the audio.
7. The device according to claim 6, wherein the adjustment unit adjusts the at least one indicator included in the audio according to a preset priority.
8. The device of claim 1, wherein the information about the state of the first speaker includes information about the emotion of the first speaker.
9. A method comprising the steps of: acquiring the voice of a first speaker and biometric information of a second speaker who is conversing with the first speaker via a voice terminal; detecting information about the state from the voice of the first speaker; detecting information about the state from the biometric information of the second speaker; and adjusting the voice based on the information about the state of the first speaker and the information about the state of the second speaker.
Citation Information
Patent Citations
System and program for voice conversion
JP2004252085A
Automatic voice recognition / voice conversion system
JP2014095753A
Voice characteristic change system and voice characteristic change method
JP2021107873A
Program, information processing device, and information processing method
JP2023105607A
Audio processing system, audio processing device, and audio processing method
JP7164793B1