Control device, control method, and control program
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- NT T INC
- Filing Date
- 2024-11-15
- Publication Date
- 2026-05-21
Smart Images

Figure JP2024040598_21052026_PF_FP_ABST
Abstract
Description
Control device, control method, and control program
[0001] The present invention relates to a control device, a control method, and a control program.
[0002] The practical application of systems in which humanoid computer graphics (CG) or robots (hereinafter referred to as agents) with bodily expressions engage in voice dialogue is progressing. In such systems, in order to aim for more human-like dialogue, methods have been devised to control the selection of nodding and other actions that accompany affirmative responses (hereinafter referred to as affirmative actions) in response to the user's utterances by the agent, as well as the speed of these affirmative actions (Non-Patent Documents 1, 2).
[0003] Conventional technology extracts text information and prosodic information from input speech, estimates appropriate actions from an action database (DB) based on this information, and outputs action control information for the agent.
[0004] For example, in conventional technology, within a speech interval, action determination is performed to estimate the time when an interjection occurs and to select an action based on text information and prosodic information derived from the input speech. Furthermore, conventional technology obtains an action identifier and action start time based on the action determination, retrieves an action command from the action database, and outputs an action command to the agent to execute the action command at the action time. In addition, conventional technology may pre-set multiple speed stages and corresponding action times for each speed stage, and output an action command to execute the action command at the action time corresponding to the speed stage.
[0005] Koji Inoue et al., "Android ERICA's Listening Dialogue System - Comparative Evaluation with Human Listening," Transactions of the Japanese Society for Artificial Intelligence, Vol. 36, No. 5H, 2021. Masato Yokota et al., "Voice-Driven Body Engagement System that Performs Reaction Actions According to Mora-Based Speech Rate," Transactions of the Japan Society of Mechanical Engineers, Vol. 89, No. 919, 2023.
[0006] For an agent to make a conversation with a human feel more friendly and cooperative, like "co-talk" (Reference 1), it is important for the humanoid agent to show more active participation in the conversation. Reference 1: Yasuko Sasaki, "A Study on the Theory of 'Co-talk'", Language and Culture and Japanese Language Education 9 47-59, 1995.
[0007] To demonstrate active engagement in a conversation, it is important for the agent to nod or give verbal cues. However, if the other person finishes speaking while the agent is nodding, and the start of a relatively meaningful utterance following the nod is delayed, it can give the impression that the agent is not actively participating in the conversation.
[0008] For example, if an agent performs a long nodding motion just before a change of speech, preventing the user from moving on to the next utterance, it can sound unnatural. Therefore, it is desirable to limit the length of these nodding motions.
[0009] The present invention has been made in view of the above, and aims to provide a control device, a control method, and a control program that enable the selection of appropriate operation of an agent and the determination of appropriate operation time.
[0010] To solve the above-mentioned problems and achieve the objective, the control device of the present invention is characterized by comprising: an estimation unit that performs speech turn estimation to estimate whether it is the agent's turn to speak based on the user's voice information and the user's video information; a determination unit that determines the operation identifier and operation start time of the agent's operation based on the user's voice information, the user's video information, and the estimation result of the speech turn estimation; and a control unit that determines the operation time of the operation determined by the determination unit based on at least the user's speech rate and outputs an operation command for the agent.
[0011] Furthermore, the control method of the present invention is a control method executed by a control device, and is characterized by including the steps of: performing speech turn estimation, which estimates whether it is the agent's turn to speak based on the user's voice information and the user's video information; determining the operation identifier and operation start time of the agent's operation based on the user's voice information, the user's video information, and the estimation result of the speech turn estimation; and determining the operation time of the operation determined in the determination step, based on at least the user's speech rate, and outputting an operation command for the agent.
[0012] Furthermore, the control program of the present invention causes a computer to perform the following steps: 1) estimate whether it is the agent's turn to speak based on the user's voice information and the user's video information; 2) determine the operation identifier and start time of the agent's operation based on the user's voice information, the user's video information, and the estimation result of the speech turn estimation; and 3) determine the operation time of the operation determined in the determination step, based at least on the user's speech rate, and output an operation command for the agent.
[0013] According to the present invention, it is possible to improve the accuracy with which the agent interprets the user's statements and enable the expression of more appropriate response actions.
[0014] Figure 1 is a diagram illustrating the response control by the control method according to the embodiment. Figure 2 is a diagram showing an example of the configuration of the control device according to the embodiment. Figure 3 is a diagram illustrating the operation identifier. Figure 4 is a diagram illustrating the operation identifier. Figure 5 is a diagram illustrating the operation identifier. Figure 6 is a diagram showing the correspondence table between the score for turn management willingness or turn change and the judgment value. Figure 7 is a diagram showing the correspondence table between the score for turn management willingness or turn change and the judgment value. Figure 8 is a diagram showing the correspondence table between the score for turn management willingness or turn change and the judgment value. Figure 9 is a diagram showing an example of a response label. Figure 10 is a diagram showing an example of the correspondence table between the output value of the agent utterance turn estimation unit, the changed operation identifier, and the operation command. Figure 11 is a diagram showing an example of the correspondence table between the output value of the agent utterance turn estimation unit, the changed operation identifier, and the operation command. Figure 12 is a diagram showing an example of the correspondence table between the output value of the agent utterance turn estimation unit, the changed operation identifier, and the operation command. Figure 13 is a diagram illustrating the calculation process of operation time. Figure 14 is a diagram illustrating the process for calculating the operating time. Figure 15 is a diagram illustrating the process for calculating the operating time. Figure 16 is a diagram illustrating the output of the determination result by the agent operation determination unit and the operation command by the agent operation control unit. Figure 17 is a flowchart showing the processing procedure of the control method according to the embodiment. Figure 18 is a flowchart showing an example of the processing procedure of the agent operation determination process shown in Figure 17. Figure 19 is a flowchart showing an example of the processing procedure of the agent operation control process shown in Figure 17. Figure 20 is a diagram showing an example of a computer in which a control device is realized when a program is executed.
[0015] Hereinafter, one embodiment of the present invention will be described in detail with reference to the drawings. However, the present invention is not limited by this embodiment. Furthermore, in the drawings, the same parts are denoted by the same reference numerals.
[0016] [Embodiment] [Overview of Control Device] The control device estimates speech turns based on prosody, text, and video, and selects a nodding action for a humanoid CG (Computer Graphics) or robot (hereinafter referred to as "agent") with bodily expression based on the speech turn estimation result.
[0017] The control device suppresses the length of the nodding motion because if the agent performs a long nodding motion just before switching to speaking, it would sound unnatural and prevent them from moving on to speaking. Figure 1 is a diagram illustrating the nodding motion control according to the control method of the embodiment.
[0018] In this embodiment, for example, if the user stops speaking in anticipation of the agent speaking (Figure 1(1)), the agent will not use a fixed-length nod as in the conventional method, but will instead use a shorter nod in anticipation of a change in the speaking turn (Figure 1(2)). In this way, the control device according to this embodiment controls the length of the nod so that the agent can start speaking immediately after the other person finishes speaking, without any pause, even while giving a nod.
[0019] [Control device configuration] The control device according to the embodiment will now be described. Figure 2 is a diagram showing an example of the configuration of the control device according to the embodiment.
[0020] The control device 10 according to this embodiment is a control device that is realized by loading a predetermined program into a computer or the like, which includes ROM (Read Only Memory), RAM (Random Access Memory), CPU (Central Processing Unit), etc., and having the CPU execute the predetermined program. The control device 10 also has a communication interface for sending and receiving various information with other devices connected via a network or the like.
[0021] The control device 10 includes a user voice input unit 11, a user video input unit 12, a user voice recognition unit 13, a user prosodic information extraction unit 14, a user speech rate calculation unit 15, an agent speech turn estimation unit 16 (estimation unit), an agent action determination unit 17 (determination unit), an agent action DB 18, an agent action control unit 19 (control unit), and an agent action output unit 20. The control device 10 takes the user's voice and video footage of the user as input and outputs an action representation of the agent responding to the user. Note that each functional unit of the control device 10 may be implemented as a single piece of hardware, or it may be configured so that multiple pieces of hardware share the role of each functional unit.
[0022] The user voice input unit 11 is implemented using a microphone. The user voice input unit 11 receives the user's voice input and outputs the user's voice information to the user voice recognition unit 13 and the user prosodic information extraction unit 14.
[0023] The user video input unit 12 is implemented using a camera. The user video input unit 12 receives incident light with an image sensor and converts it into an image. The image is of the user. The user video input unit 12 outputs the image information of the user to the agent speech turn estimation unit 16 and the agent action determination unit 17.
[0024] The user speech recognition unit 13 receives user voice information as input. The user speech recognition unit 13 extracts text information and speech segment information from the user's voice information using speech recognition technology. As speech recognition technology, the user speech recognition unit 13 uses, for example, the BERT (Bidirectional Encoder Representations from Transformers) model described in Reference 2. The user speech recognition unit 13 outputs the text information and speech segment information, which are the speech recognition results, to the user speech rate calculation unit 15, the agent speech turn estimation unit 16, and the agent action determination unit 17. Reference 2: Ishii, Ryo et al., "Simultaneous Prediction of Turn Management Intention and Actual Turn Alternation Using Multimodal Features," Institute of Electronics, Information and Communication Engineers, 2021.
[0025] The user prosodic information extraction unit 14 takes the user's voice information as input. The user prosodic information extraction unit 14 extracts prosodic information (e.g., voice features) from the user's voice information. The user prosodic information extraction unit 14 extracts voice features using, for example, the VGGish model described in Reference 1. The user prosodic information extraction unit 14 outputs the extracted prosodic information to the agent utterance turn estimation unit 16 and the agent action determination unit 17.
[0026] The user speech rate calculation unit 15 receives text information and speech segment information output by the user speech recognition unit 13 as input. The user speech rate calculation unit 15 calculates the user's speech rate for each speech segment. The user speech rate calculation unit 15 calculates the speech rate [mora / s] using, for example, the method described in Non-Patent Document 2. The user speech rate calculation unit 15 outputs the calculated speech rate to the agent operation determination unit 17. Note that mora is a unit of syllables, as will be described later.
[0027] The agent speech turn estimation unit 16 performs speech turn estimation based on the user's voice information and the user's video information to estimate whether it is the agent's turn to speak. The agent speech turn estimation unit 16 calculates the user's willingness to speak and / or listen (willingness to manage turns) or whether there is a change in speech turns, based on speech interval information, text information, prosodic information, and video information, and estimates whether it is the agent's turn to speak. The user's willingness to speak and / or listen (willingness to manage turns) and whether there is a change in speech turns can both be continuous values such as probability values, discrete values, or truth values. The agent speech turn estimation unit 16 outputs the estimation result of the speech turn estimation to the agent operation determination unit 17 and the agent operation control unit 19.
[0028] The agent action determination unit 17 determines the action identifier and start time of the agent's action based on the user's voice information, the user's video information, and the estimation result of the speech turn estimation. The agent action determination unit 17 determines the action identifier and start time of the action based on the video information output from the user video input unit 12, the speech interval information and text information output from the user voice recognition unit 13, the prosodic information output from the user prosodic information extraction unit 14, and the speech turn estimation result output from the agent speech turn estimation unit 16. The action identifier is an identifier that can distinguish between actions accompanying verbal acknowledgments (verbal acknowledgment actions) and actions other than verbal acknowledgment actions (standby actions). The start time of the action is the start time of the verbal acknowledgment action or the standby action. The agent action determination unit 17 determines that the agent's action is either a verbal acknowledgment action or a standby action.
[0029] The agent action determination unit 17 determines the agent's action based on the result of the speech turn estimation and suppresses the nodding action if the agent should not perform it. Based on the estimation result of the speech turn estimation, the agent action determination unit 17 suppresses the nodding action and changes the agent's action to a standby action if the agent should not perform it. The agent action determination unit 17 outputs the action identifier and the action start time as the action determination result to the agent action control unit 19.
[0030] The agent operation DB18 stores the agent's operation commands and standard operation times along with identifiers. The identifiers are those that can distinguish between response operations and standby operations.
[0031] The agent operation control unit 19 determines the operation time of the operation determined by the agent operation determination unit 17, based at least on the user's speech rate, and outputs an operation command for the agent.
[0032] The agent operation control unit 19 determines the length of the agent's operation time based on the determination result from the agent operation determination unit 17, the estimation result of the speech turn estimation, and the user's speech rate.
[0033] For example, the agent operation control unit 19 determines the operation time for an operation command corresponding to an operation identifier determined by the agent operation determination unit 17, based on the standard operation time stored in the agent operation DB 18 and the user's speech rate calculated by the user speech rate calculation unit 15, and outputs the operation and operation time as an operation command to the agent operation output unit 20. The agent operation control unit 19 also determines the operation time based on the standard operation time, the user's speech rate, and the estimated result of the speech turn estimation, and outputs the operation and operation time as an operation command to the agent operation output unit 20. Several examples of operation time calculation will be described later.
[0034] The agent operation control unit 19 outputs a standby operation as an operation command during periods when no operation has been determined by the agent operation determination unit 17, that is, during periods when no operation has been determined by the agent operation determination unit 17.
[0035] The agent operation control unit 19 does not output an operation command if the agent is continuing the previous operation and this previous operation is a nodding operation, but outputs an operation command if the previous operation was a standby operation and the operation command is a nodding operation. In other words, if the agent operation control unit 19 receives an input for a nodding operation that overlaps in time while a standby operation is being output, it interrupts the standby operation and outputs a nodding operation.
[0036] It is preferable that the agent operation control unit 19 sets the operation commands in the agent operation DB 18 so that the start and end of each operation are always at the same position, so that the operations appear continuous rather than intermittent. Alternatively, it is preferable that the agent operation control unit 19 gradually changes the operation using a moving average or the like for a certain period of time before and after the operation switch.
[0037] The agent operation control unit 19 preferably outputs a speech action generated by a method other than the one described above, as needed, during the output of the standby operation, by interrupting the standby operation or combining it with the standby operation.
[0038] The agent motion output unit 20 takes an operation command as input and outputs it as the motion expression of an agent (such as a CG animation or a robot).
[0039] For adjusting, evaluating, analyzing, etc. the motion of the agent, the agent motion output unit 20 can also display, together with the motion, the turn prediction result and the speech rate used in the motion determination and motion control of the agent.
[0040] [Motion identifier] The motion identifier will be described. As described above, the motion identifier is an identifier that can distinguish between a nodding motion and a standby motion.
[0041] Figures 3 to 5 are diagrams for explaining the motion identifier. The correspondence relationships in Figures 3 to 5 may be, for example, preset or updated as appropriate.
[0042] As shown in Figure 3, the motion identifier is, for example, associated with each fixed motion command. Also, as shown in Figure 4, the motion identifier may be one that can perform separate processing for adjusting each fixed motion command.
[0043] The motion command may be accompanied by parameters such as "move the neck angle down by 10 degrees and then back". Also, the motion command may be one that plays a fixed motion, such as a video, without accompanying parameters.
[0044] The standby motion includes, for example, involuntary motions. Involuntary motions are, for example, "stay still and not move at all", "blink", "slightly sway the body", etc. Also, the standby motion may include motions that are somewhat intentional, such as "cross the hands", "change the way of crossing the hands", "change the standing posture", etc.
[0045] Also, as shown in Figure 5, the motion command may combine multiple commands. Also, the motion command may include not only body movements but also voice emission.
[0046] [Processing of the Agent Speech Turn Estimation Unit] Next, the agent utterance turn estimation unit 16 will be described. The agent utterance turn estimation unit 16 takes text and speech interval information, prosodic information, and video information as input, calculates a score for the willingness to speak and / or listen (willingness to manage turns), or whether or not there is a turn change, and then outputs an estimated value for whether it is the agent's turn to speak.
[0047] For example, the technique described in Reference 2 can be used as a method for calculating the estimated value. Alternatively, other techniques may also be used as methods for calculating the estimated value.
[0048] Figures 6 to 8 show a correspondence table between the score for turn management motivation or turn change status and the judgment value.
[0049] As shown in Figure 6, the agent utterance turn estimation unit 16 may calculate the estimated value (score) as a continuous value such as 0.0 to 1.0 (score x in the left column of Figure 6), or as a discrete value such as 5 levels (score y in the left column of Figure 7). The agent utterance turn estimation unit 16 uses, for example, the technology described in Reference 2 to determine the probability that the utterance will continue during the agent's turn as score x, and the probability that the agent will start uttering due to the turn change during the user's turn as score x. The agent utterance turn estimation unit 16 also uses, for example, the technology described in Reference 2 to obtain the maximum values of the user's willingness to turn-yield and their willingness to listen as score y.
[0050] The agent utterance turn estimation unit 16 may determine score x or score y at an arbitrary threshold and output an estimated value as a truth value (determined value) (right column in Figures 6 and 7).
[0051] The agent utterance turn estimation unit 16 outputs a truth value, setting it to True when it is the turn the agent will utter, and otherwise, a larger value indicates a higher probability that the agent will utter.
[0052] Furthermore, as shown in Figure 8, the agent utterance turn estimation unit 16 may calculate multiple values (scores x and y) and combine them for an overall evaluation. The agent utterance turn estimation unit 16 may also output a combination of multiple values. For example, in the example in Figure 8, the agent utterance turn estimation unit 16 outputs both score x and the overall evaluation. The agent utterance turn estimation unit 16 may also output other combinations (scores x and y, score y and overall evaluation, or scores x, y and overall evaluation).
[0053] In the processing of the agent utterance turn estimation unit 16, any of the input information for the agent utterance turn estimation unit 16, namely text and utterance segment information, prosodic information, and video information, can be omitted, but at least one must be used. In addition, the processing of the agent utterance turn estimation unit 16 may similarly acquire text and utterance segment information, prosodic information, and video information for the agent and other users and use them as input information.
[0054] Furthermore, the agent utterance turn estimation unit 16 can also take text and utterance interval information, prosodic information, and video information as input, extract motion responses (user's head movements and gaze), estimate the level of comprehension, and then output an estimated agent utterance turn value based on the results (References 3, 4). Here, the level of comprehension indicates the degree to which the user understands the content of the conversation with the agent. Reference 3: Ishii, Ryo et al., "Prediction of the next speaker based on head movements in multi-person dialogue", Transactions of the Information Processing Society of Japan, Vol. 57, No. 4, 1116-1127, 2016. Reference 4: Shunichi Kinoshita, et al., "A Study of Prediction of Listener's Comprehension Based on Multimodal Information", Proc. the 23rd ACM IVA '23, Article No. 30, pp. 1-4, 2023.
[0055] For example, if m is a score based on motor response (head movement, gaze) and c is the level of comprehension, the output value o is expressed by equation (1). The value inside {} is the value of c with its sign reversed and normalized to between 0 and 1.
[0056]
[0057] Furthermore, according to Reference 3, score m represents the probability that an agent's utterance will continue during the agent's turn, and the probability that an agent's utterance will begin due to a turn change during the user's turn. Comprehension level c, according to Reference 4, is obtained as a score from -2 to 2 representing the user's level of comprehension. In this case, comprehension level c is positive, and a higher value indicates that the user understands better.
[0058] [Processing of the Agent Operation Determination Unit] [Example of Selected Operation Determination] Next, the processing of the agent operation determination unit 17 will be explained.
[0059] The agent action determination unit 17 estimates an interjection label for each user utterance, for example, by the method described in Reference 5. Figure 9 shows an example of an interjection label. The agent action determination unit 17 estimates the interjection label based on the input text information, prosodic information, and video information, using a machine learning model that has learned the relationship between the speaker's utterance and video and the listener's interjection labels. The machine learning model is, for example, a model in which machine learning has been performed using text information obtained by converting the speaker's utterance into text, speech features extracted from the speaker's utterance data, video information during the speaker's utterance, and the interjection labels at that time as training data. When the machine learning model receives text information obtained by converting the speaker's utterance into text, speech features extracted from the speaker's utterance data, and video information during the speaker's utterance as input, it outputs an interjection label corresponding to the input data. For example, the interjection label is associated with the interjection action and the start time of the action. Based on the estimated interjection label, the agent action determination unit 17 tentatively selects the interjection action and the start time of the action. Reference 5: Naoki Higashi et al., "Basic Study on the Generation of Diverse Responses Based on Multimodal Information," Information Processing Society of Japan Research Report, Vol.2024-GN-119, No.8, 1-6, 2023.
[0060] The agent action determination unit 17 acquires the speech turn estimation result and modifies the provisionally selected action based on the acquired speech turn estimation result.
[0061] Figures 10 to 12 show an example of a correspondence table between the output value of the agent speech turn estimation unit, the modified operation identifier, and the operation command.
[0062] For example, if the speech turn estimation result is a boolean value, the agent action determination unit 17 determines whether to change the provisionally selected action depending on whether it is true or false. For example, as shown in Figure 10, if the speech turn estimation value is true, that is, if it is the agent's turn to speak, the agent action determination unit 17 changes the action identifier to a standby action (for example, standby action 1) and determines the action command to "shake the body slightly and blink".
[0063] Furthermore, if the speech turn estimation result is a score, the agent action determination unit 17 compares the score, which is the speech turn estimation result, with a predetermined threshold and decides to change the action identifier.
[0064] For example, if the estimated speech turn score is 0.8, as shown in Figure 11, the agent action determination unit 17 changes the action identifier to a standby action and determines the action command "shake the body slightly and blink" which corresponds to a standby action (for example, standby action 1).
[0065] Furthermore, the agent operation determination unit 17 can also determine the operation time according to the output value of the agent utterance turn estimation unit 16. For example, if the utterance turn estimation value is 0.6, the agent operation determination unit 17 determines an operation command to "perform the selected operation in half the standard operation time" while keeping the operation identifier the same (row L12 in Figure 12).
[0066] The agent action determination unit 17 determines the agent's action based at least on the result of the speech turn estimation. At the same time, if the agent action determination unit 17 determines that the agent should not perform a nodding action, it changes the determined action to a waiting action to suppress the nodding action. This prevents unnatural responses such as the agent performing a long nodding action just before the turn to speak, which prevents the user from moving on to the next utterance.
[0067] The association between the speech turn estimation results shown in Figures 10 to 12 and the action identifier, or the action identifier and action time, is pre-configured and can also be changed as appropriate depending on the interaction between the user and the agent.
[0068] [Processing of the Agent Action Control Unit] Next, the processing of the agent action control unit 19 will be described. The agent action control unit 19 takes the action determination result, the speech turn estimation result, and the speech rate calculation result as input and performs the determination of the action time according to the speech rate and the output control of the standby action during the time when there is no determined action.
[0069] [Calculation of Operation Time] The agent operation control unit 19 calculates the operation time using the user's speech rate calculated by the user speech rate calculation unit 15 as input. In other words, the agent operation control unit 19 adjusts the operation time of the nodding action or the waiting action so that the operation time corresponds to the user's speech rate. Figures 13 to 15 are diagrams illustrating the operation time calculation process.
[0070] [Example of calculation of operation time 1] The agent operation control unit 19 takes as input, for example, the speech rate [mora / s] calculated by the user speech rate calculation unit 15 using the method described in Non-Patent Document 2, and calculates the operation time as illustrated in Figure 13.
[0071] Here, "mora" is a unit of syllable. A letter that can be pronounced independently constitutes 1 mora. Letters that cannot be pronounced independently are combined with the preceding letter to form 1 mora. For example, "ka," "kat," and "kaa" all constitute 1 mora.
[0072] When the calculation rules shown in Figure 13 are applied, the agent operation control unit 19 sets the operation time of the agent's response action to the standard operation time × 2.0 if the speech rate v is less than 1.5 [mora / s], in order to match the relatively slow speech rate of the user. Furthermore, if the speech rate v is between 1.5 [mora / s] and 3.5 [mora / s], the agent operation control unit 19 sets the operation time to the standard operation time, and if the speech rate v is greater than 3.5 [mora / s], the operation time is set to the standard operation time × 0.5.
[0073] Thus, according to this calculation example 1, it is possible to handle situations where the user's speaking speed is fast and the user expects quick or short interjections from the agent. Furthermore, if the user's speaking speed is slow, according to this calculation example 1, the agent's interjection speed can be slowed down to prevent it from sounding rushed.
[0074] In this way, the agent operation control unit 19 causes the agent to perform actions with an action time adjusted according to the user's speaking speed, enabling the agent to express more appropriate responses to the user.
[0075] [Example 2 for calculating operation time] Alternatively, the agent operation control unit 19 may calculate the operation time based on the estimated speech turn value (score) output by the agent speech turn estimation unit 16. For example, the output score of the agent speech turn estimation unit 16 is the probability that a turn change occurs and it becomes the agent's speech turn, and is a continuous value between 0.0 and 1.0.
[0076] The agent operation control unit 19 determines the length of the operation time using equation (2), with the probability α of a turn change occurring and it being the agent's turn to speak, and the user's speech rate v [mora / s] as variables.
[0077]
[0078] In equation (2), T0 is the standard operating time for the acknowledgment motion. C1 and C2 are correction coefficients. C1 is a coefficient that determines the length control range of the operating time. C2 is a coefficient that determines the standard for the shortest operating time.
[0079] The coefficient C1 will be described. For example, in Equation (2), the minimum value α of the probability α min , the minimum value v of the speech rate v min are substituted into the longest operation time T max when substituting into Equation (2), and the maximum value α of the probability α max , the maximum value v of the speech rate v max are substituted into the shortest operation time T min when substituting into Equation (2), and the difference between them is taken as the control width T w (constant). In this case, the coefficient C1 is obtained by Equation (3). Here, the subscripts max and min respectively indicate the maximum and minimum values that the variable can take.
[0080]
[0081] Also, C1 may be determined by obtaining C1 such that the difference between two different operation times T obtained by substituting two sets of arbitrary probability α and speech rate v into Equation (2) is equal to a certain control width (constant) of the difference in the operation time T.
[0082] The coefficient C2 will be described. The coefficient C2 determines the criteria such as the shortest operation time. For example, when the shortest operation time T min is desired to be T1 (constant), the coefficient C2 is obtained by Equation (4).
[0083]
[0084] Also, the coefficient C2 may be determined by obtaining the coefficient C2 such that the operation time T becomes equal to a certain constant, with reference to the longest operation time T max or the operation time T when arbitrary probability α and speech rate v are substituted into Equation (2).
[0085] Thus, in equation (2), the timing of the nodding action is adjusted so that the probability of a turn change in speech occurring and the agent taking the turn to speak increases, the shorter the duration of the nodding action becomes. By applying equation (2) to calculate the length of the nodding action, the agent action control unit 19 can anticipate that the turn of speech is likely to change and control the agent's nodding action to be short in duration. Furthermore, even if the user stops speaking expecting the agent to speak, the agent action control unit 19 can immediately nod in agreement, demonstrating the user the agent's active participation in the conversation.
[0086] Furthermore, equation (2) is adjusted so that the faster the user's speaking speed, the shorter the time for the nodding action becomes. The agent action control unit 19 applies equation (2) to calculate the length of the nodding action according to the user's speaking speed, so that even when the user's speaking speed is fast, the agent can nod and then immediately begin speaking after the user finishes speaking without any pause. This allows the control device 10 to demonstrate the agent's active involvement in the conversation.
[0087] Furthermore, the agent operation control unit 19 can adjust the coefficients C1 and C2 in equation (2) for each user according to, for example, the probability α of a turn change occurring and it being the agent's turn to speak, the user's speaking speed v, two different operation times, the longest operation time, the shortest operation time, etc., thereby causing the agent to output an operation in accordance with the user's speaking situation.
[0088] In other words, in calculation example 2, by setting a coefficient C1 that determines the control range of the length of the operation time, and a coefficient C2 that determines the criteria such as the shortest operation time, it is avoided that the agent's nodding time will be shorter than the user's nodding time, or that the agent's nodding will be too fast. As a result, the agent operation control unit 19 can have the agent express a more natural nodding motion, enabling smoother interaction with the user.
[0089] [Example 3 of calculation of operation time] In addition, if the agent operation determination unit 17 selects an operation using an interjection label (see, for example, Figure 9) estimated by, for example, the method described in Reference 5, the agent operation control unit 19 may change the method of calculating the operation time for each interjection label.
[0090] For example, if the response label for the selection action is not NP (Non-positive), the agent action determination unit 17 calculates the action time by applying the action time calculation example 1 (or action time calculation example 2) (Figures 14 and 15). Figures 14 and 15 illustrate the case where, if the response label is NP, the response is negative or indicates hesitation, and the agent's action time is controlled to prevent it from becoming short.
[0091] In contrast, if the response label for the selection action is NP, the agent action determination unit 17 does not apply the action time calculation example 1 or the action time calculation example 2, and sets the standard action time as the action time.
[0092] In this way, the agent operation control unit 19 causes the agent to output an acknowledgment action and a waiting action with an appropriate operation time for each acknowledgment label, enabling the agent to express more flexible responses to the user.
[0093] [Output of Operation Commands] The agent operation control unit 19 updates the operation commands to be executed with an operation time adjusted according to the user's speech rate, and outputs them to the agent operation output unit 20.
[0094] However, if the previous action is still in progress, the decision to output an action command is made according to the type of the previous action. The agent action control unit 19 does not output an action command if the previous action is a nodding action.
[0095] In response to this, the agent operation control unit 19 will interrupt only the nodding action if the previous action is a standby action.
[0096] The agent operation control unit 19 issues a command for a nodding action that overlaps in timing while the standby operation is outputting. As a result, the agent operation control unit 19 interrupts the agent's standby operation and inserts the nodding action.
[0097] In this way, the control device 10 prioritizes the nodding action, allowing the agent to immediately nod in agreement and demonstrating to the user the agent's active participation in the conversation.
[0098] Furthermore, the agent operation control unit 19 outputs a standby operation as an operation command during periods when there are no operations determined by the agent operation determination unit 17.
[0099] [Output Flow] Figure 16 illustrates the output of the determination result by the agent operation determination unit 17 and the operation command by the agent operation control unit 19. Figure 16(a) shows the output flow of the determination result by the agent operation determination unit 17. Figure 16(b) shows the output flow of the operation command by the agent operation control unit 19.
[0100] As shown in Figure 16(a), the agent operation determination unit 17 outputs the following as operation determination results to the agent operation control unit 19: the acknowledgment operation during operation time Ta (time t1 to time t3), the standby operation during operation time Tb (time t4 to time t5), and the acknowledgment operation during operation time Tc (time t6 to time t7).
[0101] In response to this, the agent operation control unit 19 adjusts the operation time Ta of the nodding action to an operation time Ta' (< Ta) according to, for example, the user's speaking speed (Figure 16 (1)). As a result, the agent outputs a nodding action for an operation time Ta' from time t1 to time t2.
[0102] The agent operation control unit 19 does not output an operation command for the standby operation with an operation time Tb because the previous operation was a response operation (Figure 16 (2)). In other words, the agent operation control unit 19 does not interrupt.
[0103] Then, the agent operation control unit 19 outputs a standby operation as an operation command for time t2 and later, if there is no operation determined by the agent operation determination unit 17 (Figure 16 (3)). As a result, the agent outputs a standby operation for the operation time Hb' from time t2 to time t6.
[0104] Then, the agent action control unit 19 interrupts the nodding action with action time Tc because the previous action is a standby action (Figure 16 (4)). At this time, the agent action control unit 19 adjusts the action time Tc of the nodding action to action time Tc' (>Tc), for example, according to the user's speaking speed (Figure 16 (5)). As a result, the agent outputs a nodding action for action time Tc' from time t6 to time t8. Note that the example in Figure 16 shows an example where the standby action is interrupted when issuing a command for a nodding action with overlapping timing. However, if the agent action control unit 19 can process in parallel, such as by moving different body parts for the standby action and the nodding action, the agent action control unit 19 may have the agent output the nodding action in parallel with the standby action. For example, the agent action control unit 19 may have the agent perform a standby action (minor body movement) and a nodding action (nodding, i.e., neck movement) in parallel.
[0105] [Processing Procedure of the Control Method] Next, the processing procedure of the control method according to the embodiment will be described. Figure 17 is a flowchart showing the processing procedure of the control method according to the embodiment.
[0106] In the control device 10, the user voice input unit 11 receives the user's voice input (step S11). The user video input unit 12 receives the input of video footage of the user (step S12).
[0107] The user speech recognition unit 13 performs speech recognition processing to extract text information and speech segment information from the user's voice information using speech recognition technology (step S13).
[0108] The user prosodic information extraction unit 14 extracts prosodic information from the user's voice information (step S14).
[0109] The user speech rate calculation unit 15 calculates the user's speech rate for each speech segment recognized by the user speech recognition unit 13 (step S15).
[0110] The agent utterance turn estimation unit 16 calculates the user's willingness to speak and / or listen, or whether it is time for a turn change, based on the utterance interval information, text information, prosodic information, and video information, and estimates whether it is the agent's turn to speak (step S16).
[0111] The agent action determination unit 17 determines the action identifier and the action start time based on the video information, speech interval information and text information, prosodic information and the estimated agent speech turn value (step S17).
[0112] The agent operation control unit 19 determines the operation time for the operation command corresponding to the operation identifier determined in step S17, according to the standard operation time stored in the agent operation DB 18 and the speech rate value calculated in step S15, and outputs the operation and operation time as an operation command to the agent operation output unit 20 (step S18).
[0113] The agent operation output unit 20 outputs the operation command as an operation representation of the agent (step S19).
[0114] [Agent Operation Determination Process] The processing procedure for the agent operation determination process (step S17) shown in Figure 17 will be explained below. Figure 18 is a flowchart showing an example of the processing procedure for the agent operation determination process shown in Figure 17.
[0115] The agent operation determination unit 17 acquires speech segment information and text information through speech recognition (step S13) (step S21). The agent operation determination unit 17 acquires prosodic information of a certain length prior to the end of the speech segment (step S22). The agent operation determination unit 17 acquires video information of a certain length prior to the end of the speech segment (step S23). Figure 17 shows an example in which time synchronization of text, prosodic information, and video information is performed by acquiring prosodic and / or video of a certain length prior to the end of the speech segment, but the method of time synchronization is not limited to this.
[0116] The agent action determination unit 17 estimates the time when an interjection action exists and makes a provisional selection of an action based on the text information, prosodic information, and video information within the speech interval (step S24). The agent action determination unit 17 performs time estimation and provisional selection of an action, for example, using the method described in Reference 5.
[0117] The agent action determination unit 17 obtains the estimation result (step S25) obtained by estimating the agent's speech turn (step S16).
[0118] As explained using Figures 10 to 12, the agent operation determination unit 17 determines the selected operation of an agent by threshold processing or the like (step S26).
[0119] [Agent Operation Control Processing] The processing procedure for the agent operation control processing (step S18) shown in Figure 17 will be described below. Figure 19 is a flowchart showing an example of the processing procedure for the agent operation control processing shown in Figure 17.
[0120] The agent operation control unit 19 obtains the operation identifier and operation start time determined by the operation determination process (step S17) (step S31).
[0121] The agent operation control unit 19 determines whether the acquired operation identifier is anything other than a standby operation (step S32).
[0122] If the operation identifier is not a standby operation (step S32: No), that is, if it is a standby operation, the agent operation control unit 19 obtains the operation command and standard operation time corresponding to the operation identifier from the agent operation DB 18 (step S33).
[0123] Next, the agent motion control unit 19 calculates the motion time by multiplying the motion standard time by a random number (step S34). In step S34, the agent motion control unit 19 may use the same method for generating involuntary motions as in Section 3.2 (Experimental Setup) of Reference 6. Alternatively, the agent motion control unit 19 may omit step S34. Reference 6: Kurima Sakai et al., "Motion Generation Corresponding to Speaker's Voice and the Additive Effect of Motion on Remotely Operated Robots", Artificial Intelligence Society of Japan Research Meeting Proceedings, SIG-Challenge-B303, pp. 7-13, 2014.
[0124] If the operation identifier is anything other than a standby operation (step S32: Yes), the agent operation control unit 19 obtains the operation command and standard operation time corresponding to the operation identifier from the agent operation DB 18 (step S35).
[0125] Next, the agent operation control unit 19 acquires the speech rate calculated in step S15 (step S36).
[0126] The agent operation control unit 19 calculates the operation time corresponding to the acquired standard operation time and speech rate using, for example, one of the calculation methods described in operation time calculation examples 1 to 3 (step S37).
[0127] After step S34 or step S37 is completed, the agent operation control unit 19 updates the operation command to execute the operation command within the operation time and outputs it to the agent operation output unit 20 (step S38). At this time, if the previous operation is still in progress, the agent operation control unit 19 does not output if the previous operation is a nodding operation, and if the previous operation is a standby operation, it interrupts only if it is a nodding operation.
[0128] Next, the agent operation control unit 19 determines whether or not the next operation (operation identifier and operation start time determined by the operation determination process (step S17)) has been acquired (step S39).
[0129] If the next operation has not been acquired (Step S39: No), the agent operation control unit 19 sets the operation identifier to a standby operation and the operation start time to repeat as the end time of the operation output immediately before (Step S40), and proceeds to Step S33. If the next operation has been acquired (Step S39: Yes), the agent operation control unit 19 terminates the agent operation control processing for this operation.
[0130] [Effects of the Embodiment] The control device 10 according to the embodiment determines the action identifier and start time of the agent's action based on the user's voice information, the user's video information, and the estimation result of speech turn estimation, which estimates whether it is the agent's turn to speak. The control device 10 takes not only voice information but also video information as input and determines the action to be performed by the agent based on the recognition of the user's action response and the estimation result of speech turn estimation, thereby enabling the selection of a more appropriate response action and the determination of the action time.
[0131] The control device 10 then determines the duration of the agent's actions based at least on the user's speaking speed and outputs an action command for the agent. In other words, the control device 10 does not use fixed-length interjections as in the past, but rather anticipates a change in the speaking turn and determines the agent's actions, while also adjusting the action duration according to the user's speaking speed. As a result, the control device 10 can select the appropriate action for the agent and determine the appropriate action duration.
[0132] Furthermore, the control device 10 determines whether the agent should perform an affirmative action or a standby action based on prosodic information derived from the user's voice information, text information extracted from the user's voice information, the user's video information, and the estimation result of speech turn estimation. Based on the estimation result of speech turn estimation, the control device 10 suppresses the affirmative action if the agent should not perform an affirmative action and changes the agent's action to a standby action.
[0133] Therefore, if the control device 10 determines that the agent should not perform a nodding action, it changes the determined action to a waiting action, thereby suppressing the nodding action itself. In this way, the control device 10 can avoid unnatural responses such as the agent performing a long nodding action just before a change of speech, which prevents the user from moving on to the next utterance.
[0134] Furthermore, the control device 10 determines the duration of the agent's nodding action based on the determination result by the agent action determination unit 17, the estimation result of the speech turn estimation, and the user's speech speed. Therefore, the control device 10 causes the agent to perform actions with an appropriate duration according to whether or not a turn change occurs and the user's speech speed. For example, with the control device 10, the agent can nod while still giving a nod, but with a restrained length so that it can start speaking immediately after the other person has finished speaking without any pause. In this way, the control device 10 enables the agent to express more appropriate responses to the user.
[0135] Furthermore, the control device 10 can suppress nodding actions and avoid unnatural responses by the agent by controlling the standby action as the agent's action during the time when the agent action determination unit 17 has not determined an action.
[0136] Furthermore, if the agent's previous action is still in progress and this previous action is a standby action, the control device 10 outputs an action command to interrupt with an acknowledgment action if the action command is an acknowledgment action. This allows the control device 10 to have the agent immediately acknowledging the user, demonstrating the agent's active participation in the conversation to the user. Therefore, the control device 10 enables the agent to express more appropriate responses to the user.
[0137] Thus, the control device 10 allows for the appropriate selection of the agent's operation and operation time, enabling the agent to express more appropriate responses to the user.
[0138] [Regarding the System Configuration of the Embodiment] The control device 10 is a functional concept and does not necessarily need to be physically configured as shown in the figure. In other words, the specific forms of distribution and integration of the functions of the control device 10 are not limited to those shown in the figure, and all or part of it can be configured by functionally or physically distributing or integrating in any unit according to various loads and usage conditions.
[0139] Furthermore, each process performed in the control device 10 may be implemented, in whole or in part, by a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), and a program that is analyzed and executed by the CPU and GPU. Alternatively, each process performed in the control device 10 may be implemented as hardware using wired logic.
[0140] Furthermore, among the processes described in the embodiments, all or part of the processes described as being performed automatically can be performed manually. Alternatively, all or part of the processes described as being performed manually can be performed automatically by known methods. In addition, the processing procedures, control procedures, specific names, and information including various data and parameters described above and illustrated may be changed as appropriate unless otherwise specified.
[0141] [Program] Figure 20 shows an example of a computer in which the control device 10 is realized when a program is executed. The computer 1000 has, for example, memory 1010 and CPU 1020. The computer 1000 also has a hard disk drive interface 1030, a disk drive interface 1040, a serial port interface 1050, a video adapter 1060, and a network interface 1070. These components are connected by a bus 1080.
[0142] Memory 1010 includes ROM 1011 and RAM 1012. ROM 1011 stores, for example, a boot program such as a BIOS (Basic Input Output System). The hard disk drive interface 1030 is connected to the hard disk drive 1090. The disk drive interface 1040 is connected to the disk drive 1100. For example, a removable storage medium such as a magnetic disk or optical disk is inserted into the disk drive 1100. The serial port interface 1050 is connected to, for example, a mouse 1110 and a keyboard 1120. The video adapter 1060 is connected to, for example, a display 1130.
[0143] The hard disk drive 1090 stores, for example, an OS (Operating System) 1091, an application program 1092, a program module 1093, and program data 1094. That is, the program that defines each process of the control device 10 is implemented as a program module 1093 in which code executable by the computer 1000 is written. The program module 1093 is stored, for example, in the hard disk drive 1090. For example, a program module 1093 for performing the same processes as the functional configuration of the control device 10 is stored in the hard disk drive 1090. Note that the hard disk drive 1090 may be replaced by an SSD (Solid State Drive).
[0144] Furthermore, the configuration data used in the processing of the above-described embodiment is stored as program data 1094 in, for example, memory 1010 or hard disk drive 1090. The CPU 1020 then reads the program module 1093 and program data 1094 stored in memory 1010 or hard disk drive 1090 into RAM 1012 as needed and executes them.
[0145] Furthermore, the program module 1093 and program data 1094 are not limited to being stored in the hard disk drive 1090; for example, they may be stored in a removable storage medium and read by the CPU 1020 via a disk drive 1100 or the like. Alternatively, the program module 1093 and program data 1094 may be stored in another computer connected via a network (LAN (Local Area Network), WAN (Wide Area Network), etc.). The program module 1093 and program data 1094 may then be read by the CPU 1020 from the other computer via a network interface 1070.
[0146] Although embodiments applying the invention made by the present inventors have been described above, the present invention is not limited by the descriptions and drawings that constitute part of the disclosure of the present invention in these embodiments. That is, all other embodiments, examples, and operational techniques made by those skilled in the art based on these embodiments are included in the scope of the present invention.
[0147] 10 Control device 11 User voice input unit 12 User video input unit 13 User voice recognition unit 14 User prosodic information extraction unit 15 User speech rate calculation unit 16 Agent speech turn estimation unit 17 Agent operation determination unit 18 Agent operation DB 19 Agent operation control unit 20 Agent operation output unit
Claims
1. A control device comprising: an estimation unit that estimates whether it is the agent's turn to speak based on the user's voice information and the user's video information; a determination unit that determines the operation identifier and start time of the agent's operation based on the user's voice information, the user's video information, and the estimation result of the speech turn estimation; and a control unit that determines the operation time of the operation determined by the determination unit, based on at least the user's speech rate, and outputs an operation command for the agent.
2. The control device according to claim 1, wherein the agent's actions include a nodding action and a waiting action which is an action other than the nodding action, and the determination unit determines the agent's action to be either the nodding action or the waiting action based on prosodic information based on the user's voice information, text information extracted from the user's voice information, the user's video information, and the estimation result of the speech turn estimation.
3. The control device according to claim 2, characterized in that the determination unit suppresses the nodding action and changes the agent's operation to a standby operation when the agent should not perform the nodding action based on the estimation result of the speech turn estimation.
4. The control device according to claim 1, characterized in that the control unit determines the length of the operation time of the agent's operation based on the determination result by the determination unit, the estimation result of the speech turn estimation, and the user's speech rate.
5. The control device according to claim 1, wherein the agent's operation includes a nodding action and a standby action which is an action other than the nodding action, and the control unit outputs the standby action as an operation command during the time when the determination unit has not determined an action.
6. The control device according to claim 1, wherein the agent's operation includes a nodding action and a standby action which is an action other than the nodding action, and the control unit does not output the operation command when the agent is continuing the previous action and the previous action is the nodding action, and outputs the operation command when the previous action is the standby action and the operation command is the nodding action.
7. A control method executed by a control device, comprising: a step of performing speech turn estimation, which estimates whether it is the agent's turn to speak based on the user's voice information and the user's video information; a step of determining the operation identifier and operation start time of the agent's operation based on the user's voice information, the user's video information, and the estimation result of the speech turn estimation; and a step of determining the operation time of the operation determined in the determination step, based at least on the user's speech rate, and outputting an operation command for the agent.
8. A control program for causing a computer to perform the following steps:
8. Performing speech turn estimation, which estimates whether it is the agent's turn to speak based on the user's voice information and the user's video information; 9. Determining the action identifier and start time of the agent's action based on the user's voice information, the user's video information, and the estimation result of the speech turn estimation; and 10. Determining the action duration of the action determined in the determination step, based at least on the user's speech rate, and outputting an action command for the agent.