Control device and control method
Patent Information
- Application Number
- PCT/JP2026/006519
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-03-26
- Filing Date
- 2026-02-24
- Publication Date
- 2026-10-01
Smart Images

Figure JP2026006519_01102026_PF_FP_ABST
Abstract
Description
Control Apparatus and Control Method
[0001] The present disclosure relates to a control apparatus that controls a response generation apparatus, and the like.
[0002] In recent years, various proposals have been made for information processing technology using language models (e.g., Patent Document 1).
[0003] “SensorLLM: Aligning Large Language Models with Motion Sensors for Human Activity Recognition”, https: / / arxiv.org / abs / 2410.10624
[0004] An object of an aspect of the present disclosure is to implement a control apparatus and the like that can improve the convenience of a response generation apparatus that generates a response to a language input.
[0005] A control apparatus according to an aspect of the present disclosure is a control apparatus that controls a response generation apparatus that generates a response to a user's language input, the control apparatus comprising: a motion acquisition unit that acquires motion data relating to the motion or movement of a measurement target from at least one sensor that acquires the motion data; and an apparatus control unit that controls the response generation apparatus based on the motion data acquired by the motion acquisition unit.
[0006] A control method according to an aspect of the present disclosure is a control method executed by a control apparatus that controls a response generation apparatus that generates a response to a user's language input, the control method comprising: a motion acquisition step of acquiring motion data relating to the motion or movement of a measurement target from at least one sensor that acquires the motion data; and an apparatus control step of controlling the response generation apparatus based on the motion data acquired in the motion acquisition step.
[0007] According to an aspect of the present disclosure, the convenience of the response generation apparatus can be improved.
[0008] This is a block diagram showing an example of a response generation device according to Embodiment 1. This is a block diagram showing an example of a response control unit included in the response generation device. This is a schematic diagram showing an example of a prompt generated by the instruction synthesis unit. This is a perspective view showing a specific example of a response generation device. This is a flowchart showing an example of the processing flow by the control unit included in the response generation device. This is a block diagram showing an example of a response generation device according to Embodiment 2. This is a schematic diagram showing another example of a prompt generated by the instruction synthesis unit. This is a block diagram showing an example of a response generation device according to Embodiment 3. This is a schematic diagram showing an example of a prompt template that serves as the basis for a prompt generated by the instruction synthesis unit. This is a block diagram showing an example of a response generation device according to Embodiment 4. This is a block diagram showing an example of a response generation device according to Embodiment 5. This is a block diagram showing an example of a response generation device according to Embodiment 6. This is a block diagram showing an example of a response generation device according to Embodiment 7.
[0009] [Embodiment 1] Hereinafter, one embodiment of the present disclosure will be described in detail.
[0010] <Configuration of the Response Generation Device> Figure 1 is a block diagram showing an example of a response generation device 1 according to Embodiment 1. Figure 2 is a block diagram showing an example of a response control unit 432 included in the response generation device 1.
[0011] The response generation device 1 is a device that generates a response to language input. The response generation device 1 may be a device that realizes an AI (Artificial Intelligence) conversation system. In this embodiment, the response generation device 1 is a device that generates a response to the user's voice input as language input. The response generation device 1 includes, for example, a sensor 2, a microphone 3, a control unit 4, a storage unit 5, an output device 6, and a power supply 7. In this embodiment, the sensor 2 is provided in the response generation device 1.
[0012] Sensor 2 is a motion sensor that measures the movement or motion of the object being measured and acquires the measurement result (motion measurement value) as motion data. The response generation device 1 only needs to be equipped with at least one sensor 2, and the number of sensors 2 is not particularly limited. The movement of the object being measured is, for example, the movement of the user. The motion of the object being measured is, for example, the movement of the response generation device 1, and includes the movement or fall of the response generation device 1 due to an applied external force. Sensor 2 measures the movement or motion of the object being measured in real time. The object being measured is the response generation device 1, or the user using the response generation device 1. Sensor 2 outputs the acquired motion data to the control unit 4.
[0013] Sensor 2 only needs to be able to measure the motion or movement of the object being measured, and may be, for example, a gyro sensor (rotational speed sensor), an acceleration sensor, or a velocity sensor. In this case, sensor 2 acquires motion data such as, for example, the direction of movement of sensor 2, the speed of movement, the acceleration, the direction of rotation, the rotational speed, and the rotational acceleration. Sensor 2 may also be a camera. Sensor 2 may be composed of multiple types of sensors (for example, a gyro sensor and a camera).
[0014] Microphone 3 is an audio input device that receives the user's voice input to the response generation device 1.
[0015] The control unit 4 is, for example, a control device comprising an action acquisition unit 41, an audio input unit 42, a device control unit 43, an audio output unit 44, and a status output unit 45. The control unit 4 only needs to include at least the action acquisition unit 41 and the device control unit 43.
[0016] The motion acquisition unit 41 acquires motion data from the sensor 2. The voice input unit 42 functions as a language input unit that accepts user voice input via the microphone 3. The voice input unit 42 converts the input voice into an audio signal. The device control unit 43 controls the response generation device 1 based on the motion data acquired by the motion acquisition unit 41.
[0017] In this embodiment, the device control unit 43 includes, for example, a motion state estimation unit 431 and a response control unit 432.
[0018] The motion state estimation unit 431 is an example of an estimation unit that estimates the state of the user or response generation device, which is the object of measurement, based on motion data. The motion state estimation unit 431 estimates the user's motion state as the user's state (estimated result).
[0019] The user's motion state may be a qualitatively described text, a quantitatively expressed value, or structured data that combines both. The user's motion state may include, for example, at least one of the user's movement state and rotation state.
[0020] The user's movement state may include, for example, at least one of the user's direction of movement, user's speed of movement, and user's movement stability. This information is provided as an example within the following curly braces {}. Here, we illustrate the case where the user's movement state is represented by qualitatively described text. • User's direction of movement: {stationary, forward, backward, left-hand direction, right-hand direction} • User's speed of movement: {stationary, walking speed, running speed, cycling speed, car speed} • User's movement stability: {steady, intermittently changing, changing}.
[0021] The motion state estimation unit 431 calculates the user's direction of movement by, for example, converting the sensor-referenced movement speed measured by the sensor 2 based on the mechanism of the response generation device 1 and the user's wearing state of the response generation device 1. The motion state estimation unit 431 converts the sensor-referenced coordinate system to a user-referenced coordinate system based on, for example, the position of the sensor 2 in the response generation device 1 and the part of the user (e.g., the neck) on which the response generation device 1 is worn. As a result, the sensor-referenced movement speed can be converted to a user-referenced movement speed, and thus the user's direction of movement can be determined.
[0022] In this embodiment, the motion state estimation unit 431 identifies the calculated user's direction of movement as either stationary, forward, backward, left-hand direction, or right-hand direction. In the user-referenced coordinate system, the ranges for forward, backward, left-hand direction, or right-hand direction are predetermined with respect to the origin (user). Furthermore, if the user's movement speed is 0 or a value close to it, the motion state estimation unit 431 identifies the user's direction of movement as stationary.
[0023] The motion state estimation unit 431 calculates the user's movement speed (user-referenced movement speed) by, for example, the coordinate transformation described above.
[0024] In this embodiment, the motion state estimation unit 431 identifies the calculated user's movement speed as one of the following: stationary, walking speed, running speed, cycling speed, or driving speed. The range of the calculated value (real number) of the user's movement speed and its correspondence with stationary, walking speed, running speed, cycling speed, and driving speed are predetermined.
[0025] The motion state estimation unit 431 determines the user's movement stability based, for example, on changes in the user's direction of movement or the user's movement speed in the most recent period.
[0026] In this embodiment, the motion state estimation unit 431 estimates that the user's movement is in a steady state if the amount of change (absolute value) of the user's direction of movement or movement speed is less than a predetermined value. Furthermore, if the amount of change (absolute value) is greater than or equal to a predetermined value and is undefined, the motion state estimation unit 431 estimates that the user's movement is in a state of intermittent change, and if it follows a certain rule, it determines that the user's movement is in a state of change. The correspondence between the amount of change and the state of change and the user's movement states, such as steady state, intermittent change, and in a state of change, is predetermined.
[0027] Furthermore, the user's rotation state may include, for example, at least one of the user's rotation direction, rotation speed, and rotation stability. This information is provided as illustrative information within the following curly braces {}. Here, we illustrate the case where the user's rotation state is indicated by qualitatively described text. • User's rotation direction: {stationary, right-hand side, left-hand side, forward / backward direction} • User's rotation speed: {stationary, slow, fast} • User's rotation stability: {steady, intermittently changing, changing}.
[0028] The motion state estimation unit 431 calculates the user's rotation direction by, for example, converting the sensor-referenced rotational angular velocity measured by the sensor 2 based on the mechanism of the response generation device 1 and the user's wearing state of the response generation device 1. The user's rotation direction can be calculated based on the relative position between the user and the response generation device 1, and the relative position between the response generation device 1 and the sensor 2, similar to the calculation of the user's direction of movement.
[0029] The motion state estimation unit 431 calculates the user's rotation speed (user-referenced rotation speed) by, for example, the coordinate transformation described above.
[0030] In this embodiment, the motion state estimation unit 431 determines whether the user is stationary, slow, or fast based on the calculated user rotation speed. The range of the calculated user rotation speed (real number) and the correspondence between stationary, slow, and fast are predetermined. Slow speed is a speed at which the user can slowly turn around, and fast speed is a speed greater than that.
[0031] The motion state estimation unit 431 determines the user's rotational stability, for example, based on the most recent change in the user's rotational direction or rotational speed. In this embodiment, similar to the user's movement stability, the motion state estimation unit 431 selects one of the following as the user's rotational stability: steady state, intermittent change, or in a state of change, based on the amount (absolute value) of change in the user's rotational direction or rotational speed and the state of that change.
[0032] By using qualitative data for the user's movement speed and rotation speed, the response control unit 432, described later, can be provided with data that is easy for it to handle, thereby improving the accuracy of the response. Furthermore, by using the user's movement stability and rotation stability, the response control unit 432 can generate a response by considering what the user's motion state is most likely to be in the next moment.
[0033] The device control unit 43 controls the response generation device 1 based on the estimation results of the motion state estimation unit 431. In this embodiment, as an example of controlling the response generation device 1 based on the estimation results of the motion state estimation unit 431, the response control unit 432 controls the generation of a response to the voice input.
[0034] As shown in Figure 2, the response control unit 432 includes, for example, a user instruction generation unit 461, an instruction synthesis unit 462, a context generation unit 463, and a response generation unit 464.
[0035] The user instruction generation unit 461 generates text representing user instructions contained in the voice from the voice signal acquired from the voice input unit 42. The text is composed of natural language. Natural language is a language that humans use on a daily basis and includes various languages such as Japanese, English, French, and German. The user instruction generation unit 461 uses VAD (Voice Activity Detection) to detect the speech segment spoken by the user from the voice signal, and then uses STT (Speech to Text) to convert the voice of the detected speech segment into text.
[0036] The instruction synthesis unit 462 acquires the text generated by the user instruction generation unit 461, the estimation results from the motion state estimation unit 431, and the context information generated by the context generation unit 463 (described later), and synthesizes this data. As a result, the instruction synthesis unit 462 generates a prompt (conversation response request) that can be interpreted by the LLM (Large Language Model) 470 of the response generation unit 464. An example of a prompt will be described later.
[0037] The context generation unit 463 generates contextual information related to the conversation history, which is a history of user instructions generated by the user instruction generation unit 461 and responses generated by the response generation unit 464. The contextual information may be information related to the context of a conversation specific to the application in which the response generation device 1 is used. For example, if the response generation device 1 is used in a route guidance application, the contextual information may include the conversation history, location information of a measurement target or a specific location identified by the interpretation of the conversation history, or the direction in which the user is moving.
[0038] The location information of the object to be measured may be acquired by a receiver (not shown) that receives GPS (Global Positioning System) signals provided in the response generation device 1. The location information of a specific place may be identified using map information from an external device. The direction in which the user is moving may be identified by GPS signals and map information, etc. Furthermore, the analysis of the conversation history may be, for example, an analysis by a small-scale language model that extracts keywords from the user's language input and generated responses (response messages) based on past conversation history and organizes the correlations between keywords.
[0039] The response generation unit 464 uses the LLM 470 to generate a response corresponding to the combination of the user's state estimated by the estimation unit and the voice input. In this embodiment, the response generation unit 464 uses the LLM 470 to generate a response corresponding to the combination of the user's movement state estimated by the movement state estimation unit 431 and the voice input acquired by the voice input unit 42. The response generation unit 464 inputs the prompt generated by the instruction synthesis unit 462 to the LLM 470, causing the LLM 470 to generate a response message.
[0040] The audio output unit 44 converts the response message generated by the response control unit 432 into speech and outputs it to the output device 6. The audio output unit 44 outputs the response message to the output device 6 using TTS (Text to Speech; speech-to-text reading technology).
[0041] The state output unit 45 may function as an output unit that outputs the estimation results of the estimation unit. In this embodiment, the state output unit 45 may output the user's motion state estimated by the motion state estimation unit 431.
[0042] The state output unit 45 outputs the estimation results of the motion state estimation unit 431 from the output device 6 using an output method that corresponds to the difference in the estimation results of the motion state estimation unit 431. The output method according to the estimation result is predetermined. An example of the output method will be described later.
[0043] The storage unit 5 stores information necessary for controlling the response generation device 1. The storage unit 5 stores, for example, conversation history, prompt templates, and the like.
[0044] The output device 6 outputs the audio output from the audio output unit 44 to the outside of the response generation device 1. Accordingly, the response generated by the LLM 470 based on the user's audio input can be presented to the user. The output device 6 may be, for example, an audio output device such as a speaker. In the present embodiment, the response generated by the LLM 470 is output as audio, but it may also be output as a response message (text). In this case, the output device 6 may be, for example, a display device, and the control unit 4 does not need to include the audio output unit 44.
[0045] Furthermore, when outputting the estimation result of the estimation unit, the output device 6 may be, in addition to an audio output device or a display device, a lighting device such as an LED (Light-Emitting Diode), or a vibrator.
[0046] For example, when the output device 6 is a lighting device, it performs lighting according to the estimation result of the estimation unit. For example, when the user's exercise state estimated by the exercise state estimation unit 431 is resting, the state output unit 45 blinks the output device 6. Furthermore, when the exercise state is walking (when the user's movement speed is a walking speed), the state output unit 45 keeps the output device 6 constantly lit.
[0047] For example, when the output device 6 is a speaker, it outputs sound according to the estimation result of the estimation unit. For example, when the user's exercise state estimated by the exercise state estimation unit 431 is resting, the state output unit 45 does not perform output from the output device 6 (sets it to silence). Furthermore, when the exercise state is walking, the state output unit 45 causes the output device 6 to output a sound only once (outputs a beep sound) at the time when the walking state is estimated for the first time (when the start of walking is detected). Furthermore, when the exercise state is running (when the user's movement speed is a running speed), the state output unit 45 causes the output device 6 to output a sound twice (output two beeps) at the time when the running state is estimated for the first time (when the start of running is detected). The state output unit 45 may increase the number of output sounds as the user's movement speed increases.
[0048] Further, when the output device 6 is a speaker, the status output unit 45 may cause the output device 6 to output the estimation result of the estimation unit via voice. When the output device 6 is a display device, the status output unit 45 may cause the output device 6 to output display information (e.g., character information) indicating the estimation result of the estimation unit.
[0049] For example, when the output device 6 is a vibrator, it outputs vibration in accordance with the estimation result of the estimation unit. For example, when the user's exercise state estimated by the exercise state estimation unit 431 is stationary, the state output unit 45 does not perform output from the output device 6 (sets no vibration). Further, when the exercise state is walking, the state output unit 45 causes the output device 6 to output a small vibration at the time point when the start of walking is detected. Furthermore, when the exercise state is running, the state output unit 45 causes the output device 6 to output a large vibration at the time point when the start of running is detected. The state output unit 45 may stepwise increase the output vibration as the user's movement speed increases.
[0050] Accordingly, the user can confirm whether the estimation result of the exercise state estimation unit 431 matches the user's actual state. Through this confirmation, the user can determine whether a member necessary for executing the processing of the exercise state estimation unit 431 is out of order. Further, by confirming the aforementioned matching, the user can determine whether the response generation device 1 can be reliably used.
[0051] It should be noted that notification of the estimation result of the estimation unit is not required. In this case, the control unit 4 does not need to include the state output unit 45, and the output device 6 does not need to include a mechanism (such as an LED and a vibrator) specialized for performing the notification.
[0052] <Example of Prompt> Figure 3 is a schematic diagram showing an example of a prompt generated by the instruction synthesis unit 462. The instruction synthesis unit 462 generates a prompt by setting, in a prompt template stored in the storage unit 5, the text generated by the user instruction generation unit 461, the estimation result of the exercise state estimation unit 431, and the context information generated by the context generation unit 463.
[0053] For example, if a user says, "Which way is the station?", the voice input unit 42 receives voice input indicating the content of this utterance via the microphone 3. When the voice input unit 42 receives voice input, the motion acquisition unit 41 transmits the motion data acquired from the sensor 2 to the motion state estimation unit 431. The voice input unit 42 also converts the received voice input into an audio signal and transmits it to the user instruction generation unit 461.
[0054] The motion state estimation unit 431 estimates the user's motion state based on the acquired motion data and transmits the result to the instruction synthesis unit 462. The instruction synthesis unit 462 embeds the user's motion state acquired from the motion state estimation unit 431 into the "User Motion State" field of the prompt template.
[0055] Furthermore, the user instruction generation unit 461 generates text representing the user's instructions included in the speech and sends it to the instruction synthesis unit 462. The instruction synthesis unit 462 embeds the text generated by the user instruction generation unit 461 into the "user utterance" of the prompt template.
[0056] Furthermore, the context generation unit 463 generates contextual information related to the conversation history when the voice input unit 42 receives voice input. In the example in Figure 3, the context generation unit 463 uses information such as "I want to go to the station" included in the previous conversation history, GPS signals, and map information to generate contextual information including the user's location, the station's location, and the direction the user is moving (the user's direction). If "○○ Station" is specified, the context generation unit 463 identifies the location of that station; otherwise, it identifies the location of the nearest station from the user's location. The instruction synthesis unit 462 embeds the text generated by the context generation unit 463 into the "contextual information" of the prompt template.
[0057] The instruction synthesis unit 462 transmits the generated prompt to the response generation unit 464. The response generation unit 464 inputs this prompt to the LLM 470 and generates a response such as "It's okay to go straight ahead."
[0058] If the prompt does not include information indicating the user's movement state, the LLM 470 cannot determine the user's direction of movement and will generate a response such as "Please go north." In this case, the user may need to separately confirm which direction north is. On the other hand, the response control unit 432 generates a response using the user's movement state, so it can generate a response that is easy for the user to understand intuitively. In this example, the response control unit 432 can generate a response that does not include the direction as described above, taking into account the fact that the user is moving north, since the user's direction of movement is forward and the user's direction is "north."
[0059] Furthermore, the LLM470 can determine if a user is confused by interpreting the user's rotational stability. In this case, the LLM470 is trained to make this determination. Based on this determination, the LLM470 can also generate a reassuring response for the user (for example, a response such as "It's okay").
[0060] <Specific Examples of Response Generation Devices> Figure 4 is a perspective view showing a specific example of a response generation device 1. As shown in Figure 4, the response generation device 1 may be a wearable electronic device that can be worn around the user's neck. The response generation device 1 is not limited to this, and may be implemented, for example, as a portable electronic device that the user can carry with them.
[0061] The response generation device 1 includes a mounting portion 15. The mounting portion 15 is a component for attaching the response generation device 1 to the user's neck. The mounting portion 15 has an open ring shape that can be hooked onto the user's neck. With this configuration, the user's movements are less likely to be hindered even when the response generation device 1 is attached.
[0062] The attachment portion 15 has a microphone 3. Specifically, the microphone 3 is located at one end of the attachment portion 15, which has an open ring shape. This position is near the user's mouth when the user is wearing the response generation device 1 around their neck. Therefore, the user can easily input voice through the microphone 3 while wearing the response generation device 1 around their neck.
[0063] The mounting portion 15 may have a camera as the sensor 2. Specifically, the camera may be positioned near the end of the ring-shaped mounting portion 15 so as to face outwards. This ensures that the camera's orientation substantially coincides with the orientation of the user's face when the user is wearing the response generation device 1 around their neck. Therefore, it becomes possible to capture images of what the user sees with the camera. The response generation device 1 may also be equipped with a gyro sensor, acceleration sensor, or velocity sensor (not shown) as the sensor 2 inside.
[0064] The attachment portion 15 has an output device 6. Specifically, speakers, which serve as output devices 6, are positioned on the left and right sides of the attachment portion 15, which has an open ring shape, when the open ring portion is facing forward. This position is near the user's ears when the user is wearing the response generation device 1 around their neck. The user can easily hear the output from the output device 6 while wearing the response generation device 1 around their neck.
[0065] <Processing Flow> Figure 5 is a flowchart showing an example of the processing flow (an example of a control method) by the control unit 4 of the response generation device 1.
[0066] First, the motion acquisition unit 41 acquires motion data from the sensor 2 (S1; motion acquisition step). In this embodiment, when the voice input unit 42 receives voice input, the motion acquisition unit 41 acquires motion data.
[0067] Next, the device control unit 43 controls the response generation device 1 based on the motion data acquired by the motion acquisition unit 41 (S2; device control step). In this embodiment, as described above, the device control unit 43 estimates the user's motion state based on the motion data acquired by the motion acquisition unit 41 and controls the generation of a response based on the estimation result and voice input.
[0068] Specifically, the motion state estimation unit 431 estimates the user's motion state based on motion data, the user instruction generation unit 461 generates text representing the user's instructions, and the context generation unit 463 generates contextual information related to the conversation history. Subsequently, the instruction synthesis unit 462 generates a prompt based on the user's motion state, the text representing the user's instructions, and the contextual information, and the response generation unit 464 generates a response to this prompt using the LLM 470.
[0069] <Effects> In this way, the device control unit 43 controls the response generation device 1 based on the estimation results of the motion state estimation unit 431, which is based on the motion data acquired by the sensor 2. As a result, the device control unit 43 can perform control in accordance with the movements of the user, who is the object of measurement for the sensor 2.
[0070] In this embodiment, the motion state estimation unit 431 estimates the user's motion state based on motion data. Therefore, the device control unit 43 can control the response generation device 1 based on the user's motion state.
[0071] In this embodiment, the device control unit 43 uses the LLM 470 to generate a response corresponding to a combination of the user's movement state and language input. Therefore, the device control unit 43 can generate a response that is in line with the user's actions. In addition, the device control unit 43 can provide instructions to the LLM 470 that are easy for the LLM 470 to interpret by estimating (classifying) the movement state from the movement data, rather than the movement data itself.
[0072] Conversational systems are used in a variety of user situations, but in conventional conversational systems, the responses were sometimes inappropriate and did not match the situation. For example, Non-Patent Document 1 discloses the use of sensor data in a conversational system, but it does not disclose the control of the response generation device based on user or response generation device operation data. Therefore, for example, a user must first inquire with the LLM about the sensor status, and based on the response, inquire with the LLM again about the content they actually want to inquire about. Consequently, it may take time to receive a response to the content that was originally inquired about, and from the perspective of real-time operation of responses, the convenience of the response generation device 1 may be impaired.
[0073] In this embodiment, as described above, the device control unit 43 generates a response based on the operation data acquired from the sensor 2, so that the user can obtain a response in real time to the content they actually want to inquire about. Furthermore, it can generate a response that is in line with the user's actions, that is, a highly accurate response that takes the user's actions into consideration. Therefore, the convenience of the response generation device 1 can be improved.
[0074] In this way, the device control unit 43 controls the response generation device 1 based on the operation data acquired from the sensor 2, thereby improving the usability of the response generation device 1.
[0075] Furthermore, if the sensor 2 is provided in the response generation device 1, the device control unit 43 can accurately acquire the movements of the user wearing the response generation device 1 or the movement of the response generation device 1.
[0076] <Modification> As described above, the sensor 2 may be composed of multiple types of sensors (for example, a gyro sensor and a camera). In this case, the motion state estimation unit 431 can estimate the user's motion state based on multiple motion data acquired from each sensor. Therefore, the state can be estimated more accurately.
[0077] For example, if the response generation device 1 is equipped with a camera in addition to a gyro sensor, acceleration sensor, or velocity sensor as the sensor 2, the motion state estimation unit 431 can more accurately estimate the user's motion state by analyzing the image captured by the camera. For example, in a situation where the user is riding on a moving object, the motion state estimation unit 431 can determine whether the user is accelerating or the moving object is accelerating, and estimate the user's motion state.
[0078] Furthermore, the response generation device 1 does not necessarily have to include at least one of the microphone 3, storage unit 5, and output device 6, nor does it necessarily have to include at least some of the functions of the control unit 4 described above. Functions that the response generation device 1 does not have can be implemented in another device, including a server or cloud, and can be connected to that other device in a communicative manner. In this case, a control system for controlling the response generation device 1 may be formed, which includes at least an operation acquisition unit 41 and a device control unit 43.
[0079] Furthermore, the control device that controls the response generation device 1 does not necessarily have to have all the functions of the control unit 4. Also, the control device may have some of the components of the response generation device 1, such as the microphone 3.
[0080] Furthermore, the motion state estimation unit 431 may estimate the user's motion state over time and store the estimated user's motion state in the memory unit 5. In this case, the instruction synthesis unit 462 can create prompts using the motion states of multiple users from the most recent period. Therefore, the response generation unit 464 can generate responses that take into account changes in the user's motion state, thus generating responses that are more in line with the user's condition.
[0081] Alternatively, the user may input language via text input. When language input is performed via text input, a text input unit should be provided that accepts text as language input via a text input device such as a touch panel or smartphone. In this case, the response generation device 1 does not need to include a microphone 3, a voice input unit 42, and a user instruction generation unit 461.
[0082] [Embodiment 2] Another embodiment of the present disclosure is described below. For the sake of convenience of explanation, components having the same function as those described in the above embodiments are denoted by the same reference numerals, and their descriptions are not repeated. The same applies to Embodiment 3 and subsequent embodiments.
[0083] <Configuration of the Response Generation Device> Figure 6 is a block diagram showing an example of a response generation device 1A according to Embodiment 2. As shown in Figure 6, the response generation device 1A differs from the response generation device 1 of Embodiment 1 in that it does not have a sensor 2. In this embodiment, the sensor 2 is installed in the user's surrounding environment or is located at a different location from the response generation device 1 and is connected to the control unit 4 in a communicative manner. The sensor 2 and the control unit 4 may be connected, for example, by wireless communication. In other words, in this embodiment, a response generation system including the response generation device 1A and the sensor 2 is constructed.
[0084] An example of the sensor 2 in this embodiment is a directional microphone. The directional microphone may be implemented as a device carried by the user. In this case, the motion acquisition unit 41 acquires ambient sounds around the response generation device 1 acquired by the directional microphone as motion data. The motion state estimation unit 431 detects specific sound patterns or acoustic events (for example, footsteps or car sounds) by analyzing the ambient sounds. Based on these detection results, the motion state estimation unit 431 estimates the user's motion state.
[0085] For example, the detection of a specific sound pattern or acoustic event may be performed by pattern matching. This detection may also be performed using a machine learning model that has been trained on combinations of ambient sounds and specific sound patterns or acoustic events as training data. Furthermore, the estimation of a user's movement state may be performed, for example, by pre-defining the correspondence between a specific sound pattern or acoustic event and the user's movement state (e.g., the user's movement speed and stability). This estimation may also be performed using a machine learning model that has been trained on combinations of specific sound patterns or acoustic events and the user's movement state as training data.
[0086] Furthermore, examples of the sensor 2 in this embodiment include a directional microphone and a speaker. In this case, the directional microphone acquires sound emitted from a speaker installed in the user's surrounding environment. The motion acquisition unit 41 acquires the sound acquired by the directional microphone as motion data. The motion state estimation unit 431 estimates the user's position and movement by estimating the arrival time and direction of the sound from the speaker. Such a method for estimating the arrival time and direction of sound is also called acoustic triangulation. The motion state estimation unit 431 estimates the user's motion state by performing this estimation over time. The speaker may also transmit its position information to the control unit 4.
[0087] Alternatively, a directional speaker may be used as the speaker. In this case, the directional speaker emits sound at a specific frequency, and the directional microphone acquires the reflected sound. The motion acquisition unit 41 acquires the reflected sound acquired by the directional microphone as motion data. The motion state estimation unit 431 estimates the user's position and movement by estimating the distance to the object. This method of estimating the distance to the object is also called echolocation. The motion state estimation unit 431 estimates the user's motion state by performing this estimation over time. In echolocation, the directional speaker may be located near the directional microphone.
[0088] Furthermore, the sensor 2 in this embodiment may be, for example, a camera installed in the user's surrounding environment. The camera tracks the user's movements by analyzing images over time. The camera may also track the user's movements using facial recognition technology or posture estimation technology. In this case, the motion acquisition unit 41 acquires the results of tracking the user's movements as motion data. The motion state estimation unit 431 estimates the user's motion state based on the results of tracking the user's movements. The camera may also transmit its position information to the control unit 4.
[0089] Furthermore, the sensor 2 in this embodiment may include, for example, a beacon installed in the user's surrounding environment. In this case, the motion acquisition unit 41 acquires location information transmitted from the beacon as motion data. The motion state estimation unit 431 estimates the user's motion state (for example, the user's direction of movement (path) and speed of movement) based on the location information acquired over time.
[0090] One of the estimation methods described above may be used, or two or more estimation methods may be used. When two or more estimation methods are used, the motion state estimation unit 431 can estimate the user's motion state, which more accurately reflects the user's actions.
[0091] Thus, even in a configuration where the sensor 2 is not provided in the response generation device 1A, the device control unit 43 can acquire the operation data acquired by the sensor 2. Therefore, similar to the response generation device 1 equipped with the sensor 2, the device control unit 43 can control the response generation device 1A based on the operation data.
[0092] <Example of a Prompt> Figure 7 is a schematic diagram showing another example of a prompt generated by the instruction synthesis unit 462. In this example, the motion state estimation unit 431 estimates the user's motion state using data transmitted from cameras installed in the user's surrounding environment as motion data. The context generation unit 463 uses data transmitted from the cameras to determine the direction in which the user is moving (the user's direction). For example, the context generation unit 463 determines the user's direction based on the image captured by the camera, the camera's location information, and map information. In this embodiment, the context generation unit 463 generates context information that includes the camera's location information instead of the user's location information. Therefore, in this embodiment, the prompt generated by the instruction synthesis unit 462 includes the camera's location information.
[0093] The instruction synthesis unit 462 transmits the generated prompt to the response generation unit 464. The response generation unit 464 inputs this prompt to the LLM 470 and, similar to Embodiment 1, can generate a response that is easy for the user to understand intuitively, such as "Just keep going straight."
[0094] Furthermore, the motion state estimation unit 431 estimates the user's rotational stability from the user's movements (practice swings) included in the image. Therefore, similar to Embodiment 1, the LLM 470 can determine that the user is in a state of confusion by interpreting the user's rotational stability. Based on this determination, the LLM 470 can also generate a response that reassures the user (for example, a response such as "It's okay").
[0095] [Embodiment 3] Figure 8 is a block diagram showing an example of a response generation device 1B according to Embodiment 3. As shown in Figure 8, the response generation device 1B differs from the response generation device 1 of Embodiment 1 in that the device control unit 43B of the control unit 4B includes a conversation state estimation unit 433 instead of a motion state estimation unit 431. The response generation device 1B is equipped with a sensor 2, but the sensor 2 may be connected to the outside of the response generation device 1B in a communication manner. The same applies to the response generation devices of Embodiment 4 and later.
[0096] The conversation state estimation unit 433 is an example of an estimation unit that estimates the state of the user or response generation device being measured based on operation data. The conversation state estimation unit 433 estimates the user's conversation state as the user's state (estimation result). The conversation state estimation unit 433 transmits the estimation result to the instruction synthesis unit 462.
[0097] The user's conversation state may be a qualitatively described text, a quantitatively expressed value, or structured data containing a combination of both. For example, the user's conversation state may include not only the user's state but also the state of the user's surrounding environment during the conversation.
[0098] The conversation state estimation unit 433 includes, for example, a conversation environment estimation unit 481, a user intention estimation unit 482, and a motion state estimation unit 483. The function of the motion state estimation unit 483 is the same as that of the motion state estimation unit 431, so its description is omitted.
[0099] The conversation environment estimation unit 481 estimates what kind of environment the user is in during their conversation. The conversation environment includes, for example, at least one piece of information such as the location where the conversation is taking place (i.e., the location where the response generation device 1B is being used), the presence or absence of a third party, and the time of day when the conversation is taking place.
[0100] Information indicating the location where the conversation is taking place may include, for example, information such as home, inside a car, inside a train, and outdoors. The conversation environment estimation unit 481 may estimate the location where the conversation is taking place by detecting a specific value from the user's motion state (for example, changes in the user's movement speed or acceleration) estimated by the motion state estimation unit 483. In this case, the correspondence between the specific value and each location is predetermined. The motion state estimation unit 483 may also estimate the user's acceleration from the change in the user's movement speed over time, or estimate the change in acceleration from the change in the user's acceleration over time. The estimation of the location where the conversation is taking place may also be performed using a machine learning model that uses combinations of specific values and locations as training data.
[0101] Information indicating the presence or absence of a third party is information indicating whether or not there is someone other than the user. The conversation environment estimation unit 481 may estimate the presence or absence of a third party from the user's motion state (for example, changes in the user's acceleration or rotation state) estimated by the motion state estimation unit 483. The motion state estimation unit 483 may also estimate changes in the user's rotation state from changes in the user's rotation state over time.
[0102] Information indicating the time of day during a conversation includes information indicating time periods such as morning, noon, evening, and night. For example, by storing the correspondence between patterns of action data acquired over time and the time periods in which the action data is acquired in the storage unit 5, the conversation state estimation unit 433 can estimate the time period by pattern matching. This estimation of the time period may also be performed using a machine learning model that uses combinations of action data patterns and time periods as training data.
[0103] The user intention estimation unit 482 estimates the user's intention regarding the response generated by the response control unit 432. Examples of information indicating the user's intention include affirmative, negative, and neutral. By storing the correspondence between action data that represents typical actions indicating the user's intention and the user's intention in the storage unit 5, the user intention estimation unit 482 can estimate the user's intention by pattern matching. Examples of such typical actions include nodding to indicate affirmative and shaking the head to indicate negative. The estimation of the user's intention may also be performed using a machine learning model that uses combinations of user actions and user intentions as training data.
[0104] In this embodiment, the state output unit 45 may output the user's conversation state estimated by the conversation state estimation unit 433 as the estimation result of the estimation unit. The state output unit 45 may output the conversation state, such as the location and time of the conversation, the presence or absence of a third party, and the user's intentions, from the output device 6 in voice. Similar to Embodiment 1, the state output unit 45 may output the conversation state from the output device 6 using an output method associated with the conversation state, not limited to voice. Furthermore, the user's movement state may be output as described in Embodiment 1.
[0105] <Example of a prompt> Figure 9 is a schematic diagram showing an example of a prompt template that serves as the basis for prompts generated by the instruction synthesis unit 462.
[0106] The instruction synthesis unit 462 generates a prompt by setting the text generated by the user instruction generation unit 461, the estimation result of the conversation state estimation unit 433, and the context information generated by the context generation unit 463 to the prompt template stored in the memory unit 5.
[0107] For example, suppose a user is inside a car and is consulting the response generation device 1B about dinner. First, when the user says, "What do you think we should have for dinner?", the voice input unit 42 receives voice input indicating this utterance via the microphone 3. When the voice input unit 42 receives voice input, the motion acquisition unit 41 transmits the motion data acquired from the sensor 2 to the conversation state estimation unit 433. The voice input unit 42 also converts the received voice input into an audio signal and transmits it to the user instruction generation unit 461.
[0108] The motion state estimation unit 483 estimates the user's motion state based on the acquired motion data and transmits the result to the instruction synthesis unit 462. In this example, the conversation environment estimation unit 481 estimates that the location of the conversation is "inside a car" and the time of day is "evening" based on the estimated user's motion state, and transmits the result to the instruction synthesis unit 462. The instruction synthesis unit 462 then embeds the data acquired from the motion state estimation unit 483 and the conversation environment estimation unit 481 into the "user motion state" and "conversation environment" fields of the prompt template, respectively.
[0109] Furthermore, the user instruction generation unit 461 generates text representing the user's instructions included in the speech and sends it to the instruction synthesis unit 462. The instruction synthesis unit 462 embeds the text generated by the user instruction generation unit 461 into the "user utterance" of the prompt template.
[0110] Furthermore, the context generation unit 463 generates contextual information related to the conversation history when the voice input unit 42 receives voice input. In this example, the context generation unit 463 uses conversation history information, GPS signals, and map information to generate contextual information including the user's location, the restaurant's location, and the direction the user is moving (the user's direction). The instruction synthesis unit 462 embeds the text generated by the context generation unit 463 into the "contextual information" (not shown) of the prompt template.
[0111] The instruction synthesis unit 462 transmits the generated prompt to the response generation unit 464. The response generation unit 464 inputs this prompt to the LLM 470 and generates a response such as, "Let's buy dinner at a hamburger shop and go home via the drive-thru." This response is output from the output device 6.
[0112] Suppose the user nods in response to this output. The user intention estimation unit 482 estimates that the user's intention is affirmative based on the motion data acquired at this time. The motion state estimation unit 483 also estimates the user's motion state from motion data other than this nod. For example, the motion state estimation unit 483 estimates that the user's direction of movement is forward and the user's speed of movement is the speed of a car. The conversation environment estimation unit 481 also estimates that the conversation environment (location) is "inside a car" based on the estimation result of the motion state estimation unit 483. The context generation unit 463 generates contextual information related to the conversation history in the same manner as above when the user's intention is estimated.
[0113] The instruction synthesis unit 462 transmits a prompt containing this information embedded in a prompt template to the response generation unit 464 without waiting for the next user utterance (i.e., without waiting for the user instruction generation unit 461 to generate text). The response generation unit 464 inputs this prompt to the LLM 470 and generates a response that takes into account that the user is inside the car, such as "Please continue straight ahead."
[0114] In this way, by estimating the conversation state based on behavioral data, it becomes possible to generate responses by extracting information from the conversational environment that has not been verbalized, or by extracting responses by understanding the user's intentions that have not been verbalized.
[0115] <Processing Flow> In this embodiment, as described above, the device control unit 43 estimates the user's conversation state based on the motion data acquired by the motion acquisition unit 41, and controls the generation of a response based on the estimation result and the voice input. This control by the device control unit 43B is an example of the control of the response generation device 1B based on the motion data acquired by the motion acquisition unit 41 (step S2 in Figure 5).
[0116] Specifically, when the voice input unit 42 receives voice input, the conversation state estimation unit 433 estimates the user's motor state based on motion data, as well as the user's conversation environment. The user instruction generation unit 461 generates text representing the user's instructions, and the context generation unit 463 generates contextual information related to the conversation history. Subsequently, the instruction synthesis unit 462 generates a prompt based on the estimation results of the conversation state estimation unit 433, the text representing the user's instructions, and the contextual information, and the response generation unit 464 generates a response to this prompt using the LLM 470. This allows the response generation unit 464 to generate a response that is intuitive to the user and takes the conversation environment into consideration. The voice output unit 44 outputs the response generated by the response generation unit 464 from the output device 6.
[0117] Subsequently, when the motion acquisition unit 41 receives an action indicating the user's intention, the conversation state estimation unit 433 estimates the user's intention. The conversation state estimation unit 433 also estimates the user's motor state based on motion data other than the action indicating the user's intention, and estimates the user's conversation environment, and the context generation unit 463 generates contextual information related to the conversation history. Then, without waiting for the generation of text representing the user's instruction, the instruction synthesis unit 462 generates a prompt based on the estimation results of the conversation state estimation unit 433 and the contextual information, and the response generation unit 464 generates a response to this prompt using the LLM 470. As a result, the response generation unit 464 can generate a response that takes into account the flow of conversation estimated from the user's actions without waiting for the user to speak.
[0118] <Modification> In this embodiment as well, the sensor 2 may be a gyro sensor, an acceleration sensor, or a velocity sensor, but it may also be a camera. When a camera is used as sensor 2, the conversation state estimation unit 433 can estimate the conversation environment, such as the location where the user is having a conversation, the presence or absence of a third party, and the time of day the conversation is taking place, by analyzing the images captured by the camera as operation data. The conversation state estimation unit 433 can also estimate the user's operation state. These estimations may be performed by pattern matching. In this case, the correspondence between the images captured by the camera and these user states is predetermined. Alternatively, this estimation may be performed using a machine learning model that uses combinations of images captured by the camera and these user states as training data.
[0119] Furthermore, sensor 2 may be a microphone. This microphone may also be microphone 3. In this case, the conversation state estimation unit 433 can estimate the user's state, including whether the user is speaking or not, in addition to the user's state described above, by analyzing the ambient sound acquired by the microphone as operation data. This estimation may also be performed using a model obtained by pattern matching or machine learning, by associating ambient sound with the user's state in advance.
[0120] Furthermore, the conversation state estimation unit 433 may estimate the conversation environment using the conversation history stored in the memory unit 5. This estimation may also be performed using pattern matching or a model obtained by machine learning, by associating the content of the conversation with the conversation environment in advance.
[0121] <Effects> In this way, the device control unit 43B controls the response generation device 1 based on the estimation result of the conversation state estimation unit 433, which is based on the operation data acquired by the sensor 2. As a result, the device control unit 43B can perform control in accordance with the actions of the user, who is the object of measurement by the sensor 2.
[0122] In this embodiment, the conversation state estimation unit 433 estimates the user's conversation state based on the operation data. Therefore, the device control unit 43B can control the response generation device 1B based on the user's conversation state.
[0123] In this embodiment, the device control unit 43B uses the LLM 470 to generate a response corresponding to a combination of the user's conversation state and language input. Therefore, the device control unit 43B can generate a response that is in line with the user's actions. Thus, the device control unit 43B, like the device control unit 43 in Embodiment 1, can improve the convenience of the response generation device 1B. In addition, the device control unit 43B can provide instructions to the LLM 470 that are easy for the LLM 470 to interpret by estimating (classifying) the conversation state from the operation data, rather than the operation data itself.
[0124] [Embodiment 4] Figure 10 is a block diagram showing an example of a response generation device 1C according to Embodiment 4. As shown in Figure 10, the response generation device 1C differs from the response generation device 1 of Embodiment 1 in that the device control unit 43C of the control unit 4C includes an operation control unit 434. In this embodiment as well, the device control unit 43C may estimate the user's movement state or conversation state based on the operation data acquired by the operation acquisition unit 41, and control the generation of a response based on the estimation result and voice input.
[0125] In this embodiment, the motion control unit 434 controls the operation of the components of the response generation device 1C based on the user's state estimated from the motion data. This control by the motion control unit 434 is an example of the control of the response generation device 1C based on the motion data acquired by the motion acquisition unit 41 (step S2 in Figure 5). First, an example will be described in which the motion control unit 434 controls the sound pickup direction of the microphone 3 based on the user's motion state estimated by the motion state estimation unit 431. Note that the user's conversation state estimated by the conversation state estimation unit 433 includes the user's motion state. Therefore, control by the motion control unit 434 can also be performed based on the user's conversation state estimated by the conversation state estimation unit 433.
[0126] In this embodiment, the motion state estimation unit 431 estimates the user's head orientation as the user's motion state based on the motion data. This estimation may be performed using pattern matching or a model obtained by machine learning, after the motion data and head orientation have been associated in advance.
[0127] The motion control unit 434 estimates the position of the sound source (the position of the mouth where speech is uttered) based on the estimated head orientation. The position of the sound source may be specified relative to the estimated head orientation. This estimation may be performed using a model obtained by pattern matching or machine learning, by first establishing a correspondence between the head orientation and the position of the sound source. In this correspondence, the position of the sound source may be specified, for example, in a coordinate system with the position of the sound source when facing forward as the origin. The motion control unit 434 may estimate the actual position of the sound source by converting this coordinate system to a user-referenced coordinate system.
[0128] The motion control unit 434 determines the direction of sound pickup by the microphone 3 based on the estimated position of the sound source. For example, if the response generation device 1C is equipped with a control mechanism (e.g., a small motor) that controls the orientation of the microphone 3, the motion control unit 434 controls the control mechanism so that the microphone 3 faces the direction of the sound source. The motion control unit 434 may generate an instruction that includes the absolute direction in which the microphone 3 should face and the amount of movement of the microphone 3, based on the estimated position of the sound source and the position of the microphone 3, and transmit it to the control mechanism.
[0129] For example, when a user is looking upwards and speaking, sensor 2 measures the user's movements, and motion state estimation unit 431 estimates that the user's head is facing upwards based on these measurements. Motion control unit 434 estimates the location of the sound source based on this estimation and directs microphone 3 upwards. This shifts the direction of sound pickup by microphone 3 upwards.
[0130] In this way, by setting the sound pickup direction of the microphone 3 to the location of the sound source in conjunction with the acquisition of motion data by the sensor 2, the microphone 3 can more accurately pick up the user's voice. As a result, the response control unit 432 can generate text representing the user's instructions more accurately, thereby improving the quality of the response. Therefore, the convenience of the response generation device 1C can be improved.
[0131] <Modification 1: Sound collection using multiple microphones 3> The response generation device 1C may be equipped with multiple microphones 3. In this case, the operation control unit 434 may increase the receiving sensitivity of at least one of the multiple microphones 3 based on the estimated location of the sound source. Based on the estimated location of the sound source and the location of each microphone 3, the receiving sensitivity of the microphones 3 may be controlled so that the closer the microphone 3 is to the sound source, the higher its receiving sensitivity. Alternatively, beamforming using multiple microphones 3 may be performed. That is, the phase of each audio signal acquired by each microphone 3 may be adjusted to increase the directivity with respect to the estimated location of the sound source. In these cases as well, similar to the case where the orientation of the microphones 3 is controlled, the microphones 3 can more accurately capture the user's voice. Note that when capturing the user's voice using multiple microphones 3 in this way, the response generation device 1C does not need to be equipped with a control mechanism to control the orientation of the microphones 3.
[0132] <Modification 2: Use of data other than motion data> The motion control unit 434 may control the direction of sound pickup by the microphone 3 based on the context information generated by the context generation unit 463. In this case, the motion state estimation unit 431 may use the context information to estimate the orientation of the head. Instead of the context information, the conversation history stored in the memory unit 5 may be used. This estimation may be performed by associating the content of the conversation with the orientation of the head in advance, using pattern matching or a model obtained by machine learning.
[0133] Furthermore, the motion control unit 434 may control the direction of sound pickup by the microphone 3 based on the text representing the user's instruction generated by the user instruction generation unit 461. In this case, the motion control unit 434 may estimate the actual sound source location using a model obtained by pattern matching or machine learning, by associating the content of the generated text with the location of the sound source.
[0134] <Other Control Examples> The motion control unit 434 may perform controls other than sound pickup control by the microphone 3. One example is described below. The motion control unit 434 may perform one of the various controls described in this specification, or it may perform two or more controls.
[0135] (Control of output device) The motion control unit 434 may control the output direction of the speaker, which is the output device 6, based on the user's motion state estimated by the motion state estimation unit 431. In this case, the motion control unit 434 estimates the position of the user's ears based on the orientation of the head estimated by the motion state estimation unit 431, instead of the position of the sound source. The motion control unit 434 determines the output direction of the speaker based on the estimated position of the ears.
[0136] Similar to microphone 3, if the response generation device 1C is equipped with a control mechanism for controlling the direction of the speaker, the operation control unit 434 controls the control mechanism so that the speaker faces the ear. Furthermore, if the response generation device 1C is equipped with multiple speakers, the output sensitivity of at least one of the multiple speakers may be increased based on the estimated ear position. Alternatively, sound may be output towards the estimated ear position by beamforming using multiple speakers.
[0137] In this way, by aligning the speaker's output direction with the ear position in conjunction with the acquisition of motion data by sensor 2, the speaker can deliver clearer sound to the user. Therefore, the convenience of the response generation device 1C can be improved.
[0138] (Camera control) The motion control unit 434 may control the imaging direction of the camera based on the user's motion state estimated by the motion state estimation unit 431. In this case, the motion control unit 434 determines the imaging direction of the camera based on the head orientation estimated by the motion state estimation unit 431, instead of the sound source position. For example, the memory unit 5 may store information relating the head orientation to the imaging direction.
[0139] The response generation device 1C may include, for example, a control mechanism for controlling the orientation of the camera. The motion control unit 434 may generate an instruction including the camera's orientation and movement amount so that the camera faces the determined imaging direction, and transmit it to the control mechanism.
[0140] In this way, by controlling the imaging direction of the camera in conjunction with the acquisition of motion data by sensor 2, it becomes possible to orient the imaging direction to match the direction the head is facing. Therefore, even if the direction the body is facing and the direction the head is facing are different, the camera can capture the scenery in the direction the head is facing. As a result, when the response control unit 432 generates text representing user instructions, for example, by taking the camera image into consideration, it becomes possible to generate the text more accurately, thereby improving the quality of the response. Consequently, the convenience of the response generation device 1C can be improved.
[0141] [Embodiment 5] Figure 11 is a block diagram showing an example of a response generation device 1D according to Embodiment 5. As shown in Figure 11, the response generation device 1D differs from the response generation device 1C of Embodiment 4 in that the device control unit 43D of the control unit 4D includes a usage state estimation unit 435 instead of a motion state estimation unit 431. In this embodiment as well, the device control unit 43D may estimate the user's motion state or conversation state based on the motion data acquired by the motion acquisition unit 41, and control the generation of a response based on the estimation result and voice input. The device control unit 43D may also perform the control performed by the motion control unit 434 of Embodiment 4.
[0142] The usage status estimation unit 435 is an example of an estimation unit that estimates the status of the user or the response generation device based on operation data. The usage status estimation unit 435 estimates the usage status of the response generation device 1D as the status of the response generation device 1D.
[0143] The usage status of the response generation device 1D may be a qualitatively described text, a quantitatively expressed value, or structured data containing a combination of these. The usage status of the response generation device 1D may indicate, for example, whether the response generation device 1D is in use or whether the user is speaking. Alternatively, the usage status of the response generation device 1D may indicate whether the response generation device 1D is being used normally, i.e., whether it is in an abnormal state.
[0144] The usage status estimation unit 435 may, for example, estimate that the response generation device 1D is in use if it determines, based on the operation data acquired over time, that the amount of change in the operation data (for example, the amount of change in the acceleration or rotational acceleration of sensor 2) is greater than or equal to a predetermined value. On the other hand, if the usage status estimation unit 435 determines that the amount of change is less than a predetermined value (i.e., if there is no change in the acceleration of sensor 2, etc., for a certain period of time), it may estimate that the response generation device 1D is not in use.
[0145] The usage status estimation unit 435 may, for example, identify patterns in the motion data acquired over time by analyzing the motion data, and estimate whether the user is speaking based on these patterns. This estimation may be performed using pattern matching by pre-storing typical motion data patterns during speech, or it may be performed using a machine learning model that has been trained on combinations of motion data patterns and speech presence / absence as training data.
[0146] The usage state estimation unit 435 may, for example, estimate whether the response generation device 1D is in an abnormal state, such as falling, moving rapidly, or undergoing irregular changes, based on changes in operation data acquired over time. This estimation may be performed by pattern matching, by associating patterns of operation data with these abnormal states in advance, or by using a model obtained through machine learning.
[0147] The motion control unit 434 controls the response generation device 1D based on the state of the response generation device 1D estimated by the usage state estimation unit 435. That is, the device control unit 43D estimates the state of the response generation device 1D based on the motion data acquired by the motion acquisition unit 41 and controls the response generation device 1D based on the estimation result. This control is an example of controlling the response generation device 1D based on the motion data acquired by the motion acquisition unit 41 (step S2 in Figure 5). This makes it possible for the device control unit 43D to perform control in accordance with the movement of the response generation device 1D that is being measured.
[0148] For example, the operation control unit 434 may control the power supply 7 of the response generation device 1D based on the estimation result of the usage state estimation unit 435.
[0149] The operation control unit 434 may, for example, turn off the power supply 7 if it is estimated to be in use, and turn on the power supply 7 if it is estimated to be in use from an unused state. Even when the power supply 7 is off, it is sufficient that the functions used in the processing of estimating the usage status of the response generation device 1D remain operational. In addition, the operation control unit 434 may, for example, put the power supply 7 into sleep mode if it is estimated to be in use, and then release the sleep mode of the power supply 7 if it is subsequently determined to be in use. That is, the operation control unit 434 may put the power supply 7 into sleep mode by turning off the power supply of only some of the components (for example, the microphone 3 and the output device 6) that the response generation device 1D has. Furthermore, the operation control unit 434 may turn off the power supply 7 of the response generation device 1D if it is estimated that the response generation device 1D is in an abnormal state.
[0150] Furthermore, the operation control unit 434 may control whether or not to output the response generated by the response control unit 432, or whether or not to execute the response processing by the response control unit 432, based on the estimation result of the usage state estimation unit 435.
[0151] For example, the operation control unit 434 may not output the response generated by the response control unit 432 if it is estimated that the user is speaking, but may output the response generated by the response control unit 432 if it is estimated that the user is not speaking. Alternatively, the operation control unit 434 may execute the response processing by the response control unit 432 if it is estimated that the user is speaking, but may not execute the response processing by the response control unit 432 if it is estimated that the user is not speaking. In this way, the operation control unit 434 can control the input / output timing of data in the response control unit 432 according to the user's state. Furthermore, the operation control unit 434 may control the power supply of the output device 6, for example, depending on whether output is enabled or disabled.
[0152] Furthermore, the operation control unit 434 may control the status output unit 45 to output the estimation result of the usage status estimation unit 435 from the output device 6. For example, the status output unit 45 may output from the output device 6 by voice or display that the response generation device 1 is in an abnormal state. Alternatively, the status output unit 45 may output from the output device 6 that the response generation device 1D is in an abnormal state using an output method associated with the abnormal state (for example, flashing or vibration).
[0153] For example, the status output unit 45 may also output the usage status of the response generation device 1D estimated by the usage status estimation unit 435, the user's speech status, or that the response generation device 1D is functioning normally. These outputs may also be performed by voice or display, or by an output method corresponding to each status.
[0154] Furthermore, the estimation result of the usage state estimation unit 435 may be used for purposes other than control by the operation control unit 434, for example, in generating a response by the response control unit 432. In this case, the response control unit 432 may generate a response corresponding to the combination of the estimation result of the usage state estimation unit 435 and the audio signal from the audio input unit 42.
[0155] <Effects> In this way, the device control unit 43D controls the response generation device 1D based on the estimation result of the usage state estimation unit 435, which is based on the operation data acquired by the sensor 2. As a result, the device control unit 43D can perform control in accordance with the movement of the response generation device 1D, which is the object of measurement by the sensor 2.
[0156] In this embodiment, the usage state estimation unit 435 estimates the usage state of the response generation device 1D based on the operation data. Therefore, the device control unit 43D can control the response generation device 1D based on its usage state. Specifically, as described above, the device control unit 43D can control the power supply 7 of the response generation device 1D, control the input / output timing of data in the response control unit 432, and notify the status of the response generation device 1D based on the operation data. Furthermore, the device control unit 43D can generate a response that is in line with the operation of the response generation device 1D or the user by generating a response corresponding to the combination of the usage state of the response generation device 1D and language input using the LLM 470.
[0157] Therefore, the device control unit 43D can improve the usability of the response generation device 1D by controlling the response generation device 1D based on the operation data acquired from the sensor 2.
[0158] Furthermore, since the status output unit 45 outputs the estimation result of the usage status estimation unit 435, the user can confirm whether the estimation result matches the actual state of the response generation device 1D. Therefore, the user can determine whether any components necessary for the usage status estimation unit 435 to perform its processing are malfunctioning. In addition, by confirming the above-mentioned match, the user can determine whether the response generation device 1D can be used with confidence.
[0159] [Embodiment 6] Figure 12 is a block diagram showing an example of a response generation device 1E according to Embodiment 6. As shown in Figure 12, the response generation device 1E differs from the response generation device 1 of Embodiment 1 in that the control unit 4E comprises a device control unit 43E which includes an motion acquisition unit 41E and a motion state estimation unit 431E.
[0160] As shown in Figure 12, in this embodiment, a response generation system is constructed that includes a response generation device 1E and an external sensor 11. The external sensor 11 is connected to the response generation device 1E in a communicative manner and acquires operational data of the object to be measured. The external sensor 11 and the response generation device 1E may be connected, for example, by wireless communication.
[0161] Sensor 2 may be referred to as the first sensor, and the motion data acquired from sensor 2 may be referred to as the first motion data. Similarly, external sensor 11 may be referred to as the second sensor, and the motion data acquired from external sensor 11 may be referred to as the second motion data. The first and second sensors may be located on different parts of the user's body. The first sensor may be located inside the response generation device, and the second sensor may be located outside the response generation device. The first and second sensors may be of the same type or different types. In other words, the first motion data and the second motion data may be of the same type or different types.
[0162] Specifically, the external sensor 11 may have the same functions as sensor 2. The external sensor 11 may be, for example, a gyro sensor, an acceleration sensor, or a velocity sensor, or it may be a camera, or it may be composed of multiple types of sensors. The external sensor 11 may be the same as sensor 2 provided by the response generation device 1E, or it may be a different sensor. In addition, the external sensor 11 acquires at least one of the following as operation data: the direction of movement of the external sensor 11, the speed of movement, the acceleration, the direction of rotation, the speed of rotation, and the acceleration of rotation.
[0163] The motion acquisition unit 41E acquires motion data from the sensor 2 (step S1 in Figure 5). At this time, the motion acquisition unit 41E also acquires motion data from the external sensor 11. Then, the device control unit 43E controls the response generation device 1E based on the motion data from the sensor 2 and the motion data from the external sensor 11 acquired by the motion acquisition unit 41E (step S2 in Figure 5).
[0164] In this embodiment, the motion state estimation unit 431E estimates the user's motion state (for example, the user's movement state, rotation state, and head orientation) based on motion data from sensor 2 and motion data from external sensor 11. Therefore, the motion state estimation unit 431E can estimate the user's motion state with greater accuracy.
[0165] The external sensor 11 is held in a location different from the user's body part (e.g., the neck) to which the response generation device 1E is attached. Therefore, the motion acquisition unit 41E can acquire motion data from the external sensor 11 in a location different from the sensor 2. Therefore, the motion state estimation unit 431E can estimate the user's motion state based on the motion data from each part of the user's body.
[0166] In this embodiment, an example in which an external sensor 11 is connected to the response generation device 1E is described, but an external sensor 11 different from sensor 2 may be connected to the response generation devices 1A to 1D of embodiments 2 to 5. For example, not only the motion state estimation unit 431 of embodiment 1, but also the conversation state estimation unit 433 of embodiment 3 may estimate the user's conversation state based on the motion data from sensor 2 and the motion data from the external sensor 11. Also, for example, the usage state estimation unit 435 of embodiment 5 may estimate the usage state of the response generation device 1D based on the motion data from sensor 2 and the motion data from the external sensor 11.
[0167] <Specific examples of external sensor locations> The external sensor 11 may be provided, for example, in earphones, earrings, piercings, or glasses. In this case, the motion acquisition unit 41E can acquire the user's head movement state as motion data. Therefore, the motion state estimation unit 431E can estimate the user's head rotation state (for example, the direction of head rotation, the speed of head rotation, or the stability of head rotation) as the user's rotation state.
[0168] Therefore, for example, the motion state estimation unit 431 (motion state estimation unit 431E) can estimate the orientation of the head more accurately. As a result, for example, the conversation state estimation unit 433 of Embodiment 3 can estimate the user's intentions with greater accuracy. Also, for example, the motion control unit 434 of Embodiment 4 can control the direction of the microphone 3, camera, or output device 6 with greater accuracy. Also, for example, the usage state estimation unit 435 of Embodiment 5 can estimate the user's speech state with greater accuracy.
[0169] Furthermore, the external sensor 11 may be provided in, for example, a ring, watch, bracelet, or smartwatch. In this case, the motion acquisition unit 41E can acquire the user's arm movement state as motion data. Therefore, the motion state estimation unit 431E can estimate the user's arm movement state, such as the magnitude and speed of arm swing, as the user's motion state.
[0170] Furthermore, the external sensor 11 may be provided, for example, in a shoe or anklet. In this case, the motion acquisition unit 41E can acquire the user's leg motion state as motion data. Therefore, the motion state estimation unit 431E can estimate the user's leg motion state as the user's motion state. Also, the external sensor 11 may be provided, for example, in a smartphone.
[0171] [Embodiment 7] Figure 13 is a block diagram showing an example of a response generation device 1F according to Embodiment 7. As shown in Figure 13, the response generation device 1F differs from the response generation device 1 of Embodiment 1 in that the control unit 4F includes a device control unit 43F which includes a response control unit 432F.
[0172] In this embodiment, the device control unit 43F estimates the user's motion state based on the motion data acquired by the motion acquisition unit 41, and based on the estimation result, selects an LLM 470 to be used for generating the response from among a plurality of LLMs 470. The device control unit 43F generates a response based on the estimation result and the voice input using the selected LLM 470. This control by the device control unit 43F is an example of the control of the response generation device 1F based on the motion data acquired by the motion acquisition unit 41 (step S2 in Figure 5).
[0173] The response control unit 432F includes, for example, a user instruction generation unit 461, an instruction synthesis unit 462, a context generation unit 463, an LLM determination unit 466, and a response generation unit 464F.
[0174] The LLM determination unit 466 determines which LLM 470 to be used by the response generation unit 464F from among a plurality of LLM 470 based on the estimation result of the motion state estimation unit 431. Specifically, the LLM determination unit 466 determines which LLM 470 to be used by the response generation unit 464F based on the content of the prompt generated by the instruction synthesis unit 462. In other words, the LLM determination unit 466 determines which LLM 470 to be used by the response generation unit 464F based on the estimation result of the motion state estimation unit 431, user instructions, and conversation history.
[0175] The LLM determination unit 466 selects at least one LLM 470 suitable for generating a response to the prompt based on the content of the prompt, and transmits the prompt to the selected LLM 470. The LLM determination unit 466 may select one LLM or multiple LLMs as the LLM 470 to be used by the response generation unit 464F.
[0176] The LLM determination unit 466 analyzes the prompt and determines which field (task) the content contained in the prompt relates to. The LLM determination unit 466 may analyze the context contained in the prompt, for example, by using a small-scale language model used in the analysis of the conversation history described above. Based on the analysis results, the LLM determination unit 466 determines the field of the content contained in the prompt and selects an LLM 470 to send the prompt by referring to the field information stored in the storage unit 5. The field information is information about the field in which each LLM excels at generating responses.
[0177] The field information may include keywords related to each field. The LLM determination unit 466 may determine which field the content of the prompt relates to by comparing the language included in the prompt with the keywords included in the field information.
[0178] The response generation unit 464F generates a response to a prompt using at least one of a plurality of LLMs 470 based on the determination result of the LLM determination unit 466. In this embodiment, the response generation unit 464F has a first LLM 471 and a second LLM 472 as LLMs 470. The first LLM 471 and the second LLM 472 may be LLMs that differ in their ability to generate a response to a prompt.
[0179] The first LLM 471 is an LLM suitable for generating responses based on a specific motion state, for example. The second LLM 472 is an LLM suitable for generating responses based on other fields unrelated to the specific motion state. The second LLM 472 generates responses in fields where the first LLM 471 cannot or has difficulty generating appropriate responses.
[0180] The first LLM 471 and the second LLM 472 are not limited to those described above, and may be an LLM suitable for generating a response based on a specific field and an LLM suitable for generating a response based on a different field. Furthermore, the LLM 470 is not limited to two, but may comprise three or more LLMs, in which case each LLM may be suitable for generating a response based on a different field from the others.
[0181] In this embodiment, an example is described in which the response generation device 1 of Embodiment 1 is equipped with an LLM determination unit 466. However, the embodiment is not limited to this, and the response generation devices 1A to 1E of Embodiments 2 to 6 may also be equipped with an LLM determination unit 466.
[0182] <Specific Examples of Response Generation> We will explain using the example of the first LLM 471 being an LLM suitable for providing responses related to driving navigation. That is, the first LLM 471 is an LLM suitable for providing responses in a specific motion state where the user is in a car (an LLM suitable for providing responses in the field of route guidance while driving a car). On the other hand, the second LLM 472 is an LLM suitable for providing responses in fields other than driving navigation.
[0183] Suppose a user, while in a car, says, "I want to go to Tokyo, how do I get there?" The voice input unit 42 receives this voice input, including the content of the utterance. The motion acquisition unit 41 acquires motion data from the sensor 2.
[0184] The motion state estimation unit 431 estimates the user's motion state based on the motion data acquired by the motion acquisition unit 41. In this example, the motion state estimation unit 431 estimates the user's motion state, specifically the user's movement speed, to be the speed of a car. In the response control unit 432F, the user instruction generation unit 461 generates text representing the user's instruction based on the voice input, and the context generation unit 463 generates contextual information related to the conversation history. Subsequently, the instruction synthesis unit 462 generates a prompt based on the user's motion state, the text representing the user's instruction, and the contextual information.
[0185] The LLM determination unit 466 determines, based on the content of the prompt, that the current conversation is related to "driving navigation". Therefore, the LLM determination unit 466 selects the first LLM 471, which is suitable for responses related to driving navigation, as the LLM 470 to be used by the response generation unit 464F, and sends the prompt to the first LLM 471. The response generation unit 464F uses the first LLM 471 to generate a response to the prompt, for example, "Turn right at the next traffic light".
[0186] In this way, the LLM determination unit 466 selects an appropriate LLM 470 for the user's movement state, so the response generation unit 464F can generate a response that takes the user's movement state into account. Therefore, the accuracy of the response by the response generation unit 464F can be improved.
[0187] In the case of a response generation device that lacks a motion state estimation unit 431 and does not estimate the user's motion state, for example, when a user requests route guidance while riding in a car, the LLM determination unit 466 may not be able to select a first LLM 471 suitable for a response related to driving navigation. In this case, the response generation unit 464F may generate a response using another LLM 470 (for example, a second LLM 472). As a result, the other LLM 470 may determine that the user is walking and generate a response that guides the user to the nearest station and transfers, or a response that guides the user to a road that is not passable by car.
[0188] On the other hand, in this embodiment, in response to the above requirements, the LLM determination unit 466 selects a first LLM 471 suitable for responses related to the driving navigation system as an appropriate LLM 470 for the user's motion state. Therefore, the response generation unit 464F can generate a response that takes into account that the user is in a car. Consequently, the response generation unit 464F can reduce the possibility of providing the user with incorrect information regarding the user's motion state.
[0189] [Example of implementation by software] The functions of the response generation devices 1, 1A to 1F (hereinafter referred to as "devices") can be realized by programs that cause a computer to function as the device, and by programs that cause a computer to function as each control block of the device (especially each part included in the control units 4, 4B to 4F).
[0190] In this case, the device includes a computer having at least one control device (e.g., a processor) and at least one storage device (e.g., memory) as hardware for executing the program. By executing the program using this control device and storage device, the functions described in each of the embodiments are realized.
[0191] The above program may be recorded on one or more computer-readable recording media, not temporary ones. These recording media may or may not be provided by the above device. In the latter case, the program may be supplied to the above device via any wired or wireless transmission medium.
[0192] Furthermore, some or all of the functions of each of the above control blocks can also be implemented by logic circuits. For example, an integrated circuit in which logic circuits functioning as each of the above control blocks are formed is also included in the scope of this disclosure. In addition, it is also possible to implement the functions of each of the above control blocks by, for example, a quantum computer.
[0193] Furthermore, each process described in the above embodiments may be performed by AI (Artificial Intelligence). In this case, the AI may operate on the control device described above, or it may operate on another device (for example, an edge computer or a cloud server).
[0194] [Summary] The control device according to Embodiment 1 of the present disclosure is a control device for controlling a response generation device that generates a response to a user's language input, comprising: an action acquisition unit that acquires action data relating to the action or movement of a measurement object from a sensor that acquires the action data said to be said to be said to be said to be said to be a device control unit that controls the response generation device based on the action data acquired by the action acquisition unit.
[0195] In the control device according to Embodiment 2 of the present disclosure, in Embodiment 1, the measurement target of the sensor is the user or the response generation device, the device control unit includes an estimation unit that estimates the state of the user or the response generation device based on the operation data, and controls the response generation device based on the estimation result of the estimation unit.
[0196] In the control device according to embodiment 3 of the present disclosure, in embodiment 2, the estimation unit estimates the user's motion state as the user's state.
[0197] In the control device according to embodiment 4 of the present disclosure, in embodiment 2 or 3, the estimation unit estimates the user's conversation state as the user's state.
[0198] In the control device according to embodiment 5 of the present disclosure, in any of embodiments 2 to 4, the estimation unit estimates the usage state of the response generation device as the state of the response generation device.
[0199] In any of embodiments 2 to 5, the control device according to embodiment 6 of the present disclosure includes a response generation unit that uses a language model to generate a response corresponding to a combination of the user's state estimated by the estimation unit and the language input.
[0200] In the control device according to embodiment 7 of the present disclosure, in embodiment 6, the device control unit determines, based on the estimation result of the estimation unit, the language model to be used by the response generation unit from among a plurality of language models.
[0201] In any of embodiments 2 to 7, the control device according to embodiment 8 of the present disclosure comprises a microphone, and the device control unit controls the direction of sound pickup by the microphone based on the user's state estimated by the estimation unit.
[0202] In any of embodiments 2 to 8, the control device according to embodiment 9 of the present disclosure includes a speaker in the response generation device, and the device control unit controls the output direction of the speaker based on the user state estimated by the estimation unit.
[0203] In any of embodiments 2 to 9, the control device according to embodiment 10 of the present disclosure includes a camera in the response generation device, and the device control unit controls the imaging direction of the camera based on the user's state estimated by the estimation unit.
[0204] In the control device according to embodiment 11 of the present disclosure, in any of embodiments 2 to 10, the device control unit controls the response generation device based on the state of the response generation device estimated by the estimation unit.
[0205] In any of embodiments 2 to 11, the control device according to embodiment 12 of the present disclosure includes an output unit that outputs the estimation result of the estimation unit.
[0206] In the control device according to embodiment 13 of the present disclosure, in any of embodiments 1 to 12, the sensor is provided in the response generation device.
[0207] In any of embodiments 1 to 12, the control device according to embodiment 14 of the present disclosure is configured such that the sensor is installed in the user's surrounding environment or is located at a different location from the response generation device and is communicatively connected to the control device.
[0208] In any of embodiments 1 to 14, the control device according to embodiment 15 of the present disclosure is configured such that the operation acquisition unit acquires first operation data as operation data from a first sensor as the sensor provided in the response generation device, and acquires second operation data as operation data from a second sensor as the sensor which is communicably connected to the response generation device.
[0209] A control method according to aspect 16 of the present disclosure is a control method performed by a control device that controls a response generation device that generates a response to a user's language input, and includes an action acquisition step of acquiring action data from a sensor that acquires action data relating to the action or movement of a measurement object, and a device control step of controlling the response generation device based on the action data acquired in the action acquisition step.
[0210] Each aspect of the present disclosure may be implemented by a computer, in which case the control program for the control device that enables the computer to implement the control device by operating the computer as each part (software element) of the control device, and the computer-readable recording medium on which the program is recorded, also fall within the scope of the present disclosure.
[0211] [Additional Notes] This disclosure is not limited to the embodiments described above, and various modifications are possible within the scope of the claims. Embodiments obtained by appropriately combining the technical means disclosed in different embodiments are also included in the technical scope of this disclosure. Furthermore, new technical features can be formed by combining the technical means disclosed in each embodiment.
[0212] This application claims priority to Japanese Patent Application No. 2025-052255, filed on 26 March 2025, and all of its contents are included herein by reference.
[0213] 1, 1A-1F Response generation device 2 Sensor 3 Microphone 6 Output device (speaker) 11 External sensor 4, 4B-4F Control unit (control device) 41, 41E Motion acquisition unit 43, 43B-43F Device control unit 431, 431E Motion state estimation unit (estimation unit) 433 Conversation state estimation unit (estimation unit) 435 Usage state estimation unit (estimation unit) 45 State output unit (output unit) 464, 464F Response generation unit 470 LLM (Language Model) 471 First LLM (Language Model) 472 Second LLM (Language Model)
Claims
1. A control device for controlling a response generation device that generates a response to a user's language input, comprising: an action acquisition unit that acquires action data relating to the action or movement of a measurement target from at least one sensor; and a device control unit that controls the response generation device based on the action data acquired by the action acquisition unit.
2. The control device according to claim 1, wherein the sensor measures either the user or the response generating device, the device control unit comprises an estimation unit that estimates the state of the user or the response generating device based on the operation data, and controls the response generating device based on the estimation result of the estimation unit.
3. The control device according to claim 2, wherein the estimation unit estimates the user's motion state as the user's state.
4. The control device according to claim 2 or 3, wherein the estimation unit estimates the user's conversation state as the user's state.
5. The control device according to any one of claims 2 to 4, wherein the estimation unit estimates the usage state of the response generation device as the state of the response generation device.
6. The control device according to any one of claims 2 to 5, wherein the device control unit comprises a response generation unit that generates a response corresponding to a combination of the user state estimated by the estimation unit and the language input using a language model.
7. The control device according to claim 6, wherein the device control unit determines, based on the estimation result of the estimation unit, the language model to be used by the response generation unit from among a plurality of language models.
8. The control device according to any one of claims 2 to 7, wherein the response generation device comprises a microphone, and the device control unit controls the direction of sound pickup by the microphone based on the user's state estimated by the estimation unit.
9. The control device according to any one of claims 2 to 8, wherein the response generation device comprises a speaker, and the device control unit controls the output direction of the speaker based on the user state estimated by the estimation unit.
10. The control device according to any one of claims 2 to 9, wherein the response generation device comprises a camera, and the device control unit controls the imaging direction by the camera based on the user's state estimated by the estimation unit.
11. The control device according to any one of claims 2 to 10, wherein the device control unit controls the response generation device based on the state of the response generation device estimated by the estimation unit.
12. The control device according to any one of claims 2 to 11, wherein the device control unit comprises an output unit that outputs the estimation result of the estimation unit.
13. The control device according to any one of claims 1 to 12, wherein the sensor is provided in the response generation device.
14. The control device according to any one of claims 1 to 12, wherein the sensor is installed in the user's surrounding environment or is located at a different location from the response generation device and is communicably connected to the control device.
15. The control device according to any one of claims 1 to 14, wherein the motion acquisition unit acquires first motion data as motion data from a first sensor as motion data provided in the response generation device, and acquires second motion data as motion data from a second sensor as motion data connected to the response generation device in a communicative manner.
16. A control method performed by a control device that controls a response generation device that generates a response to a user's language input, comprising: an action acquisition step of acquiring action data relating to the action or movement of a measurement object from at least one sensor that acquires the action data; and a device control step of controlling the response generation device based on the action data acquired in the action acquisition step.