Dialogue device, terminal device, dialogue method, and dialogue system

JP2026131332APending Publication Date: 2026-08-14HONDA MOTOR CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-02-03
Publication Date
2026-08-14

AI Technical Summary

Benefits of technology

【0006】 本発明の一態様によれば、ユーザから入力されたプロンプトに対して、内容が関連した複数の文を出力できる。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026131332000001_ABST
    Figure 2026131332000001_ABST
Patent Text Reader

Abstract

The system will be able to output multiple related sentences in response to a prompt (PR) entered by the user (P). [Solution] The dialogue device (21) that interacts with the user by voice or text comprises: a decision unit (203) that determines the policy for the content of the response to a prompt input by the user; a first generation unit (204) that generates a first sentence (SE1) based on the policy determined by the decision unit; a second generation unit (205) that generates a second sentence (SE2) based on the policy determined by the decision unit; and an output unit (206) that outputs the first sentence generated by the first generation unit and the second sentence generated by the second generation unit by voice or text.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an interactive device, a terminal device, an interaction method, and an interaction system.

Background Art

[0002] Conventionally, technologies for interacting with users are known. For example, Patent Document 1 discloses a vehicle voice interaction device that performs an interlinking process to connect "intervals" when the time until an answer to a question from a driver is output is longer than an allowable waiting time.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] However, the voice output by the interlinking process of Patent Document 1 is a voice that does not indicate a direct answer to the question from the driver. Therefore, in Patent Document 1, for the question from the driver, a voice for connecting "intervals" and a voice for a question whose content is not related to this voice are output, and the user may feel uncomfortable in the interaction. The present invention has been made in view of the above circumstances, and an object thereof is to be able to output a plurality of sentences with related content for a prompt input from a user.

Means for Solving the Problems

[0005] One aspect of the present invention is a dialogue device that interacts with a user by voice or text, comprising: a decision unit that determines a policy for the content of the response to a prompt input by the user; a first generation unit that generates a first sentence based on the policy determined by the decision unit; a second generation unit that generates a second sentence based on the policy determined by the decision unit; and an output unit that outputs the first sentence generated by the first generation unit and the second sentence generated by the second generation unit by voice or text. [Effects of the Invention]

[0006] According to one aspect of the present invention, multiple sentences related in content can be output in response to a prompt entered by a user. [Brief explanation of the drawing]

[0007] [Figure 1] Figure 1 shows the configuration of the dialogue system. [Figure 2] Figure 2 shows the configuration of the dialogue device. [Figure 3] Figure 3 shows the functional components of the processor. [Figure 4] Figure 4 is a flowchart showing the operation of the dialogue device. [Figure 5] Figure 5 shows an example of interaction between a dialogue device and a crew member. [Figure 6] Figure 6 shows an example of interaction between a dialogue device and a crew member. [Figure 7] Figure 7 shows an example of interaction between a dialogue device and a crew member. [Modes for carrying out the invention]

[0008] [1. Configuration of the Dialogue System] Figure 1 shows the configuration of the dialogue system 1000. The dialogue system 1000 is a system that communicates with the occupant P of vehicle 1 using voice. Crew member P is an example of a "user".

[0009] The dialogue system 1000 comprises a dialogue device 21 installed in the vehicle 1 and a traffic information server 2 connected to a network NW, which is a WAN (Wide Area Network). The traffic information server 2 is a server device that provides traffic information indicating the traffic conditions in the area including the current location of the vehicle 1. In Figure 1, an example is shown in which a terminal device 3 used by the occupant P is installed inside the vehicle 1. This terminal device 3 has sound output means such as a speaker and display means such as a display.

[0010] Figure 1 shows the configuration of Vehicle 1. Vehicle 1 illustrated in Figure 1 is a four-wheeled vehicle. Vehicle 1 is equipped with a driver's seat 10A, a passenger seat 10B, a rear right seat 10C, and a rear left seat 10D. In Vehicle 1 in Figure 1, the driver is seated in the driver's seat 10A, and the driver is also a passenger P.

[0011] Vehicle 1 is equipped with seating sensors 11A, 11B, 11C, and 11D that detect the weight applied to the seats. Seating sensor 11A is provided in the driver's seat 10A, seating sensor 11B is provided in the passenger seat 10B, seating sensor 11C is provided in the rear right seat 10C, and seating sensor 11D is provided in the rear left seat 10D. In the following explanation, when seat sensors 11A, 11B, 11C, and 11D are not distinguished, they will be referred to as "seat sensor 11" with the designation "11".

[0012] Vehicle 1 is equipped with a dashboard 12. The dashboard 12 is fitted with a touch panel 13, a speaker 14, and a microphone 15. The touch panel 13 consists of a display panel that displays characters and images and a touch sensor that detects contact with the display panel, which are superimposed or integrated. The speaker 14 outputs sound into the passenger compartment of vehicle 1. The microphone 15 collects sound from inside vehicle 1. The installation positions and number of the speaker 14 and microphone 15 can be changed as desired.

[0013] Vehicle 1 is equipped with a shift lever 16. The shift lever 16 is located near the driver's seat 10A.

[0014] Vehicle 1 includes an illuminance sensor 17. The illuminance sensor 17 is installed near the windshield 18 and detects the illuminance outside Vehicle 1.

[0015] Vehicle 1 includes a front camera 19A, a rear camera 19B, a right side camera 19C, and a left side camera 19D. The front camera 19A is a camera that captures the front of Vehicle 1. The rear camera 19B is a camera that captures the rear of Vehicle 1. The right side camera 19C is a camera that captures the right side of Vehicle 1. The left side camera 19D is a camera that captures the left side of Vehicle 1. Hereinafter, when not distinguishing between the front camera 19A, the rear camera 19B, the right side camera 19C, and the left side camera 19D, they are denoted by the symbol "19" and expressed as "camera 19".

[0016] Vehicle 1 includes a TCU (Telematics Control Unit) 20. The TCU 20 is a communication device that communicates with devices connected to the network NW.

[0017] Vehicle 1 includes an interaction device 21. The interaction device 21 is a device that interacts with the occupant P.

[0018] [2. Configuration of the Interaction Device] FIG. 2 is a diagram showing the configuration of the interaction device 21. Connected to the interaction device 21 are a seating sensor 11, a touch panel 13, a speaker 14, a microphone 15, an illuminance sensor 17, a camera 19, a TCU 20, a shift position sensor 22, a vehicle speed sensor 23, and a GNSS (Global Navigation Satellite System) unit 24. Note that the devices connected to the interaction device 21 are not limited to these devices, and other types of devices may also be connected.

[0019] The seating sensor 11 detects the weight at a predetermined cycle and outputs a detection value indicating the detected weight to the interaction device 21. The touch panel 13 displays various information according to the control of the interaction device 21. The speaker 14 outputs various sounds according to the control of the dialogue device 21. The microphone 15 collects sound according to the control of the dialogue device 21. The illuminance sensor 17 detects illuminance at predetermined intervals and outputs a detected value indicating the detected illuminance to the dialogue device 21. The camera 19 takes pictures according to the control of the dialogue device 21 and outputs the captured data obtained from the pictures to the dialogue device 21. The TCU20 communicates with devices connected to the network NW according to the control of the dialogue device 21.

[0020] The shift position sensor 22 detects the shift position of the shift lever 16 of the vehicle 1. The shift positions include, for example, P (parking) used when parked, R (reverse) and N (neutral) used when reversing, and D (drive) used when driving. The shift position sensor 22 outputs a detected value indicating the detected shift position to the dialogue device 21.

[0021] The vehicle speed sensor 23 detects the speed of vehicle 1. The vehicle speed sensor 23 detects the speed of vehicle 1 at predetermined intervals and outputs a detected value indicating the detected speed of vehicle 1 to the dialogue device 21 each time it detects the speed.

[0022] The GNSS unit 24 determines the current position of vehicle 1. The GNSS unit 24 generates position data indicating the current position of vehicle 1 and outputs the generated position data to the dialogue device 21.

[0023] The interactive device 21 includes a processor 200 such as a CPU (Central Processing Unit) or MPU (Micro Processor Unit), a memory 220, and interface circuits for connecting various devices. Memory 220 is an example of a "storage unit".

[0024] Memory 220 is a storage device that stores programs and data. Memory 220 stores the control program 221, the policy database 222, the inter-statement database 223, and data to be processed by the processor 200. Memory 220 has a non-volatile storage area. Alternatively, memory 220 may also have a volatile storage area and constitute the work area of ​​the processor 200. Memory 220 is composed of, for example, ROM (Read Only Memory) or RAM (Random Access Memory).

[0025] The control program 221 is a program that, when read and executed by the processor 200, causes the processor 200 to function as the functional unit shown in Figure 3.

[0026] Here, referring to Figure 3, the functional parts of the processor 200 will be explained through the explanation of policy DB222 and interposition statement DB223. Figure 3 shows the functional components of the processor 200.

[0027] As shown in Figure 3, the processor 200 functions as a speech recognition unit 201, a determination unit 202, a decision unit 203, a first generation unit 204, a second generation unit 205, and an output unit 206 by reading and executing the control program 221.

[0028] [2-1. Speech Recognition Unit] The speech recognition unit 201 performs speech recognition on prompt PR (see, for example, Figure 5) which includes the content of the occupant P's speech, based on the sound collected by the microphone 15. Here, prompt PR is input information that includes instructions and questions for generating dialogue content. The speech recognition unit 201 converts the speech prompt PR into text using its speech recognition function and outputs the converted text to the decision unit 203. The speech recognition unit 201 converts the speech prompt PR into text using an existing speech recognition function that references an acoustic model and a language model.

[0029] The speech recognition unit 201 determines whether crew member P has uttered a specific phrase indicating the start of a voice command (a so-called wake-up word or trigger word). If the wake-up word is uttered, the unit recognizes the subsequent utterance as prompt PR. Alternatively, instead of using a wake-up word, the unit may use voice activity detection to detect a voice segment of the recorded sound and recognize the utterance within that segment as prompt PR.

[0030] The speech recognition unit 201 outputs the converted text to the determination unit 203 and the second generation unit 205.

[0031] [2-2. Judgment section] The determination unit 202 determines the level of cognitive difficulty. The level of cognitive difficulty indicates the degree of difficulty in recognizing the content of the response statement SE2 (hereinafter referred to as "response statement SE2" as appropriate) to the content of the prompt PR. In this embodiment, the determination unit 202 determines the level of cognitive difficulty based on the traffic conditions at the current location of the vehicle 1, the brightness around the vehicle 1, the difficulty level of the work performed by the occupant P on the vehicle 1, and the presence or absence of passengers. The reply is an example of a "second sentence."

[0032] The worse the traffic conditions, the less likely occupant P is to be able to focus on anything other than the traffic conditions. In other words, the better the traffic conditions, the easier it is for occupant P to focus on things other than the traffic conditions. Therefore, the worse the traffic conditions, the more difficult it is expected that occupant P will be to recognize the content of response statement SE2. Accordingly, the judgment unit 202 will determine a higher level of cognitive difficulty the worse the traffic conditions are.

[0033] Furthermore, the darker the surroundings of vehicle 1, the poorer the visibility of the outside of vehicle 1, making it difficult for occupant P to focus on matters other than those related to the movement of vehicle 1. In other words, the brighter the surroundings of vehicle 1, the better the visibility of the outside of vehicle 1, making it easier for occupant P to focus on matters other than those related to the movement of vehicle 1. Therefore, the darker the surroundings of vehicle 1, the more difficult it is expected that occupant P will be to recognize the content of response statement SE2. Accordingly, the judgment unit 202 determines a higher level of cognitive difficulty the darker the surroundings of vehicle 1.

[0034] Furthermore, if there is a passenger, it is considered highly likely that a conversation will take place with the passenger, and therefore, it is considered that the passenger P will have little opportunity to focus on the response sentence SE2. For this reason, the judgment unit 202 determines that the level of cognitive difficulty is higher when there is a passenger than when there is no passenger.

[0035] Furthermore, when vehicle 1 is backing up or turning right at an intersection, the difficulty of the occupant P performing tasks on vehicle 1 is considered to be higher compared to when vehicle 1 is stopped or not backing up. Therefore, the determination unit 202 determines the level of cognitive difficulty based on the difficulty of the occupant P performing tasks on vehicle 1.

[0036] The determination unit 202 sends request information to the traffic information server 2 via the TCU 20. This request information includes the latest location data received from the GNSS unit 24. After sending the request information, the determination unit 202 receives traffic information from the traffic information server 2. Next, the determination unit 202 uses the congestion status and restriction status included in the received traffic information as parameters to obtain a numerical value indicating the degree of traffic conditions. This predetermined algorithm outputs a value indicating worse traffic conditions the more congestion, congestion length, and restrictions there are.

[0037] Furthermore, the determination unit 202 acquires the detection value from the illuminance sensor 17 as the degree of brightness around the vehicle 1. Alternatively, the determination unit 202 may acquire the degree of brightness around the vehicle 1 numerically from the captured image shown by the camera 19's captured data.

[0038] Furthermore, the determination unit 202 determines whether or not a passenger is present based on the detection value of the seat sensor 11. The determination unit 202 determines that a passenger is present if at least one of the seat sensors 11B, 11C, or 11D outputs a detection value corresponding to seating. On the other hand, the determination unit 202 determines that there is no passenger present if none of the seat sensors 11B, 11C, or 11D output a detection value corresponding to seating.

[0039] Furthermore, the determination unit 202 identifies the type of work performed by occupant P on vehicle 1 based on detection values ​​from the vehicle speed sensor, gyro sensor, and shift position sensor of vehicle 1. Next, the determination unit 202 obtains the difficulty level of the work corresponding to the identified work from a database or the like.

[0040] The determination unit 202 obtains an overall evaluation value using a predetermined algorithm based on the degree of traffic conditions, the degree of brightness around vehicle 1, the difficulty of the work performed by occupant P on vehicle 1, and the presence or absence of passengers. This predetermined algorithm outputs a higher evaluation value the worse the traffic conditions, the darker the surroundings of vehicle 1, the higher the difficulty of the work, and the higher the evaluation value when there are passengers than when there are no passengers.

[0041] When the determination unit 202 obtains an overall evaluation value, it compares the obtained evaluation value with a threshold value to determine the cognitive difficulty level corresponding to the obtained evaluation value from among the multi-level cognitive difficulty levels provided. The determination unit 202 determines that the higher the overall evaluation value, the higher the cognitive difficulty level.

[0042] In this embodiment, the cognitive difficulty level is divided into three stages: "low," "medium," and "high," with "low" being the lowest level of cognitive difficulty, "medium" being the second highest, and "high" being the highest. The determination unit 202 determines that the cognitive difficulty level is "low" if the overall evaluation value is below the first threshold. The determination unit 202 also determines that the cognitive difficulty level is "medium" if the overall evaluation value is greater than the first threshold but below the second threshold. The determination unit 202 also determines that the cognitive difficulty level is "high" if the overall evaluation value is greater than the second threshold.

[0043] When the determination unit 202 determines the level of cognitive difficulty, it outputs the determination result to the decision unit 203.

[0044] [2-3. Decision Section] The decision unit 203 determines the policy for the response to the prompt PR input by crew member P (hereinafter referred to as "policy" as appropriate). The decision unit 203 refers to the policy DB 222 and determines the policy based on the text input from the speech recognition unit 201.

[0045] Policy DB222 is a database that stores policies as information. Policies indicate what kind of response should be given to a prompt PR. Policy DB222 stores multiple policies as information, such as "negative," "affirmative," and "greeting." The "negative" policy indicates that the response to the prompt PR will be negative. Furthermore, the "affirmative" policy indicates that the response to a prompt PR should be positive in relation to the content of the prompt PR. Furthermore, the "greeting" policy indicates that the response to a prompt PR should be a greeting.

[0046] Policy DB222 associates each policy with at least one or more words and at least one or more groups of words as information. For example, the policy of "negation" is associated with groups of words such as "legal speed limit," "exceeding," and "no problem," as well as groups of words such as "stop sign," "intention to obey," and "not," and groups of words such as "ramen," "salt," and "more." Furthermore, for example, the policy of "affirmation" is associated with groups of words such as "legal speed limit" and "do not exceed," "stop sign" and "intend to obey" and "have," and "ramen" and "soy sauce" and "number one." Furthermore, for example, the "greetings" policy is associated with words such as "good morning," "hello," "good evening," and groups of words like "hi" and "long time no see."

[0047] The decision unit 203 extracts one or more words from the text input from the speech recognition unit 201. The decision unit 203 extracts one or more words from the text by, for example, referring to a dictionary in which multiple words are recorded. The dictionary referred to in this example is stored in memory 220.

[0048] The decision unit 203 extracts words from the text input from the speech recognition unit 201, then refers to the policy database 222 to identify the policy corresponding to one or more of the extracted words. For example, if the word extracted from the text is "Good morning," the decision unit 203 identifies "Greetings" as the policy. For example, if the words extracted from the text are "legal speed limit," "exceeding," and "no problem," the decision unit 203 will identify "negative" as the policy. For example, if the words extracted from the text are "ramen," "soy sauce," and "number one," the decision unit 203 identifies "affirmative" as the policy.

[0049] The determination unit 203 further determines the type of response sentence SE2 to be either a long sentence or a short sentence. A long sentence is a sentence longer than a short sentence and of a predetermined length or longer. In other words, a long sentence is a sentence with a predetermined number of characters or more. A short sentence is a sentence shorter than a long sentence and of a predetermined length or shorter. In other words, a short sentence is a sentence with a predetermined number of characters or less. For example, the predetermined number is 20, but it may be 19 or less or 21 or more. A long sentence is an example of a "sentence of first length." A short sentence is a "sentence of second length."

[0050] The decision unit 203 determines whether the response sentence SE2 is a long sentence or a short sentence based on the cognitive difficulty level determination result input from the judgment unit 202. For example, if the cognitive difficulty level determination result indicates "high", the decision unit 203 determines the response sentence SE2 to be a long sentence. Alternatively, if the cognitive difficulty level determination result indicates "low" or "medium", the decision unit 203 determines the response sentence SE2 to be a short sentence.

[0051] Furthermore, if the decision unit 203 extracts words related to driving from the text input from the speech recognition unit 201, it determines the type of response SE2 to be a long sentence. On the other hand, if the decision unit 203 extracts words related to driving from the text input from the speech recognition unit 201, it determines the type of response SE2 to be a short sentence. Here, we will explain the words related to driving. Examples of words related to driving include legal speed limits, stop signs, traffic lights, highways, oncoming vehicles, and pedestrians, and these are pre-set.

[0052] When the decision unit 203 determines a policy, it outputs the determined policy to the first generation unit 204 and the second generation unit 205. Also, when the decision unit 203 determines the type of response statement SE2, it outputs the determined type of response statement SE2 to the first generation unit 204 and the second generation unit 205.

[0053] The decision unit 203 may also determine the policy and the type of response statement SE2 as follows. In this case, policy DB222 contains instructions such as "Always deny any claims that violate the law," "Do not deny personal preferences, but respond in lengthy answers," and "Respond in lengthy answers to matters related to driving." The decision unit 203 inputs the text received from the speech recognition unit 201 and the information described in the policy DB 222 into a language model such as an LLM (Large Language Model), and determines the policy and the type of response sentence SE2 by obtaining the policy and the type of response sentence SE2 from the language model. For example, if the word extracted from the text input by the speech recognition unit 201 is "Good morning", the decision unit 203 inputs the word "Good morning" and the information described in the policy DB 222 into the language model. If the language model outputs "Long sentence, affirmative", the decision unit 203 determines "Greeting" as the policy and "Long sentence" as the type of response sentence SE2. Also, if the language model outputs "Greeting, Long sentence", the decision unit 203 determines "Greeting" as the policy and "Long sentence" as the type of response sentence SE2.

[0054] [2-4. 1st generation part] The first generation unit 204 determines whether or not to generate a filler or interjection sentence SE1. The first generation unit 204 determines not to generate an interjection sentence SE1 if the type of response sentence SE2 is a short sentence, and determines to generate an interjection sentence SE1 if the type of response sentence SE2 is a long sentence. The interlude sentence SE1 is an example of a "first sentence".

[0055] If the first generation unit 204 determines that it should generate a transition statement SE1, it generates the transition statement SE1 based on the policy input from the decision unit 203. The first generation unit 204 generates the transition statement SE1 by extracting it from the transition statement DB 223.

[0056] The interposition statement DB223 is a database that stores interposition statements SE1 as information. In the interposition statement DB223, one or more interposition statements SE1 are associated with each policy. For example, in the interlude sentence DB223, the policy "negation" is associated with interlude sentences SE1 such as "but," "no," and "not that." Furthermore, for example, in the interlude sentence DB223, the policy of "affirmative" is associated with interlude sentences SE1 such as "certainly," "I see," and "that's right." Furthermore, for example, in the transitional phrase DB223, transitional phrases SE1 such as "Hey," "Hi," and "Long time no see" are associated with the policy "Greeting."

[0057] The first generation unit 204 generates a transition statement SE1 by obtaining a transition statement SE1 corresponding to the policy determined by the decision unit 203 from the transition statement DB 223. For example, if the policy decided by the decision unit 203 is "negative", the first generation unit 204 obtains the interlude statement SE1 for "but" from the interlude statement DB 223. Furthermore, for example, if the policy decided by the decision unit 203 is "affirmative", the first generation unit 204 obtains the interlude statement SE1 meaning "That's right" from the interlude statement DB 223. Furthermore, for example, if the policy decided by the decision unit 203 is "greeting," the first generation unit 204 obtains the interlude phrase SE1 for "hi" from the interlude phrase DB 223.

[0058] When the first generation unit 204 generates a transitional statement SE1, it outputs the generated transitional statement SE1 to the second generation unit 205 and the output unit 206.

[0059] The first generation unit 204 may also generate the interlude sentence SE1 as follows: The first generation unit 204 may input the text obtained from the interlude sentence DB 223 into a language model such as an LLM, and generate the interlude sentence SE1 by obtaining an expression that the language model deems appropriate from the language model.

[0060] [2-5.Second generation part] When the first generation unit 204 outputs a filler sentence SE1, the second generation unit 205 generates a response sentence SE2 based on the text input from the speech recognition unit 201, the policy and type of response sentence SE2 input from the decision unit 203, and the filler sentence SE1 input from the first generation unit 204. On the other hand, if the first generation unit 204 does not output a filler sentence SE1, the second generation unit 205 generates a response sentence SE2 based on the text input from the speech recognition unit 201 and the policy and type of response sentence SE2 input from the decision unit 203.

[0061] For example, the second generation unit 205 generates the response sentence SE2 using a language model such as LLM. In this example, a memory device accessible to the processor 200 (such as memory 220 or a database connected to the network NW) stores a machine learning model that outputs the response sentence SE2 in response to inputs such as the prompt PR entered by crew member P, dialogue conditions such as the topic, the policy, the type of response sentence SE2, and the interlude sentence SE1. This model outputs a negative response SE2 to the prompt PR entered by crew member P when the policy "negative" is input, a positive response SE2 to the prompt PR entered by crew member P when the policy "positive" is input, and a greeting response SE2 to the prompt PR entered by crew member P when the policy "greet" is input. Furthermore, if a transitional sentence SE1 is input, this model outputs a response sentence SE2 that does not include the input transitional sentence SE1 at the beginning. Furthermore, this model outputs a response SE2 corresponding to the type of response SE2 input. In other words, if the input response SE2 is a long sentence, this model outputs a response SE2 consisting of a predetermined number of characters or more, and if the input response SE2 is a short sentence, it outputs a response SE2 consisting of fewer than the predetermined number of characters.

[0062] Furthermore, the generation of response sentence SE2 is not limited to generation by a language model such as LLM. For example, response sentence SE2 may be generated based on a set of boilerplate text or a rule base. When generation is performed based on a set of boilerplate text or a rule base, the second generation unit 205 generates a negative response sentence SE2 if the policy is "negative", generates a positive response sentence SE2 if the policy is "affirmative", and generates a response sentence SE2 that includes a greeting if the policy is "greeting". In this case, the second generation unit 205 generates a response sentence SE2 that does not include the interlude sentence SE1 at the beginning. In this case, if the type of response sentence SE2 is a long sentence, the second generation unit 205 generates a response sentence SE2 consisting of a predetermined number of characters or more, and if the type of response sentence SE2 is a short sentence, it generates a response sentence SE2 consisting of fewer than the predetermined number of characters.

[0063] When the second generation unit 205 generates the response sentence SE2, it outputs the generated response sentence SE2 to the output unit 206.

[0064] [2-6. Output Section] If the output unit 206 does not receive the interlude sentence SE1 from the first generation unit 204, but receives the reply sentence SE2 from the second generation unit 205, it outputs the reply sentence SE2 via the speaker 14.

[0065] Furthermore, if the output unit 206 receives a transition sentence SE1 from the first generation unit 204 and a response sentence SE2 from the second generation unit 205, it outputs the transition sentence SE1 and the response sentence SE2 through the speaker 14. In this case, the output unit 206 outputs the transition sentence SE1 before the response sentence SE2. After outputting the transition sentence SE1, the output unit 206 outputs the response sentence SE2 when a predetermined trigger occurs. Examples of predetermined triggers include the completion of outputting the transition sentence SE1, the elapsed of a predetermined time (e.g., 2 seconds) since outputting the transition sentence SE1, the vehicle 1 stopping after the output of the transition sentence SE1, and the cognitive difficulty level determined by the judgment unit 202 falling below a predetermined level ("low" or "medium"). The output unit 206 determines that the vehicle 1 has stopped based on the values ​​detected by the shift position sensor 22 and the vehicle speed sensor 23.

[0066] [3. Operation] Next, the operation of the dialogue device 21 in this embodiment will be described. Figure 4 is a flowchart showing the operation of the dialogue device 21.

[0067] When crew member P inputs prompt PR by voice, the voice recognition unit 201 converts prompt PR into text and outputs the converted text to the decision unit 203 (step S1).

[0068] The decision unit 203 refers to the policy DB 222 and determines the policy and the type of response statement SE2 based on the text input from the speech recognition unit 201 and the cognitive difficulty level determination result input from the judgment unit 202 (step S2).

[0069] The decision unit 203 outputs the policy and the type of response statement SE2 determined in step S2 to the first generation unit 204 and the second generation unit 205 (step S3).

[0070] Next, the first generation unit 204 determines whether or not to generate the interlude statement SE1 (step S4).

[0071] If the first generation unit 204 determines that it does not generate the interlude sentence SE1 (step S4: NO), the second generation unit 205 generates the response sentence SE2 (step S5). Then, the output unit 206 outputs the response sentence SE2 generated in step S5 (step S6).

[0072] On the other hand, if it is determined to generate a transition statement SE1 (step S4: YES), the first generation unit 204 refers to the transition statement DB223 and generates the transition statement SE1 (step S7).

[0073] Next, the second generation unit 205 generates a response sentence SE2 using the text input from the speech recognition unit 201, the policy and response sentence SE2 type generated in step S3, and the interlude sentence SE1 generated in step S7 (step S8).

[0074] Next, the output unit 206 outputs the interlude statement SE1 generated in step S7 and the response statement SE2 generated in step S8 (step S9).

[0075] The following describes specific examples of dialogue between the dialogue device 21 and crew member P with reference to Figures 5, 6, and 7.

[0076] Figure 5 shows an example of a conversation between the dialogue device 21 and crew member P. Figure 5 illustrates a case where the level of cognitive difficulty is not "high."

[0077] Figure 5 illustrates a case where occupant P inputs the prompt PR, "It's okay if we exceed the legal speed limit by 10 km / h, right?", into the dialogue device 21 via voice. This prompt PR includes words related to driving.

[0078] In Figure 5, the dialogue device 21 decides on a policy of "negation" and determines that the type of response sentence SE2 is a long sentence. Therefore, in the example in Figure 5, the dialogue device 21 outputs a filler sentence SE1 with "No." As a result, crew member P can predict that the input prompt PR will be rejected before the response sentence SE2 is output.

[0079] In the example in Figure 5, the dialogue device 21 outputs an interlude sentence SE1 saying "No," and then, when a predetermined trigger occurs, outputs a response sentence SE2 saying "I wouldn't recommend it. Let's prioritize safety." This allows occupant P to recognize that their input prompt PR has been rejected. In addition, in the example in Figure 5, occupant P can recognize that they must drive with safety as their top priority. Furthermore, because the interlude sentence SE1, which contains content related to the content of response sentence SE2, is output before response sentence SE2, occupant P can recognize the content of response sentence SE2 without feeling any discomfort.

[0080] As mentioned above, the transitional sentence SE1 is not included at the beginning of the response sentence SE2. Therefore, in the case of Figure 5, the dialogue device 21 does not output the response sentence SE2, "Well, I wouldn't really recommend it. Let's prioritize safety," after the transitional sentence SE1, "Well, well." This prevents crew member P from feeling uncomfortable with the transitional sentence SE1 being output multiple times.

[0081] Figure 6 shows an example of a conversation between the dialogue device 21 and crew member P. Figure 6 illustrates a case where the level of cognitive difficulty is not "high."

[0082] Figure 6 illustrates a scenario where crew member P inputs the prompt PR, "Soy sauce is better than salt for ramen," via voice. This prompt PR does not contain any words related to driving.

[0083] In Figure 6, the dialogue device 21 decides on a policy of "negation" and determines that the type of response sentence SE2 is a short sentence. Therefore, in the example in Figure 6, the dialogue device 21 outputs the response sentence SE2, "Salt is delicious too," without outputting an interlude sentence SE1.

[0084] Figure 7 shows an example of a conversation between the dialogue device 21 and crew member P. Figure 7 illustrates a case where the level of cognitive difficulty is "high".

[0085] Figure 7 illustrates a scenario where occupant P inputs the prompt PR via voice, "You don't seem to be planning on obeying the stop sign." This prompt PR includes words related to driving.

[0086] In the example in Figure 7, the dialogue device 21 decides on a policy of "negation" and determines the type of response sentence SE2 to be a short sentence. Therefore, in the example in Figure 7, the dialogue device 21 outputs the response sentence SE2, "No," without outputting an interlude sentence SE1.

[0087] [4. Other Embodiments] The embodiments described above are merely examples and can be modified and applied as needed.

[0088] In the embodiments described above, an example was given in which the processor 200 of the dialogue device 21 functions as a speech recognition unit 201, a determination unit 202, a decision unit 203, a first generation unit 204, a second generation unit 205, and an output unit 206. In other embodiments, the processor of a server device connected to a network NW may function as at least one of the speech recognition unit 201, a determination unit 202, a decision unit 203, a first generation unit 204, and a second generation unit 205. When the processor of the server device connected to a network NW functions as a speech recognition unit 201, the server device obtains the prompt PR entered by the occupant P from the vehicle 1. When the processor of the server device connected to a network NW functions as a determination unit 202, the server device obtains from the vehicle 1 whether there is a passenger and the brightness outside the vehicle 1. When the processor of the server device connected to a network NW functions as a decision unit 203, the server device stores the policy DB 222. Furthermore, if the processor of the server device connected to the network NW functions as the first generation unit 204, the server device stores the intermediary message DB223 and sends the generated intermediary message SE1 to the vehicle 1. Furthermore, if the processor of the server device connected to the network NW functions as the second generation unit 205, the server device sends the generated reply message SE2 to the vehicle 1.

[0089] In other embodiments, the terminal device 3 may interact with the crew member P instead of, or together with, the dialogue device 21. In this other embodiment, the processor of the terminal device 3 functions as a speech recognition unit 201, a determination unit 202, a decision unit 203, a first generation unit 204, a second generation unit 205, and an output unit 206. Also in this other embodiment, the memory of the terminal device 3 stores a policy DB 222 and a linking statement DB 223. Of the functional units of the processor of the terminal device 3, at least one of the speech recognition unit 201, the determination unit 202, the decision unit 203, the first generation unit 204, and the second generation unit 205 may be executed as a function of the processor of a server device connected to the network NW.

[0090] In the above-described embodiment, the crew member P and the dialogue device 21 or terminal device 3 communicate via voice. In other embodiments, communication may be conducted via text, or by converting gestures or sign language into text. In these other embodiments, crew member P inputs a text prompt PR to the dialogue device 21 or terminal device 3. In these other embodiments, the input text prompt PR is output to the decision unit 203 without going through the voice recognition unit 201. In these other embodiments, the output unit 206 outputs the response sentence SE2, or the interlude sentence SE1 and response sentence SE2, via information display. In these other embodiments, the output unit 206 may output the response sentence SE2, or the interlude sentence SE1 and response sentence SE2, via sound output. In another embodiment, text associated with gestures is set in advance, the processor 200 acquires the movements of the occupant P using sensors such as a camera (an in-vehicle camera such as a driver monitoring camera), and the text corresponding to the gesture obtained from the acquired movements is input to the decision unit 203. In yet another embodiment, the processor 200 acquires the movements of the occupant P using sensors such as a camera, and the processor 200 inputs the acquired movements to an image recognition AI or a visual language model to generate text representing the content of the gesture, and the generated text is input to the decision unit 203.

[0091] In other embodiments, the output unit 206 may change the output manner of the interlude message SE1 and the response message SE2 based on information relating to the crew member P or settings made by the crew member P. Examples of information relating to the crew member P include the age of the crew member P and the gender of the crew member P. For example, if the age of crew member P corresponds to the age of a child, the output unit 206 outputs the interlude sentence SE1 and the response sentence SE2 in the voice of a child or character. Furthermore, for example, if the age of crew member P corresponds to an adult age, the output unit 206 outputs the interlude sentence SE1 and the response sentence SE2 in a calm, adult voice. Furthermore, for example, if the crew member P is male, the output unit 206 outputs the interlude sentence SE1 and the response sentence SE2 in a female voice. Furthermore, for example, if the crew member P is female, the output unit 206 outputs the interlude sentence SE1 and the response sentence SE2 in a male or female voice. Furthermore, for example, if the age of crew member P corresponds to that of a child, the output unit 206 outputs the interlude sentence SE1 and the response sentence SE2 in the tone of a child or character. Furthermore, for example, if the age of crew member P corresponds to an adult age, the output unit 206 outputs the interlude sentence SE1 and the response sentence SE2 in an adult tone. Furthermore, if, for example, crew member P has set a preferred tone of voice or voice, the output unit 206 will output the interlude sentence SE1 and the response sentence SE2 in the tone of voice or voice that crew member P has set. Furthermore, information related to crew member P and information set by crew member P are stored in a storage device accessible by the output unit 206 (such as memory 220 or a database connected to the network NW).

[0092] In other embodiments, the first generation unit 204 may determine the type of interlude statement SE1 based on information relating to the crew member P or settings made by the crew member P. For example, if the age of crew member P corresponds to the age of a child, the first generation unit 204 generates a transitional statement SE1 that is used by children but not by adults. Furthermore, for example, if the age of crew member P corresponds to an adult age, the first generation unit 204 generates a transitional statement SE1 that is generally used by adults but not generally used by children. Furthermore, for example, if crew member P has set their preferred transitional sentence SE1 for each policy, the output unit 206 will generate the transitional sentence SE1 set by crew member P.

[0093] In other embodiments, the output unit 206 may differentiate the output modes of the interlude statement SE1 and the response statement SE2 based on the policy determined by the decision unit 203. For example, if the decision unit 203 determines a policy to be "negative," the output unit 206 outputs a filler sentence SE1 in a child's voice and a response sentence SE2 in an adult's voice. Furthermore, for example, if the policy decided by the decision unit 203 is "negative," the output unit 206 outputs the interlude sentence SE1 in a tone of voice that conveys a sense of fear, and the response sentence SE2 in a tone of voice that conveys a sense of calm. In this example, by attracting the attention of crew member P with the interlude sentence SE1 and outputting the response sentence SE2 in a calm tone of voice, it becomes possible to ensure that the content of the response sentence SE2 is accurately perceived. Furthermore, for example, if the decision unit 203 determines the policy to be "affirmative," the output unit 206 outputs the interlude sentence SE1 in a cheerful voice and the response sentence SE2 in an even cheerful voice. A cheerful voice is one in which the listener can imagine the speaker smiling, and the tone is not too high. Furthermore, for example, if the policy decided by the decision unit 203 is "greeting," the output unit 206 outputs the interlude sentence SE1 in a cheerful voice and the response sentence SE2 in a calm voice.

[0094] In the embodiments described above, a transitional statement SE1 was used as an example of the "first statement," and a response statement SE2 was used as an example of the "second statement." However, the "first statement" is not limited to a transitional statement SE1, and the "second statement" is not limited to a response statement SE2. For example, if multiple transitional statements SE1 are output before outputting a response statement SE2 to a prompt PR, the "first statement" may be one transitional statement SE1, and the "second statement" may be a transitional statement SE1 different from the one specified. For example, if multiple response statements SE2 are output in response to the content of a prompt PR, the "first statement" may be one response statement SE2, and the "second statement" may be a response statement SE2 different from the one specified.

[0095] In other embodiments, the robot or avatar may perform actions such as shaking its head or changing its facial expression in accordance with the policy. In this case, the policy determined by the decision unit 203 is output to the device that controls the robot or avatar. For example, if the policy is "negative," the robot or avatar may shake its head from side to side or make a stern face. Also, for example, if the policy is "positive," the robot or avatar may nod its head up and down or make a happy face. Furthermore, the robot or avatar may act in conjunction with or before the utterance of the interlude sentence SE1 and the response sentence SE2. This allows the user (for example, crew member P in the above embodiment) to predict the content of the verbal expression by having the nonverbal expression output before the verbal expression. Alternatively, the simultaneous output of a consistent verbal expression and a nonverbal expression can contribute to facilitating the user's understanding of their utterance intent.

[0096] In the embodiment described above, the level of cognitive difficulty is determined based on the traffic conditions at the current location of vehicle 1, the brightness around vehicle 1, and the presence or absence of passengers. In other embodiments, the level of cognitive difficulty may be determined based on at least one of the traffic conditions at the current location of vehicle 1, the brightness around vehicle 1, the difficulty level of the work performed by occupant P on vehicle 1, and the presence or absence of passengers. In other embodiments, the level of cognitive difficulty may be determined based on factors other than the traffic conditions at the current location of vehicle 1, the brightness around vehicle 1, and the presence or absence of passengers. For example, in other embodiments, the level of cognitive difficulty may be determined based on the biometric data of occupant P or the operation data of vehicle 1 (data on the amount of accelerator and brake pedal depression and steering wheel operation). Note that if traffic conditions are not considered in determining the level of cognitive difficulty, the dialogue system 1000 does not need to be equipped with a traffic information server 2. Note that in other embodiments described later, it is mentioned that the "dialogue system" and "dialogue device" are applicable to vehicles other than vehicle 1, but in these other embodiments described later, the level of cognitive difficulty may be determined based on the difficulty level of the work performed by the user (the person interacting with the dialogue system). In this case, the difficulty level of the task is determined based on factors such as whether the user is folding laundry or handling a knife.

[0097] In the embodiment described above, the cognitive difficulty level is determined to fall into one of three levels. In other embodiments, the cognitive difficulty level may be determined to fall into any of N levels (where N is an integer less than or equal to 2, or an integer greater than or equal to 4).

[0098] In the embodiment described above, the policy DB222 and the interrogation statement DB223 are stored in memory 220. In other embodiments, at least one of the policy DB222 and the interrogation statement DB223 may be stored in a storage device accessible by the processor 200. An example of such a storage device is a device connected to a network NW.

[0099] In the embodiment described above, the occupant P of vehicle 1 was used as an example of the "user." However, if the "dialogue device" is a device mounted on a mobile body other than vehicle 1, the "user" may be the user of that mobile body.

[0100] In the embodiment described above, the "terminal device" was exemplified as a terminal device used by crew member P. However, the "terminal device" may also be a home appliance such as a smart speaker or a smart TV.

[0101] In the embodiments described above, the "dialogue system" was exemplified as a system mounted on vehicle 1. However, the "dialogue system" is not limited to a system mounted on vehicle 1. For example, the "dialogue system" may be a system applied inside a house, or a system applied to a mobile object other than vehicle 1, or to an object other than a mobile object. For example, the "dialogue system" may be a monitoring robot or an AI speaker capable of communicating inside a house.

[0102] The processor 200 may consist of multiple processors or a single processor. The processor 200 may also be hardware programmed to implement the functions described above. In this case, the processor 200 may consist of, for example, an ASIC (Application Specific Integrated Circuit) or an FPGA (Field Programmable Gate Array).

[0103] Furthermore, the configuration of each part of the dialogue device 21 shown in Figure 2 is merely an example, and the specific implementation form is not particularly limited. In other words, it is not necessarily required that hardware corresponding to each part be implemented individually, and it is certainly possible to configure the system so that a single processor executes a program to realize the functions of each part. Also, some of the functions realized by software in the above embodiment may be implemented as hardware, or some of the functions realized by hardware may be implemented as software.

[0104] Furthermore, the operation steps shown in Figure 4 are divided according to the main processing content, and the present invention is not limited by the way the processing units are divided or their names. Depending on the processing content, it may be further divided into more steps. Alternatively, it may be divided so that one step unit includes even more processing. Also, the order of the steps may be changed as appropriate, as long as it does not impede the spirit of the present invention.

[0105] Furthermore, when the dialogue method using the dialogue device 21 described above is implemented using the processor 200, the program to be executed by the processor 200 can be configured as a recording medium or a transmission medium for transmitting this program. In other words, the control program 221 can also be implemented by recording it on a portable information recording medium. Examples of information recording media include magnetic recording media such as hard disks, optical recording media such as CDs, and semiconductor storage devices such as USB (Universal Serial Bus) memory and SSDs (Solid State Drives), but other recording media can also be used.

[0106] [5. Configurations supported by the above embodiment] The above embodiment supports the following configuration.

[0107] (Composition 1) A dialogue device that interacts with a user by voice or text, comprising: a decision unit that determines a policy for the content of the response to a prompt input by the user; a first generation unit that generates a first sentence based on the policy determined by the decision unit; a second generation unit that generates a second sentence based on the policy determined by the decision unit; and an output unit that outputs the first sentence generated by the first generation unit and the second sentence generated by the second generation unit by voice or text. According to the dialogue device of Configuration 1, by generating the first and second sentences based on a common policy, it is possible to generate the first and second sentences, which are related in content. Therefore, it is possible to output multiple sentences that are related in content in response to a prompt entered by the user. Furthermore, because it is possible to output multiple sentences that are related in content, it is possible to suppress the user from feeling a sense of incongruity during the dialogue.

[0108] (Configuration 2) The dialogue device according to configuration 1, wherein the first sentence is a filler or interjection indicating an acknowledgment, the second sentence is a response to the content of the prompt, and the output unit outputs the first sentence before the second sentence. The dialogue device of Configuration 2 can generate relevant transitional sentences and response sentences. Therefore, it can output relevant transitional sentences and response sentences in response to prompts entered by the user. In addition, since the transitional sentences are output first, it is possible to suppress the user from feeling discomfort or unease due to pauses during the dialogue.

[0109] (Composition 3) The dialogue device according to configuration 1 or configuration 2, wherein the first generation unit refers to the correspondence between the policy and the first sentence in a storage unit that stores the first sentence for each policy, and generates the first sentence corresponding to the policy. According to the dialogue device of configuration 3, it is possible to generate a first sentence corresponding to the policy, thus enabling the generation of a first sentence that aligns with the policy. Therefore, it is possible to output a sentence that aligns with the policy in the response content of the prompt.

[0110] (Composition 4) The dialogue device according to any one of configurations 1 to 3, wherein the second generation unit generates a second sentence that does not begin with the first sentence generated by the first generation unit. According to the dialogue device of configuration 4, when outputting the first and second sentences, it is possible to suppress the consecutive output of the first sentence. Therefore, it is possible to suppress the consecutive output of the same content in response to prompts entered by the user. Furthermore, because the consecutive output of the same content is suppressed, it is possible to suppress the user from feeling uncomfortable during the dialogue.

[0111] (Composition 5) The dialogue device according to any one of configurations 1 to 4, wherein the determination unit determines the type of the second sentence to be a sentence of a first length indicating a predetermined length or longer, or a sentence of a second length shorter than the predetermined length, and the first generation unit determines whether or not to generate the first sentence based on the type of second sentence determined by the determination unit, and generates the first sentence if it is determined to generate the first sentence. According to the dialogue device of configuration 5, if the length of the second sentence is short, it does not take time to generate the second sentence, and therefore, if the length of the second sentence is short, the process can be executed without generating the first sentence. Thus, the processing load on outputting to prompts can be reduced.

[0112] (Composition 6) The dialogue device according to configuration 5, wherein the determination unit determines the length of the second sentence based on the content of the prompt. According to the dialogue device of configuration 6, the first sentence can be output or not output depending on the content of the prompt, thus avoiding dialogues in which the first sentence is always output.

[0113] (Composition 7) The dialogue device according to configuration 5 or 6, further comprising a determination unit that determines a cognitive difficulty level indicating the degree of difficulty the user has in recognizing the content of the second sentence, and a determination unit that determines the length of the second sentence based on the cognitive difficulty level determined by the determination unit. According to the dialogue device of configuration 7, the higher the level of cognitive difficulty, the more difficult it becomes to understand the content of the second sentence. Therefore, by determining the length of the second sentence based on the level of cognitive difficulty, it is possible to output sentences that are easier to understand according to the level of cognitive difficulty.

[0114] (Composition 8) The dialogue device according to configuration 7, wherein the output unit changes the timing of outputting at least one of the first sentence and the second sentence based on the cognitive difficulty level determined by the determination unit. According to the dialogue device of configuration 8, by changing the timing of sentence output based on the level of cognitive difficulty, the likelihood of the content of the outputted sentence being recognized can be increased.

[0115] (Composition 9) The output unit changes the output mode of the first and second sentences based on information relating to the user or settings made by the user, as described in any one of configurations 1 to 8. According to the dialogue device of configuration 9, the output manner of the first and second sentences can be changed based on user-related information or user settings, thereby increasing the acceptability of the content of the output sentences.

[0116] (Composition 10) The output unit is an interactive device according to any one of configurations 1 to 9, which causes the output modes of the first sentence and the second sentence to differ based on the policy. According to the dialogue device of configuration 10, the first and second sentences are not output in the same manner, but rather their output manner is made different based on a policy, thereby increasing the acceptability of the content of the output sentences.

[0117] (Composition 11) A terminal device that interacts with a user by voice or text, comprising: a determination unit that determines a policy for the content of the response to a prompt input by the user; a first generation unit that generates a first sentence based on the policy determined by the determination unit; a second generation unit that generates a second sentence based on the policy determined by the determination unit; and an output unit that outputs the first sentence generated by the first generation unit and the second sentence generated by the second generation unit in voice or text. The dialogue method of configuration 11 produces the same effect as the dialogue device of configuration 1.

[0118] (Composition 12) A dialogue method for interacting with a user by voice or text, wherein a processor determines a policy for the content of the response to a prompt input by the user, generates a first sentence based on the determined policy, generates a second sentence based on the determined policy, and outputs the generated first sentence and the generated second sentence by voice or text. The dialogue method of configuration 12 produces the same effect as the dialogue device of configuration 1.

[0119] (Composition 13) A dialogue system that interacts with a user by voice or text, comprising: a decision unit that determines a policy for the content of the response to a prompt input by the user; a first generation unit that generates a first sentence based on the policy determined by the decision unit; a second generation unit that generates a second sentence based on the policy determined by the decision unit; and an output unit that outputs the first sentence generated by the first generation unit and the second sentence generated by the second generation unit by voice or text. The dialogue system of configuration 13 produces the same effect as the dialogue device of configuration 1. [Explanation of symbols]

[0120] 1...Vehicle, 2...Traffic information server, 3...Terminal device, 10A...Driver's seat, 10B...Passenger seat, 10C...Rear right seat, 10D...Rear left seat, 11,11A~11D...Seat sensor, 12...Dashboard, 13...Touch panel, 14...Speaker, 15...Microphone, 16...Shift lever, 17...Illuminance sensor, 18...Windshield, 19...Camera, 19A...Front camera, 19B...Rear camera, 19C...Right side camera, 19D...Left side camera, 20...TCU, 21...Dialogue device, 22...Shift position sensor, 23...Vehicle speed sensor, 24...GNSS unit, 200...Processor, 201...Vehicle recognition unit, 202...Determination unit, 203...Decision unit, 204...First generation unit, 205...Second generation unit, 206...Output unit, 220...Memory (storage unit), 221 ...control program, 222...policy DB, 223...interruption message DB, 1000...dialogue system, NW...network, P...crew member (user), PR...prompt, SE1...interruption message (first sentence), SE2...response message (second sentence).

Claims

1. A dialogue device that interacts with the user via voice or text, A decision unit that determines the policy for the content of the response to the prompt entered by the user, Based on the policy determined by the decision unit, a first generation unit generates a first sentence, Based on the policy determined by the aforementioned decision unit, a second generation unit generates a second sentence, The system comprises an output unit that outputs the first sentence generated by the first generation unit and the second sentence generated by the second generation unit in either voice or text format. Dialogue device.

2. The first sentence above is a filler or interjection indicating agreement, The second sentence above is a response to the content of the prompt, The output unit outputs the first sentence before the second sentence. The dialogue device according to claim 1.

3. The first generation unit is, For each of the aforementioned policies, the system refers to the correspondence between the policy and the first sentence in the storage unit that stores the first sentence, and generates the first sentence corresponding to the policy. The dialogue device according to claim 1 or 2.

4. The second generation unit is, The first generation unit generates a second sentence that does not begin with the first sentence generated by the first generation unit. The dialogue device according to claim 1 or 2.

5. The determination unit determines the type of the second sentence to be either a sentence of a first length indicating a predetermined length or longer, or a sentence of a second length shorter than the predetermined length. The first generation unit is, Based on the type of the second sentence determined by the determination unit, it is determined whether or not to generate the first sentence. If it is determined that the first sentence will be generated, then generate the first sentence. The dialogue device according to claim 1 or 2.

6. The determination unit determines the length of the second sentence based on the content of the prompt. The dialogue device according to claim 5.

7. The system includes a determination unit that determines a cognitive difficulty level indicating the degree of difficulty the user has in understanding the content of the second sentence, The determination unit determines the length of the second sentence based on the level of cognitive difficulty determined by the judgment unit. The dialogue device according to claim 5.

8. The output unit changes the timing of outputting at least one of the first sentence and the second sentence based on the cognitive difficulty level determined by the determination unit. The dialogue device according to claim 7.

9. The output unit changes the output configuration of the first and second sentences based on information relating to the user or settings made by the user. The dialogue device according to claim 1 or 2.

10. The output unit makes the output modes of the first sentence and the second sentence different based on the policy. The dialogue device according to claim 1 or 2.

11. A terminal device that interacts with the user via voice or text, A decision unit that determines the policy for the content of the response to the prompt entered by the user, Based on the policy determined by the decision unit, a first generation unit generates a first sentence, Based on the policy determined by the aforementioned decision unit, a second generation unit generates a second sentence, The system comprises an output unit that outputs the first sentence generated by the first generation unit and the second sentence generated by the second generation unit in either voice or text format. Terminal device.

12. A method of interaction with a user via voice or text, The processor, Determine the policy for the response content to the prompt entered by the user, Based on the aforementioned policy that was decided, the first sentence is generated, Based on the aforementioned policy that was decided, the second sentence is generated, The generated first sentence and the generated second sentence are output as speech or text. Methods of dialogue.

13. A dialogue system that interacts with the user via voice or text, A decision unit that determines the policy for the content of the response to the prompt entered by the user, Based on the policy determined by the decision unit, a first generation unit generates a first sentence, Based on the policy determined by the aforementioned decision unit, a second generation unit generates a second sentence, The system comprises an output unit that outputs the first sentence generated by the first generation unit and the second sentence generated by the second generation unit in either voice or text format. Dialogue system.

Citation Information

Patent Citations

  • Automatic testing system for printed circuit board

    JP1986050077A