Information processing system and information processing method

The system adjusts positive and negative sentence ratios in dialogue based on user proficiency to improve understanding and maintain variety, addressing comprehension issues in mixed emotion dialogue systems.

JP7767439B2Active Publication Date: 2025-11-11NISSAN MOTOR CO LTD +1
View PDF 10 Cites 0 Cited by

Patent Information

Application Number
JP2023544788
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-09-06
Publication Date
2025-11-11
Estimated Expiration
2041-09-06

AI Technical Summary

Technical Problem

Conventional dialogue systems struggle with mixed positive and negative emotion words, making it difficult for users to understand the content, especially when using agent devices, and limiting the dialogue expression if only positive words are used.

Method used

An information processing system that adjusts the ratio of positive and negative sentences in dialogue data based on user proficiency with the agent device, increasing positive sentences for low proficiency and allowing negative sentences for high proficiency to enhance understanding and maintain variety.

Benefits of technology

Facilitates easy user comprehension and maintains diverse dialogue expressions by tailoring sentence ratios to user familiarity, enhancing user affinity and trust with the agent.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007767439000001
    Figure 0007767439000001
  • Figure 0007767439000002
    Figure 0007767439000002
  • Figure 0007767439000003
    Figure 0007767439000003
Patent Text Reader

Abstract

Provided are an information processing system 1 and an information processing method which, while using a variety of dialogue expressions, enable communications that are easy for a user to understand. The present invention detects a user's proficiency level with respect to an agent device 5, and generates dialogue sentence data according to the user's proficiency level using response sentence information classified into positive sentences and negative sentences. In this case, the present invention changes the proportion of positive sentences and the proportion of negative sentences used in the dialogue sentence data, depending on whether the user's proficiency level is relatively low or relatively high.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an information processing system and an information processing method. [Background technology]

[0002] A dialogue system is known in which a positive emotion word expressing a positive meaning, such as "I love you," and a negative emotion word expressing a negative meaning, such as "I'm tired," are associated with a certain word and stored in a database, and when a word included in a received input sentence is stored in the database, a reply sentence is created using a combination of multiple emotion words associated with that word (see Patent Document 1). [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Application Publication No. 2017-157011 Summary of the Invention [Problem to be solved by the invention]

[0004] In the above-mentioned conventional dialogue system, the responses created contain an irregular mixture of positive and negative emotion words. However, in a typical dialogue, if the content uttered by a speaker contains a mixture of sentences with positive meanings and sentences with negative meanings, it can be difficult for the listener to understand the content. By unifying the content to sentences with positive meanings, the meaning is conveyed more directly, making it easier for the listener to understand the content.

[0005] Even in a communication device using an agent function, if sentences with positive and negative meanings are mixed, it may be difficult for the user to instantly understand what is being said by the agent, especially if the user is unfamiliar with the agent device or is unable to pay much attention to the agent device. On the other hand, if sentences with negative meanings are not used at all, the dialogue expression is limited and the agent function cannot be fully utilized. Therefore, it is desirable for a device using so-called agent functionality to be able to communicate in a way that is easy for the user to understand, using a variety of dialogue expressions.

[0006] The problem to be solved by the present invention is to provide an information processing system and an information processing method that can carry out communication that is easy for users to understand while using a variety of dialogue expressions. [Means for solving the problem]

[0007] The present invention provides Estimate the load on the user when recognizing the dialogue data, and User familiarity with the agent device , based on an operation history including a specific operation on the agent device. The response sentence information is then classified into positive and negative sentences, and the user's Load and Dialogue data is generated according to the level of proficiency. When the user's load is relatively high, dialogue data is generated so that the proportion of positive sentences used in the dialogue data is higher than when the user's load is relatively low; The above problem is solved by changing the ratio of positive sentences and negative sentences used in the dialogue data depending on whether the user's proficiency level is relatively low or relatively high. [Effects of the Invention]

[0008] According to the present invention, it is possible to communicate with the user in a way that is easy for the user to understand, using a variety of dialogue expressions. In particular, when providing information using an agent function, the dialogue expressions are changed according to the user's situation, which is expected to have the effect of making the user more likely to have a favorable impression and feel a sense of affinity with the agent. [Brief explanation of the drawings]

[0009] [Figure 1] 1 is a block diagram showing a first embodiment of an information processing system according to the present invention. [Figure 2] 2 is a diagram showing an example of the location of the agent device in FIG. 1 inside a vehicle. [Figure 3] 2 is a diagram showing an example of the configuration of an occupant information database (driver) that has been subjected to information processing by the information processing device of FIG. 1. FIG. [Figure 4] 2 is a diagram for explaining information processing of input data executed by the information processing device of FIG. 1. FIG. [Figure 5] 1. FIG. 4 is a diagram showing an example of generation of response sentence information executed by the information processing device of FIG. [Figure 6] 4 is a flowchart showing an example of an information processing procedure based on a driver's proficiency level, which is executed by the information processing device of FIG. 1; [Figure 7A] FIG. 7 is a diagram showing an example of a scene in which dialogue data is generated by performing the information processing in steps S107 to S109 in FIG. 6. [Figure 7B] FIG. 7 is a diagram showing an example of a scene in which dialogue data is generated by performing the information processing in step S106 in FIG. 6. [Figure 7C] FIG. 7 is a diagram showing an example of a scene in which dialogue data is generated by performing the information processing in step S105 in FIG. 6. [Figure 8] 7A to 7C are diagrams showing other examples of scenes for which dialogue data corresponding to each of FIGS. 7A to 7C is generated. [Figure 9] FIG. 10 is a block diagram showing a second embodiment of the information processing device system according to the present invention. [Figure 10] 10 is a flowchart showing an example of an information processing procedure based on a cognitive load of a driver, which is executed by the information processing device of FIG. 9. [Figure 11A] FIG. 11 is a diagram showing an example of a scene in which dialogue data is generated by performing the information processing in steps S207 to S209 in FIG. [Figure 11B] FIG. 11 is a diagram showing an example of a scene in which dialogue data is generated by performing the information processing in step S206 in FIG. [Figure 11C]FIG. 11 is a diagram showing an example of a scene in which dialogue data is generated by performing the information processing in step S205 of FIG. [Figure 12] FIG. 10 is a block diagram showing a third embodiment of an information processing device system according to the present invention. [Figure 13] 13 is a diagram showing an example of the configuration of an occupant information database (passengers) that has been subjected to information processing by the information processing device of FIG. [Figure 14] 13 is a flowchart showing an example of an information processing procedure based on attribute information of a fellow passenger, which is executed by the information processing device of FIG. 12; [Figure 15A] FIG. 15 is a diagram showing an example of a scene in which dialogue data is generated by performing the information processing in steps S308 to S310 in FIG. [Figure 15B] FIG. 15 is a diagram showing an example of a scene in which dialogue data is generated by performing the information processing in step S307 in FIG. [Figure 15C] FIG. 15 is a diagram showing an example of a scene in which dialogue data is generated by performing the information processing in step S306 in FIG. DETAILED DESCRIPTION OF THE INVENTION

[0010] <<First Embodiment>> Hereinafter, embodiments of the present invention will be described with reference to the drawings. FIG. 1 is a block diagram showing a first embodiment of an information processing system 1 according to the present invention, and FIG. 2 is a diagram showing the interior of a vehicle, showing an example of the installation location and operation of the agent device 5 shown in FIG. 1. The information processing system 1 according to the present invention uses an agent device 5 that provides information to a user through an agent function provided by an anthropomorphic agent 52 (hereinafter simply referred to as agent 52), specifically, through media such as voice, images, the movement of a character robot, and a combination thereof, or that interacts with the user. The information processing system 1 detects the user's level of familiarity with the agent device 5, and changes the proportion of positive sentences and the proportion of negative sentences used in the dialogue data output by the agent device 5 according to the user's level of familiarity.

[0011] In this embodiment, a user refers to a person who uses the information processing system 1. Hereinafter, an example will be described in which the information processing system 1 is applied to a driver D of a vehicle. However, the user may also be a passenger X other than the driver D (hereinafter, the driver and passenger will also be simply referred to as passengers). While the following describes an example in which the agent device 5 is provided in a vehicle, the form and installation location of the agent device 5 are not limited thereto. The agent device 5 may be any electronic device equipped with an agent function, such as a portable speaker-type electronic device or an electronic device with a display. Furthermore, the audio and video output functions of the agent device 5 described below may be incorporated into a mobile phone such as a smartphone. The anthropomorphized agent 52 is merely an example, and the agent 52 may not necessarily resemble a human, but may instead be an animal, plant, a specific character, an avatar, or an icon. The agent 52 may be provided as a physical entity, or the shape of a human, animal, plant, specific character, or the like may be displayed as an image on a display.

[0012] 1 and 2, the agent device 5 of this embodiment has a human-like character robot 52 mounted on a base 51 so that it can appear and disappear using an actuator (not shown). When the agent device 5 receives a control command from the output unit 27 and communicates with the passenger using the agent function, it appears from the base 51 as shown in the lower diagram of Fig. 2, and when communication with the passenger is finished, the character robot 52 is stored in the base 51 as shown in the upper diagram of Fig. 2.

[0013] The agent device 5 includes a speaker or other audio output unit for outputting voice and sound effects, and a display or other display unit for displaying images including text, and outputs communication information by providing voice, sound effects, text, and other images to the occupant along with the movement of the character robot 52. Note that, in this embodiment, an example will be described in which the agent is a three-dimensional object such as the character robot 52, but the agent is not limited to this and may be a two-dimensional image displayed on a display as shown in Fig. 8.

[0014] Returning to FIG. 1 , the information processing system 1 of this embodiment includes an information processing device 2 including a ROM (Read Only Memory) storing programs for executing various processes, a CPU (Central Processing Unit) as an operating circuit that functions as an information processing device by executing the programs stored in the ROM, and a RAM (Random Access Memory) that functions as an accessible storage device; vehicle sensors 3; an input device 4; and an agent device 5 as an output device. These devices are connected, for example, via a Controller Area Network (CAN) or other in-vehicle LAN, and can transmit and receive information to and from each other. In terms of the functional configuration achieved by executing the information processing program, the information processing device 2 includes an occupant identification unit 21, an occupant information database 22, a proficiency detection unit 23, an input data processing unit 24, a response sentence information database 25, a data generation unit 26, and an output unit 27, as shown in FIG. 1 .

[0015] The occupant identification unit 21 identifies the driver D based on an input signal from the vehicle sensors 3 and stores the identification at least temporarily. In the scene shown in FIG. 2 , the occupant identification unit 21 identifies the person sitting in the driver's seat as the driver D based on image data captured by an in-vehicle camera serving as the vehicle sensors 3. The identification information of the identified driver D is output to the proficiency detection unit 23. Note that the driver D may be identified using an input signal from a seating sensor provided inside the seating portion of the seat or a seat belt sensor. Alternatively, the driver D may be identified by storing the identification data of the driver D in a wireless key for the vehicle, and automatically or semi-automatically reading the identification data of the driver D when the vehicle is unlocked or started using the wireless key.

[0016] Furthermore, the occupant identification unit 21 generates riding information by associating the detection values ​​acquired from the vehicle sensors 3 with the identification information of the identified driver D, and stores the information in the occupant information database 22. The riding information is a riding history that records the time when the driver D boarded the vehicle and the time when he / she disembarked from the vehicle. The riding information may also include the usage history of on-board devices such as a navigation device and a driving assistance device acquired from the vehicle sensors 3.

[0017] Furthermore, the occupant identification unit 21 associates the detection values ​​acquired from the vehicle sensors 3 with the identified identification information of the driver D to generate usage information of the agent device 5 by the driver D, and stores it in the occupant information database 22. The usage information includes a usage history that records the time when the driver D started to use the agent device 5 and the time when he finished using it, an operation history such as a specific operation such as changing settings and canceling the operation (returning to the previous operation, canceling the operation), and an input data history such as voice information and text information input by the driver D. This usage information of the agent device 5 is used to calculate the driver D's proficiency with the agent device 5.

[0018] The occupant information database 22 stores the driver D's riding information and the driver D's use information of the agent device 5, which are generated by the occupant identification unit 21. FIG. 3 is a diagram showing an example configuration (driver) of the occupant information database 22. As shown in FIG. 3, the occupant attributes include, for example, recording the father as user 1, and storing the riding history of the father of user 1 getting on and off the vehicle as riding information. Furthermore, the usage information of the agent device 5 stores the usage history of the agent device 5 when the father of user 1 got on the vehicle, as well as the setting change, the operation history of stopping the conversation of the agent device 5, and the input data history such as the voice information and the text information input by the father of user 1. Furthermore, the occupant information database 22 may include proficiency information for each user, which is calculated based on the use information of the agent device 5. While the occupant information database 22 is configured to be included in the information processing device 2, it may also be configured to store or acquire various information by communicating with an external server.

[0019] When the proficiency detection unit 23 acquires the identification information of driver D from the occupant identification unit 21, it refers to the occupant information database 22 and calculates the proficiency using the usage information of the agent device 5 of driver D. This is because by using the usage information of the agent device 5, it is possible to calculate the proficiency according to the actual usage situation of driver D. The proficiency is an index that indicates how familiar and skilled driver D is with using the agent device 5. A higher proficiency indicates that driver D is more familiar with using the agent device 5, and a lower proficiency indicates that driver D is less familiar with using the agent device 5. The proficiency of driver D calculated by the proficiency detection unit 23 is stored in the occupant information database 22 and is also output to the data generation unit 26.

[0020] The method for calculating the proficiency of the driver D is not particularly limited, but the proficiency may be calculated by using the usage information of the agent device 5, for example, by adding a score every time the cumulative time of using the agent device 5 exceeds 50 hours or every time the cumulative number of times the agent device 5 is used exceeds 10. Also, the operation history of the agent device 5 may be referenced to record a specific operation and its cancellation (returning to the previous operation, canceling the operation), and if cancellation has not been detected a predetermined number of times, it may be determined that the driver D has become accustomed to the specific operation and a score may be added. Furthermore, a score may be added depending on the frequency with which the driver D inputs information in response to information output by the agent device 5 or the level of detail of the content input by the driver D.

[0021] The proficiency detection unit 23 may calculate the proficiency using any one of the exemplified information, or may calculate the proficiency by combining a plurality of pieces of information. In addition to the usage information for the agent device 5, familiarity with the vehicle, such as the driving time that the driver D has driven the vehicle and the usage history of on-board devices such as a navigation device and a driving assistance device, may also be taken into consideration. If the driver D is a beginner driver, the proficiency with the agent device 5 can be appropriately calculated by taking familiarity with the vehicle into consideration. Note that the method for calculating the proficiency is a method of adding up scores, but is not limited to this.

[0022] The input data processing unit 24 performs speech recognition processing on the voice information of the driver D acquired from the input device 4, and classifies the input data into positive sentences or negative sentences based on the words contained in the voice information. The input device 4 is, for example, a microphone installed in the vehicle that can accept voice input. The installation location of the input device 4 is not particularly limited, but it is preferably installed near the passenger seats. Known technologies can be applied to the speech recognition processing. In the speech recognition processing, the voice information of the driver D is digitized to generate text data (character strings), and the text data is classified into positive sentences or negative sentences based on the words contained in the text data. A positive sentence refers to a word or sentence used in a positive or active sense for the driver D, and a negative sentence refers to a word or sentence used in a negative or passive sense for the driver D.

[0023] For example, as shown in the upper diagram of FIG. 4, when speech information such as "I'm happy that my car is clean after washing it" is acquired from driver D, the input data processing unit 24 digitizes the speech information to generate text data and classifies the input data based on the word "happy" included in the text data. Since the word "happy" is used with a positive or proactive meaning for driver D, the input data processing unit 24 classifies the sentence "I'm happy that my car is clean after washing it" as a positive sentence and stores it in the response sentence information database 25. Note that only the word "happy" may be stored in the response sentence information database 25.

[0024] Similarly, as shown in the lower diagram of Figure 4, when speech information such as "I hate being stuck in traffic" is acquired from driver D, the input data processing unit 24 digitizes the speech information to generate text data and classifies the input data based on the word "I hate being stuck in traffic" contained in the text data. Since the word "I hate being stuck in traffic" is used with a negative or negating meaning by driver D, the input data processing unit 24 classifies the sentence "I hate being stuck in traffic" as a negative sentence and stores it in the response sentence information database 25. Note that only the word "I hate being stuck in traffic" may be stored in the response sentence information database 25.

[0025] In this way, the input data processing unit 24 classifies words contained in the voice information of driver D, or sentences containing these words, into positive sentences if they are used in a positive or proactive sense for driver D, and into negative sentences if they are used in a negative or passive sense for driver D.

[0026] Furthermore, the classification of positive sentences and negative sentences may reflect the preferences of driver D. For example, if speech information is acquired from driver D saying, "I'm happy because team α (the baseball team I support) beat team β (the baseball team I don't support)," information about "team α (the baseball team I support)" may be classified as a positive sentence used in a positive or proactive sense for driver D, and information about "team β (the baseball team I don't support)" may be classified as a negative sentence used in a negative or passive sense for driver D.

[0027] The positive and negative sentences based on the voice information of the driver D classified by the input data processing unit 24 are stored in the response sentence information database 25 as response sentence information to be used in generating dialogue sentence data to be output from the agent device 5.

[0028] The response sentence information database 25 stores response sentence information of positive sentences and negative sentences accumulated based on the voice information of the driver D. In addition to this response sentence information, it also stores language information such as typical examples of response sentences, vocabulary such as words and phrases, and grammatical information used by the data generation unit 26 to generate dialogue sentence data. Note that although the response sentence information database 25 is configured to be included in the information processing device 2, it may also be configured to store or acquire various information by communicating with an external server.

[0029] In addition to using the positive and negative sentences accumulated based on the voice information of driver D described above to generate dialogue data, it is also possible to use information about a specific topic obtained from an external website or other external site via a telecommunications network such as the Internet by fitting it into a typical example of a response sentence, or to use response sentence information generated using a predetermined algorithm.

[0030] 5 is a diagram showing an example of response sentence information generated using a predetermined algorithm. The algorithm for generating response sentence information is not particularly limited, but various conditions are pre-programmed, such as 1) outputting using positive or negative words / sentences, 2) outputting information about the progress of the task to the user, 3) outputting information about the user, 4) outputting information that rejects the user's behavior, 5) outputting information that rejects the user's driving behavior, and 6) outputting information about the user's preferences.

[0031] For example, 1) when outputting using positive or negative words / sentences, response sentence information is generated using predetermined words or sentences with positive meanings for positive sentences, and predetermined words or sentences with negative meanings for negative sentences. In the example shown in Fig. 5, when driver D approaches the vehicle, response sentence information is generated as "Yay, it's time for a drive" using "Yay!", which has a positive meaning, as a positive sentence. On the other hand, response sentence information is generated as "Did you forget to buy something?" using "Wait!", which has a negative meaning, as a negative sentence.

[0032] Next, 2) when outputting information about the progress of a task to a user, response sentence information is generated in which the positive sentence contains content indicating that the task has been completed, and the negative sentence contains content indicating that the task has not been completed or is in an abnormal state. In the example shown in Fig. 5, for the task of washing a car for driver D, response sentence information is generated as a positive sentence, "Thank you for washing the car," indicating that the car wash has been completed. In contrast to this, response sentence information is generated as a negative sentence, "You haven't washed the car recently," indicating that the car wash has not been completed.

[0033] Next, 3) when outputting information about the user, response sentence information is generated in which the positive sentence contains content that affirms (accepts, praises, encourages, and is positive) information related to the user, and the negative sentence contains content that denies (rejects, warns, discourages, and is negative) information related to the user. In the example shown in FIG. 5, when driver D opens the door of the vehicle, response sentence information is generated as a positive sentence, "Welcome, Mr. / Ms. XX (Driver D's name)," which indicates affirmation (acceptance) of driver D opening the door. In contrast to this, response sentence information is generated as a negative sentence, "Is there anything I can help you with?", which indicates denial (rejection) of driver D opening the door.

[0034] 4) When outputting information that denies the user's behavior, the system generates response sentence information in the form of negative sentences that contain content that more strongly denies the user's behavior than the positive sentences. In the example shown in FIG. 5, regarding driver D's sudden acceleration or deceleration, the system generates response sentence information as a positive sentence such as "This will worsen fuel efficiency," while the system generates response sentence information as a negative sentence such as "It's dangerous," which more strongly denies driver D's behavior. 5) When outputting information that denies the user's driving behavior, the system may include information related to the danger to the occupants and the vehicle in the negative sentences. For example, regarding driver D's sudden steering, the system generates response sentence information as a positive sentence such as "You'll get drunk," while the system generates response sentence information as a negative sentence such as "You're driving dangerously."

[0035] 6) When outputting information related to user preferences, response sentence information containing content that matches the user's predetermined preferences is generated as a positive sentence, and response sentence information containing content that does not match the user's preferences is generated as a negative sentence. In the example shown in Fig. 5, if driver D is rooting for team α (baseball team), response sentence information such as "team α is on a winning streak" is generated as a positive sentence using content that is favorable to team α. On the other hand, response sentence information such as "team α has conceded two runs" is generated as a negative sentence using content that is unfavorable to team α.

[0036] Returning to FIG. 1 , the data generation unit 26 determines the proportion of positive sentences and negative sentences to be used in the dialogue data based on the proficiency of the driver D obtained from the proficiency detection unit 23, and generates the dialogue data to be output from the agent device 5. As described above, if sentences with positive meanings and sentences with negative meanings are mixed in communication using the agent device 5, it may be difficult for the driver D to instantly understand what is being spoken by the agent 52. This tendency is particularly strong when the driver D is unfamiliar with using the agent device. Consistently using sentences with positive meanings conveys the meaning more directly and makes it easier for the listener to understand the content, but on the other hand, not using any sentences with negative meanings at all limits the dialogue expression.

[0037] Therefore, the information processing system 1 of this embodiment generates output data by changing the ratio of positive sentences and negative sentences used in the dialogue data based on the proficiency of the driver D. This makes it possible to communicate in a way that is easy for the user to understand while using a variety of dialogue expressions. In particular, the information processing system 1 of this embodiment increases the ratio of positive sentences used in generating the dialogue data when the driver D is inexperienced in using the agent device 5, i.e., when the driver D has a low level of proficiency, compared to when the driver D is accustomed to using the agent device 5, i.e., when the driver D has a high level of proficiency.

[0038] The term "dialogue data" refers to data output from the agent device 5 using response sentence information classified into positive sentences and negative sentences, and is a collective term for multiple pieces of response sentence information output during a predetermined period or a predetermined number of times. The term "proportion of positive sentences" refers to the number of positive sentences output relative to the total number of response sentence information outputs in the "dialogue data," and the term "proportion of negative sentences" refers to the number of negative sentences output relative to the total number of response sentence information outputs in the "dialogue data." The predetermined period is not particularly limited, but may be a fixed period such as from the start to the end of a series of dialogues between the driver D and the agent device 5, until a series of dialogues on a certain topic ends, from the time the driver D gets on the vehicle until the time the driver D gets off the vehicle, from the time the driver D starts driving the vehicle until the time the driver D stops driving, or until the trip meter installed in the vehicle is reset. The predetermined number of times is not particularly limited, but may be a fixed number of times, such as every five counts of the number of times response sentence information is output, the number of times positive sentences are output, and the number of times negative sentences are output, calculated for each user using a counter (not shown). The count records of these response sentence information may be stored in the occupant information database 22. In the following, an example will be described in which the "dialogue sentence data" is applied to a scene in which the data is output during the period from when the driver D gets into the vehicle until when he gets out of the vehicle.

[0039] For example, suppose that response sentence data using a positive sentence, "Welcome, Mr. / Ms. XX (name of driver D)," is output from the agent device 5 to driver D, who is getting into a vehicle for the first time in a while and using the agent device 5. In such a situation, it can be understood that the content is accepting of driver D, regardless of whether driver D's proficiency level is low or high. In contrast, suppose that response sentence data using a negative sentence, "Is there anything I can help you with?" is output from the agent device 5. In such a situation, if driver D's proficiency level is low, the response data may be interpreted as rejecting driver D, causing the driver to feel uncomfortable with the agent device 5. On the other hand, if driver D's proficiency level is high, driver D is accustomed to communicating with the agent device 5, and the negative sentence, "Is there anything I can help you with?", can be interpreted as supplementing the intention behind the utterance of the agent device 5. For example, if driver D has not used the agent device 5 for a while and interprets the negative sentence as being intended to make the agent device 5 seem upset, driver D may actually develop an attachment to the agent device.

[0040] However, even if the driver D has a high level of proficiency, if dialogue data using negative sentences is output continuously, the driver D may feel uncomfortable. Therefore, the data generation unit 26 may use a counter (not shown) or the like to store the number of times positive sentences and negative sentences are used in dialogue data for each user, and perform control so that negative sentences are not output frequently and continuously. Furthermore, when a restraining operation by the driver D, such as interrupting dialogue data output by the agent device 5 by speaking loudly or taking an action to inhibit output, is detected, or when the driver D turns off the power of the agent device 5, there is a risk that the driver D's receptivity to the agent device 5 has decreased, and therefore control may be performed so that negative sentences are not output for a predetermined period of time.

[0041] In this way, the impression given to the driver D by negative sentences differs depending on the driver D's proficiency. Therefore, when the driver D has a low proficiency, the proportion of positive sentences used in the dialogue data is increased and the proportion of negative sentences is reduced compared to when the driver D has a high proficiency. This makes it possible to provide communication that is easy to understand, especially for a driver D who is unfamiliar with using the agent device 5. On the other hand, when the driver D has a high proficiency, the proportion of negative sentences used in the dialogue data is increased compared to when the driver D has a low proficiency. For a driver D who is accustomed to using the agent device 5, outputting a moderate amount of negative sentences can be expected to give the driver D the impression that the agent 52 is speaking his or her true feelings, thereby building a relationship of trust and making the driver more likely to have a favorable impression and feel closer to the agent 52.

[0042] Furthermore, when the proficiency level of driver D is lower than a predetermined value, that is, when driver D has just started using the agent device 5, the data generation unit 26 may output only positive sentences as response sentence data without using negative sentences. The predetermined value is not particularly limited, but may be a value at which driver D is estimated to be unfamiliar with using the agent device 5, for example, when the cumulative usage time of the agent device 5 by driver D is less than 50 hours, when the frequency of use is less than once a week, or when a specific operation and the cancellation of that operation are detected about the same number of times. As a result, in the early stages of starting to use the agent device 5, communication is carried out in a manner that is easy for driver D to understand, making it possible to build smooth communication using the agent device 5.

[0043] The data generation unit 26 determines whether to use positive or negative response sentence information based on the driver D's proficiency level, and then generates dialogue text data using response sentence information based on the driver D's voice information stored in the response sentence information database 25, response sentence information generated by fitting information on a specific topic obtained from an external website or other external site to typical response sentence examples, and response sentence information generated using a predetermined algorithm. When outputting dialogue data using the voice function of the agent device 5, the data generation unit 26 converts the dialogue text data into voice data by voice synthesis processing and transmits this to the output unit 27 as output data to be spoken by the agent 52. Known technologies can be applied to the voice synthesis processing. When outputting the dialogue text data as character information, the data generation unit 26 transmits the generated text data to the output unit 27 as output data to be displayed by the agent 52.

[0044] When the output unit 27 receives output data from the data generation unit 26, it outputs a control signal to the speaker or other audio output unit of the agent device 5, and to the display or other display unit, and outputs dialogue data using the agent function of the agent 52.

[0045] Next, the information processing procedure of the information processing system 1 of this embodiment will be described with reference to Fig. 6 to Fig. 8. Fig. 6 is a flowchart showing the information processing procedure based on the proficiency level of the driver D, which is executed by the information processing device 2 of Fig. 1.

[0046] 6, when the ignition switch of the vehicle is turned on, the following information processing is executed. In step S102, the occupant identification unit 21 identifies the driver D based on information acquired from the in-vehicle camera serving as the vehicle sensors 3. Next, in step S103, the proficiency detection unit 23 searches the occupant information database 22 based on the identification information of the driver D, and calculates the proficiency of the driver D using the usage information of the agent device 5.

[0047] In the following step S104, if it is determined that the proficiency of driver D calculated by proficiency detection unit 23 is lower than the predetermined value, the process proceeds to step S105. If driver D's proficiency is medium, the process proceeds to step S106, and if the driver's proficiency is high, the process proceeds to step S107.

[0048] If it is determined in step S104 that driver D's proficiency level is lower than the predetermined value, then driver D is in the early stages of starting to use the agent device 5, and therefore in step S105 the data generation unit 26 generates dialogue data using only positive sentences without using negative sentences.

[0049] Fig. 7C is a diagram showing an example of a scene when step S105 is reached. This diagram shows a scene in which dialogue data is generated to be output from the agent device 5 using a voice function for driver D. As shown in Fig. 7C, when the proficiency level of driver D is lower than a predetermined value, dialogue data is generated using only response sentence information of positive sentences such as "Welcome, Mr. XX (name of driver D)," "Yay, it's time for a drive," "Fasten your seatbelt and depart safely," and "You're driving safely."

[0050] If it is determined in step S104 that driver D has a medium level of proficiency, negative sentences are also used because driver D is not yet in the early stages of starting to use the agent device 5. However, because driver D is not yet fully accustomed to using the agent device 5, in step S106 the data generation unit 26 generates dialogue sentence data so that the proportion of positive sentences is high.

[0051] FIG. 7B is a diagram illustrating an example of a scene when step S106 is reached. As shown in FIG. 7B, when driver D's proficiency level is relatively low, dialogue data is generated such that the proportion of positive sentences is higher than when driver D's proficiency level is relatively high, including response sentence information of negative sentences such as "Carelessness is the enemy" in addition to response sentence information of positive sentences such as "Welcome, Mr. / Ms. XX (name of driver D)," "Yay, you're driving," and "Fasten your seatbelt and depart safely." Note that in FIG. 7B, dialogue data is generated from four pieces of response sentence information, and the negative sentence is output fourth. However, the number of response sentence information used in the dialogue data and the output order of the negative sentences are not limited to this. Furthermore, the data generation unit 26 may generate dialogue data by appropriately referring to the count record of response sentence information stored for driver D.

[0052] If it is determined in step S104 that driver D has a high level of proficiency, then in step S107 the data generation unit 26 determines whether or not the driver D has not performed a restraining operation on the output of the agent device 5 for a predetermined period of time. The predetermined period is not particularly limited, but may be a certain period of time, such as the past week. If a restraining operation on the output of the agent device 5 has been performed for the predetermined period of time, there is a risk that the driver D's receptivity to the agent device 5 has decreased, so the process proceeds to step S106, and dialogue sentence data with a high proportion of positive sentences is generated. On the other hand, if it is determined in step S107 that a restraining operation on the output of the agent device 5 has not been performed for the predetermined period of time, the process proceeds to step S108.

[0053] In step S108, the data generation unit 26 determines whether dialogue data with negative sentences has been generated a predetermined number of times or more within a predetermined period. The predetermined period and the predetermined number of times are not particularly limited, but may be a certain frequency, such as three or more times in the past five outputs. If dialogue data using negative sentences has been generated a predetermined number of times or more within a predetermined period, there is a risk that negative sentences will be output frequently and consecutively to the driver D. In this case, in order to prevent a decrease in the driver D's receptivity to the agent device 5, the process proceeds to step S106, where dialogue data with a high proportion of positive sentences is generated. On the other hand, if dialogue data with negative sentences has not been generated a predetermined number of times or more within a predetermined period, the process proceeds to step S109, where dialogue data with a high proportion of negative sentences is generated compared to when the proficiency level is relatively low.

[0054] FIG. 7A illustrates an example of a scene when step S109 is reached. As illustrated in FIG. 7A , when driver D's proficiency level is relatively high, dialogue data with a higher proportion of negative sentences is generated compared to when driver D's proficiency level is relatively low. This is done by using both positive response sentence information, such as "Yay, you're driving!" and "Fasten your seatbelt and depart safely!", and negative response sentence information, such as "What can I do for you?" and "Carelessness is the enemy." However, the data generation unit 26 controls the processing procedures of steps S107 and S108 described above to prevent the proportion of negative sentences from becoming too high. Note that the information processing of steps S107 and S108 is not essential to the present invention and may be omitted as appropriate. Also, in FIG. 7A , dialogue data is generated from four pieces of response sentence information, and negative sentences are output first and fourth. However, the number of response sentence information used in the dialogue data and the output order of the negative sentences are not limited thereto. The data generation unit 26 may also generate dialogue data by appropriately referencing count records of response sentence information stored for driver D.

[0055] In the following step S110, the generated dialogue data is used to generate output data to be output from the agent device 5. As shown in FIGS. 7A to 7C, when dialogue data is output using the voice function of the agent device 5, the text data of the dialogue is converted into voice data by voice synthesis processing, and output data is generated. Also, as shown in FIG. 8, when dialogue data is output using the display function of the agent device 5, the text data of the dialogue is used as the output data. Note that, although FIG. 8 shows an example in which an image of the agent 52 and the text data of the dialogue are displayed on the display DP, the display form is not limited to this. Also, an example in which different display forms are used to make it easier to distinguish between positive sentences and negative sentences is shown, but the display form is not limited to this.

[0056] In step S111, the output unit 27 outputs a control signal to a speaker or other audio output unit, a display or other display unit of the agent device 5, and outputs output data using the agent function of the agent 52. Note that the vehicle sensors 3, the input device 4, etc. may be used to acquire the driver D's reaction to the output data, and if an acceptable reaction (positive reaction) to the output data output from the agent device 5 is detected, the proficiency score of the driver D may be increased.

[0057] In step S112, when the ignition switch is turned off, the above information processing is terminated. On the other hand, the information processing from step S104 to step S111 is repeatedly executed until the ignition switch is turned off.

[0058] As described above, the information processing system 1 and information processing method of this embodiment detect the driver D (user)'s proficiency with the agent device 5, and generate dialogue data according to the driver D's (user's) proficiency using response sentence information classified into positive sentences and negative sentences. In this case, the proportion of positive sentences and the proportion of negative sentences used in the dialogue data are changed depending on whether the driver D's (user's) proficiency is relatively low or relatively high, so that communication that is easy for the driver D (user) to understand can be carried out using a variety of dialogue expressions.

[0059] Furthermore, according to the information processing system 1 and information processing method of this embodiment, the proficiency of the driver D (user) is estimated from at least one of the use time, use frequency, operation state, input frequency and input contents of input data by the driver D (user) for the agent device 5. This makes it possible to calculate the proficiency of the driver D (user) according to the actual use situation of the agent device 5.

[0060] Furthermore, according to the information processing system 1 and the information processing method of this embodiment, when the driver D (user) has a relatively low level of proficiency, the data generation unit 26 increases the proportion of positive sentences used in the dialogue data compared to when the driver D (user) has a relatively high level of proficiency. This makes it possible to provide communication that is easy to understand, especially for the driver D (user) who is unfamiliar with using the agent device 5.

[0061] Furthermore, according to the information processing system 1 and information processing method of this embodiment, the data generation unit 26 does not use the negative sentences in the dialogue data when the proficiency level of the driver D (user) is lower than a predetermined value. As a result, in the early stage when the driver D (user) starts using the agent device 5, communication is carried out with content that is easy for the driver D (user) to understand, and smooth communication using the agent device 5 can be established.

[0062] <<Second embodiment>> Next, a second embodiment of the present invention will be described with reference to FIGS. 9 to 11C. In the second embodiment, the information processing device 2 of the first embodiment is provided with a load estimation unit 28. In this embodiment, the load estimation unit 28 estimates the cognitive load that the driver D places when recognizing dialogue sentence data. When generating dialogue sentence data, the data generation unit 26 changes the proportion of positive sentences and negative sentences used in the dialogue sentence data based on the cognitive load of the driver D. Note that the configuration of the information processing device 2 other than the load estimation unit 28, and the configurations of the vehicle sensors 3, the input device 4, and the agent device 5 are the same as those of the first embodiment shown in FIG. 1, and therefore the explanations of these blocks will be cited from the above-mentioned embodiment.

[0063] The load estimation unit 28 estimates the cognitive load when the driver D recognizes the dialogue sentence data. The cognitive load is an index that indicates how easily the user can recognize the output data from the agent device 5. For example, if the user is performing a task other than operating the agent device 5 and cannot focus solely on the agent device 5, the cognitive load is estimated to be high. On the other hand, if the user is not performing a task other than operating the agent device 5 and can focus solely on the agent device 5, the cognitive load is estimated to be low. Furthermore, if the user is in an environment where it is difficult for the user to recognize the output data from the agent device 5, such as when there is noise in the surroundings, the cognitive load may be estimated to be high, and if the user is in an environment where it is easy for the user to recognize the output data from the agent device 5, such as when the surroundings are quiet, the cognitive load may be estimated to be low.

[0064] In this embodiment, when the user is a driver D of a vehicle, the load estimation unit 28 detects the surrounding environment of the vehicle based on detection values ​​from the vehicle sensors 3 and estimates the cognitive load of the driver D. For example, when the vehicle is traveling on a congested road, when there are many pedestrians around the vehicle, or when the vehicle is traveling on a highway, the driver D is in a situation where he or she must concentrate on driving operations and is unable to pay much attention to the agent device 5, and therefore the cognitive load is estimated to be high. On the other hand, when the driver D is traveling on a road without congested traffic, when there are no pedestrians around the vehicle, or when the vehicle is traveling in an autonomous driving control mode (autonomous speed control mode and / or autonomous steering control mode) using a driving assistance device, and the degree to which the driver D is concentrating on driving operations is relatively low, the driver D is in a situation where he or she can pay attention to the agent device 5, and therefore the cognitive load is estimated to be low. The load estimation unit 28 outputs the estimated cognitive load of the driver D to the data generation unit 26.

[0065] The data generation unit 26 determines the ratio of positive sentences and negative sentences to be used in the dialogue data based on the cognitive load of the driver D received from the load estimation unit 28. Specifically, when the cognitive load of the driver D is high, the ratio of positive sentences that are easy for the driver D to understand is increased compared to when the cognitive load of the driver D is low. This makes it possible to communicate with the agent device 5 taking into consideration the usage state of the driver D.

[0066] Furthermore, the data generation unit 26 may output only positive sentences as dialogue data without using negative sentences when the cognitive load of the driver D is higher than a predetermined value. A case in which the cognitive load is higher than the predetermined value occurs when, for example, the driver D is driving on a mountain road with many sharp curves, or when the driver D has just been switched (overridden) from the autonomous driving control mode to the manual driving mode, and the driving operation load is estimated to be higher than during normal driving. In such cases, outputting only positive sentences can prevent the driver D from performing driving operations impeded. Furthermore, since only dialogue data that has a positive or proactive meaning for the driver D is output, the driver D is more likely to have a favorable impression of the agent 52. The process of generating dialogue data using response sentence information, the process of generating output data from the dialogue data, and the process of outputting output data using the agent function are the same as those in the first embodiment, and therefore the above description is applicable.

[0067] Next, the information processing procedure of the information processing system 1 of this embodiment will be described with reference to Fig. 10 to Fig. 11C. Fig. 10 is a flowchart showing the information processing procedure based on the cognitive load of the driver D, which is executed by the information processing device 2 of Fig. 9. Note that the information processing based on the proficiency level of the driver D of Fig. 6 and the information processing based on the cognitive load of the driver D of Fig. 10 may be appropriately combined and processed.

[0068] First, in step S201 of Fig. 10, when the ignition switch of the vehicle is turned on, the following information processing is executed. In step S202, the occupant identification unit 21 identifies the driver D based on information acquired from the in-vehicle camera serving as the vehicle sensors 3. Next, in step 203, the load estimation unit 28 detects the vehicle's surrounding environment based on detection values ​​from the vehicle sensors 3, and estimates the cognitive load of the driver D.

[0069] In the following step S204, if it is determined that the cognitive load of driver D estimated by the load estimation unit 28 is higher than a predetermined value, the process proceeds to step S205. If the cognitive load of driver D is medium, the process proceeds to step S206, and if the cognitive load of driver D is low, the process proceeds to step S207.

[0070] If it is determined in step S204 that the cognitive load of driver D is higher than a predetermined value, i.e., if it is estimated that the driving load of driver D is higher than during normal driving, then in step S205 the data generation unit 26 generates dialogue data using only positive sentences without using negative sentences.

[0071] Fig. 11C is a diagram showing an example of a scene when step S205 is reached. This diagram shows a scene in which dialogue data is generated to be output from the agent device 5 to the driver D using the voice function. As shown in Fig. 11C, when the cognitive load of the driver D is higher than a predetermined value, dialogue data is generated using only response sentence information of positive sentences such as "Watch out for sharp curves," "You're driving safely," "Pay attention to what's ahead," and "We'll arrive soon."

[0072] If it is determined in step S204 that the cognitive load of driver D is medium, driver D is in a situation where he or she cannot pay much attention to the agent device 5, so in step S206 the data generation unit 26 generates dialogue data so that the proportion of positive sentences that are easy for driver D to understand is high.

[0073] FIG. 11B is a diagram showing an example of a scene when step S206 is reached. As shown in FIG. 11B, when the cognitive load of driver D is relatively high, dialogue data is generated so that the proportion of positive response sentence information such as "Watch out for what's ahead" and "We'll arrive soon" is higher than when the cognitive load of driver D is low, while including response sentence information of negative sentences such as "Carelessness is the enemy." Note that in FIG. 11B, dialogue data is generated from three pieces of response sentence information, and the negative sentence is output first, but the number of response sentence information used in the dialogue data and the output order of the negative sentences are not limited to this. Furthermore, the data generation unit 26 may generate dialogue data by appropriately referring to the count record of response sentence information stored for driver D.

[0074] If it is determined in step S204 that the cognitive load of driver D is low, then in step S207 the data generation unit 26 determines whether or not the driver D has not performed a restraining operation on the output of the agent device 5 for a predetermined period of time. If a restraining operation on the output of the agent device 5 has been performed for a predetermined period of time, there is a risk that the driver D's receptivity to the agent device 5 has decreased, so the process proceeds to step S206, and dialogue sentence data with a high proportion of positive sentences is generated. On the other hand, if it is determined in step S207 that a restraining operation on the output of the agent device 5 has not been performed for the predetermined period of time, the process proceeds to step S208.

[0075] In step S208, the data generation unit 26 determines whether dialogue data using negative sentences has been generated a predetermined number of times or more within a predetermined period. If dialogue data using negative sentences has been generated a predetermined number of times or more within a predetermined period, there is a risk that negative sentences will be output frequently in succession to the driver D, so the process proceeds to step S206, where dialogue data with a high proportion of positive sentences is generated.

[0076] On the other hand, if dialogue data with negative sentences has not been generated the predetermined number of times or more within the predetermined period in step S208, the process proceeds to step S209, where the data generator 26 generates dialogue data with a higher proportion of negative sentences than when the recognition load is relatively high. Note that the processing procedures of steps S207 and S208 are controls to prevent the proportion of negative sentences from becoming too high, and are not essential components of the present invention, so they may be omitted as appropriate.

[0077] FIG. 11A is a diagram illustrating an example of a scene when step S209 is reached. As shown in FIG. 11A, when the cognitive load of driver D is low, it is estimated that the driver is able to focus on the agent device 5. Therefore, dialogue data is generated using both positive response sentence information such as "Watch out for what's ahead," and negative response sentence information such as "Carelessness is the enemy" and "We're finally getting there," with a higher proportion of negative sentences than when the cognitive load is relatively high. Note that in FIG. 11A, dialogue data is generated from three pieces of response sentence information, and negative sentences are output first and third. However, the number of response sentence information used in the dialogue data and the output order of the negative sentences are not limited to this. Furthermore, the data generation unit 26 may generate dialogue data by appropriately referring to count records of response sentence information stored for driver D.

[0078] In the following step S210, the generated dialogue data is used to generate output data to be output from the agent device 5. As shown in Figs. 11A to 11C, when dialogue data is output using the voice function of the agent device 5, the text data of the dialogue is converted into voice data by voice synthesis processing to generate output data. Furthermore, although not shown, when dialogue data is output using the display function of the agent device 5, the text data of the generated dialogue may be used as output data.

[0079] In step S211, the output unit 27 outputs a control signal to a speaker or other audio output unit of the agent device 5, a display or other display unit, and outputs output data by the agent function of the agent 52.

[0080] In step S212, when the ignition switch is turned off, the above information processing is terminated. On the other hand, the information processing from step S204 to step S211 is repeatedly executed until the ignition switch is turned off.

[0081] As described above, the information processing system 1 and information processing method of this embodiment further include a load estimation unit that estimates the load on the driver D (user) when recognizing dialogue data, and the data generation unit 26 increases the proportion of positive sentences used in the dialogue data when the driver D (user) has a relatively high load compared to when the driver D (user) has a relatively low load. This makes it possible to communicate with the agent device 5 while taking into consideration the driver D's (user's) usage state.

[0082] Furthermore, according to the information processing system 1 and information processing method of this embodiment, when the load on the driver D (user) is higher than a predetermined value, the data generation unit 26 does not use the negative sentences in the dialogue data, thereby preventing the driver D (user) from being hindered in performing a task. Also, since only dialogue data that has a positive or proactive meaning for the driver D (user) is output, the driver D (user) is more likely to have a favorable impression of the agent 52.

[0083] <<Third Embodiment>> Next, a third embodiment of the present invention will be described with reference to FIGS. 12 to 15C. In the third embodiment, the information processing device 2 of the second embodiment is equipped with an other-person determination unit 29. In this embodiment, the other-person determination unit 29 determines whether or not a person other than the user is present within a predetermined range. If a person other than the user is present within the predetermined range, the data generation unit 26 changes the proportion of positive sentences and negative sentences used in the dialogue data when generating dialogue data. The predetermined range is not particularly limited, but is a range within which communication with the agent device 5 can be performed. In the present embodiment, if the agent device 5 is installed in a vehicle, it is the interior of the vehicle. Note that the configuration of the information processing device 2 other than the other-person determination unit 29, the vehicle sensors 3, the input device 4, and the agent device 5 are the same as those of the first embodiment shown in FIG. 1, and therefore the description of these blocks will be incorporated in the above-described embodiment. Furthermore, the load estimation unit 28 is not an essential component of the present invention and may be omitted as appropriate.

[0084] The other person determination unit 29 determines whether or not a passenger X other than the driver D is present in the vehicle cabin based on input signals from the vehicle sensors 3. If it determines that the passenger X is present, it stores the identification information of the passenger X at least temporarily. In the scene shown in FIG. 2 , the other person determination unit 29 identifies a person sitting in a seat other than the driver's seat as the passenger X based on image data captured by an in-vehicle camera serving as the vehicle sensors 3, and stores the identification information of the passenger X. Note that the determination of the presence of the passenger X may also use input signals from a seating sensor provided inside the seating portion of the seat or a seat belt sensor. The other person determination unit 29 also generates riding information by associating the detection values ​​acquired from the vehicle sensors 3 with the identification information of the passenger X, and stores the generated riding information in an occupant information database. The riding information is information such as a riding history that records the times when the passenger X boarded the vehicle and disembarked from the vehicle.

[0085] FIG. 13 is a diagram showing an example configuration (passenger) of the occupant information database 22. As shown in FIG. 13, in the occupant attributes, for example, friend A is recorded as guest 1, and the riding history of friend A of guest 1 getting on and off the vehicle is stored. The occupant information database 22 may also include attribute information such as the riding frequency of the guest and the intimacy between the driver D and each guest estimated from voice information acquired via the input device 4. In addition, in the case of passenger X riding in the vehicle for the first time, since passenger X is not stored in the occupant information database 22, the other person determination unit 29 generates new data. The identification information of passenger X identified by the other person determination unit 29, and the riding history and attribute information of passenger X generated by the other person determination unit 29 are stored in the occupant information database 22 and output to the data generation unit 26.

[0086] The data generation unit 26 determines the proportion of positive sentences and the proportion of negative sentences to be used in the dialogue data based on the identification information and attribute information of the passenger X received from the other person determination unit 29. As described above, negative sentences are sentences that are used with a negative or passive meaning for the driver D, so if a large number of negative sentences are output when the passenger X is present, the driver D may feel uncomfortable. Therefore, when the passenger X is present, the data generation unit 26 increases the proportion of positive sentences to be used in the dialogue data compared to when the passenger X is not present. This allows appropriate communication to be carried out even when a passenger X other than the driver D is present.

[0087] Furthermore, if there is a passenger X who is not stored in the occupant information database 22, i.e., who is riding in the vehicle for the first time, the data generation unit 26 may output only positive sentences as response sentence data without using negative sentences. Even if the passenger X is stored in the occupant information database 22, if the passenger X's attribute information is lower than a predetermined value, the data generation unit 26 may output only positive sentences as response sentence data without using negative sentences. The predetermined value is not particularly limited, but may be a case where the passenger X rides infrequently or where the intimacy between the driver D and the passenger X is low. In cases where there is a passenger X who is riding in the vehicle for the first time, a passenger X who rides infrequently, or a passenger X who has low intimacy with the driver D, the driver D needs to focus on communication with the passenger X rather than with the agent device 5. Therefore, smooth communication between the driver D, the passenger X, and the agent device 5 can be achieved by outputting only easy-to-understand positive sentences from the agent device 5. Furthermore, since sentences that are used with a negative or negative meaning for the driver D are not output, the driver D is more likely to have a favorable impression of the agent.

[0088] Next, the information processing procedure of the information processing system 1 of this embodiment will be described with reference to Figs. 14 to 15C. Fig. 14 is a flowchart showing the information processing procedure based on the determination of whether or not a passenger X is present, which is executed by the information processing device 2 of Fig. 12. Note that any two of the information processing based on the proficiency level of the driver D of Fig. 6, the information processing based on the cognitive load of the driver D of Fig. 10, and the information processing based on the determination of whether or not a passenger X is present of Fig. 14 may be appropriately combined and processed, or all three of the information processing may be combined and processed.

[0089] First, in step S301 of Fig. 14, when the ignition switch of the vehicle is turned on, the following information processing is executed. In step S302, the occupant identification unit 21 identifies the driver D based on information acquired from the in-vehicle camera serving as the vehicle sensors 3. Next, in step 303, the other person determination unit 29 determines whether or not a passenger X other than the driver D is present in the vehicle based on image data captured by the in-vehicle camera. If it is determined that a passenger X is present in the vehicle, the process proceeds to step S304. On the other hand, if it is determined that a passenger X is not present in the vehicle, the process proceeds to step S308.

[0090] If it is determined in step S303 that passenger X is present in the vehicle, it is determined in step S304 whether or not the attribute information of passenger X is stored in the occupant information database 22. If the attribute information of passenger X is not stored in the occupant information database 22, the process proceeds to step S306. On the other hand, if the attribute information of passenger X is stored in the occupant information database 22, the process proceeds to step S305.

[0091] In step S305, the other person determination unit 29 determines whether the attribute information of the passenger X is lower than a predetermined value. The attribute information of the passenger X includes the frequency of the passenger X riding in a vehicle and the degree of intimacy between the driver D and the passenger X. If the attribute information of the passenger X is lower than the predetermined value, the process proceeds to step S306. On the other hand, if the attribute information of the passenger X is not lower than the predetermined value, the process proceeds to step S307.

[0092] If it is determined in step S304 that the attribute information of passenger X is not stored in occupant information database 22, i.e., passenger X is a person riding in a vehicle for the first time, or if it is determined in step S305 that the attribute information of passenger X is lower than a predetermined value, i.e., passenger X rides infrequently or the intimacy between driver D and passenger X is low, then in step S306 data generation unit 26 generates dialogue sentence data using only positive sentences and no negative sentences.

[0093] Fig. 15C is a diagram showing an example of a scene when step S306 is reached. This diagram shows a scene in which dialogue data is generated to be output from the agent device 5 to the driver D using the voice function. As shown in Fig. 15C, if there is a passenger X who is not stored in the passenger information database 22, or if there is a passenger X who is stored in the passenger information database 22 but whose attribute information is lower than a predetermined value, dialogue data is generated using only response sentence information of positive sentences such as "Thanks for washing the car," "You're driving safely," and "We'll arrive soon."

[0094] If it is determined in step S305 that the attribute information of passenger X is not lower than a predetermined value, negative sentences may be used. However, in order to avoid outputting many negative sentences that are used in a negative or passive sense for driver D, in step S307 the data generation unit 26 generates dialogue data so that the proportion of positive sentences is high.

[0095] FIG. 15B is a diagram showing an example of a scene when step S307 is reached. As shown in FIG. 15B, when a passenger X whose attribute information is higher than a predetermined value is present, dialogue data is generated so that the proportion of positive sentences such as "Thanks for washing the car" and "We'll arrive soon" is higher than when a passenger X is not present, while suppressing negative sentences such as "Carelessness is the enemy." Note that in FIG. 15B, dialogue data is generated from three pieces of response sentence information, and the negative sentence is output second, but the number of response sentence information used in the dialogue data and the output order of the negative sentences are not limited to this. Furthermore, the data generation unit 26 may generate dialogue data by appropriately referring to the count record of response sentence information stored for the driver D.

[0096] If it is determined in step S305 that there is no passenger X other than the driver D, then in step S308 the data generation unit 26 determines whether or not the driver D has not performed a restraining operation on the output of the agent device 5 for a predetermined period of time. If a restraining operation has been performed on the output of the agent device 5 for a predetermined period of time, there is a risk that the driver D's receptivity to the agent device 5 has decreased, so the process proceeds to step S307, and dialogue sentence data with a high proportion of positive sentences is generated. On the other hand, if it is determined in step S308 that no restraining operation has been performed on the output of the agent device 5 for a predetermined period of time, the process proceeds to step S309.

[0097] In step S309, the data generation unit 26 determines whether dialogue data using negative sentences has been generated a predetermined number of times or more within a predetermined period. If dialogue data using negative sentences has been generated a predetermined number of times or more within a predetermined period, there is a risk that negative sentences will be output frequently in succession to the driver D, so the process proceeds to step S307, where dialogue data with a high proportion of positive sentences is generated.

[0098] On the other hand, if dialogue data with negative sentences has not been generated a predetermined number of times or more within a predetermined period in step S309, the process proceeds to step S310, where dialogue data with a higher proportion of negative sentences is generated compared to when passenger X is present. Note that the processing procedures of steps S308 and S309 are controls to prevent the proportion of negative sentences from becoming too high, and are not essential components of the present invention, so they may be omitted as appropriate.

[0099] FIG. 15A is a diagram illustrating an example of a scene when step S310 is reached. As shown in FIG. 15A, when there is no passenger X other than driver D, dialogue data is generated using both positive sentences such as "We'll arrive soon" and negative sentences such as "You haven't washed your car recently" and "Your guard is the enemy," with a higher proportion of negative sentences than when there is passenger X. Note that in FIG. 15A, dialogue data is generated from three pieces of response sentence information, and negative sentences are output first and second, but the number of response sentence information used in the dialogue data and the output order of the negative sentences are not limited to this. Furthermore, the data generation unit 26 may generate dialogue data by appropriately referring to the count record of response sentence information stored for driver D.

[0100] In the following step S311, the generated dialogue data is used to generate output data to be output from the agent device 5. As shown in Fig. 15A to Fig. 15C, when dialogue data is output using the voice function of the agent device 5, the text data of the dialogue is converted into voice data by voice synthesis processing to generate output data. Furthermore, although not shown, when dialogue data is output using the display function of the agent device 5, the text data of the generated dialogue may be used as output data.

[0101] In step S312, the output unit 27 outputs a control signal to a speaker or other audio output unit of the agent device 5, and a display or other display unit, and the agent function of the agent 52 outputs output data.

[0102] In step S313, when the ignition switch is turned off, the above information processing is terminated. On the other hand, the information processing from step S304 to step S312 is repeatedly executed until the ignition switch is turned off.

[0103] As described above, the information processing system 1 and information processing method of this embodiment further include the other person determination unit 29 that determines whether or not a fellow passenger X (other person) other than the driver D (user) is present within a predetermined range, and when the other person determination unit 29 determines that a fellow passenger X (other person) is present within the predetermined range, the data generation unit 26 increases the proportion of positive sentences used in the dialogue sentence data compared to when it is determined that a fellow passenger X (other person) is not present within the predetermined range. This allows appropriate communication to be carried out even when a fellow passenger X (other person) other than the driver D (user) is present.

[0104] Furthermore, according to the information processing system 1 and information processing method of this embodiment, the other person determination unit 29 identifies and stores attribute information of the passenger X (other person). If the passenger X (other person) is present within a predetermined range and the attribute information of the passenger X (other person) is not stored, or if the stored attribute information of the passenger X (other person) is lower than a predetermined value, the data generation unit 26 does not use negative sentences in the dialogue data. This allows smooth communication between the driver D (user), the passenger X (other person), and the agent device 5. Furthermore, because sentences that are used with a negative or negative meaning for the driver D (user) are not output, the driver D (user) is more likely to have a favorable impression of the agent 52.

[0105] It should be noted that the above-described embodiments have been described to facilitate understanding of the present invention, and are not intended to limit the present invention. Therefore, each element disclosed in the above-described embodiments is intended to include all design modifications and equivalents that fall within the technical scope of the present invention. [Explanation of symbols]

[0106] 1. Information processing system 2...Information processing device 21...Occupant identification section 22...Crew information database 23...Proficiency detection unit 24...Input data processing section 25...Response sentence information database 26...Data generation section 27...Output section 28...Load estimation section 29...Other Judgment Section 3...Vehicle sensors 4...Input device 5...Agent device

Claims

1. An information processing system comprising: an agent device having an agent function; and an information processing device that generates dialogue data for a user, the information processing system outputting the generated dialogue data to the user using the agent function, The information processing device includes: a load estimation unit that estimates a load when the user recognizes the dialogue data; a data generation unit that generates the dialogue data using response sentence information classified into positive sentences and negative sentences; a detection unit that detects the user's proficiency with the agent device based on an operation history including a specific operation with respect to the agent device, The data generation unit When the load on the user is relatively high, a proportion of the positive sentences used in the dialogue data is increased compared to when the load on the user is relatively low; An information processing system that changes the ratio of the positive sentences and the negative sentences used in the dialogue data depending on whether the user's proficiency level is relatively low or relatively high.

2. The information processing system according to claim 1 , wherein the data generating unit does not use the negative sentence in the dialogue data when the load on the user is higher than a predetermined value.

3. 3. The information processing system according to claim 1, wherein the user's proficiency is estimated from at least one of the user's usage time, usage frequency, operation state, input frequency, and input content of input data from the user with respect to the agent device.

4. The information processing system according to any one of claims 1 to 3, wherein the data generation unit increases the proportion of positive sentences used in the dialogue data when the user's proficiency level is relatively low compared to when the user's proficiency level is relatively high.

5. 5. The information processing system according to claim 1, wherein the data generating unit does not use the negative sentence in the dialogue data when the user's proficiency level is lower than a predetermined value.

6. Further, an other person determination unit is provided to determine whether or not other people other than the user are present within a predetermined range, An information processing system according to any one of claims 1 to 5, wherein, when the other person determination unit determines that the other person exists within the specified range, the data generation unit increases the proportion of positive sentences used in the dialogue data compared to when it determines that the other person does not exist within the specified range.

7. the other person determination unit identifies and stores attribute information of the other person; The information processing system of claim 6, wherein if the other person is present within the specified range and attribute information of the other person is not stored, or if the stored attribute information of the other person is lower than a specified value, the data generation unit does not use the negative sentence in the dialogue data.

8. 1. An information processing method for causing a processor to execute a process of generating dialogue data for a user using response sentence information classified into positive sentences and negative sentences, and outputting the dialogue data to the user using an agent device, comprising: The processor: Estimating a load on the user when recognizing the dialogue data; detecting a user's level of proficiency with the agent device based on an operation history including a specific operation with respect to the agent device; generating the dialogue data so that, when the user's load is relatively high, a proportion of the positive sentences used in the dialogue data is higher than when the user's load is relatively low; generating the dialogue data so that a ratio of the positive sentences and a ratio of the negative sentences used in the dialogue data are changed depending on whether the user's proficiency level is relatively low or relatively high; The information processing method outputs the generated dialogue data to the user.

Citation Information

Patent Citations

  • Drive evaluation device

    JP2006347296A

  • Vehicle information providing apparatus

    JP2014139777A

  • Conversation system and program

    JP2017157011A

  • On-board device and driving diagnostic method

    JP2019011015A

  • Interaction device

    JP2020064144A